A profile that fails validation for a reason unrelated to identity — a non-UUID
id, a mismatched partition — armed both the notice and its degraded flag. Since
hydrateFromPersisted skips such entries silently and nothing ever repairs them,
the user got "an old browser identity choice could not be inspected" forever,
about a profile that never carried one.
Key the notice on the presence of userAgentMode instead, and use validation only
to decide whether the choice that was found is inspectable. Refusing to hydrate
an entry and finding a retired choice are now separate facts.
The old case table asserted the defect for null, 42 and 'broken', so it is
replaced by two tables stating the new contract rather than adapted to pass.
The retired per-profile userAgentMode bytes are retained on disk by design, so
every launch rediscovers them and re-arms the notice — including the launch
right after the user answers it, and every launch after that. Documented as
one-time, it was permanent.
The record already carries explicitSelection, which is exactly the fact that
should end the notice. Gate the mark at the single writer rather than deleting
the legacy key, so the retained bytes stay untouched and disk never claims a
notice is pending beside a choice the user already made.
The new test pushed the persistence suite past max-lines, so the in-memory fs
and module mocks move to a named fixture module and the retired-identity tests
move beside them in their own file.
This probe passes locally and fails on CI with an empty receipt set, an empty
CDP diagnostic list, and a fixture that still exits 0 — so the assertion message
carried nothing usable. Thread the fixture's own result and stderr into the
capture assertion so the next run says what the fixture actually did.
My first attempt anchored on the engine comment, which broke a startup fixture
whose platform comment is "(Test)" with no "(KHTML, like Gecko)" at all — the app
token survived and the ordering test went red.
Anchoring on the nearest ")" before Chrome/ and consuming only non-")" tokens
keeps the match inside that gap, so it handles a multi-word app name, a synthetic
platform comment, and an already-clean identity alike. A user agent with no such
gap is still returned unchanged.
The fixture shape is now a test case, since it is what caught the first attempt.
app.setName decides the app token in the user agent, and dev sets "Orca Dev".
The cleaner matched a single whitespace-delimited token, which cannot span that
space, so the replace failed outright and every dev build presented
"Orca Dev/1.4.203" on the wire — the exact token class that gets transplanted
sessions revoked.
Anchoring on the engine comment and consuming lazily up to Chrome/ removes any
number of app tokens. A user agent without that comment is returned unchanged
rather than mangled, because over-stripping is worse than under-stripping.
The function had no unit test at all; it was only exercised through the
real-Electron wire tests, which run with a single-token fixture name. That is
why this survived.
The wire probe server and CDP collector arrived with the cross-context coverage
and were never added to the audit list. The collector's two real call sites are
safe: the poll cancels its unread body and the version probe consumes it through
response.json(). Every hit in the probe server is inside an injected page or
worker script source string, not a call this process makes.
The channel split is asserted total, so adding browser:identity:get/set left it
short by two. They manage the host's own process-wide user-agent choice rather
than acting on a guest the reader is looking at, so they sit with the session
and profile channels, not the preview tools.
Pins the mixed-version guarantee that had no test: browser.identity.v1 is advertised when the identity store is initialized and absent when it is not. Verified discriminating -- advertising it unconditionally fails the test.
The profileCreate rejection test asserted ok:false against a runtime with no browserProfileCreate, so that assertion passed even when the retired field was accepted. It now stubs a working runtime method, making ok:false load-bearing, and asserts the runtime is never reached.
Removes the persistence fixture's dead failIdentityWrite branch on writeFileAtomically: nothing on that path calls it, so it implied a second write mechanism that does not exist. Failure is injected through node:fs, which is what the identity write actually uses.
The queue could not be falsified by any test: writeRecord is synchronous end to end, so two calls cannot interleave and removing serialization entirely left every store test green. Carrying machinery whose guard is unconstructible is what the design review told us to cut, so it is gone. If durable writes ever become async, serialization comes back with the change that makes it testable.
The test that claimed to prove serialization now states what it actually pins -- the later of two selections is the one that survives -- and the module doc no longer claims a queue that is not there.
Adds the guard that was missing on reset: two resets across separate launches must produce two distinct backups, each holding its own original bytes. Verified discriminating -- a fixed backup filename fails it.
browser.identity.v1 was static, so every host claimed it including one that never initialized the identity store, where both methods can only throw. It now follows the browser.headless.v1 precedent and is pushed at status time when the store is actually initialized.
Also covers the retired profileCreate userAgentMode field at the dispatcher rather than only at the schema, so an older client provably gets the changed-semantics rejection over the wire instead of a success with the field quietly dropped.
A corrupt or newer-version record left the identity unchangeable with no way out. An explicit reset now copies the old bytes verbatim to a fresh unique path before publishing a replacement, and refuses the whole operation if that backup cannot be written -- so the reset can never be the thing that loses the data. Nothing resets automatically.
Future-version data says update Orca rather than reporting corruption. Reset is opt-in via browser.identity.set and orca browser identity set --reset.
ProfileCreate and BrowserIdentitySet move to browser-identity-params.ts: both carry the per-profile to app-wide identity move, and browser-params.ts was over its line cap.
Also registers browser as a top-level CLI name so the Windows launch redirect covers it -- without it orca browser identity get boots the GUI and exits silently there -- and adds the canonical browser identity show alias the CLI vocabulary policy requires.
The rescued work already serialized identity writes, but the writer lived beside the pre-ready reader, so nothing stopped a second caller from writing the record directly -- which is the shape of the bug this change set removes.
browser-identity-mode-record.ts is now read-only: record shape, path, parsing and the pre-ready synchronous read. browser-identity-mode-store.ts owns every mutation behind one queue, holds the snapshot and listeners, and derives restartRequired from appliedMode vs configuredMode rather than storing it. Consumers move to the store.
The two identity RPC methods also move out of browser-core.ts into browser-identity-rpc.ts: they read and write this host's own process identity rather than driving a page, and browser-core.ts was over its line cap. The generated params catalog is byte-identical.
Finishes the interrupted edits in 7db9c54b54:
- browser-user-agent-migration-notice.ts was truncated mid-write; close the
then() callback so the file parses.
- Register browser.identity.get/set in the generated RPC params catalog so the
params type-parity gate is satisfied.
- Retire the persistence assertions for the superseded design: a
migratedNativeProfileIds event map, a notice-acknowledgement clear, and a
global persistence-failure accessor. Legacy userAgentMode bytes are retained
now, so these assert retention plus a failed notice write still hydrating.
- The in-memory fs fixture threw a codeless ENOENT, which reads as "unreadable"
rather than "missing" and made every identity write refuse. Carry the code.
- Use the segmented control's per-option disabled rather than adding a
control-level prop it does not have.
* feat(native-chat): add a message rail for jumping between your prompts
A vertical rail down the right edge of the transcript, one bar per user
message, with the bar for the turn you are reading highlighted once
scrolling settles. Hovering the rail opens a panel that previews every
prompt and jumps to it on click.
Bars are capped at 20 and sampled evenly across the thread, always
keeping both ends and the active bar, so the rail stays readable at a
glance on a long conversation.
The active bar is resolved from virtualizer offsets rather than by
scanning rendered rows: the transcript is windowed, so an off-window row
has no element to measure. The row at the scroll fold resolves to its
owning prompt through turnKey, which is what keeps your own message lit
while you read a long reply instead of going dark.
Jumps reuse the existing reveal/pin path and scrollMessageToTop, which
releases the bottom pin. Scrolling through the virtualizer directly would
leave a reader snapped back down by the next streamed token.
Ticks cover loaded history only; older prompts gain a bar once "Load
earlier messages" pages them in.
* fix(native-chat): service a rail jump once and give its pin back
The rail borrowed the diff reveal's pin to reach a row the window had left
behind, but copied only its state shape, not its consumption. The request
was never cleared and the effect depended on `slots`, which is rebuilt on
every render, so three things went wrong at once:
- every later render re-scrolled to the jumped message, dragging a reader
back there for the rest of the pane's life, and forcing the bottom pin
off each time;
- the standing request outranked `revealedDiff` in the shared pin, so
revealing a diff outside the window silently stopped mounting its row;
- the pinned row stayed mounted and measured indefinitely.
The request now carries a monotonic id, is serviced once, and is released
as soon as the scroll is issued, which hands the pin back.
The rail's scroll listener had the same churn: it listed `items` in its
deps, so a streaming turn tore the listener down and cancelled the pending
idle timer on every frame and the highlight never settled. It now
subscribes once and re-reads on a key built from the prompt ids.
Also: the hover trigger is a real button, because `asChild` discards the
primitive's focusable trigger and the panel is the only way to reach these
messages; the wheel forwarder honours line and page delta modes rather
than treating every delta as pixels; and the e2e panel assertion is exact,
since a loose bound passed at 20 rows against 20 ticks.
* fix(native-chat): make prompt rail accessible and reuse previews
* fix(native-chat): supersede prior navigation when selecting a prompt
* fix(native-chat): stop the transcript following an end it measured short
The virtualizer compensates a row's measured size change by moving scrollTop
whenever it believes the view was already at the end. It decides that from the
spacer's own height minus a container-absolute offset, so the distance it
computes is short by everything in the document outside the spacer: the
transcript's top gutter, the "load earlier" block while older history is still
pageable, and the trailing chrome. A reader sitting ~100px above the bottom
therefore measured as "at the end", and every row that settled below them
dragged them down to it.
Measured in the windowing harness with a 92px gutter and 24px of trailing
chrome: a reader parked 96px above the end is pulled to the end on the first
growth frame, scrollTop 9261 to 9357.
The same option gates following an append, but that path measures the true
document distance, so it was never wrong, only redundant. The transcript
already decides whether to follow the end from the scroll container's real
geometry, and it re-pins once the growth is in the document rather than before
it, where the library's own write is clamped. Both library end behaviours are
retired by a threshold no finite distance can meet; the prepend anchoring that
shares the option is kept.
overflow-anchor:none is restated as structural: the engine's anchoring writes
never pass through the scrollToFn adapter that attributes this pane's own
scrolls, so they would arrive unmarked and read as the reader leaving.
* fix(native-chat): preserve visible rows on first measurement
Worker ctx_cb5b1262d7fe stopped ~2h ago mid-implementation (last heartbeat
2026-09-14T22:48:06Z) leaving this uncommitted. Committed unverified to make it
recoverable; not reviewed, not necessarily green.
* feat(design-system): gate renderer UI with @shadcn/lint
Wires shadcn-ui/lint's Oxlint plugin into the two places this repo already
ratchets: the changed-lines PR gate for rules the renderer can't satisfy
today, and `pnpm lint` for the one that is already at zero.
- config/oxlint-design-system.json: no-restyle (layout allowed),
no-raw-colors, require-static-classes -- scoped to src/renderer/**/*.tsx,
run over added lines only. Measured at 10 findings across the last 60
commits (771 changed files), so it holds the line without a migration.
- config/oxlint-dead-classes.json: no-unknown-classes repo-wide, with the
renderer's plain-CSS hook namespaces allow-listed. Now at zero.
- no-inline-styles and no-arbitrary-values stay off; STYLEGUIDE says why.
Fixes the three live bugs the linter found:
- `--editor-surface` never reached `@theme inline`, so `bg-editor-surface`
generated no CSS -- 12 editor/artifact/notebook panes fell through to the
page background instead of #1e1e1e in dark mode.
- `scrollbar-none` is not a Tailwind utility and was declared nowhere, so
the remote file browser breadcrumbs showed the scrollbar they meant to
hide. Declared as a real `@utility`.
- Notebook markdown cells used `markdown-preview-body`, which no stylesheet
defines; the styled class is `markdown-body`. They rendered unstyled.
* ci: run the dead-class gate in PR CI
`pnpm lint` gained check:dead-classes, and pr-workflow-lint-parity requires
every `pnpm lint` step to have a matching step in pr.yml.
* fix(notebook): keep markdown theme selectors working
* fix(native-chat): let a reader park just above the latest message
A reader who scrolled up by less than the bottom threshold was still
classified as being at the end, so follow stayed armed and the next chunk
of stream carried them back down. One constant was answering two
different questions: how close to the end still counts as pinned, and
whether a reader's own scroll meant to stay there.
The first wants slack, because a streaming last message jitters in height
by tens of pixels. The second wants almost none, because it is a
statement of intent. Give it its own, far stricter band, and move the
choice of band into the decision rather than leaving it to the call site,
which is where the two got conflated.
Re-arming follow now requires the reader to be within 4px of the end:
enough for fractional-pixel and zoom rounding, well inside one line of
prose. The pin and the jump-to-latest affordance keep their 48px band.
* fix(native-chat): make transcript intent own end following
* fix(pty): preserve unverifiable local child reads
* fix(pty): make child-process inspection synchronous
Separate foreground and child-process sampling. Sample child processes
synchronously after confirming foreground availability, returning
unverifiable verdicts when pty reads fail. Handle both transport loss
and local read failures uniformly in the completion coordinator.
* refactor(renderer): give the IPC error reader a clamped and an unclamped shape
* fix(settings): tell a failed load apart from a genuinely empty pane
* refactor: consolidate import types and simplify failure handling
- Move filesystem import types to shared for renderer use
- Add compactIpcErrorMessage for single-line error display
- Consolidate entry failure toasts to single global slot
- Simplify account tracking and discard retry logic
* fix type
* fix: clear stale state when pane loads fail
Credential reads, account fetches, and skill scans can fail, leaving stale
data on screen. This change clears previous state when a load fails,
distinguishing load failures from genuinely empty results, and prevents
stale controls from appearing after failed re-checks.
Use readIpcErrorMessage for consistent error handling and track runtime
targets to invalidate results from old targets.
* fix(settings): show credential action when bitbucket status read fails
When the credential-read operation fails, allow users to retry by showing
"Add or replace credentials" button. Initialize the credentials dialog with
the current (confirmed) connection state instead of stale data from a failed
read, preventing outdated information from pre-populating the form.
* refactor(renderer): give the IPC error reader a clamped and an unclamped shape
* fix(composer): name the attachments a drop could not add, in one toast
* fix(composer, source-control): use one stable failure toast slot
- Replace per-worktree toast IDs with single slot that replaces on each failure
- Remove destructive retry actions; discard must confirm in dialog
- Consolidate filesystem import types to shared location
- Add compactIpcErrorMessage for string error handling
* refactor: centralize filesystem import types and clarify failure naming
Move import result types from main/ipc to shared layer so they're available
across preload and renderer. Rename uniformFailure → commonFailure and
skippedOrFailed → failureCount for clarity. Simplify preload/API type
definitions by reusing shared types directly instead of duplicating inlined
union shapes.
* Reuse single toast slot for composer drop failures
Multiple drop failures now replace the previous toast instead of
stacking, preventing notification clutter. Uses a dedicated toast ID
separate from Source Control's stage/discard notifications.
Expand saved OMP descendants lazily while preserving exact child targets for Resume and View Log. Retain expanded branches across virtual scrolling and reject late responses/cycles. Includes the independently reviewed child-workspace correction from #20629.
61 combined target/map/nesting tests and actual OMP child/grandchild storage/CLI smoke pass. Earlier hidden Electron proof covers eight generations and narrow sidebar layout. Folder-only unresolved child targets remain disabled. No live delegation or full terminal-launch proof claimed.
Addresses #12885 Scope 2.
Add Resume to eligible local OMP child history rows. Resolve lazy child targets from their own cwd and host, never an unrelated active workspace. Unresolved folder-only targets stay disabled; copy-command remains available.
Verified production map/resume resolver regression before/after; 50 focused tests and independent 40-test review, web types and code quality passed. Actual OMP storage/CLI smoke confirms distinct child/grandchild sessions. No native Windows or live SSH launch claim.
Addresses #12885 Scope 1.
Preserve the original folder locator through Git upgrade and subsequent listing, persistence, and removal decisions after proving it still names the same checkout.
Independently reviewed with 60 focused persistence/listing/removal tests and six native Windows real-Git/NTFS cases covering case/slashes, junction retention and retargeting, remote-host isolation and unrelated checkout preservation. Prior source-connected native OMP proof confirms process survival. Full PR CI passed; no rebuilt full-app after-proof claimed.
Preserve checkout files and the named branch when removing a positively attested malformed Git-file registration. Reject file/symlink targets in deferred directory deletion.
Verified exact head with 75 focused tests including actual Git malformation, preserved marker/file bytes and branch HEAD. Independent review and complete product CI passed. WSL routing is covered by unit tests; direct SSH fails safely without local recovery.
Fixes#17316
Repairs #20559, whose termination was a no-op: `detached` is a spawn-only option and `exec` ignored it, so the shell never became a group leader. Verified against real processes.
Refs #19334
fix(runtime): retry a delivery that the idle gate refused
Gates delivery at the two points where each implementation commits to typing into
the pane, rather than at each caller, and parks-and-re-offers a refusal so late
idle evidence cannot strand a queued message.
Refs #6011
An idle agent pane kept ~40% of a core busy just by being frontmost. The
xterm cursor and the native chat caret blink with no paint-containment
boundary, so Chromium treated each blink as damage to the whole pane
ancestry and re-rasterized it twice a second.
- `.xterm-container` and the native composer's input shell get
`contain: paint`, bounding blink damage to the surface that blinks.
- The mention hint gains `z-20` to match the slash picker: a contained
element becomes a stacking context and paints at z-index 0 in tree
order, which would otherwise cover the hint's drop shadow.
Also records that DECSCUSR pins `decPrivateModes.cursorBlink`, which wins
over the option in `_updateCursorBlink` — so parking `cursorBlink` does not
reliably stop a hidden pane blinking. Pre-existing, documented only.
Co-authored-by: Wooseong Kim <innocarpe@users.noreply.github.com>
* feat(runtime): stream file uploads instead of buffering whole files
Staging read each dropped file whole with readFile(), base64-encoded it
(a 4/3 expansion), and passed the string through IPC to the renderer,
which re-chunked it. Peak memory was ~2.3x the file size before a byte
moved, so a 25 MB per-file cap existed to protect the heap.
Staging now records identity only. The byte pump moves into main, where
the file handle and the runtime socket both live: 384 KiB slices (512 KiB
once base64-encoded, matching the chunk size the renderer used) appended
through the existing files.writeBase64Chunk RPC. Peak memory is one slice
regardless of file size, so the ceilings become user-safety limits on an
unattended transfer — 2 GB per file, 8 GB per drop — and over-limit errors
name both the size and the limit.
Because staging and streaming are separate calls, the staged entry carries
size, inode, device and mtime, and the streamer re-checks all four against
the pre-open lstat and against the handle it actually reads. A source
replaced or rewritten at the same size between the two calls is refused
rather than uploaded under the original name. The post-read check compares
mtime as well as size, so an in-place rewrite mid-transfer aborts before
commitUpload renames anything into place.
O_NOFOLLOW, realpath containment and stat identity are preserved, and the
pairing revision plus the runtime id ride every chunk, so a re-pair or a
replacement runtime aborts instead of appending the rest of the file to a
different host.
No wire change: files.writeBase64Chunk and its params are untouched, so
old and new hosts behave identically. The SSH import path is separate and
unchanged. The web client has no local filesystem to stream from and says
so instead of failing obscurely.
* fix(runtime): close the empty-upload and per-drop budget holes
Two gaps the first pass left open.
A zero-byte source returned before the post-transfer identity check, so a
file that gained content during the empty write's round trip committed as
an empty file at the user's chosen name. The empty chunk now falls through
to the same final check the slice loop uses.
Each staged source also started its own byte counter, so the 8 GB ceiling
capped one source rather than the drop: five 2 GB files staged cleanly at
10 GB total. The IPC handler now carries one budget across sourcePaths and
adds only what each source actually staged. The per-file ceiling is still
re-enforced where the bytes move; the drop total holds at staging because
identity enforcement means each file streams exactly the bytes measured.
* docs(runtime): name the invariants the upload helpers carry
* fix(runtime): name the source in errors and stop uploads with their window
Three problems an independent review turned up.
A dropped file's relative path is '', so the over-limit error read "'' is
3 GB, over the 2 GB per-file remote import limit" — the message this change
exists to fix, naming nothing. Errors now fall back to the file's own name;
the staged entry keeps '' so the destination path is unaffected. The
streamer had the same shape, falling back to the hidden .orca-upload-<nonce>
temp destination, a path the user never chose.
The byte loop used to live in the renderer and died with it. Moving it into
main meant closing or reloading the window left the rest of a multi-GB
transfer running, with the renderer's temp cleanup never reaching its
finally. An AbortSignal now rides the caller's lifetime and every chunk, is
re-checked per slice, and main sweeps the abandoned temp path itself when
the renderer is no longer there to do it.
Upload failures also reached the import result wrapped in Electron's
"Error invoking remote method '...'" prefix, because the throw crossed IPC
instead of happening in-renderer; extractIpcErrorMessage unwraps it.
An existing staging test asserted the empty-name message, so it encoded the
bug rather than catching it; it now asserts the file name.
* test(runtime): cover the containment check and the per-chunk host guards
The "escapes the dropped root" test only reached the lstat symlink guard,
so assertEntryInsideRoot had no coverage at all. The shape that actually
needs it is a regular file under a symlinked intermediate directory: lstat
sees a plain file, and realpath containment is the only thing that refuses
it. Disabling the guard now fails this test and nothing else.
Nothing asserted that the SSH target, connection generation and execution
host reach the writeBase64Chunk params either — the renderer tests stop at
the IPC boundary, so the streamer's half of that contract was untested.
* fix(runtime): survive a straggling append when sweeping an aborted upload
Aborting rejects the in-flight chunk locally, but the host may still apply
that append, and appends open with flag 'a' — which recreates the file the
sweep just deleted. The delete and the straggler also race: they are
separate calls on a queue that is not ordered between them.
Slices are strictly sequential, so at most one append can be outstanding.
A second pass after it has had time to land is therefore sufficient, not
merely a heuristic. The sweep moves out of filesystem-mutations.ts into its
own module so the behaviour is testable directly.
Found by an independent review pass, which also pointed out that the
"escapes the dropped root" test only reached the lstat symlink guard.
* fix(runtime): abort uploads only when the document commits, and honour manual disconnect per chunk
did-start-navigation fires before will-navigate blocks an external link or a
stray file drop, and the renderer survives those (verified against Electron 43
with a hidden window). Aborting there killed a healthy upload with a misleading
'window went away' error. did-navigate fires only once a new document has
replaced the caller.
The renderer's per-chunk calls used to go through the IPC handler that refuses
a manually disconnected environment; the loop in main made no such check, so a
disconnect mid-upload kept pushing the rest of the file. The handler now
resolves the selector to an environment id and the streamer checks it per slice.
Adds slice-boundary coverage against the real chunk schema and host write
flags, staging-to-stream on a real filesystem, and handler-level lifetime tests.
---------
Co-authored-by: Neil <neil@stably.ai>
* fix(agents): find OMP by its full project name
* test(agents): make picker baseline proof omit OMP aliases
* style(test): brace picker baseline condition
* fix(codex): settle a structured send on admission, and stop minting a colliding identity
Two sends could be written into the journal under one durable identity.
Codex coalesces a mid-turn `turn/start` into the running turn rather than
refusing it -- measured against real `codex app-server` builds 0.147.0,
0.150.1 and 0.153.4, none of which refuse and none of which fire a second
`turn/started`. The dispatch path read the turn id from the turn/start
response and stamped every accepted send `ordinal: 0`. Since a coalesced
send gets the running turn's id back, two submissions persisted the same
`providerItemId`. That string is durable, and it is the key a restore uses
to match a submission against provider history, so the second message's real
history row matched nothing and rendered as an extra bubble on replay.
On 0.147.0 it is worse than a collision: the coalesced response returns a
turn id that never starts and never completes, so the persisted key named a
turn absent from history and NEITHER message could match.
Identity is now minted from the echoed user message at `identityFor` -- the
single point that mints the journal row's own identity -- so the settled key
is by construction the one replay computes, rather than a parallel
calculation that can drift.
Dispatch returns `admitted` when the transport write completes; identity
settles on the echo through a channel that did not previously exist for
Codex. Waiters are keyed by client message id instead of being shifted off
the front of an array by arrival order, and they are cleared on session
close and child exit -- previously a timeout was the only thing that ever
ended one.
`TURN_ID_WAIT_MS` is deleted. It was never reachable on any build measured:
`readCodexTurnId` returns non-null on all three, so the 10s wait never
fired. The comment justifying it claimed older builds acknowledge before the
id exists, which no tested build does.
Three comments asserting Codex answers a mid-turn send with `turn already
running` are corrected. Their only backing was a test fixture inventing that
error string. The correction is factual only -- every changed line in
`src/main/runtime/orchestration/` is a comment, and mid-turn delivery is
still refused for both providers. Whether that policy is right is a separate
question; it was resting on a false premise.
Known gap, stated rather than implied: this prevents new collisions and does
not repair journals already written with a colliding or phantom key. Those
conversations keep duplicating on restore. Repairing them means re-matching
persisted submissions against provider history and rewriting
`providerItemId` -- which is what `journal-submission-reconciler.ts` is
written for, and it still has no production caller.
* test(codex): drop the synchronous-accept contract and the colliding `:0` from the integration fakes
Three tests in the structured-session integration suites encoded the dispatch
contract this branch replaces, and two of them pinned the defect it fixes.
They asserted `agentSession.send` answers `dispatchState: 'accepted'` carrying
`providerItemId: codex:<thread>:<turn>:0` at send time. That ordinal was never
observed; it was stamped on every accepted send, which is exactly the collision
this branch removes -- a send coalesced into a running turn is answered with the
running turn's id, so two submissions persisted one durable key.
The visible failure was a 30s timeout rather than a failed assertion. The fake
client advertised no `agent-session.pending-send-result.v1`, and without it the
host holds the reply until the send settles: a shim for clients too old to
render a pending bubble. The fake provider then echoed the user message with no
`clientId`, so nothing could correlate that echo back to the submission, and the
wait ran to its own 30s ceiling. Real Codex sends `clientId` on that echo, and
the fake now does too, which is what makes it a model of the provider rather
than a sketch of one.
The identity assertion is kept rather than dropped. Each send now asserts
`pending` with no identity at admission, then asserts the submission settles
`accepted` at `codex:<thread>:<turn>:0` once the echo lands. Same ordinal, but
earned from `identityFor` on the echo -- the key a replay recomputes -- instead
of guessed from the turn/start response. Ablated: removing `clientId` from the
two echoes leaves both submissions `pending` and fails both assertions, so the
assertion is load-bearing and not satisfied by something incidental.
Both suites' client fixtures now advertise the capability set the desktop
renderer sends in `src/main/ipc/runtime.ts`, which is what these suites mean by
a client. The older-client settlement wait keeps its own coverage in
`src/main/runtime/rpc/methods/structured-agent-session.test.ts`.
`structured-agent-session-runtime-exit.test.ts` asserts `pending` for the same
reason; it drives the host directly, so it never took the compatibility path,
and what proves delivery there is still the turn the reacquired provider starts.
The replay suite's "without dispatching it twice" property is untouched: one
`turn/start` call, one replayed ledger row.
* fix(codex): preserve unsettled dispatch correlations
* test(codex): type the dispatch fixtures instead of asserting over them
main's new casting gate (#20367 base) flags type assertions on changed
lines. Replace them with checked types: the recording sink already
satisfies its interface, both CodexSession fixtures are now annotated and
carry real collaborators, the settlement assertion compares whole
identities, and the integration helper reads submissions through the
host's public journalSnapshot instead of its private session map.
* fix(test): merge the duplicate doubt-reasons import the merge left behind
Both sides added an import from journal-dispatch-doubt-reasons and the
merge kept both statements, which the whole-repo native plugin gate
refuses under --deny-warnings.
* test(codex): a Fast mode turn is admitted, not accepted
#20506 landed its Fast mode tests against the dispatch contract this
branch replaces: a Codex send now returns admitted and settles its
identity on the provider echo. The tier assertions the test exists for
are untouched.
---------
Co-authored-by: Merge Sim <sim@local>
* fix(store): stop two no-op writes from re-running every selector in the app
zustand bails out of a `set` only when `Object.is(next, state)`. Two updaters that
mean "nothing changed" hand it a fresh reference instead:
- `setWorkspacePortScanRefreshing` wrote unconditionally — the one action in its
file that did; its four siblings all early-return `state`.
- `applyGitHubPRRefreshEvent` ended its no-op branch with `: {}`, and
`Object.assign({}, state, {})` reproduces every field unchanged while still
notifying. ~20 sibling sites in the same store already use `return state`.
Both rebuild the root and wake every subscribed selector (~2.2k per the listener
census). Renders are unaffected — the selection is unchanged — so the cost is
wasted selector evaluation, not commit pressure. The resulting state looks
identical either way, which is why it goes unnoticed; both tests therefore count
subscriber notifications rather than asserting state.
* test(store): use checked initial state in notification regression
---------
Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Neil <neil@stably.ai>
* fix(source-control): prevent text wrapping in section headers and action
Use flex layout constraints (flex-1, shrink-0) and text truncation instead of
wrapping to keep section labels and action buttons on a single line in the
right sidebar.
* test(source-control): add section action button alignment tests
Ensure View all button stays on single line with icon actions in
crowded section headers. Pin layout constraints (shrink-0, flex-wrap,
whitespace-nowrap) to prevent regression.
* Rely on Button base styles for action label wrapping
Remove redundant shrink-0 and whitespace-nowrap utilities from
section action buttons. These should be supplied by the Button
component's base variant, not duplicated at each usage site.
* feat(native-chat): add copy button to code blocks
Enable users to copy code snippets directly from chat messages via a dedicated copy button on fenced code blocks. Supports language detection and integrates with markdown rendering via a `renderCodeBlock` prop.
* i18n: add English copy code button label
* refactor: use React.isValidElement type parameters for type narrowing
- Specify props types as type parameters to React.isValidElement instead
of casting after the fact
- Allows TypeScript to narrow element.props type automatically
- Eliminates manual type assertions in extractCodeText and extractCodeFenceLanguage
* fix(source-control): surface stage, unstage and discard failures
* fix(source-control): use single slot for entry failure toasts
- Consolidate entry failures to one stable slot instead of per-worktree
- Handle stale retries inline at click time rather than via a cleanup hook
- Remove retry button from discard failures to prevent destructive accidents
* test: improve type safety and mock patterns in source-control tests
- Add proper type definitions for toast options and test data instead of using `as never`
- Replace `mock.calls.at(-1)` with safer `mock.lastCall` pattern
- Create `entry()` helper to construct typed test entries
- Add explicit type annotations to mocked functions for better IDE support
* test: extract shared toast options type for source control tests
Consolidate duplicate `ToastOptions` type definitions across three test files into a single `SourceControlToastTestOptions` type, reducing duplication and improving consistency.
* fix(source-control): separate refresh failures from mutation failures
Post-mutation refresh failures are logged separately, not surfaced as toasts
(mutation already succeeded). Use preventDefault() on retry to prevent sonner's
auto-dismiss from swallowing re-raised failures. Consolidate stage/unstage into
a shared handler to reduce duplication.
* fix(source-control): only dismiss entry failures from the owning worktre
Track which worktree owns the shared entry-failure toast slot. When a mutation
completes, only dismiss the slot if the completing worktree is the one that
raised the failure — a slow retry in one worktree should not erase a failure
another worktree has since raised into the slot.
* Remove entry mutation status refresh helper
Inlined into the caller during consolidation of failure handling and
tracking in the source-control entry mutations flow.
* Simplify entry mutation refresh without wrapper
Call refreshActiveGitStatusAfterMutation directly instead of through the
refreshEntryMutationStatus helper. This ensures refresh failures propagate
directly from the callback without being caught as mutation failures.
Remove tests that validated the wrapper's error handling.
* fix(native-chat): keep a resumed transcript pinned to its end
Follow state was recomputed from distance on every scroll event, and a pin writes scrollTop itself, so the browser reports that write back as a scroll event a frame later. Once a resumed session's later history pages and settling row heights had moved the end away from it, that echoed event read as the reader leaving and the pin was dropped for good, stranding them mid-transcript. Measured in Chromium: a pin followed by same-task growth delivers a scroll event reading 2000px from the bottom, indistinguishable from a reader scrolling up.
Pins now go through the virtualizer instead of writing scrollTop directly, so both parties resolve the end through the same maximum rather than holding rival definitions of it. Whether the reader left is now a question of provenance rather than distance: an offset this transcript wrote is never a departure. The end test reads live geometry, because the virtualizer's own isAtEnd subtracts a cached offset from a live maximum and this handler runs before that cache is refreshed. overflow-anchor:none stops the engine moving scrollTop under a settling row, which would otherwise look like the reader.
This does not make ownership singular. The virtualizer still writes autonomously from several paths and those writes stay unattributed; what this removes is the rival definition of the end, not the second writer.
* fix(native-chat): cancel stale end reconciliation
* fix(native-chat): attribute scroll ownership centrally
`AgentJournalSubmission.reason` was the only unbounded field written by
Orca's own code. `dispatchSafely` sets it from the adapter's raw
`error.message` and `journalDispatchRowBuilder` stored it verbatim, so a
provider error carrying a multi-megabyte body -- a stringified HTTP error
payload, say -- reached the row at whatever length the provider sent, and
stayed on disk at that size for the life of the journal.
It now goes through `boundInlineText` with the journal's existing inline
limit, the same idiom already applied to arbitrary text on the Claude and
Codex translation paths.
The bound must stay head-preserving. `dispatchRejectionWasTransportWriteFailure`
prefix-matches the value, and `dispatchRejectionReasonIsInternal` builds on
it, so a bound that kept the tail instead would stop classifying a clipped
transport failure and render raw provider text to the user as an ordinary
rejection notice. A test pins that, and clipping stays marked rather than
silent so a truncated reason is never presented as the provider's complete
explanation.
Rows written before this keep their full text, so readers can still meet an
unbounded reason.