Commit Graph
10949 Commits
Author SHA1 Message Date
Wooseong KimandWooseong Kim 46eb5959fa fix(ui): contain idle caret paint so agent panes stop burning CPU (#10554)
An idle agent pane kept ~40% of a core busy just by being frontmost. The
xterm cursor and the native chat caret blink with no paint-containment
boundary, so Chromium treated each blink as damage to the whole pane
ancestry and re-rasterized it twice a second.

- `.xterm-container` and the native composer's input shell get
  `contain: paint`, bounding blink damage to the surface that blinks.
- The mention hint gains `z-20` to match the slash picker: a contained
  element becomes a stacking context and paints at z-index 0 in tree
  order, which would otherwise cover the hint's drop shadow.

Also records that DECSCUSR pins `decPrivateModes.cursorBlink`, which wins
over the option in `_updateCursorBlink` — so parking `cursorBlink` does not
reliably stop a hidden pane blinking. Pre-existing, documented only.

Co-authored-by: Wooseong Kim <innocarpe@users.noreply.github.com>
2026-09-14 14:27:31 -07:00
NeilandBrennan Benson 389d672dab fix(omp): preserve saved conversation names in session history (#20636)
Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
2026-09-14 14:17:04 -07:00
mmarabelandNeil 68f0b2e835 feat(runtime): stream file uploads instead of buffering whole files (#16106)
* feat(runtime): stream file uploads instead of buffering whole files

Staging read each dropped file whole with readFile(), base64-encoded it
(a 4/3 expansion), and passed the string through IPC to the renderer,
which re-chunked it. Peak memory was ~2.3x the file size before a byte
moved, so a 25 MB per-file cap existed to protect the heap.

Staging now records identity only. The byte pump moves into main, where
the file handle and the runtime socket both live: 384 KiB slices (512 KiB
once base64-encoded, matching the chunk size the renderer used) appended
through the existing files.writeBase64Chunk RPC. Peak memory is one slice
regardless of file size, so the ceilings become user-safety limits on an
unattended transfer — 2 GB per file, 8 GB per drop — and over-limit errors
name both the size and the limit.

Because staging and streaming are separate calls, the staged entry carries
size, inode, device and mtime, and the streamer re-checks all four against
the pre-open lstat and against the handle it actually reads. A source
replaced or rewritten at the same size between the two calls is refused
rather than uploaded under the original name. The post-read check compares
mtime as well as size, so an in-place rewrite mid-transfer aborts before
commitUpload renames anything into place.

O_NOFOLLOW, realpath containment and stat identity are preserved, and the
pairing revision plus the runtime id ride every chunk, so a re-pair or a
replacement runtime aborts instead of appending the rest of the file to a
different host.

No wire change: files.writeBase64Chunk and its params are untouched, so
old and new hosts behave identically. The SSH import path is separate and
unchanged. The web client has no local filesystem to stream from and says
so instead of failing obscurely.

* fix(runtime): close the empty-upload and per-drop budget holes

Two gaps the first pass left open.

A zero-byte source returned before the post-transfer identity check, so a
file that gained content during the empty write's round trip committed as
an empty file at the user's chosen name. The empty chunk now falls through
to the same final check the slice loop uses.

Each staged source also started its own byte counter, so the 8 GB ceiling
capped one source rather than the drop: five 2 GB files staged cleanly at
10 GB total. The IPC handler now carries one budget across sourcePaths and
adds only what each source actually staged. The per-file ceiling is still
re-enforced where the bytes move; the drop total holds at staging because
identity enforcement means each file streams exactly the bytes measured.

* docs(runtime): name the invariants the upload helpers carry

* fix(runtime): name the source in errors and stop uploads with their window

Three problems an independent review turned up.

A dropped file's relative path is '', so the over-limit error read "'' is
3 GB, over the 2 GB per-file remote import limit" — the message this change
exists to fix, naming nothing. Errors now fall back to the file's own name;
the staged entry keeps '' so the destination path is unaffected. The
streamer had the same shape, falling back to the hidden .orca-upload-<nonce>
temp destination, a path the user never chose.

The byte loop used to live in the renderer and died with it. Moving it into
main meant closing or reloading the window left the rest of a multi-GB
transfer running, with the renderer's temp cleanup never reaching its
finally. An AbortSignal now rides the caller's lifetime and every chunk, is
re-checked per slice, and main sweeps the abandoned temp path itself when
the renderer is no longer there to do it.

Upload failures also reached the import result wrapped in Electron's
"Error invoking remote method '...'" prefix, because the throw crossed IPC
instead of happening in-renderer; extractIpcErrorMessage unwraps it.

An existing staging test asserted the empty-name message, so it encoded the
bug rather than catching it; it now asserts the file name.

* test(runtime): cover the containment check and the per-chunk host guards

The "escapes the dropped root" test only reached the lstat symlink guard,
so assertEntryInsideRoot had no coverage at all. The shape that actually
needs it is a regular file under a symlinked intermediate directory: lstat
sees a plain file, and realpath containment is the only thing that refuses
it. Disabling the guard now fails this test and nothing else.

Nothing asserted that the SSH target, connection generation and execution
host reach the writeBase64Chunk params either — the renderer tests stop at
the IPC boundary, so the streamer's half of that contract was untested.

* fix(runtime): survive a straggling append when sweeping an aborted upload

Aborting rejects the in-flight chunk locally, but the host may still apply
that append, and appends open with flag 'a' — which recreates the file the
sweep just deleted. The delete and the straggler also race: they are
separate calls on a queue that is not ordered between them.

Slices are strictly sequential, so at most one append can be outstanding.
A second pass after it has had time to land is therefore sufficient, not
merely a heuristic. The sweep moves out of filesystem-mutations.ts into its
own module so the behaviour is testable directly.

Found by an independent review pass, which also pointed out that the
"escapes the dropped root" test only reached the lstat symlink guard.

* fix(runtime): abort uploads only when the document commits, and honour manual disconnect per chunk

did-start-navigation fires before will-navigate blocks an external link or a
stray file drop, and the renderer survives those (verified against Electron 43
with a hidden window). Aborting there killed a healthy upload with a misleading
'window went away' error. did-navigate fires only once a new document has
replaced the caller.

The renderer's per-chunk calls used to go through the IPC handler that refuses
a manually disconnected environment; the loop in main made no such check, so a
disconnect mid-upload kept pushing the rest of the file. The handler now
resolves the selector to an environment id and the streamer checks it per slice.

Adds slice-boundary coverage against the real chunk schema and host write
flags, staging-to-stream on a real filesystem, and handler-level lifetime tests.

---------

Co-authored-by: Neil <neil@stably.ai>
2026-09-14 14:08:31 -07:00
Neil f21f81dcfc fix(agents): find OMP by its full project name (#20647)
* fix(agents): find OMP by its full project name

* test(agents): make picker baseline proof omit OMP aliases

* style(test): brace picker baseline condition
2026-09-14 14:03:11 -07:00
Neil 8d93505958 fix(terminal): retain renames before renderer pane hydration (#20619)
* fix(terminal): retain renames before renderer pane hydration

* test(terminal): keep late renames from recreating closed tabs
2026-09-14 13:57:59 -07:00
Neil ee1a0a4e2d fix(git): avoid Windows tree kills after the command has exited (#20606)
Validated and independently reviewed OMP integration fix.
2026-09-14 13:56:25 -07:00
Neil fc4519cda4 fix(omp): preserve zsh startup with global aliases (#20621)
Validated and independently reviewed OMP integration fix.
2026-09-14 13:56:22 -07:00
Neilandshahidbeig-a11y 3632311d0b fix(omp): preserve status after terminal title owner rewrite (#20610)
Validated and independently reviewed OMP integration fix.

Co-authored-by: shahidbeig-a11y <258701601+shahidbeig-a11y@users.noreply.github.com>
2026-09-14 13:56:18 -07:00
Neilandplotarmordev 59d29af402 test: add a verified OMP native-chat mock scenario (#20655)
Co-authored-by: plotarmordev <plotarmordev@users.noreply.github.com>
2026-09-14 13:43:51 -07:00
Brennan BensonandMerge Sim 955051ded0 fix(codex): settle a structured send on admission, and stop minting a colliding identity (#20138)
* fix(codex): settle a structured send on admission, and stop minting a colliding identity

Two sends could be written into the journal under one durable identity.

Codex coalesces a mid-turn `turn/start` into the running turn rather than
refusing it -- measured against real `codex app-server` builds 0.147.0,
0.150.1 and 0.153.4, none of which refuse and none of which fire a second
`turn/started`. The dispatch path read the turn id from the turn/start
response and stamped every accepted send `ordinal: 0`. Since a coalesced
send gets the running turn's id back, two submissions persisted the same
`providerItemId`. That string is durable, and it is the key a restore uses
to match a submission against provider history, so the second message's real
history row matched nothing and rendered as an extra bubble on replay.

On 0.147.0 it is worse than a collision: the coalesced response returns a
turn id that never starts and never completes, so the persisted key named a
turn absent from history and NEITHER message could match.

Identity is now minted from the echoed user message at `identityFor` -- the
single point that mints the journal row's own identity -- so the settled key
is by construction the one replay computes, rather than a parallel
calculation that can drift.

Dispatch returns `admitted` when the transport write completes; identity
settles on the echo through a channel that did not previously exist for
Codex. Waiters are keyed by client message id instead of being shifted off
the front of an array by arrival order, and they are cleared on session
close and child exit -- previously a timeout was the only thing that ever
ended one.

`TURN_ID_WAIT_MS` is deleted. It was never reachable on any build measured:
`readCodexTurnId` returns non-null on all three, so the 10s wait never
fired. The comment justifying it claimed older builds acknowledge before the
id exists, which no tested build does.

Three comments asserting Codex answers a mid-turn send with `turn already
running` are corrected. Their only backing was a test fixture inventing that
error string. The correction is factual only -- every changed line in
`src/main/runtime/orchestration/` is a comment, and mid-turn delivery is
still refused for both providers. Whether that policy is right is a separate
question; it was resting on a false premise.

Known gap, stated rather than implied: this prevents new collisions and does
not repair journals already written with a colliding or phantom key. Those
conversations keep duplicating on restore. Repairing them means re-matching
persisted submissions against provider history and rewriting
`providerItemId` -- which is what `journal-submission-reconciler.ts` is
written for, and it still has no production caller.

* test(codex): drop the synchronous-accept contract and the colliding `:0` from the integration fakes

Three tests in the structured-session integration suites encoded the dispatch
contract this branch replaces, and two of them pinned the defect it fixes.

They asserted `agentSession.send` answers `dispatchState: 'accepted'` carrying
`providerItemId: codex:<thread>:<turn>:0` at send time. That ordinal was never
observed; it was stamped on every accepted send, which is exactly the collision
this branch removes -- a send coalesced into a running turn is answered with the
running turn's id, so two submissions persisted one durable key.

The visible failure was a 30s timeout rather than a failed assertion. The fake
client advertised no `agent-session.pending-send-result.v1`, and without it the
host holds the reply until the send settles: a shim for clients too old to
render a pending bubble. The fake provider then echoed the user message with no
`clientId`, so nothing could correlate that echo back to the submission, and the
wait ran to its own 30s ceiling. Real Codex sends `clientId` on that echo, and
the fake now does too, which is what makes it a model of the provider rather
than a sketch of one.

The identity assertion is kept rather than dropped. Each send now asserts
`pending` with no identity at admission, then asserts the submission settles
`accepted` at `codex:<thread>:<turn>:0` once the echo lands. Same ordinal, but
earned from `identityFor` on the echo -- the key a replay recomputes -- instead
of guessed from the turn/start response. Ablated: removing `clientId` from the
two echoes leaves both submissions `pending` and fails both assertions, so the
assertion is load-bearing and not satisfied by something incidental.

Both suites' client fixtures now advertise the capability set the desktop
renderer sends in `src/main/ipc/runtime.ts`, which is what these suites mean by
a client. The older-client settlement wait keeps its own coverage in
`src/main/runtime/rpc/methods/structured-agent-session.test.ts`.

`structured-agent-session-runtime-exit.test.ts` asserts `pending` for the same
reason; it drives the host directly, so it never took the compatibility path,
and what proves delivery there is still the turn the reacquired provider starts.

The replay suite's "without dispatching it twice" property is untouched: one
`turn/start` call, one replayed ledger row.

* fix(codex): preserve unsettled dispatch correlations

* test(codex): type the dispatch fixtures instead of asserting over them

main's new casting gate (#20367 base) flags type assertions on changed
lines. Replace them with checked types: the recording sink already
satisfies its interface, both CodexSession fixtures are now annotated and
carry real collaborators, the settlement assertion compares whole
identities, and the integration helper reads submissions through the
host's public journalSnapshot instead of its private session map.

* fix(test): merge the duplicate doubt-reasons import the merge left behind

Both sides added an import from journal-dispatch-doubt-reasons and the
merge kept both statements, which the whole-repo native plugin gate
refuses under --deny-warnings.

* test(codex): a Fast mode turn is admitted, not accepted

#20506 landed its Fast mode tests against the dispatch contract this
branch replaces: a Codex send now returns admitted and settles its
identity on the provider echo. The tier assertions the test exists for
are untouched.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-14 13:37:13 -07:00
Neil dede24df46 fix(store): preserve state identity for no-op updater branches (#20703) 2026-09-14 13:33:19 -07:00
Neil 1ba9801574 fix(ci): stop hourly versions dropping below a tagged or already-shipped build (#20699)
* fix(ci): stop hourly versions dropping below a tagged or already-shipped build

Hourly/daily/adhoc based their X.Y.Z on GitHub releases, not git tags. When
v1.4.202 was tagged and then its GitHub release vanished, the next hourlies
shipped as 1.4.202-hourly — below both stable 1.4.202 and the 1.4.203-hourly
builds already installed, so electron-updater stopped offering updates.

Read main's v* tags and already-published channel tags instead.

* docs(ci): record that 1.4.202's release was unpublished for a bug

The leftover tag is what hourly must still honor; this was not a failed cut.
2026-09-14 13:33:04 -07:00
Neil c3372aeadc fix(ai-vault): bound streamed remote JSONL records (#20700) 2026-09-14 13:30:40 -07:00
2186a885dd fix(store): stop two no-op writes from re-running every selector in the app (#20641)
* fix(store): stop two no-op writes from re-running every selector in the app

zustand bails out of a `set` only when `Object.is(next, state)`. Two updaters that
mean "nothing changed" hand it a fresh reference instead:

- `setWorkspacePortScanRefreshing` wrote unconditionally — the one action in its
  file that did; its four siblings all early-return `state`.
- `applyGitHubPRRefreshEvent` ended its no-op branch with `: {}`, and
  `Object.assign({}, state, {})` reproduces every field unchanged while still
  notifying. ~20 sibling sites in the same store already use `return state`.

Both rebuild the root and wake every subscribed selector (~2.2k per the listener
census). Renders are unaffected — the selection is unchanged — so the cost is
wasted selector evaluation, not commit pressure. The resulting state looks
identical either way, which is why it goes unnoticed; both tests therefore count
subscriber notifications rather than asserting state.

* test(store): use checked initial state in notification regression

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Neil <neil@stably.ai>
2026-09-14 13:23:34 -07:00
Jinjing 6c70801f72 fix(source-control): prevent text wrapping in section headers and action buttons (#20046)
* fix(source-control): prevent text wrapping in section headers and action

Use flex layout constraints (flex-1, shrink-0) and text truncation instead of
wrapping to keep section labels and action buttons on a single line in the
right sidebar.

* test(source-control): add section action button alignment tests

Ensure View all button stays on single line with icon actions in
crowded section headers. Pin layout constraints (shrink-0, flex-wrap,
whitespace-nowrap) to prevent regression.

* Rely on Button base styles for action label wrapping

Remove redundant shrink-0 and whitespace-nowrap utilities from
section action buttons. These should be supplied by the Button
component's base variant, not duplicated at each usage site.
2026-09-14 13:07:27 -07:00
Jinjing d5be0d69e7 Add copy button to code blocks (#20357)
* feat(native-chat): add copy button to code blocks

Enable users to copy code snippets directly from chat messages via a dedicated copy button on fenced code blocks. Supports language detection and integrates with markdown rendering via a `renderCodeBlock` prop.

* i18n: add English copy code button label

* refactor: use React.isValidElement type parameters for type narrowing

- Specify props types as type parameters to React.isValidElement instead
  of casting after the fact
- Allows TypeScript to narrow element.props type automatically
- Eliminates manual type assertions in extractCodeText and extractCodeFenceLanguage
2026-09-14 12:54:31 -07:00
Jinjing 2acd2f4c88 Surface stage, unstage and discard failures with retry capability (#20423)
* fix(source-control): surface stage, unstage and discard failures

* fix(source-control): use single slot for entry failure toasts

- Consolidate entry failures to one stable slot instead of per-worktree
- Handle stale retries inline at click time rather than via a cleanup hook
- Remove retry button from discard failures to prevent destructive accidents

* test: improve type safety and mock patterns in source-control tests

- Add proper type definitions for toast options and test data instead of using `as never`
- Replace `mock.calls.at(-1)` with safer `mock.lastCall` pattern
- Create `entry()` helper to construct typed test entries
- Add explicit type annotations to mocked functions for better IDE support

* test: extract shared toast options type for source control tests

Consolidate duplicate `ToastOptions` type definitions across three test files into a single `SourceControlToastTestOptions` type, reducing duplication and improving consistency.

* fix(source-control): separate refresh failures from mutation failures

Post-mutation refresh failures are logged separately, not surfaced as toasts
(mutation already succeeded). Use preventDefault() on retry to prevent sonner's
auto-dismiss from swallowing re-raised failures. Consolidate stage/unstage into
a shared handler to reduce duplication.

* fix(source-control): only dismiss entry failures from the owning worktre

Track which worktree owns the shared entry-failure toast slot. When a mutation
completes, only dismiss the slot if the completing worktree is the one that
raised the failure — a slow retry in one worktree should not erase a failure
another worktree has since raised into the slot.

* Remove entry mutation status refresh helper

Inlined into the caller during consolidation of failure handling and
tracking in the source-control entry mutations flow.

* Simplify entry mutation refresh without wrapper

Call refreshActiveGitStatusAfterMutation directly instead of through the
refreshEntryMutationStatus helper. This ensures refresh failures propagate
directly from the callback without being caught as mutation failures.
Remove tests that validated the wrapper's error handling.
2026-09-14 12:45:25 -07:00
Brennan Benson d8b6151e8c fix(native-chat): keep a resumed transcript pinned to its end (#20651)
* fix(native-chat): keep a resumed transcript pinned to its end

Follow state was recomputed from distance on every scroll event, and a pin writes scrollTop itself, so the browser reports that write back as a scroll event a frame later. Once a resumed session's later history pages and settling row heights had moved the end away from it, that echoed event read as the reader leaving and the pin was dropped for good, stranding them mid-transcript. Measured in Chromium: a pin followed by same-task growth delivers a scroll event reading 2000px from the bottom, indistinguishable from a reader scrolling up.

Pins now go through the virtualizer instead of writing scrollTop directly, so both parties resolve the end through the same maximum rather than holding rival definitions of it. Whether the reader left is now a question of provenance rather than distance: an offset this transcript wrote is never a departure. The end test reads live geometry, because the virtualizer's own isAtEnd subtracts a cached offset from a live maximum and this handler runs before that cache is refreshed. overflow-anchor:none stops the engine moving scrollTop under a settling row, which would otherwise look like the reader.

This does not make ownership singular. The virtualizer still writes autonomously from several paths and those writes stay unattributed; what this removes is the rival definition of the end, not the second writer.

* fix(native-chat): cancel stale end reconciliation

* fix(native-chat): attribute scroll ownership centrally
2026-09-14 12:22:44 -07:00
Brennan Benson b4d435806f fix(native-chat): bound a dispatch reason before it reaches the journal row (#20654)
`AgentJournalSubmission.reason` was the only unbounded field written by
Orca's own code. `dispatchSafely` sets it from the adapter's raw
`error.message` and `journalDispatchRowBuilder` stored it verbatim, so a
provider error carrying a multi-megabyte body -- a stringified HTTP error
payload, say -- reached the row at whatever length the provider sent, and
stayed on disk at that size for the life of the journal.

It now goes through `boundInlineText` with the journal's existing inline
limit, the same idiom already applied to arbitrary text on the Claude and
Codex translation paths.

The bound must stay head-preserving. `dispatchRejectionWasTransportWriteFailure`
prefix-matches the value, and `dispatchRejectionReasonIsInternal` builds on
it, so a bound that kept the tail instead would stop classifying a clipped
transport failure and render raw provider text to the user as an ordinary
rejection notice. A test pins that, and clipping stays marked rather than
silent so a truncated reason is never presented as the provider's complete
explanation.

Rows written before this keep their full text, so readers can still meet an
unbounded reason.
2026-09-14 12:14:07 -07:00
Brennan Benson 1d1bca2a7b Bump mobile app.json to 0.0.50 (#20661) 2026-09-14 12:11:17 -07:00
github-actions[bot] 875b86d168 Update README downloads badge 2026-09-14 18:30:25 +00:00
Brennan BensonandMerge Sim c6a7216984 fix(native-chat): hide activity while awaiting input (#20496)
* fix(native-chat): hide activity while awaiting input

* fix(native-chat): keep approval turns cancellable

* test(native-chat): satisfy split PR quality gate

* fix(native-chat): catalog approval cancellation label

* fix(native-chat): include approval cancellation runtime label

* fix(codex): settle prompts when cancelled turns complete

* fix(codex): settle prompt registry fallbacks

* test(native-chat): cover pending interaction fallbacks

* test(native-chat): split prompt state coverage

* test(native-chat): keep prompt state isolated

* fix(native-chat): bound prompt turn backfill

* refactor(codex): centralize prompt registry bounds

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-14 10:42:38 -07:00
Jinwoo Hong eba56f2f69 feat(ai-vault-search): construct the session search indexer in the scanner service behind a setting (#20516)
* feat(ai-vault-search): persist agent-session search consent and retention

Two booleans and nothing else: `enabled` and `historyDays`, off by default
because building the index reads every transcript on the machine. No `paused` --
the PR 3 indexer is immutable, so every change is close-and-construct.

The settings IPC normalizes a write like every other field and hands the change
to the index; there is no UI for it until PR 8.

* feat(ai-vault-search): hold one indexer and engine pair per host

The object that owns a host's live index and the three recipes that change it.
The indexer is immutable, so a settings change is close-and-construct, disabling
is close with no replacement, and clearing is close, remove the database,
construct. The new instance's first sweep purges a narrowed window and admits a
widened one, so neither needs a code path.

The database sits beside the scanner's parse cache, one file per host. A runtime
with no node:sqlite can hold no index at all, which the Node 18 floor on orcad
and the relay makes a real case rather than a hypothetical one.

* feat(ai-vault): let the scanner child own the session search index

The transcript reader runs in that child, so the index consumer has to as well:
one read serves both the session list and the index. Three request operations
(search, status, reconcile) and one fire-and-forget settings message carry
everything a parent needs; main never opens the database file.

The init frame becomes a factory because it is read at every spawn, so a
respawned child sees current consent rather than the first frame's. A child
holding a running index is never idle from the parent's side, so idle retirement
is suppressed while the index is on -- retiring it would stop the reconcile loop
until some later scan happened to respawn one.

Both files this lands in were already at the max-lines ceiling, so three
collaborators move to where they belong rather than being disabled around: the
invalidation deadline into the class that owns invalidations, call cancellation
and the start requeue into the call-state module, and orcad's flag parsing into
its own file.

* feat(ai-vault-search): register a search service on every host that answers

Without a registered service a host answers no-service, which means "this host
does not have the feature" rather than "the index is off". All three hosts now
answer the second thing.

The desktop forwards to the scanner child. orcad and the SSH relay daemon have
no such child -- orcad ships only the watcher and daemon entries, and the relay's
AI Vault sidecar runs the remote scanner, which publishes nothing to the
transcript channel -- so on those two the index lives in the process that would
drive its reads, gated on a runtime that has node:sqlite at all.

The relay registers with consent off and no way to turn it on: nothing carries a
setting to a remote host yet. That is the honest state, and it is still worth
registering, because it is what tells a client the difference between off and
too old.

* test(ai-vault-search): price a warm pass over five thousand transcripts

The number the reconcile interval will be revisited against, measured rather
than argued: a warm sweep stats every file under every root, a warm cycle stats
the newest N per agent, and neither reads what the index already holds. It does
not tune the interval.

* fix(ai-vault-search): answer the casting gate without assertions

main's new type-assertion rule reaches every file this branch touches. All nine
sites drop the cast rather than carry a SAFETY: rationale: the operation guard
narrows with `in`, the sqlite probe narrows the builtin it loads, the child test
keeps the discriminated reply instead of widening it, and the settings resolver
takes `unknown` -- which is what it really reads, since a persisted profile can
hold a value no version of this code wrote.

* fix(ai-vault-search): let a refreshed scan root reach the live index

The parent re-resolves scan roots before every policy push, precisely so a
WSL distro or extra Codex home that appeared since the child spawned enters
the window. The child forwarded only the settings to a live instance and used
the roots solely in its `??=` initializer, so those roots were dropped for the
child's lifetime.

The indexer stays immutable: a structurally different root set closes the pair
and constructs a new one, the same way a changed databasePath already does.
Compare via `sameSessionSearchRoots` rather than a plain JSON compare, because
nothing fixes the key order two producers write; lists are sorted too, since
the indexer walks every root and a re-enumeration that reorders is not a
change. An unchanged set still never restarts a running index.

The orcad and relay in-process hosts resolve roots once at install and never
re-apply, so they have no such seam.

* fix(ai-vault): restart the scanner child the index is holding

Three review items.

The hold keeps a child alive for the index, but only a queued call ever
started one: `pump()` skipped a hold with an empty queue, so an idle indexing
child that crashed, or an `ensureChild()` that failed at start, left indexing
stopped until an unrelated request happened to arrive. `pump()` now starts the
child the hold requires, which is also the restart callback the fault policy
already schedules, so the existing delay and circuit bound the retry exactly as
they bound a queued call's start. `updateSessionSearch` goes through the same
seam instead of its own `ensureChild` call.

A search registers no AbortController, so a cancel sent for a search id was
added to the `cancelled` set and never consumed. Nothing can reach that today
-- no caller passes a signal and the child answers in milliseconds -- so this
is only a leak of ids: consume it when the search settles.

The orcad argument doc claimed a `--`-prefixed value stays a flag. The parser
takes the next token regardless, and orcad-launch-contract.test.ts pins that,
so the doc is what was wrong. Behaviour is unchanged.

* fix(ai-vault): recover search indexing and refresh scan roots

* fix(ai-vault): defer search refresh policy reads

* fix(session-search): stabilize paging and host enablement

* fix(session-search): refresh host roots within full sweeps

* docs(session-search): clarify initial root fallback
2026-09-14 13:38:37 -04:00
Jinwoo Hong fc525c355d refactor(mobile): send the task workspace-creation domain through typed RpcOperations (#20568)
* test(mobile): record main's task workspace-creation RPC behaviour before migrating it

28 scenarios over nine task senders, recorded from main so the step-4 migration of
the workspace-creation half of src/tasks/ has a frozen answer to compare against.
Four senders mount as plain exported functions; three are model-chained hooks
mounted the way the settings adapters mount theirs.

The 153 existing goldens change header-only (`baseline`, `recorderSha256`): any new
scenario re-digests the recorder, and the pinned baseline had drifted from main
because the source-control migration landed. Content is byte-identical on all 153 —
verified field-by-field against HEAD.

`operation-module-loader.ts` now shares src/transport/rpc-delivery-ambiguity.ts with
mounted modules instead of evaluating a second copy. The mark is a WeakSet keyed on
the rejection object, so the copy the loader built had an empty registry and every
delivery-unknown rejection read as a definite failure inside the operation under
test — worktree.create's whole replay path was unreachable. With one registry,
`tw-create-retry-ambiguous-after-drop` records the create still pending at the
reconnect wait and abandoning at exactly 20000 ms, while the unstamped-create
scenario records the same rejection surfacing at 0 ms. No existing golden moves:
no other mounted module consumes the mark.

`task-preferences-optimistic` is re-anchored above the send rather than across it,
so migrating this file does not have to move the anchor. It still kills, and for
the same reason: the preset the screen shows no longer follows the tap.

Scenarios deliberately pin the empty-message refusals (`*-refused-empty-message`,
`*-empty-message`), because a refusal with no message falls back to the screen's
copy while a transport error with no message does not, and the two paths are easy
to collapse when a call site moves behind an acceptance policy.

Goldens: 153 -> 201, 2.9M -> 3.7M.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): format the recording manifest and re-digest the goldens

`oxfmt --check` from mobile/ collapses a one-element `sites` array in each new
scenario. The JSON value is unchanged — verified by comparing both files parsed
and key-sorted — but the manifest is inside `recorderSha256`, so all 201 goldens
carry a new digest. Every other field, header and observation alike, is
byte-identical.

Re-recorded in a separate worktree at the previous commit so the goldens stay
attributable to main's product source rather than to the migration that follows.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): separate the goldens from the migration, and re-digest

The previous commit accidentally carried the product migration alongside the
manifest format, which both broke the commit that is supposed to prove parity and
left the suite red: the digest was recorded without a comment move that a lint fix
had made inside the adapter, so all 201 goldens failed their `recorderSha256`
header.

This backs the product half straight out again — the next commit re-applies it
byte-for-byte — and re-records from the pinned baseline in a separate worktree
carrying this branch's recorder, per the procedure in the recording README. Every
field except `recorderSha256` is byte-identical to the previous commit's goldens on
all 201 files, so no observation moved in either direction. The suite is green here
with main's product source, which is what makes the next commit's "no golden
changed" claim mean something.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): send the task workspace-creation domain through typed RpcOperations

12 of src/tasks/'s 37 raw-port files now send through a declared operation instead of
the raw request port: 36 references to 0, leaving 25 files and 73 references for the
provider item/detail/mutation half. No golden moved — `git show --stat` on this commit
touches nothing under mobile/rpc-foundation/, which is the parity claim.

Twenty-two operations over twenty methods, in four modules named for what they send:
workspace create (create, PR/MR base resolution, create-time capabilities), workspace
source (SSH connect/state, agent detection, repo hooks, sparse presets, ref search),
task runtime (status, ui.get/ui.set, preflight, Linear status, settings.update) and the
Smart picker's provider reads.

Two methods carry two policies each, and both pairs are named. `status.get`: the Tasks
screen cannot hydrate without it and surfaces the host's message, while create-time
capability probing degrades to "no capabilities" and creates anyway — so one throws on
refusal and one skips. `ui.set`: two sites await it, one is fire-and-forget and never
interprets the reply at all. Both pairs share one reader, so no method has two readers.
No new acceptance policy.

worktree.create keeps its delivery-unknown contract. `request` returns the transport
promise itself, so the retry loop catches the object the transport marked; two new tests
assert `toBe(marked)` in one direction and that an unmarked rejection stays unmarked in
the other, because a mark added on the way out would replay a create the host never
received. `tw-create-retry-ambiguous-after-drop` records the create still pending at the
reconnect wait and abandoning at exactly 20000 ms.

Three sites still read the raw refusal envelope before interpreting, because the code or
the message decides the route and no acceptance policy carries either through: the create
retry needs the message for `isRetryableWorktreeCreateConflict`, and the paste lookup
needs `method_not_found` to retire the slug probe host-wide. Both are documented at the
site.

The hydration barrier keeps raw requests inside its `Promise.all`. main's group rejects as
soon as one leg rejects; `startRpcOperation` + `interpretAtRpcBarrier` would wait for the
slowest peer and let a later policy surface a different error. Interpretation stays after
the `stale` guard, where it was.

`WorkspaceCreateParams` is now `RpcSendParams<'worktree.create'>` rather than
`Record<string, unknown>`, which types the builder and the operation together; every field
the three builders already sent typechecks against the host schema unchanged.
`RpcSendArguments` now also makes params optional for a method whose params type has no
required field, because `preflight.check` is such a method and main sent it none —
requiring `{}` would have put a new object on the wire.

The Mobile Tasks source-parity hashes move for the same reason bound settings requests
moved them: the method string and the envelope read leave the screen. The signature diff
is evidence rather than a re-pin — `semantics` is a pure deletion of 22 `rpc:` call
signatures and 22 method literals with nothing added, statement/declaration/render/style
counts are unchanged, and render tokens, styles and declarations are byte-identical.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): split the task workspace adapters at the sender/hook seam

The single adapter file reached 344 lines against mobile's 300-line limit. CI lints
every file, so this is red there even though the changed-code gate does not report it.
Split along the seam the recording README already draws: exported async senders that
take a client and need no React host, and the drawer's three model-chained hooks.
No adapter body changed.

Both files are inside `recorderSha256`, so all 201 goldens carry a new digest. Every
other field is byte-identical, verified file by file. Re-recorded from the pinned
baseline in a separate worktree carrying this branch's recorder, so the goldens stay
attributable to main's product source.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): give the one-key unchecked reader a name

Four readers were the same three lines: read one property off the reply, wrap
it unchecked. `rpcUncheckedMemberReader` is the one-key sibling of the existing
`rpcUncheckedPayloadReader`, so the annotation and the closure go away at each
site. The pilot's `commitCompareEntriesReader` is converted too, so the helper
has no longhand twin left to copy from.

No behaviour change: the helper composes the same `rpcReadUnchecked` over
`rpcPayloadMember`, including the property-read exception on a null result.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): record the local arm of workspace agent detection

`preflight.detectAgents` was the one migrated operation with no recorded
coverage: the ssh adapter hardcoded `connectionId: 'ssh-1'`, so the detection
effect's ternary only ever took the remote arm and the local call site could be
repointed at another method without a golden noticing.

The adapter now takes the connectionId as a parameter and registers twice;
`tasks.workspace-ssh-local` mounts the same hook with no connection, which is
the only difference the effect branches on. Recorded at the pinned baseline
with this branch's recorder laid over it, so the new golden is main's
behaviour and the migrated code has to reproduce it — it does.

Goldens: two added (`tw-workspace-ssh-local-agents` and its reply matrix). The
other 201 changed on `recorderSha256` only, because the adapter edit moves the
recorder digest every golden pins.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): drive workspace create and the Linear list to a recorded wire

Two operations passed a policy swap unnoticed, both because no golden reached
their acceptance branch.

`worktree.create`: the create hook's fixture resolved setup to a prompt, so all
three settings.task-workspace scenarios stopped before the request and the only
consumer that hands a refusal to interpret was never recorded. The adapter now
takes the setup resolution as a parameter and registers a second family that
resolves it, so createWorkspace runs to the wire. Two scenarios: a Linear item
that creates directly, and a GitHub pull request that resolves its base first,
which also puts this hook's built params — start point, generated display name,
agent launch fields — in a golden for the first time. The existing prompt
family is untouched, so its recordings still pin that branch.

`linear.listIssues`: it appeared only in a non-base scenario, and the matrix
reads the family base, so the family had no partition for it. The base now
lists assigned issues after searching.

Goldens: five added. Five moved beyond the digest, all derived from the
smart-search base that gained the list leg. The other 198 changed on
`recorderSha256` only. Recorded at the pinned baseline with this branch's
recorder laid over it, so every new golden is main's behaviour.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): drop the unreachable unmount branches from the task adapters

Nothing dispatches `unmount` to these three adapters: the only producer is
`lifecycleSchedules`, driven from a hardcoded five-id list that names no
task-workspace family, and it pushes a `remount` right after, which these
adapters would throw on. The branch read as lifecycle coverage that was never
wired up. `dispose: hook.unmount` already tears the mount down.

Goldens re-recorded at the pinned baseline because the recorder digest moved;
`recorderSha256` is the only line that changed in all 208.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-record the goldens at main's post-squash baseline

Recorded from a detached checkout of e53f1557e1 (main's unmigrated
product code) with this branch's recorder laid over it, so the parity
claim stays non-circular.

- `baseline` repinned to e53f1557e1 on all 208 goldens; main pinned
  5ec0b2698f, a pre-squash branch commit not reachable from main.
- `recorderSha256` moved on all 208 because this branch's adapters are
  in the whole-manifest digest.
- 55 task-workspace goldens re-recorded at the new baseline.
- No other line in any of main's 153 goldens changed.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-record the goldens under #20562's per-scenario digest

Baseline repinned to 50e752fc66 and all 208 goldens recorded from that
commit's unmigrated product tree with this branch's recorder laid over it.
recorderSha256 moves on every golden because the task-workspace adapters
live in the recorder directory. scenarioSha256 does not move on any of
main's 153: the manifest only adds 31 scenarios and edits none, which is
the property #20562 was built to give.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-14 13:36:16 -04:00
Brennan Benson a4c11f1889 fix(native-chat): stop a bounded tail read from moving the chat cursor past unapplied rows (#20581)
* fix(native-chat): stop a bounded tail read from moving the chat cursor past unapplied rows

A structured chat pane could latch "Working for N" forever after the agent had
finished, showing the send arrow rather than Stop, while the sidebar and
`worktree ps` correctly read idle.

The client replica has one position (`state.cursor`) and one body. Two
operations keep those consistent: replace (both from one host snapshot) and
append (rows contiguous with the cursor). The `tail-page` branch was a third
thing: it took the cursor from the journal head, the items from a bounded page
(200 items, byte-capped), then merged retained client submissions over the
page's. Under continuous journal writes the client is always slightly behind,
so the branch ran on every window focus and on every pane re-activation. When
more than a page of rows had landed since a send, that send's user item fell
off the page, its submission was not carried, the retained `pending` survived,
and the cursor jumped past the dispatch-acceptance row. Nothing re-sends it: a
batch carries only touched items and that submission is never touched again.

Delete the third operation rather than guard it. A live subscription is now the
only thing that moves the cursor, and `subscribe({ cursor })` already replays
exactly the missed rows.

- remove the window `focus` listener and the owner/transport `refresh` contract
- skip warm hydration: a retained owner subscribes at its applied cursor
- cold hydration keeps its history read, applied as the existing `snapshot`
  (replace) event rather than `tail-page`
- delete the `tail-page` action and its reducer branch
- delete `resumeCursor` and `shouldAdvanceStructuredResumeCursor`; two cursors
  with two advancement rules were how position and body drifted apart

`older-page`/`loadOlder`, the unattached-refusal grace, generation guards and
the coalescer are unchanged. No host, wire or schema change.

Also fixes a second cost of the same branch: focus during a busy turn discarded
paged-in older items, shrinking the transcript to one bounded page mid-turn.

* fix(native-chat): preserve unavailable mixed-version session fences
2026-09-14 10:28:16 -07:00
Jinwoo Hong 50e752fc66 test(mobile): pin each RPC golden to its own scenario input, not the whole manifest (#20562)
* refactor(mobile): pin each golden to its own scenario input, not the whole manifest

`recorderSha256` covered the recorder directory plus `pilot-scenarios.json`, so every golden's
header was a function of every other family's scenarios. Adding a family for one domain re-digested
all 153 goldens and put a conflict on that line in every domain branch in flight, which serialized
the step-4 fan-out.

Split the two things it conflated. `recorderSha256` now covers the recorder directory only, with
unchanged semantics: a recorder edit still forces a full, deliberate re-record. A new
`scenarioSha256` pins the scenario input that golden was recorded from — every scenario
`runRecording` consumed for it, in order — canonicalised through `captureValue` so an
explicit-undefined param stays distinct from an absent one. `goldenRecording` takes that list
instead of just its first member.

The variants are hashed rather than the base they expand from because they are what was recorded: a
matrix site, its replayed normal result and its partition replies are all visible in them without
the derivation having to be restated. `derived-goldens.ts` is that derivation, extracted from
`family-recordings.test.ts` so the digest and the recording agree by construction — a property test
that restated how a matrix or schedule expands could agree with itself and with nothing else. It
reproduces exactly the 153 golden ids on disk, and the census the suite already ran (every family
matrixed, no stale normal-result inventory entry) now reads off its output.

`golden-header-digest.test.ts` pins the four properties:

- a new family in the manifest moves zero existing goldens' headers, and derives two of its own
- editing one field of `b1` moves exactly `b1` and its family's four matrix goldens — not the two
  other `legacy-inventory` scenarios, and not the goldens that expand from `inventory-lifecycle`
- editing a recorder file still moves every golden's `recorderSha256`, and no `scenarioSha256`
- `recorderSha256` is unchanged by the manifest's contents, and no longer reads the file at all

`GOLDEN_FORMAT_VERSION` goes to 4: a version-3 header has no `scenarioSha256`, and `compareGolden`
walks the expected header's keys, so a reader that accepted one would compare that golden's own
scenarios as though they were unpinned.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-record the 153 RPC goldens for the split digest

Recorder edit, so every golden needs rewriting. Recorded from the pinned baseline
`16d1ab81d3` with this branch's recorder overlaid, per the README's procedure: main has
moved past the baseline, so recording in place would have failed the product-source fence.

Three header fields moved and nothing else did:

- `recorderSha256` 6a12160a87… -> 2fda557f58…, one value across all 153 files
- `scenarioSha256` added, 153 distinct values
- `goldenFormatVersion` 3 -> 4

No observation, checkpoint, value-pool entry, `baseline`, `lockfileSha256` or `platform` changed:

    git diff -U0 -- mobile/rpc-foundation | grep -E '^[+-]' | grep -vE '^(\+\+\+|---)' \
      | grep -vcE 'recorderSha256|scenarioSha256|goldenFormatVersion'
    0

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): certify the pilot goldens from the derivation that digests them

`pilot-recordings.test.ts` restated `[scenario]` instead of consuming `pilotGoldens`, so the claim
that a golden's `scenarioSha256` is a function of the same derivation that records the file held
only for the 75 family goldens: dropping a scenario from `pilotGoldens` left the whole suite green
and put that golden outside the header oracle. The pilot suite now iterates `pilotGoldens`, and a
census fails if the derivation and the goldens directory disagree in either direction — which also
closes the pre-existing orphan-golden gap.

Also from review: pin the cross-sibling replay that hashing the generated variants buys (a matrix
golden's `normal` partition replays a sibling's recorded reply, so editing that sibling must move
it); state the real reason for the format bump, which is the diagnosis a version check gives rather
than a rejection the byte compare already made; drop the fourth property test, which re-proved what
tests 1 and 2 and `recording-runner`'s digest test already fail on; and drop a guard in
`scenarioSha256` that its only caller reaches after an identical one.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-record the 153 RPC goldens for the review edits

Recorder files changed, so `recorderSha256` moved. Recorded from the pinned baseline with this
branch's recorder laid over it, per the README's migration-branch procedure. That one header field
is the only line that moved in all 153 files: `scenarioSha256` and `goldenFormatVersion` are
unchanged, and no observation moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): correct the recording suite's test count

Round-2 review: the README said 200 tests; the suite is 209 after the five
added here. Markdown is outside recorderSha256, so no golden moves.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-record the 153 goldens on the merged baseline

Four header fields moved and nothing else. Proven against origin/main: every
changed line in all 153 files is one of these, and the file set is unchanged.

- `recorderSha256` 70aa6f59e0 -> 58a461dbc9: this branch's recorder, and it now
  digests only the recorder directory, not the scenario manifest.
- `scenarioSha256` added, 153 distinct values over 153 goldens.
- `goldenFormatVersion` 3 -> 4 for that added field.
- `baseline` 5ec0b2698f -> e53f1557e1, the merge's repoint onto the real main
  commit. #20563's value was a branch commit the squash left unreachable, so the
  record fence's `git diff <baseline>` could not resolve it.

No checkpoint, value pool, effect or settlement byte moved, so main's recorded
behaviour is carried over intact.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): account for main's added recorder test in the suite count

The merge brought in `unhandled-recording.test.ts`, one test, so the recording
suite is 210 rather than the 209 this branch documented. Markdown is excluded
from `recorderSha256`, so no golden moves.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-14 13:10:22 -04:00
Jinwoo Hong e53f1557e1 fix(mobile): two known main bugs the RPC migration preserved (#20563)
* fix(mobile): two known main bugs the RPC migration preserved

A malformed host `error` and a null settings result both reach a property read that throws.
Both are deliberate behaviour changes; the goldens move in the follow-up commit.

`hostReplyErrorTextOrFallback` passed a truthy non-string through under a `string` annotation.
Its one caller is the in-band `git.commit` failure, and every consumer of that text is display or
prompt copy: `use-mobile-create-pr-runner` and `PrSidebarCreateEmptyState` record it as a commit
failure, `use-mobile-commit-failure-recovery` hands it to `summarizeCommitFailure`, which starts
with `raw.slice(...).replace(...)`. So no consumer needs the value, and the decision is the
fallback rather than `String(value)` — the relay handler declares
`commit(): Promise<{ success: boolean; error?: string }>`, so a non-string is a malformed reply,
and `generatedCommitMessageReader` in the same domain already reads a non-string host error as
absent. The parameter stays `unknown`, which it honestly is, and the `SAFETY` cast is gone.

`useNewWorkspaceRuntimeContext` read settings through `settingsRead`, whose reader preserves
main's `boxed!.settings` throw, so a `null` or absent result threw a TypeError out of the effect —
losing the trusted-hooks publish and the available-provider computation that follow it, not just
the settings. It now uses `optionalSettingsRead`, the operation that already reads a null or
absent result as absent settings, so the reply degrades exactly the way a reply with no `settings`
member does. Reply-side only: same method, same params, same barrier, no wire change.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-record the goldens the two bug fixes move

Baseline bumped to 3f71999237. Two goldens move an observation; the other 151 move only
`baseline` and `recorderSha256`, which `pilot-scenarios.json` is still digested into.

Observation moves, one claim each:

- `matrix-hostedreview.create-intent-git.commit-1`, partition `inner-false-object-error`:
  `settlements.run.value.error` and `state.outcome.error` go from `{"message":"inner refused"}` to
  `"Commit failed"`. A non-string in-band `git.commit` error is a malformed reply and now reads as
  the screen's copy, converging with `result-absent`, `result-null` and `outer-refused-no-message`,
  which already reported the fallback. The other ten partitions at this site are unchanged.
- `matrix-settings.workspace-context-settings.get-1`, partitions `result-null` and `result-absent`:
  the `unhandled-rejection` TypeError effect (`reading 'settings'`) is gone and `state.providers`
  goes from `[]` to `["github"]`. The effect no longer aborts the rest of the hook, so the
  provider computation runs; `state.settings` stays null because nothing was published, which is
  how a reply with no `settings` member already degraded. The other nine partitions are unchanged.

Header-only moves:

- 9 goldens of the `settings.workspace-context` family rename `namedDeltas` from
  `new-workspace-runtime-context-null-settings-typeerror` — the name now lies, the TypeError is
  fixed — to `new-workspace-runtime-context-null-settings-degrades-to-absent`.
- All 153 move `baseline` and `recorderSha256`. The digest covers `pilot-scenarios.json`, so the
  baseline bump and the rename re-digest every file.

No sender recording moved: both fixes are reply-side, and no golden's `sender` or `payloads` field
differs. The README paragraph that claimed the settings TypeError was preserved is updated, and
now records that the `ui.get` leg of the same hook still is.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): degrade a null ui.get result the way the settings leg now does

One host answers both legs of useNewWorkspaceRuntimeContext, so fixing only
settings.get left the likelier failure in place: a null or absent ui.get result
still threw `reading 'ui'` out of the effect, skipping the provider commit.

Review follow-ups on the same files: reply() returns the literal uncast and the
stub client is FakeSession, dropping two assertions and their SAFETY disables;
the degradation cases now assert absolute state instead of comparing mounts.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-record the goldens the ui.get leg fix moves

Observational move, 1 golden:

- matrix-settings.workspace-context-ui.get-1: the `result-absent` and
  `result-null` partitions drop their `reading 'ui'` unhandled-rejection effect
  and their state commits `providers: ["github"]` instead of `[]`, because the
  effect no longer throws before the provider commit.

Header-only moves, 153 goldens: `baseline` to the fix commit and `recorderSha256`,
which covers `pilot-scenarios.json` and so re-digests on the delta rename.

The delta is renamed `new-workspace-runtime-context-null-settings-degrades-to-absent`
-> `new-workspace-runtime-context-null-results-degrade-to-absent` (9 goldens): it
now covers both reads, not just settings. README updated to match.

No sender recording moved: resolving the value pool across all 153 goldens shows
`sender` and `payloads` byte-identical everywhere.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): name the ui.get result shape so the changed cast carries a rationale

The inline union wrapped over four lines and tripped the changed-code casting gate
as a new assertion; a named alias keeps the cast on one line under a SAFETY note.
No behaviour change.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the golden baseline to the cast-rationale commit

Header-only: `baseline` on all 153 goldens. The re-record is inert — no golden
moves observationally and no field other than `baseline` changes.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): pin the ui.get trust blank, and correct two stale acceptance comments

The goldens cannot catch a regression to `if (uiResult?.result)`: the scenario's
success reply is `{"ui":{}}`, so every partition of
matrix-settings.workspace-context-ui.get-1 records the same `trust:{}` state. The
new case answers once with real trust and again with a null result on a fresh
client, which is the only shape where skipping the blank is observable —
trustedOrcaHooks gates the setup-hook approval prompt in
use-new-workspace-create-submit.ts, so a stale value would skip it.

settingsRead's comment still claimed workspace context, which this branch moved to
optionalSettingsRead; both comments now name their real callers.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the golden baseline to the trust-blank commit

The record fence rejected the previous pin ("Product sources or lockfile differ
from the pinned main baseline"), so the branch was no longer re-recordable.

Header-only: `baseline` and `recorderSha256` on all 153 goldens — the digest
covers pilot-scenarios.json, whose only edit is that baseline. The re-record is
inert: 0 goldens move observationally and no other field changes.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): name the right operation per refuse-after-data probe

Three of the five probes read through optionalSettingsRead, not settingsRead:
repo metadata and resume metadata already did, and workspace context does as of
this branch. The sentence now splits them and states why the split does not move
what the probes record.

Markdown is excluded from recorderSha256, so no re-record.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): pin the recorder's unhandled-rejection capture

This branch removed the last two goldens that recorded an unhandled-rejection
effect, so nothing exercised unhandled-recording.ts any more: gutting the emit to
`void captureError(error)` leaves all 153 goldens comparing clean. The unit test
drives a detached rejection through the window and asserts both the effect and
the listener restore. README says so where it describes the capture.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-digest the goldens for the new recorder test

Header-only: `recorderSha256` on all 153 goldens, which covers every non-markdown
file under rpc-recording/ and so moves for the added test file. The re-record is
inert: 0 goldens move observationally and no other field changes.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-14 11:48:28 -04:00
github-actions[bot] 93c3702463 Update README downloads badge 2026-09-14 12:38:13 +00:00
nireak b87a6c0f23 fix(pty): pace the EAGAIN write retry so a stalled reader can't saturate the daemon thread (#15319)
node-pty's CustomWriteStream retries an EAGAIN write with setImmediate, which
re-attempts within microseconds. A pty whose child has stopped draining stdin
keeps that branch EAGAIN-ing, so the retry becomes a busy-loop on the thread
that owns every pty on the runtime. Measured against this commit's parent on
macOS arm64: 121,316 EAGAIN/s at 101.6% CPU, versus 805/s at 4.1% with the
retry paced to 1ms.

The delay is 1ms rather than longer because the cost lands on readers that
drain in bursts -- what an agent does between event-loop ticks. Delivering 2MB
to a reader that drains 20ms out of every 100ms: 689ms unpaced, 907ms at 1ms,
1414ms at 5ms. 1ms keeps essentially all of the CPU saving without the
delivery regression.

clearImmediate -> clearTimeout in dispose() is required, not cosmetic: once the
handle is a Timeout, clearImmediate does not cancel it and a pending retry can
fire after dispose. The disposal guards that make that harmless (_fd = -1, queue
drop) are already on main; this mirrors them into src/unixTerminal.ts so the
TypeScript twin no longer drifts from the compiled lib.

Scope: this fixes the CPU saturation. It does not stop other terminals from
being serviced -- a second live pty kept answering echo round-trips throughout
the storm in every configuration tested (1 and 8 stalled writers, macOS and
Linux, 8 CPUs and 1), with throughput down ~20-50% rather than hung. The
"every terminal froze" symptom in #11178 has another cause and that issue
stays open.

Upstream chose setImmediate deliberately (microsoft/node-pty#831, #833) to fix
large-paste latency, and rejected polling POLLOUT because it reports writable
rather than flushed. That reasoning targets a per-write delay in an interactive
terminal; this delays only the EAGAIN branch in a long-lived daemon. Pastes to
a draining reader are unaffected (0-3 EAGAINs per MB in every arm).

Verified: patch applies to a pristine node-pty@1.1.0 tarball, the patched
src/unixTerminal.ts compiles byte-identical to the patched lib/unixTerminal.js,
patch_hash matches the file, and on Windows the changed code never executes
(WindowsTerminal, 0 EAGAINs on a 300KB conpty write).
2026-09-14 02:32:08 -07:00
Neil 8e26d516d8 perf(browser): dispatch coordinate pointer input in process instead of one subprocess per event (#20593) 2026-09-14 01:45:34 -07:00
Bjorn RunakerandNeil 01f8aa8d96 fix(grok): stop SessionStart orca-status hook hanging for 10s (#20090)
* fix(grok): stop SessionStart orca-status hook hanging for 10s

Grok writes one JSON hook payload and waits for the process to exit
without closing stdin. The POSIX hook used `cat`, which waits for EOF,
so SessionStart deadlocked until Grok's 10s timeout:

  session_start hook (global/orca-status) failed, ignored: timed out after 10000ms

Read one JSON object with raw_decode instead; return as soon as the
object is complete. Fall back to cat when Python is missing.

* fix(grok): decode the hook payload incrementally instead of per chunk

The JSON stdin reader strict-decoded the whole accumulated buffer after
every read and caught only json.JSONDecodeError. A multi-byte character
split across two os.read calls therefore raised UnicodeDecodeError, which
is a ValueError but not a JSONDecodeError, so the interpreter died — after
consuming stdin. The `python3 || python || cat` chain then handed the next
reader a truncated stream, and `cat` blocked on a pipe the caller never
closes, reinstating the exact 10s SessionStart timeout this reader exists
to avoid. Measured: a CJK+emoji payload written byte by byte hung until it
was killed; a 200KB payload split mid-character arrived as 4 bytes.

Hold the decoder across reads, treat any ValueError as "not complete yet",
and guard the whole program so a non-zero exit implies stdin was never
read — only then is the `||` fallback safe. Emit the object's own text
rather than re-serialising it, which was rewriting non-ASCII as \uXXXX.
Skip leading whitespace, which raw_decode does not. Separate the
first-byte wait (5s) from the idle wait (1.5s) so a writer that is merely
late is no longer dropped.

Also unset a HOME that does not exist before spawning the interpreter:
macOS resolves /usr/bin/python3 through an Xcode stub that re-runs its
whole tool lookup without a reachable cache, costing 6.6s per spawn and
overrunning Grok's budget on its own. This is what made the existing
large-payload and empty-PATH lifecycle cases fail.

The Python program now lives in a shell variable instead of being inlined
twice, which halves the generated script and keeps it readable.

---------

Co-authored-by: Neil <neil@stably.ai>
2026-09-14 01:45:00 -07:00
manuaudioandClaude Opus 5 170dbdb874 fix(ai-vault): ignore non-absolute env overrides for agent scan roots (#13118)
Six scan roots took a directory from an environment variable and used it
verbatim. A relative value is resolved by whichever Orca process reads it —
main sits at `/` when Finder-launched, the terminal daemon chdirs itself to
the user data dir, the AI Vault service inherits main's cwd — so one value
names a different directory in each, and walkSessionFiles walks it with no
depth cap, no entry cap and no time budget, about once a minute per the
session-list cache TTL.

The agent CLIs do accept a relative home (verified against real Grok 1.0.30:
`GROK_HOME=myhome grok du` creates `<grok-cwd>/myhome`), but they resolve it
against their own per-terminal cwd, which no Orca reader shares. Falling back
to the default home is therefore not a lost configuration — it replaces an
unbounded walk of an arbitrary tree with a bounded read of a known one, and
matches what readGrokHomeEnvelope, skill-provider normalizedRoot and
absoluteConfiguredDir already do with the same values.

Add resolveAbsoluteDirOverride and apply it to CODEX_HOME, COPILOT_HOME,
OPENCLAW_STATE_DIR, DEVIN_HOME, KIMI_CODE_HOME and GROK_HOME. It takes an
explicit platform so the Windows shapes are provable from a POSIX CI box:
`C:\...`, `C:/...` and UNC roots are kept, while the drive-relative `C:foo`
and bare `C:` fall back. Tilde expansion stays out of it — Grok creates a
literal `~` directory rather than expanding one — so absoluteConfiguredDir
keeps its own Pi/Prime-specific expansion and delegates the absolute check.

isAbsolute is syntactic only, so `/..` still collapses to `/`. That is fine
for read-only discovery; these roots never gate renderer-supplied paths.

Tests assert at the call sites, not just on the helper: the four
session-scanner-agent-sources roots are module-level consts evaluated at
import time, so they are exercised through AI_VAULT_AGENT_SOURCES with
vi.stubEnv plus vi.resetModules. Reverting any one of the six call sites
fails them (11-33 cases each).

Closes #13082

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 01:44:53 -07:00
2ed89b8781 fix(github): name an unfiltered empty project view instead of blaming a filter (#20588)
* fix(github): skip Projects search index for unfiltered views

Empty query still used items(query:\$q), which routes through GitHub's
Projects search index and can return totalCount 0 while the board is full
during index lag. Omit the query argument when the view filter is empty.

Fixes #12648.

* docs(github): drop the false stable-shape claim for empty project filters

Unfiltered item fetches omit items(query:) so boards skip search-index
lag. The View.filter field is still '' when GitHub returns null.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(github): name an unfiltered empty project view instead of blaming a filter

The search-index workaround in this branch was a no-op. Live introspection of
ProjectV2.items shows `query` is declared `String = ""`, so omitting the
argument and sending `$q = ""` coerce to the identical resolver input; GitHub
applies declared defaults for omitted args (verified against its own endpoint).
There is no non-search item field on ProjectV2 and ProjectV2View has no `items`
at all, so no request shape can dodge the index. Revert the branching query
construction and the module it added.

What the user actually reported in #12648 is the copy: a view with no filter
rendered "No items match this view's filter", which reads as data loss when a
freshly populated board momentarily comes back empty. Word the empty state from
the view's own filter — the filter message only when there is a filter, and an
honest "no items yet" plus a transience hint when there is not — and share the
one implementation between the table and roadmap surfaces.

Refs #12648.

---------

Co-authored-by: bbingz <zzb@gxsmjx.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-14 01:44:47 -07:00
Brennan Benson 539d4d1f32 fix(native-chat): resume structured chats cleanly after restart (#20509)
* fix(native-chat): retire provider ownership on restart

* fix(native-chat): stop showing a restart eviction as a provider death

Restarting Orca turned a resumable structured chat into a user-visible
`Provider exited: recorded pid absent on host`. Quit never released the durable
lease, so restart probed the recorded pid, adjudicated the session evicted, and
wrote a synthetic status row against a chat that was perfectly resumable.

The fix is the missing teardown phase plus the missing fence check: quit now
evicts every provider child this host owns — stopping it, settling its journal
and handing the lease back — and the release compare-and-swaps on the fence it
expected. Restart then finds a released lease and reopens the chat silently.

What the user sees is decided by the typed death evidence rather than the shape
of a settlement id: only an `exit-observed` death writes copy, and that copy now
carries its cause so an auth failure and an OOM kill do not read alike. The
reassuring wording stays. Historical synthetic rows are filtered out of the
render projection, which needs no schema change and leaves every real
provider-exit row alone.

Also:
- Bound the new eviction phase well below the quit deadline; a quit that dies
  mid-eviction leaves the lease unreleased, which is the original bug.
- Scope the interruption verdict to work that was mid-response. A provider that
  died while waiting on an approval interrupted nothing.
- Keep host bookkeeping in step with the adapter: the provider-child flag clears
  when the child is proven stopped, not seven steps later.
- Drop the router's duplicate shutdown gate and acquisition drain — both
  adapters already own theirs — and latch the router closed so a late acquire
  cannot fan a session back out to closed adapters.
- Attach the real cause to the settlement failure a quit reports, and remove a
  recovery-ticket field that was hardcoded at its only construction site.

* fix(native-chat): scope the legacy status filter to the copy it retires

The read-time filter hid every status row carrying a `restart-eviction:`
identity. That identity is still minted, so a genuine provider death settled
under it would have been dropped from every rendered page. Match the retired
`Provider exited` copy as well, so only the legacy rows are hidden.

Three smaller corrections alongside it:

- The settlement retry path now applies the same unfinished-work check the
  live exit path uses, so a provider that died waiting on an approval no
  longer gets told a response was in progress.
- Bound the exit reason before composing the outcome copy, so a stderr dump
  in the reason cannot push the "you can continue" sentence past the row's
  byte cap.
- Correct the teardown comment: tail rows are protected by eviction's own
  per-session ordering, and `closeAll` is a backstop for children eviction
  never took, including one whose eviction was refused.

* fix(native-chat): retire legacy status rows at the projection source

The read-time filter that hides the retired `Provider exited …` rows ran on the
way OUT of the page builder, after the paging math had already measured the
unfiltered timeline. A backward window landing entirely on those rows returned
an empty page that still reported `hasOlder: true` with a null `window.oldest`,
so the renderer's backfill loop re-asked from the same anchor forever. Its only
no-progress guard compares `window.oldest?.sequence` to the anchor, and
`undefined === n` never breaks. The live subscription opens behind that loop, so
the transcript never finished loading either.

Filter where items ENTER the page pipeline instead: the reduced snapshot gets
one renderable timeline, the forward path gets one renderable batch, and the
window bound, effective limit, `hasOlder`, `window.oldest` and `nextCursor` are
all computed over that single array. A window with nothing left behind it now
reports end-of-history.

Also restore the eviction retry contract. Clearing `hasProviderChild` as soon as
the adapter proves the child gone is honest, but it is a different fact from the
wind-down this host still owes. A retry after a step aborted between the two was
reading "no child here" and skipping both the dead-generation settlement and the
lease release the aborted attempt had promised to repeat. The obligation is now
tracked separately and cleared only by a release that actually landed.

And rename the filter to the copy it retires: it drops only rows carrying the
retired `Provider exited` text, not restart-eviction status rows in general.

* fix(native-chat): read the wind-down a close owes from the live child

An eviction recorded "nothing owed" whenever it ran over a session with no
provider child of its own, and the retry then read that record in preference
to the child in front of it. A session suspended to an agent terminal is
exactly that shape, and the trip back to native re-acquires into the SAME
session object rather than replacing it, so the next close skipped both the
dead-generation settlement and the lease release — leaving the record claiming
a live owner this host had just stopped, and a pending send unsettled.

The obligation is now derived the way the quit sweep already derived it, from
one shared predicate: a live child always owes a wind-down, and a remembered
`false` only carries the obligation forward, never cancels it.

Also drops a memoization in the history page that could never hit. Its key was
the snapshot's items array, which the reducer rebuilds on every `snapshot()`
call, so each backward page allocated a fresh key; the one reader that does
share a snapshot across pages reads forward and never calls it. The comment
claimed a multi-page read filtered once, which was not true of either path.

Tests: the handoff round trip that strands the lease, and the quit sweep
picking up an eviction whose close retry never came.

* chore(native-chat): scope three helpers to their file and pin the teardown order

retryUnexpectedExitSettlement, hasUnfinishedStructuredAgentSessionWork and
isRetiredProviderExitStatusItem each have no consumer outside the file that
defines them, so they no longer advertise an external contract.

The quit-path phase list documents its order as load-bearing, but nothing
asserted it. Pin the phase names so evict-owned-sessions cannot drift out of
its slot between drain-attaches and flush-event-sinks.

* fix(native-chat): stop the router reporting a stop it never observed

`closeAll` cleared the route table and set one boolean, after which that boolean
was the only surviving evidence about any session. Two call sites then spent it:
`releaseAcquisition` and the stop path each turned a route-lookup MISS into
reported success. Eviction reads a `true` from the stop path as proof the
provider child is gone and releases the durable lease on it, so a session the
router never routed could have its lease handed back on the strength of "I have
no record, but everything is closed."

Loss of contact is not evidence of process death. The fix keeps the evidence
instead of the inference: adapter shutdown only resolves once every child is
proven stopped, so `closeAll` now marks each routed session `stopped` rather
than forgetting it. A routed session still answers `true` from its own retained
proof; a session with no route answers `false`, which leaves it indexed for a
real retry. `releaseAcquisition` drops its short-circuit and asks the adapters,
which answer from their own session maps.

The acquire-side latch is unchanged: once closed, the router stays closed and
refuses new work.

Behaviour that changed: a post-`closeAll` stop for a session the router never
routed, or one the host already acknowledged as released, now reports unproven
instead of proven. That matches what the same call already answered before
`closeAll`, and no real flow reaches it — quit evicts every owned session before
`closeAll` runs, and eviction only asks the adapter for sessions whose provider
child this host acquired through the router.

* test(native-chat): ratchet the retired provider-exit copy out of production

The retirement filter hides a status row on two facts: a restart-eviction item
id and copy that opens with the retired prefix. The identity half is still
minted today, so the filter cannot tell a new producer's row from the legacy row
it exists to hide — any future writer of that copy would be dropped from every
transcript with no trace. Until now that safety property lived only in a doc
comment.

Scan the shipped tree for a string literal that OPENS with the retired prefix,
which is exactly what the filter's `startsWith` reads. Comments are stripped
first, so prose about the retirement is not a producer, and the filter's own
constant is exempt. Tests are excluded: writing the copy is how the filter is
exercised.

* revert(native-chat): drop the read-time retired provider-exit filter

Fix forward instead. The lifecycle change in this branch stops any new
`Provider exited: <reason>` row from being written; rows a previous build
already persisted stay in those transcripts and age out with them. A
permanent read-time filter for a cosmetic, shrinking set was not worth its
maintenance cost, and its paging seam was the only place a backward window
could land entirely on hidden rows.

Removes the filter module and its test, restores agent-session-history-page.ts
to its pre-branch form, and drops the tests that only existed to prove the
filter did not over-match or wedge the backfill loop.

The copy ratchet stays and now carries the whole guarantee: with no filter in
front of it, any production writer that resurrects the retired prefix reaches
the user's transcript directly.
2026-09-14 00:23:20 -07:00
Brennan Benson 4634d2c03b fix(native-chat): let the provider reopen a Claude turn it resumed itself (#20518)
* fix(native-chat): let the provider reopen a Claude turn it resumed itself

A Claude turn could only be opened by Orca's own send echo, while any
`result` frame closed it. The provider resumes work on its own — a
background task reports in and wakes the agent after `result` settled the
turn — and nothing Orca sent ever arrives to reopen one, so the session
projected `idle` for the rest of the work. The model's own output is the
evidence a turn is running, the way Codex's `turn/start` is, so it opens
one; whichever opened it, the next `result` settles it.

Subagent frames still open nothing: children outlive the turn that spawned
them, and their work is their parent turn's, never a turn of its own.

* fix(native-chat): bracket a resumed turn around the output that opened it

A resumed turn was published after the frame's own rows, so its first tool
call sat above the turn record. Every reader that scans back to the turn
record and stops — the active-tool reader behind the sidebar's tool line,
and the turn-window activity selector — looked straight past it, and the
row showed working with no tool until a second call landed.

The turn now opens before its frame is journaled. A send's turn keeps its
existing order: the user echo is that turn's anchor and is written first.

Prompt journaling moves to its own module, verbatim, to keep the translator
clear of the line cap.

* refactor(native-chat): declare the turn open at each content site

The resumed-turn rule was a predicate that re-derived whether a frame had
produced anything, duplicating work the frame handler had already done. The
content sites know: each one now calls an idempotent ensureTurnOpen before it
journals, and the guard against reopening a live turn lives in that one place.

Behaviour is unchanged; claude-turn-opening.ts is left owning only the send
echo, which is the one opener that anchors its turn to a user row.

* fix(native-chat): gate both turn edges on root-ness

The reopen path already refused nested output; the guard now reads before the
already-running check so both edges state root-ness first. The close path had
no nesting check at all, so a child's result would have ended the turn that
spawned it.

The two edges read parent_tool_use_id differently on purpose, and both fail
towards not over-claiming: opening needs proof of root-ness, so an absent field
opens nothing; closing needs proof of nesting, so an absent field still closes.

No real Claude stream has been observed carrying a nested result — the session
that prompted this work has none in any subagent stream — so the close-side
guard is symmetry, not a demonstrated fix.

* fix(native-chat): stop provider output reopening a turn nothing can close

Self-audit found two paths the reopen rule opened where no event could ever
settle the turn it created, leaving the row working for the life of the
session. Both now suppress reopening until an accepted send lifts it:

- a frame arriving after the session ended, when no event will settle anything
- a turn the provider failed, or the user stopped, where the next thing the
  provider says is not a resumption

Each has a one-lever ablation: removing the suppression read alone fails
exactly those two tests, and both fail as working-instead-of-idle.

* fix(native-chat): open a resumed turn from its first streamed delta

Streamed deltas short-circuit before the frame handler, so a resumed turn
whose first output is streamed text — the common case, since partial messages
are a pinned launch contract — kept reading idle while its partial text was
already journaled and visible. The streamed path now opens the turn too.

Also from the same audit:
- a nested result no longer swallows its own failure diagnostic; the turn gate
  now guards only settlement, and the provider-fallback row is written either way
- root-ness treats an absent parent_tool_use_id as root, so a build that omits
  the field cannot silently stop opening turns
- the suppression latch only ever sets on a failed result; a later clean result
  cannot lift it, and only an accepted send does

The opener moves into claude-turn-opening.ts so both entry points share one
root-then-suppression-then-idempotency order.

* fix(native-chat): preserve resumed-turn lifecycle semantics
2026-09-14 00:19:39 -07:00
Brennan Benson 0c812843ef feat(native-chat): show a drop target on the whole chat pane (#20561)
* feat(native-chat): show a drop target on the whole chat pane

Dropping a file into a chat only worked if you hit the input box, and nothing
on screen said so. A drag aimed at the transcript fell through to the terminal
behind the chat, which pastes the paths into the hidden TUI.

The chat pane shell is now the drop surface. While a drag carrying files is
over it, the pane dims behind a card naming what the drop will do; the
composer's existing attach logic — workspace, execution-host and SSH-owner
checks included — runs unchanged from the wider element.

Two routes reach the composer and they are widened differently:

- an in-app drag is claimed in the renderer, so the pane calls the composer's
  handlers through a claim the composer publishes to the surface around it.
- an OS drag is delivered by the preload drop route, which consumes the event
  at `document` before React sees it. The pane widens that route by publishing
  the composer's scope key as the nearest drop marker, and learns the drag
  ended from a document-level listener rather than a React drop.

The surface lives in the chat portal because that is the one place wrapping
both the bridge and structured panes, so neither pane root grows a second copy
of this wiring.

No composer mounted (a question card owns the input region) means no marker and
no overlay, so that drag stays the terminal's. A guarded composer still answers
for the drag — that refusal is what keeps it out of the terminal — but the pane
does not invite a drop it will refuse.

* Fix native chat drop ownership across session surfaces

* Keep native drop completion listeners ready before hover renders
2026-09-13 23:35:14 -07:00
Brennan BensonandMerge Sim 06c83e24c0 fix(workspaces): recover a real terminal when activation throws for a planless agent create (#20190)
* fix(workspaces): recover a real terminal when activation throws for a planless agent create

#20175's activation-failure recovery re-seeds the workspace surface, but passed
callerProvidesSurface: true whenever an agent was selected. Activation is the
caller that would have provided that surface, and it just threw. With the flag
set and no startup plan, worktree-initial-terminal-seeding takes its zero-tab
pre-seed branch: it queues setup/issue commands and returns null instead of
creating a terminal. The recovery then recovers nothing and the freshly created
workspace opens with no tabs.

Dropping the flag from the recovery call is the whole fix. The create-time call
keeps it: there, the agent pane really is the surface the caller provides.

The enclosing object literal becomes a ternary because oxlint's
unicorn/no-useless-spread rejects a lone conditional spread.

* test(workspaces): reuse activation recovery coverage

* test(workspaces): keep recovery assertion focused

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-13 23:29:51 -07:00
Brennan BensonandMerge Sim ad49aae600 fix(workspaces): re-seed a gate-reported empty workspace after an agent selection (#20182)
* fix(workspaces): re-seed a gate-reported empty workspace after an agent selection

#19940 routed two different questions through one predicate. An agent selection
justifies skipping the *pre-emptive* shell at create time, but it was also
suppressing the async activation gate's fail-closed re-seed. The gate returns
`empty` only after adoption, structured inventory and resume all produced
nothing, so it is positive evidence the agent surface never arrived (dead PTY,
unreadable census, null startup plan) — exactly when the re-seed is needed.

Suppressing it there left the workspace with zero tabs and no recovery:
ensureWebRuntimeWorktreeTerminalAfterWake returns early for a plain local
workspace, so nothing else seeds one.

`gatedEmptyOutcomeReseedSuppressed` now answers the re-seed question on its own,
honouring only an explicit `providesInitialSurface` caller promise — restoring
v1.4.200's behaviour for that path. The create-time skip keeps using the
agent-inclusive `activationProvidesInitialSurface` unchanged.

* refactor(workspaces): make gated reseed policy explicit

* fix(workspaces): harden gated empty reseeding

* chore(workspaces): split duplicate-folder routing fix

* refactor(workspaces): keep reseed helper outside caller census

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-13 23:29:13 -07:00
OrcaWinandOrca Worker 243f443155 fix(session-search): read oversized numeric file IDs on Windows (#20551)
Co-authored-by: Orca Worker <orca-worker@localhost>
2026-09-13 23:25:51 -07:00
Neil 55b3392018 fix(terminal): drop the agent gutter from copied selections (#19770) (#20545)
* fix(terminal): drop the agent gutter from copied selections (#19770)

xterm selections are screen cells, not logical text. Agent CLIs paint
their messages behind a fixed left gutter, so every copied line carried
that gutter into the clipboard and pasted replies came out indented.

Terminal clipboard writes now drop the run of spaces that *every*
selected line shares, so relative indentation (nested bullets, fenced
code, YAML) survives and only the gutter is lost. A selection that
starts mid-line, or that includes any column-0 line, has a shared run of
zero and is copied verbatim.

Applied at every terminal clipboard seam: the Cmd/Ctrl+C shortcut, the
pane context menu's Copy, right-click-to-copy, the app menu's Copy,
copy-on-select, the X11 primary selection, the dashboard popout's
preview terminal, and mobile's selection Copy button.

New "Trim Gutter on Copy" terminal setting (default on) restores the
old verbatim-cell behaviour.

* fix(terminal): honour the gutter-trim setting on mobile copy

Mobile stripped the gutter unconditionally, so turning "Trim Gutter on
Copy" off left one surface still rewriting the clipboard. Mobile now
mirrors the desktop preference through the existing settings.get RPC —
a host predating the setting sends no key, which reads as on, matching
the desktop default.

Also folds the single-use gutter helpers into their callers so the
shared module exposes one function.

* refactor(terminal): parse each selection line once in the gutter rule

Also locks the Windows subtlety with a test: a blank CRLF row is '\r',
which reads as a zero-indent content row and would cancel the gutter
unless the CR is split off first.

* fix(terminal): publish the gutter-trim setting to paired clients

settings.get is an explicit allowlist projection, not the whole settings
object, so terminalCopyTrimsGutter never reached mobile: the client read
the key as absent, which means "older host", which means on. Mobile
therefore always trimmed and the desktop opt-out was inert.

Adds the field to the projection and a test that fails if it is ever
dropped again — absence is indistinguishable on the client from an old
host, so a silent regression here has no other signal.

* chore: drop unrelated formatter drift from this branch

A repo-wide `pnpm format` swept a quote-style change in pnpm-workspace.yaml
and a blank line in source-tree-walk.test.ts into this branch; neither is
related to the gutter fix.

* fix(terminal): trim the gutter on native copy events too

xterm binds its own DOM `copy` listener that writes raw screen cells
(CoreBrowserTerminal `_initGlobal`). Orca's own chords never reach it —
they preventDefault in keydown — but Ctrl+Insert is a Chromium copy
accelerator on Windows/Linux and is not in `terminal.copySelection`'s
bindings, so it still copied the gutter. Orca binds Shift+Insert for
paste on those platforms, which makes the asymmetry worse.

A capture-phase listener on the xterm element now writes the trimmed
text, closing the class rather than the one chord: any native copy event
— assistive tech, execCommand — lands on the same path. Installed for
both terminal panes and the dashboard popout's preview terminal.
2026-09-13 23:01:47 -07:00
Neil 3763103084 fix(hooks): report a timed-out hook as unverifiable and terminate its process tree (#20559)
## In plain terms

Orca lets a project define scripts that run at certain moments — one when a workspace is set up,
one just before it is deleted. Those scripts get a time limit. When the limit ran out, Orca asked
the script to stop and then believed whatever the script said on its way out — so a script written
to shut down politely could be cut off halfway through its work and still report that it had
finished. Anything relying on that answer was relying on a guess.

Now the verdict comes from the clock, not from the script: if it ran out of time, that is what is
reported, whatever exit code it managed on the way out. Orca also stops the script's *children*
rather than just the script, so a background process it started can no longer outlive it.

Split out of #20153 so the gate that consumes this answer is reviewed separately. `Refs #19334`
rather than `Fixes`, because it does not close the issue on its own.

## The bug

`exec({ timeout })` sends SIGTERM and then reports what the child did. A hook that traps SIGTERM
and exits 0 therefore comes back with a **null error** — success — despite having been cut off.

```js
exec("trap 'exit 0' TERM; sleep 5", { timeout: 200 }, (err) => …)  // err === null
```

That is not an `exited` vs `unverifiable` nicety: it is a failed hook reported as a passing one.
Realistic triggers are ordinary — a Node wrapper with a graceful `process.on('SIGTERM')`, an rsync
wrapper that cleans up on signal.

## What changed

**`runHook` owns the deadline.** The verdict comes from running out of time rather than from the
corpse's exit code, and it is settled *at* the deadline rather than whenever the child gets around
to dying — a hook that traps the signal and keeps running must not hold its caller open.

**A timeout withholds the exit code.** So does a spawn failure, where `exec` reports a *string*
code (`ENOENT`); the `typeof code === 'number'` guard is what keeps a hook that never ran out of the
"exited" verdict. Callers that distinguish "exited N" from "outcome never observed" can now trust
that distinction:

| failure mode | `error.code` | signal | verdict |
| --- | --- | --- | --- |
| non-zero exit | `23` | — | `exited 23` |
| command not found | `127` | — | `exited 127` |
| killed | `null` | SIGKILL | outcome not observed |
| deadline expired | *(the deadline, not the exit)* | SIGTERM→SIGKILL | outcome not observed |
| deadline expired, hook traps SIGTERM and exits 0 | `0` | — | outcome not observed |
| spawn failure | `"ENOENT"` *(string)* | — | outcome not observed |

**Termination reaches the process group.** The script is a shell and the work is its children, so
signalling only the shell leaves a `sleep` or an `rsync` alive holding the pipes open. SIGTERM
first, then SIGKILL after a grace.

**One `classifyHookProcessResult`** now serves the native and WSL branches, which had been mapping a
finished process to a hook verdict by hand, identically. That duplication predates this change.

## Terminating the tree, and a test that could not fail

The escalation went wrong once in review, in a way worth recording.

A first attempt skipped the SIGKILL when the *direct child* had already exited — a dead child needs
no signal. That is correct about the child and wrong about the group: a hook that backgrounds a
server typically loses its shell leader to the first SIGTERM while the server keeps running, so the
skip fired in exactly the case the escalation exists for. The escalation now probes the **group**
with signal 0: `ESRCH` means nothing is left to kill, anything else gets the signal.

**The residual trade-off, stated rather than implied.** Signalling by negative pid names whatever
group owns that pid *now*. Once the leader is reaped its pid can be recycled, and a probe cannot
distinguish a surviving descendant from a stranger that inherited the number. Killing a runaway hook
is both the likelier event and the one the deadline promises, so the group is signalled whenever it
answers; the remaining window is pid wraparound inside the grace.

**A test that cannot fail is worse than no test.** The first regression test drove `runHook` with
`process.kill` intercepted — and passed against *both* the broken and the fixed version, because
with signals intercepted nothing dies, so the child never reached the exited state the bad guard
keyed on. It was false assurance, not coverage. `terminateHookTree` is therefore exported and the
regression pinned directly against it: it fails on the old version
(`expected [] to deeply equal [[-4242, 'SIGKILL']]`) and passes on this one.

## Behaviour change for `setup` hooks

Both hook kinds share `runHook`, so this is not confined to archive hooks. **A setup hook that
backgrounds a long-running server now has that server SIGTERM'd — then SIGKILL'd — with the rest of
its process group when the deadline expires, where previously it was orphaned and survived.**
Arguably the better behaviour, since an orphaned server is a leak, but it is a real change and
should be a decision rather than a discovery.

## Evidence

Against real shells and real signals, because this bug is invisible to a mock
(`hook-archive-timeout-observation.test.ts`, through `runHook` itself rather than an extracted
helper):

```
✓ fails a hook that traps SIGTERM and exits zero, despite its zero exit
✓ settles at the deadline even when the hook refuses to die
✓ passes a hook that finishes inside its deadline
✓ reports an observed non-zero exit as the exit it is
```

Plus `hooks-archive-exit-observation.test.ts` for the wiring — including the string-`ENOENT` case —
and `hook-archive-termination-safety.test.ts` for the escalation branching.

## Checks

`pnpm tc` · `oxlint src` · `oxfmt --check` · 113 tests across `src/main/hooks*`. The classification
table above is measured against real `exec`, not reasoned.
2026-09-13 22:36:01 -07:00
Brennan Benson 3cd60e76e9 feat(agent-status): run-identity types for keying rows by agent instead of pane (#20531)
* feat(agent-status): add run identity types

* fix(agent-status): harden run identity codecs
2026-09-13 22:06:26 -07:00
JinjingandJinwoo-H d2d32691ef perf(persistence): skip redundant whole-state flushes on terminal reattach (#20137)
* perf(persistence): add pty-binding fast lane to skip redundant flushes

Terminal pane reattachment currently clones the session and serializes the
entire 9.2 MB app state even when the binding is already in place and durable.
Add an early-return fast path that skips this work when all nine predicates
hold: no split, binding matches in-memory and on-disk, incarnation matches,
no tombstone, and generation counter proves durability.

Includes one-line fix in `writeToDiskSync` to record hash-matched sync flushes
as durable, so the fast path doesn't stay parked behind a stale generation.

Adds `persistence.pty-binding` observability spans (local NDJSON, unsampled for
mutations, budgeted for fast-lane hits) to measure eligibility rates before
and after. Includes ratchet test to ensure every binding writer bumps the
generation. Diagnostic tools and full investigation notes from September 7,
2026 capture that identified the 59–100 ms no-op binds and measured a real
terminal keystroke queued 117 ms behind one such call.

* perf(persistence): add pty-binding fast lane to skip redundant flushes

Rapid rebinds of already-durable PTY bindings (e.g., remounting panes)
were unnecessarily expensive because they cloned and flushed the entire
document state every time. Detect when a binding hasn't changed since the
last durable write and skip to return immediately, eliminating main-thread
cost on that path.

* perf(persistence): record binding.origin on the pty-binding span

Fresh spawns always flush, so a fast-lane rate over all calls is diluted
by however many terminals the user opened. Each caller knows whether it
is a spawn, a reattach, a split, or a relay reattach; pass that through
as metadata and record it so the reattach hit rate can be read from the
trace file. Never branched on.

* fix(persistence): keep the tab row on its first pane when a sibling pane binds

A tab row names one PTY, but a split tab holds several panes. The
renderer keeps the row on the first pane and refuses to let later
split-pane spawns steal it, since a remount reattaches the tab to
whatever the row says. Main overwrote it with whichever pane was binding,
and the renderer's next publish put it back, so every sibling reattach
was a state change and could never take the fast lane. On the real
profile that is 38% of panes.

Rewrite the row only when it names nothing useful: null, the PTY this
leaf is replacing, or a PTY no leaf holds. The fast-lane predicate
compares against the same rule.

* perf(persistence): record durable pty-binding flushes per pane

The global write generation is held back by any unrelated dirty
state, causing bindings unchanged for minutes to appear unpersisted
despite being on disk. Track per-pane durability to skip redundant
flushes.

* docs(persistence): describe the per-pane durability record

The durability section still described the global generation check as the
whole story and claimed there was no binding durability cache. Record the
measurement that motivated the per-pane record, and why retiring one needs
no cooperation from other binding writers.

* docs(perf): consolidate every measured Orca performance issue into one register

Folds the findings from all related debug sessions into the live lag
investigation: the persistence/main-thread work (P1-P11), host contention
(H1-H5), git and subprocess load on main (G1-G8), renderer and terminal
rendering (R1-R8), the terminal daemon session leak from the deleted
debug-orca-perf-issue worktree (D1-D9), and the Cmd-J palette review (C1-C6).

Keeps the measurement behind each claim, records what is fixed versus open,
and restates what the 117 ms keystroke delay still does not explain.

* fix: address performance review findings

* fix: satisfy diagnostic probe lint

* chore: keep investigation artifacts out of performance PR

* fix: run lag probe regression tests with Vitest

* perf(persistence): replace pane receipts with global durability check

* refactor(persistence): remove redundant binding review machinery

* test(persistence): satisfy current assertion-free quality gate

---------

Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
2026-09-14 01:04:43 -04:00
Jinwoo Hong 7d98c8e2f3 refactor(mobile): send the source-control domain through typed RpcOperations (#20544)
* test(mobile): record main's source-control RPC behaviour before migrating it

45 scenarios over 11 source-control senders, recorded from main so the step-4
migration has a frozen answer to compare against. Adapters mount the real
exported senders as plain functions, so no React host or device is needed.

The 73 existing goldens change header-only (`baseline`, `recorderSha256`): any
new scenario re-digests the recorder, and the pinned baseline had drifted from
main in `src/shared` so recording required bumping it. Content is byte-identical
on all 73 — verified field-by-field against HEAD.

Scenarios deliberately pin the empty-message cases (`sc-*-refused-empty-message`,
`sc-*-rejected-empty-message`), because a refusal with no message falls back to
the screen's copy while a transport error with no message does not, and the two
paths are easy to collapse when a call site moves behind an acceptance policy.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): send the source-control domain through typed RpcOperations

13 of the domain's 14 files now send through a declared operation instead of the
raw request port: 44 references to 0. The holdout is use-mobile-git-requests.ts,
whose single reference is a `(method: string, params)` dispatcher that five other
hooks feed `{ method, params }` action steps at runtime; typing it is a step model
change, not a call-site move, so its line stays at 1.

Fifteen operations over fourteen methods. Two of them read git.status, and that is
deliberate: the Changes screen publishes the host payload verbatim while
hosted-review preparation reads the normalized projection, which returns null when
`entries` is not an array and drops entries missing a path. Sharing the projecting
reader would change what the Changes list renders, so both are named.

Four loads still read the refusal envelope before interpreting, through
readMobileGitRefusal: two degrade to a capability-missing screen, one retries a
selector that is not visible yet, and one falls back from files.openDiff to
files.open. `isMobileGitUnavailable` consults the code *and* the message and no
acceptance policy carries either through, so the alternative was parsing a code out
of a message. No new acceptance policy was added.

Every migrated site keeps two error paths where it had two: a refusal with no
message falls back to the screen's copy, a transport rejection surfaces its own
message verbatim and keeps its delivery-unknown mark. Collapsing them into one catch
is what would have turned an unknown mutation into a failed one.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-digest the goldens after a lint fix in the new adapter

recorderSha256 only, all 118 files; every recorded observation is byte-identical.
Re-recorded from c57de48fd0 in a separate worktree so the goldens stay attributable
to pre-migration product source — recording from this branch would have made the
parity claim circular.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): drive the reply matrix over every scripted reply, and fail closed

The matrix picked its driven request from a hardcoded prefix list and `continue`d
past any family the list did not name. That was 10 of 23 families — every one the
source-control migration added — with no red test to say so, which is why that
migration's mutation evidence came down to single hand-written scenarios.

`replyMatrixSites` now takes every completion step in a family's base scenario:
61 sites instead of 13, one golden per site, no judgement about which request is
the "real" one and nothing to edit when a domain is added. A family that scripts
no reply throws, a repeated request name throws, and a census test asserts every
family in the manifest has a matrix. A variant's downstream replies are marked
`optional` and answered only if the request is outstanding, so a diverged reply
that ends the chain records the truth instead of failing on an unsent request.

The `normal` partition replays the first fulfilled reply the family records for
that request, rather than a payload the test file invented per family. Absent and
null do not count — each is already a partition — so four sites with no other
recorded success are inventoried in REPLY_MATRIX_NORMAL_RESULT_INVENTORY with a
reason each, and an entry whose family later records a success fails.

Two partitions added: a refusal and a transport rejection with no message. That
is the axis that separates a refusal falling back to the screen's copy from a
transport drop surfacing its empty message verbatim; without it the two paths
produce the same text and collapsing them is invisible. Every source-control
family carried a hand-written `*-empty-message` scenario for exactly that.

13 hand-written scenarios the matrix now covers are deleted: 8 `*-empty-message`
cases plus sc-history-rejected, sc-commit-message-null-result, sc-eligibility-
refused, sc-create-stops-on-push-refusal and sc-base-ref-rejected. Kept, with
reasons, are the ones the matrix cannot reach: a different action or action args
(sc-review-commit-*, sc-prefill-*, sc-create-{refused,rejected}-empty-message,
sc-prerequisite-{publish,force-with-lease,skipped}), a payload shape rather than
an envelope shape (sc-review-status-entries-not-array, sc-create-existing-review),
and multi-request combinations (sc-base-ref-{unavailable,repo-fallback},
sc-reveal-timeout).

Goldens: 118 -> 153. All 92 survivors changed by their `recorderSha256` line only;
no recorded observation moved. Re-recorded from the pinned baseline in a separate
tree so the goldens stay attributable to pre-migration product source.

`matrix-hostedreview.create-intent-git.commit-1` fails on this branch, and it is
a true positive: `hostReplyErrorTextOrFallback` stringifies a non-string in-band
host error where main returned `result?.error || fallback` and passed the object
through. Left failing — the fix is a product change, documented in the README.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): keep main's in-band commit error pass-through

The expanded reply matrix caught a real divergence the nine original partitions
missed. Main returned `result?.error || 'Commit failed'`, passing a truthy
non-string straight through under a `string` annotation; the migrated helper
stringified it to "[object Object]".

Stringifying is arguably better — downstream does `result.error.replace(...)`,
which throws on an object and merely looks ugly on a string. But this migration's
contract is that no behaviour changes, and shipping an unannounced improvement
inside a refactor is exactly what the parity evidence exists to prevent. Restores
the pass-through; the latent throw is its own ticket.

No host sends this today (`git.commit` is typed `{success, error?: string}`), but
nothing validates it and mixed client/host versions are normal.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-digest the goldens for the merged recorder

Main inverted two guards in family-recordings.test.ts and pilot-recordings.test.ts.
No behaviour change, but both files are inside recorderSha256, so all 153 goldens
failed the header check after the merge.

Re-recorded from 16d1ab81d3 in a separate worktree carrying main's product source
and this branch's merged recorder, so the goldens still capture main's behaviour
rather than the migration's. `baseline` moves from 7ce8e18d07 to 16d1ab81d3 because
main touched src/shared/skills*.ts, which the record guard compares; that change
moved no recorded observation. Every field except `baseline` and `recorderSha256`
is byte-identical across all 153 files.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): intern each observation entry instead of the whole field

Making the reply matrix fail closed took the goldens from 118 files / 1.53 MB to
153 / 5.35 MB, because a per-site golden replays the chain across 11 reply
partitions and every checkpoint's sender, payloads, settlements and effects
re-state the whole history that came before them. Format version 2 pooled those
fields whole, so the shared prefix was stored once per checkpoint, and once per
partition again.

Version 3 pools each entry of a list or map field instead. `golden-value-pool.ts`
declares the container per field rather than sniffing it from the value, so a
projection that changes one fails loudly instead of silently switching encodings.
153 files / 5.35 MB becomes 153 / 2.78 MB; the family that drove this,
hostedReview.create-intent, 2.0 MB over 12 sites becomes 792 KB.

This is a re-encoding, not a re-observation. Every one of the 153 goldens resolves
to the recording its version 2 file resolved to, checked field by field, and every
header field except recorderSha256 and goldenFormatVersion is byte-identical. The
three mutations this branch's coverage rests on fail exactly as before: the
gitStatusProjectionRead acceptance policy 16 (13 matrix, 3 hand-written),
interpret inside the request chain 5 (all matrix), and the rewrapped transport
rejection 4 (all matrix).

It also makes diffs smaller, which is the opposite of what version 2's note
predicted when it rejected this. Adding a timeoutMs to the first git.status of the
create-intent chain touches the same 16 goldens either way, but version 2 moves
17,100 lines / 1.03 MB and version 3 moves 3,764 / 0.20 MB, because a changed
entry no longer rewrites every field value containing it.

`readGolden` now also refuses a pool entry that does not hash to its own key, and
one no checkpoint reads. Content addressing is what makes an entry shared between
checkpoints safe to share; an unreferenced entry would be content in the file that
nothing compares.

Recorded from 16d1ab81d3 with this branch's recorder laid over it, per the
README's flow. The record fence is unchanged.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): make the matrix census and inventory checks able to fail

Three review findings, all in the recorder, none in product code.

The family census pushed every family unconditionally, so it could never differ
from the manifest keys; it now records a family only when a site generated a
test, which is independent of replyMatrixSites throwing on an empty list.
REPLY_MATRIX_NORMAL_RESULT_INVENTORY was only consulted for a live site, so a
stale entry retired silently; a new assertion fails on any entry that names no
live (family, request). Both verified by mutation: an empty site list and a
renamed inventory request each fail the suite. The value pool resolves hashes
with Object.hasOwn so a malformed golden cannot read an inherited key.

Re-recorded from 16d1ab81d3 with this recorder laid over main's product source,
per the README. All 153 goldens move on recorderSha256 only.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-14 01:02:28 -04:00
Brennan Benson 33149fcde5 fix(claude): install SessionEnd for capable versions (#20530) 2026-09-13 21:59:57 -07:00
Brennan Benson c287a5d9b7 feat(native-chat): add provider-aware Fast mode (#20506)
* feat(native-chat): add provider-aware fast mode

* chore: drop unrelated formatter churn from the merge

pnpm format reflowed pnpm-workspace.yaml quoting and a source-scan test
that this PR does not otherwise touch.

* fix(native-chat): review fixes for provider-aware fast mode

Review pass over the Fast mode work.

Claude reads its model catalog once per option write. The admit check, the
effort guard and the Fast guard each took their own `list_models`, so a model
write with Fast on paid two round trips for one list and let two guards answer
from two different catalogs. The guards are now pure over a single read.

Claude no longer refuses a Fast enable when the catalog identified nothing at
all. An empty list is not evidence against a model -- the same rule the model
admit-check already applies -- so a CLI that cannot answer would otherwise have
Fast refused on every model. A catalog that did list the model and stayed silent
about Fast is still not positive evidence and keeps refusing.

Codex refuses a direct `serviceTier` write instead of accepting one the next
turn discards. The turn derives the tier from `fastMode`; the key still restores
so a session persisted before Fast existed migrates.

Both option surfaces return a cached snapshot again. `SessionOptionsSurface` is
read through `useSyncExternalStore`, whose contract is a stable snapshot, and
rebuilding it per call breaks that for any consumer wired that way.

Also records two decisions that were emergent rather than stated: routing
Standard when Fast is on but no tier is named yet, and what a readback
disagreement does and does not prove.

Quality gate: merges the duplicate imports static analysis flagged, adds SAFETY
rationales for two pre-existing casts the changed-code gate now sees, and drops
a new assertion in favour of a checked narrowing.

* fix(native-chat): read Claude Fast state from the session frame

A fresh Claude session reports `fastModeState` while the settings readback still
has no `fastMode` boolean, so the two are not redundant -- the frame answers at a
moment the boolean has none. The picker fell back to "value unknown" and asked
the user to disambiguate what the provider had already reported, and the state it
reported had no reader at all.

Falls back to the frame only when neither a pick nor the settings readback
answers. `cooldown` throttles routing rather than clearing the pick, so it reads
as on; reading it as off would flip a control nobody touched.

Display only. The launch seed is untouched: an unset Fast preference still seeds
nothing, which its own guard continues to pin.

* perf(native-chat): skip the model catalog read when turning Fast off

Turning Fast off needs no support evidence, so the read only cost a
round trip — and restore replays a stored `false` on every acquire.

Also narrows the alias-matcher comment: the effort and admit guards
match on alias and resolved id only, so calling it the sole matcher
overstated it.

* fix(native-chat): clear a Claude Fast block once the child stops reporting it

The child omits fast_mode_disabled_reason entirely when nothing blocks Fast
and never sends a null, so requiring the key back latched the first reason
for the session's life: switching to a model that disallows Fast and back
retired the control for good, leaving a session running Fast with no way to
turn it off. A frame that reports state without a reason is the all-clear.

* test(native-chat): cover the mobile structured option hook

useMobileStructuredAgentOptions gained generation fencing, a pending-write
guard and a post-write options refresh with no test file. Pins the concurrency
contract and the fast mode round trip:

- a superseded options read is dropped instead of overwriting newer state
- an overlapping write is refused and the pending guard is released after
- an accepted same-fence write reads options back and applies the result,
  and a different-fence write does not
- a boolean fastMode pick reaches the wire encoded and is remembered decoded
- no Fast row when session support, catalog support or the model capability
  is missing

Each behaviour was ablated against the production logic to confirm it fails
without it. No production code changed.

* feat(native-chat): render a boolean session option as one toggle

On and Off were two radio rows under a header repeating the option name,
so a binary choice cost three lines and two clicks to read. It is now a
single switch row that owns its label, on desktop and mobile.

An unknown value keeps its caption: a switch cannot say "unset".

* fix(native-chat): resolve a boolean option's display value at the producer

A boolean session option reached the UI in three states while its control had
only two, so the renderer apologised for the gap with a "Current value unknown"
caption beside a switch that had already collapsed to off. For `thinking`, whose
catalog default is on, that caption sat next to a switch asserting the opposite
of what every composed dispatch assumes.

One expression fed both the displayed value and the option's provenance. Split
them: the boolean descriptor now always carries a value, resolved to the same
`values[id] ?? defaultValue` that buildNativeChatSessionOptionCommand already
composes, while `valueSource` is untouched and still records whether anything
confirmed it. `kind.currentValue` is required on the boolean arm so the third
state cannot come back.

The launch path is unaffected: resolveAgentSessionOptionLaunch and
buildNativeChatSessionOptionCommand build the composed `--model` argument from
the caller's picks and the catalog, never from a descriptor.

Both surfaces mark an unconfirmed value instead of captioning it, and the two
reasons stay distinct — `default` says the catalog value is what a launch will
send, `unreported` says nothing has told us anything. Only `unreported` is
reachable in the structured lane, where the agent may be routing a tier we have
never been told about, so the two never share a label.

* fix(native-chat): let assistive tech read the option value marker

The marker was aria-hidden next to an explicit aria-label, so the label
already won the accessible name and hiding it only cost screen reader
users the default-vs-unreported distinction that sighted users get. It is
now referenced by aria-describedby, which keeps the name Fast mode.

Mobile's summary row said "Not set" for a boolean while the sheet behind
it showed the switch on, so the two screens disagreed. A boolean always
has a value; the summary states it and the sheet's marker qualifies it.

* chore(i18n): drop the On/Off option strings the switch row retired

Replacing the On/Off radio pair removed the only call sites for these two
keys. i18next cannot rebuild a key with no call-site default, so leaving
them in the catalogs forced them into the boot bundle as dead weight.
Removing them shrinks it by two entries instead.
2026-09-13 21:58:32 -07:00
Brennan Benson 2ce252f471 fix(grok): announce a completion once, when Grok is actually finished (#20523)
* fix(grok): announce a completion once, when Grok is actually finished

Orca pinged on every Grok turn-end. Grok runs turns the user never asked for:
when a background task finishes it wakes itself, does a little work, and ends
another turn. One request produced several pings.

Grok already reports, on every turn-end, whether it still has work outstanding.
Read that instead of trying to classify which turns are "real":

  backgroundTasks absent          -> silent, this is the session-end tail
  StopFailure / StopCancelled     -> announce, a failure is never hidden
  stopHookActive                  -> silent, a Stop hook is keeping it working
  a shell task or subagent running -> silent, the work is not done
  otherwise                       -> announce

Nothing here knows what an auto-wake turn is. A turn that ends with work
outstanding stays quiet; the later turn where that work is finally done is the
one that announces. That is also why this survives the case where Grok completes
a user's goal inside one of those turns — prefix-based suppression would have
silenced it.

Monitors and scheduled entries are deliberately not counted as outstanding work.
They can run indefinitely, so counting them would suppress a user's completion
permanently, and a lost ping is worse than an extra one.

Also registers StopCancelled, which Grok fires instead of Stop on a user
interrupt, a declined permission, --max-turns, or a no-progress bail-out. Orca
never subscribed to it, so those turns were reported as successes.

Also removes a stale notification matcher that searched for prose the shipping
binary never sends; the typed notification kind is matched instead, and neither
idle_prompt nor task_complete is treated as a completion.

Needs-input behaviour (permission prompts and ask_user_question waits) is
unchanged and stays ungated by background work.

* fix(grok): never hide a failed or cancelled turn behind the background-work gate

The announce predicate checked field-absence before terminal outcome. Grok's
StopFailure and StopCancelled payloads carry no background inventory at all, so
the absent-field branch — added so the session-end tail stays silent — fired
first and silenced every failure and every cancellation.

That inverted the rule it was meant to serve. Before this series a cancelled turn
at least surfaced as a (wrong) success; gated this way it surfaced as nothing.

Terminal outcome is now checked first, so a failure or cancellation announces
regardless of what other fields the payload happens to carry.

The existing tests passed straight through the bug because they built failure
payloads with a backgroundTasks field Grok never sends for those events. They now
model the real payload shapes, verified against the provider's payload
definitions and the captured envelopes.

* fix(grok): settle completion from provider lifecycle state

* fix(grok): fence stale turn ends without prompt ids
2026-09-13 21:09:01 -07:00
dngur6344andNeil 9cf0a6c37f perf(remote): avoid repeated capability probes during file imports (#14555)
* perf: avoid repeated remote import capability probes

* test: cover cold remote import compatibility probe

* fix(remote): fence imports across runtime reconnects

* fix(remote): bind import proof to connection

* fix(remote): fence import routing by runtime identity

* test(remote): remove unsafe import fixture assertions

- type remote RPC mocks at declaration so call arguments stay checked
- narrow upload params before reusing generated temp paths

---------

Co-authored-by: Neil <neil@stably.ai>
2026-09-13 20:51:15 -07:00
96b450fae8 fix(ssh): bound relay incumbent lsof probe (#18304)
* fix ssh relay incumbent probe timeout

* test ssh relay unconfirmed probe termination

* fix(ssh): preserve connect evidence while bounding lsof

* fix: preserve uncertainty when relay holder enumeration fails

* fix(ssh): supervise lsof helpers and preserve partial holder evidence

* test(ssh): prevent GC racing unconfirmed probe cleanup

* refactor(process): keep POSIX lsof supervision in process owner

* fix(ssh): confirm census cleanup and handle probe startup signals

* fix(ssh): keep lsof holder evidence usable on hosts with unstat-able mounts

lsof warns to stderr about mounts it cannot stat, and any stderr byte forced
the holder enumeration to 'unavailable' — making the 'exited' verdict, and so
husk reaping, unreachable on those hosts. Pass -w to suppress the warnings.

* Revert "fix(ssh): keep lsof holder evidence usable on hosts with unstat-able mounts"

This reverts commit 10d7b0db37.

Passing -w is correct in isolation, but it activates a previously dormant
path: in the SSH docker e2e lsof warns about the container filesystem, so
every probe used to degrade to 'unavailable'. Suppressing the warning lets
holder enumeration succeed and the takeover/reaping path run for the first
time there, and 'e2e / ssh docker watcher isolation' then hung to the job
timeout (52m, vs 28m passing on the parent commit).

The underlying gap is real and still open: on hosts where lsof always warns,
the 'exited' verdict and husk reaping stay unreachable. Fixing it needs the
reaping path understood in that environment, not just the flag.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Neil <neil@stably.ai>
2026-09-13 20:51:12 -07:00
Jinwoo Hong 3ab2a1b91c refactor(orchestration): derive delivery eligibility from messages (#19837)
* fix(orchestration): retire read deliveries and clarify mailbox recovery

* fix(orchestration): simplify delivery recovery and update nudge contracts

* test: align orchestration check help expectation

* refactor(orchestration): derive delivery eligibility from messages

* fix(orchestration): validate live consumers and simplify batch revocation

* refactor(orchestration): keep deliveries.status and derive eligibility without a column drop

The outstanding_deliveries view now reads status = 'outstanding' plus unread
membership, so v41 only drops uniqueness from idx_deliveries_one_outstanding
and adds the view and trigger. Older binaries can still open the database.
Removes the column-drop migration, the v40 test fixture and hasColumn guards,
the fenced skew probe, and the unrelated nudge-text change.

* docs(orchestration): drop delivery storage reference

The compatibility caveat it existed to explain no longer applies; the view
and index comments carry the remaining rationale.

* docs: revert unrelated formatter churn

* test(orchestration): verify historical database downgrade round trip
2026-09-13 23:24:04 -04:00