A first-hand Claude exit is not published where it is observed. `handleExit`
re-enters the close ladder and persists the transcript cursor before it emits
`ended`, and only that emission reaches the runtime's recovery chain. So the
runtime's `waitForRecovery` — whose whole job is to drain an in-flight recovery
before teardown stops children — returns immediately for an exit that is still
climbing the ladder, and nothing outside the adapter can tell an observed exit
from a published one.
The integration test for fenced host reconciliation had no handle on that
barrier, so it bounded-polled the lease for 100ms instead. Measured under 16x
local concurrency, publication alone takes 77-204ms: 19/24 runs failed.
Retain the ladder-then-settle tail on the exit record and expose
`drainObservedExits`, fold it into `waitForRecovery`, and export the barrier so
a caller that needs the settled lease can await it. Codex publishes inside its
own exit callback and needs nothing. The test now awaits the barrier: 0/24
under the same load, and it fails on an idle machine without the drain.
* fix: stop a handoff flow from outliving the host that owns its session
A structured handoff runs on the session's serialized chain and nothing in
production awaited it. When the client-side deadline for the switch expired
first, teardown dropped the session map out from under a live flow, and the
flow's own failure notification then threw `agent_session_ownership_unknown`
out of a status publish — an unhandled rejection, plus journal rows written
into a directory that was already being removed.
Three fixes, each with a regression test that fails without it:
- The status publish is a notification, not a mutation: it now reads the fence
without requiring an attached session, so an evicted or torn-down session
makes it a no-op instead of a throw.
- `track` used `.finally`, which forwards a rejection onto a promise nobody
awaits. `drain` settles flows through `allSettled`, so the bookkeeping chain
is now settle-only and cannot resurface one.
- Host teardown drains in-flight handoffs before dropping the session map.
`drain` existed for exactly this and was never wired up.
The integration test's `vi.waitFor` is dropped rather than widened: the request
enqueues the flow on the session's serialized chain before it returns, so the
status read is already ordered behind it. The poll only added a wall-clock
deadline that a loaded runner missed.
* fix: bound the handoff drain so a wedged flow cannot hold the quit open
Every Electron E2E spec that boots the app has been failing on
`workspaceSessionReady did not become true`, and the app itself has been
launching to a blank white window: the renderer threw
`ReferenceError: process is not defined` while evaluating a shared chunk,
so React never mounted and no startup step ever ran.
`agent-completion-poll-interval.ts` (renderer) imported one constant,
`PROCESS_TABLE_SNAPSHOT_MAX_STALENESS_MS`, out of
`shared/process-table-snapshot-reader.ts` — a `node:child_process` /
`node:fs/promises` module whose dependency evaluates `process.platform` at
module scope to pick `ps` columns. The renderer runs sandboxed with
contextIsolation, where `process` is undefined, so that module-scope read
threw and took the whole chunk with it. Introduced by #18742; #18780 added a
second module-scope read next to the first.
The constant now lives in `shared/process-table-snapshot.ts`, the
environment-neutral half of the pair, and the reader re-exports it so host
callers are unchanged. The two `ps` column sets read the platform behind a
`typeof process` guard, which defuses the same landmine for any future
renderer import of that module — only hosts ever run the argv.
The regression test walks the renderer import graph (lazy routes included)
from all three entries and refuses any module that reaches a `node:` builtin.
It fails on the pre-fix import with the full 10-hop chain from `main.tsx`.
`worktree-base-divergence-real-git.test.ts` builds cap-sized histories (100 and
101 commits). Every `git commit` detaches `git maintenance run --auto`, whose
commit-graph task arms at 100 new commits, so the fixture reliably spawns a
background `git commit-graph write --split` that keeps creating
`.git/objects/info/commit-graphs` entries after the synchronous exec returns.
The `afterEach` recursive remove is then deleting `.git/objects` underneath a
live writer and dies with ENOTEMPTY — which is how "counts drift in both
directions" failed on main.
Reuse the existing `GIT_FETCH_SKIP_AUTO_MAINTENANCE_CONFIG_ARGS` (it already
covers modern maintenance and legacy auto-gc, so it holds at the Git 2.25
baseline) in the fixture's git helper. Traced spawns of
`git commit-graph write` over a full run of this file: 4 before, 0 after.
The production path under test only runs `rev-list` and `merge-base`, neither of
which triggers auto-maintenance, so there is nothing to fix outside the fixture.
* fix(native-chat): say when a structured launch fell back to a terminal
A definitive refusal already opened a terminal instead of the requested
structured chat, but said nothing — indistinguishable from the bug where the
wrong surface opens. Notify at message severity, since nothing failed.
Also stop putting the raw error in the failure toast's description: it carried
errnos and absolute paths straight into the UI. The detail moves to a warn log
and the toast gets catalog copy, matching how the coded refusals already read.
* fix(native-chat): avoid overstating terminal fallback
---------
Co-authored-by: Merge Sim <sim@local>
#18796 made every SSH Codex background launch wait for the shell-ready marker,
but the client cannot see the remote shell. On a host that never publishes one --
fish, sh, Windows, or a relay predating #18796 -- no marker arrives and delivery
falls back at 1.5s where it used to write at 50ms.
The relay already computes whether it armed the marker; publish that as an
optional `shellReadyArmed` on the spawn reply and let the client skip a wait it
now knows is pointless. Absent stays UNKNOWN and keeps the client's own guess, so
an older host behaves exactly as before; false is only ever an answer a host gave.
It rides every reply, false included, or absent would stop meaning "old host".
A host that did not arm the marker did not arm bracketed paste either, so the
released path still submits raw.
* perf(terminal): cheap-tier process inspection for anchored local agent panes
Every idle local pane's completion cadence ran a full whole-host `ps` (with
`tty=` and `command=`, 0.34-0.50s on a 1,900-process Mac, 1.15s on Linux)
purely to build `foregroundProcessEvidence` that the renderer then discards
for local ids. Add a cheap tier (same job-control columns, no tty/command,
0.03s) gated so that it introduces no user-facing trade-off:
- Only a pane whose last FULL capture proved a recognized agent may take the
cheap tier. Panes with no anchor always take the full capture, so start
discovery keeps today's exact behaviour.
- The cheap tick compares a per-pane fingerprint (root shell pid+start, tpgid,
every descendant's pid+start+pgid+job-control state). Any change, a changed
node-pty foreground name, an unreadable capture, or an incarnation mismatch
escalates to the full capture. A recognized agent's exit is always a pid
vanishing, which the fingerprint always sees.
- A cheap answer OMITS evidence rather than fabricating a tty-less fence.
Remote/restore consumers never send `steadyState`, so they keep the full
capture unchanged.
- `steadyState` is a new optional request field; an old daemon ignores it and
answers with the full capture.
Measured (8 idle panes, 60s, idle cadence, forks counted by column set):
30 full -> 1 full + 29 cheap.
* fix(terminal): route the cheap ps capture through runProcess
The cheap-tier reader imported node:child_process directly, which the
child-process import-boundary and windowsHide ratchet tests reject (CI shards
1/8 and 3/8). Use Orca's single spawn entry point instead; it pins windowsHide
and encodes argv. Map its result onto the capture-error vocabulary:
outputTruncated -> capture_truncated, timedOut -> capture_timeout, non-zero
exit -> ps_exit_<code>. Tests mock at the runProcess seam.
* fix(perf): refuse a pane fingerprint when any descendant start marker is missing
`buildPaneProcessFingerprint` rejected only a missing root start marker; a missing descendant
marker was stamped as `?`. Two captures that both failed to read the same descendant therefore
compared equal, which removes the pid-reuse protection the fingerprint exists to provide: a
recycled pid could make a vanished agent look unchanged, and the cheap tier would keep serving
its name instead of escalating.
Reachable on Linux, where `readLinuxProcStartTime` legitimately returns null when a process
exits between the `ps` capture and the `/proc/<pid>/stat` read.
Every subtree member now needs a start marker or the fingerprint is refused, which sends the
caller to the full capture — the same conservative default every other uncertain path takes.
Reported by CodeRabbit on #18780. The two new tests fail against the previous code with
`expected '4242@2400#4300:|4300@?:4300:+' to be null`.
`readRelayDir` issued one `stat` per symlinked entry and awaited them all in a single
`Promise.all`. A pnpm `node_modules` is hundreds-to-thousands of package symlinks in one
directory, so expanding it over SSH put that many stats in flight at once, saturating libuv's
four-thread pool and delaying every other relay filesystem operation — including the
interactive reads `fs-list-files-scan-coordinator` exists to protect.
The probes now run through `forEachWithConcurrency` at 8, the cap every other bounded probe in
this codebase already uses (`GIT_COMMON_SNAPSHOT_CONCURRENCY`,
`PRUNABLE_EXISTENCE_PROBE_CONCURRENCY`, `SPARSE_CHECKOUT_DETECTION_CONCURRENCY`).
Results and ordering are unchanged: every symlink still resolves to its target's kind, and
`sortDirEntries` still runs afterwards. The new test builds a 60-symlink directory and asserts
the same 60 stats happen with exactly 8 in flight at peak — the probes overlap, and never past the cap.
* fix(agent-session): refuse a pre-commit structured create with an envelope
The create route refused by throwing, which reaches a client as a generic
transport error indistinguishable from a lost answer — so desktop parked the
launch as visibility-unknown with no chat and no terminal. Convert the whole
pre-commit span, everything before `attach`, into a refusal envelope carrying a
code, and name the definitive-refusal allowlist the fallback decision needs.
* fix(agent-session): gate legacy fallback on definitive refusals
* fix(mobile): preserve unknown structured create outcomes
---------
Co-authored-by: Merge Sim <sim@local>
* perf(terminal): tighten the partial-escape-tail benchmark and equivalence test
* perf(terminal): spell the ESC gate the same way as the sibling ingest gates
* test(terminal): differential-fuzz the ESC-free partial-escape-tail gate against the unguarded fold
* test(terminal): make the escape-tail fuzz exhaustive at symbol depth, and cap the fold expectation
Two review findings on the differential fuzz, both about the test faithfully modelling the
function it guards.
The odometer generated strings by symbol depth but the caller filtered on `chunk.length`, which
is the UTF-16 code-unit count. An astral symbol is two code units, so every depth-4 string
containing one was silently skipped and the corpus was not exhaustive at depth 4 the way the
test name claimed. The generator now yields `{ depth, text }` and the caller filters on depth.
That restores the missing strings and takes the pinned corpus from 516,566 to 593,468 - exactly
the count CodeRabbit derived for the intended corpus.
The pairing assertion in the sibling suite compared the capped `advancePartialEscapeTail`
against an uncapped `extractPartialEscapeTail(pending + chunk)`. It passed only because no
pairing in that corpus crosses MAX_PARTIAL_ESCAPE_TAIL_LENGTH; it would have stopped modelling
the function the moment one did. The cap now lives in the expectation, matching the fuzz
oracle.
Re-verified the fuzz still fails on a wrong guard: mutating the gate to a bracket check fails
all four tests with a `gate diverged` assertion on a lone ESC chunk.
Reported by CodeRabbit and pullfrog on #18748.
Four renderer projections scanned a collection inside a loop over another collection, each on a
path that reruns per keystroke or per store write. All four now build the index once per stable
input, which is what the surrounding code already does for its other lookups.
`workspace-kanban-search.ts` called `searchWorktrees`, the convenience wrapper that builds the
palette document index inline. The board's filter hook memoized the whole call on the query, so
every character re-normalized and re-segmented every indexed field of every worktree — and did
it again on every agent-status tick while a query was active, since those churn board
identities. `buildWorkspaceBoardPaletteDocuments` splits out, memoized on
`[worktrees, repoMap]`; only the match reruns per keystroke. This is the shape
`worktree-jump-palette-document-index.ts` already provides for Cmd-J.
`useTabGroupItemProjections` resolved each editor tab against `state.openFiles` — the global
list across every worktree — and each `tabOrder` entry against the group's tabs, with a `.find`
per element. Both are now `Map` lookups, alongside the `terminalTabById` index the same file
already built. The memo key `groupTabs` gets a new identity on any unified-tab write, so this
ran on title, label and colour changes.
`buildSourceControlTree` rebuilt every ancestor path with `segments.slice(0, i + 1).join('/')`
per segment, making tree construction O(files x depth^2) in characters copied — on the path the
Source Control file filter rebuilds per keystroke. The path now accumulates.
`worktree-header-section-boundaries.ts` ran a full `findIndex` over the render rows for every
header row, plus an `indexOf` over the bucket ordering, in two `useMemo`s keyed on `renderRows`
— so it recomputed on every sidebar row-model change, not just during a drag. One indexing pass
each, first-match-wins to match `findIndex`/`indexOf`. The successor index is keyed per bucket
because a repo or group id can appear in more than one bucket ordering; a flat id-keyed map would
pick whichever bucket was iterated first. `worktree-header-section-boundaries.test.ts` pins that
and the first-wins duplicate-header case.
Measured by `pnpm bench:renderer-quadratic-scans`. Three scenarios time the production export
against a reproduction of the pre-change function; the tab-group scenario is modelled on both
sides because the projection lives inside a React hook. Each asserts before/after agree first.
| projection | drives | scale | before | after | |
| --- | --- | --- | --- | --- | --- |
| workspace board filter (per keystroke burst) | production | 300 worktrees x 12 keystrokes | 15.4 ms | 4.4 ms | 3.5x |
| tab-group projections (per unified-tab write) | modelled | 60 tabs x 120 open files | 1.23 ms | 0.34 ms | 3.6x |
| source-control tree build (per filter keystroke) | production | 5000 changed files | 6.2 ms | 3.9 ms | 1.6x |
| sidebar header boundaries (per row-model rebuild) | production | 80 repos x 600 rows | 2.0 ms | 0.8 ms | 2.6x |
The tree and sidebar wins are smaller than the scans they remove because the rest of each
function (tree finalize/sort, per-row size estimation) is linear and now dominates.
4,537 existing sidebar, tab-group and right-sidebar tests pass unmodified.
`advancePartialEscapeTail` runs once per PTY chunk, on the main thread, for every terminal —
visible, hidden or parked — inside `HeadlessEmulator`'s write path. It unconditionally
concatenated the pending tail with the whole chunk and then walked the result one code unit at
a time through a VT500 state machine.
`extractPartialEscapeTail` only leaves `ground` on an ESC byte, so with no pending tail and no
ESC in the chunk the answer is always ''. Taking that case up front skips both the full-chunk
concat and the walk. `String.prototype.includes` is a native scan, so the gate costs
essentially nothing on the chunks it does not short-circuit.
This is the same gate its two neighbours on the very same ingest path already apply —
`TerminalOscCwdTitleScanner.scan` and `TerminalMouseModeMirror.scan`, both carrying a comment
citing their measured share of a 2.2x ingest regression. This call was simply missed.
Measured by `pnpm bench:terminal-partial-escape-tail` over 640 x 16 KB chunks (10 MB), median
of 7 rounds:
| stream shape | before | after | |
| --- | --- | --- | --- |
| ESC-free (build logs, `cat`, piped output) | 40.3 ms | 0.19 ms | 213x |
| SGR-coloured output (gate does not apply) | 30.6 ms | 32.1 ms | 1.0x |
The benchmark proves equivalence over a 226-case corpus before timing, and a new unit test
pairs every pending-tail state the scanner can be left in against every chunk shape, asserting
the gate is indistinguishable from the unconditional fold.
The structured-session read owner emits per SDK event with no coalescing, so a working turn
re-renders `NativeChatMessageList` tens of times a second. `MessageRow` was a plain function
component, so every one of those frames re-ran `nativeChatProseToMarkdown`, an image-block
scan, and a provider-frame lookup for every message in the window — up to the 300-message
pagination limit — even though only the streaming tail had changed.
`MessageRow` moves to its own module and is memoized, and its four derivations fold into the
`useMemo` that already keyed on `message.blocks`. All of its props are primitives or
already-stable references (`onScrollMessageToTop` is a `useCallback`, `onLinkClick` comes
through `useShallow`), so the memo holds for every settled row. The move also takes
`NativeChatMessageList` back under the max-lines limit rather than bumping it.
`NativeChatToolRun` gets the same treatment: it was unmemoized and made four independent
passes over its block list per render.
`useNativeChatTurnStatus` allocated a fresh slice of the whole current turn and ran a nested
`.some()` over it on every render; both now memoize on `[messages, latestUserIndex]`.
Measured by the new `NativeChatMessageList.stream-render.perf.test.tsx`, a 120-message
transcript with a 20-frame streaming turn:
| markdown rebuilds per stream frame | before | after |
| --- | --- | --- |
| 120-message transcript | 120 | 1 |
The test counts real per-row work rather than a render counter, and it fails on the
pre-change component (`expected 120 to be less than 12`), so it is a genuine guard.
Pure memoization on unchanged props: no visual, ordering, scroll-anchoring, or lifecycle
change. All 997 native-chat tests pass unmodified.
* perf(terminals): let idle panes share one process-table capture instead of forking their own
Every visible local pane runs an agent-completion cadence that resolves through
`getStrictProcessTableSnapshot`, and the inspection queue already collapses every
shared-observation task enqueued in the same tick onto a single whole-host `ps`. Independent
±10% jitter per pane defeated that: the jitter was re-rolled on each reschedule, so panes
drifted permanently apart, each landing in its own tick and each missing the snapshot's 500 ms
TTL. Four idle panes cost four captures where one would have served all of them.
Idle panes now aim at a deadline grid anchored at the epoch. The pull-forward is clamped to the
snapshot TTL, so no interval is ever longer than its tier and none is more than 500 ms shorter:
a pane off the grid walks onto it over at most `tier / TTL` steps, costs at most one extra
inspection in total, and no inspection is ever delayed.
Scoped deliberately. A pane with a foreground agent, or one still inside the 10 s post-activity
hot window, keeps its exact interval and its own phase, so the bounded hot cadence is unchanged.
The error-backoff path keeps its jitter, where spreading retries across panes is the point.
Measured by `pnpm bench:agent-inspection-cadence` — whole-host `ps` captures over 60 s at the
2 s idle tier, median of 21 rounds:
| visible panes | before | after | reduction |
| --- | --- | --- | --- |
| 1 | 29 | 29 | 0% |
| 2 | 42 | 30 | 29% |
| 4 | 62 | 31 | 50% |
| 8 | 82 | 32 | 61% |
`process-table-snapshot-reader.ts` measures the `command=` column at 1.15 s of work for 1,948
processes, so these are captures a quiet app was paying for continuously.
All 4,202 existing terminal-pane tests pass unchanged, including the no-evidence cadence suite
that pins the relaxed and hot intervals.
* test(terminals): report n/a instead of dividing by a zero baseline in the cadence benchmark
A window shorter than one cadence tier leaves the baseline capture count at zero, and the
reduction line then divided by it and printed a meaningless percentage. Reported by CodeRabbit
on #18742.