A first-hand Claude exit is not published where it is observed. `handleExit`
re-enters the close ladder and persists the transcript cursor before it emits
`ended`, and only that emission reaches the runtime's recovery chain. So the
runtime's `waitForRecovery` — whose whole job is to drain an in-flight recovery
before teardown stops children — returns immediately for an exit that is still
climbing the ladder, and nothing outside the adapter can tell an observed exit
from a published one.
The integration test for fenced host reconciliation had no handle on that
barrier, so it bounded-polled the lease for 100ms instead. Measured under 16x
local concurrency, publication alone takes 77-204ms: 19/24 runs failed.
Retain the ladder-then-settle tail on the exit record and expose
`drainObservedExits`, fold it into `waitForRecovery`, and export the barrier so
a caller that needs the settled lease can await it. Codex publishes inside its
own exit callback and needs nothing. The test now awaits the barrier: 0/24
under the same load, and it fails on an idle machine without the drain.
* fix: stop a handoff flow from outliving the host that owns its session
A structured handoff runs on the session's serialized chain and nothing in
production awaited it. When the client-side deadline for the switch expired
first, teardown dropped the session map out from under a live flow, and the
flow's own failure notification then threw `agent_session_ownership_unknown`
out of a status publish — an unhandled rejection, plus journal rows written
into a directory that was already being removed.
Three fixes, each with a regression test that fails without it:
- The status publish is a notification, not a mutation: it now reads the fence
without requiring an attached session, so an evicted or torn-down session
makes it a no-op instead of a throw.
- `track` used `.finally`, which forwards a rejection onto a promise nobody
awaits. `drain` settles flows through `allSettled`, so the bookkeeping chain
is now settle-only and cannot resurface one.
- Host teardown drains in-flight handoffs before dropping the session map.
`drain` existed for exactly this and was never wired up.
The integration test's `vi.waitFor` is dropped rather than widened: the request
enqueues the flow on the session's serialized chain before it returns, so the
status read is already ordered behind it. The poll only added a wall-clock
deadline that a loaded runner missed.
* fix: bound the handoff drain so a wedged flow cannot hold the quit open
Every Electron E2E spec that boots the app has been failing on
`workspaceSessionReady did not become true`, and the app itself has been
launching to a blank white window: the renderer threw
`ReferenceError: process is not defined` while evaluating a shared chunk,
so React never mounted and no startup step ever ran.
`agent-completion-poll-interval.ts` (renderer) imported one constant,
`PROCESS_TABLE_SNAPSHOT_MAX_STALENESS_MS`, out of
`shared/process-table-snapshot-reader.ts` — a `node:child_process` /
`node:fs/promises` module whose dependency evaluates `process.platform` at
module scope to pick `ps` columns. The renderer runs sandboxed with
contextIsolation, where `process` is undefined, so that module-scope read
threw and took the whole chunk with it. Introduced by #18742; #18780 added a
second module-scope read next to the first.
The constant now lives in `shared/process-table-snapshot.ts`, the
environment-neutral half of the pair, and the reader re-exports it so host
callers are unchanged. The two `ps` column sets read the platform behind a
`typeof process` guard, which defuses the same landmine for any future
renderer import of that module — only hosts ever run the argv.
The regression test walks the renderer import graph (lazy routes included)
from all three entries and refuses any module that reaches a `node:` builtin.
It fails on the pre-fix import with the full 10-hop chain from `main.tsx`.
The same-cap wave validator approves 19 cells (c7-c26 plus the Asia cells
c27-c29), but the canary script it drives hard-rejected anything outside the
16 US capacity cells, so the first Asia same-cap canary failed closed at
isolate. Give the canary an explicit --approved-cells switch that selects the
same-cap allowlist, and pass it from the four same-cap job invocations. With no
switch the behaviour is unchanged, so the US-only capacity workflow keeps its
scope.
`worktree-base-divergence-real-git.test.ts` builds cap-sized histories (100 and
101 commits). Every `git commit` detaches `git maintenance run --auto`, whose
commit-graph task arms at 100 new commits, so the fixture reliably spawns a
background `git commit-graph write --split` that keeps creating
`.git/objects/info/commit-graphs` entries after the synchronous exec returns.
The `afterEach` recursive remove is then deleting `.git/objects` underneath a
live writer and dies with ENOTEMPTY — which is how "counts drift in both
directions" failed on main.
Reuse the existing `GIT_FETCH_SKIP_AUTO_MAINTENANCE_CONFIG_ARGS` (it already
covers modern maintenance and legacy auto-gc, so it holds at the Git 2.25
baseline) in the fixture's git helper. Traced spawns of
`git commit-graph write` over a full run of this file: 4 before, 0 after.
The production path under test only runs `rev-list` and `merge-base`, neither of
which triggers auto-maintenance, so there is nothing to fix outside the fixture.
* fix(native-chat): say when a structured launch fell back to a terminal
A definitive refusal already opened a terminal instead of the requested
structured chat, but said nothing — indistinguishable from the bug where the
wrong surface opens. Notify at message severity, since nothing failed.
Also stop putting the raw error in the failure toast's description: it carried
errnos and absolute paths straight into the UI. The detail moves to a warn log
and the toast gets catalog copy, matching how the coded refusals already read.
* fix(native-chat): avoid overstating terminal fallback
---------
Co-authored-by: Merge Sim <sim@local>
#18796 made every SSH Codex background launch wait for the shell-ready marker,
but the client cannot see the remote shell. On a host that never publishes one --
fish, sh, Windows, or a relay predating #18796 -- no marker arrives and delivery
falls back at 1.5s where it used to write at 50ms.
The relay already computes whether it armed the marker; publish that as an
optional `shellReadyArmed` on the spawn reply and let the client skip a wait it
now knows is pointless. Absent stays UNKNOWN and keeps the client's own guess, so
an older host behaves exactly as before; false is only ever an answer a host gave.
It rides every reply, false included, or absent would stop meaning "old host".
A host that did not arm the marker did not arm bracketed paste either, so the
released path still submits raw.