Commit Graph
457 Commits
Author SHA1 Message Date
Neil fb52c0602a fix(terminal): release xterm's DEC 2026 render hold instead of waiting out its 1s timeout (#23920)
* fix(terminal): release xterm's DEC 2026 render hold instead of waiting out its 1s timeout

xterm paints nothing while DEC mode 2026 (synchronized output) is open and only
force-flushes after 1000ms. Codex wraps every draw in mode 2026, so any byte gap
or chunk split that loses the closing \x1b[?2026l freezes the pane for a full
second and then repaints in one burst.

Orca never emitted \x1b[?2026l anywhere, and three paths could destroy a TUI's:
the per-PTY pending cap drops buffered output wholesale (mode 2031 was already
salvaged there, 2026 was not), main sliced pending data at a blind 16KB offset
that can land inside an open frame or sever the 8-byte marker, and the renderer's
backlog warnings replace a queued tail that may hold the close.

- salvage the 2026 latch across dropped output, mirroring the existing 2031
  salvage, and append the release on both delivery sites
- ground 2026 in RESET_AFTER_BYTE_GAP and the replay baseline, and in both
  backlog warnings, so every drop path is self-healing
- make main's 16KB flush split frame-aware instead of a blind byte offset
- lift the synchronized-output scanner into shared/ so main and the renderer
  use one implementation

Closing a frame early costs one premature repaint; leaving it open costs a
second of blank screen, so the asymmetry favours always closing.

Also adds the reproduction this needed: the pre-existing typing bench observes
the xterm BUFFER, which the parser fills while rendering is held, so it scored
these freezes as fast echoes.

* fix(terminal): stop the renderer's queue drain cutting inside an open DEC 2026 frame

takeQueuedChunk sliced a queued chunk at a blind byte offset to fit the 16KB
coalescing budget, which can strand a frame's closing \x1b[?2026l in the residual
until a later drain. Same defect as main's flush split, same fix: reuse the
frame-aware split helper.

Usually masked because the drain coalesces adjacent chunks and reassembles what
main split, but not when the budget boundary falls inside a frame.

* fix(relay): keep the SSH path's bounded slice outside an open DEC 2026 frame

pty-handler split pending output at a byte offset with a surrogate-pair guard but
no synchronized-output awareness, so a frame straddling the 16KB wire slice had
its closing \x1b[?2026l stranded in the remainder — the same defect just fixed on
the local path, on the path AGENTS.md requires us to consider.

Placed before the surrogate guard so that guard keeps the final say, and floored
at 2 so frame alignment can never walk a healthy slice into the guard's
decrement and then into the chunkChars <= 0 pause-and-retry path.

Also drops a dead `splitAt === 0` branch in takeQueuedChunk: both callers pass a
positive limit and the helper never returns 0 for one.

The two new split tests were each confirmed to fail without their fix.

* test(terminal): sweep the DEC 2026 split helper over escape-sequence shapes and every limit

Covers OSC 52, DCS, repeated open/close markers and limits 1..len+3, asserting the
result never exceeds the limit, never reaches 0, and stays byte-exact. Also pins
that a buffer beginning inside an open frame degrades to the blind offset rather
than doing something worse, and documents that callers do not thread latch state.

* fix(terminal): ground DEC 2026 on the daemon slice, the recovery replays, and the process boundary

Four more sites could strand the latch, found by sweeping every path that drops,
splits, or replays terminal bytes.

- daemon-stream-data-batcher: the 64KB bulk-write slice used a surrogate-only
  clamp, and its remainder is HELD until 'drain' — "seconds for multi-MB
  backlogs" per the file's own note. A frame straddling that boundary parked its
  \x1b[?2026l behind the hold, blanking the pane past xterm's 1s timeout once per
  frame for as long as the backlog lasted. This is the default daemon-backed pane
  path, so it is the one users actually hit. The new
  clampToSafeBulkWriteSplitIndex frame-aligns first and surrogate-clamps last,
  and lives in daemon-stream-data-split alongside the policy it belongs to.
- replay-data-drain and remote-runtime-terminal-binary-snapshots wrote a bare
  \x1b[2J\x1b[3J\x1b[H, which does not clear mode 2026 — so on the SSH/remote
  reconnect path, the very event most likely to sever a frame, the whole replay
  could paint nothing.
- ipc-pty-attach: trimIncompleteTerminalControlTail can cut a half-written
  \x1b[?2026l while its opening marker survives in the replayed prefix.
- PROCESS_BOUNDARY_GROUND: the "process that armed these modes is gone" ground
  omitted 2026, the last unexplained gap in that file. A disable, so it still
  satisfies the recovery barrier's ownership scan (only ?25h may be an enable).

Recovery-path expectations updated where they pin the emitted bytes. Deliberately
NOT touched: apply-reattach-payload and ssh-snapshot-prepaint already ground via
buildSnapshotReplayPrologue.

Still unfixed, deferred with reason: terminal-output-frame-chunks.ts splits the
remote wire on accumulated UTF-8 byte width and needs a different shape than the
char-index helper; desktop clients reassemble in main's pending buffer, so the
exposure is mobile/web only.

* fix(terminal): emit the DEC 2026 release before the mode-2031 tail, and stop claiming the drop path writes it

Two corrections from adversarial review of the earlier commits.

1. Ordering bug I introduced. getDroppedMode2031RendererData ends with
   `state.tail`, which extractPrivateModeScanTail deliberately retains as an
   INCOMPLETE private-mode sequence so the next chunk can resolve it. Appending the
   2026 release after it put an ESC behind a dangling CSI, aborting it and silently
   losing whatever mode spanned the drop boundary. The release now goes first.

2. The drop-path release does not reach xterm in the dominant case, and the comment
   now says so instead of implying otherwise. live-data-callback's droppedOutput
   branch discards `data` and salvages only queries
   (salvageRendererQueriesFromDiscardedRestoreData handles CPR/DA1/OSC colour;
   \x1b[?2026l is not a query), so for hidden panes and visible panes outside
   foreground-restore backpressure the synthesized release was dropped. The grounded
   snapshot replay releases the latch instead.

   I tried writing it through writePtyOutputToXterm there and reverted: it consumes
   the pending hidden-output snapshot and broke
   pty-connection-hidden-snapshot-resize-signals ("re-restores a skipped alt frame"),
   so the release rides the restore rather than perturbing that state machine.
   Residual gap, documented: a cap-dropped pane whose restore never arrives.

The salvage is still load-bearing on the fall-through path, so it stays.

* fix(terminal): release DEC 2026 on the reattach clears, floor the split, and correct the freeze framing

Remaining findings from adversarial review.

- apply-reattach-payload's three bare-clear branches (:63 daemon snapshot, :229
  relay replay, :269 cold restore) had no release anywhere in their sequence: I
  checked all seven POST_REPLAY_* profiles reachable via chooseReattachReplayReset
  and none contains \x1b[?2026l. Only the buildMainModelSnapshotReplayWrites branch
  was grounded, so covering the streamed replay path and not the main reattach path
  was inconsistent. Verified no production code matches these clear strings — the
  three test updates are mock equality, and each was confirmed to fail without the
  source change.
- clampToSafeBulkWriteSplitIndex could return 0 (('\u{1F600}aaaa', 1) — alignment
  returns 1, the surrogate clamp decrements to 0), which would leave a zero-length
  slice that never shifts the batcher's queue entry and spin its drain loop.
  Unreachable from today's only caller, but it is exported with an unstated
  precondition. Floored at 1.
- Frame alignment could halve per-PTY flush throughput: main re-queues the
  remainder with eligibleRound = round + 1, so the shortfall cannot be refilled in
  the same round, and aligned size is floor(W/F)*F — 50% worst case in the 8-16KB
  band, which is exactly the full-screen redraw burst that reaches the pending cap.
  Alignment is now rejected below half the window, preferring throughput and
  letting the reset profiles release the latch.

Framing corrected throughout: bufferRows records a row range and clears nothing, so
the pane freezes on its last painted frame — it does not go blank. The real trade is
"stale but coherent for <=1s" versus "immediate partial frame", and
RESET_AFTER_BYTE_GAP (written alone, with no repaint behind it in the same write) is
the one site that can newly flash a partial frame. Said so at the constant instead
of implying the release is free.

* fix(terminal): rename the shape-flagged symbols the anti-slop audit rejects

CI's anti-slop gate rejects "shape" in symbol names as structural rather than
domain language: `shapes` -> `outputSamples`, and
`writeCodexShapedEchoProbeScript`/`codexShapedEchoProbeScript` ->
`writeCodexEchoProbeScript`/`codexEchoProbeScript`.
2026-09-29 20:27:30 -07:00
Jinwoo Hong fb67d5d7c3 fix(runtime): stop a busy Codex 0.150-0.157 pane reading as tui-idle (#23805)
* fix(runtime): stop a busy Codex 0.150-0.157 pane reading as tui-idle

The startup header box (OpenAI Codex / model: / directory:) stays on
screen and in the tail for the whole session, so as tier-1 evidence it
settled tui-idle mid-turn. For a codex pane it now counts only in the
quiet lane, held to the same quiescence as the composer.

* fix(runtime): keep a restored Codex pane's header as tier-1 readiness

A restored or reattached pane has no lastOutputAt, so the quiet lane that
now holds a Codex header can never fire and the wait sat pending until
timeout, where main settled it. Gate the Codex tier-1 veto on the output
clock rather than the agent name, and share one settled-prompt helper.

* test(runtime): read no screen in the restored Codex pane test

* test(runtime): pin the clock in the clockless Codex header cases

* test(runtime): name the screen-readiness comparison for what it proves
2026-09-29 22:28:55 -04:00
Brennan Benson ad2e1b5efa fix(terminal): restore the mouse format with mouse tracking, so phone swipes don't type into Codex (#23946)
* fix(terminal): restore the mouse encoding with mouse tracking in every snapshot

Swiping to scroll Codex from the phone on a Windows host typed legacy
`ESC [ M` mouse reports into the Codex composer (#23818). SerializeAddon
re-arms mouse tracking (?1000h/?1002h/?1003h) but never the SGR encoding
(?1006h/?1016h). Any snapshot taken from a desktop pane's xterm (the
runtime seeds its headless model from it after a reattach, and serves it
to remote viewers when no model exists) therefore restored "tracking on,
legacy encoding", and the phone encoded wheel events as X10 bytes, which
ConPTY hands to Codex as keystrokes.

serializeWithAbsoluteCursor, the one wrapper every Orca snapshot producer
uses, now appends the encoding xterm itself parsed, read from xterm's
mouse state service. The daemon/runtime headless model reads tracking and
encoding from xterm too, so its regex mirror of the DECSET stream is
deleted (one source of truth; one less regex pass per PTY chunk).

Mixed versions: no wire field changes. A new host's snapshot carries an
extra DECSET that old desktop and phone clients already parse; an old
host's snapshot restores exactly as before. With tracking off the encoding
alone sends no reports, so the wheel still scrolls scrollback.

* test(terminal): pin the mouse-encoding read against the renderer xterm build

* fix(terminal): type the xterm mouse-state read behind named shapes
2026-09-29 19:23:25 -07:00
Brennan Benson 59ef74876f fix(terminal): the terminal's owner answers colour queries for the terminal's whole life (#23925)
* fix(terminal): the PTY owner answers OSC 10/11 for the terminal's whole life

Codex and Claude's `theme: auto` ask the terminal for its foreground and
background colours (OSC 10/11) and pick their colours from the reply. Orca
answered in the process that owns the PTY only for agent launches and only
for 5 s; after that the query was handed to whichever viewer was attached.
On Windows ConPTY the owner kept swallowing the query but stopped answering
it, so a Codex started from an older shell tab lost its message shading
(#22332). On a headless `orca serve` host no viewer existed yet, so a Codex
started before anyone attached got no reply at all (#22500).

The owner (in-process provider, terminal daemon, SSH relay) now answers
every OSC 10/11 query for the PTY's whole life and strips it, so no
downstream view ever sees one to answer twice. It answers from, in order:
the host-wide viewer theme pushed to that process, the creating viewer's
colours sent at spawn (now for every PTY, not only agents), and Orca's
default dark theme. The desktop pushes its renderer theme to every owner
on change and on (re)connect: a daemon request gated on protocol v38, and
an SSH relay notification that older relays ignore. The answer-once rule,
the 5 s colour window and the colour-authority handoff are removed; Kitty
keyboard queries keep their startup window.

Viewer-side answerers (renderer xterm, main's hidden-pane model responder,
the mobile webview) stay as the fallback for older owners, which still hand
queries off; they are never reached for a new owner.

* fix(terminal): answer OSC 10/11 with the colours the pane is really painted with

Review follow-ups to the lifetime PTY-owner colour answerer.

- The theme catalog moves to src/shared so the renderer and the PTY owners
  read one source; the owner's last-resort default is derived from it
  rather than copied.
- Main seeds every owner from the host's saved theme settings (light or
  dark, custom themes, colour overrides) at startup, so a headless host
  and a desktop pane that queries before the renderer's first push are
  not told dark to a light-theme user. The renderer's push replaces it.
- Colours an app sets with OSC 10/11, and clears with OSC 110/111, are
  tracked per terminal and reported back, as a viewer paints them; a theme
  change drops them, as a viewer's theme apply does.
- A terminal a paired client created with its own colours answers with
  those, not the host's theme (`colorSource: 'remote-viewer'` on the
  spawn intent), so a light client on a dark host is told light.
- After the 5 s startup window a reply's echo is watched for 512 bytes
  instead of 256 KB, and a torn query candidate is released after 500 ms
  rather than held indefinitely.

* fix(terminal): keep the long echo watch for relayed replies; one theme lookup

The 512-byte post-startup echo watch now applies only to replies the PTY
owner produced itself. A viewer's reply relayed through
answerLiveQueryReply keeps the 256 KB watch, because a cooked-mode app can
keep printing after it queries and the echo then trails that output.

The renderer's getTerminalTheme now calls the shared lookupTerminalTheme,
so the custom-vs-built-in theme lookup exists once.

* perf(terminal): scan colour overrides in one pass over each PTY chunk

Two indexOf searches per OSC went quadratic on long runs of ST-terminated
hyperlinks, and the tracker now sees every chunk of every terminal.

* fix(terminal): one host viewer colour value, set by whichever viewer acted last

A paired client's colours reached the host only as frozen spawn colours on
terminal.create, tagged remote-viewer. UI-started agent sessions on a headless
host answered OSC 10/11 with the host's saved theme, and a client's theme flip
never reached panes it had created.

The host now holds one viewer colour value that every PTY owner answers with.
The desktop renderer's push, a window focus on the host, the new
terminal.setViewerColors RPC, and terminal.create colours from older clients
all set it; equal values do not re-notify daemons or relays. The remote-viewer
tag (colorSource / terminalColorQuerySource / spawnFromRemoteViewer) is gone;
it never shipped in a release.

* fix(terminal): paired clients push their terminal colours on connect, change and focus

The renderer publisher now hands each published fg/bg to subscribers. A new
remote-runtime-terminal-color-push module calls terminal.setViewerColors on
every host this client is connected to when it connects (or the host restarts),
when the colours change, and when the window gains focus. A host that answers
method_not_found or forbidden is not asked again until it reconnects.

The app shell also republishes terminal view attributes on settings and system
theme changes, so a theme change reaches main and paired hosts with no
terminal pane open.

* fix(terminal): a host with its own window answers OSC 10/11 with its own theme

Round 1 kept one host-wide viewer colour value set by whichever viewer acted
last, so a paired client's push (reconnect after sleep, a dusk theme flip)
took over the host desktop's own panes until its window regained focus.

The value is now derived: this host's renderer colours when a local window
has pushed, otherwise the last paired client's push (terminal.setViewerColors
or terminal.create colours), otherwise the saved theme. Only a headless host
takes a client's theme. The window-focus reassert and the identical-re-push
takeover are gone; owners are notified only when the derived value changes.

* perf(terminal): scan only OSC starts for colour queries once the Kitty window closes

The PTY owner answers OSC 10/11 for the terminal's whole life, and it tried
every ESC as a query start: a 240 KB SGR-heavy read cost about 2 ms and 256 KB
of bare ESC about 15 ms, long after startup.

Once the Kitty query window closes only an OSC colour query can match, so the
scan jumps between ESC ] starts, plus a trailing lone ESC so a query torn right
after its ESC still resolves on the next read. Output and replies are
unchanged.
2026-09-29 19:06:57 -07:00
Jinwoo Hong 26bb7c23f1 fix(terminal): run Orca's cmd.exe, path-named and setup-gated Codex launches without the shared server (#23933)
* fix(terminal): give plain shells and cmd.exe Codex launches --no-daemon

Plain bash, zsh and fish tabs were never wrapped, so a typed codex skipped the
shell function that adds --no-daemon. Wrap them (bash keeps its prompt and
DEBUG trap untouched unless Orca asked for command markers), add --no-daemon
host-side where no function can run (cmd.exe, path-named binaries), and move
new tabs to a v38 terminal daemon so they get the new wrappers.

* fix(terminal): keep plain bash a login shell and give the setup gate the codex function

Plain bash and Git Bash tabs launch exactly as before again: the rcfile
wrapper would have made every one a non-login shell. Plain tabs on the
user's configured shell args stay unwrapped on both transports. The
wait-for-setup gate's bash -lc now defines the codex function, so a
sequenced Codex launch gets --no-daemon from the binary it actually runs.

* fix(terminal): define the setup gate's codex function after setup finishes

Setup can be what puts codex on PATH, so defining the function before the
marker wait found no binary and skipped --no-daemon.

* refactor(terminal): fold the SSH/WSL guard into the Codex launch planner and bound the gate test

* fix(terminal): honour the pane's env deletions in the Codex opt-out check

Also pin the setup-gate test's fake codex ahead of path_helper's PATH.

* revert(terminal): launch plain zsh and fish tabs exactly as on main

Drops the always-wrap for plain zsh and fish, the configured-args guard
that only served it, and the v38 daemon bump: the daemon's launch configs
and generated wrappers are byte-identical to main again. Keeps the
host-side --no-daemon for cmd.exe and path-named launches and the setup
gate's codex function.
2026-09-29 21:40:53 -04:00
Neil 7980ab9942 test: remove mock echoes and duplicate contracts that only reading finds (#23953)
* test: remove mock echoes and duplicate contracts that only reading finds

Two veins in one wave, both requiring the production path to be read rather than
pattern-matched.

Mock echoes (36 strongest candidates reviewed, 2 real): the flagged shape —
literals shared between a mock factory and an expect matcher — is almost always
a test feeding an input and asserting a transform. The two genuine echoes are in
`pty-management.test.ts`, where the handler returns
`getDaemonFolderAccessMismatch(identity)` verbatim, so asserting the mock's own
`evidence('allowed')` object and its `null` proved only the mock. One case's own
comment conceded the handler makes no decision. The branch-exercising cases in
that file stay.

Semantic sweep of 80 files no detector flagged, 27 junk cases removed. What it
found has no mechanical signature:
- a handler that is literally `() => getComputerUsePermissionStatus()`, so
  `resolves.toBe(result)` guarded nothing;
- "does not mutate a stale registration off Linux" whose refusal came from a
  DIFFERENT guard — `cli.ts` has no platform check, and the test's own mock made
  `resolveAppImageRuntimeIdentity` return null, which a sibling case already owns;
- "keeps probing a host that is still retained" asserting a no-op, because
  `retainScopes` only cancels queued probes and stores no state;
- two cases comparing against `referenceAllowedRoots`, a verbatim copy of the
  pre-change algorithm kept in the test file — expected values produced by the
  thing under test. The third such case stays: it calls the old algorithm to
  COUNT its work (100_000 containment checks vs 100), a real bound on the
  authorization hot path;
- a remote-folder refusal that came from the fake provider's own rejection
  message rather than any Orca guard;
- "advertises each capability once", where a duplicate entry is inert in
  production because membership is `includes`;
- an Antigravity `scaffold self-check` built on a hand-typed five-line screen —
  the exact fixture shape docs/reference/antigravity-readiness-evidence.md blames
  for five failed detector attempts — whose banner had already drifted to
  `Antigravity CLI 1.0.3` against a real captured `1.2.0`. The raw-capture
  provenance guard and the transcript checklist ratchet in that file stay.

No production file is touched and no test file is deleted.

* test: retire duplicate daemon and filesystem cases the sweep found

Continues the semantic sweep into src/main/ipc filesystem handlers and
src/main/daemon. 16 cases removed across 13 files; no production file touched.

The recurring shape is a case that reaches the same branch as its sibling by a
different-looking route:
- `parseArgs` is a flag-scanning loop, so "handles flags in any order" asserts the
  identical result object as the in-order case, and "throws with no args" lands on
  the same `Usage:` throw the two missing-flag cases already reach;
- an ENOENT case with "no code at all" and its sibling with a numeric transport
  code both fall through to the same `ENOENT_MESSAGE.test` branch;
- "establishes connection with hello handshake" and "receives stream events" are
  strict subsets of cases that require two matching authenticated sockets and a
  frame split inside a multibyte character.

Two were vacuous rather than duplicated: a fixture self-comparison asserting
`String.normalize` gives different NFC and NFD spellings, and a
`not.toThrow` batch case whose 200k events are truncated by
`MAX_BATCHED_WATCHER_EVENTS = 5_000` long before the argument spread it was
written to exercise.

One whole-file deletion was reversed: `freebuff-detection.test.ts` looked like a
table restatement, but the sibling detection test covers only dependency ordering
and platform gating, not freebuff/codebuff independence — and those two names
share a suffix, so it is the only guard against one shadowing the other.
2026-09-29 17:35:12 -07:00
Jinwoo Hong 8f4a740b90 fix(daemon): roll Codex no-daemon shell launch into a fresh v37 daemon (#23907)
Terminal daemons survive app updates, so new tabs keep spawning from the
old v36 daemon and never get #23900's Codex shell function. Bump to v37
so new tabs move to a fresh daemon; v36 owners stay attachable.
2026-09-29 14:04:51 -04:00
Jinwoo Hong 9420d49bcb fix(terminal): run Codex in Orca terminals without the shared background server (#23900) 2026-09-29 10:22:27 -07:00
Neil 31012aeb09 test: remove assertion-free probes, copied inventories and export-shape checks (#23816)
Second audit wave, targeting three more junk patterns:

- assertion-free cases that run code and assert nothing, so they pass no
  matter what the code does;
- inventory literals re-typed from a production declaration, where the only
  way the assertion can fail is someone editing one of the two copies;
- export key-set and export-shape loops (`typeof x === 'function'` over every
  export) that restate what TypeScript already enforces.

Yield is much smaller than wave 1 on purpose: the assertion-free scanner has
a high false-positive rate, because many flagged blocks assert through a
shared helper or their oracle is "this must not throw". Those were kept.

`mobileWebCheckArgs` in `config/scripts/run-mobile-web-app-checks.mjs` is
de-exported — after the inventory comparison went away, nothing outside the
module read it.
2026-09-29 02:21:47 -07:00
Brennan Benson 36c473ea1a fix(runtime): wait out Codex 0.157's startup screen, and stop at Codex's startup dialogs, before typing a worker brief (#23745)
* fix(runtime): wait for Codex's live chat before typing a worker brief

Codex 0.157 draws a provisional startup screen (header reads model: loading) and
discards typed input while it starts its shared daemon behind it; a fresh Codex
home makes that window seconds long, so worker-start pasted briefs that were
truncated or never submitted. Codex 0.158 dropped the header labels Orca matched,
so worker-start stopped seeing Codex as ready at all.

Readiness now requires Codex's live chat on both layouts: the provisional header
vetoes a text match unless the live status row is already painted, and 0.158's
greeting layout counts once that status row appears. Codex 0.158's model
announcement dialog is reported as a blocked prompt instead of receiving the brief.

* fix(runtime): recognise Codex's provisional screen from the text copy and the screen probe

Live worker-start on a fresh Codex 0.157 home still typed during the daemon
start: Codex leaves its alternate screen for that window, so the live screen
showed no header and the screen-based veto never fired. The text copy keeps the
provisional header until the live chat paints its status row, so the veto now
reads it there. The tui-idle visible-screen probe used the bare text rule on the
rendered screen; it now goes through the same body rule.

* test(daemon): register the new Codex captures' known serializer divergences

The serialize round-trip replay picks up every fixture under
runtime/__fixtures__, and the three new Codex 0.157/0.158 captures showed
48/9/8 "new-fail" checkpoints against an expected 0, turning CI red. They
are the existing live-pen colour leak on restored cells, the same class as
the other Codex and DSH entries; this branch changes no serializer code.

* fix(runtime): keep Qoder off the Codex screen probe change; drop an unbacked row filter

- The tui-idle visible-screen probe now classified Qoder panes with
  isQoderComposerReady, which skips the working veto evaluateTuiIdle applies
  first. Qoder paints its composer mid-turn, so an adopted Qoder pane whose
  hooks said "working" settled the wait immediately. Only Codex and unknown
  panes take the body rule there; every other agent keeps its old verdict.
- The live status-row check skipped rows containing "waiting for startup",
  a string Codex 0.157/0.158's TUI never prints. The line-folded text copy
  keeps a whole screen on one line, so the filter could only ever veto the
  real status row. It is now a bounded includes() with no split.
- Lowercase the wait text once in isKnownReadyPromptBody.
- Restore the per-frame "screen never takes a settled header away" check,
  guarded on the provisional veto, instead of checking the final frame only.

* fix(runtime): stop reporting Codex 0.158's model announcement once it is answered

The announcement's choices stay in the text copy after the user answers it,
and the existing dismissal check needed the model:/directory: labels that
0.158's header lacks, so tui-idle waits and the agent-status query kept
reporting codex-model-migration-prompt over a live chat. Codex repaints its
whole screen, header included, when a startup dialog closes, so the header
after the dialog now marks it answered.

Also corrects the live-chat marker comments: the middle dot also comes from
the daemon session's agents hint row and the warnings notice, not only the
status row.

* docs(runtime): note that Codex startup dialogs also draw the live-footer dot

* fix(runtime): recognise Codex 0.157/0.158 startup dialogs by the rows they really print

Codex 0.157 and 0.158 no longer print `Press enter to continue/confirm` on
their startup dialogs; they print key rows instead (`enter continue · esc
skip`, `enter confirm · esc skip`, `enter/esc continue · ctrl+c quit`). The
update, hooks-review and model-migration matchers still required the old
wording, so none of these dialogs was reported as blocked. On 0.157 the
dialog's `·` also satisfied the live-footer check, so a tui-idle wait read
the update dialog as ready and worker-start would type the brief into it,
where Enter picks "Update now" (npm install -g, Codex exits). On 0.158 the
wait timed out instead of reporting blocked.

The matchers now accept the old wording or the new row, tolerating the
spaces the line-folded text copy drops around `·`. Each one matches from
the dialog's first `·` (for the update dialog that is its title row,
`Update available · 0.157.0 → …`), so the dialog is blocked from the same
character that would otherwise make the provisional header read as live.
The retired-model notice without choices has a catalog-supplied heading
(`GPT-5.4 is no longer available`), so it is matched by its own key row.
No new blocked-reason value. The startup-dialog matchers move to
startup-dialog-blocked-signals.ts to keep terminal-wait-detection.ts under
the line limit.

Backed by six real captures (update available, hooks review, retired model
without choices, each on 0.157.0 and 0.158.0), replayed frame by frame and
through a tui-idle wait; the serializer round-trip replay registers their
existing live-pen colour divergences.

* fix(runtime): keep reporting Codex's retired-model notice after a relaunch in the same pane

The retired-model notice is matched by its key row alone, and the matcher
took the first `enter/esc continue ·` in the live window while every other
startup-dialog matcher takes the last. Quitting Codex from the notice and
relaunching it in the same pane leaves the old copy ahead of the new
launch's header, so the header read as having dismissed the new notice:
0.157 then read ready and a worker brief would be typed into the dialog.

Take the last key row, and replay each captured dialog quit-and-relaunched
to pin all six.

* fix(runtime): match Codex startup dialogs by the rows the text copy keeps intact

Codex 0.157+ paints each startup dialog over its startup screen by cell
diff, so Orca's line-folded text copy can drop letters and spaces from a
heading: #23765's 0.157.1 capture reads `Updat available`. The update
matcher needed `update available`, so on that capture tier 1 read the
dialog's own `·` as the live chat's footer and a tui-idle wait settled
ready on the update dialog, whose Enter picks "Update now".

Match each dialog from its first `·` by rows Codex prints as fixed
literals: `available · <version>` and `enter continue · esc skip`
(update), `enter confirm ·` (hooks review), `enter/esc continue|confirm ·`
(model notices, which also covers 0.158's new-model announcement, so
its choice-text matcher goes). Legacy `Press enter to …` wording still
matches. Add the new rows to the blocked-signal prefilter, and replay
#23765's 0.157.1 update capture in the dialog suite.

* fix(runtime): don't name Codex's mid-session pickers a hooks review

Codex's rate-limit reset popup (and its other pickers) end their key row
with `enter confirm · esc back`, which the hooks-review row matched now
that it no longer needs the heading. Exclude `esc back` instead of
requiring `esc skip`, so a half-painted hooks-review row still blocks.
2026-09-29 01:16:20 -07:00
Jinwoo Hong c8af48d8a4 fix(runtime): settle a quiet Codex composer as ready on every version (#23765)
* fix(runtime): settle a quiet Codex composer as tui-idle on every version

Codex 0.158 dropped `model:`/`directory:` from its startup header, which both
Codex readiness rules require, so `worker-start --agent codex` timed out; an
idle Codex pane after a turn also had no readiness signal once the header left
the screen.

Generalize the Muse tier-1b lane into a quiet-ready-screen lane: a Codex (or
agent-unknown) pane whose live screen shows the empty composer placeholder,
no `to interrupt)` status row, no header `loading`, and no dialog wording in
its live window, settles once the stream has been quiet for the tui-idle
quiescence window. Additive only: the tier-1 rules and the Muse rule are
unchanged. Fixtures: codex 0.150.1-0.158.0 captures at 120x40, including
chunk-timed turns.

* refactor(runtime): anchor the Codex quiet lane to the empty composer line

Move the Codex screen rules into codex-terminal-readiness.ts and the quiet-screen
body beside isKnownReadyPromptBody. The composer rule now matches only the
`› Ask Codex to do anything` line and drops its dialog markers: every Codex
dialog replaces the composer, and an answer ending "Would you like to…?" above a
live composer must not hold the lane forever. The quiet lane checks quiescence
before reading the screen. Trim the redundant startup and untimed turn fixtures.

* fix(runtime): read Codex's busy row above the composer and scope the lane to codex panes

* fix(runtime): read only Codex's live status row above the composer
2026-09-29 00:39:01 -04:00
OrcaWinandm4air 813aff8f8a fix(opencode): stop OpenCode 2 loading a stale plugin from the retired shared hooks dir (#23500)
* fix(opencode): stop OpenCode 2 loading a stale plugin from the retired shared hooks dir

Before 1.4.209 Orca pointed OPENCODE_CONFIG_DIR at <userData>/opencode-hooks/shared
and wrote a server()-only status plugin there. 1.4.209 moved the plugin to OpenCode's
global config dir and 1.4.210 added the v2 setup() export, but nothing rewrote the
old file. Shells, daemon-persisted panes and OpenCode 2 background services that
still carry that OPENCODE_CONFIG_DIR load only that dir under OpenCode 2 (it replaces
the global dir), so the v2 loader rejects the stale plugin with "Plugin must export a
default definition with an id and an effect or setup function" and pane status dies.

- Refresh the plugin in the retired shared dir (only when it already exists and its
  content differs) so OpenCode processes started later from old shells load the dual
  v1/v2 export. Runs on OpenCode pane spawns and on any spawn that inherits the
  retired dir, even with agent status hooks off.
- Drop an inherited OPENCODE_CONFIG_DIR / ORCA_OPENCODE_* marker that points at the
  retired dir when building a new pane env, so new panes use global discovery.

Limitation: an OpenCode 2 background service already running from an old pane keeps
its cached copy of the stale module even after the file is rewritten (verified with
opencode2 v2.0.18). It must be restarted (`opencode service restart`); a restart from
a new Orca pane then picks up the global config because the env is stripped.

* fix(opencode): harden legacy plugin repair and inherited config cleanup

* fix(opencode): preserve daemon-owned user config during legacy cleanup

* fix(opencode): sanitize inherited sources and repair unseen legacy copies

* test(opencode): update shared PTY mocks for legacy repair

* test(opencode): annotate shared repair mock signature

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-28 14:55:42 -07:00
+7 1f6f8523ab feat(terminal): add Reset Terminal that clears leftover input modes on the host and pane (#23602)
* test(native-chat): await the async history and journal snapshot in three tests (#23560)

#22835 made history() and journalSnapshot() async; tests from #23502 and #22944 still call them synchronously, so the typecheck job is red on every PR while main pushes do not run it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(usage): show ZCode Coding Plan quota on current main (#23520)

Shows the ZCode Coding Plan quota in the status bar alongside the Claude and Codex usage readouts, reading the key from the user's own ZCode config.

Credentials are scoped tightly: the host must be an exact match in the allowlist, HTTPS on port 443 only, `redirect: 'error'`, and the key is checked for CR/LF before it reaches a header. The key itself is never stored or logged — account identity is an HMAC.

Both JSON inputs (a user-edited config file and the remote quota response) are narrowed at runtime rather than asserted, and the request cancels an unread response body on the error path so it cannot trip the undici parser crash (orca#8695).

Co-authored-by: guanbear <guanbear@users.noreply.github.com>

* fix(mobile): paired clients re-derive a kept terminal after a cold restore (#23109)

* fix(mobile): paired clients re-derive a kept terminal after a cold restore

A renderer frame published before a cold-restored terminal's PTY registered
was fenced to an empty tab list and recorded as accepted, and the renderer
never resends unchanged content. When registerPty binds a surface the
accepted frame fenced out, re-merge that frame so the fence reads current
state.

* test(mobile): drive the live desktop window through the runtime's desktop seam

* test(mobile): the re-derive path never flushes the store synchronously

* test(mobile): a re-derived frame must not bring back a surface the host retired after accept

* fix(mobile): a re-derived frame changes membership only for the registering surface

The replay re-ran the whole accepted frame, so a surface the host retired
after accept (a phone close whose remote PTY is still exiting, or a closed
chat tab) came back. Every other surface now keeps the host's current
decision; the removal repair is extracted from the terminal retirement
helper so non-terminal tabs are removed the same way.

* test(mobile): a re-derived frame must not drop or disown a phone-created terminal the desktop has not published

* fix(mobile): a re-derived frame does not infer renderer retirements from its older frame

* revert(mobile): drop the replay of a fenced renderer frame

Reverts the production parts of a88e1eaa0a, f5b99d0003 and 5af1c6d97a: the kept
renderer frame, rederiveFencedRendererSurface and its registerPty call, and the
mergeRendererMobileSnapshot / removeMobileSessionSnapshotTabs extractions. The
fence will instead read the host's saved membership record. The test file stays
and is rewritten for that mechanism.

* fix(mobile): the paired-list fence admits a terminal the saved session still lists

After a cold restore the in-memory mobile snapshot and PTY table start empty,
so in a repo with host-authoritative terminal membership the fence dropped a
restored terminal whose renderer frame arrived before its PTY registered, and
paired clients never listed it. The fence now also admits a surface the host's
saved workspace session still lists (tab under the worktree, leaf in its
layout), and registerPty pushes the listing so pending-handle turns ready at
once. A restored pane whose PTY never returns is listed as pending-handle, as in
repos that are not host-authoritative.

* test(mobile): keep the desktop window stub's type assertion on its SAFETY line

* fix(mobile): coalesce the registration push for a listed restored terminal

registerPty pushed the paired list immediately on every registration that backs a
listed surface. The desktop's graph sync after a spawn already publishes the same
pending-handle to ready flip on the 50 ms coalescing window, so each restored pane
cost two pushes, and a restore of N panes cost N immediate full-list pushes per
client. The touch now rides the same coalescing window, which still covers a
registration no graph change follows.

The test's "unchanged" sync dropped the graph's tab, which is itself a change, and
its no-extra-push assertion ran before any coalesced push could fire; both are
fixed, and a restore of two panes is asserted to push once.

The fence comment no longer claims the new clause keeps pending leaves out of the
graph: once the surface is listed, its leaves pass the shared predicate through
that listing, as any listed surface's do.

* fix(mobile): read saved membership only from the worktree's own session partition

For a runtime-host workspace, emptying the owning partition re-routes session reads to a single
other partition that still lists the worktree. If that older copy lists a surface a retirement just
removed, a lagging renderer frame could re-admit it. The saved-membership check now reads only the
partition the worktree's host names, so it never trusts a fallback copy.

* test(mobile): pin the own saved partition for every workspace host kind

Also correct the immediate-emit comment: only an exit bypasses the window; a registration's ready flip coalesces.

* Fix terminal focus when Cmd+J wakes a workspace (#23546)

* fix: retain workspace terminal focus through wake restoration

* test: reset CPU throttling after wake focus assertion

* fix: require terminal textarea readiness before claiming focus

* test: configure React act environment for dialog regression

* fix(mobile): size a terminal's first subscribe from the document's reported cell box (#23080)

* fix(mobile): size a terminal's first subscribe from the document's reported cell box

#22960 sent phone dims on a terminal's first subscribe by opening a throwaway
empty terminal (init 80x24 ""), awaiting its ready and measuring, behind a
per-document first-subscribe mark whose lifetime was tied to web-ready. That
cost a second xterm/WebGL instance and ~150 ms per open, plus lifecycle state.

The document now measures the cell box without a terminal (xterm 6's
CharSizeService strategy, rounded as the renderer rounds it) for every
text-size preset and reports it with its viewport in web-ready; a table,
because the text scale only reaches the document after that notify. Each
init's ready reports the box xterm actually laid out, which replaces the
probe's entry. The controller answers fitDimensions/measureFitDimensions from
that table and the view's layout with no message; without a table it asks the
document as before.

The session seeds an unmeasured viewport synchronously in subscribeToTerminal,
so the first subscribe carries dims by construction. Deleted: the empty init,
its awaitReady gate, deferFirstSubscribeUntilViewportMeasured and the
subscribedDocuments mark. The fit pass is unchanged and still covers a
document that reports no cell box.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): read the reported cell box through in-narrowing, not Reflect.get

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): correct the probe's cell-box guess from the box xterm lays out

The web-ready probe is a guess: building the WebGL addon creates no context,
so a context that fails on load lands on the DOM renderer, whose width is not
snapped and depends on the column count. Before, a ready box that differed was
only logged; the first subscribe had carried the wrong column count, the host
echoed it, the fit pass saw the viewport equal to the host's dims, and the grid
stayed slightly shrunk. The store also kept the WebGL width after a context loss.

The document now reports the box xterm laid out whenever it changes (from
onRender, which covers a renderer swap and a DPR change that
onDimensionsChange does not fire for, and at ready). The store replaces the
guess; when that changes the current text size's entry, the view calls
onCellBoxChange with xterm's grid and the session re-fits, running the bounded
fit pass if the dims moved (one resubscribe). Equal boxes do nothing.

The RN layout box now survives a document reload; the document's own viewport
only stands in until the view reports a layout (on the page, web-ready arrives
first). The mismatch console.log is gone, and the probe's rounding names the
xterm version it copies.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): correct each cell-box guess at most once, so a DOM renderer cannot loop

On the DOM renderer the cell width is the rounded canvas width divided by the
column count, so every re-init at new cols reported a new box. Each one counted
as a correction, a floor over floats could flip the fit between two sizes, and
each flip landed converged, which reset the resubscribe budget: an unbounded
series of full-snapshot resubscribes.

Only the first laid-out box for a guessed text size may be a correction; later
reports still update the store, so fits stay truthful, but never resubscribe on
their own. The fit's floor gains a 1e-6 epsilon so floating-point error at an
exact boundary cannot flip a column or row. New document tests pin the render
report after a renderer swap and the report at ready for a paused renderer.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): make xterm the only terminal cell measurer

The document builds its real terminal before web-ready, at the app's
text scale, and reports the box xterm laid out; the first init reuses
that terminal. The page-side prediction, the per-scale guess table and
the once-per-document correction are gone. The app remembers the box
per text scale for its lifetime, so a later open at a known scale
subscribes with phone dims at once. A box that changes at the same grid
(renderer swap, pixel ratio) refits the open terminal in place.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): keep commands queued before the terminal WebView first loads

A subscribe sized from the stored cell box can queue init before the
native WebView reports its first load start, which cleared the queue
and left the terminal blank. Only a reload now drops queued commands.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): re-init a document that lacks the subscription's init, and fit one frame width

- Web-ready now says whether the document holds the terminal's latest init
  (a reload before the first ready drops a queued one); the session
  resubscribes any initialized terminal whose document lacks it.
- One grid fit, shared by the app and the document, fed the unrounded frame
  width React Native laid out; it keeps exact fits whole at fractional pixel
  ratios. The document's viewport-width fits are gone.
- The page builds every document at the scale the view mounted with, as the
  native WebView does.
- A new document's first cell box is compared against the grid the
  subscribe fitted from the stored box.
- The terminal built before ready stays hidden until its first init.
- The cell-box census matches glyph-measurement techniques, not names;
  the store's unused clear() is gone.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): build the terminal before ready only for the view shown at mount

A session mounts one terminal view per tab, and each built xterm and a
WebGL context before ready: 20 tabs made 20 contexts at load, past the
~16 a page (or Android's shared WebView renderer) holds, and native logged
32 context losses. Only the view shown when it mounts builds early now;
the rest build at their first init as before. Deferring the WebGL addon
instead would change the reported box: the DOM renderer lays out 7.8x15
where WebGL lays out 7.667x15 at the same font.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): write a WebView document's start values into its page, not an injected script

Android ran the pre-content injected script after the document's own in
1 of 22 documents on the emulator; that document started with no text
scale or shown flag and built a terminal it should not have. The values
now sit in the page ahead of the document script, one source object per
start pair so a render never reloads the WebView.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): the pre-ready terminal measures and reports while hidden

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): measure only the laid-out frame, and refit on a new grid, not a new width

- A measure needs both of the frame's dimensions from React Native; the
  document's viewport-height fallback is gone, and before the first layout
  the handle answers no fit without asking the document.
- A frame width change that still fits the PTY's grid from the stored box is
  a no-op, so sub-pixel layout jitter no longer re-measures. The width ref is
  written in that effect rather than during render (react-doctor).
- One "last grid" ref: the last reported grid, or the one a subscribe fitted
  from the stored box.
- The page render rig measures through the frame it laid out, as the session
  does, and lets the replay's fit settle before its resize-refit witness.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): let only the current terminal document's ready flush

A reload kept the WebView and its onMessage, so the old document's late
web-ready flushed the queue into the reloading view and the new document
got a second init. Each document now gets its own view (keyed on a
generation the controller owns), every notify carries the generation of
the view that received it, and a web-ready from a replaced document
flushes nothing and stamps nothing.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): drop every notify from a replaced terminal document

One rule at the receive boundary: a notify from any generation but the
current one is dropped, whatever its type, not only web-ready.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): make fitDimensions a pure question; name each generation counter

- fitDimensions no longer records the grid. A width change to a new grid
  asked it first, so the DOM renderer's report of that grid's box read as
  "same grid, new box" and refit again. Only the first-subscribe seed
  (seedFitDimensions) records the grid the document's first report is
  checked against.
- viewGeneration counts the views, readyGeneration counts web-readies.
- replaceDocument no longer resets the load flag; the load-start reset
  stays as the guard for a view that reloads itself.
- The name-based lifecycle census is replaced by a behavioural test.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): typecheck the handle mocks, drop the unused cell-box get

- The two handle mocks carry both fitDimensions and seedFitDimensions,
  and the fake-timer acts return nothing, so the three test files check
  under tsconfig.test.json again.
- terminalCellBoxes.get had no product caller; the store's tests assert
  through fit.
- The load-start comment says what the controller does now.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): hold the grid the document has, ignore a replaced view's load start, dispose a failed pre-ready terminal

- The document reports a new grid even with an unchanged box, so an
  in-place reflow on WebGL is held before a later renderer swap at that
  grid; the swap then refits. The app's apply paths do not hold the grid
  themselves: the DOM renderer's box follows cols, and a grid held on
  apply would read its own box as a renderer change and loop. One
  writer (holdGrid) holds the seeded or reported grid.
- A load start from a view a replacement unmounted is ignored, as its
  notifies already are.
- A terminal whose open throws before ready is disposed, not only
  unreferenced.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): ignore every native event from a replaced terminal view

One wrapper binds each WebView lifecycle event (load start, error, HTTP
error, render process gone, content process terminated) to the view's
generation, so a replaced view's late event cannot reset, replace or
put an error over the current document.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): a DOM seed refits once on its first report, not on the refit's own

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): subscribe a terminal only after its document is ready

The document still builds its terminal before ready and reports the cell
box xterm laid out in web-ready; the app now subscribes after that ready
and fits from that box, so nothing is sent to a document before it is
ready. Everything that made a pre-ready subscribe safe goes: the
app-lifetime box store, the seed fit, the per-document view generations
and their event filtering, the init tracker and the hasInit resubscribe.
The native view reloads in place again and web-ready keeps main's reload
rule. Boxes are kept per view; the grid a document last reported still
guards the in-place refit against the DOM renderer's cols-dependent box.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): hold one reported cell box and the grid the subscribe fitted

The controller keeps only the box the current document last reported,
not a per-text-size store: the document re-reports on a scale change.
The subscribe after ready fits from that box and holds the grid it
fitted, so the DOM renderer's first report at that grid (a new box)
refits once in place and converges; refit and apply paths hold nothing.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): fit only a ready box at the app's scale; forget a reloaded document's box and grid

A reload keeps the document's mount scale, so a ready after a text-size
change reports a box at the old scale; that box no longer sizes the first
subscribe, which then takes the no-box path. A readiness reset drops the
old document's box and held grid, so the new document's first DOM report
at the same grid does not refit.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): give the terminal document its frame at init, and fit text scale over it only

A subscribe sized from the ready box sends no measure, so the document
had no frame when the text size changed and reported the pre-refit row
pitch. The app's init now carries the frame it laid out, in the fields a
measure uses; the router takes it from either. The text-scale fit reads
only that frame, with no viewport fallback.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): say why a frameless text-scale change skips the resize

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): one cell box per terminal notify, not an array

web-ready and cell-metrics carry `cellBox: {fontScale, cellWidth,
cellHeight} | null`; the document's `laidOutCellBox` returns one or
null and the parser validates one object. The text-scale match moves
from web-ready into `handle.fitDimensions`, the one place a box is fitted.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): fit terminals in the app from the reported box; drop the measure round trip

The app already holds the box the document reported, so the refit and
the fit pass await the init's ready and call `handle.fitDimensions`
instead of posting `measure` and waiting on `measure-result`. The
document's measure, its retries, and the measure promise and timeout go.
The document still resizes locally on a text-size change, so every grid
the app sends (init, resize, reflow) carries the laid-out frame it was
fitted to. `holdSubscribedGrid` replaces `subscribeFitDimensions`, so
the only fits are `fitDimensionsFromCell` and `handle.fitDimensions`.
The render rig reads its fit from the ready box. The recorder adapter
mounts the new handle with the same recorded effects; the goldens it
mounts move on their adapterSha256 header only.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): keep the terminal frame in one ref, and notify a new width imperatively

The session held the frame in a height ref, a width ref, a width state
and the refit's own width ref. It now holds one `terminalFrameRef`
({width, height} | null until the first layout; a hidden 0x0 layout
keeps the last box). onLayout notifies a new width imperatively, as it
does height, and the refit's notify skips a width whose fit is the grid
the PTY has. `terminal-frame-width-refit.ts`, the width state and its
effect go. The subscribe's layout gate reads "no frame yet" directly.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): subscribe a held-back terminal on the frame's first layout only

`handleTerminalFrameLayout` ran on every onLayout; it now runs once, when
the frame first has a size. Later layouts only notify a new width.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): size the first subscribe inline in subscribeToTerminal

`sizeTerminalViewportFromCellBox` wrapped five lines in a 37-line
module; the subscribe now fits the ready box against the frame, holds
that grid and records the diagnostic itself. The helper's tests fold
into the subscription tests, which move to the subscription's name.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): drop the unreachable font-size guard on the reported cell box

xterm 6.1.0-beta.303 updates the render service's cell box in the same
task that sets `options.fontSize`: CharSizeService.measure fires
onCharSizeChange, and RenderService.handleCharSizeChanged runs the
renderer's `_updateDimensions` (DomRenderer.ts:359, WebglRenderer.ts:229).
`term.onRender` fires from RenderService._renderRows after the rows
are drawn (RenderService.ts:213, CoreBrowserTerminal.ts:538), and the
document writes its text scale and the font size in one task
(text-scaling.ts applyTextScale, terminal-init.ts init). So no report
can read a box between the font and the scale; the guard and its test
go. A new test pins the real order: no report when the font is set,
the new box at the new scale on the next render.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): one start seam, no source cache, the reported box as an object

- `useState` already pins each view's WebView source at mount (a new
  test re-renders at another text scale and gets the same object), so
  the module-level `webViewSources` Map goes.
- `initialTextScale` and `buildsTerminalBeforeReady` become one
  `start(): { textScale, shown }` seam.
- `reportedCellBox` holds the last reported box and grid, not a string key.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the RPC recordings to this branch and re-record

The terminal refit now fits in the app from the reported box and reads
one frame ref, so the recorder's terminal adapter mounts the new handle
(`awaitReady` + `fitDimensions`) and options (`terminalFrameRef`),
keeping its recorded effects. `baseline` is repinned to 21954dbd2f, the
last commit to touch a fenced path, and every golden is re-recorded.
Proof by class against HEAD: 787 header-only, 0 body moved, 0 added,
0 deleted. Header keys moved: `baseline` on all 787, and `adapterSha256`
on the 14 goldens `terminal-mount-adapters.ts` mounts (query-reply 3,
accessory-raw-send 4, takeover-report 4, viewport-refit 3). No recorded
traffic moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): hold the reported cell box and its grid in one ref

The controller kept the box in `cellBoxRef`, the grid in a string
`lastGridRef` and wrote it through `terminal-held-grid.ts`. One
`heldRef` now holds `{ cellBox, grid }`, as the document's own
`reportedCellBox` does: web-ready writes the box, every cell-metrics
report writes both, `holdSubscribedGrid` writes the grid, and a
readiness reset clears it. Same write points, so the one-refit bound
holds; the DOM-loop and refit-once tests pass unchanged.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the RPC recordings to the hold-rule commit and re-record

H (f00bebba48) touched a fenced path after the last repin, so
`baseline` moves to it and every golden is re-recorded. Against the
corpus before this branch's refreshes (21954dbd2f): 787 header-only,
0 body moved, 0 added, 0 deleted; `baseline` on all 787 and
`adapterSha256` on the 14 goldens `terminal-mount-adapters.ts` mounts.
Against the previous refresh: `baseline` only. No recorded traffic moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the RPC recordings to the main merge and re-record

The merge (b3b1b0def2) is the last commit to touch a fenced path, so
`baseline` moves to it and every golden is re-recorded. Against
97b5bb2b9a: 787 header-only, 0 body moved, 0 added, 0 deleted;
`baseline` on all 787, and `adapterSha256` on the 14
session.diff-review-actions goldens whose adapter #22951 edited. Against
origin/main: 787 header-only, 0 body moved/added/deleted; `baseline` on
all 787 and `adapterSha256` on this branch's 14 terminal goldens.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): the terminal document holds the grid and decides each refit

The document already kept the last reported box and grid; the app kept
a mirror of both to decide the refit. Now the document decides: its
`cell-box` notify carries `{ cellBox, refit }`, sent only when the box
changes, with `refit` a box that changed at a kept grid. web-ready
records the pre-ready terminal's box at its 80x24 grid, and the first
init that reuses that terminal holds the init's grid, so the DOM
renderer's first report refits once, as the subscribe's hold did. A
re-init no longer clears the record, so a new renderer at the same grid
still refits. The app keeps one `cellBoxRef` and `holdSubscribedGrid`,
`heldRef` and the grid on the notify go. The one-refit, DOM-loop and
renderer-swap tests move to the document with the same scenarios.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): one init options object, and a frame on every grid

`init` takes `{ cols, rows, data, preserveScroll, oscLinks, frame }`
instead of six positionals, and `init`, `resize` and `reflow` (handle
and messages) require `frame: TerminalFrame | null`. The refit's reflow
check reads `!dims` alone, and the controller's test file is named for
the `cell-box` notify it now covers.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): one notifyTerminalFrame for the frame's layout

The frame's onLayout made four calls and held the classification
itself. It now calls `notifyTerminalFrame({ width, height })`, and the
session's terminal-webview hook keeps the one frame ref, notifies the
height, subscribes the document held back for the first layout, and
notifies a later width change. `handleTerminalFrameLayout` is named for
what it does: `subscribeIntendedActiveTerminal`. The layout tests move
to that hook.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the RPC recordings to the round-8 head and re-record

85d421963c is the last commit to touch a fenced path. Against
a676c1b65a: 787 header-only, 0 body moved/added/deleted, `baseline`
only. Against origin/main: 787 header-only, 0 body moved/added/deleted;
`baseline` on all 787 and `adapterSha256` on this branch's 14 terminal
goldens. No recorded traffic moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): name the init option initialData, as the message does

The init option `data` becomes `initialData`, the message field's name,
so the controller passes it through unrenamed. The `preserveScroll` why
stays on the message type only, and the document test's title names the
three grids that carry the frame.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the RPC recordings to the round-9 head and re-record

486566c82b is the last commit to touch a fenced path. Against
3371c39715: 787 header-only, 0 body moved/added/deleted, `baseline`
only. Against origin/main: 787 header-only, 0 body moved/added/deleted;
`baseline` on all 787 and `adapterSha256` on this branch's 14 terminal
goldens. No recorded traffic moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* ci: rerun checks against main with #23560 landed

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(ai-vault): show ZCode CLI session history (#23513)

Surfaces ZCode CLI session history in AI Vault, so past ZCode sessions show up next to the other agents' instead of being invisible.

ZCode stores sessions in the same SQLite shape OpenCode uses, so this reuses the existing OpenCode lister and parser rather than adding a second scanner — the worker only varies the agent it stamps on each row. Discovery covers the native home and any WSL homes.

SQLite rows are narrowed at runtime rather than asserted: the statement API returns untyped column values, so the declared row shape is only a claim until something checks it, and a drifted schema or a database written by another tool reaches the same code.

Co-authored-by: guanbear <guanbear@users.noreply.github.com>

* test(mobile): repin the RPC recording corpus to main after #23080 (#23565)

#23080 squash-merged a corpus pinned to its branch commit 486566c82b, which
the squash left unreachable from main. Repin baseline to main's tip and
re-record; every golden moves only its baseline header.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* perf(ci): diff against the merge commit's first parent so PR checkouts can be shallow (#23562)

Every changed-path gate asked git for `--merge-base "$BASE_SHA" "$HEAD_SHA"`,
which needs the event payload's base SHA to be in the local graph. That is the
only reason two jobs cloned all 8127 refs' history. On a pull_request checkout
HEAD is already the merge commit, so its first parent is the base side and no
merge base has to be computed. config/scripts/git-pull-request-diff-base.mjs
resolved that for the two Node gates; the workflow's inline gates now use the
same helper through a small CLI rather than open-coding it.

code_paths gates all 22 jobs, so its checkout is charged to the start of every
one of them: measured 20.7s to 1.6s, keeping blob:none because its sparse tree
is ~7 files and leaves no blobs to refetch. Static analysis drops the filter
instead, since populating all 30,226 files makes blob:none force a second
promisor fetch: 23s to ~11s.

Verified on a real merge ref. At depth 50 the old and new forms produce
identical changed-file sets. At depth 2 the new form still works and the old one
fails with `fatal: bad object`, which is the failure a stale base would have
caused once the checkout stopped being complete.

Also drops the dead resolveBase + merge-base prelude in the changed-code gate,
whose result resolvePullRequestDiffBase already discarded on every PR.

* feat(mobile): the keyboard covers the page like a native screen, and the shell says its height (#23110)

The shell no longer shortens the WebView for the keyboard; it publishes the keyboard height like the safe-area insets, so native's keyboard lift, refit hold and dismiss key run on the page unchanged. Keyboard and inset arithmetic read the shell's OS through a host-os seam. One page-version floor (manifest pageVersion, shell floor 1) replaces per-feature accept negotiation; a page below the floor gets the existing update wall, a desktop with no bundle keeps native screens. iOS shell drops the form accessory bar and its own keyboard observers. Native session screens untouched.

* fix(agent-session): wait for in-flight session-store writes before teardown returns (#23545)

* fix(agent-session): stop lease renewal before the renewal's write lands

Clearing the renewal interval only cancelled the next tick. A tick already
past its guard still had a whole-file store transaction to commit, and the
store's transaction lock re-creates the store directory before it writes, so
that commit could land after host teardown had finished releasing everything
it touches.

`stop()` now resolves once the tick in flight has finished writing, and host
teardown's stop-lease-renewal phase waits for it. The three test harnesses
that model a host vanishing without a clean quit shared a copy of the same
incomplete shutdown; they now share one helper that waits.

The symptom was a CI flake: the refusal-oracle spec removes its temp
directory in `afterEach`, and a renewal landing mid-removal put the store
directory back, so the removal failed with ENOTEMPTY on the temp root.

* fix(agent-session): wait for the delivery loop's restart when abandoning a host

The abandon helper disposed the delivery loop and moved on. Disposing only
stops the loop's NEXT step: a step already past that check keeps going, and
the restart it runs for an accepted send reserves an owner, which is a store
commit. The store re-creates its own directory before every commit, so that
commit put the directory back under the temp-directory removal the test does
next, and the removal failed with ENOTEMPTY.

Quit already waits for exactly this work, in its drain-attaches phase — every
attach is registered with the task queue from enqueue. The helper now runs
the same drain, in quit's order, so it waits for both producers that reach
the store after the last awaited call returns.

Adds a regression test that holds the loop's restart inside its provider
acquisition and asserts abandoning does not return until it lands.

* chore: re-trigger PR checks

The push to 2ecc9c6e emitted no pull_request event, so the matrix never ran.

* fix(sidebar): clip worktree card content to its border (#23566)

* fix(runtime): answer terminal.subscribe at once for a pane the desktop already has mounted (#23512)

* fix(runtime): answer terminal.subscribe at once for a pane the desktop already has mounted

A mobile subscribe to a PTY with no headless model asked the renderer to
mount its tab and waited for a newer serializer settle. The renderer drops
mount requests for tabs it already has mounted, so a reattached daemon PTY
whose restored provider snapshot outranked the live renderer held the reply
for the full 3 s deadline. A live renderer screen now proves attachment and
is adopted directly; an unmounted pane still requests the mount and waits.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(runtime): treat any renderer answer, even a blank screen, as an attached pane

A fresh shell that has printed nothing has a registered serializer and an
empty screen; requiring non-empty data sent it back through the dropped
mount request and the 3 s wait. A blank screen skips the wait but does not
replace the chosen snapshot, so a parked pane cannot erase provider history.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(runtime): treat a mounted pane with unsettled output as attached

The stable renderer snapshot returned null both when no renderer answered
and when output advanced under every retry, so a desktop pane printing
continuously still took the dropped mount request and the 3 s wait. It now
returns a typed outcome (settled, moving, absent); moving skips the mount
and publishes the chosen snapshot, and late recovery still requires settled.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(runtime): adopt only a renderer-ordered screen when probing a mounted pane

The attachment probe read the terminal before knowing it would adopt, which
can reach the provider snapshot on the unmounted path; the read now follows
the decision. A seq-less renderer screen would replay every buffered chunk
on top of itself, so the probe keeps the chosen snapshot for it. The probe,
adopt and mount wait move into their own module to stay under the line cap.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(runtime): adopt a seq-less renderer screen when no output is pending

The seq gate only prevents a double replay of buffered output, so a settled
non-blank screen without a seq is safe when nothing is pending. That keeps
the better screen for a pane right after a deferred cold restore, before it
is renderer-ordered. The rule now applies after the mount wait as well.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(runtime): decide renderer attachment from the host's serializer flag

The mounted-pane probe re-derived attachment by serializing the renderer up to six
times, which cost ~7.5 s for a registered but unresponsive renderer on a busy PTY.
The host already holds that fact in the serializer readiness flag. The flag is never
cleared when a pane closes over a live PTY, so one null serializer answer falls back
to the mount wait: worst case is the old 3 s plus one 750 ms serialize.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(runtime): let the serializer's answer alone prove renderer attachment

serializeRendererTerminalBuffer already answers null when the host's serializer flag is
unset, so the separate flag accessor was redundant. The numeric-seq adopt test now
replays only the byte past the seam.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* Revert "perf(ci): shard the anti-slop audit across processes instead of one JS runtime (#23543)" (#23575)

This reverts commit fae0ae7a46.

* perf(ci): cache pnpm verification records on Linux (#23568)

* perf(ci): pilot pnpm verification record caching on Linux

* test(ci): review pnpm verification record in mobile cache audit

* fix(native-chat): the host writes chat failures for a person, with a typed fact beside them (#23116)

* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* fix(native-chat): name a chat write by its target, not the owner generation

A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.

Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.

Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.

* fix(native-chat): every journal append reaches the chats that are open

A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.

A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.

* test(native-chat): an epoch replacement reaches the open chat

* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

* perf(native-chat): a publish behind a delivered commit reads nothing

Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.

* test(native-chat): state why the teardown test's fake journal is safe to cast

* docs(native-chat): say mutation admission checks only the writer lease

* docs(native-chat): drop the send rebase from comments that still described it

* fix(native-chat): a message is accepted, then delivered

A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".

A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.

Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.

A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.

Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.

* fix(native-chat): settle queued messages only for the child that ended

A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.

A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.

The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.

* fix(native-chat): an adoption that fails to import keeps the conversation open

The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.

* perf(native-chat): the recovering open reads the journal once

Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.

* fix(native-chat): an attach that fails after indexing its child leaves no child behind

A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.

* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer

The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.

A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.

* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down

The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.

* fix(native-chat): a message rejected while its chat was closed reads as not sent

A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.

* test(orchestration): name why the readiness settlement fakes are cast

* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent

* docs(native-chat): drop the fence from the admission the send effects run behind

* docs(native-chat): give the fence move on release the reason that still holds

* docs(native-chat): stop citing a write fence check in launch and mailbox comments

Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.

* refactor(native-chat): the provider child is its own record

A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.

- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
  patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
  own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
  the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
  dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
  is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
  the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.

* fix(native-chat): the delivery loop alone settles a message its start or child failed

A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.

- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
  reads how it ended: a Stop continues; anything else writes one failure row and rejects every
  queued message with the same words, then stops. A child still starting whose start the adapter
  says did not land fails the same way. The exit, eviction and the settlement retry only settle
  the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
  nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
  failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
  closed, with or without a child, and a start the loop already has in flight is waited for so the
  child it produces is stopped rather than left behind.

* refactor(native-chat): a stopped child ends on the one reading of its stop

The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.

* feat(native-chat): the host says it accepts a send before any agent has it

The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.

* refactor(native-chat): an attach never opens a journal of its own

The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.

* fix(native-chat): a moved fence resends nothing on a host that accepts first

The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.

The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.

* refactor(native-chat): a child's end says whether the user or the host stopped it

The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.

* fix(native-chat): a chat whose only work is a queued message is not offered for resume

A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.

* test(native-chat): type the queued-message fixtures in the resume-offer tests

* fix(native-chat): a start that dies while a message waits on it is that message's failed start

Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.

* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now

* test(native-chat): pin what a failed start settles, and what a resume offer names

A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.

* test(native-chat): the failed-start pins fail on what the message became, not on a timeout

* test(orchestration): the preamble's host stub is typed, not cast

The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.

* feat(native-chat): a typed failure fact beside every failure sentence

Adds the shared vocabulary the host writes a failure with: a closed failure kind, a
provider diagnostic that says who it is for (a person, or a log), and a refusal cause
beside the refusal code. Status rows gain an optional failure fact and rejected
submissions an optional rejection fact; the dispatch row carries it, the reducer reads
it field by field, and the projection forwards it. Older rows and older readers are
untouched: every field is optional and the schemas stay open.

* fix(native-chat): durable failure rows and rejection reasons are written for a person

Every host writer that records a failure now writes a sentence for a person beside a
typed fact, instead of embedding a refusal's message, an exception or a composed exit
string. A provider's own words travel as a separate diagnostic from the places Orca
composes them - the Claude and Codex exit stderr (a log), Codex's JSON-RPC message,
Claude's compact_error and Codex's turn error (for a person) - and are never inferred
from a string afterwards. Not signed in and oversized history are typed at the
adapter that detects them.

Covers start and restart failures, the delivery loop, dispatch rejections (content,
queue-full, write failures, provider refusals), cancel and answer confirmation rows,
compaction, the rewind fallback, and not_delivered, which released clients printed
as it was. Two leaks close on the way: a settlement retry no longer writes Orca's
probe evidence into the exit row, and an attach or journal-sink failure is recorded
as Orca's fault rather than as the provider stopping. The legacy rejection markers
and the reasons on sends in doubt stay byte-identical.

* feat(native-chat): refusals name their cause, and a failed start is worded in one place

A refusal now carries an optional cause beside its code: one closed enum of the situations
a chat write can meet, set at every emitter a structured-chat write reaches. Returned
refusals build it with refuse(code, cause, message). Store and host paths that raised a
bare Error(code) now throw AgentSessionRefusalError, whose message is still the code and
which has no code property; the RPC error mapper handles it before any other passthrough,
keeps today's wire code and message byte-identical, and adds { refusal: { code, cause } }
to the error's data. The hold throws it, and restart-resume files the cause beside the
unchanged reason. The operation ledger stores the cause beside the code, so a replay names
the same situation as the first answer. The store fallback copy picks its words and cause
by situation, so a stale replay or a moved lease no longer reads as a latched owner.

Every failed start is worded by structuredAgentSessionStartFailure(cause, context), which
returns the row sentence and the typed fact together; the delivery loop, the exit
settlement and the dispatch that met a starting child all call it. Provider diagnostics
are capped at the lease record's 512 characters wherever a fact is built.

* refactor(native-chat): one reader of why a submission was rejected

classifyDispatchRejection(submission) returns { category, verdict, kind? }. It reads the
typed rejection fact when the row carries one this build can place, and the legacy
markers otherwise - all six, including not_delivered, which released clients printed as
it was. The verdict is null only for a withdrawal, a host restart and a closed chat;
write failures, a full queue and not-delivered stay failures. It replaces
dispatchRejectionReasonIsInternal and every string comparison against the markers: the
outbox reconcile, the rejection notice, and the send disposition, where a replayed
Stop-withdrawn send no longer surfaces as a failed send.

The journal reducer's echo-aliasing guard reads the narrow isWriteFailureSubmission,
which matches the typed kind or the legacy prefix in any dispatch state, so its behaviour
on legacy unknown rows is unchanged.

* test(native-chat): pin provider diagnostics where they are composed

The Claude exit status and stderr, the Codex stderr tail and Codex's JSON-RPC message are
each checked at the place Orca composes its own error around them, so the typed detail
is proven to come from the provider's value and never from Orca's wording.

* fix(native-chat): an attachment Orca refuses says which limit it broke

The content check's refusals (20 images, 5 MB per image, 20 MB in total, supported types) are
written for a person, but the rejection writer replaced them all with one generic sentence.
Each refusal now carries its own sentence, in MB rather than bytes, and the writer records it
beside kind attachmentInvalid. Only an attachment that could not be read keeps the generic
sentence.

* fix(native-chat): the chat tab table refuses with a typed cause

Showing a chat tab refused with bare Error('agent_session_conflict') and
Error('agent_session_identity_required'), the only chat-reachable refusals still thrown without a
cause (opening a chat from history can reach the second when the chat is removed mid-open). Both
now throw the typed refusal; wire code and message are unchanged.

* fix(native-chat): a compaction Codex refuses up front keeps Codex's words

When Codex refused thread/compact/start, the adapter passed on only Orca's wrapped error text and
dropped Codex's own message, so the chat's row read just "Compaction failed." The refusal now
carries Codex's message as the failure detail, as a compaction that fails later already did.

* fix(native-chat): an unreadable chat record no longer promises an update fixes it

The recordUnreadable copy said a newer Orca saved the chat, but the store marks a record
unreadable for damage and key mismatches too, where updating does nothing. The sentence now says
Orca can't read it and gives both next steps.

* fix(native-chat): a restart or a close leaves released clients a sentence, not a marker

host_restarted_before_delivery and provider_closed_before_delivery are not in the markers released
desktop and mobile builds hide, so they printed raw on every rejected message a restart or a chat
close left. Neither marker has shipped. New rows carry a sentence plus kind hostRestarted or
chatClosed, as not_delivered already did; the classifier still reads both markers, and the
verdict for both stays no-failure.

* fix(native-chat): log why an attachment could not be read

The rejected message now says only that the attachment couldn't be read, so the error that said
why (a missing file, a permission, or an unexpected throw) went nowhere. It is logged instead. The
content rejection moves beside the content check that owns its errors.

* refactor(native-chat): drop an unused thrown-refusal cause reader

It had no callers, and its comment claimed it looked through wrappers, which it did not.

* fix(native-chat): a compaction the provider never confirmed is recorded as unconfirmed, not failed

* fix(native-chat): an empty or non-user message is not recorded as a bad attachment

* fix(native-chat): an undelivered preamble's error ends in one period

* refactor(native-chat): a compaction ends as compacted, failed, or unconfirmed, never an unlabelled error

* refactor(native-chat): a failure's sentence is written only from its fact

A writer could choose its sentence and its kind separately, so five writers
put hand-written words beside a fact that said something else. One shared
constructor, agentSessionFailureWords(fact, { surface, agentName }), now
makes both, and the journal types refuse anything else: a status row, a
rejected message or a conversation command that carries a fact must carry
the sentence that constructor branded. Persisted rows keep their shape.

- The sentence table and the restart table move to src/shared, as the English
  default a client copy table can reuse.
- A rejection's legacy markers come from the same constructor. A write
  failure is now the bare `provider_write_failed` marker, which released
  clients already hide; its error goes to the log.
- An image Orca refuses carries which check it failed (and the limit) in the
  fact instead of a sentence; an empty message gets its own kind, and a
  non-user message is Orca's fault.
- The exit row, the interrupted compaction, the /clear failure and the rewind
  placeholder no longer carry their own words beside a fact: the exit row
  says what its fact says, and the placeholder carries no fact.

* fix(native-chat): a start that failed without an observed exit no longer blames the provider

Every untyped start error was recorded as "The provider stopped before it
finished starting.", so an Orca fault, a failed spawn or a close that ended
a start was blamed on the provider. Only an exit the adapter observed says
so now: the Claude adapter marks the error it saw the child exit with, and
anything else is a new `startFailed` kind, "<Agent> couldn't start.", keeping
the provider's diagnostic when the error carried one. A child gone with no
end observed, and a /clear whose new conversation was refused, read the
same way.

* fix(native-chat): a view never restarts a chat whose last start failed

A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.

* fix(native-chat): a chat a terminal agent still holds says to quit that agent

The merged base words a restart a terminal agent's claim refused with the
refusal's own message, which names the process. That message is Orca's text
and never reaches a durable row here, so the row read only "<Agent>
couldn't restart." and lost the one step that frees the chat. The restart
sentence now derives it from the refusal's cause: a claimConflicted refusal
adds "This chat is still open in a terminal agent. Quit that agent to
continue the chat here." The live refusal still names the process.

The restart-resume ledger test now sets up a claim the base still refuses:
a terminal owner that is proven running.

* refactor(native-chat): keep the changed files inside their lint limits

The legacy-marker lookup is a table, not a non-exhaustive switch; the
preamble tests read the error without a cast; and the Codex history refusal
goes through a named constructor so its file stays under the line limit.

* test(native-chat): start the child the loop waits on with an attach, not a second view

A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.

* fix(native-chat): settle a gone generation's turn wherever a conversation opens

A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.

* test(native-chat): prove the next child's start settles the turn an earlier child left

The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.

* test(native-chat): count a failed start's rows by row, not by text

Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.

* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget

* refactor(native-chat): refusals carry details keyed by their code

A refusal named its situation with one flat cause list shared by every
code, so nothing stopped a site pairing a code with a situation that code
never means, and the loose fields released clients read (fence, revision,
resolution, verdict, rewind reason) were written by hand at each emitter.

A refusal is now one variant per code with optional details: a reason that
code lists plus that code's own facts. refuse(code, details, message)
rejects a reason the code does not list at compile time, and it is the one
place the loose top-level fields are copied from details, so released
clients read exactly what they read before. A site that cannot name its
situation uses refuseUnclassified, which carries facts but no reason, the
same as an older host; there is no catch-all reason.

Thrown refusals put { refusal: { code, details } } in the RPC error's data
(wire code and message unchanged). The operation ledger, the restart-resume
record and a restart failure's embedded refusal keep details beside the
code and read them back against it; a row an unreleased build wrote with a
cause parses and reads as naming none.

* test(native-chat): pin the open's and the send's start and row counts, however the view binds

Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.

* test(native-chat): a restart-resume failure keeps its refusal details

The recovery capsule reads a failure's details back against the refusal
code in its reason: facts the code does not list and a reason another code
owns are dropped, and a record an unreleased build wrote with a cause
still parses, naming nothing.

* test(native-chat): import the failure words once in the provider child test

The merge left two imports of the same modules.

* fix(native-chat): settle a gone generation's turn at every open but an acquisition's

The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.

* test(native-chat): hold the create's start open until the views bind

The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.

* fix(native-chat): a start the provider refused or Orca broke no longer says the provider stopped

A failed start whose cleanup proved the child gone was typed as
`providerStartFailed` whatever failed it: Codex refusing to resume a thread,
a timeout, or Orca's own store fault. The chat then read "The provider
stopped before it finished starting.", which was untrue, and the provider's
own words were dropped. A refused restart also inferred the same from an
`exited` verdict, which only says nothing runs now.

Only an exit the adapter observed names that situation now; anything else
is refused with no reason, keeping its verdict, and the chat reads
"<Agent> couldn't restart.". The provider's words travel host-side from
where the acquisition failed into the start-failure fact's detail, never
onto the refusal, and the sentence does not quote them. What failed is
logged once where the start failed.

* fix(native-chat): a refused /clear start keeps the situation it named

The replacement start that /clear makes built its own start-failure fact,
so a refusal that named its situation, such as not being signed in or a
history too large to restore, was recorded as a bare "couldn't start". It
now takes its fact from the same start-failure function as every other
start, as a new session that failed to start.

* fix(native-chat): the journal schema and comments describe a refusal's details, not its cause

The persisted failure fact's schema still described `refusal.cause`, which
this branch replaced with `details`. It now describes `details` as an
optional open object; a row an earlier build wrote with a `cause` still
parses. A comment and three test descriptions that still named the cause
now name the details.

* fix(native-chat): a Claude child that exits while being acquired still reads as the provider stopping

Now that only an exit the adapter observed says the provider stopped, the
exit the Claude adapter saw during acquisition has to be marked where it is
seen, as the exit after acquisition already is. Without the mark, a Claude
CLI that exited at spawn read "Claude couldn't restart." instead of "The
provider stopped before it finished starting."

* fix(native-chat): word a failed chat start's refusal as its start failure

A chat whose agent failed to start answered the create with the raw error: the launch strip read "Chat could not be started. claude stream-json exited (code 1): claude: not signed in", and the ledger replayed the same text. The first answer and the replay now carry the sentence the chat's start-failure row reads as ("The provider stopped before it finished starting.", "Claude couldn't start.", the not-signed-in and history-too-large sentences), and the raw error goes to the log. A store refusal's code and the unproven-exit marker are unchanged.

* fix(native-chat): show Claude's API retries as one sentence row

While Claude retried a refused request (a 429, say), the chat gained one red row per attempt reading "rate_limit", with the raw retry frame behind Details. Each retry run now writes one warning row that later attempts revise in place: "Claude is rate-limited and retrying." for a rate limit (error `rate_limit` or status 429), and "Claude hit a temporary problem and is retrying." otherwise. The row carries a `providerRetrying` fact with the provider's error type and status, and the frame as a log detail capped at 512 characters.

* fix(native-chat): tell the user to run /clear again when its new conversation can't start

When /clear's replacement conversation failed to start, the result told the user to "send your message again", which would go into the old conversation. The failure words now take the command the start was for, so a failed /clear reads "Codex is not signed in for the selected account. Sign in, then run /clear again.", "Codex couldn't start. Run /clear again." or "The provider stopped before it finished starting. Run /clear again." A message send keeps its wording.

* test(native-chat): expect a failed Claude create to be refused in a sentence

The runtime suites asserted the CLI's stderr reached the create refusal; it now goes to the log and the refusal reads as the chat's start failure.

* fix(native-chat): name the agent that stopped starting and say how to retry a failed start

A start the provider ended now reads "Claude stopped before it finished starting." (or "The agent ..." when the chat's agent is unknown) instead of naming "the provider". A start or restart that failed with a chat left to retry now ends in "Send your message to try again.", or "Run /clear again." for /clear; released clients print only this sentence. A message rejected at dispatch because its child died while starting names the same agent as the start's row.

* refactor(native-chat): drop the composer's second error formatter

After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.

* test(native-chat): pin the reason on a message rejected while its chat was closed

The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.

* fix(native-chat): answer a create refused before spawn in the words its replay reads

A create that failed before any process started threw Orca's own error text as its first answer,
while its replay from the operation ledger read the generic start sentence. The two refusals a
person can act on, a launch whose Anthropic sign-in variables override the managed Claude account
and a Claude account switch in progress, are now typed where they are thrown and worded by the
shared failure constructor, so the first answer and the replay say the same thing. Every other
pre-spawn failure reads the generic start sentence, with its own text in the log. The first answer
keeps its wire code; a message that is itself a code is unchanged.

* fix(native-chat): say how to retry after an agent stopped before it finished starting

"<Agent> stopped before it finished starting." gave no next step outside /clear, unlike every other
failed start. It now ends "Send your message to try again.", the same step a start or restart that
could not run gives; after /clear it still says "Run /clear again."

* fix(native-chat): say how to reach a Claude chat when a WSL Claude account blocks it

A Claude chat that restarts while a Claude account is added in WSL and no Windows Claude account is selected was refused before spawn with Orca's own text as its first answer, and the generic start sentence on replay. The refusal is now typed where the account gate throws, and both answers read the same sentence: choose or add a Windows Claude account in Claude Accounts settings, then send the message again. Account settings that cannot be read name no situation and keep the generic sentence.

* fix(native-chat): say the reason a host names for a refused chat write

A refused Stop, answer, setting, goal, command or queued message now reads the refusal's reason as
well as its code. A code stands for several situations, so the code alone could only say what did
not happen; with the reason, the notice says why and, where the person has a step to take, what it
is: "The agent is still responding. The command didn't run. Wait for the agent to finish
responding, or stop it." The phone uses the same words.

The notice table keys on code, then reason, then the kind of write. Every reason of every code has
an entry, so a reason the host adds does not compile until it has words; a reason whose honest
words are its code's keeps the code's row. A refusal with no reason, or one this build does not
know, reads exactly as before, which is what an older host gets. A start that failed reuses the
failure row's own sentence rather than a second one.

A queued message keeps the reason, the rewind reason and the owner's verdict with its saved
failure, and a rejected one keeps the host's typed fact without its provider detail. Nothing that
moves with the owner or comes from the provider is saved; entries saved before this load as they
were. The saved failure moves to its own module beside the words chosen from it.

* fix(native-chat): retry a refused send under a new id only once its agent is proven gone

A send refused because Orca could not tell who owns the chat keeps its operation id, since the
first attempt may still land. When the refusal also says the agent process has exited, nothing can
run that attempt, so the message moves to its Retry row under a new id instead of holding the queue.

The owner's verdict is read as a floor: a saved `exited` is final, and any other saved verdict never
changes the id on its own. Only a verdict re-derived from the current lease can raise it to
`exited`, and nothing lowers a saved `exited`. Today no host sends a verdict on a send refusal, so
nothing a person sees changes; the rule is in place for the saved verdict a reload reads back.

* fix(native-chat): say a /clear that never finished did not finish, instead of that it cleared the chat

A send into a chat whose /clear started but never committed was refused as if the conversation had been cleared: "This conversation has been cleared. Your message was not sent. Open the current conversation to continue." The clear never finished, so its new conversation may not exist and there is nothing to open. That refusal now has its own reason and reads "The last /clear didn't finish. Your message was not sent. Start a new chat to continue.", which is the only way on today. Its message for released clients says the same: "The last /clear didn't finish. Start a new chat to continue." Only a committed /clear still says the conversation was cleared.

* fix(native-chat): a send the provider never received after a restart has no verdict

Restart reconciliation rejects a crash-stranded send the provider's history
proves it never received. Nobody failed that send, but the verdict table
treated it as a failure. Each rejection kind now has its verdict in one
exhaustive table, so a new kind does not compile until its verdict is chosen;
no verdict for a withdrawal, a host restart, a chat close, or this lost send.

A rejection whose kind this build cannot place, such as one a newer host
added, now reads as undelivered with no verdict instead of falling back to the
reason beside it: all it proves is that the message did not happen. The host
keeps such a fact's kind when it reads the row back, rather than dropping it
and letting the reason decide.

Only kinds that can be why a message was not sent may reject one, by type:
compaction, cancel/answer confirmation and provider-retry kinds stay on status
rows. The one dispatch-row builder takes its input from the type that makes a
rejected row carry its fact.

* fix(native-chat): a send refused after its agent exited keeps its place in the queue

When Orca cannot tell who owns a chat but the refusal says the agent process
has exited, the next attempt may use a new id, since nothing can run the old
one. It no longer marks the message as rejected: nothing recorded it, so it
still holds the head of the queue, and later messages wait behind it instead
of being sent ahead of it.

* fix(native-chat): stop reading a provider's words from an error that contains itself

A cleanup that aggregates errors restarted the depth count for each one, so an
aggregate error that contains itself recursed until the host ran out of stack.
One depth bound now covers both the cause chain and the aggregated errors.

* fix(native-chat): say a chat whose history can't be read can't continue, and to start a new one

A read of a chat's history is refused with `agent_session_journal_unreadable` only when the chat's journal file is corrupt or not a database, which no retry can change. The notice table had no words for a read at all, so a pane had nothing to show but the raw code. Reading a chat's history is now its own request kind, `read-history`, and that refusal on it reads "This chat's history couldn't be read, so it can't continue here. Start a new chat to continue.", whether the host names the reason or raises the bare code. A write refused under the same code keeps its words, because its cause is any failed open, which can clear. Any other read refusal says only "This chat's history couldn't be loaded." `agentSessionReadHistoryRefusalParts(code, details)` gives a pane those words from a read error.

The notice sentences move to their own module so the table stays within its size limit.

* fix(native-chat): decide what every way a child ends means for queued messages in one table

What a child's end means for the messages queued behind it was an if-chain: a user's Stop was checked in one place, a host stop in another, and every other cause, including one added later, fell through to "the provider exited". It is now one table over every end cause, so a new cause does not compile until someone says whether it fails what is queued and how. A user's Stop still fails nothing, a host stop is still Orca's fault, and an exit, a failed attach or an eviction still carry the failure the end recorded. Nothing a person sees changes.

* test(native-chat): a close that stops the child and then fails rejects what is queued by how the child ended

When a chat closes, stops its agent, and then fails a later step, the chat stays open with its messages still queued. The delivery loop then rejects them by the way the child ended: an eviction during startup reads as a failed start, otherwise as the provider having stopped, and either counts as a failure. The cases are rows over the end cause, so another way a chat closes is one more row.

* test(native-chat): type the refusal a persisted-schema test admits

* fix(native-chat): say whether a chat's history is damaged or just couldn't open

A write refused because the chat's journal would not open named one reason, `journalUnreadable`, for every failed open, and a read of the history took that same reason as final. So the words depended on what was asked, not on what happened: a busy or permission-denied open could tell a person to start a new chat, and a damaged one could read as something that clears.

The host now decides at the refusal which it was. `journalCorrupt` is set only when SQLite itself reports the journal damaged or not a database (SQLITE_CORRUPT or SQLITE_NOTADB, extended codes included, read from the driver's result code and never from message text, through any `cause` chain). Every other failed open is `journalUnavailable`. A corrupt history reads "This chat's history couldn't be read, so it can't continue here. Start a new chat to continue.", with "Your message was not sent." before the step on a send. One that couldn't open reads "Orca couldn't open this chat's history right now. Try again.", likewise on a send. A host that names no reason gets "Orca couldn't read this chat's saved history.", which promises neither, because damage can't be proven from the code alone.

The refusal's message, which released clients print for a send, is now that person sentence instead of the open error's own text; the error is logged instead. `journalUnreadable` is replaced outright: no released build wrote it, and a stored one reads as a refusal with no reason.

* docs(native-chat): say why a history that couldn't open names its retry step

* fix(native-chat): a failed start's row keeps the words its rejected messages carry

When the delivery loop settles a failed start before the child's exit is published, it writes the
start's error row and rejects every queued message with the adapter's startup answer. The exit
settlement then rewrote the same row from the exit event, so the row could say one thing while the
rejected messages, which are terminal, said another. The exit now leaves a start's row it finds
already written.

* fix(native-chat): a Claude start Orca itself failed no longer says Claude stopped

Every error that ended a Claude session was marked as an exit the adapter
observed, including a start Orca failed while the CLI was still running: a
saved option whose restore lost its answer, an init frame naming another
session, or a journal write fault. Those read "Claude stopped before it
finished starting." although Claude never stopped on its own. The mark now
stays where the child's exit is seen (the connection's exit callback), so
those starts read "Claude couldn't start." with any diagnostic beside it,
and a real exit before the start lands still says Claude stopped.

* fix(native-chat): name the agent that stopped, and blame Orca for its own closes

A chat that lost its agent mid-response said "The provider stopped…", and a
started Claude session that Orca itself closed after a journal fault said the
same, as if Claude had exited on its own.

The exit row and the rejected-message reason now name the chat's agent ("Claude
stopped while this response was in progress…", "Codex stopped before this
message was sent."), or "The agent" when the name is unknown; the stale-state
settlement now passes the agent name too. After a Claude start has landed, the
ended event reports providerExited only when the child's own exit was observed;
any other close is Orca's fault and reads as one.

* test(native-chat): pin the sidebar verdict to the rejection classifier for every kind and legacy marker

* fix(native-chat): blame Orca, not Codex, when Orca closes the Codex child

A Codex chat that Orca itself closed (a journal sink that could not take a
frame, or a forced close) said "Codex stopped while this response was in
progress", as if Codex had exited on its own.

Orca's own close path now reports hostFault. providerExited is left to the
app-server connection's exit callback, which the connection withholds while Orca
is closing the child, so it only ever reports the child's own exit.

* fix(native-chat): a chat whose history is damaged reads "Unable to load this chat."

* fix(native-chat): route the conversation-outlives-agent writers through the typed refusals

Three writers that arrived with the merge wrote refusals the old way:

- An operation that starts the agent itself, such as a goal change, turned any error the start
  threw into a refusal whose message was Orca's own error text, which released clients print.
  It now logs the error and says only that the agent couldn't restart.
- An option picked while the chat is at rest, for a key the provider would not accept, is refused
  with the rejected-option reason, like the same pick on a running agent.
- A restart continuation whose agent was refused a start filed the rejected message's sentence as
  the failure's reason. It files the refusal's code with its details again, which is what the
  restart-failure guidance keys on.

* fix(native-chat): a start Orca stopped because it never finished reads as that

The idle sweep now stops an agent whose start never finished and rejects the messages waiting on
it. The chat read "Orca ran into a problem, so this didn't go through. Try again." for that,
because every host stop was worded as Orca's own fault. It now reads "Codex never finished
starting, so Orca stopped it." in the chat's row and on each rejected message, carried as its own
failure kind so newer clients can tell it apart. The message counts as failed, like any start that
did not land.

* test(native-chat): pin the merged close and host-stop rows to their typed facts

The merge left two expectations on the old words: the close tests looked for the marker a close
used to write, and the host-stop test for the host-fault sentence. A close now writes "The chat
closed before this message was sent." with its fact, and a host stop the hostStopped sentence the
constructor gives, whatever reason the stop carried. Also folds the conversation command's two
imports from send preparation into one.

* fix(native-chat): a read of a chat this host cannot open says why

Reads now reach a chat through one accessor, which refused a missing record and a provider this
host does not run as a bare code with nothing beside it. Revealing the same chat already names
those reasons, so a client could tell "this chat no longer exists" and "update Orca" apart there
but not on the history or subscribe read that follows. The accessor now throws the same typed
refusals. The wire code and message are unchanged; the reason rides only in the error's data,
which released clients ignore.

* test(native-chat): a Claude retrying past the idle window keeps its conversation open

Every api_retry frame publishes the journal, and that publish is the activity the idle sweep reads, so a retry run revised into one row still renews the clock on each attempt.

---------

Co-authored-by: Claude <noreply@anthropic.com>

* perf(ci): stop duplicating shared Linux download caches (#23578)

* perf(ci): pilot pnpm verification record caching on Linux

* test(ci): review pnpm verification record in mobile cache audit

* perf(ci): share Electron downloads and clean closed PR caches

* test(ci): retain cache ordering checks for restore-only consumers

* perf(ci): limit archive sharing rollout to primed Linux hosts

* perf(ci): run static analysis on the free ARM runner (#23576)

Measured 128s against 172s on ubuntu-latest, with every compute-bound step
faster: type-aware 24s to 15s, anti-slop 28s to 19s, localization extraction
67s to 46s, and the orcad terminal smoke 39s to 14s. Checkout and the install
were unchanged at 12s and 16s.

The toolchain resolves on arm64: both lint engines ship linux-arm64 bindings
(@oxlint/binding-linux-arm64-gnu, @oxlint-tsgolint/linux-arm64), and
build-orcad-bun.mjs derives its target from process.arch. The orcad smoke
booting and round-tripping a real PTY is the evidence that node-pty compiled
and that Bun, the bundled ripgrep and @parcel/watcher all resolved.

The runner is free for public repositories, the same one the typecheck job
already uses.

* fix(orchestration): reland process-incarnation reap for stale worker terminal handles (#23583)

* fix(orchestration): reland process-incarnation reap for stale worker terminal handles

Relands the orchestration part of #18790, which #22601 reverted in full
because that squash commit also bundled an unannounced agent. Only the
worker terminal-reap fix returns; no agent catalog changes.

When a worker's saved terminal handle stops resolving while its process is
still running, worker-show/read, worker-stop and worker-release now re-mint
a live handle from the recorded process incarnation (exact pty id +
incarnation id, same host scope) and act on it, instead of reporting the
terminal missing and leaving the process running on the host.

Fixes STA-8493

* test(orchestration): spy on the public incarnation methods instead of casting the runtime

* feat(agents): add Freebuff launch and sidebar status support (#23567)

Add Freebuff launch support and execution-host status reporting for the sidebar, including running, question, blocked, and settled states. Validate against captured CLI transcripts and real rendered sidebar evidence.

Cross-referenced community implementations #17065, #20839, and the Freebuff portion of #18790. Preserve their agent/catalog/mobile/documentation coverage and add canonical status publication and regression tests.

Co-authored-by: Harkaran Brar <18134082+harkaranbrar7@users.noreply.github.com>
Co-authored-by: Prarambha369 <98906077+Prarambha369@users.noreply.github.com>
Co-authored-by: Lesley Murfin <260182349+LesleyMurfin@users.noreply.github.com>

* perf(ci): move six more jobs to the free ARM runner (#23594)

* perf(ci): move six more jobs to the free ARM runner

Follows the static-analysis move, which measured 172s to 128s. Each of these was
checked for an x86 requirement rather than assumed portable.

pr.yml:
  cross-version-wire        source-only, tagged checkout plus in-process vitest
  managed_hook_node18       Node 18 publishes linux-arm64; the per-platform
                            runtime files are read as data, so host arch is moot
  codex_index_heal_contract @openai/codex ships @openai/codex-linux-arm64
  shell_contracts           fish 4.x is published for noble/arm64, zsh is in the
                            arm64 archive, so the fatal fish-4 gate still holds

mobile.yml:
  verify                    209s of its 298s is Vitest; no Android SDK, emulator,
                            gradle, Hermes or Watchman, no docker, no artifacts.
                            Gemfile.lock lists the generic `ruby` platform, so
                            frozen bundler resolves without an aarch64-linux entry
  recording-pin             pure Node plus git; the golden comparison masks
                            `platform`

Left on x86 deliberately:
  package                   builds --x64 targets, its docker gates are
                            --platform linux/amd64, and it is where the glibc
                            floor check runs. node-pty's .symver pin is
                            arch-specific, so flipping would validate the arm64
                            pin and stop validating the shipped x64 one
  orcad_browser             Google ships no Linux arm64 Chrome
  mobile_web_app            same Chrome wall; its render check fails closed
  git_compatibility         its cache key carries runner.arch and the warmer is
                            x86, so flipping alone means a cold `make git` every
                            run
  relay_integration         no technical blocker, but x86 relay coverage is a
                            documented placement and the reusable workflow has no
                            per-job runner input
  e2e and the ssh lanes     the ssh jobs would silently retarget the tested
                            remote from linux-x64 to linux-arm64

* perf(ci): move xterm_patch_sync to ARM too

The patch check rebuilds 4 packages x 2 builds and byte-compares against the
checked-in bundles. Ran it on darwin-arm64: exit 0, in sync at 1362743 bytes,
with the full fetch-and-rebuild path exercised rather than a short-circuit. The
bundles were generated on Linux x64 and reproduce byte-for-byte on a different
arch and a different OS, so the output is host-independent.

* Fix status bar on smaller screen (#23587)

* fix(status-bar): adapt layout to smaller screens with responsive density

Replace fixed breakpoints with a responsive density system that measures the status bar's content and picks the roomiest density level that fits. At progressively tighter widths, the bar sheds usage mini bars, segment labels, and finally collapses calm usage chips into a "+N" overflow indicator — keeping every feature one click away in the Usage popover. Usage-first collapse order ensures urgent providers stay visible longer.

* fix(status-bar): use overflow-clip-margin to preserve focus rings

Adds a 3px overflow-clip-margin to status bar containers so focus rings
and badge dots aren't clipped when content overflows on smaller screens.

* feat: add first-class Qoder CLI support (#23581)

feat: add first-class Qoder CLI support

Integrate Qoder launch, identity, canonical hook status, trust and resume.
Verify with captured Qoder 1.1.64 transcripts and hidden Electron sidebar checks.

Builds on and cross-reviews #7502, #8611, #9655, #12910, #13311 and #15291.

Co-authored-by: dalveytech-vincent <vincent@dalveytech.com>
Co-authored-by: Eridanus117 <45489268+Eridanus117@users.noreply.github.com>
Co-authored-by: xingqingzzp-gif <xingqingzzp-gif@users.noreply.github.com>
Co-authored-by: jyang2004 <jyang2004@users.noreply.github.com>
Co-authored-by: yunqian <yunqian@alibaba-inc.com>
Co-authored-by: huzhening.hzn <huzhening.hzn@alibaba-inc.com>

* ci: skip unrelated installs and share xterm build dependencies (#23607)

* ci: pilot shared xterm installed dependencies

* ci: bound xterm cache production to verified main entries

* ci: benchmark xterm reuse on the production ARM runner

* ci: avoid installing Orca dependencies for standalone xterm checks

* ci: use Node-only setup in the production xterm job

* ci: remove completed xterm benchmark workflow

* ci: use faster gzip for temporary Linux test packages (#23609)

* ci: benchmark faster Linux package compression

* ci: pass compression options through typed builder configuration

* ci: retain original configuration for benchmark baseline

* test: preserve release settings in CI compression configuration

* ci: normalize generated changelog dates in package comparison

* ci: remove completed Linux compression benchmark

* Update README downloads badge

* fix(deps): take Electron 43.7.5 so detached webviews stop blanking browser tabs (#23586)

Electron 43.7.0 threw 'Invalid guestInstanceId' from <webview>'s
disconnectedCallback for a loaded guest (electron/electron#53989), so a
webview React removed and re-inserted kept a dead guest id: the tab went
blank, reload did nothing, and the destroyed listener never fired.
43.7.4 (electron/electron#54097) returns early when the guest is gone.

Raises the runtime floor test to 43.7.4 so a downgrade cannot re-ship it.

Fixes STA-8757

* docs(readme): remove the TestFlight link (#23660)

* docs: update GitHub star history chart (#23661)

* fix(claude): withdraw a follow-up queued behind a stopped turn instead of dropping it silently (#23553)

* fix(claude): withdraw a follow-up queued behind a stopped turn instead of dropping it silently

Orca's SessionStart hook proves most Claude starts before the turn's
system/init, the only frame that advertises interrupt_cancel_queued_v1.
Capabilities were read once from the start proof, so they stayed empty:
Stop never asked Claude to cancel its queue, and the one-at-a-time
fallback withdrew the follow-up without settling it. The send stayed
pending and the chat read Working until the child exited.

Capabilities are now one derived value, taken from whichever report
names any and never cleared by one that names none. The fallback settles
each send the CLI confirms it withdrew.

* test(claude): pin that only a confirmed cancel_async_message counts as withdrawn

* test(native-chat): await the async journal snapshot in the queued-stop test

* test(native-chat): wait for the queued-stop states instead of sleeping

A fixed 20 ms sleep does not cover the host's asynchronous journal writes on a loaded runner.

* fix(codex): a message whose turn was stopped before Codex took it is withdrawn, not stuck (#23618)

* fix(native-chat): land a late settlement from a streamed turn's end after that turn's rows

A settlement that says a streamed turn ended waits for the session's event
sink to drain before writing its dispatch row. The journal reducer still
refuses to overwrite an accepted or rejected send.

* fix(codex): settle a send from the end of the turn Codex answered it into

The turn/start answer names the turn that holds a send. The send's echo
entry now keeps that binding, in memory only. If the bound turn is
interrupted without echoing the send, the send is withdrawn: Codex clears a
turn's pending input on interrupt, so the model never saw it. If the turn
fails first, the send is rejected in Codex's words. A completed turn settles
nothing, since Codex records pending input when it finishes and the echo is
still due. An answer read after its turn already ended is settled by that
end. The echo is still the acceptance and carries the item key.

* test(codex): a send settles from the end of the turn Codex answered it into

The fake Codex keeps 0.157's turn bookkeeping, and can deliver the turn/start
answer after turn/started or after turn/completed. The tests cover:
- a Stop before any echo withdraws the send, and the working rule reads idle;
- a steered follow-up is withdrawn when the turn is interrupted;
- a failed turn rejects the send in Codex's words;
- a completed turn leaves the send to its echo;
- a normal echo and a late echo;
- two steered sends in one turn;
- an answer read after the turn ended;
- a timed-out answer;
- child-thread turns;
- how a binding dies.

* refactor(native-chat): drop the stream flush before a late turn-end settlement

Nothing reads the order of a dispatch row against the turn's terminal row:
the reducer keeps a settled send terminal and the working state is derived
from both. The echo acceptance on the same path never waited either, and the
wait could drop the settlement on a failed sink barrier.

* fix(codex): settle a failed turn's sends at its end, not at its error

Codex keeps a failed turn's pending input and records it after the error
frame, before turn/completed. Settling at the error rejected a steered
follow-up the model had in fact received, so a Retry would send it twice.

* docs(codex): say a completed turn echoes what it took before it ends

Codex records a completed turn's pending input before `turn/completed`, so
a bound send that turn never echoed is left for recovery, not awaiting an
echo. The comments and one test title said the echo was still due.

* test(codex): settle a send whose answer is read after its turn failed or completed

A failed turn that ended before the answer rejects the send in Codex's words,
once; a completed one leaves it admitted and still armed for its echo.

* refactor(codex): read a failed turn's reason with the typed thread-fact reader

* Show a resting Claude chat's effort instead of a blank picker (#23106)

* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* fix(native-chat): name a chat write by its target, not the owner generation

A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.

Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.

Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.

* fix(native-chat): every journal append reaches the chats that are open

A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.

A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.

* test(native-chat): an epoch replacement reaches the open chat

* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

* perf(native-chat): a publish behind a delivered commit reads nothing

Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.

* test(native-chat): state why the teardown test's fake journal is safe to cast

* docs(native-chat): say mutation admission checks only the writer lease

* docs(native-chat): drop the send rebase from comments that still described it

* fix(native-chat): a message is accepted, then delivered

A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".

A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.

Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.

A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.

Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.

* fix(native-chat): settle queued messages only for the child that ended

A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.

A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.

The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.

* fix(native-chat): an adoption that fails to import keeps the conversation open

The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.

* perf(native-chat): the recovering open reads the journal once

Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.

* fix(native-chat): an attach that fails after indexing its child leaves no child behind

A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.

* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer

The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.

A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.

* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down

The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.

* fix(native-chat): a message rejected while its chat was closed reads as not sent

A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.

* test(orchestration): name why the readiness settlement fakes are cast

* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent

* docs(native-chat): drop the fence from the admission the send effects run behind

* docs(native-chat): give the fence move on release the reason that still holds

* docs(native-chat): stop citing a write fence check in launch and mailbox comments

Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.

* refactor(native-chat): the provider child is its own record

A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.

- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
  patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
  own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
  the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
  dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
  is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
  the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.

* fix(native-chat): the delivery loop alone settles a message its start or child failed

A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.

- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
  reads how it ended: a Stop continues; anything else writes one failure row and rejects every
  queued message with the same words, then stops. A child still starting whose start the adapter
  says did not land fails the same way. The exit, eviction and the settlement retry only settle
  the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
  nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
  failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
  closed, with or without a child, and a start the loop already has in flight is waited for so the
  child it produces is stopped rather than left behind.

* refactor(native-chat): a stopped child ends on the one reading of its stop

The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.

* feat(native-chat): the host says it accepts a send before any agent has it

The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.

* refactor(native-chat): an attach never opens a journal of its own

The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.

* fix(native-chat): a moved fence resends nothing on a host that accepts first

The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.

The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.

* refactor(native-chat): a child's end says whether the user or the host stopped it

The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.

* fix(native-chat): a chat whose only work is a queued message is not offered for resume

A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.

* fix(native-chat): the conversation outlives its agent

Opening a chat no longer starts its agent. A conversation is reached through one host
accessor that opens its journal at rest, and a send is what starts the agent, through
the delivery loop. One idle sweep, every five minutes, stops an agent that has been
quiet for thirty minutes and owes no work, then drops an open journal handle that is
only a cache. Its record, tab, status row and readers stay.

- hold and release are no-ops; hold still builds the host for shipped mobile builds.
- The holders, the holds, the release clock and the exit respawn are deleted.
- Options, the model list, the goal and the context meter answer at rest; a model pick
  at rest is recorded as intent for the next start.
- Compact, rewind, clear and goal changes start the agent first. A send does too when
  a rewind is still in doubt after the conversation opens.
- Orchestration routes mail and group addresses on ownership (the record plus the chat
  tab), not on whether the process runs. An open dispatch keeps its worker running.
- The restart continuation is a send; Resume all holds each slot until the message is
  handed over or rejected.
- A read error never replaces a loaded transcript, and shows the host's own words.

* test(native-chat): type the queued-message fixtures in the resume-offer tests

* fix(native-chat): a start that dies while a message waits on it is that message's failed start

Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.

* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now

* test(native-chat): pin what a failed start settles, and what a resume offer names

A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.

* test(native-chat): the failed-start pins fail on what the message became, not on a timeout

* fix(native-chat): a restart offer ends when the chat's agent starts again

The offer used to end only when the chat's newest user message changed,
because opening a chat started its agent and that start could not be told
apart from real activity. Opening a chat starts nothing now, so the host
reads the fact it already publishes: a chat's status row goes from not
host-owned to host-owned exactly when its agent is started. At that edge the
offer and any failure record for the chat are withdrawn, unless the start is
a resume action's own (its continuation is the oldest undelivered message).

A continuation and a message racing to be first are decided at acceptance:
the continuation is refused, quietly and with nothing filed, when any other
message was accepted since the restart. A failed continuation start leaves
the offer retryable, and each resume action sends its own message id.

Deleted: the newest-user-message comparison, its journal reader, the
continuation filter, and the failure ledger's own "answered by the chat"
check. The marker still carries its message id for one release, so the
previous build can read it.

* fix(runtime): end a transcript stream when its client unsubscribes

Desktop: the IPC subscription controller was dropped as soon as the streaming
handler returned, which for most streams is right after it binds. A later
runtime:unsubscribe then found nothing to abort, so the host kept the subscriber
and derived and sent every publish to a channel no one listened to. The controller
now lives until the renderer unsubscribes, resubscribes the same id, or goes away.

Mobile: disposing an agentSession.subscribe stream now sends agentSession.unsubscribe
with the stream's frame id, so the host ends that subscriber and leaves a sibling
stream on the same socket running. The direct path now passes the frame id the relay
path already passed.

* test(orchestration): the preamble's host stub is typed, not cast

The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.

* fix(native-chat): one fact ends a restart offer: the chat moved on since the restart

The offer is live while no other message has been accepted in the chat since the
restart and its agent has not proved a start since. The offer list, the resume's
reservation check and the continuation's acceptance check all read that one fact,
so a message whose start then failed withdraws the offer too, and a stale click
finds nothing to act on.

The fact is read off the conversation's open handle, which the restart closed, so
it is retired durably whenever it may have changed: a message accepted, a start
proven. A close and reopen within the same run therefore cannot bring the offer
back. A continuation rejected before it reached the agent does not count, so a
retry after a failed start still runs.

Deleted: the quit-time gate on withdrawal, which changed nothing because the
withdrawal and the quit's own offer write share one queue; the per-action
"withdrawn" flag and the separate acceptance check it paired with.

* test(native-chat): an older build reads the restart offer this build records

The offer lives in a file the previous release reads after a downgrade. Pin that
against the pinned release's own capsule, and run the lane when the marker or the
capsule changes.

* fix(native-chat): read a restart offer against where the journal stood when it was taken

"Since the restart" was read off the conversation's open handle, which the idle
sweep closes: after a reopen, a message the user had already sent looked older
than the handle and the withdrawn offer came back.

The offer now records the journal position (epoch and sequence) at the moment
it is taken, and a message accepted after that position, or a journal on another
epoch, means the chat moved on. That is derived from the journal, so it holds
across any number of closes and reopens. An older build's offer has no position;
only a start withdraws it. Because the message half is now durable, the offer is
no longer rewritten in the recovery file on every accepted message; a proven
start still writes it, since only the host that saw the start knows of it.

* test(native-chat): wait for the listing's retire write before reading the recovery file

* fix(native-chat): keep the terminal-backed chat's read error over its local echoes

Messages winning over a read error is right for the structured chat, whose read retries and whose
messages came from the transcript. The terminal-backed view assembles its list from local echoes
too (a launch prompt, a pending send), so a failed read there showed only those bubbles and no
error. Only the structured pane now keeps messages over an error.

* fix(native-chat): a start retries the exit settlement a failed journal write left owed

An agent exit whose journal settlement write failed releases the lease latched until a retry lands.
Reopening the chat used to be that retry; with reveal now only opening the journal, nothing retried
it before the next app launch, and every send was refused. The start the send needs now runs the
retry first, where the attach would.

* perf(native-chat): answer the owner check without opening the chat

Worktree activation calls agentSession.handoffStatus for every chat tab in the worktree, and the
answer comes from the session record alone. Reaching it through the accessor opened each resting
chat's journal (a full read, the crash-boundary write and a restored status publish), then kept it
open for the idle window. It now checks the record and the adapter's support, as before this series,
and opens nothing.

* fix(native-chat): a read waiting on the session lock opens nothing once quit began

The accessor checked for quit before queueing the open, so a read queued behind a session task ran
its open after teardown had begun and indexed a journal no teardown step would close. The check now
runs at the open itself.

* fix(native-chat): read a failed resume's chat before calling it retryable

Whether a failed resume is retryable is the offer's own rule: the chat has not moved on since the
restart, read from its journal. The failure list read it only for a chat already open, so once the
idle sweep closed a chat the user had moved on in, its failure showed Retry again, and the click did
nothing. The list now opens the failed chats first, as the offer list does.

* test(native-chat): type the provider event sink the settlement test reaches for

* fix(native-chat): say the structured read keeps trying only where it does

The structured pane's "Orca keeps trying to load it" line never showed: the view state filled in an
untranslated fallback whenever the read error had no text, and the empty state prefers any message.
The view state now leaves the message out, so the structured pane shows that line and the
terminal-backed pane its own translated one. Mobile's structured lane does not resubscribe after an
error frame, so it no longer makes the claim.

* test(native-chat): await the send's settlement instead of polling for the start

The at-rest send tests polled for the provider start with vi.waitFor's one-second default, which a
loaded machine outran. They now await the host's own settlement of the message.

* fix(native-chat): a restart offer resumes any time after the quit, and knows its own continuations

The continuation's message id was dated by the quit, and the ledger refuses a new id dated more than
a day back, so Resume or Retry a day after quitting was always refused (on main too). It is now
dated by the resume action.

Telling a rejected continuation from the user's own message read the operation ledger, whose rows
expire after about a day; after that a failed resume stopped being retryable. The offer now
records the continuation each action sends on its own capsule entry, bounded to the newest 16, so
the ids end with the offer. The ledger read is deleted.

* fix(orchestration): route no mail to a structured worker its orchestration released

A structured worker is routed on ownership, and a resting worker's lease is released, so ownership
held while its chat tab stayed listed. A worker the coordinator abandoned and then released, found
at rest by the release, therefore still took peer mail and @worktree: broadcasts, and each one
restarted its agent. Routing now also reads the orchestration's own resource row: once it is
released, direct mail, group addressing and worker-show's addressable answer drop the worker, as
they would a terminal worker whose terminal closed. The chat tab stays, and nothing new is stored.

* fix(native-chat): a failed retry names the user's prompt, not Orca's continuation

A resume's continuation is written to the chat before its start, so after a failed attempt the chat's
newest user message is that rejected continuation. A second failure then showed Orca's own restart
text as the chat's prompt. A retry now keeps the prompt its first failure named.

* fix(orchestration): read the released row optionally, as the authority does

worker-show's observation called the row lookup directly, which a runtime double without it threw on
and failed the structured tab-retirement release.

* fix(native-chat): the status bar drops a restart offer the chat moved on from

The renderer re-read the host's restart offer only when a failed chat showed activity, so after a
message withdrew a pending offer the host answered no chats while the status bar kept counting one,
and clicking it opened nothing. The same watch now covers pending offers: a status change in an
offered chat asks the host again, once.

* fix(native-chat): a resting Claude chat shows the effort its next start runs

A running Claude child reports the effort it applies. With no effort pick of
its own, that is what the CLI runs for the model when none is sent, so the
catalog write-through records it as the model's defaultEffort. The store keeps
a known default through a listing that names none, and the resting options
read answers the pick, else that default, as a live child does. Nothing is
saved as the chat's pick.

* test(native-chat): a roster of idle or finished children does not keep an agent awake

The sweep reads owed background work through the shared child-work liveness that upstream's
release clock adopted; a child that went idle or finished is not work the agent still owes.

* fix(orchestration): a task dispatched into a resting structured worker keeps it running

The sweep's open-dispatch check read only the worker-start dispatch that owns the worker's terminal
resource, so a task later dispatched to the same worker (orchestration dispatch --to, which writes a
dispatch with no worker row) did not count: after thirty quiet minutes the worker was stopped while
that task was open, and its coordinator read exited. Any unsettled dispatch addressed to the worker's
process incarnation now counts, derived from the existing rows.

* docs(native-chat): comments stop describing the hold this PR removed

Eight comments still justified orderings and teardown choices by a viewer or dispatch hold that
pinned the provider child. Nothing holds any more; the orderings stand for the binding's redrive
subscription and parked mail, and a chat's agent runs from a send until the idle sweep rests it.
Comment-only.

* fix(native-chat): a restart offer keeps the start its own continuation made

Whose start ended an offer was decided at read time, from whether the offer's continuation was
still the queued message. Once the provider refused that continuation, the child it had started
read as someone else's start, so the offer ended and its failure showed no Retry. The delivery
loop now records which queued message a start is for on the in-memory child, and the child's end
carries it; the offer counts a start as its own when that message is one of its continuations.

* fix(native-chat): an agent gets a full idle window after its owed work ends

The sweep measured quiet only from the last journal row, so once a subagent, command, monitor or
dispatch that had outlived the window ended, the agent was stopped at the next tick. A child can
read done before the lead's wake-up turn writes anything, and stopping in that gap loses the
wake-up. The sweep now counts owed work it observes as activity, which gives the agent the full
window afterwards, as the release clock it replaced did.

* test(claude): the options-read fixture runs a live child

The fixture marked its conversation running with a hasProviderChild field the
session type does not have, so the read took the at-rest path and refused a
session with no record. It now carries a child, which is what the read checks.

* test(native-chat): host tests reach its collaborators through a typed seam

The rest-test rig and three test files read the host's private members with
Reflect.get and cast the result. The host now exposes one test-only accessor,
collaboratorsForTests(), and the subscribers class a subscriberCountForTests()
beside its existing retainedActivityCountForTests(), so the tests are checked
against the real types and the casts are gone.

* refactor(orchestration): one owner answers a structured worker's custody

Routing, group addressing, worker-show and the idle sweep each composed their own reading of
whether orchestration still holds a structured worker, so each new obligation or retirement state
had to be added to every reader. structured-worker-custody now derives both answers from the
worker-terminal list state coordinators see in worker-list: addressable is owned and not released,
and owed work is an active custody or an unsettled task dispatched to the same incarnation. The
owner's state is read through the remote dispatch attachment too, as the terminal transfer lookup
already does. Behaviour is unchanged; a settled worker awaiting its coordinator still rests.

* refactor(orchestration): owed work is an open dispatch on the worker's incarnation

A supervised worker's own dispatch context stays open exactly while the worker is active, so the
separate active-custody branch only repeated it. Owed work is now one fact, which also states the
policy that a worker awaiting its coordinator's decision may rest, and both custody decisions are
written once at the top of the module.

* fix(native-chat): a restart offer knows its continuations by a tag in their id

The offer recorded each continuation id in a list on its capsule entry, capped at 16, and a running
action's id in memory. Both could disagree with the journal: past the cap an old rejected
continuation read as the chat moving on, and a crash during a retry restored the failure's older
entry, which lacked the retry's id. Each continuation id now carries a tag derived from the offer
(its teardown and chat), then the action's own part, so any continuation of this offer, queued or
rejected, is recognised from the journal row and the marker alone. The persisted list, its cap and
the in-memory action map are deleted; the agent-start withdrawal keeps an offer whose own
continuation the start was for, read against the stored marker.

* test(runtime): the legacy-worker reveal test judges its stale snapshot inside the wait

The tui-idle probe reads through readTerminal, which now awaits the structured
worker check before the PTY read, so the probe's snapshot request starts a
microtask later. vi.waitFor missed it on its first check and polled again at
50 ms, the same moment the wait's own 50 ms timeout fired. The stale snapshot
then resolved after the wait had already timed out, so the test passed without
judging it, and the rejection landed before any handler was attached. Vitest
reported that as an unhandled error and failed the shard.

Polling every 1 ms sees the request within a few ms, so the snapshot is judged
while the wait is still pending.

* fix(native-chat): a view never restarts a chat whose last start failed

A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.

* fix(native-chat): the idle sweep reads owed work every tick

Owed work counted as activity, but the sweep read it only once the idle window had elapsed, so it
refreshed the clock at most once a window. Work that ended just before the next read left the
agent to be stopped at that read, moments after the work ended, which is the gap the refresh was
meant to cover. The sweep now reads owed work on every tick for a started agent, so the window
always runs from the last tick that saw work owed.

* fix(native-chat): a continuation handed to the agent stays sent

The offer read its own continuation as not reaching the agent while its dispatch was pending, which
also covered one already handed over and still unanswered. When the wait for that answer ended first,
the failure it filed read as retryable, and a retry sent a second continuation to an agent that may
have acted on the first. Only a continuation still queued, or rejected, is now read as unsent.

* test(native-chat): start the child the loop waits on with an attach, not a second view

A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.

* fix(native-chat): settle a gone generation's turn wherever a conversation opens

A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.

* test(native-chat): prove the next child's start settles the turn an earlier child left

The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.

* test(native-chat): count a failed start's rows by row, not by text

Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.

* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget

* test(native-chat): the interrupted create's own retry continues again

The merge of main's lease-latch fix replaced that test's retry of the interrupted create, under its
own operation id, with a fresh start whose result nothing read. That fresh start passes with the
released-reservation continuation deleted, so the case the fix exists for went untested. The retry
and its assertion are main's again.

* docs(native-chat): three comments that still had views starting agents

A start with nothing queued now comes from a command, goal change or rewind; an interrupted compaction
left alone would refuse every send, so no agent would ever start to finish it; and a current host
raises the unattached read refusal only once quit began, with the attach window belonging to an older
host.

* test(native-chat): pin the open's and the send's start and row counts, however the view binds

Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.

* test(native-chat): a reader's open settles the turn a failed exit settlement left running

An exit whose settlement write failed leaves its turn running in the open journal. PR 1's open now
settles it, and this pins the two reads that reach it here: a reader reopening a chat the idle
sweep closed, and a read that opens the chat before the restart restore reaches it.

* test(native-chat): the view-start test's starting window outlasts two subscriptions on a loaded runner

A subscription reads the conversation before it returns, so under load the two views took longer
than the create child's 300 ms start, which then exited before the test checked that it had not.
The child now takes a second to fail.

* fix(native-chat): settle a gone generation's turn at every open but an acquisition's

The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.

* test(native-chat): hold the create's start open until the views bind

The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.

* test(native-chat): a read that reaches a crashed chat before the startup reconcile settles its turn

On desktop the chat on screen at relaunch reads before startup reconciles the leases, while the
dead process's lease still reads live. The open settles the turn it left running anyway, and the
restore that follows finds it settled.

* refactor(native-chat): drop the composer's second error formatter

After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.

* test(native-chat): pin the reason on a message rejected while its chat was closed

The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.

* docs(native-chat): drop the removed dispatch hold from six comments

A worker's session no longer takes a dispatch hold, and no release clock
rests a chat by visibility; the agent-launch comments, the abandon test,
the teardown test and the refusal census still said so.

* test(native-chat): rest the owner-status chat through the idle sweep, not a hold

The activation-gate test from #22808 put its chat at rest by holding and
releasing it, and passed the release-clock grace. This branch deleted both,
so the case threw before it reached its assertions. It now moves the host's
clock past the idle window and lets the sweep stop the agent and close the
conversation, then asserts the same owner answer and activation gate.

* fix(native-chat): show the structured pane's retrying line when a read fails

The read transport always hands the pane the host's words, so the error
state's "Orca keeps trying to load it" line, which showed only when there
were none, was never seen: the pane showed the host's text twice, as its
subtitle and on the status line under it. The structured pane now always
says its read keeps retrying, and the host's text stays on the status line.
The terminal-backed chat is unchanged.

* fix(native-chat): keep a resting Codex chat's unsaved effort blank, as its live child shows it

The resting options read fell back to the catalog's default effort for every
provider. A live Codex child answers only the effort its thread reported, which
is none when Codex's config names none, so a resting Codex chat showed Medium
and then went blank once it started.

* test(native-chat): read the resting Claude chat from the record its start persisted

The at-rest test built the record by hand, so nothing proved the model id a start
persists is the listed row the catalog learns the default under.

---------

Co-authored-by: Claude <noreply@anthropic.com>

* fix(release): stop the release policy from deleting pipeline-cut releases (#23669)

* fix(release): stop the release policy from deleting pipeline-cut releases

The policy judged a release by who created the release object. Cut Release
reuses an existing draft, so a CI-built v1.4.216 whose draft a person had
created was deleted (tag included) when its notes were edited, and Latest
fell back to v1.4.214 because v1.4.215 was also published by a person.

- Authorize a desktop release when its annotated tag was created by the
  release pipeline and points at its `release: vX` commit, not only by author.
- Only delete on `published`; an edit never deletes a release or tag.
- Pick Latest from the highest authorized stable using the same check.
- Move the policy into config/scripts/release-policy.mjs with tests.

* fix(release): load the policy module from the tagged commit

Release events run the workflow file from the tag's commit, so checking out
the default branch could pair an old workflow with a newer module.

* fix(claude): write only the hook events and statusLine the user's Claude accepts (#23614)

* refactor(claude): name the Claude version module after the hook events it gates

Pure move of claude-session-end-hook-capability.ts and its tests; the next
commit turns its one-event SessionEnd floor into a per-event version table.

* fix(claude): write only the hook events the resolved Claude knows

Claude 1.0.81 through 2.1.100 validate settings.json `hooks` against a
closed event enum and discard the whole file on one unknown name, so
Orca's install made Claude <= 2.1.77 silently ignore the user's env,
permissions and hooks. Each managed event now carries the first Claude
release that knows it (pinned to per-release enums read from the npm
packages), and install, status and the SSH/WSL relay installer write
only the events the resolved Claude accepts. An unresolved version gets
the set every tabled Claude knows; a downgrade removes only Orca's own
entry for an event the older Claude would reject.

* refactor(claude): move the managed Claude hook events into their own module

hook-settings.ts is at its line limit; the event list and its version gate
move out whole so the next change has room.

* fix(claude): an unresolved Claude version never removes Orca's hook entries

A failed or timed-out version probe is no evidence of an old Claude, so it
must not strip StopFailure, PermissionRequest and the other newer events a
version-aware install wrote. With the version unknown, install adds only
the set every tabled Claude knows and leaves every other entry exactly as
it is; only a known version that lacks an event retires Orca's entry.

* fix(claude): gate the core hook events on the Claude release that added them

Claude validates hooks against a closed event list from 1.0.23, not 1.0.81.
The table treated SessionStart, UserPromptSubmit, Stop, SubagentStop,
PreToolUse and PostToolUse as known by every resolved version, so a Claude
from 1.0.23 to 1.0.61 was still sent names it rejects, and it dropped the
whole settings file. Pin each to its first release from the packed enums and
keep the unresolved-version set as its own policy.

* fix(claude): write Orca's statusLine only for a Claude that knows it

Claude 1.0.49 through 1.0.66 also reject any unknown top-level settings
key, and statusLine joined that schema only in 1.0.64. Orca wrote its
statusLine for every Claude, so 1.0.49 to 1.0.63 still dropped the whole
settings file even with the event gate. Gate statusLine on 1.0.64, pinned
by the packed schemas; a known older Claude has Orca's own statusLine
removed along with the opt-out marker, so an upgrade re-adds it. An
unresolved version is now assumed to be 1.0.64, which knows the same core
events and keeps the statusLine install it had before.

* test(claude): a user statusLine opt-out survives a downgrade and upgrade

Retiring Orca's statusLine for a Claude older than 1.0.64 forgets the
install marker only when Orca's own statusLine was removed. Pin that, so
a user who deleted Orca's statusLine is not opted back in by an upgrade.

* test(claude): check the whole written settings file against each strict schema

Claude 1.0.49 through 1.0.66 discard the whole settings file over any
top-level key their schema lacks. The fixture recorded only whether each
release knew statusLine, so a new top-level key Orca wrote would pass every
test. Record each release's top-level keys instead (statusLine is derived
from them), add the hook enums for every packed release in that window,
and check that a real install and a downgrade write only keys and events
each strict release accepts.

* fix(native-chat): a failed Codex turn keeps its failure and "Worked for" (#23514)

* fix(native-chat): a Codex turn's first terminal settlement is final

Codex follows a turn-ending `error` with a failed `turn/completed` for the
same turn. The error settled the turn and dropped its start and attributed
send; the completion then re-settled it from nothing, so the record lost its
start time and flipped from completed/failure to interrupted/failure. A
failed turn lost its "Worked for" and read as interrupted.

Both ends now go through one settlement that refuses a turn already in the
recent-turns window, so every later end (error then completion, a duplicate
completion) adds nothing and overwrites nothing.

* test(native-chat): the settles-once fixture is an ordinary failed turn

The fixture was labelled and worded as a failed compaction; this change
covers the ordinary-turn path, so its frames now read as one.

* test(native-chat): a failed Codex turn keeps its record when its two ends coalesce in the queue

* fix(terminal): stop guessing that apps died and wiping their keyboard modes (#23584)

* fix(terminal): stop guessing that apps died and wiping their keyboard modes

The renderer wiped xterm's Kitty keyboard flags on every Ctrl+C, every live
reattach, and every Windows agent turn end, though the app usually survives.
xterm then encoded keys in legacy form while the pane mirror Orca's shortcut
policy reads still held the negotiated flags, so Cmd+C, Shift+Enter,
Option/Alt and IME commits disagreed with each other and with the app.

- Delete the Ctrl+C wipe, the ConPTY agent-idle wipe, and the mirror reset on
  every PTY exit (it also ran on unverified host-loss exits).
- Live reattach profiles no longer reset Kitty; every replay epilogue instead
  re-asserts the mirror's flags (pop-all, then the host-proven set; a bare pop
  while unproven), so a revealed xterm gets the live app's flags back.
- Route every renderer-originated mode write through one scanning writer so
  xterm and the mirror always parse the same bytes: confirmed-shell reset,
  hibernate, cold restore, and a full process-boundary ground at fresh spawn.
- Read Kitty flags as 0 where the protocol is withheld (ConPTY), since xterm
  ignores CSI u there but the mirror still scans it.
- The dashboard popout restores snapshot flags as bytes so its xterm agrees.

* style(terminal): separate the kitty restore builder from the pen reset

* fix(terminal): let the kitty mirror own the withheld-protocol rule

Review follow-ups for the stop-guessing change:

- The mirror takes a `kittyKeyboard` option from the xterm's advertisement
  and ignores CSI u when withheld, as xterm does, replacing a per-reader
  helper that any new reader could skip. Daemon/headless users keep the
  default.
- Replay epilogues are writers (`writeReplayEpilogue`,
  `writeReattachReplayReset`) that take the sync or async xterm writer, so
  nothing that looks like a builder mutates the mirror.
- An abandoned hidden restore re-asserts the mirror's kitty flags after the
  byte-gap reset: its discarded chunks were already scanned.
- Tests pin a non-zero host restore (epilogue ends `=31u`, mirror 31), a
  withheld pane staying at 0, and the restart-in-place ground landing after
  the mirror reset; the epilogue test helper is now an exact builder.

* refactor(terminal): one epilogue writer and one scanned ground per boundary

- Reattach callers write `chooseReattachReplayReset(...)` through
  `writeReplayEpilogue`, dropping the second writer from the session.
- Fresh spawn and cold restore rely on their scanned ground alone: it leaves
  the mirror known at 0 with a proven baseline, so the extra reset() was dead.

* feat(terminal): add Reset Terminal that grounds input modes on the host and the pane

An app that crashes where the host sees no command end (no shell hooks, after
exec bash, inside a manual ssh) keeps its Kitty, mouse, paste and focus modes.
Reset Terminal grounds them in every model that can re-arm them: the pane's
xterm and Kitty mirror, the daemon emulator, records and lifecycle scanner, the
relay replay buffer, and main's headless model, so park/reveal, reattach and
reload stay grounded.

Each holder grounds its own model at the request, like Clear Screen. A zero-raw
in-stream span would be dropped by SSH credit delivery and is ambiguous under
snapshot-seq dedup. The host request is a sibling of clear, not a field on it,
because an older host would ignore the field and clear scrollback instead.

* test(e2e): Reset Terminal grounds an unhooked crash's modes through park and reveal

* test(runtime): pin Reset Terminal's headless snapshot to a proven 0 kitty flag

* fix(terminal): ground main's provider mode tracker in arrival order on Reset Terminal

onPtyData scans live bytes into the tracker on arrival, so a ground queued on
the headless write chain could land after a newer ?1049h and leave the tracker
grounded while the emulator stays on the alternate screen. Also pass the
serializer registry's clear/reset hooks as an options object.

* fix(terminal): ground in-flight provider snapshot captures on Reset Terminal

onPtyData feeds both the settled provider mode tracker and any in-flight
capture's live tracker; the reset grounded only the first, so a capture whose
request predated the reset could republish the pre-reset alternate screen.
Share one scanProviderModeTrackers for both. Cover the pty:resetInputModes IPC
handler and the runtime controller's host-window request.

---------

Co-authored-by: GuanBear <123guan@gmail.com>
Co-authored-by: guanbear <guanbear@users.noreply.github.com>
Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Harkaran Brar <18134082+harkaranbrar7@users.noreply.github.com>
Co-authored-by: Prarambha369 <98906077+Prarambha369@users.noreply.github.com>
Co-authored-by: Lesley Murfin <260182349+LesleyMurfin@users.noreply.github.com>
Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com>
Co-authored-by: dalveytech-vincent <vincent@dalveytech.com>
Co-authored-by: Eridanus117 <45489268+Eridanus117@users.noreply.github.com>
Co-authored-by: xingqingzzp-gif <xingqingzzp-gif@users.noreply.github.com>
Co-authored-by: jyang2004 <jyang2004@users.noreply.github.com>
Co-authored-by: yunqian <yunqian@alibaba-inc.com>
Co-authored-by: huzhening.hzn <huzhening.hzn@alibaba-inc.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-09-28 15:49:28 -04:00
Jinwoo Hong 3eac4d93d3 fix(terminal): stop guessing that apps died and wiping their keyboard modes (#23584)
* fix(terminal): stop guessing that apps died and wiping their keyboard modes

The renderer wiped xterm's Kitty keyboard flags on every Ctrl+C, every live
reattach, and every Windows agent turn end, though the app usually survives.
xterm then encoded keys in legacy form while the pane mirror Orca's shortcut
policy reads still held the negotiated flags, so Cmd+C, Shift+Enter,
Option/Alt and IME commits disagreed with each other and with the app.

- Delete the Ctrl+C wipe, the ConPTY agent-idle wipe, and the mirror reset on
  every PTY exit (it also ran on unverified host-loss exits).
- Live reattach profiles no longer reset Kitty; every replay epilogue instead
  re-asserts the mirror's flags (pop-all, then the host-proven set; a bare pop
  while unproven), so a revealed xterm gets the live app's flags back.
- Route every renderer-originated mode write through one scanning writer so
  xterm and the mirror always parse the same bytes: confirmed-shell reset,
  hibernate, cold restore, and a full process-boundary ground at fresh spawn.
- Read Kitty flags as 0 where the protocol is withheld (ConPTY), since xterm
  ignores CSI u there but the mirror still scans it.
- The dashboard popout restores snapshot flags as bytes so its xterm agrees.

* style(terminal): separate the kitty restore builder from the pen reset

* fix(terminal): let the kitty mirror own the withheld-protocol rule

Review follow-ups for the stop-guessing change:

- The mirror takes a `kittyKeyboard` option from the xterm's advertisement
  and ignores CSI u when withheld, as xterm does, replacing a per-reader
  helper that any new reader could skip. Daemon/headless users keep the
  default.
- Replay epilogues are writers (`writeReplayEpilogue`,
  `writeReattachReplayReset`) that take the sync or async xterm writer, so
  nothing that looks like a builder mutates the mirror.
- An abandoned hidden restore re-asserts the mirror's kitty flags after the
  byte-gap reset: its discarded chunks were already scanned.
- Tests pin a non-zero host restore (epilogue ends `=31u`, mirror 31), a
  withheld pane staying at 0, and the restart-in-place ground landing after
  the mirror reset; the epilogue test helper is now an exact builder.

* refactor(terminal): one epilogue writer and one scanned ground per boundary

- Reattach callers write `chooseReattachReplayReset(...)` through
  `writeReplayEpilogue`, dropping the second writer from the session.
- Fresh spawn and cold restore rely on their scanned ground alone: it leaves
  the mirror known at 0 with a proven baseline, so the extra reset() was dead.
2026-09-28 15:21:53 -04:00
8b410b4893 feat: add first-class Qoder CLI support (#23581)
feat: add first-class Qoder CLI support

Integrate Qoder launch, identity, canonical hook status, trust and resume.
Verify with captured Qoder 1.1.64 transcripts and hidden Electron sidebar checks.

Builds on and cross-reviews #7502, #8611, #9655, #12910, #13311 and #15291.

Co-authored-by: dalveytech-vincent <vincent@dalveytech.com>
Co-authored-by: Eridanus117 <45489268+Eridanus117@users.noreply.github.com>
Co-authored-by: xingqingzzp-gif <xingqingzzp-gif@users.noreply.github.com>
Co-authored-by: jyang2004 <jyang2004@users.noreply.github.com>
Co-authored-by: yunqian <yunqian@alibaba-inc.com>
Co-authored-by: huzhening.hzn <huzhening.hzn@alibaba-inc.com>
2026-09-28 02:59:50 -07:00
400e4e7957 feat(agents): add Freebuff launch and sidebar status support (#23567)
Add Freebuff launch support and execution-host status reporting for the sidebar, including running, question, blocked, and settled states. Validate against captured CLI transcripts and real rendered sidebar evidence.

Cross-referenced community implementations #17065, #20839, and the Freebuff portion of #18790. Preserve their agent/catalog/mobile/documentation coverage and add canonical status publication and regression tests.

Co-authored-by: Harkaran Brar <18134082+harkaranbrar7@users.noreply.github.com>
Co-authored-by: Prarambha369 <98906077+Prarambha369@users.noreply.github.com>
Co-authored-by: Lesley Murfin <260182349+LesleyMurfin@users.noreply.github.com>
2026-09-28 02:32:41 -07:00
Neil 45f3512a33 feat(agents): add first-class DeepSeek Harness (dsh) support (#22468)
* feat(agents): add first-class DeepSeek Harness (dsh) support

Register DSH as a supervised Orca agent: catalog entry and detection for its
dsh-tui profile, status/question hooks through DeepSeek's own Claude-Code hook
bridge, composer-ready prompt delivery, session resume, headless Source Control
AI, and title identity that no longer collides with Gemini's.

* fix(dsh): reach Orca through DSH's credential scrub and stop reading its title as Gemini

DSH runs command hooks through its own shell executor, which drops every env var whose
name contains KEY, TOKEN, SECRET or PASSWORD — taking ORCA_PANE_KEY and
ORCA_AGENT_LAUNCH_TOKEN with it, so every hook exited without posting. Mirror both onto
scrub-safe aliases at spawn and restore them at the top of the DSH hook script.

Its title collided too: DSH rests on the same glyph Gemini works on, so a resting DSH
pane was relabelled Gemini CLI and reported working forever. Defer both the Gemini
classifier and the title status detector on DSH's whale, in the base module both copies
of that classifier read.

* test(mobile): repin the session-route closure for the DSH agent icon

* fix(dsh): address review — never splice user rows, cover remote panes, keep the diff off argv

- findManagedDshPatchRegion paired an orphan start marker with a later block's end, so a
  truncated write made install/remove delete the user's own rows. Pair each end with the
  nearest preceding start; regression test fails without the fix.
- The relay PTY env builder never applied the scrub-safe aliases, so remote DSH status
  silently never appeared even with the remote hook installed.
- Source Control AI sent the whole diff on argv; send it over stdin with DSH's '-' marker.
- dsh-tui/dst already chose the interactive profile, so a workspace folder named 'web' or
  'plugin' no longer marks a live agent pane non-interactive.
- Isolate USERPROFILE as well as HOME so a Windows run cannot edit the real home.
- Drop the duplicate README badge and revert an incidental doc reformat.

* refactor(dsh): share the managed-hooks reader and tighten the new modules

Reuse before reimplementing: readManagedDshHookEvents was a near-verbatim copy of Muse's,
with byte-identical private helpers. Both now call one readManagedHookEventsFromJson.

Also: one readTextOrAbsent instead of two spellings of the same read (dropping an
existsSync TOCTOU), one status() builder instead of four inline literals, rmSync(force)
instead of exists-then-unlink, and a redundant empty-string guard before JSON.parse.
The patch-file transforms lose their index juggling for a predicate plus a filter.

* fix(dsh): refuse a flow-style patch file, keep its mode, and stop the relay inheriting a pane

- applyManagedDshPatch matched only an exact `[]`, so `[] # keep empty` or a non-empty
  flow sequence got a block entry appended after it — invalid YAML that would leave DSH
  unable to load the user's own patch layer either. It now strips the token from an empty
  sequence (keeping a trailing comment) and returns null for a non-empty one; install
  reports that and changes nothing.
- The patch rewrite dropped an owner-only file to the umask default (CWE-732); pass
  preserveMode.
- The relay PTY env never dropped inherited pane identity the way the local and daemon
  builders do, so a spawn that specified none could inherit the relay's own and every
  agent's hook would report against that pane.

* fix(dsh): keep the flow-style refusal in every status read, and scope the mode test to POSIX

A refused patch file carries no managed region, so getStatus() fell through to a bare
not_installed with detail null — the actionable 'rewrite it as a block sequence' message
only ever reached the one-shot install() return. Export the predicate and check it first,
behind one shared message constant.

The owner-only mode assertion cannot hold on Windows, where chmod only toggles the
read-only attribute and mode & 0o777 reads 0o666 for any writable file.

* docs(readme): restore the DeepSeek Harness badge lost in the rebase

* test(mobile): repin the session-route closure to the measured 4221

Measured, not derived: 4220 without the DSH icon entry, 4221 with it. Two of the three
modules above main's 4218 pin are not this change's — they arrived with the mobile work
after #22570 and were never repinned; the changelog records that split explicitly.

* fix(dsh): settle tui-idle on the agent's own hook, so supervised workers see it ready

Reported by a tester on the adhoc build: `terminal wait --for tui-idle` ran to its 90s
timeout against an already-ready DSH composer, so a supervised worker never sees the agent
as ready.

Every existing tier reads the title, and DSH deliberately carries no title status: its rest
prefix is Gemini's working glyph, so the detector reports none. A fresh first-party `done`
is better evidence than any title anyway — it is the agent's own account of its own turn,
and normalizeDshEvent drops subagent events, so it is the lead's. Scoped to DSH: for agents
whose hooks report child turns, a mid-turn `done` is the #6011 class this file prevents.

* test(daemon): record the DSH transcript's true-colour I2 divergences

Adding the dsh-tui capture to __fixtures__ enrolled it in the serialize replay sweep, where
it reports 10 I2 divergences and failed the unlisted-transcript default of 0.

Every one is the same shape — visible-grid row=0, a 24-bit background the round trip does
not restore to default — which is DSH's whale intro painting whole rows of true colour.
Verified as an upstream limitation rather than a regression by replaying against the
previous build (build-serialize-addon-at-ref.mjs --ref origin/main): I1 and I3 both hold.

* fix(dsh): return the new tui-idle verdict from the first-party done lane

Main refactored isTuiIdleSatisfied into evaluateTuiIdle, which returns a verdict rather
than a boolean. The DSH lane still returned `true`; it is tier-1 positive evidence, so it
returns READY_STRONG like the title/body lane above it. Re-verified the regression test
still fails without the lane.

* test(relay): pin the scrub-safe pane-identity aliases on the relay spawn path

The relay builds a remote pane's env itself, so the alias mirroring there had no
test: removing the call left every suite green while remote DSH status silently
vanished. Both cases fail without it.

* docs(dsh): point the hook service at the integration reference

The reference doc had no inbound link from anywhere in the repo.
2026-09-27 22:44:18 -07:00
Jinjing 708123b868 Improve PTY device error messages with localization support (#23538)
* Improve PTY device error messages with localization support

- Extract error hints to shared module for reuse across host and renderer
- Change multi-line hints from space to newline separator for readability
- Add localization of resource-limit hints in the renderer
- Prevent duplicate issue requests when toast renders its own link
- Handle legacy hint formats from older hosts

* Prevent duplicate PTY allocation hints on legacy messages

- Extract hint-detection logic into hasPtyAllocationHint() helper
- Check for both current and legacy PTY allocation hint variants
- Prevents duplication when messages already contain legacy hints
2026-09-27 22:43:37 -07:00
Jinwoo Hong 5219b8ada9 fix(relay): reset input modes a dead program left on in SSH terminals (#23488)
* fix(relay): ground input modes a dead app left armed on SSH terminals

SSH relay PTYs now run the same recovery barrier as the local daemon:
between startup ingress and the replay buffer/publish sink, it pauses at
an OSC 133;D that closes a command which left input modes or the
alternate screen armed, proves the shell owns the PTY foreground on the
execution host, and on proof injects the process-boundary ground ahead of
the prompt. Teardown flushes held bytes through releaseRelayIngress.

The barrier now holds only the D marker's terminator and carries the
ground on it as one transformed emission over exactly one raw unit.
Previously the ground was a zero-raw emission, which source-credit
delivery accepts but never sends, so live output and replay diverged.
Every emission now covers at least one raw unit; refuted, timed-out,
overflowed and flushed episodes release the terminator unmodified.

* fix(terminal): keep the held 133;D terminator inside its emission's raw span

Treat a trigger end below 1 as unsplittable instead of trusting the
scanner's clamp, arm the bail timer before queueing the terminator, and
build the relay's startup ingress unconditionally beside its barrier.

* fix(terminal): hold the whole 133;D mark so a mid-proof snapshot ends on an escape boundary

The scanner now reports where the unclean-death D mark starts; the
barrier releases everything before it and holds the mark itself (bounded
to 4K, so the grounded span stays inside a relay source frame). A
snapshot taken mid-proof therefore has no open OSC for consumers that do
not restore the pending escape tail. The relay defines PTY liveness once.
2026-09-28 01:27:07 -04:00
Brennan Benson 2119730ec0 fix(terminal-wait): unattended launches report Claude's trust dialog instead of timing out or typing into it (#22927)
* fix(terminal-wait): recognise Claude's workspace trust dialog as a blocking prompt

Claude's first-launch "trust this folder?" dialog parks the cursor above its
options, and the host's line tail drops the lines below it, which are the only
ones the trust matcher knew ("trust this folder", "Enter to confirm"). An
unattended Claude launch into a fresh folder therefore waited out its whole
budget and reported a timeout instead of the blocking prompt. The dialog's
opening question ("... one you trust?") survives in the tail, so it is now
recognised, pinned by a captured transcript of the real dialog.

* fix(terminal-wait): read the rendered screen for blocked prompts on the tui-idle poll

Claude's workspace-trust dialog parks the cursor on its highlighted option with a
cursor-up, and the host line tail deletes every retained row below the cursor, so
"Yes, I trust this folder" and "Enter to confirm" never reach the blocked-prompt
detector. In a live launch the dialog also arrives in 1024-byte reads, and the tail's
plain path blanks each line that ends in a carriage return before the newline, so
the opening question does not survive either. An unattended launch waited out its
whole budget and reported a timeout.

The runtime already feeds every PTY chunk into its own headless emulator. The tui-idle
poll now also runs the existing blocked-prompt rules over that emulator's visible
screen (no provider or host round trip), after the tail checks and before the
quiet-foreground idle fallback. The "one you trust" phrase is dropped: it only matched
when the whole dialog arrived in one chunk, wrapped away on narrow panes, and could
match prose.

Replays three live Claude 2.1.280 captures through the runtime: the dialog in one
chunk and in 1024-byte reads, a 60-column pane, and the dialog answered with "Yes",
which must report ready rather than blocked.

* fix(terminal-wait): skip the rendered-screen blocked check while the agent reports working

The screen check runs on every tui-idle poll, and the automation observer holds a tui-idle
wait open for a whole agent turn. A working Claude whose screen showed dialog wording (a diff
of the detector, say) was reported blocked where main kept waiting. The dialogs only the
screen reveals are start-up ones painted before any title, so a working title now vetoes it.

* fix(terminal-wait): settle weak tui-idle evidence only after a clean screen read

A shell auto-title (oh-my-zsh's `claude`, fish's `claude <cwd>`) names Claude before
its workspace trust dialog paints, and the line tail loses that dialog. The wait took
the bare name as rest, settled ready, and the launch typed its brief into the dialog.

Every tui-idle settle site now asks one evaluator for a verdict: blocked, strong ready,
working, weak ready or pending. The pre- and post-registration checks and the
title-change resolvers settle only blocked or strong ready; weak ready is left to the
poll, which settles it only once the rendered screen shows no blocker. A name-only
Claude title is held to the same quiet window as Codex and Devin, since Claude
announces rest with its own explicit title.

* test(serialize): record the answered Claude trust capture's known serializer divergences

The captured answered-dialog transcript added by this PR is replayed by the
serialize round-trip suite and diverges at 13 checkpoints. It diverges
identically on origin/main and on the pre-#22586 addon build (13 both-fail,
0 regressions): the live SGR pen leaks into the alt buffer, and an alt buffer
first entered after a shrink keeps hidden scrollback. Pin the count like the
other known captures and correct the comment, which called these upstream.
2026-09-27 20:59:02 -07:00
Jinwoo Hong 433986fa3b fix(runtime): read Codex readiness from the live screen (#23475)
* fix(runtime): read Codex readiness from the live screen

Codex 0.157 runs in embedded mode when Orca passes `-c model_reasoning_effort=…`,
shows a startup warning, and repaints its header by cell diff
(`ESC[5;3Hdir ESC[5;7Hctory:`). The tui-idle body check read the line-folded
tail, which drops those cursor moves and reads `dirctory:`, so worker-start
timed out at agent_readiness while Codex sat idle at its prompt.

For Codex (or unknown) panes whose live emulator screen shows the Codex header,
the screen now decides readiness: `model:` and `directory:` present and neither
still `loading`. Otherwise today's text rules apply unchanged. All six tui-idle
satisfaction sites share one helper so they cannot disagree.

Fixes STA-8628 / #23241.

* refactor(runtime): read the Codex screen lazily and only for Codex panes

Pass the screen as a thunk so non-Codex panes never build the grid, hoist the
pane agent lookup, and share the unblocked-ready check with the Muse rule.

* fix(runtime): let the Codex screen only add readiness

A grid out of step with the PTY (size mismatch or a resize mid-paint) can garble
Codex's header. Keep every verdict the text rules give today and consult the
screen only when they say not ready, so no flow that settles today can stop.

* fix(runtime): read Codex readiness from the header box only

Chat below the header can mention "OpenAI Codex" or "model: loading", so the
screen rule now reads model/directory/loading only inside the header box.

The serialize round-trip sweep picks up every runtime fixture; record the
pre-existing header-border attribute divergence the new Codex 0.157 recordings
expose, which this change does not touch.
2026-09-27 18:48:23 -04:00
Jinwoo Hong 618a8b0758 fix(terminal): keep a dead app's input modes recoverable after a refuted proof (#23474)
A command's armed input modes were demoted at the first OSC 133;D whether or
not the foreground proof confirmed, and the alternate-screen trigger was spent
the same way. A full-screen agent's nested command shells leak their own D onto
the main PTY; the refuted proof for that stray D used up the only trigger, so
the app's real death later went unrecovered.

Every D now re-asks while a command's mode (or the alternate screen) is still
up; only the confirmed ground or the app's own disable clears it. Proofs that
can never succeed (wsl.exe, no Windows job reads) already refute immediately
without spawning anything. The accepted cost is one process read plus a bounded
output hold per D while a command-owned mode stays up and the proof keeps
refuting: a live agent leaking D, a stopped job, a nested subshell, or an rc
that execs another shell.
2026-09-27 18:17:04 -04:00
NeilandHarshul Rathod bdb897b735 Respect disabled OpenCode variants and refresh WSL settings safely (#23328)
Respect disabled OpenCode variants, preserve explicit config ownership, and refresh WSL guest settings safely across reconnects.

Based on Harshul Rathod proposal #22805.

Co-authored-by: Harshul Rathod <harshulrathod1640@gmail.com>
2026-09-26 22:48:35 -07:00
OrcaWinandm4air 4b6fe95943 fix(windows): preserve relocated terminals and native process scans (#22872)
* fix(windows): ship the process-table addon to the relocated daemon host

The Windows terminal daemon runs from a copy of the app under
%LOCALAPPDATA%\Orca\daemon-host\<version>. That copy took node-pty but not
@vscode/windows-process-tree, so the daemon's bare require of the addon found
nothing and every process-table read (foreground tracking, descendant sweeps)
fell back to a powershell.exe Get-CimInstance scan (#16905).

- Copy the addon's runtime files (package.json, lib/, the .node binary) into
  the host; the ~25MB of gyp intermediates beside them are filtered out.
- Treat a host missing those files as unmaterialized, so hosts built before
  this are rebuilt, and skip relocation if the install itself lacks them.
- Log the daemon's native/CIM capability at startup and warn once when the
  process table falls back to CIM.

Revives #19525 on current main.

* test(windows): locate update-survival loss before relaunch

* test(windows): preserve daemon tree before update-survival proof

* test(windows): distinguish Electron exit from launcher close timeout

* test(windows): verify process exit when inherited pipes delay close

* test(windows): trace installer process checks in isolated survival runs

* fix(windows): probe process-query capability before installer sweep

* fix(windows): match installer probe and process-check profile behavior

* fix(windows): use NSIS separators for the process-check include

* test(windows): dismiss session-search overlay in survival harness

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 14:31:07 -07:00
OrcaWinandm4air f6631ddeee fix(daemon): release terminal attach cancellation listeners (#23191)
* fix(daemon): release terminal attach cancellation listeners

* fix(daemon): preserve attach wait settlement ordering

* fix(i18n): register existing diff note fallback strings

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 13:55:22 -07:00
Neil dd46f51278 perf(terminal): avoid intermediate history frame buffers (#23005) 2026-09-26 12:50:24 -07:00
OrcaWinandm4air 8b7a2a5393 Wait for foreground job readiness before testing Ctrl-Z (#23158)
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 01:28:55 -07:00
OrcaWin 6fc3cdcad6 Bundle Bun for headless Orca and profile persistence (#22635)
Bundle a pinned, verified Bun runtime for headless Orca so existing Node launch commands can hand off before opening a profile. Keep desktop execution on Electron.

Add the Bun SQLite adapter and terminal backend, bounded shutdown, process inspection and cross-platform artifact qualification. Keep future managed SSH deployment separate from current production launch paths.
2026-09-25 22:49:06 -07:00
OrcaWin 82412dab8b Persist profile state in SQLite with background writes (#22612)
Migrate profile state to SQLite and move writes and backups into a background worker. Acknowledge terminal, SSH and automation changes only after durable saves. Preserve JSON import, recovery, rollback and compatibility exports.

Validate migration, worker failures, maintenance, cross-profile moves and terminal lifetime races with unit, integration and end-to-end coverage.
2026-09-25 22:47:33 -07:00
OrcaWinandOrca Worker 58d1ff3b6a Provide the Orca CLI automatically in managed WSL terminals (#22761)
* Provide the Orca CLI automatically in managed WSL terminals

* Simplify managed WSL CLI provisioning

Never block a shell on CLI availability, keep the shared WSL login-shell
builder unchanged, provision from PTY env assembly only, drop the error
variable and command probing, and reuse the existing WSLENV helper.

* Scope the managed WSL CLI to WSL terminals

Provision only for WSL panes and publish the directory through
addOrcaWslInteropEnv, so daemon terminals keep inherited WSLENV entries and
non-WSL builds never see the variable. Write the bridge with a UTF-8 BOM so
Windows PowerShell 5.1 keeps non-ASCII user-data paths, give the dev bridge the
dev launcher's app-launch env, and drop the unused skill-setup wiring and
runtime capability.

* Tighten the managed WSL CLI bridge and setup

Launch the bridge child exactly like the registered bridge (no hidden
window or output relay; verified through WSL with Node and Electron), give
the dev bridge the dev launcher's NODE_OPTIONS stash, clear the guest-only
directory before starting Windows processes, collapse setup into one
function, warn once, and guard WSL env routing with tests.

* Harden managed WSL CLI quoting and inheritance

PowerShell also ends single-quoted strings at typographic quotes, so a
user-data path such as O'Brien with a curly apostrophe broke the managed
bridge. Fix the shared quotePowerShellLiteral and reuse it. Drop an inherited
ORCA_WSL_CLI_DIR on the daemon path, remove the unreachable PATH dedupe, and
cover failed setup with a stale caller value.

* Cover the managed WSL CLI in zsh and on POSIX CI

Add a live zsh case that reaches a real prompt, a POSIX test that runs the
PATH restore snippet in bash and zsh under set -u, and a null result for
unwritable user data. Say what a failed write actually costs, and document
per-spawn write logging and older-daemon behaviour.

* Keep system bashrc out of the PATH restore test

CI runners make bash -c read /etc/bash.bashrc, which fails under set -u.

---------

Co-authored-by: Orca Worker <orca-worker@localhost>
2026-09-25 21:55:21 -07:00
Brennan Benson 2796a3ac15 fix(claude): prove a stopped chat's child processes gone when they exit with it (#22918)
* fix(claude): prove a stopped chat's child processes gone when they exit with it

Stopping a Claude chat snapshots its child processes, closes Claude, and then verifies each child is gone before the stop counts as proven. The verifier only accepted a child as gone after that child had appeared in one of its own process-table reads. When Claude exits gracefully it takes its short-lived children with it before the first read, so none of them was ever seen again. Every read confirmed them absent, and the verdict was still "unverifiable" after the full 3.5 s window. Measured live: 37 complete reads, target absent from all, verdict unverifiable, on every idle stop.

The snapshot is itself a table read that saw each child alive, so it now counts as the first sighting. An absence counts only from a read that started after the child was last seen, which keeps what the old rule protected against: a shared or in-flight read begun before the snapshot cannot list a child forked since. Two such absences prove a child gone. Live, the same stop now proves the tree gone in about 150 ms.

The daemon's terminal shutdown uses the same verifier and gets the same rule. Test reads that reused one capture stamped with the snapshot's own time now stamp each read when it starts, as real scans do.

* fix(claude): keep the latest sighting and count the final read as an absence

A matching read that started earlier but resolved later could move a target's
last sighting back and let an older absence count; the sighting now only moves
forward. The read after the deadline now records its absences the same way the
polling loop does, so a second qualifying absence there proves the target gone.
2026-09-25 14:57:31 -07:00
Jinwoo Hong 2c7609bf6a fix(terminal): serialize only the visible width after a column shrink (#22586)
* fix(terminal): serialize only the visible width after a column shrink

xterm does not reflow the alternate buffer (or a normal buffer under
pre-21376 ConPTY), so after a shrink each line keeps its old length.
SerializeAddon walked every non-final row to line.length, so any snapshot
taken after a shrink carried stale right-hand cells that wrapped into extra
rows on replay; restores repainted that garbage and a differential TUI such
as OpenCode never cleared it.

Clamp the row walk and the wrap-boundary lookups to the terminal's columns
in Orca's addon-serialize source patch, and regenerate the bundles, maps and
lockfile hash per docs/reference/xterm-patch-regeneration.md.

* test(terminal): read shrink-snapshot fixtures through public APIs

Drops the private-terminal casts the casting gate flags; the normal-buffer case
now drives a plain pre-21376 ConPTY terminal and its SerializeAddon directly.

* fix(terminal): blank a wide glyph clipped by a column shrink when serializing

After a non-reflowing shrink a width-2 glyph can have its lead half in the last
column and its trailing half past the grid. Serializing the lead half makes the
replay wrap it to the next row and shift every row below, so serialize that
cell as a blank and keep the row exactly the grid's width. A glyph ending
exactly at the edge is unchanged.

* test(mobile): move the session closure pin past the main agent status modules

#22452 added src/shared/main-agent-status.ts and src/shared/agent-turn-outcome.ts,
which agent-status-types.ts imports, so the session route's closure grew by two
local modules (4218 -> 4220). That change was src/shared-only, so its own CI never
ran this suite; main has been at 4220 since, and any PR that fires the mobile web
app job fails on the stale pin. Measured on 4064653740 and on origin/main 3ea15dd0a2.

* test(terminal): differential serialize round-trip fuzz and transcript replay

Seeded VT streams (text, CJK/emoji/combining, SGR, cursor/edit ops, scroll
regions, DECAWM/IRM, alt-screen variants, DECSC, and shrink-heavy resizes)
drive a source terminal in three modes: reflowing normal buffer, alternate
buffer, and a non-reflowing pre-21376 ConPTY normal buffer. At each checkpoint
every SerializeAddon build under test serializes it, and each output is
replayed into a fresh terminal of the same size and compared cell by cell,
plus cursor, active buffer and modes.

CI runs 25 seeds per mode against a pinned list of pre-existing divergences,
and replays the committed PTY transcripts (the existing agent fixtures plus new
vim, less, pico, and OpenCode captures) under four resize schedules. Point
ORCA_OLD_SERIALIZE_ADDON at a previous patched build to also check byte
identity when no line is wider than the grid, and that no checkpoint regresses.

* test(terminal): build the differential serialize baseline from any git ref

config/scripts/build-serialize-addon-at-ref.mjs reverse-applies the patch that
produced the installed @xterm/addon-serialize dist, applies the ref's patch,
and verifies each step against the patches' blob hashes, so the fuzz can use
origin/main (or any fix commit) as its baseline without a second install.
ORCA_NEW_SERIALIZE_ADDON swaps in a built dist for the build under test, and a
seed-pinned test replays the nine I3 regressions found against origin/main.

* fix(terminal): serialize a clipped wide glyph as a width-1 blank

The stand-in for a wide glyph clipped by a column shrink came from getNullCell(),
whose width is 0. _nextCell skipped it as a wide trailer, and the row-end wrap
check counted the width-0 _backgroundCell as content, so a soft wrap after the
clipped column was taken as natural and replayed one column early
(conpty seed 1149: `abcdefghi中WRAPPED` at 12 -> 10 cols replayed as
`abcdefghiR`/`APPED`). Blank the cell in place instead: width 1, no codepoint,
its own attributes, so it counts as one empty column and forces the wrap.

Differential sweep vs origin/main, 7000 cases per mode: I1 0 byte diffs, seed
1149 fixed; the remaining I3 regressions are the trailing background-row seeds.

* test(terminal): neutral paths in the serialize fixtures and usage comment

The OpenCode transcript carried this machine's lane paths in its footer; replace
them with same-length neutral paths so the recorded cursor layout is unchanged.

* fix(terminal): keep trailing background-only rows when serializing

Without scrollback, _serializeString trims rows after the last content cursor.
A row made only of background-colored blanks emits its erase in _rowEnd but
never moved that cursor, so two or more such rows at the bottom were dropped
(4x3 `r1\r\n\e[48;5;157m\e[J\e[0m\e[3;1H` replayed with the last row blank).
Track the erase separately and extend the kept rows to it, except when the
cursor is wrap-pending: relative moves back from those rows cannot re-create
that state, and doing so regressed normal 508, alt 1425/6647, conpty 4699.

Harness: I1 now exempts checkpoints with a background row after the last text
row, the one place this fix changes bytes on purpose (scope helpers move to
serialize-grid-variant-scope.ts); conpty seed 5 leaves the pinned pre-existing
list. Sweep vs origin/main, 7000 cases per mode: I1 0, I3 regressions 0;
fixed/both-fail normal 967/3914, alt 2628/8316, conpty 6657/2937 (was
323/4558, 2296/8648, 5466/4128).

* test(terminal): narrow the I1 carve-out to where the serializer keeps background rows

Trailing background-only rows change bytes only when the serialized range has no
scrollback (the trimming path) and the cursor is not wrap-pending; checkpoints with
scrollback or a wrap-pending cursor are held to byte identity again. The 7000-per-mode
sweep against origin/main stays at I1 0 and I3 0.

* test(native-chat): pin that screen scrapers ignore kept background rows

The serializer now keeps trailing background-only rows, so a painted TUI's screen
read as text ends in \r\n\x1b[NX rows. Serialize the same frame with and without
them and check the Claude option scrape, the empty-prompt check and the fork
transcript read the same thing.

* test(terminal): read xterm core internals through a checked parser, not Reflect.get

The fuzz oracle reached xterm's private _core with Reflect.get, which the
low-evidence gate rejects. Narrow _core, writeSync and the DECSTBM bounds with
in/typeof checks into one named XtermCoreInternals shape instead.

* test(terminal): keep captured serialize transcripts byte-exact on Windows checkouts
2026-09-25 01:18:14 -04:00
Jinwoo Hong f559c0588a fix(terminal): ground a program that dies with input modes armed on the normal screen (#22739)
* fix(daemon): rebase durable checkpoints on the live terminal

A durable checkpoint was folded from the previous checkpoint plus recorded
output, so it inherited that checkpoint's modes forever. After a daemon
restart killed a full-screen TUI and a new process started inline, the chain
kept the dead TUI's alt screen and mouse tracking (?1049h ?1003h ?1006h)
while the live emulator was clean. Every reattach and getBufferSnapshot
served the stale chain, the renderer re-armed mouse tracking, and wheel
scrolling went to a program that never asked for it: scrolling froze.

Each full checkpoint is now the live snapshot verbatim (screen, layout,
alt frame, modes, owner) with only the normal-buffer rows live has evicted
taken from the durable replay. A checkpoint can no longer carry a dead
process's modes, and checkpoints already poisoned on disk heal on the next
compaction.

- The first fold after a cold restore replays the same seed segments live
  was given, so rows line up even over a dead TUI's alt screen.
- Idle zero-record folds keep the disk copy when it already agrees with
  live, so quit and relaunch bursts don't replay every session.
- Held teardown bytes are already in the drained records and the live
  snapshot, so they are no longer replayed twice or appended as a tail.
- The bounded getBufferSnapshot path honors the requested depth even when
  the live window is deeper, without phantom link rows.
- The fold's ownership scanner and frame merge are removed; owner and
  frame come from live.

* fix(terminal): one process-boundary ground for every known or proven boundary

Three copies of the "the process that armed these modes is gone" reset had
drifted: the cold-restore seed cleared only pen and mouse, the recovery
barrier used the renderer's dead-TUI profile, and the cold-restore payload
had none. A cold restore therefore left the dead process's focus reporting,
bracketed paste, application cursor and keypad modes armed in the live
emulator, the first checkpoint, and main's mirror. And the seed wrote the
dead process's torn escape after the reset, so the new shell's first bytes
could complete it (for example retitling the pane).

PROCESS_BOUNDARY_GROUND replaces them: CAN, leave the alt screen without
moving the normal-buffer cursor, every mouse protocol and encoding off,
focus/paste/app-cursor/keypad off, cursor shown and style reset, kitty
popped, SGR reset, grounded DECSC. It stays inert for the lifecycle
scanner. The seed, the recovery barrier, and the cold-restore payload all
use it, and the seed no longer carries the torn tail.

The first fold after a cold restore now always rebases on live, because
focus and keypad are not in TerminalModes and the zero-record shortcut
could not see them differ.

* fix(terminal): ground a program that dies with input modes armed on the normal screen

The daemon's in-stream crash detector only fired when a program died with
the alternate screen up. A normal-buffer program that armed mouse tracking,
focus reporting, keypad or kitty keyboard flags and exited without
disabling them was cleaned up only in the renderer, so the daemon kept the
modes and re-armed them on the next reattach, mobile included (#13077's
garbage-at-the-prompt family).

The lifecycle scanner now tracks armed input modes (mouse protocols and
encodings, ?1004, ?66, and kitty flags as per-screen stacks that mirror
xterm's main/alt swap and its 16-entry cap). ?2004 and ?1 are excluded:
shells arm them at their own prompts. Modes armed when a command starts
(OSC 133;C) count as the shell's, so a prompt that leaves modes on never
triggers. At OSC 133;D the trigger is now "alt screen or a program-armed
input mode", still one-shot and still gated by the shell proof, and the
existing PROCESS_BOUNDARY_GROUND is recorded through the stream so live
and durable history change together.

WSL panes spawn wsl.exe, which the shell proof does not recognise, so the
detector never grounds them; a test pins that and the renderer keeps
covering them. The mouse-leak e2e now keeps its arming process alive until
the live pane is checked, because the daemon grounds a proven exit.

* fix(terminal): keep shell- and host-armed input modes through the process-boundary ground

ConPTY arms focus reporting (?1004h) before the first prompt, and the live
recovery ground cleared it for the rest of the pane. The barrier now re-arms
the modes that were on at OSC 133;C right after the ground, so only the dead
program's modes are reset.

* fix(terminal): re-assert only modes the shell or host armed outside a command

A mode a program leaked past a refuted proof was still on at the next OSC
133;C, so the baseline snapshot re-armed it after a later ground. Record who
armed each mode instead: only enables outside a command (before any marker,
or between 133;A/D and C) form the baseline.

* fix(daemon): keep OSC links and kitty flags through durable checkpoint folds and trims

Stop seeding persisted OSC link ranges into the fold replay: they index the
base buffer, so rows evicted by pending output left a link on the wrong text.
The serializer already writes OSC 8 into the ANSI the fold replays.

Re-apply kitty keyboard flags when replaying a snapshot for trimming, since
rehydrateSequences omits them.

Bound a smaller restore request by trimming the committed checkpoint instead
of re-reading disk and rebasing the live window at a smaller depth.

* test(daemon): follow the isFirstTake rename in the process-boundary ground suite

* refactor(daemon): drop the unreachable deep-live branch from the durable fold

The live window's override cap now derives from the restore depth, so live can
never be deeper than the fold. pendingRecords and isFirstTake are required.

* test(daemon): pass pendingRecords to the process-boundary ground fold

* fix(terminal): reset alt-screen kitty flags in the process boundary ground

Kitty keyboard stacks are per screen, so resetting only after ?1049l left
a dead TUI's alt-screen flags for the next alt-screen app. Also drop the
inert CAN from the ground (every site grounds after complete bytes) and
correct two stale comments.

* fix(terminal): track input-mode ownership in one map

Each armed mode now has one owner: host (before any marker, or a prompt a
133;C proved), prompt (unproven until C), command, or stale (left past a D).
Host arming is sticky, 133;D demotes command modes (the one-shot), and the
ground re-asserts only host modes. Fixes a D without C triggering on host
modes, an ESC c mid-command turning later enables into host modes, and a
program's repeated host enable dropping host ownership. The reattach e2e now
keeps the arming program alive so only the reattach reset can disarm it.

* refactor(terminal): stop treating kitty flags as host state

fish, the one shell that pushes kitty flags at its prompt, pops them before
running a command and re-pushes at the next prompt, so the ground never needs
to restore them. Only host private modes are re-asserted now.

* fix(terminal): keep host input-mode ownership across RIS

ConPTY answers a mid-command ESC c by re-sending ?1004h, which reset() had
recorded as the command's, so the ground turned host focus reporting off for
the rest of the pane. RIS now drops only non-host ownership.

* fix(terminal): let only the host own focus reporting and leave it in the ground

Host ownership covered every mode armed before the first marker, so a tmux
that died with mouse on had it re-armed by the ground. And the re-assert's
?1004h enable made the runtime's ownership mirror revoke, so remote owners
never settled on Windows. Only ?1004 can be host-owned now, and the ground
skips its ?1004l instead of turning it off and back on, so injected bytes
carry no enables.
2026-09-25 00:50:04 -04:00
Jinwoo Hong fe46138716 fix(terminal): one process-boundary ground for every known or proven boundary (#22735)
* fix(daemon): rebase durable checkpoints on the live terminal

A durable checkpoint was folded from the previous checkpoint plus recorded
output, so it inherited that checkpoint's modes forever. After a daemon
restart killed a full-screen TUI and a new process started inline, the chain
kept the dead TUI's alt screen and mouse tracking (?1049h ?1003h ?1006h)
while the live emulator was clean. Every reattach and getBufferSnapshot
served the stale chain, the renderer re-armed mouse tracking, and wheel
scrolling went to a program that never asked for it: scrolling froze.

Each full checkpoint is now the live snapshot verbatim (screen, layout,
alt frame, modes, owner) with only the normal-buffer rows live has evicted
taken from the durable replay. A checkpoint can no longer carry a dead
process's modes, and checkpoints already poisoned on disk heal on the next
compaction.

- The first fold after a cold restore replays the same seed segments live
  was given, so rows line up even over a dead TUI's alt screen.
- Idle zero-record folds keep the disk copy when it already agrees with
  live, so quit and relaunch bursts don't replay every session.
- Held teardown bytes are already in the drained records and the live
  snapshot, so they are no longer replayed twice or appended as a tail.
- The bounded getBufferSnapshot path honors the requested depth even when
  the live window is deeper, without phantom link rows.
- The fold's ownership scanner and frame merge are removed; owner and
  frame come from live.

* fix(terminal): one process-boundary ground for every known or proven boundary

Three copies of the "the process that armed these modes is gone" reset had
drifted: the cold-restore seed cleared only pen and mouse, the recovery
barrier used the renderer's dead-TUI profile, and the cold-restore payload
had none. A cold restore therefore left the dead process's focus reporting,
bracketed paste, application cursor and keypad modes armed in the live
emulator, the first checkpoint, and main's mirror. And the seed wrote the
dead process's torn escape after the reset, so the new shell's first bytes
could complete it (for example retitling the pane).

PROCESS_BOUNDARY_GROUND replaces them: CAN, leave the alt screen without
moving the normal-buffer cursor, every mouse protocol and encoding off,
focus/paste/app-cursor/keypad off, cursor shown and style reset, kitty
popped, SGR reset, grounded DECSC. It stays inert for the lifecycle
scanner. The seed, the recovery barrier, and the cold-restore payload all
use it, and the seed no longer carries the torn tail.

The first fold after a cold restore now always rebases on live, because
focus and keypad are not in TerminalModes and the zero-record shortcut
could not see them differ.

* fix(daemon): keep OSC links and kitty flags through durable checkpoint folds and trims

Stop seeding persisted OSC link ranges into the fold replay: they index the
base buffer, so rows evicted by pending output left a link on the wrong text.
The serializer already writes OSC 8 into the ANSI the fold replays.

Re-apply kitty keyboard flags when replaying a snapshot for trimming, since
rehydrateSequences omits them.

Bound a smaller restore request by trimming the committed checkpoint instead
of re-reading disk and rebasing the live window at a smaller depth.

* test(daemon): follow the isFirstTake rename in the process-boundary ground suite

* refactor(daemon): drop the unreachable deep-live branch from the durable fold

The live window's override cap now derives from the restore depth, so live can
never be deeper than the fold. pendingRecords and isFirstTake are required.

* test(daemon): pass pendingRecords to the process-boundary ground fold

* fix(terminal): reset alt-screen kitty flags in the process boundary ground

Kitty keyboard stacks are per screen, so resetting only after ?1049l left
a dead TUI's alt-screen flags for the next alt-screen app. Also drop the
inert CAN from the ground (every site grounds after complete bytes) and
correct two stale comments.
2026-09-24 23:38:53 -04:00
Jinwoo Hong 1c1b7829ec fix(daemon): rebase durable checkpoints on the live terminal (#22732)
* fix(daemon): rebase durable checkpoints on the live terminal

A durable checkpoint was folded from the previous checkpoint plus recorded
output, so it inherited that checkpoint's modes forever. After a daemon
restart killed a full-screen TUI and a new process started inline, the chain
kept the dead TUI's alt screen and mouse tracking (?1049h ?1003h ?1006h)
while the live emulator was clean. Every reattach and getBufferSnapshot
served the stale chain, the renderer re-armed mouse tracking, and wheel
scrolling went to a program that never asked for it: scrolling froze.

Each full checkpoint is now the live snapshot verbatim (screen, layout,
alt frame, modes, owner) with only the normal-buffer rows live has evicted
taken from the durable replay. A checkpoint can no longer carry a dead
process's modes, and checkpoints already poisoned on disk heal on the next
compaction.

- The first fold after a cold restore replays the same seed segments live
  was given, so rows line up even over a dead TUI's alt screen.
- Idle zero-record folds keep the disk copy when it already agrees with
  live, so quit and relaunch bursts don't replay every session.
- Held teardown bytes are already in the drained records and the live
  snapshot, so they are no longer replayed twice or appended as a tail.
- The bounded getBufferSnapshot path honors the requested depth even when
  the live window is deeper, without phantom link rows.
- The fold's ownership scanner and frame merge are removed; owner and
  frame come from live.

* fix(daemon): keep OSC links and kitty flags through durable checkpoint folds and trims

Stop seeding persisted OSC link ranges into the fold replay: they index the
base buffer, so rows evicted by pending output left a link on the wrong text.
The serializer already writes OSC 8 into the ANSI the fold replays.

Re-apply kitty keyboard flags when replaying a snapshot for trimming, since
rehydrateSequences omits them.

Bound a smaller restore request by trimming the committed checkpoint instead
of re-reading disk and rebasing the live window at a smaller depth.

* refactor(daemon): drop the unreachable deep-live branch from the durable fold

The live window's override cap now derives from the restore depth, so live can
never be deeper than the fold. pendingRecords and isFirstTake are required.
2026-09-24 23:09:36 -04:00
mmarabel 7a71e20860 fix(pi): isolate status ownership in new terminals (#22717)
* fix(pi): isolate status ownership in new terminals

* docs(pi): explain terminal ownership boundaries

* docs(pty): clarify environment rescrubbing
2026-09-24 18:47:39 -07:00
Neil 0b16a31e6e fix(runtime): budget explicit terminal close for the daemon's immediate-kill verdict (#22385)
* fix(runtime): budget explicit terminal close for the daemon's immediate-kill verdict

Explicit terminal close (worker-release, worker-stop, `orca terminal close`)
gave the daemon kill RPC a fixed 2s deadline. The daemon's immediate kill
captures descendants, SIGTERMs them with a 2.5s verification window, then
waits up to 8s for the root's physical exit. An agent that runs exit hooks
after SIGTERM (Muse: ~3s) outlived main's 2s timeout, so close reported the
PTY unverifiable and worker-release returned release_unknown even though the
daemon confirmed the exit ~200ms later.

Derive the close budget from the daemon's own immediate-kill reply budget
(now in an import-free module) plus 2s for the post-kill inventory check. A
process that exits within the daemon's budget is released; a wedged process
or unreachable host still times out as unverifiable.

* fix(runtime): budget the force-kill retry and exercise an expired close deadline

* test(runtime): drop tautological deadline-expiry assertion
2026-09-22 22:44:14 -07:00
Jinwoo HongandDavid Bebawy 7c46a69049 feat(telemetry): report the macOS daemon's code identity on adoption and folder-denial events (#22171)
* feat(daemon): import the macOS process code-identity probe from PR #21826

Takes `daemon-mac-code-identity.ts` and its test verbatim from David Bebawy's
community PR #21826 (stablyai/orca). The probe asks Security.framework, via
`codesign --display --verbose=1 +<pid>`, where a live process's code lives on
disk — the question Node cannot answer, and the one that decides whether tccd
can still resolve a running daemon's code identity after an app update.

Imported unchanged here so the adaptation that follows is reviewable as a diff
against the author's original.

Co-authored-by: David Bebawy <david.ayad2@gmail.com>

* feat(telemetry): report the daemon pid's macOS code identity on the two adoption events

Community PR #21826 argues that macOS terminal daemons lose Documents/Desktop/
Downloads access after an update because the daemon's own executable is
unlinked — Squirrel parks the outgoing bundle under a ShipIt staging directory
and later deletes it — so tccd can no longer map the daemon pid to on-disk
code. Today's `spawner_path_class` and `tcc_attribution` read the binary that
forked the daemon, which an in-place update deletes and recreates, so neither
can see that state.

This adds the detector as a measurement only. `code_identity` rides on
`daemon_adopted` and `daemon_pty_cwd_denied`, the two events that already
describe an adopted daemon, so denied daemons can be cross-tabbed against
healthy ones. Nothing reads the verdict: no replacement, no notice, no UI.

The probe is David Bebawy's, narrowed from a path-carrying union to the closed
enum the wire allows, and memoised per pid so one codesign spawn answers for a
whole daemon generation. Off macOS, or with no pid, it reports `probe-failed`,
which keeps both schemas strict and non-optional.

Co-authored-by: David Bebawy <david.ayad2@gmail.com>

* fix(telemetry): read the daemon's code identity fresh on every adoption event

The probe memoised its verdict per pid and never expired it, so
`daemon_pty_cwd_denied` reported whatever the probe saw at adoption rather than
what was true at the denial. That breaks the measurement in both directions: a
transient codesign failure during startup pinned `probe-failed` for the rest of
the run, and the `parked` to `unresolvable` transition became invisible.
Squirrel leaves the parked bundle in place until the next update, which can be
days, so a daemon adopted as `parked` and denied as `unresolvable` is the exact
crossover this study exists to catch, and the cache hid it.

Now every ask runs its own codesign. Only concurrent asks about the same pid
share a probe, and that entry is cleared as soon as it settles, so nothing
survives to be reported later. Both events are rare enough that one spawn each
is not worth a cache.

* fix(telemetry): drop the dead existence check from the code-identity probe

The classifier stat'd the path codesign displayed and called a missing one
unresolvable. That path is unreachable: once the executable is unlinked,
`codesign --display` prints no `Executable=` line at all and exits 1 with
"No such file or directory", which the fallback below already classifies as
unresolvable. Verified directly on Darwin 25.5 against a signed binary deleted
out from under a running pid.

All the branch actually covered was the window between codesign reading the
path and this process stat'ing it, and it paid for that with a synchronous
stat on the main thread.

* fix(telemetry): never classify a timed-out codesign probe as a verdict

`runProcess` kills the child at the deadline and reports `timedOut`, but the
runner type dropped that field, so a codesign killed mid-display could still
have printed an `Executable=` line and been read as `resolved` or `parked`.
A half-written display proves nothing about where the daemon's code lives.

The runner result now carries `timedOut`, and a timed-out probe returns
`probe-failed` before the output is looked at.

* docs(telemetry): state what each code-identity verdict actually asserts

A reviewer read `resolved` as a claim that the executable sits inside the
installed app and asked for that to be validated. It is not that claim, and we
are not making it: proving containment needs the pid record's spawner path, and
deciding anything from where the code lives is #21826's proposed behaviour
rather than this measurement.

The enum doc now spells out all four verdicts in the terms the probe can
actually support, and says plainly why `resolved` stops at "exists and is not
parked". A matching note sits beside the parked-path pattern.

* docs(telemetry): stop asserting how long a parked bundle survives

The probe's rationale claimed Squirrel keeps the parked bundle "until the next
update". A reviewer claimed the opposite, that it is deleted at the end of the
same install. Neither holds up against this Mac's ShipIt log: the install moves
the outgoing bundle to a TMPDIR ShipIt directory and logs no removal of it at
all, and the one "Couldn't remove owned bundle" line names the incoming
download staging copy, not the parked one. Every parked bundle from the last
two days is nevertheless gone now.

So the rationale in the probe doc, the enum doc, and the reprobe test comment
now assert only what is established: the outgoing bundle is moved aside at
install and disappears later on a schedule we have not pinned down. That is
already enough to justify the design, since one pid's verdict can change
within an app run, which is exactly why every ask reads fresh.

* feat(telemetry): report readable TCC-gated spawns as the code-identity control

`daemon_pty_cwd_denied` gives code_identity's hit rate on denials, but a
readable spawn emitted nothing, so an `unresolvable` adoption with no denial
could not be told apart from a user who never opened a terminal in Documents,
Desktop, or Downloads. The false-positive rate that gates #21826's
auto-replacement was unmeasurable.

`daemon_pty_cwd_readable` now fires when a daemon reads a TCC-gated cwd, once
per daemon and folder class per app run, with the same origin properties as
the denial event. The read-out becomes a 2x2 of code_identity against
readable/denied on protected-folder spawns. Fire-and-forget on the spawn path
like the denial emit, and no app-side directory read.

* refactor(telemetry): one emitter and schema for both cwd verdicts, no dedupe state

The once-per-daemon dedupe on `daemon_pty_cwd_readable` was keyed before the
probe ran, so a daemon first seen readable while `parked` never reported again
once it turned `unresolvable` — the one cell that would count most against
#21826. It also counted per daemon while denials count per spawn, so the 2x2
mixed units.

Readable now reports every spawn, like denied, and both events share one
emitter (`trackDaemonPtyCwdVerdict`) and one schema. The TCC-folder gate lives
in the verdict branch. The origin fields are one shape spread into both
schemas. The codesign probe calls `runProcess` directly and tests mock it,
replacing a test-only runner parameter. The repeated "never cached" rationale
is now said once.

* fix(telemetry): rename the shared origin schema fields for the anti-slop gate

no-shape-in-symbol-names rejects daemonOriginShape; the fields are event props.

---------

Co-authored-by: David Bebawy <david.ayad2@gmail.com>
2026-09-22 22:10:51 -04:00
OrcaWinandm4air ba742a86bb fix(linux): release orphaned processes when their owner exits (#22247)
* fix(linux): release orphaned processes when their owner exits

* fix(linux): handle inhibitor errors until streams close

---------

Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
2026-09-22 05:10:03 -07:00
OrcaWinandm4air 632ae1320b fix(daemon): reap terminal descendants during shutdown (#22232)
Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
2026-09-22 04:32:26 -07:00
Jinwoo Hong 7650abe224 fix(macos): tell the user when Orca's terminal service can't read their folder, and walk them through the fix (#21923)
* fix(macos): tell the user when Orca's terminal service can't read their folder

On macOS, a terminal daemon that survived an app update can be refused access to
a workspace under Documents, Desktop, or Downloads while the Orca app itself can
still read it. Terminals opened there die with "Operation not permitted" and
nothing on screen explains why. The daemon has reported `cwdReadableByDaemon` on
every create since #18043 and main has emitted `daemon_pty_cwd_denied` on proven
divergence since then; the field data says 1,438 users hit it in 21 days. What
was missing was the notice.

The verdict itself moves off `access()`. A grant-less probe on an affected
machine showed a TCC mode where `access(R_OK|X_OK)` passes on `~/Documents` and
`opendir` still fails, so the check now does what a shell listing its cwd does:
`opendirSync`, one `readSync`, `closeSync`. Only EPERM/EACCES reads as denial —
a missing path, a non-directory, or an unexpected error still reads as readable,
so a non-permission failure can never masquerade as one. The same probe is what
the app side compares with, through one oracle shared by the telemetry emitter
and the notice, so the spawn path reads the directory once.

Proven divergence now also records evidence in main: one entry, keyed by the
daemon's pid, start time and launch nonce, carrying an opaque digest of that
identity and the folder class. No path leaves main. The existing focus-time
`macTccAttribution` poll carries it to the renderer, which raises a second toast
latched per daemon scope: dismissed stays dismissed, and a restart mints a new
identity so the poll returns null and the toast clears with no post-restart
probe. If the replacement daemon is denied too, about 31% of cases, the next
spawn re-records under the new scope and the notice returns, now with the
re-allow sentence doing the work.

No new IPC channel, no daemon protocol field, no polling change, and nothing new
on the spawn path beyond one `opendir`. `daemon_folder_access_notice` counts
shown, dismissed and open_manage_sessions against `daemon_pty_cwd_denied` as the
denominator; `shown` is emitted from main the first time a scope leaves the IPC
handler, so the renderer carries no telemetry plumbing for it.

* fix(macos): clear folder-access evidence only when the same folder class reads back

A readable spawn in ~/code said nothing about a Documents denial but was
hiding the notice; retire the evidence only when the daemon reads a folder
of the class it was denied on.

* fix(macos): say what a terminal-service restart actually does

The Manage Sessions restart confirmation still described the product as it was
before agents resumed themselves: it promised panes showing "Process exited"
that the user reopens by hand, and mentioned legacy-protocol sessions nobody
outside the daemon code can act on. Open terminals and agents come back on
their own now, so the old copy made a routine remedy sound like data loss.

It also called the thing a "daemon". The same restart is about to be offered
from a user-facing fix dialog, so both surfaces now say "terminal service", and
the confirm button is just "Restart".

The new body adds the one fact the old one never stated: terminals on remote
hosts are not affected. Translations of the two changed strings are dropped so
the five non-English locales fall back to English rather than keep showing copy
that is now wrong.

* feat(macos): give the denied-folder notice a fix the user can follow

The folder-access toast told the user their terminal service could not read
Documents and then handed them a paragraph: restart from Manage Sessions, and
if that does not work, re-allow Orca in System Settings. Both halves were
guesses. Roughly a third of restarts do not fix it, and the user had no way to
know which case they were in before spending every open terminal on finding
out.

Main can now answer that. `daemon-folder-access-probe.ts` forks a short-lived
child of the app binary the same way the daemon itself is forked, runs one
opendir/readdir/closedir against the denied path, and prints a single JSON
line. macOS attributes a TCC grant to the process that forked the child, so a
child of the app running now answers exactly the question the running daemon
cannot: would a replacement daemon get in? The child goes through the shared
child-process wrapper, never a shell, with a 3s deadline, a 1KB output cap and
an environment scrubbed to PATH/HOME/TMPDIR. Every failure — timeout, bad
output, spawn error — reads as `unknown`, never as a verdict.

That answer rides out as `restartWillHelp` on the evidence the existing
focus-time poll already carries, and the toast becomes a title and two buttons:
Fix… and Not now. Fix opens a dialog with the two real steps. When the grant is
already in place, step one is shown as done and Restart is live. When it is
not, step one is open and Restart is disabled until it completes — which it
does by itself, because the poll re-probes while the answer is still no, and
returning from System Settings is the moment that lands. An unanswered probe
never accuses the user of a missing grant; it leaves both steps open.

Restart calls the management API directly rather than stacking the Manage
Sessions confirmation on top, since the dialog already states the consequence.
Success replaces the steps with a done line and takes the toast down; failure
says so inline and leaves the button usable.

System Settings opens through the existing developer-permissions pane opener,
which takes an id rather than a URL, with Files and Folders added to it. The
event's action enum now also counts fix_opened, settings_opened,
restart_clicked and — emitted from main when a replacement daemon's first spawn
lands in the folder class the previous one was denied on — whether the restart
actually worked.

* fix(macos): let the folder-access notice return after a poll that read no daemon

A daemon identity reads as null during any reconnect blip, and the poll reports that as
"no mismatch". The notice dismissed itself and then never showed again for that daemon,
because the once-per-daemon latch still held its scope. Only "Not now" should latch.

* fix(macos): say what the folder-access notice costs the user

One line read like a stray warning. The toast now says who is blocked and what fails,
and still leaves the steps to the fix dialog.

* fix(macos): give the folder-access toast one action and the X, like every other toast

"Fix" is the only button; the X dismisses. Sonner fires onDismiss for programmatic
dismissals too, so the post-restart takedown now goes through the store and the hook,
and only a user's X is counted as dismissed.

* fix(macos): keep the fix dialog's steps a checklist and put the one action in the footer

Buttons inside each step made the list look like a form, and a footer Close duplicated
the X. The footer now carries the active step's action, with a ghost Cancel; a probe
that could not answer says so under step 1 instead of showing a check.

* fix(macos): let the checklist show the fix landed instead of saying so

A hedged sentence addressed to the user read like chat. On success both steps check
off and the footer offers Done; the unanswered-probe helper is a status, not advice.

* chore(i18n): drop the fix dialog's unused close key

* Revert "chore(i18n): drop the fix dialog's unused close key"

This reverts commit 365915df48.

* chore(i18n): drop the fix dialog's unused close key

* fix(macos): tell step 1 what to do when the folder toggle is already on

Users who need step 1 usually find Orca already allowed in System Settings; the grant
is recorded but not honoured for the daemon. Re-toggling re-records it.

* fix(macos): drop the unverified toggle instruction from step 1

Nothing has been confirmed to fix a grant that is already on, so the step says only
what the probe knows.

* refactor(macos): share the tccutil reset and bundle-id read behind one module

Clearing a macOS TCC row is about to have a second caller: the daemon
folder-access fix (STA-7948) needs the exact `tccutil reset` the computer-use
helper already issues. Extract both it and the PlistBuddy bundle-id read into
src/main/macos-tcc-reset.ts so the two remedies cannot drift apart.

The extracted calls go through runProcessSync rather than a fresh
node:child_process import: the spawn chokepoint's ratchet holds the direct
importer count at a pin, and a new module with its own spawnSync would raise it.
Behaviour is unchanged except that both calls now carry a 10s bound, and the
computer-use test asserts the same argv against the chokepoint's options.

* feat(macos): offer a permission reset when restarting the terminal service cannot help

About a third of the users who see the folder-access notice are still denied by
a freshly forked daemon even though Orca itself is allowed under Files and
Folders, so the restart the dialog offers cannot fix anything for them. That
state previously had one action: open System Settings, where the toggle they
would look for is already on.

The denied state now offers "Reset permission". Main clears Orca's TCC row for
that folder class with tccutil, then reads the folder from the app itself so
macOS raises its prompt against Orca rather than the daemon, then forces a
fresh-daemon re-probe that bypasses the poll's reuse interval. The dialog
re-renders from that verdict: allowed turns step one green and offers Restart,
still denied says so, and a refused reset points back at System Settings.

Nobody has confirmed this remedy on an affected machine, which is why main emits
the re-probe's verdict as reset_outcome_allowed/still_denied/unknown. Those
three, plus reset_clicked, are the evidence that decides whether the feature
stays.

* fix(macos): say what the permission reset does, and keep System Settings as the fallback

Step 1 was labelled like a Settings task while the button did something else, with two routes
in the footer for one step. The denied state now names the step for what the reset does,
explains it under the step, and shows System Settings only after a reset fails or leaves
things blocked.

* fix(macos): count a folder-access restart only against evidence that survived

The stored denial is the prior denial, so a second copy of it outlived the
one event that retires it: a daemon that read its own folder back cleared the
entry but left the copy, and the next daemon's first denial was then reported
as a restart that had never happened.

Track the outcome on the entry itself, drop the spawn-path probe (ten denied
terminals forked ten probe children the focus-time poll re-runs anyway), and
stop emitting `shown` from a getter the reset path calls for data. The
renderer's toast latch is what decides a scope is shown, so it emits it.

Both accessors now read one identity-matched entry.

* refactor(macos): name the folder-access verdict instead of encoding it as a tri-state

`restartWillHelp: boolean | null` re-encoded a verdict the probe already
returns as a named union, so every reader had to remember that `false` meant
"Orca itself must be re-allowed" and `null` meant "no answer".

`freshDaemonAccess: 'allowed' | 'denied' | 'unknown'` says it, end to end
through main, the IPC payload, the preload mirror and the dialog. The reset's
outcome event becomes a lookup. No user-visible string changes.

* refactor(macos): give the folder-access notice one latch instead of three

Two refs in the hook and a field in the store tracked the same fact, and the
dialog reached the hook through a store field plus an effect just to take its
own toast down before sonner echoed the dismissal back.

The store now holds the visible scope and the scopes the user closed, and
exposes the three things that happen to a notice: it is shown, someone else
retires it, or the user dismisses it. The dialog calls retire directly and the
effect is gone. `settingsIsFallback` loses an argument that was always true at
its only call site, so it becomes the local it always was.

* refactor(macos): stop blocking main on the tccutil reset

Two spawnSync calls with a ten-second timeout sat inside an async IPC handler,
so clearing a TCC row held main's event loop for as long as either binary took.

Both now run through runProcess. The computer-use caller that shared them was
already async, so it awaits them.

* test(macos): run the folder-access probe script against real paths

Every other test mocks the spawn away, so the minified child script — the one
piece that duplicates enumerateDirectoryOnce's errno mapping — had no oracle.
It now runs against a temp directory, an absent path, a file, and a directory
whose mode withholds it, which is skipped for root and on Windows.

* refactor(macos): read the folder-access entry through one identity match

All four callers that ask "is this evidence still this daemon's?" now go
through the same private accessor, so the rule the canonical path depends on
lives in one place.

* fix(macos): keep folder evidence through a failed health read, and make a forced re-probe always probe

A rejected attribution-health read nulled the folder evidence on the same poll, which the
renderer read as "cleared". A forced refresh after a reset returned early on an older
settled verdict. The dialog also closes when a reset finds the evidence gone, and stops
showing the unverified helper once the restart is done.

* fix(macos): name the folder in the access-notice scope

One daemon denied two protected folders kept one scope, so the toast, the
fix dialog, and the tccutil reset could each be about a different folder.

* refactor(macos): derive the folder-access dialog from the latest verdict

The store held an `open` flag and a mismatch frozen at the moment the toast
was raised, so the dialog could open on a stale verdict and its remedy state
could survive a close. It now keeps the latest verdict and the scope the user
opened, and the dialog is shown only while the two agree.

* fix(macos): offer the permission reset only where there is a row to reset

A workspace symlinked out of Documents or on an external volume can be denied
too, and the dialog offered a reset that main refuses. One shared list of the
TCC-backed folder classes now decides both.

* fix(macos): give the permission prompt's read a deadline

An unanswered macOS sheet blocks the app's folder read for as long as the user
ignores it, and the fix dialog is modal and busy until that read returns. The
wait now ends after a minute and reports an unknown outcome rather than
probing under the sheet.

* fix(macos): count the folder-access notice once per scope

A reconnect blip reports no daemon, which takes the toast down and lets the
same scope raise it again. Both raises counted as separate notices, inflating
the denominator behind the affected-user rate. The two latches are now one
map from scope to phase, and the count follows first insertion.

* fix(macos): drop the restart warning once the restart is done

Step two ticked green while its helper still warned that open terminals and
agents would restart, which had already happened.

* fix(macos): keep the folder-access toast up when the fix dialog opens

Sonner deletes a toast after its action button runs unless the handler
prevents the event, and it does so without calling onDismiss. Clicking Fix
therefore took the notice off screen while the scope stayed latched as
visible, so cancelling the dialog left no way back to it.

* refactor(preload): reuse the shared daemon cwd class instead of copying it

The five folder classes were hand-mirrored in preload behind a comment saying
preload cannot depend on main-only modules. The enum lives in src/shared,
which preload already imports from elsewhere, so the copy could drift.

* refactor(macos): close the fix dialog when its evidence disappears

A null verdict left the opened scope set, so the same scope coming back
remounted a checklist nobody had opened. Clearing it on a null verdict also
makes the dialog's scope key redundant, so it goes.

* fix(macos): let each fix-dialog button report its own work

The footer swaps the reset for a restart as soon as a poll says the grant
landed, which can happen while the reset is still running. Both buttons read
their label off the dialog being busy at all, so the restart button appeared
spinning as "Restarting…" for a restart nobody had started.

* fix(macos): clear the reset failure once the permission is granted

"Couldn't reset the permission" stayed on screen after the user granted it in
System Settings and the probe read allowed, contradicting the ticked step
above it. Its sibling line was already gated on the same verdict.

* fix(macos): end the folder-access remedy with the evidence it is about

Two ways out were missing. A reset that cleared the evidence closed the dialog
but left the toast on screen, because only the poll retired it; the store now
retires the notice whenever a verdict comes back null, so both callers get it
and the hook's own branch goes. And the opened scope survived a verdict for a
different scope, so the original one returning later reopened the dialog with
nobody having asked for it.

* fix(macos): keep the folder prompt off main's spawn path

The app-side readability check moved from accessSync to opendir when the
notice was added. TCC lets accessSync through but gates opendir, so on a
machine that has never granted Orca the folder, spawning a terminal there
raised the macOS sheet and froze main until the user answered it. The read is
async now and the spawn no longer waits for it. The blocking variant keeps a
name that says so, and the reset module's own copy of the read is gone.

* fix(macos): only say a folder is still blocked when something re-read it

Two paths reached "Still blocked after the reset." with no verdict behind it:
an unanswered prompt, where the reset returns the verdict stored before it
ran, and a re-probe that could not answer. The reset now returns the same
access it reports to telemetry, and the line waits for a real denial.

* refactor(macos): let the folder-access refresh decide when to skip itself

The poll handler re-implemented the refresh's own two guards, a null entry
and a settled allowed verdict, so each had to be kept in step by hand.

* fix(macos): stop the daemon blocking on its own folder read

The daemon reads the requested cwd before forking a shell to report whether
it can list it. That read is the one macOS gates, so on a folder the daemon
is refused it could hold the daemon's event loop behind a prompt. It is
awaited now, which leaves the blocking enumerator with no callers.
2026-09-22 01:09:26 -04:00
88f2f01061 fix(daemon): escape the terminal daemon into its own systemd scope so a service restart no longer kills every live PTY (#19430)
* fix(daemon): escape the terminal daemon into its own systemd scope so a service restart no longer kills every live PTY

Root cause: daemon-launched-child.ts forks the detached terminal daemon with
detached: true, which escapes the POSIX process group (setsid) but never the
systemd cgroup. Every PTY the daemon owns is itself an undetached direct
child of the daemon (native-pty-spawn.ts). Under a combined systemd unit
(Type=simple, KillMode=mixed, per docs/reference/headless-linux-server.md),
a systemctl restart/stop SIGKILLs every process still in the cgroup at the
stop timeout -- the daemon and every live terminal -- even though the
codebase already has a fully-built adoption/reattachment path for a
surviving daemon (orcad-entry.ts's refreshRestoredOrchestrationAuthority +
reconcileLegacyWorkerTerminals, gated on daemonOwnsFreshPersistentPtys()).
That path never fires today because the daemon never survives long enough.

Fix: when systemd is actually supervising the process and the OS user has a
reachable systemd --user manager (isDurableDaemonScopeSupported(), Linux
only), launch the daemon via systemd-run --user --scope so it lands in a
cgroup that is a sibling of the service unit's cgroup, not a descendant of
it. A systemctl restart of the combined unit then never reaches it. Any
failure of the scoped launch (no reachable bus, D-Bus policy rejection,
etc.) falls back transparently to the existing plain fork() launch, so
every platform/environment without this capability is unaffected.

The daemon self-detects its own resulting cgroup scope via /proc/self/cgroup
(detectOwnCgroupScopeUnit()) rather than trusting the launcher's intent, and
publishes it as cgroupUnit in its pid record and orcad's health/readiness
payload (health.terminalDaemon.cgroupUnit), so a running deployment can be
observed to confirm the fix actually engaged.

No new session registry is added: the existing daemon pid-record + adoption
protocol (publishDaemonPidFile, daemon-pid-record-quarantine.ts's
dead-record reclaim, refreshRestoredOrchestrationAuthority) already
implements durable, crash-safe reattachment for a surviving daemon -- it
was simply never exercised against a full unit restart before now.

Proven via a systemd-in-Docker recovery test: a live PTY session's shell
process, its daemon, and the daemon's cgroup scope were all confirmed
unchanged across a real systemctl restart of a Type=simple/KillMode=mixed
unit, while the main process pid changed (confirming the unit actually
restarted) and the new process's health payload recognized the surviving
daemon as adopted and live. A fresh write into the same PTY post-restart
reached the same running shell. Ordinary terminal create/work/release and
the #18789/#18790 worker-release reap-fix regression tests are unaffected.

Fixes stablyai/orca#19408

* fix(daemon): probe the real per-UID XDG_RUNTIME_DIR before trusting the process's own env

isDurableDaemonScopeSupported()/buildDurableDaemonScopeCommand() trusted the current
process's own XDG_RUNTIME_DIR env var first, falling back to /run/user/<uid> only when
that var was unset entirely. On mtl-02, orca-serve@factory.service's RuntimeDirectory=
hardening directive makes systemd export XDG_RUNTIME_DIR=/run/orca_serve/factory into the
unit's process -- a private scratch dir that shares the env var's name but has nothing to
do with the user session bus. /proc/<pid>/environ on that host confirmed exactly that path
plus DBUS_SESSION_BUS_ADDRESS=disabled:, while the real bus was reachable the whole time at
/run/user/985 (confirmed via systemctl --user is-system-running with that dir exported by
hand). The probe treated the hardened override as authoritative, found no bus socket there,
and reported unsupported on every launch -- so the cgroup-escape fix from #19408/#19430
never actually engaged on real hardware, even though tonight's factory deployment picked it
up.

Fix: resolveUserRuntimeDir() now always tries the conventional /run/user/<uid> path first
(computed independently via getuid(), never trusted from env), checking for a genuinely
connectable bus socket via statSync(...).isSocket() rather than a bare existsSync. It falls
back to the process's own XDG_RUNTIME_DIR only when that canonical path has no reachable
bus -- covering hosts that legitimately have no /run/user/<uid> at all but do have a
working bus wherever their own environment points. buildDurableDaemonScopeCommand() now
explicitly sets XDG_RUNTIME_DIR to whichever path this resolution picked, rather than
inheriting the spread env's (possibly hardened-wrong) value.

Both isDurableDaemonScopeSupported() and buildDurableDaemonScopeCommand() gained an
injectable canonicalRuntimeDir parameter (defaulting to the real computed path) so tests
can exercise the hardened-override scenario deterministically with a real, connectable
AF_UNIX socket fixture instead of the live host's actual runtime directory.

Docker's stock jrei/systemd-ubuntu test container never had this hardening directive, so
this gap was structurally invisible to the container-based verification in #19430 -- only
caught against real mtl-02 hardware.

* fix(daemon): report the daemon's own pid over the ready handshake, not systemd-run's

The launcher used to infer the daemon's identity pid from the immediate
spawned child (`child.pid`). On the durable-scope path that child is
`systemd-run --user --scope`, not the daemon, so the launcher was asserting
an identity it had no authority over.

`DaemonReadyIdentity` now carries a required `pid` populated from
`process.pid` inside the daemon itself, and `daemon-launched-child.ts` takes
`launchedIdentity.pid` from that self-report. Both sides of the
`holdDaemonAdoptionLease` pid comparison therefore originate inside the
daemon process, which is the idiom this branch already uses for cgroup
membership (`detectOwnCgroupScopeUnit` reads `/proc/self/cgroup` rather than
trusting what the launcher intended).

Note on the reported consequence: `systemd-run --scope` registers its *own*
pid on the transient scope unit and then `execvpe()`s the target command --
same pid, no intermediate process -- so adoption did not in fact fail on
systemd >= 206 (verified against systemd 255.4-1ubuntu8.17 and current main,
`src/run/run.c` `start_transient_scope()`). The fix stands on its own merits:
it removes a silent dependency on that exec-vs-fork implementation detail,
which a `systemd-run` shim earlier in PATH or any future systemd change would
have broken with no diagnostic.

`terminateLaunchedDaemonChild` was audited and deliberately left on
`child.pid`: for the same execve-preserves-pid reason that pid is either
still systemd-run mid-scope-setup (killing it correctly aborts the launch) or
already the daemon, so it targets the right process either way.

Regression coverage: `daemon-launched-child-identity.test.ts` pins the
identity source, and `daemon-ready-identity.test.ts` gains pid-validation
cases. Ready-message fixtures across the `daemon-init-*` suites were updated
for the now-mandatory field.

Addresses:
https://github.com/stablyai/orca/pull/19430#discussion_r3953722704
https://github.com/stablyai/orca/pull/19430#discussion_r3954346518

* test(daemon): assert cgroupUnit in the pid-file parse contract

`parseDaemonPidFile` returns `cgroupUnit` on every branch as of the
durable-scope commit on this branch, but five exhaustive `toEqual`
assertions in daemon-health.test.ts still described the pre-scope shape, so
they failed on the branch independently of any later change.

Adds the field to those expectations. Deliberately not relaxed to
`toMatchObject`: asserting the full parsed shape is what makes these tests
catch a field silently dropped from the pid-file contract.

* refactor(daemon): resolve the canonical user runtime dir at one point

The per-UID path cannot change for a live process, so compute it once into a module
const instead of threading the same default call through three signatures, and drop
the try/catch around a getuid() that cannot throw once it exists. Trims the module
prose to the non-obvious facts and corrects the pid-file record comment: an unscoped
daemon writes null; only records no daemon wrote are absent.

* test(daemon): clean up the cgroup-scope fixtures and assert a verdict

The cgroup fixture tracked only the file it wrote, leaking one temp dir per case.
Drains both fixture lists with splice so the pop-may-be-undefined guards go away,
and replaces a not-throw/typeof-boolean pair with the verdict it was circling:
no resolvable runtime dir means unsupported.

* refactor(daemon): share the detached child options across both launch paths

cwd, detached and stdio were repeated in the fork and systemd-run branches, which
left the two comments explaining them hovering over the env block instead. Names
them once so each branch carries only its own delta.

* refactor(daemon): validate the ready pid like every other field

typeof-first narrows the value, so the two 'as number' casts the isSafeInteger check
needed disappear and the pid guard reads like the startedAtMs guard below it.

* fix(daemon): don't retry the launch unscoped after losing the endpoint race

A scoped attempt that lost the endpoint to another daemon was retried unscoped: a
second doomed fork, a misleading 'cgroup-scope launch failed' warning, and the same
DaemonEndpointUnavailableError the caller was already going to adopt on. Rethrows it
instead, since no launch mode can win a race that is already lost.

Also drops a private alias for DaemonChildSpawnOptions and the two 'as number' casts
on child.pid in the startup-failure cleanup.

* fix(daemon): unlink the pid record by the pid the daemon published

The record holds the daemon's self-reported pid, so match on that rather than on the
immediate child's, which is the systemd-run wrapper's until it execs.

* fix(daemon): route the scope launch through the child-process chokepoint

The two files this PR added imported `node:child_process` directly, which
`child-process-import-boundary.test.ts` fails on deterministically: the
offender count went 155 -> 157 against a pin of exactly 155. Raising the pin
or listing the files is what that test explicitly forbids, and the allowlist's
own note says a split "moved the import, it did not add one" -- so the fix is
to get both new files off the module and put the count back at 155.

- `daemon-cgroup-scope.ts`: the `systemd-run --version` probe now uses
  `runProcessSync` instead of `execFileSync`, so it gets the shared spawn
  decisions. Kept synchronous deliberately: `launchDaemonChild` attaches the
  readiness listener in the same tick it is called, and an await before the
  spawn moves the child past that tick. A non-zero exit is data rather than a
  throw here, so the verdict now checks `code === 0 && !timedOut`.
- `daemon-launched-child-spawn.ts`: the scoped launch uses `spawnProcess`, and
  the long-standing unscoped launch keeps `fork` semantics through a new
  `forkProcess`.
- `src/shared/child-process/fork-process.ts`: the fork arm of the chokepoint.
  `spawnProcess` cannot express a Node child with an IPC channel started from
  a module path under an overridden `execPath`, and the existing launch tests
  are written against `fork`'s contract, so a spawn rewrite would have changed
  module resolution, `execPath` and `execArgv` at once. It passes
  `windowsHide: true` -- the flag every other call site in that directory
  sets, reachable via an assertion because `ForkOptions` omits it -- which
  keeps `windows-console-visibility.test.ts` at its pin of 65 too.

Both ratchets pass with both pins and both allowlists untouched.

Docs: `orcad-operations.md` and `headless-linux-server.md` still described the
limitation this PR removes as permanent. Both now describe the durable-scope
survival path and its preconditions (systemd as PID 1, a reachable user bus /
`loginctl enable-linger`, `systemd-run` on PATH), and scope the old text to
the unscoped-fallback case, pointing at `health.terminalDaemon.cgroupUnit` as
the way to tell the two apart on a running host.

* fix(daemon): seal the cgroup capability probe from the host and correct KillMode=mixed docs

The capability probe consulted the host's own /run/systemd/system marker and
spawned the real systemd-run binary, so the hermetic unit tests could only pass
on a systemd host (and fail closed otherwise, even with faked bus sockets).

- Thread systemdBootPath and runVersionProbe as test seams through
  isDurableDaemonScopeSupported, defaulting to the real boot marker and
  systemd-run --version probe in production.
- Narrow the injected probe to the ProcessResult slice it consumes.
- Cover: no-systemd-boot, non-zero probe exit, and probe-timeout cases.
- Correct KillMode=mixed semantics in the docs: the cgroup-wide SIGKILL fires
  the instant the main process exits, not after TimeoutStopSec; document the
  Docker-container caveat and add KillMode=mixed to the multi-service template.

* fix(daemon): satisfy assertion checks in scoped launch

* fix(daemon): satisfy anti-slop and console guards

* test(serve): update shutdown docs assertions for daemon scope

* fix(daemon): migrate adopted legacy scopes

* docs: qualify restart safety by daemon scope

* docs(daemon): qualify Upgrade restart prose with durable scope caveat

Align the Upgrade section in docs/reference/headless-linux-server.md with
the earlier preservation section and docs/reference/orcad-operations.md:
a service restart terminates live processes only when running under the
unscoped fallback, and stops should be treated as destructive unless
health.terminalDaemon.cgroupUnit names an orca-daemon-*.scope.

Update the shutdown workflow test assertion in
config/scripts/headless-serve-shutdown-workflow.test.mjs to match.

* fix(daemon): harden legacy scope migration

---------

Co-authored-by: Lesley Murfin <260182349+LesleyMurfin@users.noreply.github.com>
Co-authored-by: m4air <m4air@Mac.localdomain>
2026-09-21 17:23:30 -07:00
Neil 4feaaf5c5c feat(terminal): configure interactive Unix shell arguments (#21904)
* feat(terminal): configure interactive Unix shell args

* fix(settings): clarify Unix shell argument defaults

* fix(settings): improve Unix shell argument guidance

* fix(settings): simplify shell argument guidance

* fix(settings): clarify empty shell args

* fix(settings): explain empty shell args

* feat(settings): make shell argument modes explicit

* fix(settings): keep no args inside custom mode

* fix(terminal): apply configured shell args on the renderer spawn path

The renderer's pty:spawn handler builds options in ipc/spawn-options, not
the runtime controller, so the configured profile never reached a terminal
pane. The local launch plan also dropped the args whenever shellOverride
was set -- which the spawn path always fills from terminalDefaultShell.

Both spawn paths now share one resolver.

* chore(i18n): allowlist the new terminal shell argument strings

Matches how the sibling Terminal shell settings strings are already handled.
2026-09-21 01:35:58 -07:00
OrcaWinandm4air 4085e1cf60 fix(memory): release stale session registries (#21734)
* fix(memory): bound session and lifecycle registries

* fix(memory): bound transient filesystem registries

* fix(memory): cap path and locale caches

* fix(memory): bound runtime recovery registries

* fix(memory): bound host mirror gap verdicts

* fix(memory): bound shell startup env cache

* fix(memory): bound gitlab host context cache

* fix(memory): release removed ssh generations

* fix(memory): expire cloud refresh replay guards

* fix(memory): release retired plugin generations

* fix(memory): bound plugin log key retention

* fix(memory): bound automation authority generations

* fix(memory): bound native chat enrichment cache

* fix(memory): bound web session tracking generations

* fix(memory): bound codex credential absence paths

* fix(memory): bound WSL canonical path cache

* fix(memory): bound sparse checkout cache

* fix(memory): bound shared directory cache

* fix(memory): bound advertised URL scan snapshots

* fix(memory): bound automation manager cache

* fix(memory): bound web session reorder intents

* fix(memory): bound web session focus intents

* fix(memory): bound web session handoffs

* fix(memory): bound automation dispatch tokens

* fix(memory): bound host mirror waiters

* fix(memory): bound retained session activity

* fix(memory): bound retained session activity

* fix(memory): bound web session close intents

* fix(memory): bound cloud session cache

* fix(memory): bound WSL home cache

* fix(memory): bound SSH capability cache

* fix(memory): bound trust grant cooldowns

* fix(memory): bound WSL auth drain state

* fix(memory): bound Linear workspace credential cache

* fix(memory): bound local Git capability cache

* fix(memory): bound WSL Git environment cache

* fix(memory): bound WSL Git environment cache

* fix(memory): bound WSL preflight cache

* fix(memory): keep hot cache entries warm

* fix(memory): preserve generation fences across eviction

* fix(memory): close remaining eviction fences

* fix(memory): align evicted upstream generations

* fix(memory): trim successful capability probes

* fix(auth): retain expired refresh replay evidence

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
2026-09-20 14:41:50 -07:00
84d827a6ab fix(daemon): pause producers when stream backlogs grow (#20947)
* fix(daemon): pause producers when stream backlogs grow

* fix(daemon): reset stream backpressure on socket replacement

* docs(daemon): point retention audit at current reproducer

* test(daemon): validate stream retention audit outcomes

* fix(daemon): bound the stream producer stall and leave a visible gap

Stream backpressure pauses a session's PTY with no deadline: the only
un-pause comes from the consumer draining, so a half-open peer that stops
reading without closing freezes the shell for the rest of the session.

Arm a 60s watchdog on the false->true stream-pause transition (not on the
re-assertions refresh() makes for neighbouring sessions). On fire, mark the
session stall-released: it becomes keep-tail droppable, its backlog is
thinned behind a dataGap, and the producer runs again. The existing dataGap
path makes the renderer restore that pane from the daemon's snapshot, so the
user sees the terminal jump to current rather than sit frozen. The mark
clears once the session's last byte leaves the daemon, restoring ordinary
pausing. Nothing here reports a process exit - loss of contact with a
consumer is not evidence about the child.

Also enable TCP keepalive on the stream socket so a genuinely dead peer
closes and onStreamDisconnected clears the pause.

* test(daemon): put each casting SAFETY: directive on one line

`oxlint-disable-next-line` covers only the line directly after it, so a
rationale wrapped onto a second comment line suppressed nothing and the
casts failed the changed-code quality gate. Drop the remaining JSON.parse
cast for an annotated binding.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Neil <neil@stably.ai>
2026-09-19 17:51:59 -07:00
09073086a8 feat(terminal): inline images via @xterm/addon-image (perf-first) (#19512)
* feat(terminal): inline images via @xterm/addon-image, perf-first

Add opt-in inline terminal images (SIXEL, iTerm2 IIP, Kitty graphics)
through @xterm/addon-image, designed to keep idle terminals unaffected.

Performance:
- The addon (base64-inlined wasm decoders + protocol handlers) loads off
  the boot critical path via a deferred loader that mirrors the WebGL
  addon: primed after first paint only when the setting is on, read back
  synchronously at attach, with a 3-attempt cap so a transient failure
  never disables images for the session and a missing chunk never
  refetches per pane. renderer-boot-graph guards against eager import.
- enableSizeReports:false so the addon never sets windowOptions and
  double-answers Orca's own CSI 14t/16t responder.
- Perf-tuned decode/storage limits (storageLimit, sixel/iip/kitty size
  caps) in one place.

Correctness:
- Orca's DA1 handler wins over the addon's (last-registered-first), and
  the default DA1 response never advertised Sixel (;4), so DA1-detecting
  tools (chafa, img2sixel, viu, timg) never emitted it. The winning
  handler now appends ;4 while the setting is on, resolved per query so a
  live toggle changes the next DA1; idempotent against the ConPTY
  response that already lists it.
- ORCA_IMAGE_PROTOCOL=kitty is exported to spawned shells (local, daemon,
  relay/SSH) and forwarded across the WSL boundary, so image-capable
  agents can pick an encoder. Unknown image sequences are swallowed by
  xterm when the addon is detached, so this never garbles output.
- Settings toggle (default on) gates rendering and DA1 advertisement.

Cross-checked against community PRs #7775, #11706, and #19201 at the end;
credited below.

Co-authored-by: s546126 <s546126@users.noreply.github.com>
Co-authored-by: XRX193 <XRX193@users.noreply.github.com>
Co-authored-by: lmsh7 <lmsh7@users.noreply.github.com>

* fix(terminal): bound inline image memory and classify Kitty replies

* fix(terminal): bound image decode and release image resources on cleanup

* fix(terminal): address image addon review feedback

* test(terminal): stub setPaneInlineImagesEnabled in appearance manager fakes

* fix(terminal): evict unplaced kitty payloads before displayed images

Byte-budget eviction dropped the oldest transmitted blob regardless of
placement, so a new upload could erase a visible image while abandoned
blobs still held budget. Unplaced payloads now go first and displayed
ones only when that is not enough. The incoming image is always stored,
so an oversized one overshoots the cap by one payload instead of being
dropped after the protocol already acked OK.

* fix(terminal): gate DA1 Sixel on real addon attachment; claim SSH image spec in CI

- DA1 advertised Sixel from the setting alone, so a pane whose lazy addon
  chunk was still loading (or had failed all three attempts) told
  feature-detecting tools to emit DCS that nothing could render. Track the
  attached decoder per terminal and require it before setting the ;4 bit.
- tests/e2e/terminal-inline-images-ssh.spec.ts was Docker-gated but claimed
  by no lane runner, so pr-e2e-gate-contract failed and the spec would have
  self-skipped green forever.
- Reject non-positive PNG IHDR dimensions before decode: they are parsed with
  signed shifts, so a dimension >= 0x80000000 came back negative and slipped
  past the pixel-limit comparison.
- One resolveTerminalInlineImagesEnabled() for the default-on setting; the
  four call sites mixed '?? true' with '!== false', which disagree on null.
- One readInlineImageResources() walk of the addon internals instead of two
  copies that could drift against the patched dependency.
- Isolate the deferred-attach drain per pane; make the zoom-invariance and
  backing-storage e2e assertions fail when the feature is dead.

* refactor(terminal): one lazy xterm addon loader for webgl and image

terminal-image-addon-loader was a structural clone of the webgl one — same
memo, attempt cap, and .then(ok,err)-clears-memo recovery. Both now wrap
createLazyXtermAddonLoader; each keeps its literal import() specifier so the
bundler still splits the chunk (verified against a fresh build: addon-image
stays out of the boot graph).

* refactor(terminal): name openTerminal's addon flags; pin image addon limits

Two adjacent optional booleans could be swapped without a type error once
inline images added the second one.

* docs(terminal): state the real per-pane image ceiling; drop test ordering dependency

storageLimit:32 reads like the pane's budget but keys three pools — decoded
pixels, retained encoded Kitty blobs, and pending WASM decoders — so the worst
case is ~98 MB per pane with no cross-pane governor. Say so at the constant.

pane-inline-images.test.ts's deferred case needed to run first; it now takes a
fresh module instead, and the rest prime in beforeAll. Verified by running the
file with that test moved last.

* fix(terminal): satisfy rebased static analysis gate

* fix(terminal): complete casting gate cleanup

* fix(terminal): recover failed image addon loads

* fix(terminal): bound image decoder allocations

---------

Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
Co-authored-by: s546126 <s546126@users.noreply.github.com>
Co-authored-by: XRX193 <XRX193@users.noreply.github.com>
Co-authored-by: lmsh7 <lmsh7@users.noreply.github.com>
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
Co-authored-by: Neil <neil@stably.ai>
2026-09-18 16:32:49 -07:00
Neil 1fa6fac17c fix(daemon): answer the per-pty snapshot predicate for the pty it was asked about (#21381)
canProvideAuthoritativeBufferSnapshot is contracted as "whether this exact PTY can
return a sequence-safe provider snapshot" (pty-provider-contract.ts), and two of the
three layers already route it per id: DaemonPtyRouter forwards to adapterFor(id), and
DegradedDaemonPtyProvider forwards to the provider that owns the session. The daemon
adapter was the leaf that discarded the id and returned supportsAuthoritativeBufferSnapshots
— a negotiated protocol version, which is a fact about the connection, not about a pty.

That is reachable, not theoretical. getProviderForPty falls back to the local provider
for any id it cannot place, so a remote-runtime id (whose pty lives on another machine)
resolves to the local daemon adapter, and pty:getAuthoritativeBufferSnapshotCapabilities
answered `true` for a session this daemon has never owned. The renderer caches that as a
definitive per-pty verdict, and because the leaf discarded the id it could not tell it had
been asked about something it does not own.

Today the wrong answer is masked: allowOrdinaryParkRestore short-circuits remote and SSH
ptys before the cached verdict is read, so nothing consults it. This closes the gap before
something relies on it — a caller reaching for a per-pty answer should not be handed a
confident one that is wrong.

Not touching that short-circuit. It is deliberate: SSH bytes transit the client's own main
process into its headless mirror, so those panes have a local copy the predicate says
nothing about, and the direct-SSH lane was confirmed to repaint from a daemon-backed
restore with the park capture disabled entirely. Routing SSH around a daemon-snapshot
predicate is correct, and removing the short-circuit would disable SSH parking for no
correctness gain.

The existing protocol-compatibility test asserted `true` for a made-up session id, which
encoded the bug. It now spawns a real session, so it still proves the protocol-version
gate without depending on an unowned id reading as supported.
2026-09-17 23:21:39 -07:00
Neil 691d9692e6 fix(pty): stop detached OMP tools on immediate terminal close (#20642)
* test(omp): add opt-in owned PTY closure probe

* fix(pty): sweep detached tools on immediate unrecognized shell close

* test(omp): create close probe evidence root in fresh worktrees

* test(pty): account for asynchronous immediate descendant cleanup

* test(pty): reject inconclusive descendant cleanup probes
2026-09-17 20:58:02 -07:00
d04b05b5c8 Detach retained CI and terminal tails from oversized strings (#20960)
* fix(memory): detach retained CI and terminal tails from oversized strings

* fix(terminal): detach retained error and reattach string slices

* fix(terminal): release oversized recent-output backing strings

* fix(terminal): release backing strings held by PTY detectors

* fix(memory): own bounded Claude background task labels

* fix: detach retained terminal mode scan tails

* fix: own retained plugin worker output strings

* fix: own incomplete OSC 133 carry strings

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
2026-09-17 20:34:19 -07:00
OrcaWinandm4air 98998b18ad fix: release retired shared daemon owner metadata (#21162)
Co-authored-by: m4air <m4air@Mac.localdomain>
2026-09-17 20:12:31 -07:00