Commit Graph
11997 Commits
Author SHA1 Message Date
Neil e6fbbdf684 perf(ci): cache pnpm verification records on Linux (#23568)
* perf(ci): pilot pnpm verification record caching on Linux

* test(ci): review pnpm verification record in mobile cache audit
2026-09-28 01:50:30 -07:00
Neil 3dd7d29455 Revert "perf(ci): shard the anti-slop audit across processes instead of one JS runtime (#23543)" (#23575)
This reverts commit fae0ae7a46.
2026-09-28 01:47:07 -07:00
Jinwoo Hong a174a086d7 fix(runtime): answer terminal.subscribe at once for a pane the desktop already has mounted (#23512)
* fix(runtime): answer terminal.subscribe at once for a pane the desktop already has mounted

A mobile subscribe to a PTY with no headless model asked the renderer to
mount its tab and waited for a newer serializer settle. The renderer drops
mount requests for tabs it already has mounted, so a reattached daemon PTY
whose restored provider snapshot outranked the live renderer held the reply
for the full 3 s deadline. A live renderer screen now proves attachment and
is adopted directly; an unmounted pane still requests the mount and waits.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(runtime): treat any renderer answer, even a blank screen, as an attached pane

A fresh shell that has printed nothing has a registered serializer and an
empty screen; requiring non-empty data sent it back through the dropped
mount request and the 3 s wait. A blank screen skips the wait but does not
replace the chosen snapshot, so a parked pane cannot erase provider history.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(runtime): treat a mounted pane with unsettled output as attached

The stable renderer snapshot returned null both when no renderer answered
and when output advanced under every retry, so a desktop pane printing
continuously still took the dropped mount request and the 3 s wait. It now
returns a typed outcome (settled, moving, absent); moving skips the mount
and publishes the chosen snapshot, and late recovery still requires settled.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(runtime): adopt only a renderer-ordered screen when probing a mounted pane

The attachment probe read the terminal before knowing it would adopt, which
can reach the provider snapshot on the unmounted path; the read now follows
the decision. A seq-less renderer screen would replay every buffered chunk
on top of itself, so the probe keeps the chosen snapshot for it. The probe,
adopt and mount wait move into their own module to stay under the line cap.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(runtime): adopt a seq-less renderer screen when no output is pending

The seq gate only prevents a double replay of buffered output, so a settled
non-blank screen without a seq is safe when nothing is pending. That keeps
the better screen for a pane right after a deferred cold restore, before it
is renderer-ordered. The rule now applies after the mount wait as well.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(runtime): decide renderer attachment from the host's serializer flag

The mounted-pane probe re-derived attachment by serializing the renderer up to six
times, which cost ~7.5 s for a registered but unresponsive renderer on a busy PTY.
The host already holds that fact in the serializer readiness flag. The flag is never
cleared when a pane closes over a live PTY, so one null serializer answer falls back
to the mount wait: worst case is the old 3 s plus one 750 ms serialize.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(runtime): let the serializer's answer alone prove renderer attachment

serializeRendererTerminalBuffer already answers null when the host's serializer flag is
unset, so the separate flag accessor was redundant. The numeric-seq adopt test now
replays only the byte past the seam.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-28 04:28:46 -04:00
Neil 8db968451f fix(sidebar): clip worktree card content to its border (#23566) 2026-09-28 01:26:22 -07:00
Neil 1bfbc155d1 fix(agent-session): wait for in-flight session-store writes before teardown returns (#23545)
* fix(agent-session): stop lease renewal before the renewal's write lands

Clearing the renewal interval only cancelled the next tick. A tick already
past its guard still had a whole-file store transaction to commit, and the
store's transaction lock re-creates the store directory before it writes, so
that commit could land after host teardown had finished releasing everything
it touches.

`stop()` now resolves once the tick in flight has finished writing, and host
teardown's stop-lease-renewal phase waits for it. The three test harnesses
that model a host vanishing without a clean quit shared a copy of the same
incomplete shutdown; they now share one helper that waits.

The symptom was a CI flake: the refusal-oracle spec removes its temp
directory in `afterEach`, and a renewal landing mid-removal put the store
directory back, so the removal failed with ENOTEMPTY on the temp root.

* fix(agent-session): wait for the delivery loop's restart when abandoning a host

The abandon helper disposed the delivery loop and moved on. Disposing only
stops the loop's NEXT step: a step already past that check keeps going, and
the restart it runs for an accepted send reserves an owner, which is a store
commit. The store re-creates its own directory before every commit, so that
commit put the directory back under the temp-directory removal the test does
next, and the removal failed with ENOTEMPTY.

Quit already waits for exactly this work, in its drain-attaches phase — every
attach is registered with the task queue from enqueue. The helper now runs
the same drain, in quit's order, so it waits for both producers that reach
the store after the last awaited call returns.

Adds a regression test that holds the loop's restart inside its provider
acquisition and asserts abandoning does not return until it lands.

* chore: re-trigger PR checks

The push to 2ecc9c6e emitted no pull_request event, so the matrix never ran.
2026-09-28 01:08:45 -07:00
Jinwoo Hong ba858ee446 feat(mobile): the keyboard covers the page like a native screen, and the shell says its height (#23110)
The shell no longer shortens the WebView for the keyboard; it publishes the keyboard height like the safe-area insets, so native's keyboard lift, refit hold and dismiss key run on the page unchanged. Keyboard and inset arithmetic read the shell's OS through a host-os seam. One page-version floor (manifest pageVersion, shell floor 1) replaces per-feature accept negotiation; a page below the floor gets the existing update wall, a desktop with no bundle keeps native screens. iOS shell drops the form accessory bar and its own keyboard observers. Native session screens untouched.
2026-09-28 04:07:35 -04:00
Neil 080c562898 perf(ci): diff against the merge commit's first parent so PR checkouts can be shallow (#23562)
Every changed-path gate asked git for `--merge-base "$BASE_SHA" "$HEAD_SHA"`,
which needs the event payload's base SHA to be in the local graph. That is the
only reason two jobs cloned all 8127 refs' history. On a pull_request checkout
HEAD is already the merge commit, so its first parent is the base side and no
merge base has to be computed. config/scripts/git-pull-request-diff-base.mjs
resolved that for the two Node gates; the workflow's inline gates now use the
same helper through a small CLI rather than open-coding it.

code_paths gates all 22 jobs, so its checkout is charged to the start of every
one of them: measured 20.7s to 1.6s, keeping blob:none because its sparse tree
is ~7 files and leaves no blobs to refetch. Static analysis drops the filter
instead, since populating all 30,226 files makes blob:none force a second
promisor fetch: 23s to ~11s.

Verified on a real merge ref. At depth 50 the old and new forms produce
identical changed-file sets. At depth 2 the new form still works and the old one
fails with `fatal: bad object`, which is the failure a stale base would have
caused once the checkout stopped being complete.

Also drops the dead resolveBase + merge-base prelude in the changed-code gate,
whose result resolvePullRequestDiffBase already discarded on every PR.
2026-09-28 00:37:26 -07:00
Jinwoo Hong c8b3f2084f test(mobile): repin the RPC recording corpus to main after #23080 (#23565)
#23080 squash-merged a corpus pinned to its branch commit 486566c82b, which
the squash left unreachable from main. Repin baseline to main's tip and
re-record; every golden moves only its baseline header.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-28 03:30:28 -04:00
GuanBearandguanbear ca6c2091e5 feat(ai-vault): show ZCode CLI session history (#23513)
Surfaces ZCode CLI session history in AI Vault, so past ZCode sessions show up next to the other agents' instead of being invisible.

ZCode stores sessions in the same SQLite shape OpenCode uses, so this reuses the existing OpenCode lister and parser rather than adding a second scanner — the worker only varies the agent it stamps on each row. Discovery covers the native home and any WSL homes.

SQLite rows are narrowed at runtime rather than asserted: the statement API returns untyped column values, so the declared row shape is only a claim until something checks it, and a drifted schema or a database written by another tool reaches the same code.

Co-authored-by: guanbear <guanbear@users.noreply.github.com>
2026-09-28 00:28:36 -07:00
Jinwoo Hong 2077956254 fix(mobile): size a terminal's first subscribe from the document's reported cell box (#23080)
* fix(mobile): size a terminal's first subscribe from the document's reported cell box

#22960 sent phone dims on a terminal's first subscribe by opening a throwaway
empty terminal (init 80x24 ""), awaiting its ready and measuring, behind a
per-document first-subscribe mark whose lifetime was tied to web-ready. That
cost a second xterm/WebGL instance and ~150 ms per open, plus lifecycle state.

The document now measures the cell box without a terminal (xterm 6's
CharSizeService strategy, rounded as the renderer rounds it) for every
text-size preset and reports it with its viewport in web-ready; a table,
because the text scale only reaches the document after that notify. Each
init's ready reports the box xterm actually laid out, which replaces the
probe's entry. The controller answers fitDimensions/measureFitDimensions from
that table and the view's layout with no message; without a table it asks the
document as before.

The session seeds an unmeasured viewport synchronously in subscribeToTerminal,
so the first subscribe carries dims by construction. Deleted: the empty init,
its awaitReady gate, deferFirstSubscribeUntilViewportMeasured and the
subscribedDocuments mark. The fit pass is unchanged and still covers a
document that reports no cell box.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): read the reported cell box through in-narrowing, not Reflect.get

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): correct the probe's cell-box guess from the box xterm lays out

The web-ready probe is a guess: building the WebGL addon creates no context,
so a context that fails on load lands on the DOM renderer, whose width is not
snapped and depends on the column count. Before, a ready box that differed was
only logged; the first subscribe had carried the wrong column count, the host
echoed it, the fit pass saw the viewport equal to the host's dims, and the grid
stayed slightly shrunk. The store also kept the WebGL width after a context loss.

The document now reports the box xterm laid out whenever it changes (from
onRender, which covers a renderer swap and a DPR change that
onDimensionsChange does not fire for, and at ready). The store replaces the
guess; when that changes the current text size's entry, the view calls
onCellBoxChange with xterm's grid and the session re-fits, running the bounded
fit pass if the dims moved (one resubscribe). Equal boxes do nothing.

The RN layout box now survives a document reload; the document's own viewport
only stands in until the view reports a layout (on the page, web-ready arrives
first). The mismatch console.log is gone, and the probe's rounding names the
xterm version it copies.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): correct each cell-box guess at most once, so a DOM renderer cannot loop

On the DOM renderer the cell width is the rounded canvas width divided by the
column count, so every re-init at new cols reported a new box. Each one counted
as a correction, a floor over floats could flip the fit between two sizes, and
each flip landed converged, which reset the resubscribe budget: an unbounded
series of full-snapshot resubscribes.

Only the first laid-out box for a guessed text size may be a correction; later
reports still update the store, so fits stay truthful, but never resubscribe on
their own. The fit's floor gains a 1e-6 epsilon so floating-point error at an
exact boundary cannot flip a column or row. New document tests pin the render
report after a renderer swap and the report at ready for a paused renderer.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): make xterm the only terminal cell measurer

The document builds its real terminal before web-ready, at the app's
text scale, and reports the box xterm laid out; the first init reuses
that terminal. The page-side prediction, the per-scale guess table and
the once-per-document correction are gone. The app remembers the box
per text scale for its lifetime, so a later open at a known scale
subscribes with phone dims at once. A box that changes at the same grid
(renderer swap, pixel ratio) refits the open terminal in place.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): keep commands queued before the terminal WebView first loads

A subscribe sized from the stored cell box can queue init before the
native WebView reports its first load start, which cleared the queue
and left the terminal blank. Only a reload now drops queued commands.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): re-init a document that lacks the subscription's init, and fit one frame width

- Web-ready now says whether the document holds the terminal's latest init
  (a reload before the first ready drops a queued one); the session
  resubscribes any initialized terminal whose document lacks it.
- One grid fit, shared by the app and the document, fed the unrounded frame
  width React Native laid out; it keeps exact fits whole at fractional pixel
  ratios. The document's viewport-width fits are gone.
- The page builds every document at the scale the view mounted with, as the
  native WebView does.
- A new document's first cell box is compared against the grid the
  subscribe fitted from the stored box.
- The terminal built before ready stays hidden until its first init.
- The cell-box census matches glyph-measurement techniques, not names;
  the store's unused clear() is gone.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): build the terminal before ready only for the view shown at mount

A session mounts one terminal view per tab, and each built xterm and a
WebGL context before ready: 20 tabs made 20 contexts at load, past the
~16 a page (or Android's shared WebView renderer) holds, and native logged
32 context losses. Only the view shown when it mounts builds early now;
the rest build at their first init as before. Deferring the WebGL addon
instead would change the reported box: the DOM renderer lays out 7.8x15
where WebGL lays out 7.667x15 at the same font.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): write a WebView document's start values into its page, not an injected script

Android ran the pre-content injected script after the document's own in
1 of 22 documents on the emulator; that document started with no text
scale or shown flag and built a terminal it should not have. The values
now sit in the page ahead of the document script, one source object per
start pair so a render never reloads the WebView.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): the pre-ready terminal measures and reports while hidden

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): measure only the laid-out frame, and refit on a new grid, not a new width

- A measure needs both of the frame's dimensions from React Native; the
  document's viewport-height fallback is gone, and before the first layout
  the handle answers no fit without asking the document.
- A frame width change that still fits the PTY's grid from the stored box is
  a no-op, so sub-pixel layout jitter no longer re-measures. The width ref is
  written in that effect rather than during render (react-doctor).
- One "last grid" ref: the last reported grid, or the one a subscribe fitted
  from the stored box.
- The page render rig measures through the frame it laid out, as the session
  does, and lets the replay's fit settle before its resize-refit witness.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): let only the current terminal document's ready flush

A reload kept the WebView and its onMessage, so the old document's late
web-ready flushed the queue into the reloading view and the new document
got a second init. Each document now gets its own view (keyed on a
generation the controller owns), every notify carries the generation of
the view that received it, and a web-ready from a replaced document
flushes nothing and stamps nothing.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): drop every notify from a replaced terminal document

One rule at the receive boundary: a notify from any generation but the
current one is dropped, whatever its type, not only web-ready.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): make fitDimensions a pure question; name each generation counter

- fitDimensions no longer records the grid. A width change to a new grid
  asked it first, so the DOM renderer's report of that grid's box read as
  "same grid, new box" and refit again. Only the first-subscribe seed
  (seedFitDimensions) records the grid the document's first report is
  checked against.
- viewGeneration counts the views, readyGeneration counts web-readies.
- replaceDocument no longer resets the load flag; the load-start reset
  stays as the guard for a view that reloads itself.
- The name-based lifecycle census is replaced by a behavioural test.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): typecheck the handle mocks, drop the unused cell-box get

- The two handle mocks carry both fitDimensions and seedFitDimensions,
  and the fake-timer acts return nothing, so the three test files check
  under tsconfig.test.json again.
- terminalCellBoxes.get had no product caller; the store's tests assert
  through fit.
- The load-start comment says what the controller does now.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): hold the grid the document has, ignore a replaced view's load start, dispose a failed pre-ready terminal

- The document reports a new grid even with an unchanged box, so an
  in-place reflow on WebGL is held before a later renderer swap at that
  grid; the swap then refits. The app's apply paths do not hold the grid
  themselves: the DOM renderer's box follows cols, and a grid held on
  apply would read its own box as a renderer change and loop. One
  writer (holdGrid) holds the seeded or reported grid.
- A load start from a view a replacement unmounted is ignored, as its
  notifies already are.
- A terminal whose open throws before ready is disposed, not only
  unreferenced.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): ignore every native event from a replaced terminal view

One wrapper binds each WebView lifecycle event (load start, error, HTTP
error, render process gone, content process terminated) to the view's
generation, so a replaced view's late event cannot reset, replace or
put an error over the current document.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): a DOM seed refits once on its first report, not on the refit's own

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): subscribe a terminal only after its document is ready

The document still builds its terminal before ready and reports the cell
box xterm laid out in web-ready; the app now subscribes after that ready
and fits from that box, so nothing is sent to a document before it is
ready. Everything that made a pre-ready subscribe safe goes: the
app-lifetime box store, the seed fit, the per-document view generations
and their event filtering, the init tracker and the hasInit resubscribe.
The native view reloads in place again and web-ready keeps main's reload
rule. Boxes are kept per view; the grid a document last reported still
guards the in-place refit against the DOM renderer's cols-dependent box.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): hold one reported cell box and the grid the subscribe fitted

The controller keeps only the box the current document last reported,
not a per-text-size store: the document re-reports on a scale change.
The subscribe after ready fits from that box and holds the grid it
fitted, so the DOM renderer's first report at that grid (a new box)
refits once in place and converges; refit and apply paths hold nothing.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): fit only a ready box at the app's scale; forget a reloaded document's box and grid

A reload keeps the document's mount scale, so a ready after a text-size
change reports a box at the old scale; that box no longer sizes the first
subscribe, which then takes the no-box path. A readiness reset drops the
old document's box and held grid, so the new document's first DOM report
at the same grid does not refit.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): give the terminal document its frame at init, and fit text scale over it only

A subscribe sized from the ready box sends no measure, so the document
had no frame when the text size changed and reported the pre-refit row
pitch. The app's init now carries the frame it laid out, in the fields a
measure uses; the router takes it from either. The text-scale fit reads
only that frame, with no viewport fallback.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): say why a frameless text-scale change skips the resize

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): one cell box per terminal notify, not an array

web-ready and cell-metrics carry `cellBox: {fontScale, cellWidth,
cellHeight} | null`; the document's `laidOutCellBox` returns one or
null and the parser validates one object. The text-scale match moves
from web-ready into `handle.fitDimensions`, the one place a box is fitted.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): fit terminals in the app from the reported box; drop the measure round trip

The app already holds the box the document reported, so the refit and
the fit pass await the init's ready and call `handle.fitDimensions`
instead of posting `measure` and waiting on `measure-result`. The
document's measure, its retries, and the measure promise and timeout go.
The document still resizes locally on a text-size change, so every grid
the app sends (init, resize, reflow) carries the laid-out frame it was
fitted to. `holdSubscribedGrid` replaces `subscribeFitDimensions`, so
the only fits are `fitDimensionsFromCell` and `handle.fitDimensions`.
The render rig reads its fit from the ready box. The recorder adapter
mounts the new handle with the same recorded effects; the goldens it
mounts move on their adapterSha256 header only.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): keep the terminal frame in one ref, and notify a new width imperatively

The session held the frame in a height ref, a width ref, a width state
and the refit's own width ref. It now holds one `terminalFrameRef`
({width, height} | null until the first layout; a hidden 0x0 layout
keeps the last box). onLayout notifies a new width imperatively, as it
does height, and the refit's notify skips a width whose fit is the grid
the PTY has. `terminal-frame-width-refit.ts`, the width state and its
effect go. The subscribe's layout gate reads "no frame yet" directly.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): subscribe a held-back terminal on the frame's first layout only

`handleTerminalFrameLayout` ran on every onLayout; it now runs once, when
the frame first has a size. Later layouts only notify a new width.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): size the first subscribe inline in subscribeToTerminal

`sizeTerminalViewportFromCellBox` wrapped five lines in a 37-line
module; the subscribe now fits the ready box against the frame, holds
that grid and records the diagnostic itself. The helper's tests fold
into the subscription tests, which move to the subscription's name.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): drop the unreachable font-size guard on the reported cell box

xterm 6.1.0-beta.303 updates the render service's cell box in the same
task that sets `options.fontSize`: CharSizeService.measure fires
onCharSizeChange, and RenderService.handleCharSizeChanged runs the
renderer's `_updateDimensions` (DomRenderer.ts:359, WebglRenderer.ts:229).
`term.onRender` fires from RenderService._renderRows after the rows
are drawn (RenderService.ts:213, CoreBrowserTerminal.ts:538), and the
document writes its text scale and the font size in one task
(text-scaling.ts applyTextScale, terminal-init.ts init). So no report
can read a box between the font and the scale; the guard and its test
go. A new test pins the real order: no report when the font is set,
the new box at the new scale on the next render.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): one start seam, no source cache, the reported box as an object

- `useState` already pins each view's WebView source at mount (a new
  test re-renders at another text scale and gets the same object), so
  the module-level `webViewSources` Map goes.
- `initialTextScale` and `buildsTerminalBeforeReady` become one
  `start(): { textScale, shown }` seam.
- `reportedCellBox` holds the last reported box and grid, not a string key.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the RPC recordings to this branch and re-record

The terminal refit now fits in the app from the reported box and reads
one frame ref, so the recorder's terminal adapter mounts the new handle
(`awaitReady` + `fitDimensions`) and options (`terminalFrameRef`),
keeping its recorded effects. `baseline` is repinned to 21954dbd2f, the
last commit to touch a fenced path, and every golden is re-recorded.
Proof by class against HEAD: 787 header-only, 0 body moved, 0 added,
0 deleted. Header keys moved: `baseline` on all 787, and `adapterSha256`
on the 14 goldens `terminal-mount-adapters.ts` mounts (query-reply 3,
accessory-raw-send 4, takeover-report 4, viewport-refit 3). No recorded
traffic moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): hold the reported cell box and its grid in one ref

The controller kept the box in `cellBoxRef`, the grid in a string
`lastGridRef` and wrote it through `terminal-held-grid.ts`. One
`heldRef` now holds `{ cellBox, grid }`, as the document's own
`reportedCellBox` does: web-ready writes the box, every cell-metrics
report writes both, `holdSubscribedGrid` writes the grid, and a
readiness reset clears it. Same write points, so the one-refit bound
holds; the DOM-loop and refit-once tests pass unchanged.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the RPC recordings to the hold-rule commit and re-record

H (f00bebba48) touched a fenced path after the last repin, so
`baseline` moves to it and every golden is re-recorded. Against the
corpus before this branch's refreshes (21954dbd2f): 787 header-only,
0 body moved, 0 added, 0 deleted; `baseline` on all 787 and
`adapterSha256` on the 14 goldens `terminal-mount-adapters.ts` mounts.
Against the previous refresh: `baseline` only. No recorded traffic moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the RPC recordings to the main merge and re-record

The merge (b3b1b0def2) is the last commit to touch a fenced path, so
`baseline` moves to it and every golden is re-recorded. Against
97b5bb2b9a: 787 header-only, 0 body moved, 0 added, 0 deleted;
`baseline` on all 787, and `adapterSha256` on the 14
session.diff-review-actions goldens whose adapter #22951 edited. Against
origin/main: 787 header-only, 0 body moved/added/deleted; `baseline` on
all 787 and `adapterSha256` on this branch's 14 terminal goldens.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): the terminal document holds the grid and decides each refit

The document already kept the last reported box and grid; the app kept
a mirror of both to decide the refit. Now the document decides: its
`cell-box` notify carries `{ cellBox, refit }`, sent only when the box
changes, with `refit` a box that changed at a kept grid. web-ready
records the pre-ready terminal's box at its 80x24 grid, and the first
init that reuses that terminal holds the init's grid, so the DOM
renderer's first report refits once, as the subscribe's hold did. A
re-init no longer clears the record, so a new renderer at the same grid
still refits. The app keeps one `cellBoxRef` and `holdSubscribedGrid`,
`heldRef` and the grid on the notify go. The one-refit, DOM-loop and
renderer-swap tests move to the document with the same scenarios.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): one init options object, and a frame on every grid

`init` takes `{ cols, rows, data, preserveScroll, oscLinks, frame }`
instead of six positionals, and `init`, `resize` and `reflow` (handle
and messages) require `frame: TerminalFrame | null`. The refit's reflow
check reads `!dims` alone, and the controller's test file is named for
the `cell-box` notify it now covers.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): one notifyTerminalFrame for the frame's layout

The frame's onLayout made four calls and held the classification
itself. It now calls `notifyTerminalFrame({ width, height })`, and the
session's terminal-webview hook keeps the one frame ref, notifies the
height, subscribes the document held back for the first layout, and
notifies a later width change. `handleTerminalFrameLayout` is named for
what it does: `subscribeIntendedActiveTerminal`. The layout tests move
to that hook.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the RPC recordings to the round-8 head and re-record

85d421963c is the last commit to touch a fenced path. Against
a676c1b65a: 787 header-only, 0 body moved/added/deleted, `baseline`
only. Against origin/main: 787 header-only, 0 body moved/added/deleted;
`baseline` on all 787 and `adapterSha256` on this branch's 14 terminal
goldens. No recorded traffic moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): name the init option initialData, as the message does

The init option `data` becomes `initialData`, the message field's name,
so the controller passes it through unrenamed. The `preserveScroll` why
stays on the message type only, and the document test's title names the
three grids that carry the frame.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the RPC recordings to the round-9 head and re-record

486566c82b is the last commit to touch a fenced path. Against
3371c39715: 787 header-only, 0 body moved/added/deleted, `baseline`
only. Against origin/main: 787 header-only, 0 body moved/added/deleted;
`baseline` on all 787 and `adapterSha256` on this branch's 14 terminal
goldens. No recorded traffic moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* ci: rerun checks against main with #23560 landed

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-28 03:26:49 -04:00
Neil 0de5e4d84d Fix terminal focus when Cmd+J wakes a workspace (#23546)
* fix: retain workspace terminal focus through wake restoration

* test: reset CPU throttling after wake focus assertion

* fix: require terminal textarea readiness before claiming focus

* test: configure React act environment for dialog regression
2026-09-28 00:22:52 -07:00
Brennan Benson 6835b9b4e3 fix(mobile): paired clients re-derive a kept terminal after a cold restore (#23109)
* fix(mobile): paired clients re-derive a kept terminal after a cold restore

A renderer frame published before a cold-restored terminal's PTY registered
was fenced to an empty tab list and recorded as accepted, and the renderer
never resends unchanged content. When registerPty binds a surface the
accepted frame fenced out, re-merge that frame so the fence reads current
state.

* test(mobile): drive the live desktop window through the runtime's desktop seam

* test(mobile): the re-derive path never flushes the store synchronously

* test(mobile): a re-derived frame must not bring back a surface the host retired after accept

* fix(mobile): a re-derived frame changes membership only for the registering surface

The replay re-ran the whole accepted frame, so a surface the host retired
after accept (a phone close whose remote PTY is still exiting, or a closed
chat tab) came back. Every other surface now keeps the host's current
decision; the removal repair is extracted from the terminal retirement
helper so non-terminal tabs are removed the same way.

* test(mobile): a re-derived frame must not drop or disown a phone-created terminal the desktop has not published

* fix(mobile): a re-derived frame does not infer renderer retirements from its older frame

* revert(mobile): drop the replay of a fenced renderer frame

Reverts the production parts of a88e1eaa0a, f5b99d0003 and 5af1c6d97a: the kept
renderer frame, rederiveFencedRendererSurface and its registerPty call, and the
mergeRendererMobileSnapshot / removeMobileSessionSnapshotTabs extractions. The
fence will instead read the host's saved membership record. The test file stays
and is rewritten for that mechanism.

* fix(mobile): the paired-list fence admits a terminal the saved session still lists

After a cold restore the in-memory mobile snapshot and PTY table start empty,
so in a repo with host-authoritative terminal membership the fence dropped a
restored terminal whose renderer frame arrived before its PTY registered, and
paired clients never listed it. The fence now also admits a surface the host's
saved workspace session still lists (tab under the worktree, leaf in its
layout), and registerPty pushes the listing so pending-handle turns ready at
once. A restored pane whose PTY never returns is listed as pending-handle, as in
repos that are not host-authoritative.

* test(mobile): keep the desktop window stub's type assertion on its SAFETY line

* fix(mobile): coalesce the registration push for a listed restored terminal

registerPty pushed the paired list immediately on every registration that backs a
listed surface. The desktop's graph sync after a spawn already publishes the same
pending-handle to ready flip on the 50 ms coalescing window, so each restored pane
cost two pushes, and a restore of N panes cost N immediate full-list pushes per
client. The touch now rides the same coalescing window, which still covers a
registration no graph change follows.

The test's "unchanged" sync dropped the graph's tab, which is itself a change, and
its no-extra-push assertion ran before any coalesced push could fire; both are
fixed, and a restore of two panes is asserted to push once.

The fence comment no longer claims the new clause keeps pending leaves out of the
graph: once the surface is listed, its leaves pass the shared predicate through
that listing, as any listed surface's do.

* fix(mobile): read saved membership only from the worktree's own session partition

For a runtime-host workspace, emptying the owning partition re-routes session reads to a single
other partition that still lists the worktree. If that older copy lists a surface a retirement just
removed, a lagging renderer frame could re-admit it. The saved-membership check now reads only the
partition the worktree's host names, so it never trusts a fallback copy.

* test(mobile): pin the own saved partition for every workspace host kind

Also correct the immediate-emit comment: only an exit bypasses the window; a registration's ready flip coalesces.
2026-09-28 00:11:45 -07:00
GuanBearandguanbear 076759e728 feat(usage): show ZCode Coding Plan quota on current main (#23520)
Shows the ZCode Coding Plan quota in the status bar alongside the Claude and Codex usage readouts, reading the key from the user's own ZCode config.

Credentials are scoped tightly: the host must be an exact match in the allowlist, HTTPS on port 443 only, `redirect: 'error'`, and the key is checked for CR/LF before it reaches a header. The key itself is never stored or logged — account identity is an HMAC.

Both JSON inputs (a user-edited config file and the remote quota response) are narrowed at runtime rather than asserted, and the request cancels an unread response body on the error path so it cannot trip the undici parser crash (orca#8695).

Co-authored-by: guanbear <guanbear@users.noreply.github.com>
2026-09-28 00:07:09 -07:00
Jinwoo Hong 1b4587956a test(native-chat): await the async history and journal snapshot in three tests (#23560)
#22835 made history() and journalSnapshot() async; tests from #23502 and #22944 still call them synchronously, so the typecheck job is red on every PR while main pushes do not run it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-28 03:03:15 -04:00
Brennan BensonandClaude 7a24d3d335 fix(native-chat): the conversation outlives its agent (#22835)
* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* fix(native-chat): name a chat write by its target, not the owner generation

A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.

Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.

Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.

* fix(native-chat): every journal append reaches the chats that are open

A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.

A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.

* test(native-chat): an epoch replacement reaches the open chat

* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

* perf(native-chat): a publish behind a delivered commit reads nothing

Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.

* test(native-chat): state why the teardown test's fake journal is safe to cast

* docs(native-chat): say mutation admission checks only the writer lease

* docs(native-chat): drop the send rebase from comments that still described it

* fix(native-chat): a message is accepted, then delivered

A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".

A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.

Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.

A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.

Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.

* fix(native-chat): settle queued messages only for the child that ended

A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.

A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.

The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.

* fix(native-chat): an adoption that fails to import keeps the conversation open

The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.

* perf(native-chat): the recovering open reads the journal once

Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.

* fix(native-chat): an attach that fails after indexing its child leaves no child behind

A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.

* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer

The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.

A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.

* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down

The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.

* fix(native-chat): a message rejected while its chat was closed reads as not sent

A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.

* test(orchestration): name why the readiness settlement fakes are cast

* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent

* docs(native-chat): drop the fence from the admission the send effects run behind

* docs(native-chat): give the fence move on release the reason that still holds

* docs(native-chat): stop citing a write fence check in launch and mailbox comments

Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.

* refactor(native-chat): the provider child is its own record

A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.

- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
  patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
  own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
  the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
  dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
  is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
  the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.

* fix(native-chat): the delivery loop alone settles a message its start or child failed

A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.

- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
  reads how it ended: a Stop continues; anything else writes one failure row and rejects every
  queued message with the same words, then stops. A child still starting whose start the adapter
  says did not land fails the same way. The exit, eviction and the settlement retry only settle
  the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
  nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
  failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
  closed, with or without a child, and a start the loop already has in flight is waited for so the
  child it produces is stopped rather than left behind.

* refactor(native-chat): a stopped child ends on the one reading of its stop

The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.

* feat(native-chat): the host says it accepts a send before any agent has it

The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.

* refactor(native-chat): an attach never opens a journal of its own

The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.

* fix(native-chat): a moved fence resends nothing on a host that accepts first

The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.

The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.

* refactor(native-chat): a child's end says whether the user or the host stopped it

The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.

* fix(native-chat): a chat whose only work is a queued message is not offered for resume

A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.

* fix(native-chat): the conversation outlives its agent

Opening a chat no longer starts its agent. A conversation is reached through one host
accessor that opens its journal at rest, and a send is what starts the agent, through
the delivery loop. One idle sweep, every five minutes, stops an agent that has been
quiet for thirty minutes and owes no work, then drops an open journal handle that is
only a cache. Its record, tab, status row and readers stay.

- hold and release are no-ops; hold still builds the host for shipped mobile builds.
- The holders, the holds, the release clock and the exit respawn are deleted.
- Options, the model list, the goal and the context meter answer at rest; a model pick
  at rest is recorded as intent for the next start.
- Compact, rewind, clear and goal changes start the agent first. A send does too when
  a rewind is still in doubt after the conversation opens.
- Orchestration routes mail and group addresses on ownership (the record plus the chat
  tab), not on whether the process runs. An open dispatch keeps its worker running.
- The restart continuation is a send; Resume all holds each slot until the message is
  handed over or rejected.
- A read error never replaces a loaded transcript, and shows the host's own words.

* test(native-chat): type the queued-message fixtures in the resume-offer tests

* fix(native-chat): a start that dies while a message waits on it is that message's failed start

Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.

* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now

* test(native-chat): pin what a failed start settles, and what a resume offer names

A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.

* test(native-chat): the failed-start pins fail on what the message became, not on a timeout

* fix(native-chat): a restart offer ends when the chat's agent starts again

The offer used to end only when the chat's newest user message changed,
because opening a chat started its agent and that start could not be told
apart from real activity. Opening a chat starts nothing now, so the host
reads the fact it already publishes: a chat's status row goes from not
host-owned to host-owned exactly when its agent is started. At that edge the
offer and any failure record for the chat are withdrawn, unless the start is
a resume action's own (its continuation is the oldest undelivered message).

A continuation and a message racing to be first are decided at acceptance:
the continuation is refused, quietly and with nothing filed, when any other
message was accepted since the restart. A failed continuation start leaves
the offer retryable, and each resume action sends its own message id.

Deleted: the newest-user-message comparison, its journal reader, the
continuation filter, and the failure ledger's own "answered by the chat"
check. The marker still carries its message id for one release, so the
previous build can read it.

* fix(runtime): end a transcript stream when its client unsubscribes

Desktop: the IPC subscription controller was dropped as soon as the streaming
handler returned, which for most streams is right after it binds. A later
runtime:unsubscribe then found nothing to abort, so the host kept the subscriber
and derived and sent every publish to a channel no one listened to. The controller
now lives until the renderer unsubscribes, resubscribes the same id, or goes away.

Mobile: disposing an agentSession.subscribe stream now sends agentSession.unsubscribe
with the stream's frame id, so the host ends that subscriber and leaves a sibling
stream on the same socket running. The direct path now passes the frame id the relay
path already passed.

* test(orchestration): the preamble's host stub is typed, not cast

The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.

* fix(native-chat): one fact ends a restart offer: the chat moved on since the restart

The offer is live while no other message has been accepted in the chat since the
restart and its agent has not proved a start since. The offer list, the resume's
reservation check and the continuation's acceptance check all read that one fact,
so a message whose start then failed withdraws the offer too, and a stale click
finds nothing to act on.

The fact is read off the conversation's open handle, which the restart closed, so
it is retired durably whenever it may have changed: a message accepted, a start
proven. A close and reopen within the same run therefore cannot bring the offer
back. A continuation rejected before it reached the agent does not count, so a
retry after a failed start still runs.

Deleted: the quit-time gate on withdrawal, which changed nothing because the
withdrawal and the quit's own offer write share one queue; the per-action
"withdrawn" flag and the separate acceptance check it paired with.

* test(native-chat): an older build reads the restart offer this build records

The offer lives in a file the previous release reads after a downgrade. Pin that
against the pinned release's own capsule, and run the lane when the marker or the
capsule changes.

* fix(native-chat): read a restart offer against where the journal stood when it was taken

"Since the restart" was read off the conversation's open handle, which the idle
sweep closes: after a reopen, a message the user had already sent looked older
than the handle and the withdrawn offer came back.

The offer now records the journal position (epoch and sequence) at the moment
it is taken, and a message accepted after that position, or a journal on another
epoch, means the chat moved on. That is derived from the journal, so it holds
across any number of closes and reopens. An older build's offer has no position;
only a start withdraws it. Because the message half is now durable, the offer is
no longer rewritten in the recovery file on every accepted message; a proven
start still writes it, since only the host that saw the start knows of it.

* test(native-chat): wait for the listing's retire write before reading the recovery file

* fix(native-chat): keep the terminal-backed chat's read error over its local echoes

Messages winning over a read error is right for the structured chat, whose read retries and whose
messages came from the transcript. The terminal-backed view assembles its list from local echoes
too (a launch prompt, a pending send), so a failed read there showed only those bubbles and no
error. Only the structured pane now keeps messages over an error.

* fix(native-chat): a start retries the exit settlement a failed journal write left owed

An agent exit whose journal settlement write failed releases the lease latched until a retry lands.
Reopening the chat used to be that retry; with reveal now only opening the journal, nothing retried
it before the next app launch, and every send was refused. The start the send needs now runs the
retry first, where the attach would.

* perf(native-chat): answer the owner check without opening the chat

Worktree activation calls agentSession.handoffStatus for every chat tab in the worktree, and the
answer comes from the session record alone. Reaching it through the accessor opened each resting
chat's journal (a full read, the crash-boundary write and a restored status publish), then kept it
open for the idle window. It now checks the record and the adapter's support, as before this series,
and opens nothing.

* fix(native-chat): a read waiting on the session lock opens nothing once quit began

The accessor checked for quit before queueing the open, so a read queued behind a session task ran
its open after teardown had begun and indexed a journal no teardown step would close. The check now
runs at the open itself.

* fix(native-chat): read a failed resume's chat before calling it retryable

Whether a failed resume is retryable is the offer's own rule: the chat has not moved on since the
restart, read from its journal. The failure list read it only for a chat already open, so once the
idle sweep closed a chat the user had moved on in, its failure showed Retry again, and the click did
nothing. The list now opens the failed chats first, as the offer list does.

* test(native-chat): type the provider event sink the settlement test reaches for

* fix(native-chat): say the structured read keeps trying only where it does

The structured pane's "Orca keeps trying to load it" line never showed: the view state filled in an
untranslated fallback whenever the read error had no text, and the empty state prefers any message.
The view state now leaves the message out, so the structured pane shows that line and the
terminal-backed pane its own translated one. Mobile's structured lane does not resubscribe after an
error frame, so it no longer makes the claim.

* test(native-chat): await the send's settlement instead of polling for the start

The at-rest send tests polled for the provider start with vi.waitFor's one-second default, which a
loaded machine outran. They now await the host's own settlement of the message.

* fix(native-chat): a restart offer resumes any time after the quit, and knows its own continuations

The continuation's message id was dated by the quit, and the ledger refuses a new id dated more than
a day back, so Resume or Retry a day after quitting was always refused (on main too). It is now
dated by the resume action.

Telling a rejected continuation from the user's own message read the operation ledger, whose rows
expire after about a day; after that a failed resume stopped being retryable. The offer now
records the continuation each action sends on its own capsule entry, bounded to the newest 16, so
the ids end with the offer. The ledger read is deleted.

* fix(orchestration): route no mail to a structured worker its orchestration released

A structured worker is routed on ownership, and a resting worker's lease is released, so ownership
held while its chat tab stayed listed. A worker the coordinator abandoned and then released, found
at rest by the release, therefore still took peer mail and @worktree: broadcasts, and each one
restarted its agent. Routing now also reads the orchestration's own resource row: once it is
released, direct mail, group addressing and worker-show's addressable answer drop the worker, as
they would a terminal worker whose terminal closed. The chat tab stays, and nothing new is stored.

* fix(native-chat): a failed retry names the user's prompt, not Orca's continuation

A resume's continuation is written to the chat before its start, so after a failed attempt the chat's
newest user message is that rejected continuation. A second failure then showed Orca's own restart
text as the chat's prompt. A retry now keeps the prompt its first failure named.

* fix(orchestration): read the released row optionally, as the authority does

worker-show's observation called the row lookup directly, which a runtime double without it threw on
and failed the structured tab-retirement release.

* fix(native-chat): the status bar drops a restart offer the chat moved on from

The renderer re-read the host's restart offer only when a failed chat showed activity, so after a
message withdrew a pending offer the host answered no chats while the status bar kept counting one,
and clicking it opened nothing. The same watch now covers pending offers: a status change in an
offered chat asks the host again, once.

* test(native-chat): a roster of idle or finished children does not keep an agent awake

The sweep reads owed background work through the shared child-work liveness that upstream's
release clock adopted; a child that went idle or finished is not work the agent still owes.

* fix(orchestration): a task dispatched into a resting structured worker keeps it running

The sweep's open-dispatch check read only the worker-start dispatch that owns the worker's terminal
resource, so a task later dispatched to the same worker (orchestration dispatch --to, which writes a
dispatch with no worker row) did not count: after thirty quiet minutes the worker was stopped while
that task was open, and its coordinator read exited. Any unsettled dispatch addressed to the worker's
process incarnation now counts, derived from the existing rows.

* docs(native-chat): comments stop describing the hold this PR removed

Eight comments still justified orderings and teardown choices by a viewer or dispatch hold that
pinned the provider child. Nothing holds any more; the orderings stand for the binding's redrive
subscription and parked mail, and a chat's agent runs from a send until the idle sweep rests it.
Comment-only.

* fix(native-chat): a restart offer keeps the start its own continuation made

Whose start ended an offer was decided at read time, from whether the offer's continuation was
still the queued message. Once the provider refused that continuation, the child it had started
read as someone else's start, so the offer ended and its failure showed no Retry. The delivery
loop now records which queued message a start is for on the in-memory child, and the child's end
carries it; the offer counts a start as its own when that message is one of its continuations.

* fix(native-chat): an agent gets a full idle window after its owed work ends

The sweep measured quiet only from the last journal row, so once a subagent, command, monitor or
dispatch that had outlived the window ended, the agent was stopped at the next tick. A child can
read done before the lead's wake-up turn writes anything, and stopping in that gap loses the
wake-up. The sweep now counts owed work it observes as activity, which gives the agent the full
window afterwards, as the release clock it replaced did.

* test(claude): the options-read fixture runs a live child

The fixture marked its conversation running with a hasProviderChild field the
session type does not have, so the read took the at-rest path and refused a
session with no record. It now carries a child, which is what the read checks.

* test(native-chat): host tests reach its collaborators through a typed seam

The rest-test rig and three test files read the host's private members with
Reflect.get and cast the result. The host now exposes one test-only accessor,
collaboratorsForTests(), and the subscribers class a subscriberCountForTests()
beside its existing retainedActivityCountForTests(), so the tests are checked
against the real types and the casts are gone.

* refactor(orchestration): one owner answers a structured worker's custody

Routing, group addressing, worker-show and the idle sweep each composed their own reading of
whether orchestration still holds a structured worker, so each new obligation or retirement state
had to be added to every reader. structured-worker-custody now derives both answers from the
worker-terminal list state coordinators see in worker-list: addressable is owned and not released,
and owed work is an active custody or an unsettled task dispatched to the same incarnation. The
owner's state is read through the remote dispatch attachment too, as the terminal transfer lookup
already does. Behaviour is unchanged; a settled worker awaiting its coordinator still rests.

* refactor(orchestration): owed work is an open dispatch on the worker's incarnation

A supervised worker's own dispatch context stays open exactly while the worker is active, so the
separate active-custody branch only repeated it. Owed work is now one fact, which also states the
policy that a worker awaiting its coordinator's decision may rest, and both custody decisions are
written once at the top of the module.

* fix(native-chat): a restart offer knows its continuations by a tag in their id

The offer recorded each continuation id in a list on its capsule entry, capped at 16, and a running
action's id in memory. Both could disagree with the journal: past the cap an old rejected
continuation read as the chat moving on, and a crash during a retry restored the failure's older
entry, which lacked the retry's id. Each continuation id now carries a tag derived from the offer
(its teardown and chat), then the action's own part, so any continuation of this offer, queued or
rejected, is recognised from the journal row and the marker alone. The persisted list, its cap and
the in-memory action map are deleted; the agent-start withdrawal keeps an offer whose own
continuation the start was for, read against the stored marker.

* test(runtime): the legacy-worker reveal test judges its stale snapshot inside the wait

The tui-idle probe reads through readTerminal, which now awaits the structured
worker check before the PTY read, so the probe's snapshot request starts a
microtask later. vi.waitFor missed it on its first check and polled again at
50 ms, the same moment the wait's own 50 ms timeout fired. The stale snapshot
then resolved after the wait had already timed out, so the test passed without
judging it, and the rejection landed before any handler was attached. Vitest
reported that as an unhandled error and failed the shard.

Polling every 1 ms sees the request within a few ms, so the snapshot is judged
while the wait is still pending.

* fix(native-chat): a view never restarts a chat whose last start failed

A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.

* fix(native-chat): the idle sweep reads owed work every tick

Owed work counted as activity, but the sweep read it only once the idle window had elapsed, so it
refreshed the clock at most once a window. Work that ended just before the next read left the
agent to be stopped at that read, moments after the work ended, which is the gap the refresh was
meant to cover. The sweep now reads owed work on every tick for a started agent, so the window
always runs from the last tick that saw work owed.

* fix(native-chat): a continuation handed to the agent stays sent

The offer read its own continuation as not reaching the agent while its dispatch was pending, which
also covered one already handed over and still unanswered. When the wait for that answer ended first,
the failure it filed read as retryable, and a retry sent a second continuation to an agent that may
have acted on the first. Only a continuation still queued, or rejected, is now read as unsent.

* test(native-chat): start the child the loop waits on with an attach, not a second view

A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.

* fix(native-chat): settle a gone generation's turn wherever a conversation opens

A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.

* test(native-chat): prove the next child's start settles the turn an earlier child left

The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.

* test(native-chat): count a failed start's rows by row, not by text

Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.

* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget

* test(native-chat): the interrupted create's own retry continues again

The merge of main's lease-latch fix replaced that test's retry of the interrupted create, under its
own operation id, with a fresh start whose result nothing read. That fresh start passes with the
released-reservation continuation deleted, so the case the fix exists for went untested. The retry
and its assertion are main's again.

* docs(native-chat): three comments that still had views starting agents

A start with nothing queued now comes from a command, goal change or rewind; an interrupted compaction
left alone would refuse every send, so no agent would ever start to finish it; and a current host
raises the unattached read refusal only once quit began, with the attach window belonging to an older
host.

* test(native-chat): pin the open's and the send's start and row counts, however the view binds

Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.

* test(native-chat): a reader's open settles the turn a failed exit settlement left running

An exit whose settlement write failed leaves its turn running in the open journal. PR 1's open now
settles it, and this pins the two reads that reach it here: a reader reopening a chat the idle
sweep closed, and a read that opens the chat before the restart restore reaches it.

* test(native-chat): the view-start test's starting window outlasts two subscriptions on a loaded runner

A subscription reads the conversation before it returns, so under load the two views took longer
than the create child's 300 ms start, which then exited before the test checked that it had not.
The child now takes a second to fail.

* fix(native-chat): settle a gone generation's turn at every open but an acquisition's

The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.

* test(native-chat): hold the create's start open until the views bind

The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.

* test(native-chat): a read that reaches a crashed chat before the startup reconcile settles its turn

On desktop the chat on screen at relaunch reads before startup reconciles the leases, while the
dead process's lease still reads live. The open settles the turn it left running anyway, and the
restore that follows finds it settled.

* refactor(native-chat): drop the composer's second error formatter

After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.

* test(native-chat): pin the reason on a message rejected while its chat was closed

The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.

* docs(native-chat): drop the removed dispatch hold from six comments

A worker's session no longer takes a dispatch hold, and no release clock
rests a chat by visibility; the agent-launch comments, the abandon test,
the teardown test and the refusal census still said so.

* test(native-chat): rest the owner-status chat through the idle sweep, not a hold

The activation-gate test from #22808 put its chat at rest by holding and
releasing it, and passed the release-clock grace. This branch deleted both,
so the case threw before it reached its assertions. It now moves the host's
clock past the idle window and lets the sweep stop the agent and close the
conversation, then asserts the same owner answer and activation gate.

* fix(native-chat): show the structured pane's retrying line when a read fails

The read transport always hands the pane the host's words, so the error
state's "Orca keeps trying to load it" line, which showed only when there
were none, was never seen: the pane showed the host's text twice, as its
subtitle and on the status line under it. The structured pane now always
says its read keeps retrying, and the host's text stays on the status line.
The terminal-backed chat is unchanged.

* test(native-chat): wait for a send's background start before the refusal oracle removes its store

An accepted send wakes the delivery loop, which starts the agent in the background. The oracle's teardown disposed the loop but did not wait for that start, so its lease write could create a temp file in the store directory while the directory was being removed, failing the test with ENOTEMPTY about one run in four. The teardown now drains tracked starts before it closes the journals.

* fix(native-chat): a start a message waited on gets one failure row, the delivery loop's

When a queued message's start failed, two writers could report it under the same row: the delivery loop, when the adapter settled the start without proving it, and the exit settlement, when the child's exit landed. The last one won, so the chat's row could name a different cause than the one the message was rejected with, or be written twice.

The exit settlement now writes the start's row only when no message is queued and the loop has not already recorded that start. A start for a command, goal change or rewind, with nothing queued, still gets its row from the exit.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-27 23:46:14 -07:00
Brennan Benson 5e5f4f6603 fix(native-chat): keep a Codex ask's questions in the order it asked them (#23502)
* fix(native-chat): keep a Codex ask's questions in the order it asked them

Codex journals every question of one ask in a single write, so the
questions share a timestamp. The transcript list sorted rows by
timestamp and broke ties by row id, and a question's id ends in the
question id the model chose, so answered, cancelled and still-pending
rows of one ask came out in the alphabetical order of those ids.

The transcript projection now breaks timestamp ties by the order its
source gave: the journal's order for structured sessions. The session
assembler keeps its id tie-break, since the sources it merges share no
order of their own.

* refactor(native-chat): name the id tie-break comparator for what it does

Two shared comparators differed only in whether they break timestamp ties by id,
under near-identical names; spell the id tie-break in the name.

* fix(native-chat): order structured chat rows by their journal position

The desktop list and the host's conversation outline sorted structured rows
by timestamp. The journal's contract is that the sequence orders the
timeline and the timestamp is the provider's clock: a Codex ask writes all
of its questions in one write (one sequence, one timestamp), and a row
recovered after a crash carries an earlier clock at a later sequence.

The host now records each item's place within the write that created it,
keeps it across revisions like the sequence, and sends it as an optional
field. Rows projected from the journal carry that position, and both the
desktop list and the outline order journal rows by it. Rows outside the
journal keep their rules: rank for the streaming and pending tail, outbox
sends after every journal row, and terminal-backed chats keep time then id.

* fix(native-chat): keep a refused send at its journal place, and keep list positions off worker reads

A send the host journalled before the provider refused it is shown from the outbox, and it
sorted after every journal row, so it dropped below whatever the agent wrote after it. It now
takes the journal position of the submission the host recorded.

Structured worker reads and their archives projected journal rows through the same
projection, so they returned the list-only journal position on every message. The worker
payload bound now drops it.

* test(native-chat): a failed restart's row draws below the message it failed

Since a message is accepted before its delivery starts the agent, a restart that fails is
journalled after the message, and the refused message keeps that journal place in the chat.
Both host paths now pin it through the chat's own projection: a start the host could not make,
and a restarted child that exits before proving its start.

Folds the journal reducer's batch item write onto fewer lines, which the merge of main pushed
past the file's line limit.
2026-09-27 23:44:36 -07:00
Brennan Benson 56e691344f fix(native-chat): show the "Working for" bar while a turn runs (#23537)
* fix(native-chat): show the "Working for" bar while a turn runs

The turn bar under the user's message only rendered once a turn settled, so a
running turn had no bar, only the spinner line at the tail. The running turn
now draws the same bar with a live clock, and it settles in place to "Worked
for". The tail line keeps its spinner but drops the clock, so the clock is said
once: activity text, else "Thinking", else "Working…". Mobile gets the same split.

* fix(native-chat): keep the turn bar mounted through the settle

When a turn ends without a host-recorded duration (a send folded into a
running turn, or a host that stamps no start), the local duration is stamped
one pass after the working flag drops. For that pass the turn status resolved
to nothing, so the live "Working for" bar unmounted and remounted as "Worked
for" - invisible on desktop (layout effect) but a painted blink on mobile,
where the stamp runs in a passive effect.

The shared selector now keeps the just-ended turn's running status until its
duration lands, unless the host says it never saw the end, and mobile reads
the active turn's row from that selector the way desktop does.
2026-09-27 23:39:43 -07:00
Neil 8494b2422d perf(ci): take i18next-cli 1.74.1 so extraction stops re-scanning every key to find its leaves (#23550)
1.65.0 decided "is this key a leaf" by scanning all extracted keys per key, so the
check cost ~13,200 squared startsWith calls. 1.74.1 replaced that with a prefix
Set. Extraction drops from 22.1-23.3s to 16.9-17.4s locally, and the gate from
~23s to 17.8s.

It also extracts 17 keys 1.65.0 missed, all of which are already committed in
en.json, so they move out of the orphan list rather than needing new
translations: extracted 13,191 -> 13,208 and orphans 1,567 -> 1,550. No key is
lost and no existing value changes, so no catalog is regenerated here.
2026-09-27 23:32:47 -07:00
Neil 9179b93ebf ci: reduce repeated runner work and validate affected-test selection (#23540)
* ci: stage heavy checks and measure affected-test selection

* fix(ci): exercise the real Git boundary in unit selection planning

* Harden review cancellation and CI demand reporting
2026-09-27 23:25:20 -07:00
Neil fae0ae7a46 perf(ci): shard the anti-slop audit across processes instead of one JS runtime (#23543)
config/oxlint-anti-slop.json turns every native oxlint category off and runs its
rules through jsPlugins, so oxlint's threaded Rust engine does no work and the
pass is one JS runtime per process. Measured, it does not scale with --threads:
11.68s at 4 threads against 12.38s at 16. Parallelism has to come from more
processes, so the audit now splits its file set across them.

Locally, 11.72s single pass against 3.83s at 4 shards (3.1x) and 2.91s at 8.

Sharding is sound because every anti-slop rule is a single-file analysis; the
only mutable module state is a WeakMap keyed on each file's own Program node.
Verified by running both shapes with all 18 rules enabled: 392,398 findings from
one pass and from the shard union, identical as sorted multisets.
2026-09-27 23:15:19 -07:00
Brennan Benson c53ed030b2 fix(terminal): one intentional-stop register, and per-run spawn and input facts (#22989)
* fix(terminal): one intentional-stop register and per-run spawn and input facts

Main now keeps one register of PTY stops it made on purpose, with an owner
count and a kind: reversible (sleep, hibernation) or replaced (a restart
handing the pane to a new process). It replaces four separate markers, and
every exit path reads it, so a hibernated or restarted pane keeps its tab and
binding through the exit instead of being retired and grafted back.

Main also records two facts per process run, keyed by incarnation: whether the
run was a fresh spawn (from the spawn-commit origin, which now counts a cold
restore as a reattach) and when a client first sent it input a person
produced. Renderer keystroke, paste, drop and quick-command writes carry a
userInput flag; terminal.send, stream input and dispatched agent prompts are
recorded by the runtime, excluding terminal query replies.

* test(terminal): a stopped process's synthetic and late provider exits keep its pane

One process can reach the runtime's exit handler twice: main's synthetic exit
when the kill reply overtakes the stream, then the provider's own exit, which
certifies the death. Both must read the same stop, a later process on the
same id must not, and the stop is forgotten when the duplicate-exit window
closes.

* test(terminal): a runtime worktree sleep keeps its tabs and wake bindings

* fix(terminal): keep every overlapping stop kind on one exit

A sleep and a restart that stop the same process now each keep their
label: the exit carries both renderer flags, and a runtime kill during
the sleep still records its SSH stop as reversible. A later stop of an
already-stopped process joins its entry, so a failed repeat cannot erase
the stop that landed.

The cold-restore rule moves into the run facts, so the binding span keeps
labelling a cold restore as a spawn.

* test(terminal): IME and Hangul commits reach the PTY as user input

Drives a real xterm through the pane's input handling: a macOS
input-source substitution, an iPadOS Hangul syllable, and a composition
the route delivers after a pane switch each reach onData flagged as user
input, which is what tags the write main records.

* test(terminal): state why the IME provenance test's canvas stub is safe

* fix(terminal): a new process ends an unpinned stop, and focus reports are not input

- A landed stop that no exit pinned to an incarnation now ends when a new
  process commits on the same id, so that process's exit is not read as
  the stop. Both spawn-commit funnels report through one runtime method.
- A spawn commit with no incarnation starts its run facts clean, since it
  cannot be told from a new process.
- Stream input that is only focus reports no longer counts as user input;
  the desktop renderer already excludes them through xterm's own signal.
- The IME provenance test uses typed fakes: the pane input and composition
  route installers now take only the fields they read.

* chore: restore pnpm-lock.yaml to the base revision

* fix(terminal): typing in the dashboard preview is a run's user input

The dashboard popout's terminal preview writes to the PTY through its own
runtime path, which recorded no input, so a pane typed into only from the
preview read as untyped. It now records before the write, after the
mobile-driver check. One classifier, shared with terminal.send, decides
which provenance-free bytes nobody typed: a whole terminal reply or only
focus reports, which also covers the focus-in a desktop renderer sends
when it reattaches a remote pane.

* test(terminal): the IME provenance test removes the navigator stubs it adds

happy-dom serves navigator.userAgent from the prototype, so the test found no
own descriptor to restore and left the iPad user agent in place. The pane-switch
case then ran as a Mac pane only because it followed the Hangul case. Deleting
the own-property stub restores the default platform for every case.

* fix(terminal): every PTY write names its input kind, and one record point reads it

A run's first input was recorded by opposite defaults: the renderer's
pty:write recorded nothing unless a writer opted in, terminal.send recorded
everything that was not a reply, and main's own controller writes recorded
nothing. Each unclassified producer silently took its transport's default, so
the worktree-create draft counted from the renderer and not from main, and
mailbox pointers never counted.

Every host write entry point now takes a required kind: driving, launch or
query-reply. That covers the runtime controller's write and settled write,
the terminal writer, sendTerminal, sendTerminalAgentPrompt, pty:write and
pty:writeAccepted, the renderer transport and the runtime input helpers. The
fact is recorded once, just before the provider write, in the controller
funnel and the pty:write funnel: only driving bytes count, and a payload that
is only a reply or focus reports never does. The per-producer records are
gone.

Launch writes (create-time drafts and follow-ups, the agent launch prompt,
restored and cold-restore startup commands, the SSH background launch) do not
count. Mailbox pointers, dispatch, plugins and agent-team sends do. The
runtime mixins are unchecked by the compiler, and the pane session is an any
bag, so two source scans fail on any write there that leaves out its kind.

* fix(terminal): type the pane session's transport so the compiler checks every write's kind

The pane-connection session is an `any` bag, so a transport write there that
left out its input kind compiled, and only a text scan over the renderer
caught it. Declaring `transport: PtyTransport` on the session puts those writes
under the compiler, so the scan is gone. One hidden-delivery guard now narrows
a null PTY id itself instead of relying on an untyped predicate.

The main-process scan over `@ts-nocheck` runtime files missed five files whose
directive sits on the second line, below a lint directive, and missed a write
through a local alias of the PTY controller. It now finds both.

* fix(terminal): record a runtime spawn's run facts only after its binding save succeeds

The runtime spawn funnel reported the commit before its host-session binding save, so a spawn discarded for a failed save still recorded run facts and cleared a landed stop. Report it after the save, and separately on the adopted return, which makes no save. Also align tests and fixtures with main: the removed stop-owner map, required write input kinds, SQLite-backed stores and async binding saves.

* test(terminal): pin where the renderer spawn funnel records its commit

The merge moved the renderer funnel's spawn-commit report into the publish step, after the binding save, but no test covered it: deleting the call or moving it back before the save left every suite green. Cover both: a committed spawn records its run facts and supersedes an unpinned landed stop, and a spawn discarded for a failed binding save records nothing and leaves the stop.

* fix(terminal): record a runtime spawn's commit only once its registration succeeds

The runtime funnel reported each commit before registering the process, so a spawn that exited during start still recorded run facts and could end a landed stop, while the renderer funnel reports only after registration. Report after registration on both the ordinary and the adopted branch, so only a committed, registered run has facts and an unknown run keeps reading as not fresh.

* test(terminal): pin that a runtime adoption supersedes a landed stop no exit pinned

* fix(terminal): read an unconfirmed explicit stop's reversibility from the intentional-stop registry
2026-09-27 23:13:24 -07:00
Neil 67581dd090 perf(ci): stop the orcad smoke idling 15s and the static job fetching mobile packages it skips (#23541)
The shutdown race in the orcad terminal smoke never cleared its losing timer, so
the process sat on a live 15s timer after PASS had already printed. Measured
locally: 21.4s -> 6.52s, with the round trip and the shutdown assertion intact.

Static analysis also asked for the mixed root+mobile pnpm store (537 MB, 8.6s to
restore) on every run, while installing mobile dependencies only when the diff
needs them. Most runs paid 216 MB for packages they never linked.
2026-09-27 22:51:59 -07:00
Neil 45f3512a33 feat(agents): add first-class DeepSeek Harness (dsh) support (#22468)
* feat(agents): add first-class DeepSeek Harness (dsh) support

Register DSH as a supervised Orca agent: catalog entry and detection for its
dsh-tui profile, status/question hooks through DeepSeek's own Claude-Code hook
bridge, composer-ready prompt delivery, session resume, headless Source Control
AI, and title identity that no longer collides with Gemini's.

* fix(dsh): reach Orca through DSH's credential scrub and stop reading its title as Gemini

DSH runs command hooks through its own shell executor, which drops every env var whose
name contains KEY, TOKEN, SECRET or PASSWORD — taking ORCA_PANE_KEY and
ORCA_AGENT_LAUNCH_TOKEN with it, so every hook exited without posting. Mirror both onto
scrub-safe aliases at spawn and restore them at the top of the DSH hook script.

Its title collided too: DSH rests on the same glyph Gemini works on, so a resting DSH
pane was relabelled Gemini CLI and reported working forever. Defer both the Gemini
classifier and the title status detector on DSH's whale, in the base module both copies
of that classifier read.

* test(mobile): repin the session-route closure for the DSH agent icon

* fix(dsh): address review — never splice user rows, cover remote panes, keep the diff off argv

- findManagedDshPatchRegion paired an orphan start marker with a later block's end, so a
  truncated write made install/remove delete the user's own rows. Pair each end with the
  nearest preceding start; regression test fails without the fix.
- The relay PTY env builder never applied the scrub-safe aliases, so remote DSH status
  silently never appeared even with the remote hook installed.
- Source Control AI sent the whole diff on argv; send it over stdin with DSH's '-' marker.
- dsh-tui/dst already chose the interactive profile, so a workspace folder named 'web' or
  'plugin' no longer marks a live agent pane non-interactive.
- Isolate USERPROFILE as well as HOME so a Windows run cannot edit the real home.
- Drop the duplicate README badge and revert an incidental doc reformat.

* refactor(dsh): share the managed-hooks reader and tighten the new modules

Reuse before reimplementing: readManagedDshHookEvents was a near-verbatim copy of Muse's,
with byte-identical private helpers. Both now call one readManagedHookEventsFromJson.

Also: one readTextOrAbsent instead of two spellings of the same read (dropping an
existsSync TOCTOU), one status() builder instead of four inline literals, rmSync(force)
instead of exists-then-unlink, and a redundant empty-string guard before JSON.parse.
The patch-file transforms lose their index juggling for a predicate plus a filter.

* fix(dsh): refuse a flow-style patch file, keep its mode, and stop the relay inheriting a pane

- applyManagedDshPatch matched only an exact `[]`, so `[] # keep empty` or a non-empty
  flow sequence got a block entry appended after it — invalid YAML that would leave DSH
  unable to load the user's own patch layer either. It now strips the token from an empty
  sequence (keeping a trailing comment) and returns null for a non-empty one; install
  reports that and changes nothing.
- The patch rewrite dropped an owner-only file to the umask default (CWE-732); pass
  preserveMode.
- The relay PTY env never dropped inherited pane identity the way the local and daemon
  builders do, so a spawn that specified none could inherit the relay's own and every
  agent's hook would report against that pane.

* fix(dsh): keep the flow-style refusal in every status read, and scope the mode test to POSIX

A refused patch file carries no managed region, so getStatus() fell through to a bare
not_installed with detail null — the actionable 'rewrite it as a block sequence' message
only ever reached the one-shot install() return. Export the predicate and check it first,
behind one shared message constant.

The owner-only mode assertion cannot hold on Windows, where chmod only toggles the
read-only attribute and mode & 0o777 reads 0o666 for any writable file.

* docs(readme): restore the DeepSeek Harness badge lost in the rebase

* test(mobile): repin the session-route closure to the measured 4221

Measured, not derived: 4220 without the DSH icon entry, 4221 with it. Two of the three
modules above main's 4218 pin are not this change's — they arrived with the mobile work
after #22570 and were never repinned; the changelog records that split explicitly.

* fix(dsh): settle tui-idle on the agent's own hook, so supervised workers see it ready

Reported by a tester on the adhoc build: `terminal wait --for tui-idle` ran to its 90s
timeout against an already-ready DSH composer, so a supervised worker never sees the agent
as ready.

Every existing tier reads the title, and DSH deliberately carries no title status: its rest
prefix is Gemini's working glyph, so the detector reports none. A fresh first-party `done`
is better evidence than any title anyway — it is the agent's own account of its own turn,
and normalizeDshEvent drops subagent events, so it is the lead's. Scoped to DSH: for agents
whose hooks report child turns, a mid-turn `done` is the #6011 class this file prevents.

* test(daemon): record the DSH transcript's true-colour I2 divergences

Adding the dsh-tui capture to __fixtures__ enrolled it in the serialize replay sweep, where
it reports 10 I2 divergences and failed the unlisted-transcript default of 0.

Every one is the same shape — visible-grid row=0, a 24-bit background the round trip does
not restore to default — which is DSH's whale intro painting whole rows of true colour.
Verified as an upstream limitation rather than a regression by replaying against the
previous build (build-serialize-addon-at-ref.mjs --ref origin/main): I1 and I3 both hold.

* fix(dsh): return the new tui-idle verdict from the first-party done lane

Main refactored isTuiIdleSatisfied into evaluateTuiIdle, which returns a verdict rather
than a boolean. The DSH lane still returned `true`; it is tier-1 positive evidence, so it
returns READY_STRONG like the title/body lane above it. Re-verified the regression test
still fails without the lane.

* test(relay): pin the scrub-safe pane-identity aliases on the relay spawn path

The relay builds a remote pane's env itself, so the alias mirroring there had no
test: removing the call left every suite green while remote DSH status silently
vanished. Both cases fail without it.

* docs(dsh): point the hook service at the integration reference

The reference doc had no inbound link from anywhere in the repo.
2026-09-27 22:44:18 -07:00
Jinjing 708123b868 Improve PTY device error messages with localization support (#23538)
* Improve PTY device error messages with localization support

- Extract error hints to shared module for reuse across host and renderer
- Change multi-line hints from space to newline separator for readability
- Add localization of resource-limit hints in the renderer
- Prevent duplicate issue requests when toast renders its own link
- Handle legacy hint formats from older hosts

* Prevent duplicate PTY allocation hints on legacy messages

- Extract hint-detection logic into hasPtyAllocationHint() helper
- Check for both current and legacy PTY allocation hint variants
- Prevents duplication when messages already contain legacy hints
2026-09-27 22:43:37 -07:00
Jinwoo Hong 5219b8ada9 fix(relay): reset input modes a dead program left on in SSH terminals (#23488)
* fix(relay): ground input modes a dead app left armed on SSH terminals

SSH relay PTYs now run the same recovery barrier as the local daemon:
between startup ingress and the replay buffer/publish sink, it pauses at
an OSC 133;D that closes a command which left input modes or the
alternate screen armed, proves the shell owns the PTY foreground on the
execution host, and on proof injects the process-boundary ground ahead of
the prompt. Teardown flushes held bytes through releaseRelayIngress.

The barrier now holds only the D marker's terminator and carries the
ground on it as one transformed emission over exactly one raw unit.
Previously the ground was a zero-raw emission, which source-credit
delivery accepts but never sends, so live output and replay diverged.
Every emission now covers at least one raw unit; refuted, timed-out,
overflowed and flushed episodes release the terminator unmodified.

* fix(terminal): keep the held 133;D terminator inside its emission's raw span

Treat a trigger end below 1 as unsplittable instead of trusting the
scanner's clamp, arm the bail timer before queueing the terminator, and
build the relay's startup ingress unconditionally beside its barrier.

* fix(terminal): hold the whole 133;D mark so a mid-proof snapshot ends on an escape boundary

The scanner now reports where the unclean-death D mark starts; the
barrier releases everything before it and holds the mark itself (bounded
to 4K, so the grounded span stays inside a relay source frame). A
snapshot taken mid-proof therefore has no open OSC for consumers that do
not restore the pending escape tail. The relay defines PTY liveness once.
2026-09-28 01:27:07 -04:00
Jinwoo Hong 0e871a9a71 fix(terminal): clear the Kitty input mirror when a shell is confirmed (#23483)
* fix(terminal): clear the Kitty input mirror when a shell is confirmed

A confirmed return to the shell (a host foreground read proving the agent is
gone) wrote the Kitty keyboard reset into xterm but not into the renderer's
Kitty mode tracker. Shortcut policy reads the tracker, so after an agent died
with Kitty flags armed on a host that does not ground the stream itself,
Shift+Enter and other chords stayed CSI-u encoded for the plain shell.

Scan the same reset bytes through the tracker before xterm parses them, so
both records apply identical input. The tracker doc now states its rule: fed by
application output and host-proven process boundaries, never renderer guesses.
The Ctrl+C and reattach resets are unchanged.

* test(terminal): pin that the confirmed-shell reset leaves the parked main screen's flags
2026-09-28 01:27:02 -04:00
Brennan BensonandClaude 85067494a1 fix(native-chat): a request that failed reads as failed (#22944)
* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* fix(native-chat): name a chat write by its target, not the owner generation

A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.

Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.

Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.

* fix(native-chat): every journal append reaches the chats that are open

A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.

A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.

* test(native-chat): an epoch replacement reaches the open chat

* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

* perf(native-chat): a publish behind a delivered commit reads nothing

Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.

* test(native-chat): state why the teardown test's fake journal is safe to cast

* docs(native-chat): say mutation admission checks only the writer lease

* docs(native-chat): drop the send rebase from comments that still described it

* fix(native-chat): a message is accepted, then delivered

A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".

A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.

Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.

A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.

Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.

* fix(native-chat): settle queued messages only for the child that ended

A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.

A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.

The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.

* fix(native-chat): an adoption that fails to import keeps the conversation open

The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.

* perf(native-chat): the recovering open reads the journal once

Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.

* fix(native-chat): an attach that fails after indexing its child leaves no child behind

A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.

* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer

The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.

A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.

* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down

The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.

* fix(native-chat): a message rejected while its chat was closed reads as not sent

A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.

* test(orchestration): name why the readiness settlement fakes are cast

* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent

* docs(native-chat): drop the fence from the admission the send effects run behind

* docs(native-chat): give the fence move on release the reason that still holds

* docs(native-chat): stop citing a write fence check in launch and mailbox comments

Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.

* refactor(native-chat): the provider child is its own record

A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.

- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
  patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
  own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
  the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
  dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
  is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
  the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.

* fix(native-chat): the delivery loop alone settles a message its start or child failed

A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.

- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
  reads how it ended: a Stop continues; anything else writes one failure row and rejects every
  queued message with the same words, then stops. A child still starting whose start the adapter
  says did not land fails the same way. The exit, eviction and the settlement retry only settle
  the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
  nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
  failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
  closed, with or without a child, and a start the loop already has in flight is waited for so the
  child it produces is stopped rather than left behind.

* refactor(native-chat): a stopped child ends on the one reading of its stop

The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.

* feat(native-chat): the host says it accepts a send before any agent has it

The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.

* refactor(native-chat): an attach never opens a journal of its own

The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.

* fix(native-chat): a moved fence resends nothing on a host that accepts first

The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.

The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.

* refactor(native-chat): a child's end says whether the user or the host stopped it

The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.

* fix(native-chat): a chat whose only work is a queued message is not offered for resume

A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.

* test(native-chat): type the queued-message fixtures in the resume-offer tests

* fix(native-chat): a start that dies while a message waits on it is that message's failed start

Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.

* fix(native-chat): a request that failed reads as failed

A structured chat whose only message the agent's start refused read as a
green finish, and a cancelled structured turn did too: the host published a
verdict only for turn records, and structured rows carried no `interrupted`.

The host projection now reads the session's latest request: its turn's
outcome, or `failure` for a send the agent or its start refused. A send
that was withdrawn, or left undelivered by a restart or a close, fails
nobody and makes nothing listable. The ingest publishes `interrupted` as the
hook lanes do, and every reader decodes the verdict through one accessor, so
a failure reads Failed on the dot, the rollups, history and `worktree ps`,
behaves like a cancellation in every clean-finish policy, and notifies as
"failed".

* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now

* test(native-chat): a verdict change republishes the mobile status projection

* refactor(native-chat): the store's retention trigger keeps its flag compare

A verdict change always moves the completion clock the same check already
reads, so a second verdict compare there caught nothing new.

* test(native-chat): a user message the provider journaled keeps its session listed

* test(native-chat): pin what a failed start settles, and what a resume offer names

A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.

* test(native-chat): the failed-start pins fail on what the message became, not on a timeout

* fix(native-chat): a late provider-session update keeps a failed recovery record failed

A provider-session heartbeat that rewrites a completed recovery record kept
its interrupted flag but dropped the outcome it was copied with, so a live
failed checkpoint read as a clean finish until the next status write.

* test(orchestration): the preamble's host stub is typed, not cast

The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.

* test(native-chat): the terminal-bell check asserts the renamed verdict field

The bell notification test still checked for agentInterrupted, which no
longer exists, so it could not catch a verdict leaking into a bell dispatch.

* fix(native-chat): a failed turn ranks like a completion for attention

Attention readers (completion time, Smart Sort, sticky retention, Cmd+J
Recent) now demote only a turn the user stopped. A failure is news the
user has not seen, so it keeps its completion time, ranks in the Done
class, stays retained after its pane goes away, and a retained failure
reads failed in the worktree rollup instead of done. Clean-finish
policy (hibernation, pane ownership, the value moment) still treats a
failure like a stop.

The retention trigger compares verdicts again: success -> failure no
longer moves the completion clock.

* fix(native-chat): a failed main agent reads failed while its subagents still work

The verdict is now read from the main agent's own state, not the folded
row: a main agent that is done and failed has a verdict even while its
subagents keep the row working. Without mainAgent (history, worktree ps,
older hosts) the old combined-done rule stands.

Display marks the verdict through agentVerdictDisplayMark: a failure
outranks every combined state on the agent's dot, label, tab badge,
dashboard and activity rows; a stop marks only a done row, so a
successful or stopped main agent with live subagents still reads
working. Subagent rows keep their own state. The worktree card, terminal
tab and Cmd+J rollups share one pane fold and rank a pending question,
then failed, then working, monitoring, interrupted and done.

worktree ps publishes the main agent's outcome on a working row, and the
mobile mirror reads it. The store's change check, the paired-client
mirror's equality and its epoch now see a verdict change on a working
row, which otherwise moves no state or clock and left the worktree card
reading working. Clean-finish policy is unchanged: a working row is never
hibernated and has no completion time.

* docs(native-chat): the worktree ps outcome comment no longer claims old hosts send it

The field is new: an old host sends no outcome at all, so a reader falls
back to interrupted. The removed clause said old hosts send it on done
rows, which never shipped.

* docs(native-chat): the status-store listing rule names provider-journaled user messages

* fix(native-chat): a refused send notifies failed through the completion feed

The host's completion feed followed only the newest turn, so a send the
agent or its start refused, which creates no turn, read Failed on its row
but sent no notification. The feed now follows the session's latest
request, read from the projection the status feed already makes for the
commit: a turn keeps its id, a refused send is named by its journal item
key. It announces only while the session is idle, as the row reports a
verdict, so queued sends refused one commit at a time notify once, and a
withdrawn send falls back to a request already announced.

* fix(native-chat): every copy of a row carries the main agent's own status

History entries, sleep records and `worktree ps` rows carried a flattened
top-level `outcome`, copied under different gates and without the main agent's
clock. They now carry `mainAgent` (state, outcome, stateStartedAt), the type
the live row already persists and sends, and every copy site takes it with
`interrupted` through one function, `agentVerdictFields`.

- The accessor reads `mainAgent` then the legacy flag; the mobile mirror
  matches it line for line.
- Sleep records admit `mainAgent` with `normalizeMainAgentStatusField`, so a
  malformed value drops the field, never the record.
- Mobile dates a main agent that failed under live subagents by its own clock,
  as desktop does, and its row equality compares `mainAgent`.
- The activity feed reads a history entry's own `mainAgent` instead of
  rebuilding one; the sync key and history equality compare it.

* test(native-chat): pin the worktree ps verdict across host and phone versions

Pairs the real v1.4.212 host and phone row reader with this build: an old phone
reads a new host's rows by `interrupted`, a new phone reads an old host's rows
(no `mainAgent`) the same way, and a new phone reads a failure under live
subagents as Failed, dated by `mainAgent.stateStartedAt`. The release checkout
now carries the phone's self-contained row reader, and the lane runs when the
`worktree ps` row producers change.

* test(mobile): name the parity table's row for its role

* fix(native-chat): a request that settles while the user is asked something notifies once

The completion edge waited for an idle session, and a pending prompt (including a
subagent's approval) is not idle. Structured chat has no other attention producer,
so a main turn that finished while a subagent waited on the user sent nothing
until the prompt was answered.

The edge now waits only on owed work (a running turn or an unanswered send), which
the projection reports even beneath a pending prompt. A request that settles with
a prompt pending announces once; the renderer words it "needs input" from the
host status mirror's `attention`, and answering the prompt keeps the same request
identity, so it does not announce again. The wire shape is unchanged.

* fix(native-chat): the completion says when the user is being asked

A request that settles while a prompt waits on the user was worded "needs input"
from the renderer's status-feed mirror. Remote clients receive the status and
completion streams over separate sockets, so they can arrive in either order and
the wording could be wrong both ways.

The host already knows at emit time, so the completion now carries an optional
`awaitingUser: true` in that case and omits it otherwise. The renderer words the
notification from that field alone and no longer reads the status mirror. Old
clients ignore the field and word by outcome; old hosts never send it.

* fix(worktree-status): a departed agent's failure yields to live work on the worktree card

A retained failed agent has no expiry, so ranking it with a live failure pinned the card to Failed over other panes' live work. It now ranks below working, monitoring and permission, and above every finished outcome.

* docs(agent-status): a departed agent's failure ranks below live work on the worktree card

* fix(native-chat): a view never restarts a chat whose last start failed

A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.

* test(native-chat): start the child the loop waits on with an attach, not a second view

A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.

* fix(native-chat): settle a gone generation's turn wherever a conversation opens

A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.

* test(native-chat): prove the next child's start settles the turn an earlier child left

The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.

* test(native-chat): count a failed start's rows by row, not by text

Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.

* test(cross-version): load the phone row readers without mobile's toolchain

Vite transforms a file against its nearest tsconfig, and mobile/tsconfig.json
extends expo/tsconfig.base.json, which the root-only cross-version lane never
installs. The worktree ps verdict suite imported the current phone row reader
from mobile/ directly, so CI failed with TSConfckParseError before any test ran.

The harness now imports a copy of the working-tree reader placed under the
checkout cache, where the root tsconfig applies, as it already does for the
release checkout's copy. Both readers are still the real files.

* test(cross-version): keep the checkout path-guard message and justify the copy import's cast

* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget

* test(native-chat): pin the open's and the send's start and row counts, however the view binds

Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.

* fix(native-chat): settle a gone generation's turn at every open but an acquisition's

The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.

* test(native-chat): hold the create's start open until the views bind

The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.

* refactor(native-chat): drop the composer's second error formatter

After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.

* test(native-chat): pin the reason on a message rejected while its chat was closed

The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.

* fix(native-chat): a send the provider never received after a restart has no verdict

Restart reconciliation rejects a crash-stranded send that is absent from a
trustworthy provider history with reason 'not_delivered'. Nobody failed that
send, but the verdict allowlist did not name it, so after a crash the chat
read Failed, was listed, and could notify "failed". Give the reason a shared
constant (persisted value unchanged), add it to the no-verdict set, and treat
it as an internal marker so the Retry row no longer shows the raw string.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-27 22:23:49 -07:00
Alex-wangyang 47b8408e28 fix(zcode): wait for composer before first worker dispatch (#23374)
* fix: wait for ZCode composer before first worker dispatch

* fix: preserve readiness across renderer terminal adoption
2026-09-27 22:02:19 -07:00
Jinwoo Hong d33541c335 test(mobile): repin the RPC recording corpus to main after #22951 (#23535)
#22951 re-recorded the corpus with baseline set to its own branch commit,
which the squash left unreachable from main, so the RPC recording pin job
fails. Repin baseline to main's tip and re-record every golden; only the
baseline header moves.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-28 01:00:23 -04:00
Jinwoo Hong c8f9a65ff8 test(e2e): keep the Kitty-arming app alive in the Option-composed spec (#23495)
The host grounds Kitty flags a finished command left armed, so a bare
printf arm flipped back to 0 and raced the flags poll. Arm with a live
cat foreground and pass per-test flags into setup instead of re-arming.
2026-09-28 01:00:04 -04:00
Brennan Benson 934a2d44a0 fix(codex): a native chat's thread opens on the model the chat chose (#23532)
* fix(codex): a native chat's thread opens on the model the chat chose

* fix(codex): a resumed thread keeps its own saved model, provider and effort
2026-09-27 21:54:37 -07:00
Brennan Benson 78771646af fix(mobile): show one review sheet at a time so the review screen never freezes (#22951)
* fix(mobile): close the review sheet before opening Send Notes

On iPhone, Review Actions > Send Unsent Notes and Review Complete > Send Notes
opened the Send Notes sheet while their own sheet was still on screen. iOS
cannot present a second native sheet until the first has unmounted, so Send
Notes never appeared and every later tap on the review screen was swallowed
until the app restarted. Send Unsent Notes now uses the action sheet's
closeBeforePress, and the Review Complete drawer opens Send Notes from its
onAfterClose, the same sequencing the action sheet already uses.

* test(mobile): drive Send Unsent Notes through the real action sheet

The overflow test only checked the closeBeforePress flag, and nothing tested that the
action sheet actually defers such an action until it has closed. Press the real row and
assert Send Notes opens only from the sheet's after-close callback.

* fix(mobile): ignore a close request on a drawer that is already hiding

Android Back during a drawer's close animation restarted the hide, which
cancelled it, so the drawer never unmounted and its invisible Modal kept
swallowing every tap. Send Notes, which now opens after the review sheet
closes, never appeared either.

* refactor(mobile): give the review screen one sheet state so sheets cannot stack

The review screen kept five independent sheet flags that any caller could
set at any time. iOS cannot present a sheet while another is still on
screen (even mid-close), so any overlap froze every tap. Two openers could
still produce one: Review Complete appearing after the mark-reviewed save
landed on a sheet opened meanwhile, and the Send Notes list reopening a
sheet the user had already dismissed.

The five flags become one reducer that mounts at most one sheet. Switching
sheets closes the current one and shows the next only after its drawer
reports it has finished closing; Review Complete waits for the user's
sheet instead of closing it; a late send list only fills a Send Notes that
is still shown or queued. The two per-call-site sequencers this branch
added are removed in favour of it.

* test(mobile): repin the RPC goldens for the review sheet state

The review-actions recording adapter now drives the screen's sheet reducer
instead of the removed Send Notes setter, which moves adapterSha256 on the
14 goldens recorded through it. baseline is repinned to the refactor commit
so the recorder's product fence matches; no recorded body changed (only the
baseline and adapterSha256 header fields move).

* fix(mobile): settle a review sheet that closed before it was ever shown

A drawer mounts only once a commit shows it, so a sheet closed or displaced in
the same batch it opened in (e.g. Review Complete landing in the same frame as
a tap that opens another sheet) never sends onAfterClose. The sheet state then
waited on it forever and every later sheet on the screen stayed queued.

Track which sheet's drawer a commit actually showed and settle a closing sheet
that never reached the screen instead of waiting for a close it cannot send.

* fix(mobile): keep a queued Send Notes when a late Review Complete lands

A second Mark Reviewed save resolving while Review Complete was closing
toward Send Notes reopened Review Complete and dropped the Send Notes the
user had just asked for. A background opener now yields to any queued
user sheet, including behind a closing sheet of its own kind.

* refactor(mobile): present review sheets through one keyed drawer

iOS cannot present a native Modal while another is still presented, even
during its close animation. The review screen now renders all five sheets
through one KeyedBottomDrawer that alone decides what is presented: a
request for a different sheet hides the current one, and the next is
mounted only after its hide finished and a commit without any Modal has
landed. A request replaced before it was shown is never mounted.

The screen's sheet state now records only what the user asked for
(`requested`, plus a background Review Complete in `deferred`), so the
presented-sheet bookkeeping, the per-drawer close callbacks and the
screen-side mounted-sheet inference are gone.

BottomDrawer becomes a constant-key adapter over the same drawer, so the
app has one mount/close lifecycle. A hide that finished just before a
reopen is now ignored instead of latching, which used to swallow the next
close and leave an invisible Modal eating taps.

* test(mobile): repin the RPC goldens for the keyed review drawer

The review action adapter now drives the requested-sheet state, which
moves that family's adapterSha256 on its 14 goldens; the repin to the
keyed-drawer commit moves `baseline` on all 787. No golden body moved.
2026-09-27 21:38:29 -07:00
Jinwoo Hong 52dc32f9ea fix(mobile): the page pushes its terminal frame into the document from RN layout (#23079)
* fix(mobile): keep one mounted terminal frame under every session branch

The frame was keyed so react-native-web would attach its onLayout: a View
that gains onLayout after mount is never observed, and unkeyed the frame
reused the loading branch's View. One contentFrame View now wraps every
branch and carries the handler from the first mount, so no key is
needed. terminalFrame stays on the terminal branch, so nothing else is
clipped.

The parity pin moves: the key string leaves, one View and one style
reference arrive.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): push the page terminal's box into its document from RN layout

The page's document sized itself through a ResizeObserver on its host,
which duplicated RN layout and needed three rules in fit-scale: skip a
0-wide box, skip the last fitted box, and forget that box on any fit
request. The host View's onLayout now pushes into the mount, which is
the page's counterpart of the WebView's window resize.

react-native-web still lays a display:none screen out as 0x0 and its
return as the old box, so the mount treats neither as a change and keeps
answering the last real box while hidden, as a WebView keeps its size.
A fit asked for while hidden therefore lands at once, and fittedBox and
all three rules go. The zero-width wait in applyFitScale stays for a
host that mounts under a hidden screen before its first layout.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): hold a page terminal's fit while its host is hidden

A fit asked for under a covering screen read the last box, but where
xterm cannot measure cells on a display:none host (its DOM measure, with
no OffscreenCanvas) the retry loop ran out and committed scale 1, and
the show that followed was not a change, so 1 stuck.

The page's rect now says when the host is hidden, and the document holds
any fit asked for then as fitPending instead of committing. The mount
reports the same box coming back as 'shown', distinct from 'resized',
and the document runs a held fit on it and otherwise does nothing, so
pan and zoom still survive a plain hide and show. Native's window
resize reports 'resized' and is never hidden.

The mount also seeds its box from the host, so a first layout that
lands before the mount is not later mistaken for a resize. The hidden
text-scale case now asserts the resized column count, which stale
cells miss.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): read the page terminal's grid once xterm has sized its cell

The init render check read the grid the moment .xterm-screen existed,
but xterm sizes its one-cell helper textarea only on a cursor move or
resize, after the replay drains. Under full-suite load the read won
that race and measured a 0-wide cell. In the failing runs the fit had
committed at 390/560 on the cell-width gate, so the page was right and
the read was early. It now waits for a sized cell: 8/8 under four-way
parallel load, where it was 4/8.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): pin that a held page terminal fit runs once

A held fit that lands on show must be spent: a second hide and show
re-running it would reset the pan and zoom the user set in between. The
case now hides and shows again and expects no new transform, which a
commit that stops clearing fitPending fails.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): hold the refit a narrow hidden text-scale change owes

A text-scale change whose new cells leave fewer than MIN_FIT_COLS in
the box skips the grid resize and returned before any fit. Base cleared
fittedBox there so the next box refit; with that gone, a hidden host
shown at the same box kept the old scale under the larger font. The
branch now asks for the fit while the host is hidden, which holds it
until show. A shown host, and every native one, returns as before.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): fit a text-scale change too large to resize the grid

A visible viewport too narrow for the new cells (280 px at 200%, 18
columns) skipped the resize and returned without a fit, so the larger
text overflowed the fit made for the smaller cells. The branch now fits
unconditionally; applyFitScale already holds the fit while hidden.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-28 00:06:35 -04:00
Brennan Benson 2119730ec0 fix(terminal-wait): unattended launches report Claude's trust dialog instead of timing out or typing into it (#22927)
* fix(terminal-wait): recognise Claude's workspace trust dialog as a blocking prompt

Claude's first-launch "trust this folder?" dialog parks the cursor above its
options, and the host's line tail drops the lines below it, which are the only
ones the trust matcher knew ("trust this folder", "Enter to confirm"). An
unattended Claude launch into a fresh folder therefore waited out its whole
budget and reported a timeout instead of the blocking prompt. The dialog's
opening question ("... one you trust?") survives in the tail, so it is now
recognised, pinned by a captured transcript of the real dialog.

* fix(terminal-wait): read the rendered screen for blocked prompts on the tui-idle poll

Claude's workspace-trust dialog parks the cursor on its highlighted option with a
cursor-up, and the host line tail deletes every retained row below the cursor, so
"Yes, I trust this folder" and "Enter to confirm" never reach the blocked-prompt
detector. In a live launch the dialog also arrives in 1024-byte reads, and the tail's
plain path blanks each line that ends in a carriage return before the newline, so
the opening question does not survive either. An unattended launch waited out its
whole budget and reported a timeout.

The runtime already feeds every PTY chunk into its own headless emulator. The tui-idle
poll now also runs the existing blocked-prompt rules over that emulator's visible
screen (no provider or host round trip), after the tail checks and before the
quiet-foreground idle fallback. The "one you trust" phrase is dropped: it only matched
when the whole dialog arrived in one chunk, wrapped away on narrow panes, and could
match prose.

Replays three live Claude 2.1.280 captures through the runtime: the dialog in one
chunk and in 1024-byte reads, a 60-column pane, and the dialog answered with "Yes",
which must report ready rather than blocked.

* fix(terminal-wait): skip the rendered-screen blocked check while the agent reports working

The screen check runs on every tui-idle poll, and the automation observer holds a tui-idle
wait open for a whole agent turn. A working Claude whose screen showed dialog wording (a diff
of the detector, say) was reported blocked where main kept waiting. The dialogs only the
screen reveals are start-up ones painted before any title, so a working title now vetoes it.

* fix(terminal-wait): settle weak tui-idle evidence only after a clean screen read

A shell auto-title (oh-my-zsh's `claude`, fish's `claude <cwd>`) names Claude before
its workspace trust dialog paints, and the line tail loses that dialog. The wait took
the bare name as rest, settled ready, and the launch typed its brief into the dialog.

Every tui-idle settle site now asks one evaluator for a verdict: blocked, strong ready,
working, weak ready or pending. The pre- and post-registration checks and the
title-change resolvers settle only blocked or strong ready; weak ready is left to the
poll, which settles it only once the rendered screen shows no blocker. A name-only
Claude title is held to the same quiet window as Codex and Devin, since Claude
announces rest with its own explicit title.

* test(serialize): record the answered Claude trust capture's known serializer divergences

The captured answered-dialog transcript added by this PR is replayed by the
serialize round-trip suite and diverges at 13 checkpoints. It diverges
identically on origin/main and on the pre-#22586 addon build (13 both-fail,
0 regressions): the live SGR pen leaks into the alt buffer, and an alt buffer
first entered after a shrink keeps hidden scrollback. Pin the count like the
other known captures and correct the comment, which called these upstream.
2026-09-27 20:59:02 -07:00
Jinwoo Hong cdf9f6e938 fix(mobile): native desktop-mode keyboard lift stays on screen; metrics carry the row pitch (#23070)
* fix(mobile): the session hears a row-pitch change in the keyboard metrics

Carried from 2eab7633013 and 62113b7ac30 with the `rowPitch` metric they need (from
f7d78405c7d, without its page-only lift). The document reports the drawn row pitch, cell height
times fit scale, and reports again when a fit commits a new scale. The metrics handler compared
four fields and dropped an update that moved only the pitch, so desktop display mode kept the
phone's 15 px pitch; it now compares every field.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): native desktop display mode lifts only what the keyboard hides of the grid

Carried from 097f5d8d33f, native rule only. Desktop display mode draws 40 rows about 270 dp
tall, and a lift by the covered strip (312 on Pixel_API_37) moved the whole grid under the
header. The lift is now also capped by the hidden strip of the drawn grid. A sweep shows it
equals base's arithmetic whenever the drawn grid fills the frame, which phone mode always does.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): pin the desktop-mode anchor and the fit commit's re-emit by behaviour

A caret mid-grid in desktop display mode lifts only enough to clear its row, which the anchor
term decides; without it the lift was the whole hidden strip and no test noticed. The fit
commit's re-emit is now asserted by committing a fit and reading the pitch it reports, rather
than by the order of two lines in the source.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): give the fit test's viewport the whole rect type

The tests typecheck ratchet refused the double: the seam's rect carries `left` and `top` too.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): report keyboard-avoidance pitch after a pinch release

A pinch moves the drawn row pitch through userScale. A release onto the same
text-size preset skipped the refit, so no metrics were emitted and a write made
mid-gesture left the lift reading the transient pitch. applyTextScale now says
whether it scheduled the refit, and the release emits when it did not.

The text-scale frame also emitted a transient pitch after its resize, before
the fit it schedules commits and emits the final one. Drop that emit so a
preset change reports once.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): report text-scale pitch only through the fit commit

The pinch release emitted metrics itself whenever applyTextScale said no
refit was scheduled. A release onto a preset a settings change had just
applied saw an unchanged font size and emitted the pre-refit pitch, one
frame before the refit reported the real one.

A same-size applyTextScale now schedules the fit too, so every text-scale
change reports through commitFitScale. A pending refit supersedes that fit
by token. applyTextScale returns nothing and the release emits nothing.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-27 23:55:21 -04:00
Brennan BensonandClaude bfe476f922 fix(native-chat): a message is accepted, then delivered (#22821)
* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* fix(native-chat): name a chat write by its target, not the owner generation

A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.

Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.

Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.

* fix(native-chat): every journal append reaches the chats that are open

A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.

A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.

* test(native-chat): an epoch replacement reaches the open chat

* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

* perf(native-chat): a publish behind a delivered commit reads nothing

Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.

* test(native-chat): state why the teardown test's fake journal is safe to cast

* docs(native-chat): say mutation admission checks only the writer lease

* docs(native-chat): drop the send rebase from comments that still described it

* fix(native-chat): a message is accepted, then delivered

A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".

A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.

Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.

A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.

Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.

* fix(native-chat): settle queued messages only for the child that ended

A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.

A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.

The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.

* fix(native-chat): an adoption that fails to import keeps the conversation open

The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.

* perf(native-chat): the recovering open reads the journal once

Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.

* fix(native-chat): an attach that fails after indexing its child leaves no child behind

A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.

* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer

The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.

A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.

* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down

The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.

* fix(native-chat): a message rejected while its chat was closed reads as not sent

A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.

* test(orchestration): name why the readiness settlement fakes are cast

* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent

* docs(native-chat): drop the fence from the admission the send effects run behind

* docs(native-chat): give the fence move on release the reason that still holds

* docs(native-chat): stop citing a write fence check in launch and mailbox comments

Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.

* refactor(native-chat): the provider child is its own record

A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.

- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
  patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
  own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
  the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
  dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
  is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
  the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.

* fix(native-chat): the delivery loop alone settles a message its start or child failed

A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.

- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
  reads how it ended: a Stop continues; anything else writes one failure row and rejects every
  queued message with the same words, then stops. A child still starting whose start the adapter
  says did not land fails the same way. The exit, eviction and the settlement retry only settle
  the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
  nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
  failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
  closed, with or without a child, and a start the loop already has in flight is waited for so the
  child it produces is stopped rather than left behind.

* refactor(native-chat): a stopped child ends on the one reading of its stop

The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.

* feat(native-chat): the host says it accepts a send before any agent has it

The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.

* refactor(native-chat): an attach never opens a journal of its own

The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.

* fix(native-chat): a moved fence resends nothing on a host that accepts first

The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.

The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.

* refactor(native-chat): a child's end says whether the user or the host stopped it

The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.

* fix(native-chat): a chat whose only work is a queued message is not offered for resume

A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.

* test(native-chat): type the queued-message fixtures in the resume-offer tests

* fix(native-chat): a start that dies while a message waits on it is that message's failed start

Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.

* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now

* test(native-chat): pin what a failed start settles, and what a resume offer names

A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.

* test(native-chat): the failed-start pins fail on what the message became, not on a timeout

* test(orchestration): the preamble's host stub is typed, not cast

The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.

* fix(native-chat): a view never restarts a chat whose last start failed

A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.

* test(native-chat): start the child the loop waits on with an attach, not a second view

A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.

* fix(native-chat): settle a gone generation's turn wherever a conversation opens

A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.

* test(native-chat): prove the next child's start settles the turn an earlier child left

The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.

* test(native-chat): count a failed start's rows by row, not by text

Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.

* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget

* test(native-chat): pin the open's and the send's start and row counts, however the view binds

Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.

* fix(native-chat): settle a gone generation's turn at every open but an acquisition's

The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.

* test(native-chat): hold the create's start open until the views bind

The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.

* refactor(native-chat): drop the composer's second error formatter

After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.

* test(native-chat): pin the reason on a message rejected while its chat was closed

The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-27 20:38:00 -07:00
Jinwoo Hong 992bad5375 refactor(agent-resume): give each resume-note reader its own named check (#23497)
* fix(terminal): stop a resume note from blanking a live remote agent pane

A host-mirrored remote terminal pane holding a client resume note with
origin 'live' and state 'done' (the idle anchor kept after every finished
turn) loaded blank and looped attach/disconnect several times a second.

- Mirrored web-terminal tabs ignore client resume notes at load; the host
  answers liveness at connect, so a note must not divert the attach.
- The empty-reattach retire rule never fires for remote PTYs: a remote
  disconnect only closes this viewer's stream, so the retry lands on the
  same live PTY and loops.
- The handler gets its own sleep-evidence check with the pre-#16308
  meaning, so the wake sweep's widening no longer leaks into it.

* test(agent-resume): pin how a finished turn's idle anchor reaches parking and the hibernation wake

#16308 let a live done note count as passive for the wake sweep. The park
exemption and the suppressed-exit wake shared that check without being
reviewed for it: a finished turn's idle anchor now lets its hidden tab park,
and any suppressed exit over it arms an in-place resume on reveal. Pin both
before the check is split so the refactor cannot change them silently.

* refactor(agent-resume): give each resume-note reader its own named check

isPassiveCompletedHibernationEvidence answered two different questions after
#16308 widened it for the wake sweep. Replace it with checks named for what
each reader asks, same answers as before:

- isFinishedTurnOwingNoResume: the wake sweep, background wake, preserved-pane
  ownership and the park exemption (a finished turn owes no resume).
- noteArmsHibernatedPaneWake: the suppressed-exit wake arm and the mobile wake
  latch, which must agree with each other.

The reattach handler keeps isHibernationDoneRecord from #23491.

* refactor(agent-resume): name the activation check for what it decides

isFinishedTurnOwingNoResume also returned true for worktree-sleep done
notes, which a pane mount does resume from (#9648). Its callers all ask
whether workspace activation leaves the note alone.
2026-09-27 23:06:36 -04:00
Jinwoo Hong a8797138ef fix(terminal): stop a resume note from blanking a live remote agent pane (#23491)
A host-mirrored remote terminal pane holding a client resume note with
origin 'live' and state 'done' (the idle anchor kept after every finished
turn) loaded blank and looped attach/disconnect several times a second.

- Mirrored web-terminal tabs ignore client resume notes at load; the host
  answers liveness at connect, so a note must not divert the attach.
- The empty-reattach retire rule never fires for remote PTYs: a remote
  disconnect only closes this viewer's stream, so the retry lands on the
  same live PTY and loops.
- The handler gets its own sleep-evidence check with the pre-#16308
  meaning, so the wake sweep's widening no longer leaks into it.
2026-09-27 22:52:06 -04:00
Brennan Benson 0b2dd0a99b fix(runtime): let the phone end a terminal stream by the request that opened it (#23006)
* fix(runtime): register phone terminal subscriptions when the request arrives

* fix(runtime): don't let an already-aborted terminal subscribe take the stream slot

* fix(runtime): let the phone end a terminal stream by the request that opened it

* fix(mobile): describe the relay sibling check accurately on this base
2026-09-27 18:51:14 -07:00
Brennan Benson bc7a17fef2 fix(runtime): register phone tab-list streams when the request arrives (#23045)
* fix(runtime): register phone tab-list streams when the request arrives

* fix(runtime): keep a worktree-wide tab unsubscribe off later subscribes

Tab-list subscribes now register when they arrive, so an older phone's
worktree-wide session.tabs.unsubscribe (no request id) could end a same-worktree
subscribe that arrived while the unsubscribe was still resolving the worktree.
Capture the registration version when the unsubscribe is dispatched, as
terminal.unsubscribe already does, and sweep only streams registered by then.

* test(runtime): cover a request abort during tab-list stream setup
2026-09-27 18:27:18 -07:00
Brennan Benson bf9194126e fix(mobile): release streams whose ready arrives after a replayed cancel (#22945)
* fix(mobile): release streams whose ready arrives after a replayed cancel

* test(mobile): cover a replayed browser stream replaced before its ready
2026-09-27 18:21:30 -07:00
Jinjing 5a8237a7a5 Browser tab placement review (#23462)
* Implement source-following placement for browser tabs

- Add afterTabId and executionHostId parameters to track and anchor new tabs after their source
- Resolve source browser pages to unified tab wrappers for background link opens
- Implement anchor-based insertion that respects pinned tab boundaries
- Stage source-adjacent rows before host RPC for paired browser creation
- Ensure duplicated tabs and background links place directly after their source
- Add comprehensive test coverage for placement scenarios across single/split groups

* Preserve tab strip scroll anchor when tabs are added

Keep the viewed tab stable on screen when other tabs are inserted around it.
Records the active tab's on-screen position before insertions and restores it
after, so the user's focused tab doesn't jump unexpectedly. Matches the behavior
of VS Code and Chrome.

* Refine browser tab placement: fix host handling and sort-order gaps

- Avoid reordering unified tabs when the computed order matches the current order
- Do not substitute execution host for browser tab wrappers; let createUnifiedTab apply its active-workspace fallback instead
- Always apply sort values after tab insertion to prevent sortOrder gaps from preview replacement
- Fix import path for paired browser tab creation and clarify registration comments

* refactor(tests): use paired-browser-tab-creator registry for browser tab

- Replace direct web-runtime-session mock with paired-browser-tab-creator pattern
- Add test-rig file naming convention to localization audit skip list
2026-09-27 18:08:09 -07:00
Neil dd90b3183f fix: identify partial-clone repositories from the correct remote (#23504) 2026-09-27 18:05:13 -07:00
Brennan Benson 6c56c0a3dc fix(browser): a failed SSH route keeps its card while the host redials (#23465)
* fix(browser): a failed SSH route keeps its card while the host redials

A browser route that already failed swapped its "SSH connection unavailable"
card for "Connecting" on every dial of its host, and re-ran prepare once per
dial cycle, so the card flickered and its buttons detached mid-click. Only a
route that is still preparing now waits on a dialing host; a failed route
keeps its card until the host actually connects, which re-derives it.

Retry and Try anyway land on preparing together with the new attempt, so a
press while the host dials waits for the connect instead of starting a
prepare the effect immediately cancels.

The escape-hatch e2e now makes the host truly unreachable before the
disconnect; it passed before only because the card stayed latched over a
host the terminal had already reconnected.

* fix(browser): an unrouted SSH route waits for a dialing host too

Only a failed or ready route is exempt from the host wait; a route that just
became routed (or still shows another target's page) has no answer for this
target, so it must not start a prepare that its own preparing write cancels.

The escape-hatch spec now reads the settled failure from the renderer store:
main's ssh:getState drops the entry on disconnect and on a failed connect, so
its status is null there and never matches the failure pattern.
2026-09-27 17:57:21 -07:00
Brennan Benson d6336be8db ci(e2e): run the SSH browser route e2e when its source changes (#23498)
A change to the SSH workspace browser route, its gate card, or the host-connection
phase it waits on selected no e2e specs, so the Docker SSH browser spec never ran
on the PR that changed it.
2026-09-27 17:51:16 -07:00
Brennan Benson 7a3735fcd2 fix(native-chat): let a multi-question ask record list its questions (#23451)
* fix(native-chat): let a multi-question ask record list its questions

A cancelled or pending structured ask with several questions showed only
"Asked: 2 questions", and its questions appeared nowhere in the transcript.
The row's subject now carries the question texts instead of a count, and
the row unfolds them as a list below its toggle unless the record already
lists them with their answers.

* fix(native-chat): keep a question list's open state off its first question

A pending Codex group is keyed by its first question, which becomes its own row once answered. Opening the list left that row expanded, and folding it could drop the toggle under the pointer. The list and a lone question now remember their open state separately.
2026-09-27 17:21:47 -07:00
Brennan Benson 89cf55dfc8 fix(agent-launch): report a launch that failed before spawning as failed, with its cause (#22913)
* fix(agent-launch): settle a launch that failed before spawning as failed, with its cause

A launch into an existing workspace whose terminal create threw before the
spawn request left this process (agent disabled, no launch command, runtime
unavailable) created nothing, yet agent.launchReplay recorded and answered it
as agent_session_operation_unknown. createTerminal now reports when it hands
the spawn to the pty controller; a failure before that point settles the
ledger row as failed and returns the original error. After the request leaves,
the outcome stays unknown: an SSH or daemon spawn whose reply was lost may
still have started.

* test(agent-launch): expect the spawn-dispatch hook on the launch's terminal create

* test(agent-launch): drive the pre-spawn failure with a missing launch command

A disabled agent is moving to a check made before either launch route runs,
so the tests use a failure that stays inside the terminal build.

* fix(agent-launch): keep the not-started verdict on the launch, not the shared error

A failed pane spawn rejects the same error object into the spawner (after its
request left) and into a concurrent create waiting on that pane (before its
own). Marking the error object globally let the waiting launch's verdict clear
the spawner's, recording a launch that may have started an agent as failed.
The launch now owns its dispatch tracker and carries the decision on its
execution error instead of re-deriving it from the error.

Also names the test's launch parameter type for the anti-slop audit.

* refactor(agent-launch): move launch failure classification into its own module

Rebasing onto the caller-selection change took agent-launch.ts past its line
limit; the failure-code helpers are a self-contained concern.
2026-09-27 16:37:04 -07:00
Brennan Benson da57f47353 fix(native-chat): say a refused chat write in plain words, and keep it with the write it describes (#22999)
* fix(native-chat): say a refused chat write in plain words, and keep it with the write it describes

A send's refusal now lives on the queued message and goes when that message is sent again or delivered, so a resend that succeeds after an agent restart no longer leaves a red line under the composer. Stop, answers, option and goal changes report a refusal once as a toast; a conversation command answers inline. One shared table turns every refusal code into copy with a next step, on desktop and mobile, instead of showing the host's diagnostic text.

* test(native-chat): pin the refused-then-restarted send in the order the live app sees it

* fix(native-chat): tell the phone to send a refused message again, not to press Retry

The phone puts a refused message back in the composer and has no Retry
control, so "Retry to send it again." named something that is not there. A
phone send now reads "Your message was not sent. Send it again."

Also types the rejection-cause test's empty submissions without a cast.

* fix(native-chat): keep why a message failed as a fact, and say only what is true

A queued message saved the words of its failure to local storage, so the copy
lived in users' data. It now saves the fact: the refusal code (with the host's
words only for a failed restart, which the host writes for people), the
provider's rejection reason, or that the host could not be reached. The Retry
row chooses the words when it shows the message. A saved failure this build
cannot read is dropped, and the row says only that the message was not sent.

The words are true for every host path behind each code:
- A refusal whose code does not say why (a cleared conversation, a pending
  question, a provider's own rejection) says only what did not happen, with no
  next step that could repeat the refusal.
- An unsettled owner no longer claims the agent was restarting; it says Orca
  could not confirm which agent process owns the chat.
- A code from a newer host says only what did not happen.

Desktop translates each sentence whole, with the shared English as the
fallback the phone shows as is, so the two never say it differently. The phone
no longer shows transport or host text when a request fails without a
refusal.

* fix(native-chat): never show a host refusal message, including a failed restart's

Every refusal code has at least one host path that writes its message for a
log or carries a marker. A replayed refusal, for one, says "Operation <id> was
already refused: <code>." because the ledger stores no message. A failed
restart's message also embeds the resume's own refusal text, which can be the
ledger's. So no code's message is safe to show, and the failed-restart
exception goes. The Retry row now says "The agent couldn't restart. Your
message was not sent." with no next step, because some restarts need a new
chat. The cause is still in the chat's own status row. The saved failure keeps
only the code.

The test lists one such emitter per code, so a code that later becomes safe to
show has to be argued against that list.

* fix(native-chat): offer only a next step that works, on the phone too

A message the agent turned away on the phone said "Retry to send it again", but the phone has no Retry control; it now says "Send it again." like every other refused phone message, built from the same sentences as the other notices.

The capacity refusal said Orca was handling too many requests "for this chat" and offered a retry. The limit is counted across every chat over a day, so an immediate retry would likely be refused again. It now says so, with no next step.

* chore: restore pnpm-lock.yaml

A local install rewrote it and it was committed by mistake.

* fix(native-chat): say only what is true for every host path behind a refusal

No refusal notice says how to try again any more: the control that sent
the write already does that, and a sentence per code was false for some
of the host situations the code covers. The phone keeps "Send it again."
where a resend under the same operation id can go through, and an older
host still gets "Update Orca, then try again."

A cause is named only for codes whose every emitter means it. The owner
codes and identity_required now say only what did not happen, the
outcome-unknown sentence no longer claims the write itself is in doubt
(a send behind an unconfirmed rewind is refused outright), and a request
that failed without a refusal is recorded as `failed` rather than
`unreachable`, since most such failures are host or compatibility errors.
The census test now pins the cause allowlist and the absence of retry
sentences, and the unused sentences leave every catalog.

* fix(native-chat): don't say a chat write failed when its request may have run

A desktop Stop, answer, option, goal or conversation command whose request threw
said the write did not happen, for every error. A timeout or a lost connection can
come after the host ran the write, so a timed-out /compact said "The command didn't
run." while it ran. Only an RPC error the host answers before running the method
proves that; any other failure now says Orca couldn't confirm what happened.

The phone already drew this line; both clients now read it from one shared rule.

* fix(native-chat): say a refused background-task stop is about the task, not the agent

A background-task stop goes through agentSession.cancel, so a refusal of it
said "The agent wasn't stopped." although the agent was never asked to stop.
The write kind now reads the cancel's scope and names the task or tasks.

* fix(native-chat): say a Stop refused over a moved-on question did not stop the agent

A Stop pressed while a question or approval is pending names that prompt, so the host can refuse it
as already resolved or stale. The notice then spoke only about the question and never said the agent
kept running.

* test(native-chat): check every refusal notice cell against one spec

Replaces the per-rule loops with one table of every failure (each refusal
code, a failed or unconfirmed request, and a code from a newer host) by
every write kind, and five rules each cell must keep: a certain refusal
says that write did not happen, once; a write that may have run never
says it did not; only "Update Orca" and the phone's resend where it can
work say to try again; a cause is named only for allowlisted codes; no
cell is empty or shows the host's text.

* fix(native-chat): the launch path drops a message's old failure when it sends it again

The first message of a new chat is sent by the launch path, which staged it
without removing the failure saved by an earlier attempt. A message that failed
through the outbox and was then refused again through the launch path showed the
earlier reason in its Retry row. Both paths now stage through one shared step
that removes it.
2026-09-27 16:34:42 -07:00
NeilandClaude b61797dce3 fix(agent-trust): bound the remaining Codex trust writes a launch waits on (#23380)
* fix(agent-trust): bound the remaining Codex trust writes a launch waits on

#23148 capped the trust write on the agentTrust:markTrusted IPC handler and the
worktree-remote startup path, but three call sites still awaited
markCodexProjectTrusted with no bound: the main-process worktree startup used by
orchestration worker-start, Codex quick-launch preparation, and Codex session
resume preparation. All three share the per-config.toml lane with hook installs
and app-server trust grants, which has no cap on queue depth, and the SSH writer
adds a resolveHome round trip plus unbounded SFTP over a possibly half-open
link. A wedged lane left the user clicking "start agent" with nothing happening
and no error.

Each site now routes through the existing awaitAgentTrustWriteWithinDeadline
helper with unchanged semantics for its two current callers. Abandoning at the
deadline never cancels the write, writes nothing else, and never substitutes a
local fallback, so the workspace stays untrusted and Codex raises its own trust
prompt. The await still sits ahead of the runtime-home resolution and hook
repair that the PTY spawn waits on, so the ordering #23148 established is kept.

Reliability only; there is no speed gain. The win is that a launch cannot hang.

Verified with 2,755 agent-trust/Codex/startup/runtime tests, the 53-test
reliability-gate command, the gate manifest check, typecheck:node, changed-code
quality, oxlint and oxfmt. Red-green: with the three sites reverted to a bare
await the new bounding tests hang to the 30s vitest timeout (3 fail/14 pass);
restored, 17 pass.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(reliability-gates): stop the trust-preflight gate claiming a bounded launch

The gate's performance budget said a stuck predecessor "degrades to the
agent's own trust prompt instead of an indefinite caller wait." That is false
at two of this branch's three new sites. In startup/codex-launch-preparation.ts
the next awaited call after the abandoned write is
codexHookService.prepareRuntimeHomeForLaunch; in
startup/codex-session-resume-launch.ts it is installForLaunchPrep or
refreshRuntimeUserHooksForLaunchPrep. Both reach
runExclusivelyForRuntimeAndSystemTrustConfig, which takes the same
runExclusivelyForCodexTrustConfig queue - FIFO per config.toml, no depth cap
and no timeout - on the runtime home and on ~/.codex/config.toml that
markCodexProjectTrusted itself takes. A wedged lane still stalls those two
launches one step later, with the abandoned write holding its queue slot ahead
of the hook step. Only markLocalWorktreeTrusted has nothing after its Codex
branch and so is bounded end to end.

The budget, invariant, oracle and coverage notes now say what is true: all five
markCodexProjectTrusted call sites are bounded, so no trust write hangs a
caller indefinitely, but the bound is on the write and not on the launch. A new
knownGaps entry records the residual stall and that the new launch-prep tests
mock codexHookService wholesale, so nothing in these suites can catch it. The
already-complete-write case stays labelled a timer control, not a red.

Wording only; no production code, test or count changed. The cited command
still reports 6 files / 53 tests, and the gate manifest check passes for 138
gates.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(reliability-gates): name the first unbounded re-entry, not the hook-service call

The trust-preflight gate cited codexHookService.prepareRuntimeHomeForLaunch as
the immediate next step after the abandoned write in
startup/codex-launch-preparation.ts. Tracing the module, the first awaited step
is ensureRealHomeHooksIfSelected. It is conditional: only when the target is not
WSL and runtimeHome.isHostSystemDefaultRealHomeSelected(launchEnv) is true does
it call ensureRealHomeCodexHookState, which chains behind that module's own
serial ensure promise and then acquires runExclusivelyForCodexTrustConfig
directly on ~/.codex/config.toml - the same lane markCodexProjectTrusted's inner
acquire takes. When the real home is not selected that call returns without
awaiting the lane, runtimeHome.prepareForCodexLaunchAsync is synchronous on the
non-WSL path, and only then is prepareRuntimeHomeForLaunch the first re-entry.
So the gate could miss a launch stall one step earlier than the one it named.

startup/codex-session-resume-launch.ts had the same omission: its hook branch
takes ensureRealHomeCodexHookState when the resume home is ~/.codex, and only
otherwise installForLaunchPrep or refreshRuntimeUserHooksForLaunchPrep. The
coverage notes also understated the mocking - codex-launch-trust-write-deadline
mocks ensureRealHomeCodexHookState as well as codexHookService and defaults the
real-home selection to false, so neither lane is driven by any suite here.

The performance budget, coverage notes and the residual-stall knownGap now name
the first re-entry on each path and say when it applies, rather than only the
later hook-service call.

Wording only; no production code, test or count changed. The cited command
still reports 6 files / 53 tests, and the gate manifest check passes for 138
gates.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(reliability-gates): scope the trust-preflight invariant to the local writes

The invariant claimed every Codex trust write a launch waits on is bounded or
abandoned at a deadline. The entry's own knownGaps contradicted that: the remote
write reached through markRemoteWorktreeTrusted in
runtime/runtime-worktree-agent-startup.ts is awaited with no timer, and it is a
Codex write, not an adjacent one - it reads TUI_AGENT_CONFIG[agent].preflightTrust,
which is 'codex' for the codex agent, and markRemoteAgentWorkspaceTrusted then
branches into markRemoteCodexProjectTrusted. Its one production caller,
markRemoteWorkspaceTrustedForAgent, is reached with a connectionId from seven
launch entry points; five await it directly before the createTerminal that spawns
the agent, and two gate the launch command they hand back for their caller to
spawn. It chains a session.resolveHome round trip plus SFTP realpath/read/mkdir/
write, none with a timeout, over a possibly half-open link - the same hang this
branch bounds locally.

The invariant now claims only the five local markCodexProjectTrusted sites and
names the remote exception inline instead of leaving it to knownGaps. Audited the
rest of the entry against the code in the same pass: performanceBudget said "no
trust write can hold its caller indefinitely" (now scoped to those five); the
"all five awaited Codex trust writes" tally is now "all five awaited local" and
points at the unbounded sixth; the remote gap no longer calls that path merely
"preset-agnostic and out of Codex scope"; oracle and coverageNotes now record
that nothing here drives markRemoteWorktreeTrusted, since the remote-preset
suite only checks what markRemoteAgentWorkspaceTrusted writes, never how long it
may take; and surfaces gains the remote writer that suite actually covers.

Re-verified the rest rather than assuming it: the launch-prep re-entry chain,
the FIFO no-cap no-timeout trust-config queue, the IPC handler capping both its
remote and local writes with no cross-branch fallback, and that
upsertProjectTrustLevel and upsertProjectTrustLevelInContent remain the only
producers of a project trust_level, so no further writer needs covering. The
already-complete-write case stays labelled a timer control, not a red.

Wording only; no production code, test or count changed. The cited command still
reports 6 files / 53 tests, and the gate manifest check passes for 140 gates.

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-27 16:30:30 -07:00
Jinjing 29847641ab test(e2e): fix failures and improve stability (#23480)
- Narrow toolbar width and set window minimums for consistent testing
- Add node_modules symlink to fixture for ESM import resolution
- Exercise manual paging and fix button selector
- Preserve repo filters in reveal workflow
- Adjust timing strategy and increase test timeout
2026-09-27 16:17:06 -07:00