Commit Graph
773 Commits
Author SHA1 Message Date
Jinwoo Hong 1762a138f7 feat(mobile): slide the page's host stack on push and Back (#24268)
expo-router's Stack on web renders native-stack's web view, which flips display and ignores animation. The page's host stack now keeps expo-router's StackRouter under its public Navigator and draws the slide with the Web Animations API; a popped screen stays mounted until it has slid out. Native is a pure move.
2026-10-01 01:19:35 -04:00
Jinwoo Hong e9ec63168f fix(mobile): hold-to-dictate, repeat keys and the browser long-press survive the page's long-press (#24277)
On the OTA page a held press died ~500 ms in: the WebView's long-press selected nearby text and that selection's selectionchange/touchcancel ended the press. Page text is now unselectable unless it opts in (as native), hold surfaces declare onLongPress, the browser pane refuses contextmenu termination, and the chat mic's swapped icons no longer steal the touch target.
2026-10-01 01:18:24 -04:00
Brennan Benson 3ab3c9239f fix(native-chat): the working line shows only what the agent is doing now (#24218)
* fix(native-chat): the working line shows only what the agent is doing now

A chat's live "Working…" line could show an old notice, such as "Claude hit a
temporary problem and is retrying.", long after the agent had moved on and was
running new commands. When the host had no live activity for the turn, the line
fell back to the newest status row in the turn, and any status row qualified:
retry warnings, "Context compacted", "Cancellation requested.", and notification
summaries. Those rows record the past and already appear in the transcript.

The line now reads only the host's live, per-turn activity, which is never saved
and is cleared at turn boundaries. Without it, the line says Thinking or
Working…. Desktop and mobile share the selector, so both change.

* test(codex): guard that a subagent's compaction never becomes the parent's live activity
2026-09-30 16:44:29 -07:00
Brennan Benson b4b708c2c4 fix(native-chat): a Codex stream retry is one warning row that updates in place (#23684)
* fix(native-chat): a provider's own retry progress is quoted in its retry row

A retry row whose fact carries a detail the provider wrote for a person now
quotes it, the same way a rejected message or failed compaction does, so the
row says how the retry is going. A log detail still stays out of the sentence.

* fix(native-chat): a Codex stream retry is one warning row that updates in place

An error Codex says it will retry used to fall through to the generic frame
row: red, and a new row for every attempt. It now writes one providerRetrying
row per retry run, warning-toned, revised by each attempt with Codex's own
progress sentence. A run is the retry frames of one turn with nothing else the
thread journals between them; every attempt still publishes, so the idle sweep
keeps seeing activity. Errors Codex will not retry are unchanged.

* fix(native-chat): a Codex retry row says it is retrying and keeps the frame behind Details

The quoted retry sentence leads with "is retrying", which holds for any
provider's progress text. The Codex retry row also keeps the whole bounded
frame behind the row's Details, as the generic row did, so Codex's
additionalDetails stays available to diagnose a retry.

* test(native-chat): a Codex retry re-handled after backpressure keeps its one row

Pins the run being opened before the write: a first attempt whose publish is
refused and is handed back must revise the row it already wrote, not open a
second run. Also stops the fixture claiming Codex sends an idle thread status
beside each retry, which the app server does not do.

* perf(native-chat): a Codex frame with no retry run open is not classified

Ending a retry run classified every non-retry frame, and classifying walks the
whole payload: every streaming delta, and every large item/completed, paid a
walk about as costly as parsing the frame. Only a thread with a run open needs
the answer, so the classification now runs only there.

* fix(native-chat): each Codex retry attempt is its own warning row, with what failed on its second line

A stream error Codex says it will retry is written as its own warning row
with a providerRetrying fact, under the same per-frame identity every
Codex frame row gets. The host no longer tracks retry runs or rewrites
one row in place, so there is no run state to open, end or clear, and no
frame has to be classified to end a run. Every attempt publishes, which
keeps renewing the idle clock while Codex retries.

Codex's additionalDetails, which its own UI shows under the progress
message, is kept on the fact as the retry's cause and printed on the
row's second line. Errors Codex will not retry are unchanged.

* fix(native-chat): a transcript draws only the latest row of a provider retry run

The shared structured message projection, which both the desktop and the
mobile transcript read, collapses a run of retry rows into its latest
row. A run is retry rows from the same agent with no other drawn row
between them; a row that draws nothing, or a queued send drawn after the
conversation, does not split it. The earlier attempts stay in the
journal.

* test(native-chat): the retry-run render test uses the message list's current props

* fix(native-chat): a Codex frame row is named for its connection, so a later one never revises it

Frame rows were named provider-frame:codex:<n> from a counter that starts
over with every connection, so the first frame row after a reconnect in the
same session revised an earlier connection's row in place, at its old spot.
Each connection's frame rows now carry the acquisition generation, minted
once before the translator is built: provider-frame:codex:<generation>:<n>.
Rows already written keep their identities.

* fix(native-chat): agents retrying at once each keep one row, read from the row's own agent

* test(native-chat): a reconnect's Codex rows are named for the acquisition that received them

* fix(native-chat): a retry run is one agent's, so another agent's row never splits it

Each agent's rows are drawn apart: the session's own rows are the conversation,
and a subagent's rows open in that subagent's section. Splitting a run on any
other row in the flat list left two adjacent retry rows on screen whenever
another agent wrote between two attempts: a subagent finishing a command while
the session reconnected, or the session working while a subagent reconnected.

A run is now per agent: an agent's retry rows with none of its own other rows
between them, drawn as its latest.

* test(native-chat): a subagent's retry run is checked in its own section, and the run rule's words say same agent

* test(mobile): the retry-rows test typechecks, so the test ratchet keeps checking it
2026-09-30 15:05:22 -07:00
Brennan BensonandClaude 24edf0f64b fix(agent-status): a turn a crash cut off reads Interrupted, an unproven end Couldn't confirm (#23467)
* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* fix(native-chat): name a chat write by its target, not the owner generation

A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.

Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.

Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.

* fix(native-chat): every journal append reaches the chats that are open

A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.

A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.

* test(native-chat): an epoch replacement reaches the open chat

* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

* perf(native-chat): a publish behind a delivered commit reads nothing

Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.

* test(native-chat): state why the teardown test's fake journal is safe to cast

* docs(native-chat): say mutation admission checks only the writer lease

* docs(native-chat): drop the send rebase from comments that still described it

* fix(native-chat): a message is accepted, then delivered

A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".

A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.

Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.

A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.

Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.

* fix(native-chat): settle queued messages only for the child that ended

A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.

A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.

The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.

* fix(native-chat): an adoption that fails to import keeps the conversation open

The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.

* perf(native-chat): the recovering open reads the journal once

Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.

* fix(native-chat): an attach that fails after indexing its child leaves no child behind

A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.

* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer

The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.

A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.

* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down

The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.

* fix(native-chat): a message rejected while its chat was closed reads as not sent

A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.

* test(orchestration): name why the readiness settlement fakes are cast

* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent

* docs(native-chat): drop the fence from the admission the send effects run behind

* docs(native-chat): give the fence move on release the reason that still holds

* docs(native-chat): stop citing a write fence check in launch and mailbox comments

Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.

* refactor(native-chat): the provider child is its own record

A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.

- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
  patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
  own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
  the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
  dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
  is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
  the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.

* fix(native-chat): the delivery loop alone settles a message its start or child failed

A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.

- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
  reads how it ended: a Stop continues; anything else writes one failure row and rejects every
  queued message with the same words, then stops. A child still starting whose start the adapter
  says did not land fails the same way. The exit, eviction and the settlement retry only settle
  the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
  nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
  failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
  closed, with or without a child, and a start the loop already has in flight is waited for so the
  child it produces is stopped rather than left behind.

* refactor(native-chat): a stopped child ends on the one reading of its stop

The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.

* feat(native-chat): the host says it accepts a send before any agent has it

The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.

* refactor(native-chat): an attach never opens a journal of its own

The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.

* fix(native-chat): a moved fence resends nothing on a host that accepts first

The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.

The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.

* refactor(native-chat): a child's end says whether the user or the host stopped it

The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.

* fix(native-chat): a chat whose only work is a queued message is not offered for resume

A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.

* test(native-chat): type the queued-message fixtures in the resume-offer tests

* fix(native-chat): a start that dies while a message waits on it is that message's failed start

Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.

* fix(native-chat): a request that failed reads as failed

A structured chat whose only message the agent's start refused read as a
green finish, and a cancelled structured turn did too: the host published a
verdict only for turn records, and structured rows carried no `interrupted`.

The host projection now reads the session's latest request: its turn's
outcome, or `failure` for a send the agent or its start refused. A send
that was withdrawn, or left undelivered by a restart or a close, fails
nobody and makes nothing listable. The ingest publishes `interrupted` as the
hook lanes do, and every reader decodes the verdict through one accessor, so
a failure reads Failed on the dot, the rollups, history and `worktree ps`,
behaves like a cancellation in every clean-finish policy, and notifies as
"failed".

* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now

* test(native-chat): a verdict change republishes the mobile status projection

* refactor(native-chat): the store's retention trigger keeps its flag compare

A verdict change always moves the completion clock the same check already
reads, so a second verdict compare there caught nothing new.

* test(native-chat): a user message the provider journaled keeps its session listed

* test(native-chat): pin what a failed start settles, and what a resume offer names

A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.

* test(native-chat): the failed-start pins fail on what the message became, not on a timeout

* fix(native-chat): a late provider-session update keeps a failed recovery record failed

A provider-session heartbeat that rewrites a completed recovery record kept
its interrupted flag but dropped the outcome it was copied with, so a live
failed checkpoint read as a clean finish until the next status write.

* test(orchestration): the preamble's host stub is typed, not cast

The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.

* test(native-chat): the terminal-bell check asserts the renamed verdict field

The bell notification test still checked for agentInterrupted, which no
longer exists, so it could not catch a verdict leaking into a bell dispatch.

* fix(native-chat): a failed turn ranks like a completion for attention

Attention readers (completion time, Smart Sort, sticky retention, Cmd+J
Recent) now demote only a turn the user stopped. A failure is news the
user has not seen, so it keeps its completion time, ranks in the Done
class, stays retained after its pane goes away, and a retained failure
reads failed in the worktree rollup instead of done. Clean-finish
policy (hibernation, pane ownership, the value moment) still treats a
failure like a stop.

The retention trigger compares verdicts again: success -> failure no
longer moves the completion clock.

* fix(native-chat): a failed main agent reads failed while its subagents still work

The verdict is now read from the main agent's own state, not the folded
row: a main agent that is done and failed has a verdict even while its
subagents keep the row working. Without mainAgent (history, worktree ps,
older hosts) the old combined-done rule stands.

Display marks the verdict through agentVerdictDisplayMark: a failure
outranks every combined state on the agent's dot, label, tab badge,
dashboard and activity rows; a stop marks only a done row, so a
successful or stopped main agent with live subagents still reads
working. Subagent rows keep their own state. The worktree card, terminal
tab and Cmd+J rollups share one pane fold and rank a pending question,
then failed, then working, monitoring, interrupted and done.

worktree ps publishes the main agent's outcome on a working row, and the
mobile mirror reads it. The store's change check, the paired-client
mirror's equality and its epoch now see a verdict change on a working
row, which otherwise moves no state or clock and left the worktree card
reading working. Clean-finish policy is unchanged: a working row is never
hibernated and has no completion time.

* docs(native-chat): the worktree ps outcome comment no longer claims old hosts send it

The field is new: an old host sends no outcome at all, so a reader falls
back to interrupted. The removed clause said old hosts send it on done
rows, which never shipped.

* docs(native-chat): the status-store listing rule names provider-journaled user messages

* fix(native-chat): a refused send notifies failed through the completion feed

The host's completion feed followed only the newest turn, so a send the
agent or its start refused, which creates no turn, read Failed on its row
but sent no notification. The feed now follows the session's latest
request, read from the projection the status feed already makes for the
commit: a turn keeps its id, a refused send is named by its journal item
key. It announces only while the session is idle, as the row reports a
verdict, so queued sends refused one commit at a time notify once, and a
withdrawn send falls back to a request already announced.

* fix(native-chat): every copy of a row carries the main agent's own status

History entries, sleep records and `worktree ps` rows carried a flattened
top-level `outcome`, copied under different gates and without the main agent's
clock. They now carry `mainAgent` (state, outcome, stateStartedAt), the type
the live row already persists and sends, and every copy site takes it with
`interrupted` through one function, `agentVerdictFields`.

- The accessor reads `mainAgent` then the legacy flag; the mobile mirror
  matches it line for line.
- Sleep records admit `mainAgent` with `normalizeMainAgentStatusField`, so a
  malformed value drops the field, never the record.
- Mobile dates a main agent that failed under live subagents by its own clock,
  as desktop does, and its row equality compares `mainAgent`.
- The activity feed reads a history entry's own `mainAgent` instead of
  rebuilding one; the sync key and history equality compare it.

* test(native-chat): pin the worktree ps verdict across host and phone versions

Pairs the real v1.4.212 host and phone row reader with this build: an old phone
reads a new host's rows by `interrupted`, a new phone reads an old host's rows
(no `mainAgent`) the same way, and a new phone reads a failure under live
subagents as Failed, dated by `mainAgent.stateStartedAt`. The release checkout
now carries the phone's self-contained row reader, and the lane runs when the
`worktree ps` row producers change.

* test(mobile): name the parity table's row for its role

* fix(native-chat): a request that settles while the user is asked something notifies once

The completion edge waited for an idle session, and a pending prompt (including a
subagent's approval) is not idle. Structured chat has no other attention producer,
so a main turn that finished while a subagent waited on the user sent nothing
until the prompt was answered.

The edge now waits only on owed work (a running turn or an unanswered send), which
the projection reports even beneath a pending prompt. A request that settles with
a prompt pending announces once; the renderer words it "needs input" from the
host status mirror's `attention`, and answering the prompt keeps the same request
identity, so it does not announce again. The wire shape is unchanged.

* fix(native-chat): the completion says when the user is being asked

A request that settles while a prompt waits on the user was worded "needs input"
from the renderer's status-feed mirror. Remote clients receive the status and
completion streams over separate sockets, so they can arrive in either order and
the wording could be wrong both ways.

The host already knows at emit time, so the completion now carries an optional
`awaitingUser: true` in that case and omits it otherwise. The renderer words the
notification from that field alone and no longer reads the status mirror. Old
clients ignore the field and word by outcome; old hosts never send it.

* fix(worktree-status): a departed agent's failure yields to live work on the worktree card

A retained failed agent has no expiry, so ranking it with a live failure pinned the card to Failed over other panes' live work. It now ranks below working, monitoring and permission, and above every finished outcome.

* docs(agent-status): a departed agent's failure ranks below live work on the worktree card

* fix(native-chat): a view never restarts a chat whose last start failed

A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.

* test(native-chat): start the child the loop waits on with an attach, not a second view

A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.

* fix(native-chat): settle a gone generation's turn wherever a conversation opens

A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.

* test(native-chat): prove the next child's start settles the turn an earlier child left

The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.

* test(native-chat): count a failed start's rows by row, not by text

Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.

* test(cross-version): load the phone row readers without mobile's toolchain

Vite transforms a file against its nearest tsconfig, and mobile/tsconfig.json
extends expo/tsconfig.base.json, which the root-only cross-version lane never
installs. The worktree ps verdict suite imported the current phone row reader
from mobile/ directly, so CI failed with TSConfckParseError before any test ran.

The harness now imports a copy of the working-tree reader placed under the
checkout cache, where the root tsconfig applies, as it already does for the
release checkout's copy. Both readers are still the real files.

* test(cross-version): keep the checkout path-guard message and justify the copy import's cast

* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget

* test(native-chat): pin the open's and the send's start and row counts, however the view binds

Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.

* fix(agent-status): a turn a crash cut off reads Interrupted, an unproven end Couldn't confirm

When the provider gave no verdict, the structured status projection now derives
one from the newest turn's lifecycle: interrupted -> interruption, unverifiable ->
unconfirmed. Nothing new is journaled, the completion feed stays provider-only, and
the legacy interrupted flag stays a user stop only. Every verdict reader handles
both arms explicitly.

* fix(native-chat): settle a gone generation's turn at every open but an acquisition's

The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.

* fix(native-chat): a folded turn a crash cut off reads Interrupted after N

The settled-turn timing now carries the turn's verdict, derived by the same
agentTurnVerdict the status row uses. A turn that ended interrupted with no
provider verdict heads its fold 'Interrupted after N'; a user's stop keeps
'Worked for N'.

* test(native-chat): hold the create's start open until the views bind

The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.

* fix(native-chat): the user's close of a chat records the turn it cuts short as their cancellation

The expected-close settle writes outcome cancellation when the user aimed the
stop at this chat: a Stop while the agent starts, the chat's tab closed (the
agentSession.close RPC, or session.tabs.close with a user reason), or /clear.
A quit, an idle eviction, a worktree teardown or an orchestration stop leaves the
turn with no verdict, so it still reads Interrupted.

* test(native-chat): a Claude turn a newer send superseded reads Interrupted

The supersede fires for any send Orca dispatched, the user's or another agent's,
and nothing at that site records the sender, so the turn keeps no verdict and
folds as Interrupted after N.

* fix(native-chat): the user's close records cancellation on the turn the provider settled on its way out

The Codex and Claude adapters settle their open turn as interrupted, with no
verdict, while the host stops them. The expected-close settle then found no
running turn, so a user's close of a mid-turn chat read Interrupted. The stop
now reads the running turns before it reaches the provider and records the
user's cancellation on each one it cut short, unless the provider gave a
verdict of its own. The close-verdict test's adapter now settles its turn on
close the way the real adapters do.

* test(native-chat): update the close and settled-turn expectations for the host-observed verdict

agentSession.close now passes the user's word to the host, and a settled
interrupted turn with no provider verdict carries `interruption`. Also merge a
duplicate import the code-quality gate rejects.

* refactor(native-chat): drop the composer's second error formatter

After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.

* test(native-chat): pin the reason on a message rejected while its chat was closed

The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.

* fix(native-chat): the user's close cancels a turn whose start landed as the provider stopped

The close read which turns were running before the stop. A turn whose start was
still in flight (a send echo not yet journaled) was absent from that read, so the
provider's verdict-less settle on the way out left it Interrupted. The close now
reads which turns were already over instead, and records the user's cancellation
on every other turn the stop left running or interrupted with no verdict. A turn
cut off earlier, or finished during the stop, keeps its end.

* fix(native-chat): a send the provider never received after a restart has no verdict

Restart reconciliation rejects a crash-stranded send that is absent from a
trustworthy provider history with reason 'not_delivered'. Nobody failed that
send, but the verdict allowlist did not name it, so after a crash the chat
read Failed, was listed, and could notify "failed". Give the reason a shared
constant (persisted value unchanged), add it to the no-verdict set, and treat
it as an internal marker so the Retry row no longer shows the raw string.

* refactor(native-chat): the adapter settles the turn a stop cuts with the stop's typed cause

The host hands its stop's cause to the adapter's close. Each adapter settles its own
open turn on 'ended' through one mapping, turnVerdictForChildEnd, and the host's
dead-generation fallback uses the same mapping for any turn no adapter settled. A
user's close or stop of this chat is their cancellation; a quit, eviction, teardown
or an exit the adapter saw first is news.

Deletes the snapshot-and-diff reconstruction (endedTurnItemIds,
userStoppedTurnRevisions, settleUserStoppedTurns) and the requestedByUser flag.
host.close now takes a required cause.

* test(native-chat): a user's close drops the chat's status row like an eviction

* test(native-chat): expect the eviction cause on the host closes of idle release, worker stop and worker discard

The stop's typed cause now travels into host.close and the adapter's close, so these
three non-user closes assert the 'evict' they pass.

* refactor(native-chat): every stop names its cause, so none defaults to the user's cancellation

stopStructuredAgentSessionAgentUnderSerialize defaulted its ending to 'user-stop', which now
settles the cut turn as the user's cancellation. Every caller already passes a cause; the
parameter is now required, and a type-level test fails to compile if the default returns.

* test(native-chat): pin who a chat's session.tabs.close speaks for, older clients' reasonless close included

The mapping lived inline in a type-unchecked file, and only the explicit user reason had a test: an older client's reasonless close, or a lifecycle echo read as the user's, stayed green. It is now one exhaustive, type-checked function with a case per reason.

* style(mobile): draw the unconfirmed dot in the theme's status amber, not an inline hex

* docs(agent-status): the main agent's outcome also carries the host-observed end, interruption or unconfirmed

* fix(native-chat): a chat the user closed while its agent started is not a failed start

A still-starting child the user's close cut counted as a failed start, since only 'user-stop' was excluded: a start-failure row, and queued messages rejected as a provider failure. Whether an ending fails its start is now one exhaustive switch, shared by the failed-start read and the delivery loop's handover: a user's stop or close never does; an exit, a failed attach and the host's own stops still do.

* fix(native-chat): a chat the user closed closes its queued messages, and starts no agent for them

After 876b6989f1 a user's close of a still-starting chat went on like a Stop, so when the close did not complete the delivery loop started a new agent for the message queued behind it. A child's end now has three dispositions, not a failed-start boolean: a user's Stop lets the queue go on, the user's close closes what was queued before it, and any other end fails it. The close is the one a completed close does (the provider-closed rejection, no verdict), applied at the top of each delivery step and ordered against the close so a later send still goes on.

* fix(activity): a crash-cut turn draws the interrupted glyph; only a user's Stop keeps the done check

The Activity page drew every Interrupted row with the done check, which #2569 chose for a user's Stop. With a crash now reading Interrupted, that put a green check on a turn nobody asked to stop. The row's glyph is now an exhaustive switch over the verdict: a cancellation keeps the done check, and an interruption draws the existing interrupted dot. An unconfirmed end already drew its own glyph.

* fix(activity): the Interrupted group header draws the done check only when every row is a user's Stop

A user's Stop and a crash share the Interrupted group, and its header took its newest row's glyph, so a Stop newer than a crash put a green check over the crash. The header is now folded over the group's rows: the done check only when every row draws it, the interrupted dot otherwise.

* fix(native-chat): a failed close of what the user closed starts no agent for it

Rejecting the messages a user's close left queued swallowed a journal write failure, so the delivery step went on to start an agent for a message in a chat the user closed. The rejection now reports whether it landed, and a step whose rejection failed stops instead; the next wake re-derives and retries it. Also pins that the ordering against the close holds only within its epoch, since a later epoch's sequences restart.

* test(native-chat): the idle sweep's stop is an eviction, so its close carries that cause

Main's idle sweep now stops an idle agent through the conversation lifetime, which this branch gives the 'evict' cause; its expectations name it.

* fix(native-chat): a retried stop keeps the cause of the stop it finishes

A user's Stop or close whose wind-down failed after the child was proven gone was finished by the idle sweep as an eviction, so the turn it cut read Interrupted. The owed wind-down now carries its stop's cause, and a retry with no child settles with it.

* test(native-chat): the idle sweep's close of a retrying Claude chat carries the eviction cause

Main's new test expected the adapter close with the session id alone; every stop now names its cause, and the idle sweep's is 'evict'.

* fix(status): a user's Stop marks done on the tab and sidebar; red Interrupted is only a turn cut short by something else

The tab, the worktree card and the sidebar rows drew a Stop with the same red dot as a crash. The
verdict mark now maps a cancellation to done, still saying "Interrupted by user" in the row text,
and the mobile mirror follows. The Activity page keeps grouping a Stop under Interrupted with the
done check, as before.

* test(cross-version): a new phone reads a user's Stop as done; an old phone still draws it interrupted

* fix(native-chat): a user's Stop inside a live Claude chat reads as their cancellation

Stopping a running Claude turn interrupts it and keeps the session, so the turn's end comes from
the CLI's result frame. Claude CLIs before 2.1.91 send that frame with no terminal_reason, and later
ones may still omit it, so the user's own Stop was recorded as a failure with an error row.

Orca now records the stop on the open turn when it sends the interrupt. An error result for that
turn reads as the user's cancellation whatever reason the CLI gives. The stop belongs to that one
turn, so it cannot reach the next, and it is withdrawn when the CLI refuses the interrupt.

* docs(agent-status): a user's stop marks done; name the tab close cause by its type

The reference still said a stop marks a row interrupted and ranks between live work and an
unconfirmed end. A cancellation now marks done, and only a turn cut short by something else ranks
as interrupted. The runtime's tab close restated the close cause's union; it now uses the type.

* test(native-chat): a proven crash reads as an interruption on the status feed and in the chat

A crash the relaunch proves now settles its turn interrupted, and the status feed works the verdict
out from that record, so the restart test expects interruption for a proven crash and unconfirmed
for one it cannot prove, never a cancellation. A chat read before the proof lands reports
unconfirmed, then interruption and a folded "Interrupted after 27s" once the proof revises it.

* fix(status): a user's Stop reads Interrupted, and a turn anything else cut short reads Failed

The verdict mark now maps a cancellation, the user's own Stop, to interrupted, and an interruption,
a turn cut short by a crash or a killed agent, to failed, the same as a failure, which outranks live
subagent work. An unconfirmed end is unchanged. This applies to every agent, in a terminal or a chat,
on the tab, the sidebar rows and worktree card, the dashboard row, Cmd+J and the phone. A Stop is
not news, so the rollups rank it below an unconfirmed end, and notifications word an interruption
"failed". Recording is unchanged.

* fix(activity): group a user's Stop under Interrupted and a crash with failures

A user's Stop draws the interrupted glyph and sits alone in Interrupted, and a turn anything else cut
short sits in Failed, titled "Agent failed". Every row in a status group now draws the group's own
glyph, so the header is the group's status and the rule that folded a Stop's done check into the
header is gone. Interrupted ranks below an unconfirmed end, as in the sidebar.

* fix(status): draw a user's Stop in the muted tone, not the fault red

The interrupted dot, which now means only a user's Stop, draws in the muted foreground token on the
agent rows, the sidebar card and the phone. Red stays for a failure or a turn cut short by anything
else, and green for a finish.

* fix(native-chat): fold a stopped turn as "Interrupted after N" and a failed or crash-cut one as "Failed after N"

The settled turn header now follows the verdict mark: a user's Stop reads "Interrupted after N", and
a failure or a turn anything else cut short reads "Failed after N", under the new key
components.native-chat.status.failedAfter in all six catalogs and the boot catalog. Desktop and
phone share the one description, so they agree.

* docs(agent-status): describe the Interrupted and Failed marks

The reference and the phone's turn bar still described a user's stop as done and a crash as
interrupted. A fault now reads failed, a user's stop reads interrupted in the muted tone, and the
rollups rank an unconfirmed end above a stop.

* test(status): a crash the relaunch recovers marks failed

The recovery test still expected a recovered interruption to mark interrupted; it now marks failed,
as a failure does. Formatting only elsewhere.

* fix(native-chat): record a turn a newer request replaced as superseded, and show it Interrupted

A Claude turn that a newer send replaced before its result arrived was recorded as interrupted with
no verdict, which reads as a turn cut short by something else, now "Failed". It is now recorded with
its own outcome, `superseded`, where the replacement is detected. That outcome names no sender, so a
dispatch from another agent is never recorded as the user's Stop, and it sets no legacy flag.

Every reader handles it in an exhaustive switch: it draws the muted Interrupted mark with the plain
text "Interrupted", folds as "Interrupted after N", and attention demotes it with a Stop, through
the renamed agentTurnEndedOnRequest. Older builds read an arm they do not know as no verdict, which
is what this turn carried before, so their rows keep reading done; the cross-version suites pin an
older desktop's journal and status readers and an older phone.

* refactor(status): name the attention predicate for a turn ended on purpose

agentTurnEndedOnRequest becomes agentTurnEndedOnPurpose: a user's Stop or a newer request's
replacement, never a fault. The Claude turn-end comment no longer says a replaced turn carries no
verdict.

* test(native-chat): a turn cut off by a restart or by quitting Orca reads Failed after N

On main a restart-cut turn shows the done tick. Pin the chat's turn bar and
the tab's mark for both cuts, through the recovery settlement and the quit's
child-end mapping, and pin the quit's turn bar through the host's own quit.

* style(native-chat): format the superseded turn-bar expectations

* fix(claude): a Stop that names no turn is the user's stop of the open turn

The chat's Stop button names no turn. Claude's conversation Stop recorded the
user's stop only for a named turn, so an older CLI's error result after that
Stop read Failed. It now records it on the open turn through the same intent,
dropped when Claude refuses the interrupt and never carried to the next turn.

* fix(native-chat): keep the attach context's publishStatus required

The lifetime context type makes publishStatus optional, so the attach context
that spreads it no longer satisfied its own type once its duplicate
publishStatus went. The host's lifetime context is now inferred, and checked
with satisfies, so the spread carries the member it always sets.

* fix(activity): rank a user's Stop below live work in the status grouping

The Activity page's status grouping put the Interrupted group (a user's Stop, or a
turn a newer request replaced) above Working and Monitoring, so a Stop still sorted
like news there while the sidebar, worktree card and Cmd+J rank it below live work.
It now follows live work and stays above Done; Failed and Couldn't confirm keep
their places above live work.

* docs(agent-status): say which turn outcomes the journal records and which are derived

The journal now records superseded as well as the provider's verdict and a stop;
interruption and unconfirmed are derived on read. The resume row no longer claims
interrupted renders red.

* test(native-chat): the idle sweep's held-send rest closes with the evict cause, like its siblings

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-30 15:02:44 -07:00
Brennan Benson afa81dc3ad fix(native-chat): chat failure messages appear in the app's language (#23674)
* fix(native-chat): a read whose history will not open is refused with its reason

History, subscribe, snapshot and options reads reach a chat through one accessor, whose open had no
catch: a journal that would not open reached every client as a runtime error carrying the storage's
own text (a path, "file is not a database"). The accessor, and the options read's own open, now throw
the classified journal refusal: journalCorrupt when SQLite reports damage, journalUnavailable
otherwise. The storage text goes to the log only.

The wire code stays runtime_error and the message becomes the bare code, as for every thrown
refusal; the reason rides in the error's data.

* refactor(native-chat): the idle sweep's stop of a hung start carries no hand-written reason

The sweep passed an English sentence as the stop's reason. It lived only in memory and nothing read
it: the delivery loop words the error row and the rejection from the hostStopped fact. Dropped, with
the display-name lookup that built it.

* fix(native-chat): a refusal the host throws is worded from its data, never its message

A thrown agent-session refusal reaches the client as runtime_error with the bare code as its message
and the typed refusal in error.data. Stop, answers, options and goals, a launch's held option pick,
the option picker's failure toast, and the Retry line of a chat that could not start now word it
from that refusal through the shared notice table. What each caller decides about the outcome is
unchanged: only the words move.

The Retry line of a failed start no longer prints the host's message or a thrown error's text; it
keeps the refusal as a fact and says the cause and step its reason names, or only that the chat
could not be started. The option toast keeps a local option surface's own sentence.

* fix(native-chat): an unreadable history is worded from its refusal, and damage stops the retry

The structured chat's read failure showed the host's text on the status line, and the pane always
said Orca keeps trying. The read transport now takes the refusal from the error's data (a stream
payload or a thrown RPC error), the reducer keeps it beside the failure text, and the pane and the
status line word it through the notice table, once: on the pane when the failure took it, else
beside the transcript that stays.

A damaged journal says "Unable to load this chat." and the read stops reconnecting for that run;
reopening the chat reads again. An open that can clear names its cause without "Try again", since
the pane retries on its own. A failure that names no reason keeps today's generic line. Finality
comes from the refusal's reason, never its message, which is the bare code for both.

* fix(native-chat): a rejected message is worded from its stored fact

A message the host recorded and then rejected keeps the host's typed fact beside its reason, but
the Retry words re-read the reason alone. Now the fact decides: a hand-over failure says Orca
couldn't reach the agent, a kind whose reason may be a legacy marker gets its fact's own sentence
(a full queue now says so instead of only "not sent"), and any other kind shows the sentence the
host wrote for it, which carries the agent's name and any words the provider wrote for a person. A
row with no fact reads as before.

* fix(native-chat): each message that did not go through says why on its own row

The structured chat showed one Retry strip under the transcript for whichever single entry it
picked, so a second failed message had no reason and no Retry of its own. The terminal-backed
chat's existing per-row delivery marker now carries a notice and an optional Retry, and the
structured pane derives one per message from the outbox on each render: every rejected message,
and the one the queue stopped on (read through the drain's own rule, so a Retry never names a
message waiting behind it). Each is worded from that message's stored failure. The single strip is
deleted. Nothing new is stored, and the shared message projection is untouched.

* fix(native-chat): say each chat failure's words where its own control already acts

Three wording rules for the desktop chat:

- A chat that could not start shows Retry beside its reason, so the reason stops at its cause
  where the Retry is the step: a reason whose action is to retry, and a start failure's "send
  your message again". Any other step stays (quit the terminal agent, start a new chat). The
  start-failure sentences take a retryControl context for this; what the host writes is unchanged.
- A history that couldn't open right now still reconnects, so the pane keeps "Orca keeps trying
  to load it" under its cause. Only a damaged history, which no retry reads past, drops it.
- A read failure that names no reason while the transcript is shown is only the pane
  reconnecting: the status line says "Reconnecting to this chat…" in muted text, not an error.
  New key components.native-chat.state.reconnecting, hand-translated for es/fr/ja/ko/zh.

* fix(native-chat): a rejected message offers Retry only once the queue is moving

Each rejected message's row offered its own Retry even while the queue was stopped on another
message. Any Retry clears the stopped queue, so pressing a rejected message's Retry also sent the
message the queue was holding, which the user had not retried; behind a message whose delivery is
unconfirmed, the retried one instead went back into the queue with no notice and waited there.

While the queue is stopped, only the message it stopped on offers Retry, as the single Retry strip
this replaced did. A rejected message keeps its words on its row and gets its Retry back once the
queue moves.

* fix(native-chat): a chat whose history will not open logs once, not on every reconnect

A reader reconnects every 750 ms while a journal open can clear, and each attempt logged the
failure with its full stack. The read door now logs a session's failure once until that session
opens, closes, or fails differently; every attempt is still refused with its reason.

* fix(native-chat): a message's own Retry is its resend step, so its notice stops at the cause

A rejected message offering Retry read "Claude stopped before it finished starting. Send your
message to try again." beside that button. Its row now takes the rule the launch strip already
follows: beside its own Retry the words leave out sending or trying again, worded from the stored
fact with the chat's agent name. The stored fact keeps less than the host wrote from (a refusal,
the provider's words), so a reason it cannot rebuild exactly is kept as written. A rejected
message without a Retry, while the queue is held, keeps the step. What the host writes and the
phone's notice are unchanged.

* fix(native-chat): a message's Retry sends only that message, never the one the queue is held on

Retry released the queue's refusal hold whichever message it was pressed on. While a queued message waited ahead of a held one, every rejected message offered Retry, and pressing it also sent the held message the person had not retried. Retrying an unconfirmed message ahead of a held one did the same. Retry now releases the hold only for its own message.

* fix(native-chat): a not-signed-in failure beside Retry still says to sign in first

Beside a Retry the notice dropped the whole next step, so a chat that could not start because the agent was not signed in read only the cause. Pressing Retry without signing in fails the same way again. The words now keep the sign-in step and leave out only the resend, which the Retry button is.

* test(native-chat): a rejected message's hidden Retry only avoids waiting unseen

* fix(native-chat): every Retry beside a notice leaves out the retry step the same way

A message the queue stopped on worded its refusal with no agent name and with its retry step, beside its own Retry, while a rejected message next to it named the agent and left the step to the button. The launch strip and the history pane each had their own copy of the same rule. One wording context now goes through the one notice table for every surface: a Retry beside the words, or a pane that reconnects on its own, is the step for a reason whose action is to retry, and every other step stays. What the phone and the host write is unchanged.

* test(native-chat): read the sent message id without a type assertion

* fix(native-chat): a rejected message is worded from the journal's own fact, never by comparing sentences

A message the host recorded and then rejected kept only the rejection's kind on the message, so its notice was reworded from that smaller copy only when it rebuilt the host's sentence word for word. A different agent name, an older host's wording, or anything the copy dropped (why a start failed, the provider's own words) left the host's sentence in place, beside a Retry that repeated its resend step. The notice now reads the journal's own rejection for that message, found by id, with the pane's agent name and Retry, and shows the provider's words only when they were written for a person. The message's smaller copy words it only when that journal row is not loaded. Nothing new is stored.

* fix(native-chat): a message rejected before a restart retries under a new id the first time

Whether a Retry needed a new message id was remembered in memory for one message, or read from the journal row when it was loaded. After a restart, or for an older message whose row was not loaded, the first Retry resent under the old id, the host answered with the same settled rejection, and nothing visibly happened. The message now says so itself: one the host recorded and rejected always retries under a new id, including after a restart. A refusal that already gave the message a fresh id, and a message whose delivery is unconfirmed or in flight, keep their id as before.

* fix(native-chat): a rejected message older than the loaded history keeps the provider's words

When the journal row that rejected a message is not loaded, the message's own copy of the
fact has no provider detail or start refusal. For the kinds worded from those, the row now
shows the sentence the host wrote for the person instead of a thinner rebuilt one.

* test(native-chat): the chat pane words a rejected message from its loaded journal row

Nothing covered the pane handing the journal's rows to the per-message notices, so a pane that stopped passing them would quietly fall back to the message's smaller copy of the rejection and show the host's sentence, resend step and all. The new case renders the pane with a rejected message whose journal row is loaded and checks that it reads that row's refusal in the chat's own agent name.

* fix(native-chat): a chat whose history won't load says why in one line

A read the host refused for a named reason put its sentence under the generic
"Could not load conversation" title, so a damaged history read as two lines
saying the same thing. The pane's own sentence now takes the title's place; a
history that can come back keeps its line saying Orca keeps trying. A failure
that names nothing keeps the generic title.

* fix(native-chat): a message a failed start rejected says only that it was not sent

When an agent stopped before it finished starting, the chat showed the start's
red row ("Claude stopped before it finished starting. Send your message to try
again.") and then repeated that cause under every message the start rejected.
Each of those messages now reads "Your message was not sent." beside its Retry.

The match is made on typed facts, not on the words: the host writes the start's
row and the rejection of its queued messages from the same failure fact, and the
row is keyed by the start. The pane finds the loaded start-failure rows by that
key and shortens a message's notice only when its loaded journal submission was
rejected with the same fact. Any other rejection, or one whose submission or row
is not loaded, keeps its full notice. The row key moves to a shared module so the
host that writes it and the pane that reads it use one definition.

* fix(native-chat): a chat whose history keeps failing to open retries less often

A read the host kept refusing (its history store could not be opened right now)
reopened every 750 ms for as long as the chat stayed open, about 40 opens every
30 seconds. Each reconnect now waits twice as long as the last, from 750 ms up
to 30 seconds, and never gives up; the first read that delivers anything starts
the wait over at 750 ms. A damaged history still stops reconnecting at once.

Reset happens on a delivered read, not on connect: a local subscribe resolves
before the host's open refuses, so resetting there would keep the 750 ms loop.

* test(native-chat): the pane harness types its journal rows without a cast

* test(native-chat): the admission test passes no start-failure rows to the notices

* test(native-chat): import the journal types once

* fix(native-chat): a remote chat reads again as soon as its host is back

The read retry doubles its wait up to 30 s during an outage, and nothing
reset it when the remote runtime reconnected, so the transcript could
lag the reconnect by up to 30 s. The read now watches the runtime
status store's contact-regained edges (hostContactEpoch for a
same-runtime return, connectionGeneration for a new runtime session)
and, when one lands, runs a waiting retry immediately with the wait
reset to its base.

* refactor(native-chat): build each failure sentence from whole pieces

Every sentence agentSessionFailureWords writes is now assembled from a
table of whole English pieces, so a reader can supply its own words for
each piece. The host still fills them in English, byte for byte as
before.

* fix(native-chat): translate the failure sentences desktop notices show

A refused start and a rejected message now carry their failure fact to
the notice instead of its English sentence, and desktop words that fact
through translate keys whose English defaults are the host's own
pieces. The host keeps writing English into rows and reasons, and a
host sentence with no fact beside it still shows as written.

* fix(native-chat): the history pane says only that Orca keeps trying

When the pane's title already says Orca couldn't open this chat's
history right now, the line under it no longer repeats that the
transcript could not be read; it says only that Orca keeps trying to
load it. The pane with no named reason keeps its two-part line.

* test(native-chat): type the failure pieces a refusal notice shares

* test(native-chat): the Chinese failure words use no Japanese-only characters

* fix(native-chat): every history pane that says it didn't load says only that Orca keeps trying

A pane whose title is a code's own words ("This chat's history couldn't
be loaded.") now gets the short retrying line too. Only the pane with no
named reason keeps the two-part line.

* test(native-chat): a provider's words with nesting and markup stay as written in a translated notice

* fix(native-chat): French and Spanish say a withdrawn message was withdrawn before the agent began working on it

* fix(native-chat): Japanese and Chinese notices run their sentences on without a space

A notice joined its sentences with a space in every language, so Japanese and
Chinese read "Claude 无法启动。 请重新发送消息。" with a stray gap after the full
stop. The failure-sentence builder now takes the joiner alongside its words, and
desktop joins in the UI language: no space in Japanese and Chinese (including a
plugin pack that declares either), one space elsewhere. The host and the phone
keep English, joined with a space as before.

* fix(native-chat): a failed /clear or /compact says why in the app's language

The line under the composer printed the host's English sentence although the
result carries the typed failure beside it. It now words that failure the way
the host does (the chat's agent and /clear for a failed /clear, nothing for
/compact), in the app's language; an older host that sends no failure keeps
its sentence.

* fix(native-chat): Spanish says a rate limit, and Korean says a withdrawn message was never processed

The Spanish retry notice said the agent hit a usage limit, a different thing
from the rate limit the English names. The Korean withdrawn-message notice
said the agent had not started, which reads as the agent not launching; it now
says the agent had not begun processing the message, as the other languages do.

* refactor(native-chat): one rule says which words already say the history didn't load

* fix(native-chat): an image size limit says its unit the way the reader's language does

* test(native-chat): the option picker's i18n stand-in knows the reader's locale, which a refusal notice now reads

* fix(native-chat): a sentence a language pack left in English keeps its space

Sentences were joined by the UI language: no space in Japanese and Chinese,
one elsewhere. A plugin pack for a Chinese or Japanese variant that predates
the failure words falls back to English for them, so a notice read
"您的訊息未傳送。Claude couldn't start.Send your message to try again."

Each gap now follows the sentence before it: none after a full-width 。!?,
one space after anything else. The built-in Japanese and Chinese catalogs end
every sentence in 。, so they read as before, and English is unchanged. Since
the rule no longer needs the language, one joiner serves every surface and
the failure-sentence builder no longer takes one alongside its words.

* fix(native-chat): a failed /clear this build only partly understands shows the host's own sentence

The line under the composer words a failed /clear or /compact from the fact
the host sends beside its sentence. The reader drops any part this build
cannot place, such as a refusal code a newer host added, and the rest of the
fact can then give different advice: "Run /clear again." where the host said
"Start a new chat to continue."

When any part the host sent did not survive the read, the line now shows the
host's sentence as written, the same as for a host that sends no fact. A fact
this build reads whole is still worded in the app's language.

* fix(native-chat): a failure fact this build reads only in part shows the host's sentence everywhere

The previous check compared only a fact's top-level parts, so a known refusal
code carrying a reason a newer host added still counted as read: the reader
dropped the reason and the notice re-worded what was left, which can advise
differently from the host ("Run /clear again." against "Start a new chat to
continue.").

One shared reader now answers whether this build read the whole fact: it reads
the fact and keeps it only when the read equals what arrived, at every depth.
Every place that chooses between wording a fact and showing the host's text
uses it: the line under the composer after /clear or /compact, a rejected
message's notice from the journal's fact, and the smaller copy a rejected
message keeps for when its journal row is not loaded. Matching a rejected
message to the start row that already says why still uses what this build can
read, since that is identity, not wording. Facts this build reads whole are
worded as before.

* fix(native-chat): a failed /compact names /compact as its next step on desktop too

The host names the command a failed start was waiting on, for /clear and /compact alike.
The line under the composer re-worded only /clear with it, so a /compact whose start failed
read "The agent couldn't restart. Send your message to try again." in the reader's
language. It now words every command the host answers with the agent and the command the
host used, so desktop English matches the host and French or Japanese keep /compact.

* fix(native-chat): a /compact whose start failed no longer says the operation was not confirmed

A /compact on a chat whose agent is not running starts it first, and that start takes a new
lease, so the chat's fence moves before the command's reply arrives. The write settles as one
for a fence this pane no longer shows, and the composer line read that as "Conversation
operation was not confirmed." Such a write now says nothing there, as every other write
already does: the chat's own start-failure row says why, and the command's message is
rejected in the journal. The sentence it printed is gone from the catalogs.

* fix(native-chat): keep the message outbox within its line limit after main's growth

* fix(native-chat): a command's own reply is kept when its start moved the fence

A /clear or /compact on a chat whose agent is at rest starts the agent first, and that start
moves the chat's fence before the command's reply arrives. Every reply from an earlier fence was
discarded, so a /clear whose new chat failed to start said nothing at all, and a /compact that
started left "/compact" in the composer.

A conversation command's reply is now kept while the pane still shows the chat it was sent for;
a closed pane or another chat still drops it, and every other write keeps the fence rule. The
line under the composer says nothing only when the failure is the chat's own start and that
start's loaded row already says why, the rule a message that start rejected already follows.

* fix(native-chat): a returned queued message shows the host's sentence for a fact read in part

The caption under a returned queued message re-worded its failure from whatever this build could
read of the fact. A newer host's fact with a refusal code or reason this build drops read as a
shorter sentence with different advice. It now re-words only a fact read whole, and otherwise
shows the host's own sentence, as every other surface that re-words a fact does.

* test(native-chat): a command reply after any fence move, for /clear and /compact

The fence-move tests now state the rule as the code has it: a conversation command's reply is kept
whenever the pane still shows its chat, whatever moved the fence. They cover a /clear that
completed, and a failed start for each command in French with the fact that command really meets
(a /clear's new chat fails to start; a /compact's chat fails to restart).

* fix(native-chat): only the reply a pane still waits on outlives a fence move

A conversation command's reply was kept across a fence move whenever the pane still showed the
same chat. A reply the pane had stopped waiting on, because it left the chat and came back or a
newer command replaced it, was applied as if it answered the current one.

The pane now remembers the one command request it waits on; only that request's reply is kept
after the fence moves, and closing the pane or showing another chat forgets it. Tests now drive the
fence move during the request itself rather than through /clear starting an agent, and cover a
/clear that stops a running agent.

* fix(native-chat): a newer-Orca history error keeps a whole retry line, and the phone shows the host's words for a fact it reads in part

When a chat's history was saved by a newer Orca, the pane's title reads "Chats were saved by a newer
Orca. Update Orca to keep using them." and the line under it said only "Orca keeps trying to load
it.", with nothing for "it" to mean (in French and Spanish the pronoun also disagreed with "chats").
That title names every chat rather than this one, so the pane keeps the full line: "The transcript
could not be read. Orca keeps trying to load it."

On the phone, a returned queued message whose failure fact this build reads only in part was
re-worded from what it could read, dropping advice the host gave; it now shows the host's own
sentence, as the desktop card already does.

The comment on the reply the pane waits on now says what forgets it: a newer command, or disabling
the pane.

* fix(native-chat): a command reply this build can't place shows the host's own words

A newer host can answer a conversation command with a command name this build doesn't know. This
build re-worded that reply from its failure fact as if it were a command it knew, naming the
command in its own words, or said nothing when a loaded start row matched the fact. It now shows
the host's sentence as written, as it already does for a fact it can read only in part.
2026-09-30 14:15:07 -07:00
Neil 85f8d6b5f5 test: retire long-tail cases whose assertion is decided by the test itself (#24132)
Resumes the backlog sweep at a chunk size that actually gets read. Six auditors, 84 files
each, and all six read their full scope case-by-case against production — the first wave
where every chunk closed with no gap. 33 case declarations removed across 22 files, 1 test
file deleted, 826 lines gone. No production code touched.

This wave exists because a conclusion of mine was wrong. I had recorded that yield collapsed
~36x and that deletion was no longer the high-value work. I was dividing cases removed by
files IN SCOPE while the fraction auditors actually READ fell from 100% to about 4%, because
I kept handing them 300-800 files. Recomputed against files read, yield has been flat at 4-7
per 100 with no downward trend. This wave came in at 8.2.

The most instructive removal looked like the most valuable test in scope.
`orchestration-worker-release-reap-fixed.func.test.ts` cites a production bug by two
identifiers, describes orphaned PTYs accumulating until `TasksMax=4096` aborts processes on
EAGAIN, and advertises itself as the functional tier wiring the real orchestration RPC
surface, the real `OrchestrationDb` and the real release modules. Deleting it leaves no
reference to that bug anywhere in `src`.

It still had to go: its fake runtime performed the fence it asserted —

    if (pty.incarnationId !== inc) { return null }
    handleTable.set('term_reminted', { ptyId, epoch: rendererGraphEpoch })

— so the case checking that a reused ptyId with a mismatched incarnation does not resolve was
checking a decision its own spy made twenty lines earlier. The real fence is owned by
`orca-runtime-terminal-handle-incarnation.test.ts:257`, and the other two cases replay
`orchestration-worker-release-incarnation-fallback.test.ts` (which uses a plain
`mockReturnValue` rather than reimplementing the remint) and `worker/worker-release.test.ts:23`.
"Integration test" and "wires real modules" describe the scaffolding, not the asserted step.

Other removals: a self-comparison disguised by an alias, where
`export const getIssueOwnerRepo = getOwnerRepo` makes a case asserting the two "agree" into
`f(x) === f(x)`; four cases whose `vi.mock` of `resolveIssueSource` made both the preference
value and the topology inert; five verdict-precedence cases owned by a verdict-agnostic block;
three call-shape probes on one-line store pass-throughs whose real contracts are driven by
behavioural neighbours; and a `export type _Ref = [...]` declaration whose own comment admits
it exists only to preserve test-only module-surface references.

Kept after checking production rather than shape. An auditor found two near-identical
ten-reconnect loops and kept both: one uses a test-local live-lease filter, the other the
shipped `sshRemotePtyLeaseAllowsReattach` predicate, and the file's own comment explains the
duality is deliberate "so the two cannot drift". Another kept a paths-alignment case that
looks like a validator tested against its own list, because adding a generated file without
registering its path does fail it — and `shellReadyWrappersExist` uses that registered list to
decide whether a partial tree needs regeneration.

Production duplication is now confirmed four times over, and it is why mirrored tests exist:
`createUpdateWorktreeLineage`/`createAssignWorktreeParent` differ by one `console.error`
string; `terminal-path-tap.ts` and `document/path-tap.ts` carry hand-maintained copies of
`matchFilePathAtColumn` under a docblock reading "keep the two in sync". In those cases both
test sides are load-bearing and the duplication belongs on a refactor list.

`mobile/tests-typecheck-baseline.txt` loses one entry. Trimming
`relay-host-signed-out-verdict.test.ts` made it typecheck clean, so the ratchet required
pruning its grandfathered entry — the file graduates from exempt to enforced. Baseline is now
124 entries, down from 125.

Verified: 690 test files / 7,560 cases pass across the touched desktop areas; the modified
mobile files pass (162 cases); `check-tests-typecheck-ratchet.mjs` OK (898 files in program,
124 grandfathered); `check-reliability-gates.mjs` 140 gates; the deleted file is absent from
the gate manifest, `cloud/package.json` and the mobile baseline; nothing under
`mobile/src/test-support/rpc-recording/` or `mobile/rpc-foundation/goldens/` touched.
2026-09-30 04:58:31 -07:00
Neil a781a602a8 test: retire duplicate cases that replay an owner across a re-export or provider shim (#24114)
Resolves 208 candidate pairs where the same case title appears verbatim in two or more
files, produced by a repo-wide scan calibrated against a known positive. 46 case
declarations removed across 32 files, 798 lines gone. No file deleted whole, no
production code touched.

The headline result is the measurement, not the deletions: across the three buckets that
reported in detail, the signal ran roughly 86% false-positive (3/42, 9/42, and the rest).
It has good recall and poor precision, and it reorders a reading queue rather than
replacing one. Calibrating a detector against a known positive proves recall, not
precision.

What the deletions were:

- Duplicate invocation through a re-export shim. `native-chat-tool-summary.ts` is a
  ten-line `export {...} from '../../../../shared/native-chat-tool-summary'`, and
  `agent-status.ts:161` is `export { isExplicitAgentStatusFresh } from
  './pane-agent-evidence'`. Cases on the shim side were byte-equivalent to the owner's
  with no rendering or transport hop.
- Provider-local replays of a shared helper: three `repository-ref` providers that are
  each `createRemoteRefProbeCache(parseXRef)` and contribute nothing to transient
  handling; two `local-pty` and `daemon/session` tables replaying
  `shell-startup-output-scanner`, whose owner additionally checks every split point.
- A reader-side replay of store policy. `runtime-worktree-agent-rows-structured.test.ts`
  asserted an attention-to-blocked mapping; the reader contains zero `attention` or
  `blocked` tokens and copies `state` through. The mapping lives in
  `structuredAgentSessionAgentStatus`. Consistent with
  `docs/reference/agent-status-store.md`: readers keep only presentation policy.
- Constructor-only subclass duplication: the shared capability-cache case is covered by
  `codex-app-server-capability-cache.test.ts`, whose ten cases include the identical
  title plus all four risks `docs/reference/git-compatibility.md` names — first fallback,
  later cached call, concurrent probes, per-host isolation.
- A private predicate duplicated at a real boundary, varying only a path passed straight
  into the shared predicate.

Why most pairs were KEPT, because the false positives are principled rather than noise:

- Two independent execution hosts. `src/relay/git-handler-*` and `src/main/git/*` are
  separate Git implementations that cannot import each other and hold separate capability
  caches, exactly as the compatibility doc requires; the repo already ships
  `status-branch-line-total-relay-parity.test.ts` to pin the duality deliberately. Neither
  side's argv, timeout or cache regression is visible to the other.
- Deliberately duplicated production siblings: Codex vs Claude (different account fields,
  different CLIs, different wire protocols), gitea vs bitbucket (`/pulls/42` vs
  `/pullrequests/42`), gl-utils vs gh-utils (separate in-flight maps). Same contract
  shape, different implementations — an identical title is the correct naming.
- Shared-predicate consumers: one side tests the predicate, the other tests a caller's
  wiring to it. A caller that forgot to call the predicate passes the shared test.

In a codebase with intentional provider and host symmetry, identical test titles are
expected, and the signal cannot distinguish "copied" from "parallel by design" because
both produce the same prose. Only reading both bodies separates them.

Verified: 6,968 desktop test files pass; the three modified mobile files pass (39 cases);
`check-reliability-gates.mjs` 140 gates; nothing under
`mobile/src/test-support/rpc-recording/` or `mobile/rpc-foundation/goldens/` touched.

62 local failures across 12 files were each accounted for and none is caused by this
change: `browser-manager-tab-identity`, `browser-manager-viewport-ownership`,
`session-scanner-codex-workers` and `managed-hook-script-refresh` all fail identically in
a pristine `origin/main` worktree; five `mobile-web-app-*-render` tests need Playwright
browsers this machine lacks; `structured-agent-session-restart-ownership` and
`ssh-remote-commands` pass in isolation and fail only under concurrent load.
2026-09-30 03:30:23 -07:00
Neil cef66fbab8 test: retire long-tail cases whose input cannot reach the behavior they name (#24101)
Sweeps the triage-only backlog: 2,269 files that earlier waves saw and skipped for
size, reconstructed from the unread lists five waves of auditors disclosed. 35 case
declarations removed across 15 files, 2 test files deleted, 487 lines gone. No
production file touched.

These are large integration suites, so the junk here is individual cases buried among
real coverage rather than whole bad files. The dominant defect was again a case whose
input cannot reach the behavior its title names:

- `resume-sleeping-agent-session-remote-compat.test.ts` (deleted) — two cases titled
  for "transport-level host authority on a capable host" and "host authority is not
  known". `resume-sleeping-agent-session.ts` has no host-authority or capability
  concept at all, and its only read of `origin` is
  `if (!record.origin && record.state === 'done')`, unreachable for both rows. Both
  executed one identical path. The surviving contract is owned by
  `resume-sleeping-agent-session-execution-host-scope.test.ts`, which drives a real
  host catalog.
- `project-group-header-drag.test.ts` (deleted) — four cases setting
  `data-project-group-header-id`, which the predicate never reads. Its subject,
  `isProjectGroupHeaderActionTarget`, is byte-identical to `isRepoHeaderActionTarget`
  apart from the function name and imports the same `REPO_HEADER_ACTION_SELECTOR`, so
  all four cases were a strict subset of `project-header-drag.test.ts` using identical
  `data-repo-header-*` fixtures.
- `remote-worktree-history-cleanup.test.ts` — "repeats idempotent cleanup through the
  PTY owner" against a six-line best-effort forward with zero dedupe state. The case
  called it twice and asserted the mock recorded two calls, which is arithmetic over
  the test's own loop; nothing about idempotence was established.

Also removed:

- Runtime assertions of type-level facts, where production already makes the check at
  a stronger boundary: `const adapterSatisfiesPort: AdapterIsPort = true` followed by
  `expect(...).toBe(true)` — unconditionally true, while
  `createExpoGenerationFileSystem(): GenerationFileSystem` is explicitly annotated and
  passed into `createGenerationStore` at a typed call site. And a case named "does not
  typecheck" whose runtime assertion is a length check on its own literal, declaring
  its own local annotation so it could never notice the production annotation
  weakening.
- Private predicate tests duplicated at a real boundary: four `repo-slug-cache` cases
  delivered by `repo-slug-index.test.ts`, which drives the same resolution through the
  hook, the real store and the preload bridge, while the cache-level versions hand-seed
  the internal map and break on a cache-key format change.
- Duplicate invocations owned at the shared boundary, including commit and push
  recovery cases owned by `src/shared/source-control-recovery-agent-command.test.ts`.

Kept deliberately, verified rather than assumed: the production duplication behind the
deleted drag test was left alone, because `REPO_HEADER_ACTION_SELECTOR` ends in generic
`button, a, input, textarea, select`, so genuine action targets inside a group header
still match — it is an unspecialised copy-paste, not a live bug, and collapsing two
functions is a refactor. Reported instead.

Auditors' probes produced 20, 11 and 13 candidate hits for the signature-versus-title
shape across their chunks; every one was inspected and every one was genuine coverage.
No deletion in this wave rests on a probe alone.

Coverage is partial and stated as such: of 2,269 files, roughly 100 were read
case-by-case and the remainder reviewed at title-plus-import level. Each auditor listed
its own unread set. The largest remaining surfaces are `src/main/agent-hooks` (95),
`src/main/claude` (100), `src/renderer/src/lib/pane-manager` (62) and the 20 largest
sidebar suites.

Verified: 2,583 desktop test files / 25,636 cases pass, plus one pre-existing
`it.fails` marker; the two modified mobile files pass (57 cases);
`check-reliability-gates.mjs` 140 gates; `check:code-quality:changed` 0 new findings.
Both deleted files confirmed absent from the gate manifest, `cloud/package.json` and
`mobile/tests-typecheck-baseline.txt`.
2026-09-30 02:45:32 -07:00
Neil b99462ac1c test: retire mobile, cloud, config and e2e cases their input cannot reach (#24077)
Completes the first pass over every test area in the repository. Sweep over
`mobile/src`, `config/scripts`, `cloud/`, and `tests/` (1,494 files in scope, with
the 24 files under `mobile/src/test-support/rpc-recording/` deliberately excluded).
31 case declarations removed across 17 files, 2 test files deleted, 356 lines gone.

What went, by pattern:

- Cross-boundary replays of a shared helper. A whole mobile file re-ran
  `extractPendingAsk`/`parseAskFromStatus`/`formatAskAnswer`, all owned by
  `src/shared/native-chat-ask.test.ts`, `native-chat-ask-fifo.test.ts` and the
  renderer's interactive-prompt suite — one case title was verbatim identical to the
  owner's, and the owners' inputs are supersets. The mobile file imported the shared
  module directly and exercised no mobile transport, lifecycle or rendering.
- A case whose input cannot reach the behavior its title names: "arms it on Android
  while the drawer is open", where `use-back-claim.ts` has zero
  Platform/OS references, so flipping the mocked OS changes only shadow styles.
- Identity copiers, including one asserting `prSidebarRenderBranch(state) ===
  state.kind` against a production body that is `return state.kind`. The function
  stays; it has three live callers.
- A test of the runtime rather than the product: a case asserting Node's own
  `EventEmitter` crash contract on a bare emitter, with zero production code in the
  path. The guard it documents is exercised behaviourally by the case after it.
- Duplicate invocations, one of them provable rather than eyeballed: with
  `MODULE_SCOPE_ENV_WRITER_PIN = 0`, `files.size <= 0` is strictly implied by the
  sibling's `expect(offenders).toEqual([])`, since a non-empty `offenders` forces
  `files.size >= 1`. The pin's own doc says it may only ever be decreased from 0, so
  it could never become a meaningful bound either. Its policy guidance survives as a
  comment; the file's real ratchet and its regex self-test both stay.
- Expected values produced by the test's own arithmetic, and a p95 case strictly
  implied by a sibling that already pins exact p95 and exact max over a wider range.

One production line goes: the `export` keyword on `assignmentCleanupSteps` in
`cloud/apps/relay/src/assignment-cleanup-steps.ts`. The function itself stays and is
still called internally; only the test-only export was orphaned.

Kept deliberately: everything a gate cites, checked by case title and not only by
file path; a gate-cited case that does not deliver its claim (reported instead — see
below); a cross-version wire cell whose ledger is never invoked, left under the
raised bar for wire coverage; and every limit, bound, quota and provenance guard.

Nothing under `mobile/src/test-support/rpc-recording/` or
`mobile/rpc-foundation/goldens/` was touched — those bytes feed a `recorderSha256`
digest pinning 398 golden recordings.

Verified: `mobile` vitest over the modified mobile files (8 files, 50 cases);
`mobile/scripts/check-tests-typecheck-ratchet.mjs` OK (898 files in program, 125
grandfathered, none @ts-nocheck); relay suite 799 passed; `check-reliability-gates.mjs`
140 gates; both deleted files confirmed absent from the gate manifest,
`cloud/package.json` and `mobile/tests-typecheck-baseline.txt`.

Seven local failures were investigated and none is caused by this change: five
`mobile-web-app-*-render` tests drive `playwright-core` chromium/webkit and need
browsers this machine lacks, `release-checkout.unit.test.ts` needs cross-version git
refs, and `e2e-worker-env-isolation.unit.test.ts` fails identically with its HEAD
content restored — it recurses `tests/e2e` with symlink-following `statSync` and no
depth guard.
2026-09-30 02:05:30 -07:00
Brennan Benson 75719bd38b fix(terminal): keep the mouse report format in desktop pane snapshots so phone swipes never type escape text (#23943)
* fix(mobile): preserve host mouse modes for terminal scrolling

* revert(terminal): drop the out-of-band mouse-modes channel

The earlier commit sent the host's mouse tracking and encoding beside
each snapshot and stream frame, and made the phone trust that over what
its replayed bytes say. Where the snapshot text already carries the
encoding this is redundant, and where the text is wrong (a desktop pane
snapshot never writes ?1006h) the host's own mirror is seeded from that
same text, so the side channel asserts the wrong encoding too.

Remove the shared type, the host publication, the wire fields and the
phone's host-modes authority. The phone keeps its guard: while mouse
tracking is on but no replayed byte proved the encoding, a wheel scrolls
locally instead of sending a guessed legacy report.

* fix(terminal): carry the mouse encoding in desktop pane snapshots

Swiping a phone terminal running Codex typed `[M`-style mouse bytes into
Codex's prompt instead of scrolling (#23818). Codex turns on mouse
tracking with the SGR encoding (?1006h). A desktop pane snapshot is made
by xterm's serialize addon, which writes the tracking modes but never the
encoding, so the phone replayed "tracking on, legacy encoding" and sent
legacy `ESC[M` reports that Codex does not parse. The host's own terminal
model is seeded from the same snapshot, so it lost the encoding too.

The pane now mirrors its xterm's mouse encoding from the parser (1006,
1016 and a full reset, with xterm's own set/reset rules) and its snapshot
ends with the matching DECSET, after the alternate-screen switch that
readers keep. Programs on the default encoding get nothing appended.

The phone keeps a guard for hosts without this fix: while a wheel-
reporting tracking mode is on but no replayed byte proved the encoding,
a wheel or swipe scrolls locally and a tap or drag sends no mouse report
(the tap still focuses the keyboard). Click-only tracking (?9h) keeps
its arrow-key scrolling on the alternate screen, and an explicit ?1006l
still sends legacy reports.

* fix(terminal): state the default mouse encoding so legacy mouse apps keep phone input

Snapshots now say ?1006l while a program tracks the mouse with the default
encoding (daemon rehydrate and desktop pane serializer), and the phone
treats a tracking enable seen in live output as proof of its encoding.
Legacy-encoding programs such as vim with mouse=a keep phone scrolling,
taps and drags; replay-only tracking with no stated encoding still sends
nothing. An unproven drag falls back to local selection, and an empty pane
snapshot stays empty.

* chore(terminal): type the mouse-encoding tracker inputs for the low-evidence audit
2026-09-30 01:05:26 -07:00
Brennan BensonandHarshul Rathod 707d3dc96b fix(chat): decode Claude pastes and report terminal delivery uncertainty (#23788)
* fix(chat): decode Claude pastes and track terminal delivery uncertainty

Keep queued prompts pending while the existing agent status reports work, and check fresh history after a later idle fact. Preserve draft text and distinguish write rejection from unconfirmed delivery.

Co-authored-by: Harshul Rathod <harshulrathod1640@gmail.com>

* fix(native-chat): break the observed-send import cycle and keep renderer tests out of main

The observed-send path imported the clear helpers from native-chat-runtime-send,
which imports it back. Move the input-clear layer into its own module both use.
The Claude paste decoder test imported the renderer pending module from
src/main, which the node typecheck project cannot see; the echo-retirement
assertions now live in a renderer test.

* fix(native-chat): still submit a Claude chat send whose write acknowledgment was lost

A remote write whose acknowledgment is lost (timeout, dropped link) is not a
refusal, but the observed path stopped there and never sent Enter, leaving a
body that did land sitting unsubmitted in Claude's input line until the next
send's clear wiped it. Continue to the next write without re-sending the bytes,
as the unobserved path always did.

* perf(native-chat): keep terminal Chat pending delivery from re-rendering every row

The delivery notices were merged into a new Map on every render, which
invalidated the transcript row context and re-rendered every memoized row on
each stream update. The pending hook also wrote a fresh array on every status
ping and prune pass even when nothing changed, and the phone mapped its pending
list on every render, rebuilding the chat list data. Memoize the merged
notices, skip no-op pending writes, and memoize the phone's rendered pending
list. The phone also skips a transcript read when no send is due.

* fix(native-chat): never flag a queued Claude send, and flag one an idle Claude never starts

Two gaps in when terminal Chat calls a Claude send "Delivery unconfirmed":

A prompt sent while Claude is mid-turn is queued, and Claude folds it into the
running turn as a queued-command record. The transcript reader drops those
records, so once the turn ended the prompt Claude did run read as unconfirmed,
inviting a duplicate resend. A send made while the agent is busy is now never
checked; it keeps the pending behaviour it had before.

A prompt sent to an idle Claude that never starts a turn (Claude exited to the
shell, or the paste went nowhere) left the status at the same idle fact
forever, so the check never ran and the bubble stayed pending. An idle agent
starts a turn on a delivered prompt at once, so a send whose idle status is
unchanged after the existing 20 s bound is now checked against a fresh
transcript read.

* fix(native-chat): add the delivery notice strings to the English catalog

The Dismiss action's translate key was missing from en.json, which fails the
localization catalog and extraction gates. The desktop "Message not sent" and
"Delivery unconfirmed" notices were hard-coded English; route them through
translate with the same wording.

* fix(mobile): sync the held-send refs after commit instead of during render

Moving the acknowledgment-loss hold into its own hook made its render-time ref
writes new lines, which the React Doctor changed-lines gate blocks. Held sends
report after commit, so syncing those refs in a layout effect keeps them
current where they are read.

* fix(native-chat): report only definite terminal Chat send outcomes

The delivery rule inferred "Delivery unconfirmed" from "the turn ended and the
transcript has no matching row". Claude records a prompt sent mid-turn only as a
queued-command attachment, which the transcript reader drops, so that rule
flagged prompts Claude had answered. It also never fired for an idle Claude that
lost the write, because no newer turn arrives.

Keep only facts the transport reports:
- a refused write reads "Message not sent", keeps its text, and can be dismissed;
- a lost write acknowledgment holds the echo for 20 s, the phone's existing
  rule, then reads "Delivery unconfirmed" unless its row has landed.
An ordinary send, including one Claude queues mid-turn, stays pending as before.

Remove the agent-status subscription, the status-epoch origin, the fresh
500-row transcript read, the confirmed state and the no-status clock. The phone
already implements this rule, so its changes revert to main; only a test for
old-host paste envelopes remains.

* fix(i18n): translate the terminal Chat delivery notices

Add the Dismiss, "Message not sent" and "Delivery unconfirmed" strings to the
es, fr, ja, ko and zh catalogs, reusing each catalog's existing Dismiss wording.

* fix(native-chat): let a resend replace its failed terminal Chat echo

A "Message not sent" or "Delivery unconfirmed" echo kept its transcript
occurrence, so resending the same text numbered the resend as the second
copy: the one landed row retired the failed echo and pinned the resend
below the reply forever. Appending a send now drops a failed echo with the
same content first.

* fix(native-chat): unwrap a Claude paste that quotes pasted_content tags

The envelope parser refused any body containing a pasted_content tag, so a
pasted prompt that itself quotes one (a transcript excerpt, or code that
handles these tags) kept its wrapper and its echo stayed pinned below the
reply. Claude's per-paste id exists to disambiguate exactly that; only a
same-id tag inside the body is now ambiguous. Wrappers without an id keep
the strict rule.

* test(native-chat): pin which terminal Chat sends observe write outcomes

Only a Claude chat send (text or images) reports a refused or unacknowledged
write to its pending echo; other agents and slash commands keep the
unobserved write path exactly as before.

* fix(native-chat): keep failed terminal Chat sends through Stop

Stop cleared every optimistic echo, including a "Message not sent" or
"Delivery unconfirmed" bubble whose send had already settled. Stop cannot
affect that send, and the bubble is the only place its text stays copyable,
so it now survives until the user dismisses or resends it. Also moves the
observed-send import below the file header comment.

---------

Co-authored-by: Harshul Rathod <harshulrathod1640@gmail.com>
2026-09-30 01:04:21 -07:00
Brennan Benson ca7c14db08 fix(mobile): start + menu, quick command and diff-note agents through agent.launch (#22954)
* fix(mobile): start + menu, quick command and diff-note agents through agent.launch

The session screen's + menu, agent quick commands and diff notes' New agent
session now ask the host to start the agent with agent.launchReplay, so the
host picks chat or terminal from the desktop's default and delivers any
prompt. Hosts without the launch capabilities keep today's paths.

The phone's pending tab choice is one value (a tab, a terminal by handle, or a
launched surface) instead of two refs, and a launched chat is found by its
session id in the next snapshot rather than a predicted tab id. A launched
surface waits a bounded number of snapshots for its tab.

* test(mobile): add the + menu and diff-note launch scenarios to the recording corpus

* test(mobile): repin bridged-parity tallies for the four launch goldens; drop test casts

The corpus grows from 790 to 794 goldens; all four new ones replay identically.

* fix(mobile): show a refused agent launch as a toast beside open tabs

The inline create error renders only in an empty session, so a host refusal
(for example a disabled agent) from the + menu in a session with tabs showed
nothing. Always toast the failure: the caller's own copy when it gave one,
otherwise the host's reason.

* test(mobile): type the launch reply helper with the shared launch outcome types

* fix(mobile): record a launched agent's tab as this device's pick on the host

A launch carries no navigation, so the phone selected the new tab only
locally while the host kept this device on the tab it had before. Leaving
the session and coming back, or a reconnect that reset the screen, reopened
that old tab. The "+" terminal path this replaced asked the host to select
the tab for the caller.

When a launched surface's tab lands in a snapshot, activate it for the
caller exactly as a tap does. The resolver now names the landed tab in
place of the unused `missed` flag. Route parity re-pinned for the new
activation body, identity payload and strings.

* fix(mobile): land on a launched agent's tab without a 500 ms wait or a blank pane

The host publishes a launched tab before it replies, so the tab list the
phone already holds usually has it by the time the reply arrives. The
launch paths still waited for a refetch 500 ms later, leaving the phone on
the old tab for that long after every launch. Read the tab list at once.

On hosts without agent.launch, the chat path also unsubscribed the open
terminal and cleared its handle before the chat's tab landed, while the old
terminal tab stayed selected: a blank pane until the next tab list. Leave the
open tab live until the chat lands, as the launch path does; applying that
tab list tears the old terminal down.

Route parity re-pinned for the two bodies.

* fix(mobile): keep a tab the user picked while a prompted launch was still replying

A quick command or review-notes launch now waits for the host to deliver the prompt, which can take up to a minute. The launched tab shows up in the tab row well before that, so a user who tapped another tab meanwhile was pulled back onto the launched one when the reply arrived, and that pick was recorded on the host.

The launch now remembers which tab the phone was on when it started and only takes focus if the phone is still there when the reply lands. Any move made in between, by a tap or by the computer navigating this phone, wins. Session route parity re-pinned for the handleCreateTerminal body only.

* fix(mobile): name a launched agent's tab before asking, and land on it when it is listed

A "+" menu, quick-command or review-notes launch now reserves its tab before it asks the host: a fresh pane key (tab and leaf UUIDs) and, for an agent the host may start as a chat, a session id. Both are minted once per launch and sent unchanged on every replay, since the host's replay fingerprint covers them.

The phone arms its pending selection with that reservation before sending, so it lands on the terminal (matched by pane halves) or chat (matched by session id) as soon as the tab is listed. For an agent whose prompt is pasted after start, that is long before the reply, which waits for delivery. Landing also frees the "+" lock; the lock holds the create's id, so an older launch's reply cannot free a newer one's. The reply now only adds its own handle or session id (an older host ignores the reservation), starts the fallback countdown, and reports prompt delivery.

A tab the user picks mid-launch replaces the pending selection, so the launch-start tab check is gone. A reservation the host refuses as already taken reads "Couldn't start the agent. Try again." on the first send, and as unconfirmed after a replay. The mobile UUID fallback now yields a v4 UUID, because a pane key's leaf must be one. The host launch path moved to new-tab-agent-host-launch.ts; session route parity re-pinned for that move and the landing's lock release.

* test(mobile): expect the launch reservation in the four launch scenarios

The four launch scenarios now expect the pane key and session id the phone sends (the scripted ids come first, so the operation id moves from ...001 to ...004).

* fix(mobile): don't say an agent may not have started while the user is looking at it

When a launch's reply was lost after its tab had already landed, the phone said "Couldn't confirm the agent started", although the listed tab proves it did. Now a listed tab narrows the doubt to the prompt or notes ("The agent started, but couldn't confirm the notes were sent."), the notes stay unsent, and a bare launch says nothing. Only the nested-function parity pin moves, for handleCreateTerminal passing the tab list to the launch.

* test(mobile): read the launch's sent reservation through the host's params schema

The anti-slop audit rejects Reflect.get; parsing with AgentLaunchReplay also
asserts the host accepts the params the phone sent.

* test(mobile): check the launch reservation against the host without importing its schema

Mobile code may import the params contract only as types. The phone's tests
now read the sent reservation by narrowing, a chat reservation is checked
through the real host dispatcher, and the older-host drop is pinned host-side.

* fix(mobile): don't send the same review notes to a second new agent

The "+" lock is now freed when the launched tab lands, but review notes are
only cleared when the launch's reply confirms delivery, which for a prompted
launch can take up to a minute. In that window "Send review notes to AI" still
offered the same notes, and choosing a new agent session started a second
agent with them.

The notes a new agent session is being started with are now held from the tap
until that launch settles: the Send button no longer counts them, the sheet no
longer offers them, and a stale tap on the old sheet starts nothing. Notes the
host did not deliver become sendable again once the reply arrives.

* test(mobile): record the + menu and diff-note launch goldens

4 added (+ as a terminal, + as a chat, notes delivered, notes not delivered). 10 existing create-terminal goldens move only because the recorded state now shows one pending selection instead of two refs; their requests are unchanged.
2026-09-30 00:42:52 -07:00
Brennan Benson 21124db4d5 refactor(native-chat): a subagent's rows live in its own section, not in the parent's conversation (#23752)
* refactor(native-chat): a subagent's rows live with that subagent, not in the conversation

A subagent's rows were drawn in its parent's conversation, each captioned with
the subagent's name. They now belong to the subagent: the transcript projection
keeps the session's own rows as the conversation and each subagent's rows apart,
keyed by the agent id its roster entry already carries, folded on their own.

Desktop: a subagent's rows open in a section under the roster row that names it,
from that agent's roster entry, and are windowed like any other rows. A subagent
no loaded roster names opens where its first row happened, inside the section of
the agent that spawned it or in the conversation. Its edits still count in the
turn they were made, and revealing one opens the sections around it.

Mobile shows the conversation, with each spawn's roster line. Worker reads and
structured terminal reads serve the worker's own rows.

Removes what the move makes redundant: the per-row caption and its copy, the
producer check in the tool fold and the turn answer, the per-agent frontier
interleaved in the conversation, worker-text subagent tags, and the agent id on
worker-read messages.

* refactor(native-chat): a diff target names the sections its row sits in

Revealing a subagent's edit opens the sections around it from the target the
rollup already holds, instead of looking the row up at click time. The section
head keeps to the agent's name and dot; its state in words stays on the roster
entry. The worker page test stubs the host through its module rather than a cast.

* fix(native-chat): a working subagent's section is open; a worker page windows its own rows

A subagent's section is open while its agent works and closes once it settles,
the way the turn's own live run does; a section the reader opened or closed by
hand keeps that choice. A subagent another subagent spawned opens inside that
one's section, so a working grandchild shows inside its working parent. Openness
is derived from the roster's state and the reader's choices; nothing stores an
automatic open.

A worker page is now the newest page of the worker's own rows. The host windows
the read over them before the limit, so a subagent's burst can no longer crowd
the worker's rows off the page, and "older" still means older worker rows. The
scope is an in-process argument of the host's history read; no wire request
carries it.

* fix(native-chat): a subagent section head names the turn it sits in, for the outline rail

* fix(native-chat): a subagent section's rows sit in the turn the section is shown in, for the outline rail

A background subagent's rows written during a later turn carried that later
turn onto their slots, so scrolling through its section lit the later turn's
rail tick and then snapped back. The rollup still counts each edit in the turn
it was made; only the slot, which the rail reads, takes the shown turn.

* fix(mobile): Load earlier reads past pages that hold only a subagent's rows

Mobile draws only the session's own rows, so an older page made entirely of a
subagent's rows landed as nothing: the reader tapped Load earlier, saw the
spinner, and got the same transcript back. One load now reads on (up to 8 pages)
until a page holds a row of the session's own, then applies the pages in order.

* test(mobile): stub the RPC client the way the other structured-session hook tests do

* perf(native-chat): order subagent rows for the changed-files rollup once per change to them

The rollup flattened and re-sorted every subagent row on each update, including
every token the parent streamed. The ordering now keys on the projection's
subagent rows, which keep their identity while only the conversation changes.

* refactor(native-chat): order subagent rows in the sections hook, keeping the list under its line limit

* fix(mobile): a transcript whose newest page is only a subagent's rows reads back on its own

Opened while a subagent is busy, the newest page can hold nothing but that
subagent's rows. Mobile draws none of them, so the reader saw an empty chat with
a Load earlier button, and an empty list cannot be scrolled to page. The hook now
reads back once from each such head, and the read runs on to the session's own rows.

* fix(native-chat): count the live window in the session's own rows, so a subagent's burst keeps its roster

The live window kept the newest 1,024 rows of every agent. A subagent writing
more than that trimmed its own spawn's roster row and the prompt, and its
section fell back to a closed, unnamed header. The window now keeps the newest
1,024 of the session's own rows and everything after, with an 8,192-row cap on
every agent's rows as the memory backstop. A transcript with no subagent rows
trims exactly as before.

* fix(agent-session): window history pages by the session's own rows, with a subagent's rows riding along

A history page held the newest 200 rows of every agent, so a subagent's burst
could fill a page on its own: the phone opened on an empty chat and "Load
earlier" landed nothing. A page now starts at the oldest of the newest `limit`
rows of the session's own and serves every row from there, so the subagent's
rows come with the conversation they happened in. The page stays contiguous,
the cursor still names its first row, and the byte bound still applies. A
transcript with no subagent rows gets the same pages as before.

Clients already take a page larger than its limit: both reducers raise their
retained window to the page's size. The mobile read-on and read-back stay for
older hosts.

* test(agent-session): a page reaches back to the start rather than leaving a subagent-only page

* fix(native-chat): an own-row trim takes a trimmed roster's subagent rows with it

The live window trimmed to just after the own row it dropped, so a subagent
whose roster row went kept its rows at the top as an unnamed section until
the parent wrote again. Trim to the oldest own row kept instead; it still
fires only once an own row passes the limit, so a paged-in run of subagent
rows at the head stays until then. With no subagent rows nothing changes.

* perf(native-chat): cap the live window at 4,096 rows, bounding each delta's re-derivation

Every live batch re-derives the transcript over every retained row. On the
largest real window (7,374 rows) that cost 7-8 ms a delta on desktop against
0.6 ms at the old 1,024-row window, and held about 26 MB of row content.
4,096 halves both. The most rows any local journal puts between a roster and
its subagent's last row, with the parent inside its own-row limit, is 3,005,
so no observed subagent loses its roster to the lower cap.

* fix(native-chat): a subagent section opens only while its roster is the running scope's live frontier

A section used to open whenever its roster said the subagent was working, anywhere
in the transcript and whether or not the session was running, so a background
subagent's section stayed open and grew mid-transcript while the parent moved on.

It now opens by default only while the session runs and the roster row naming the
subagent is the newest thing the parent produced, user rows aside. Newer parent
output closes it even while the subagent still works; the roster row keeps
showing that live state. A subagent still working is a running scope of its own
for the sections it spawned; a settled one closes its scope. Derived every
render, no latch; the reader's own open or close still wins.

* fix(native-chat): name a subagent's section from a client roster the window never trims

A section took its name and state from a roster row in the loaded window. Once a
burst trimmed that row, or the row sat on an older page, the section fell back to
an unnamed, closed "Subagent" header.

The shared reducer now keeps a roster keyed by agent id, folded from every roster
row and revision the client receives: pages, older pages and live batches,
including revisions of roster rows outside the window, which live batches already
carry. The first roster naming an agent wins and its revisions update it; a
removed roster row drops its entries; it is rebuilt on every page that replaces
the window and bounded to 512 agents. Sections take their name, state and
live-frontier place from it; placement stays under the loaded roster row, else
at the section's first loaded row. Only a subagent no roster ever named stays
unnamed.

* feat(agent-session): a history page names the subagents whose roster row is older than it

A page is a contiguous run of the journal whose older-page cursor is its first
item, so it cannot pull an older roster row in without skipping the rows between.
When a page held a subagent's rows but not the roster row naming it (about 11% of
the moments a reader could open a session on local journals), that subagent drew
as an unnamed "Subagent" header.

History and hydration pages now carry an optional `subagentRoster`: the first
roster entry naming each subagent whose rows are on the page and whose roster row
is not, with the row's id, sequence and revision; bounded to 64 entries and
16 KB. Items and cursor are unchanged. The client seeds its roster from it.

Rule 1 in docs/reference/remote-wire-compatibility.md: an optional field on an
existing frame, no capability gate. An older client ignores it (the released
reducer reads a page with it exactly as one without); against an older host the
field is absent and the section falls back to an unnamed header.

* Revert "fix(native-chat): an own-row trim takes a trimmed roster's subagent rows with it"

This reverts commit 22078656b2.

Its only purpose was to stop a subagent whose roster row an own-row trim had
dropped from showing at the head of the window as an unnamed section. The client
roster now names that section whatever the window holds, so the cut is back at
just after the own row the limit passes. The retention test that pinned the
unnamed-section case now asserts the section at the head keeps its name.

* chore(native-chat): state the retention limits' own reasons, now that no name depends on the window

Own-row retention keeps the conversation a reader sees from being crowded out by
rows drawn as a one-row section on desktop and not at all on mobile; the
every-agent cap bounds memory and each live delta's re-derivation. Neither is
about keeping a roster row loaded any more.

* fix(native-chat): hold the roster fold's draft map where type narrowing can see closure writes

* fix(native-chat): a roster row's newer revision replaces it in the client roster too

A revision that stops naming an agent (the host drops an entry it learns is not a
subagent, or re-keys a provisional one) left the client roster holding the old
entry, often still "working", with nothing to re-derive it. The section then read
as working forever and could auto-open, while a fresh read of the same journal
left it unnamed. The fold now drops an entry when a newer revision of the row that
named it no longer does, before any roster takes it over.

* fix(native-chat): a parent's spawn and wait calls keep the subagent they name open

A subagent section auto-opened only while its roster row was the running session's
newest row, so any later row closed it: a Codex wait on the agent, or the parent's
text before its next spawn call. Now a row that is part of delegating to a subagent
keeps that subagent open:

- a Codex collab call (spawn, wait, resume, message, close) opens each agent its
  receiver thread ids name; one naming none is ordinary output;
- a Claude spawn call names no agent, so it counts toward the roster announcing it;
- a roster row at the frontier opens its most recently added agent, not all of them.

A roster or call naming only agents one subagent spawned is that subagent's output,
so a grandchild's roster, which the host journals as the session's row, no longer
closes the spawner's section.

* fix(native-chat): a parent's call right after the roster closes its subagent's section

A parent's tool calls after a roster row fold into the tool run drawn above
the roster, so the roster stayed the newest drawn row and its section stayed
open while the parent was already reading or running commands. The fold now
records the newest journal position among the rows it merged, and the live
frontier orders rows by that newest part. The layout is unchanged. A spawn
call folded there still counts as part of the roster announcing it.

* fix(native-chat): a Codex call naming several subagents delegates to the first

A Codex collab call that names several agents opened every one of their
sections. It now counts as delegating to the first agent it names, so one
section opens, the same as a call naming one agent.

* fix(native-chat): closing a roster's list closes the sections under it

Collapsing a roster row's list of subagents left their open sections drawn,
so the section's own head became the only way to close them. And the list's
open state lived in the row, so a row the window unmounted came back
collapsed.

The transcript now holds each roster list's open state beside the section
choices. A closed list hides every section it anchors; each section keeps
its own open or closed choice for when the list reopens. With no choice from
the reader, a list is open while a section under it is open. Closing a
section from its entry keeps the list open, and revealing a subagent's edit
opens the list it sits under.

* perf(native-chat): a reveal finds the roster lists it opens with one set lookup per entry

* fix(native-chat): a subagent's roster entry heads its own rows

An open section drew the agent's name twice: its entry in the roster's list,
then a separate section head above its rows. The entry is now the head. The
roster row draws its entries through the first open one, that agent's rows
follow, then the entries after it, each run in its own windowed slot. A
section no loaded roster row holds (an older page, a grandchild, an unnamed
agent) keeps its own head.

A roster list is open while the live frontier or a reader's choice is on
one of its agents, unless the reader closed the list, so closing an agent
from its entry no longer needs to pin the list open.

The section emitter moves to its own module, and the trailing-run
predicates it shares with the slot builder to theirs, to keep the slot
builder under its line limit.

* fix(native-chat): the entries after an open subagent's rows set in its roster's type

The roster row's list inherits the system row's small muted type; the entries that
follow an open section sit outside that row, so they now carry the same type.

* docs(native-chat): a current host can also serve a page of only a subagent's rows

A page is bounded by bytes after it is windowed by the session's own rows, so a
burst that fills the bound yields a page, or an opening page, with none of the
session's own rows. Mobile's read-on and read-back therefore serve current hosts
too, not only older ones; the comments said otherwise. The retention comment
still described a closed section as a row of its own; it now sits behind its
roster entry.

* fix(native-chat): a section's prose keeps its copy/timestamp controls inside the section

An assistant row's hover controls (copy, scroll-to-top, timestamp) hang 20px
below the row into the gap before the next one (`-mb-5`). Inside a subagent's
section that put them below the section's left border, and on the section's
last row they touched the parent's next row with no gap.

Inside a section the controls now stay in flow, so the border covers them and
the next row sits the normal gap below. The row-height estimate reserves the
same 20px for a section's prose so windowing does not jump on measure.

* test(agent-session): state each appended row's turn scope, as the journal now requires

* refactor(native-chat): the client's journal retention policy lives in its own module
2026-09-29 23:47:02 -07:00
Brennan Benson 75040eba5a test: open, seed and read the agent-session record store through one test harness (#23986)
* test: open, seed and read the agent-session record store through one harness

Tests that open the durable agent-session record store, seed it, or read
back what it persisted now go through agent-session-record-store-test-harness.ts
instead of calling AgentSessionRecordStore.open or touching agent-sessions.json
themselves. A later change that moves the store into the chat database then
changes the harness instead of every test. No production code changes.

Tests whose subject is the JSON file itself (its .bak recovery, salvage,
schema versions, permissions, and what older builds read back) keep reading
and writing the file directly; the storage move rewrites or deletes them.

* test: address the record-store harness by the host's state directory

The harness took the store's own folder, so each caller picked one
(join(root, 'store'), or 'agent-sessions' where a test read the store the
runtime owns). A later change that moves the store into the state
directory's journal database could not tell those apart, and would have
had to edit every caller again.

Every harness function now takes the state directory, the one the test's
journal database and recovery capsule already live in, and keeps the
store in the same subfolder the runtime uses. Callers pass that directory;
store-only tests pass their temp directory unchanged. Format tests that
share a directory with harness calls take the file path from
testAgentSessionStoreFilePath.

The folder name moves from a private constant in the runtime to
AGENT_SESSION_STORE_DIR_NAME beside the store's file name, so the harness
shares it without importing the runtime. Its value and every path built
from it are unchanged.
2026-09-29 23:38:43 -07:00
Brennan Benson e594cb06af test(mobile): record RPC goldens without a pinned commit, and check recorded requests against the desktop's params rules (#23732)
* test(mobile): add rpc:diff to decode what a recording change moved

The RPC recording goldens are content-addressed JSON, so their raw git diff is
pool hashes. `pnpm --dir mobile rpc:diff [<base>]` decodes both sides and prints,
per golden, the checkpoint, field and JSON path that moved with both values,
grouped across checkpoints, plus added and removed goldens. `--summary <file>`
appends a Markdown report capped for GitHub's step-summary limit.

It reads any pooled format, so it can prove the next commit's format change
moves no recorded value. Checkpoints are matched by occurrence because an id can
repeat within one golden.

This commit adds files under the recorder directory, which moves the header
digest every golden pins; the next commit removes that header.

* test(mobile): record RPC goldens without a pinned commit or input digests

Every golden carried a pinned `baseline` commit plus digests of the recorder,
its mount adapter and its scenario, and the record script refused to run unless
the product tree matched the pin. So every behaviour change repinned to its own
branch commit and rewrote all ~790 files, the squash made that commit
unreachable, and main's pin job stayed red until a hand-made repin pull request
landed (22 of them in 12 days). The digests could only fail when an input moved
and the recording did not, which is exactly the change that carries no
information; every run already re-derives each golden from the current tree and
compares it.

Format 6 keeps the format version, operation, family, named deltas, the value
pool and the recording. Removed: the pin and fence, the three digest modules and
their test, the pin guard and its CI job, and the dead scenario `version` field
(the manifest reader now refuses `baseline` and `version` with a message).

- `pnpm --dir mobile rpc:record [<golden-id>...] [--prune]` records all or some
  goldens; orphans are listed, and deleted only with `--prune`. Every derived
  test title now starts with its golden id so an id selects it.
- `compareGolden` reports every difference in one failure (identity fields by
  name, the checkpoint list, each checkpoint/field/path grouped), keeps the
  final byte compare, and ends with the command to re-record that golden.
- `unhandled-recording.test.ts` now drives a detached rejection through
  `runRecording` into a checkpoint and the cleanup checkpoint; no golden carries
  one, and disconnecting the capture passed every suite before.
- Seam rules that existed only to keep a digest honest are gone; the
  mutant-reachability, register-completeness and one-exposure rules stay.
- CI: `Mobile tests on main` runs the whole mobile suite on every merge that
  touches mobile/, src/shared/, the root lockfile or the host RPC paths, since
  `verify` never runs on main. A new `Mobile RPC Recording Replay` workflow
  replays the recordings on pull requests that touch src/shared/ or the root
  lockfile without touching mobile/. `verify` writes the `rpc:diff` report to
  the job summary.

Proof: `rpc:diff` against the parent reports no recorded behaviour moved; each
golden only loses its ten header lines.

* test(mobile): check every recorded request against the host's params contract

The goldens script the host's replies, so a scenario could record a success
for a request the real host would refuse, and a desktop change that tightens a
params schema moved no golden at all.

`recorded-request-params.test.ts` parses every distinct request the corpus puts
on the wire with the host dispatcher's own `parseRpcRequestParams` and the
schema `rpc-params-catalog.generated.ts` binds to that method. It fails on a
method the host lacks, params it refuses, params sent to a method that takes
none (the dispatcher never reads them), and keys the schema silently strips
unless an inventory entry gives the reason; a stale entry fails too. Each rule
is also shown firing on a made-up request, since the corpus has no instance of
three of them. It imports the desktop dispatcher, so it sits beside the other
Node-side tests outside the RN test program, and the params-contract boundary
now exempts test files, which are never bundled.

It found twelve requests the host would refuse, all from invented fixture
values, not product code, fixed at their source:
- git.branchDiff sent `base-oid`/`head-oid`/`merge-base` where the host needs
  full object ids (diff-review and source-control adapters, and the branch
  compare replies in the manifest that feed them);
- an iOS push registration without `apnsEnvironment`, which a real iOS token
  always carries (`push-token.ts`); the adapter now defaults to `production`;
- `settings.update` given Linear's `assigned` filter as a GitHub preset, which
  the product type forbids; the scenario now picks `my-issues`;
- GitLab `projectRef` as a string where the host and the product type take
  `{ host, path }` (7 methods, 5 adapters and the manifest).

46 goldens move, and a decoded comparison of every one of them shows no change
other than those substitutions; `rpc:diff` lists them.

* ci(mobile): detect a mobile change without a SIGPIPE-prone grep pipe

Under the runner's pipefail, grep -q exiting on its first match SIGPIPEs git
diff on a long file list, so a large pull request touching mobile/ read as
uncovered and replayed the recordings a second time.

* test(mobile): drop comments that still describe the golden header and digests

Eleven adapters justified an import rule by the header a golden no longer
carries, and that rule's test is gone. The census failure now names the
rpc:record and --prune commands.

* test(mobile): refuse a golden that keeps a key no recording writes

Decoding dropped unknown top-level keys, so an old header left behind by a
hand-resolved merge conflict passed every compare unseen.

* ci(mobile): summarize RPC recording changes after a failed test step too

* test(mobile): stream rpc:record output instead of capturing it

A captured run stayed silent for its whole duration and clipped its tail,
where the failure summary sits, past 8 MB.

* test(ci): let the Ruby-gate contract skip the always-run RPC summary step

fef088d8f4 gave the summary step an `if: ${{ !cancelled() }}`, and this test lists every gated step in `verify` and expects each to be gated on the Ruby scope.

* test(mobile): replay only a golden file that is exactly what rpc:record writes

Replay compared two re-encodings of decoded values, so anything decoding drops (a leftover header
key, a hand edit) sat in the committed file uncompared; a key allow-list covered one case of that.
Replay now passes only if the file text equals the formatted golden for the run, sharing one
formatter with writeGolden, and keeps the field-level report as the failure message. The allow-list
goes; the value-based compareGolden stays for the bridged run, which has no file.

* test(mobile): end a corrupt or hand-edited golden's failure with the re-record command

A hand edit to a pooled value failed in decode with only "Golden value <hash> does not hash to its
pool key": no golden id and no command to fix it. readGolden now prefixes parse and decode failures
with the golden id and ends them with the rpc:record command. The rpc:diff header also said it
always exits 0; it exits non-zero when git or a golden cannot be read, and now says so.

* ci(mobile): run Mobile Checks on every src/shared and root lockfile change

Replaces the replay-only workflow: mobile imports hundreds of shared modules, so a shared edit can
move a golden or break mobile's typecheck, and the full job catches both before merge. A root
lockfile-only change skips the Ruby release checks, which read no root Node dependency.

* test(mobile): list or prune orphaned goldens even when the recording run fails

Orphans come from the manifest, not the run, so a failed or timed-out rpc:record still reports
them; the exit code stays non-zero. README: say what a failed replay reports (first differing
path per field, capped groups) and what rpc:diff compares with and without a base.

* test(mobile): pin the RPC recording goldens to LF so a CRLF checkout still replays

Replay now requires the committed golden text to equal exactly what rpc:record writes, which is
LF. A Windows checkout with core.autocrlf=true converted every golden to CRLF and failed all 790
with "holds the same recording but is not the file rpc:record writes for it".
2026-09-29 23:26:55 -07:00
Neil 45c63a66e9 test: delete the source-grep tests an earlier detector's regex missed (#23976)
A rebuilt detector found 195 source-grep candidates where the original found 111.
The gap was one over-specific regex: the first scanner required a literal `.ts`
path inside `readFileSync(...)`, so every test that built its path from variables
(`join(dirname, '..', 'foo.tsx')`) was invisible to it. Roughly 84 files of a
pattern an earlier wave reported as cleared had in fact survived.

Deleted whole, every case asserting on production source text:
- `app-startup-routing.test.ts` (27 cases) — exact import statements
  (`"import('../components/UpdateCard').then"`), relative-path spelling, and
  `indexOf` source ordering. A file move or a `lazy()` refactor breaks it.
- `pull-request-page-host-boundary.test.ts` (13) — `toContain` on whole argument
  expressions concatenated across 20+ component files.
- `SmartWorkspaceNameField-source-boundaries.test.ts` (7) — placeholder copy, a
  Tailwind class string, and `not.toContain` on an already-deleted symbol.
- `github-project-repo-list-load.test.ts` (9) — `indexOf` statement ordering
  inside `loadTasks`.
- `github-enterprise-slug-routing-boundary.test.ts` (4) —
  `toContain('host: githubProjectHost(parsed?.slug.host)')`.
- `web-viewport-shell.test.ts` (3) — a regex demanding exact CSS selector-list
  ordering and whitespace.
- `agent-catalog-links.test.ts` (1) — restates two `homepageUrl` literals straight
  out of `agent-catalog.ts` with nothing in between.

Trimmed, keeping only what nothing else can reach:
- `desktop-startup-ordering.test.ts` 549 -> 66 lines, retaining the three cases
  named as `assertionRefs` by the `ssh-filesystem.stream-inactivity-lifecycle` and
  `agent-browser.owner-boundary-cleanup` gates; 15 source-order greps went.
- `ResourceUsageStatusSegment.session-polling.test.ts` keeps its census that no
  `setInterval` exists and `listSessions()` is called exactly once — an added poll
  multiplies a global daemon scan and no behavioral test sees it. The
  `indexOf('if (!open)')` ordering pair and four `not.toContain` lines went.
- `agent-skill-installed-command-callers.test.ts` 231 -> 86, keeping the
  `readdirSync` census that discovers every `<AgentSkillSetupPanel` caller and
  asserts set equality against the allowlist, so a new panel host cannot silently
  show a default Update action.

Also in this wave, from the renderer lib/runtime sweep: 22 cases whose routing
signal the production path never reads — verified by mutation, stripping
`connectionId`, the WSL preference and the UNC path from four of them left all 29
tests passing — plus braille-spinner rows collapsed onto one regex range, copied
`WELL_KNOWN_LABELS` rows, and a whole `resolveAiVaultResumeStartupShell` describe
whose four darwin/linux fixtures all return before the login shell is read.

`config/reliability-gates.jsonc` drops the two `app-startup-routing.test.ts`
references; the manifest still validates for 140 gates.
2026-09-29 19:55:50 -07:00
Brennan BensonandClaude 8a38a7a3e6 feat(mobile): mid-turn messages wait as cards above the phone composer (#23736)
* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* fix(native-chat): name a chat write by its target, not the owner generation

A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.

Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.

Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.

* fix(native-chat): every journal append reaches the chats that are open

A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.

A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.

* test(native-chat): an epoch replacement reaches the open chat

* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

* perf(native-chat): a publish behind a delivered commit reads nothing

Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.

* test(native-chat): state why the teardown test's fake journal is safe to cast

* docs(native-chat): say mutation admission checks only the writer lease

* docs(native-chat): drop the send rebase from comments that still described it

* fix(native-chat): a message is accepted, then delivered

A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".

A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.

Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.

A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.

Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.

* fix(native-chat): settle queued messages only for the child that ended

A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.

A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.

The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.

* fix(native-chat): an adoption that fails to import keeps the conversation open

The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.

* perf(native-chat): the recovering open reads the journal once

Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.

* fix(native-chat): an attach that fails after indexing its child leaves no child behind

A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.

* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer

The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.

A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.

* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down

The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.

* fix(native-chat): a message rejected while its chat was closed reads as not sent

A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.

* test(orchestration): name why the readiness settlement fakes are cast

* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent

* docs(native-chat): drop the fence from the admission the send effects run behind

* docs(native-chat): give the fence move on release the reason that still holds

* docs(native-chat): stop citing a write fence check in launch and mailbox comments

Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.

* refactor(native-chat): the provider child is its own record

A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.

- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
  patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
  own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
  the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
  dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
  is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
  the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.

* fix(native-chat): the delivery loop alone settles a message its start or child failed

A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.

- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
  reads how it ended: a Stop continues; anything else writes one failure row and rejects every
  queued message with the same words, then stops. A child still starting whose start the adapter
  says did not land fails the same way. The exit, eviction and the settlement retry only settle
  the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
  nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
  failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
  closed, with or without a child, and a start the loop already has in flight is waited for so the
  child it produces is stopped rather than left behind.

* refactor(native-chat): a stopped child ends on the one reading of its stop

The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.

* feat(native-chat): the host says it accepts a send before any agent has it

The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.

* refactor(native-chat): an attach never opens a journal of its own

The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.

* fix(native-chat): a moved fence resends nothing on a host that accepts first

The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.

The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.

* refactor(native-chat): a child's end says whether the user or the host stopped it

The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.

* fix(native-chat): a chat whose only work is a queued message is not offered for resume

A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.

* test(native-chat): type the queued-message fixtures in the resume-offer tests

* fix(native-chat): a start that dies while a message waits on it is that message's failed start

Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.

* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now

* test(native-chat): pin what a failed start settles, and what a resume offer names

A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.

* test(native-chat): the failed-start pins fail on what the message became, not on a timeout

* test(orchestration): the preamble's host stub is typed, not cast

The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.

* fix(native-chat): a Stop that names no turn stops what the conversation has in flight

Between handing a message to the agent and the agent opening its turn, there is no turn id a
client could name, so a Stop in that gap was refused as "already finished" while the agent went
on to answer. A cancel's turn id is now an optional precondition instead of its target: with
none, the host withdraws what is queued and, when the journal still reads working, asks the
adapter to stop whatever the child has in flight. Claude's interrupt is session-scoped, so it
is guarded by fence and acquisition generation rather than a turn identity. Codex interrupts
the turn its latest turn/start answered with until the journal shows one.

A cancel that names its turn behaves exactly as before.

* fix(native-chat): Stop is there from the moment a message is sent

The composer showed Stop only once the agent had opened a turn, so for the second or two after a
send the chat read "thinking" with no way to stop it. Against a host that takes a Stop naming no
turn, Stop now shows whenever the chat reads working (a turn, a queued message, or a handed-over
one still unanswered) or this client still has a message on its way. Pressing it, or Escape,
first drops every outbox entry the journal does not hold yet, so nothing goes out after the
Stop, then sends the conversation-wide cancel. A send already on its way reaches the host ahead
of the cancel, which withdraws it there. Against an older host Stop still needs a running turn.

The unconfirmed-send probe moves into its own hook so the outbox hook stays in budget.

* fix(native-chat): Stop before a turn is gated on its own host capability

A host that accepts sends first (agent-session.accepted-send.v1) can still predate the cancel
that names no turn and would refuse it as invalid, since clients and hosts ship independently.
Hosts that take that cancel now advertise agent-session.conversation-stop.v1, and the renderer
shows Stop before a turn opens, and sends the no-turn cancel, only to a host advertising it.
Every other host keeps a Stop that needs, and names, a running turn.

The host capability probe the accepted-send hook used is generalized so both read one path.

* test(native-chat): a build advertises conversation stop exactly where its cancel may name no turn

* fix(native-chat): a view never restarts a chat whose last start failed

A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.

* test(native-chat): start the child the loop waits on with an attach, not a second view

A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.

* fix(native-chat): settle a gone generation's turn wherever a conversation opens

A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.

* test(native-chat): prove the next child's start settles the turn an earlier child left

The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.

* test(native-chat): count a failed start's rows by row, not by text

Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.

* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget

* test(native-chat): pin the open's and the send's start and row counts, however the view binds

Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.

* fix(native-chat): settle a gone generation's turn at every open but an acquisition's

The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.

* test(native-chat): hold the create's start open until the views bind

The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.

* fix(native-chat): Stop reads the one working rule every session list reads

While Claude retries a rate-limited request it never echoes the message, so no
turn opens: the sidebar read Working from the unanswered send while the composer
showed Send. The chat's working state, the host's session-list status and the
host's no-turn Stop check now call one shared rule instead of three copies.

* test(native-chat): a rate-limit retry pins only that no turn opens, not how its rows are kept

* fix(native-chat): Stop leaves a message waiting on its Retry, and does not show for one

A send that failed holds the queue until the user retries it, and one the host restarted under is
parked the same way. Stop counted both as still on their way, so it showed in an idle chat and
could never go away, and pressing it dropped the failed message along with its Retry.

* test(native-chat): the chat's Stop and a session list read the main agent alike over their own copies

The chat reduces its stream and a list reads the status feed. Driven through the real host for a
rate-limit retry with no turn, a subagent still running after the main turn, and the handed-over
child exiting.

* refactor(mobile): the chat reads the main agent's working state through the shared rule

Behaviour is unchanged: the same two terms, now from the one function the host projection and the
desktop chat read.

* fix(codex): a Stop naming no turn never interrupts an earlier turn

It fell back to the id an earlier turn/start answered with when the latest start went unanswered,
or when the journal showed a compaction Codex had not started, and reported that as stopped.

* fix(native-chat): a Stop naming no turn never says a turn had already finished

When the provider found nothing left to stop, for instance a turn that ended between the host's
check and the interrupt, the chat got "The provider had already finished this turn." for a turn
the Stop never named. It now ends quietly, as a Stop with nothing in flight does.

* fix(native-chat): one Stop the host could not settle no longer refuses every later one

A Stop naming no turn has one operation key per session. When the host could not settle one, it
answered every later Stop under the same id as unknown until the id expired. Once the host says
so, the next press is a new Stop; transport doubt still replays the same id.

* refactor(native-chat): drop the composer's second error formatter

After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.

* test(native-chat): pin the reason on a message rejected while its chat was closed

The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.

* test(native-chat): read Stop operation ids without a cast

* fix(native-chat): a Stop whose answer was lost no longer swallows the next one

A Stop that names no turn has one operation key per chat. When its answer was lost in transit, the
chat kept the id, so every later Stop replayed it; the host answers a replay as already handled, so
for up to a day Stop stopped nothing. The id is now dropped once the call settles, however it
settles. A second press while the first is still on its way still shares its id.

* refactor(native-chat): a Stop naming no target keeps its operation id only for its own call

The chat kept each write's operation id per payload across calls, and dropped it only on some
settle paths. That is right for a write naming what it acts on, but a Stop naming no turn, and a
stop of every background task, share one payload with every later one, so any path that kept the id
made the next Stop replay as already handled and stop nothing. One path was still open: an answer
that arrived after the chat moved to a new fence.

Whether a write names its target is now decided once, before its id is picked. One that names none
keeps its id only while its call is in flight, so a press made meanwhile joins it, and releases it
when the call settles, however it settles. The release runs only while the key still holds that
call's id, so a joined call settling late cannot drop a newer one's. This replaces the per-path
exceptions for a thrown call.

* test(native-chat): read the Stop fences without a cast

* test(native-chat): pin the new id for a named cancel the host could not settle

After the Stop naming no turn moved to a per-call id, the only test of the unknown-refusal release
was gone, and the half that stays, for a cancel naming its turn, could be removed with every test
green.

* fix(native-chat): a Stop pressed after a new message stops it, even while the last Stop is unanswered

A Stop naming no turn shared its operation id with any press made while it was still in flight. The
host runs a chat's writes in order, so a message sent between two presses was accepted after the
first Stop ran, and the second press replayed that Stop as already handled and left the message
running, although the chat had already withdrawn it from the outbox.

A write naming no target now gets a new id on every press and is never kept, so each Stop acts on
whatever is running when the host reaches it. A write naming its target keeps its id exactly as
before. A double press can ask the provider to stop the same turn twice, which it tolerates.

* fix(native-chat): Stop no longer blinks off as Claude opens the turn for a message

Claude's echo of a sent message both answers the send and opens its turn. The echo settled the send
first, so the host published the message as answered one frame before the turn it opened, and for
that frame the chat read nothing running: Stop turned back into Send, and Working blinked off in
every session list, for tens of milliseconds on each turn.

The echo now settles the send after the turn it opens has been emitted, so the running turn is
published first.

* fix(native-chat): a message a Stop withdrew comes back to its sender's composer

A Stop withdraws every message the host holds but has not run, and S also
drops the ones this client had not handed over yet. Either way the message
left the chat and its text survived only in a hidden journal row and the
in-memory ArrowUp history.

The sending client now puts the withdrawn text and images back in that
pane's composer, after whatever is typed there. Withdrawn is read from the
rejection reason through one shared check, which the outbox reconcile now
uses too. The composer is written before the entry leaves storage, so a
failure between the two repeats the text instead of losing it, and an entry
storage no longer holds is never given back again, so a replay, a second
view or a remount restores it once. Only this client's outbox holds the
entry, so other viewers still see the message disappear. A failed Stop
withdraws nothing on the host and gives nothing back.

* fix(native-chat): withdrawn text put back during an IME composition is not lost

While the IME owns the field, the composer ignores a programmatic draft, and
the next composed keystroke wrote the draft without the restored text, after
its outbox entry had already been dropped. The composer now holds text
appended mid-composition, keeps it in the cache after each composed write,
and shows it once the composition settles, the way attachments that land
mid-composition already wait for it.

* test(native-chat): pin that only a withdrawn message comes back to the composer

* test(native-chat): set up the composer's window API for every describe in the composition-race file

* docs(native-chat): note that the withdrawn check reads the legacy reason until a typed category lands

* test(native-chat): pin that text put back mid-composition shows once, even beside a mid-composition clear

* feat(native-chat): host-owned queued-message draft store in the session journal

A queued mid-turn message is a draft row in the session's journal.db,
created idempotently at every writable open with no user_version bump so a
downgrade stays writable. Consume converts one draft into an ordinary
submission inside the journal writer's own transaction (exactly-once), and
a standing writer hook returns a consumed draft only when a committed row
newly settles its current consumed submission to a non-withdrawn rejection
— the same decision the reducer folds rows through. Open-time repair
re-derives returned state behind the stored fact; retention never prunes a
row whose refusal could still return it.

* feat(native-chat): queued-messages wire contract, dark capability, and send classifiers

The send result becomes a union: today's submission arm unchanged, plus a
capability-gated queued arm only clients that sent delivery:'queue-if-active'
ever receive. Whole-list queuedMessages fields ride the subscribe events and
history pages; Stop gains withdrawQueued with the withdrawn bodies in its
result; clear's result carries withdrawn drafts too. Both classifiers treat
queued as accepted/spent. agent-session.queued-messages.v1 is defined but
deliberately NOT advertised: the rollout prerequisites (Claude fold receipt,
integrated Codex steer matrix) are not in this host.

* feat(native-chat): queue a capable mid-turn send as a draft, drain it at turn end, and let Stop and clear return its text

A send carrying delivery:'queue-if-active' while the session owes work — or
behind an actionable backlog — becomes a host-held draft instead of a
submission. A serialized drain woken by journal commits, draft mutations and
conversation opens re-derives its gates from live facts (streamed-event
barrier first, backlog never a gate) and converts the oldest actionable
draft through the exactly-once consume; from that instant today's delivery
pipeline runs unchanged. Stop pauses the withdrawable frontier at the stop
step (a process-level pause set that survives handle eviction and, via the
per-process host instance, restarts), then withdraws it with the text in the
result for capable clients; /clear does the same for the superseded source.
The draft list publishes whole per emit with identity dedup, rides only the
final catch-up page, and attaches to history pages. queuedMessageSend
overrides queue policy only; queuedMessageDelete hands the body back.
Replays for all of it answer from op-stamped tombstone receipts.

* test(native-chat): pin mid-turn queueing against the real host

Accept (working/backlog/text-only/budget/replay), the one-per-settle drain,
returned cards with N1 overtake and the N4 re-send loop, Stop withdraw with
tombstone replays, the process-level pause across evict/reopen, Delete
receipts, /clear returning the withdrawn text, and publication (hydration,
unchanged-cursor insert, same-frame consume, identity dedup).

* test(native-chat): read the queued receipt ids before the wait closures

* chore(native-chat): SAFETY rationales on the sqlite row casts and a cast-free mobile narrowing

* fix(native-chat): queued-draft bookkeeping never costs a publish, an open, a clear or a history read

- Cache the draft list per draft-table revision. The drain re-checks on every
  journal publish, so each streamed delta was running a SELECT and parsing
  every draft body the handle had ever written (tombstones included).
- Open-time repair/prune failures are reported and skipped; they no longer
  fail opening the chat.
- /clear on a source with no drafts answers exactly as before: no empty
  `withdrawnQueued`, no empty write transaction, no extra publish. A draft read
  failure after the committed clear no longer turns it into a refusal.
- History pages read drafts through the same guarded reader as subscribers.
- Publication moves to its own module; the held-draft rule lives with the
  pause state; one pending-prompt check; drop an export nothing calls.
- Tests: restart-held drafts, pre-consume failure pause + Send retry, failed
  open repair, clear with no drafts.

* fix(native-chat): a Stop that withdraws a consumed draft's send gives its text back

A queued draft converted into a submission leaves the sender's outbox, so when
a Stop withdrew that submission before the agent received it, the text had no
holder: the draft stayed `dispatched` forever and nothing restored it.

- The returned-card rule now follows every effective `rejected` settlement of
  a consumed draft's submission, a Stop's withdrawal included, with the
  withdrawal reason stored as the fact (`dispatchWasWithdrawn`). The writer
  hook and the open-time repair share the rule, so no rejected submission can
  leave its draft `dispatched`.
- A capable Stop withdraws the cards it returned itself along with its
  frontier, stamped with its caller-scoped key: the text comes back once in
  `withdrawnQueued` and replays from the tombstone. An old client's Stop
  leaves a returned card.
- Stop's draft steps move to structured-agent-session-queued-stop.ts.
- Tests: Stop between consume and the agent's receipt for both client kinds,
  its replay, a crash after the withdrawal, restart in the window, and the
  repair of a hookless withdrawal.

* perf(native-chat): the queued-draft drain takes no serialized step while the agent works

The drain was woken by every journal publish and, with a draft waiting, queued
a serialized step (streamed-event flush included) per publish, only to find the
session still working. During a streamed turn that is one step per delta,
contending with Stop and every other mutation for the session's queue.

The pre-check now also skips while the session is working. Whatever ends the
work is itself a commit that schedules again, and the step still re-reads every
gate after its flush, so no wake is lost.

- Test: queued sends during a turn take no drain step; settling the turn drains.

* fix(native-chat): a clear withdraws queued text only for a caller that can take it back; paused reasons are markers

An older client running /clear had its source's waiting and returned drafts
withdrawn and their text returned in a `withdrawnQueued` field it does not
read, so the text was lost. Clear now mirrors Stop: `withdrawQueued: true` on
`agentSession.conversationCommand` (strict params, sent only when the
queued-messages capability is advertised) withdraws the drafts and returns
their text once, replaying from the tombstones. Without it the source keeps
its cards: the supersession fence already blocks the drain, and Delete still
hands the text back.

A paused card's reason was host-authored English on the wire. It is now a
typed marker (`send_failed`) the client localizes, like `returnedReason`; a
client treats an unknown marker as a plain pause.

- Tests: an old client's clear leaves the cards and its replay stays
  field-free, then Delete returns the text; a capable clear returns the text
  once and replays it; the paused marker.

* fix(native-chat): a draft pause that commits no journal row still reaches live subscribers

A pause writes no journal row, so it reaches subscribers only on the next
publish. Two pauses had none behind them: the drain's pre-consume failure
(the session is idle by then, so nothing else commits) and an old client's
Stop that interrupted nothing. A live card kept reading as waiting, with no
failure marker, until some unrelated commit arrived.

The drain now publishes after pausing a draft it failed to convert, and an
old client's Stop publishes when it paused a frontier.

- Tests: a failed conversion and an idle old-client Stop each reach a live
  subscriber as a paused card; both fail without the fix.

* fix(native-chat): a failed clear wakes the queued drain, a failed Stop withdrawal still publishes its pause

A conversation command can settle on the record alone (a retried clear that
fails), so drafts held behind its prepared phase waited for an unrelated
journal commit; the command controller now re-derives the drain when any
command finishes. A capable Stop whose withdrawal write failed never
published the pause it set, and a publish failure after a committed
withdrawal (Stop or clear) dropped the bodies from the answer; publishing now
happens outside the withdrawal and can no longer discard its result. Tests
reset the process-level pause set between cases: operation ids repeat per
test, so a shuffled order held later tests' drafts.

* refactor(native-chat): the draft store notifies through the journal's commit listener, the hold is a stored row fact, and one typed gate decides every queue hold

R1: every standalone draft-table transaction that changed rows (insert,
withdraw, hold, open-time repair) fires the journal's own commit listener
after COMMIT, so a draft or hold change publishes and wakes the drain through
the same path a journal row does — no call site can forget. All hand-written
publish/wake plumbing for draft changes is deleted; wakeQueuedDrain survives
only as the record-input wake (a conversation command can settle on the
record alone).

R2: the process-level pause set becomes a hold_reason column on the draft row
(pre-ship, so no migration): holds survive eviction and restart, keep their
send-failed marker across restarts, die with the session's journal, and are
cleared by consume and withdraw in their own UPDATE. The host-instance
derivation stays the one restart mechanism.

R3: one typed structuredQueueHold (blocked | command | prompt | working)
consumed by admission, the drain step and Send-now, with each caller's
override set written beside it. A capable send during a late-result /compact
now queues instead of being refused (PLAN §3.1); the dead prepared-command
branches and the drain's duplicated gate list are gone. prompt outranks
working so Send-now's one override cannot swallow it.

R4: one isUnsettledQueuedMessage predicate for the withdrawable/budget
filters.

Loop 4: a replayed send whose draft was refused answers with the returned
card, never the rejected submission, so the text cannot render twice. Rewind
completion was verified to publish after the record clears (the rewind path's
own publish; the open path's recovery precedes the open snapshot).

* fix(native-chat): a Stop with no drafts writes nothing, and a failed hold still lets a capable Stop withdraw

The stored hold turned Stop's in-memory pause into a draft-table write, so
every Stop (drafts or not, capability advertised or not) opened a BEGIN
IMMEDIATE/COMMIT. An empty hold now returns before the serialized write.

A hold that threw also emptied the frontier, so a capable Stop withdrew only
returned cards and left the waiting drafts unheld to auto-send after the
interrupt. The frontier is read once and survives a failed hold.

* fix(native-chat): a capable Stop with no drafts writes nothing

The empty-hold guard from the previous fix did not reach withdraw, so every
capable Stop still opened a write transaction after the interrupt, and a
closed handle turned its empty answer into a missing field. The draft store
now answers an empty withdraw without a transaction, for every caller.

* refactor(native-chat): Stop and /clear never withdraw queued drafts; no text rides the wire back

Adopt the host-owned-queue model end to end: a Stop holds the waiting
frontier ('stopped') for EVERY client and interrupts — the cards stay
published as paused, Send-now overrides per card, and the pause dies when
the user next starts a turn (an ordinary dispatched send lifts 'stopped'
holds in the same serialized step; 'send_failed' holds still need their
explicit Send). /clear carries the source's unsettled drafts to the
replacement session as born-held rows — identical for every client
version — then tombstones the source. Delete answers with no body: the
card leaving the published list is the outcome.

Removed (never shipped; the capability was dark and unadvertised, so no
wire compatibility is affected): CancelParams.withdrawQueued and its
refine, ConversationCommandParams.withdrawQueued,
CancelResult.withdrawnQueued, ConversationCommandResult.withdrawnQueued,
AgentSessionWithdrawnQueuedMessage, the Delete result body,
settleStopQueuedWithdrawal and the cancel finisher,
withdrawClearedSourceQueuedMessages, replayWithdrawnQueuedMessages, and
cancelPlan's tombstone replay. This also removes the defect where a
withdrawal took every row regardless of which client sent it (a phone
Stop pulled desktop-typed text): nothing moves text anymore, so a Stop
from one client can never relocate another client's drafts.

Hold and carry writes are bookkeeping: a failure is logged and never
gates the interrupt or the clear.

* feat(native-chat): a restart hold lifts like a Stop's, and paused cards say why

The user's next dispatched send lifts every stop-shaped hold in one
UPDATE: stored 'stopped' rows, and restart-held rows (host_instance
mismatch), which are adopted into the running instance — the same fact
the derivation reads, so no second copy of the hold exists. 'send_failed'
still requires its explicit Send. Publication now marks stop/restart
holds with pausedReason 'stopped' (an additive optional value on a dark
capability), so clients can caption them "sends after your next
message" and keep "couldn't send" for 'send_failed'.

* fix(native-chat): only a client's own send lifts a Stop's queue pause

The lift ran for every accepted host send, so orchestration mail, a
restart continuation and a launch prompt released drafts the user had
stopped (and adopted restart-held rows into the running instance). The
client-facing agentSession.send RPC now marks its sends as the user's
own; host-internal senders leave the pause alone. Also drops comments
still describing the withdrawn return-text rule.

* fix(native-chat): a Stop's queue pause lifts when the user's send starts its turn

The pause lifted as soon as the host accepted a user send, so a send the
provider then refused (a failed child start, a refused turn/start) had
already released the stopped drafts into the same failure. The host now
remembers a client's own send, in memory, until the provider answers it:
acceptance lifts the stop-shaped holds, a refusal forgets it with the
holds intact, and a later Stop supersedes it. Nothing is persisted, so a
restart between the send and its turn start leaves the cards held for the
user's next send rather than sending them unasked.

* fix(native-chat): a consumed draft's turn starting lifts a Stop's queue pause

Drafts are only ever a client's own sends, so a drained draft or a
Send-now is a user send for the pause: its submission joins the same
in-memory set a direct send uses, and the provider accepting it lifts the
stop-shaped holds. Before, a message typed while a stopped turn wound
down drained as a draft and left the older stopped cards held, so their
"sends after your next message" caption was false. A refused consumption
lifts nothing, a later Stop still clears the set, and orchestration mail
and restart continuations still never lift.

* fix(native-chat): queue a capable send behind a /compact and re-scope /clear's carried drafts

- A text send with queue-if-active during a /compact in flight is admitted on the
  compact's side lane as a held draft instead of being refused; it may only become
  a draft, so one the gate no longer holds is refused rather than dispatched.
- Drafts /clear carries to the replacement are fingerprinted for the replacement
  session, so the provider's echo folds into the sent bubble.
- The in-memory set of user sends awaiting their turn is capped; sends settling
  unknown no longer grow it without bound.
- Correct the userSend comment: the renderer's launch prompt goes through the
  client RPC and does set it.

* feat(mobile): render host-queued drafts as cards with Send-now/Edit/Delete, and restore withdrawn text once across reload

* feat(mobile): /clear withdraws queued drafts on capable hosts, and reason markers map to readable copy

* fix(mobile): an empty queued-draft publication never churns the held empty list

* fix(mobile): relaunch recovery replays results only — a persisted Stop or /clear never re-executes

* fix(mobile): a relaunch never reissues a Stop or /clear, a lost withdrawal answer is re-asked, and the send journal stays readable by older builds

Relaunch recovery no longer recomputes the host's replacement-session id: that
copy of the host's derivation had already drifted (full digest vs the host's
40-hex slice), so it could never fire. A previous process's Stop and /clear
handles are now released once per pane per process — the host keeps
unwithdrawn drafts visible as cards — and only an Edit is finished. The sweep
runs once per process in one serialized journal step, so a remount can no
longer drop the handle of this process's own in-flight Stop and lose its text.

A withdrawing Stop, /clear or Edit whose answer is lost is re-asked under the
same operation id (bounded): once the host committed, the card is gone and only
that answer carries the text back.

`delivery` is no longer a persisted send-journal field — an older build (or an
older host's page, which reads the same key) would find the strict schema
unreadable and refuse every structured send. It is part of the intent key
instead; the immediate key is unchanged, and a retained entry under either key
replays exactly as first sent.

An identical send whose retained id replays as a withdrawn draft goes out under
a fresh id instead of vanishing. Also: the send callback is stable across
streamed frames, card busy state is per card, card actions carry a button role,
and render-time ref writes moved into layout effects.

* fix(mobile): a lost-answer queued send never reads as unconfirmed, the restore journal works inside the page, and its handles die

- A send whose answer was lost but which the host holds as a queued-draft card no longer warns "Delivery unconfirmed" after 20 s: a card that was not on screen at send time with the send's text counts as delivery, like the transcript echo does.
- The queued-draft restore journal key joins the page storage allowlist and writes through the mirrored path, so a Stop, /clear or Edit from the page-served session screen keeps its replay handle; a write the store dropped without rejecting is refused up front so the caller restores directly instead of holding a handle no store kept.
- Seeing a published draft named by a journaled send's operation id spends that entry: the draft is the host's receipt of the send, so a draft later withdrawn elsewhere no longer leaves an entry nothing will ever settle.
- Withdrawn text is restored at most once even if removing the journal entry fails after the restore ran.
- A returned card offers Edit, as on desktop.
- Test outcome annotations use MobileNativeChatSendOutcome, which now includes 'queued'.

* fix(mobile): withdrawn queued text reaches its pane even after the screen closed, and a lost answer is re-asked in the background

A Stop, /clear or Edit answered after the session screen closed wrote into a
composer that no longer existed, while the journal entry was removed: the text
was lost. Restored text is now owed by composer scope and the pane's composer
takes it once whenever that scope is active.

A lost answer was re-asked inline up to three times, each with the full command
budget, so a /clear could hold the composer for minutes. The first answer now
returns at once and re-asks run in the background on the 1 s / 2 s / 4 s
schedule with a short budget each. Restoration stays once per operation id,
also when a retry of the same id races the re-ask.

* refactor(mobile): queued drafts stay host-owned — Stop and /clear never pull text back to the phone

A Stop or /clear no longer withdraws queued drafts and ships their text back
over the wire. Cards stay on the host as paused cards (captioned 'Paused —
sends after your next message') with Send / Edit / Delete, identical on every
device, and the host carries them across a /clear itself. Edit copies the
card's shown text into the composer first and then issues a plain Delete, so
no RPC outcome can lose it; a failed Delete leaves the card visibly beside the
copy.

With no text in flight there is nothing to make exactly-once: the AsyncStorage
restore journal and its page-storage allowlist entry, the background re-ask
schedule, the relaunch release/replay pass, the restore-once guards, and the
owed-composer-text buffer are all deleted. Cancel and conversationCommand go
back to main's plain requests, so old and new hosts see one Stop and one
/clear behaviour.

Kept: capability gating, the cards and captions, Send-now, the 'queued' send
outcome, the queued-card settlement of unconfirmed sends, the delivery-in-key
send-journal fix, and the queuedMessages frame tracking.

* fix(mobile): only a 'stopped' hold promises "sends after your next message"

The host now publishes pausedReason 'stopped' for Stop, /clear-carry and
restart holds. An absent or unknown marker is a hold whose release rule this
build does not know, so it reads as a plain "Paused" instead of promising
that the next message resumes it. Drops an unused type re-export.

* fix(mobile): a paused queued card's action reads Send, not Send now

"Send now" names jumping the running turn. A paused card (a Stop, /clear
or restart hold, or a failed conversion) waits on no turn, so like a
returned card its action and accessibility label say plain Send, matching
desktop.

* test(mobile): find the queued card's Text nodes by name so the test typechecks

* fix(mobile): retire a replayed send the host refuses by shape; 44pt queued-card targets

A retained delivery send that an older host's strict schema turns away can never be
accepted, so keeping its operation id refused every later send of the same text. The
host's own request refusal now retires it; any other doubt still keeps the id.
Queued-card actions now touch as 44pt targets while drawing as a 32pt text row.

* fix(mobile): a waiting queued card's action reads Steer, as on desktop

The same action read "Send now" on the phone and "Steer" on desktop. Its accessibility
label now matches the desktop hint: "Send now without waiting for the turn to end".
A paused or returned card keeps plain Send.

* fix(mobile): focus the composer after Edit moves a queued message into it

Edit copied the card's text into the composer but left it unfocused, so the user had
to tap the field to keep typing. The composer now takes an input ref and the chat view
focuses it on the next frame after Edit.

* fix(mobile): only a schema rejection retires an in-doubt send; an auth refusal keeps its id

An unauthorized answer to a replay says nothing about whether the first
attempt was delivered, so retiring its id there could send the message twice.

* fix(native-chat): a returned queued card carries the typed rejection fact, like a rejected submission

A consumed draft the agent never ran comes back as a returned card. The card
kept only the rejection's sentence, while its submission now also records the
typed fact a client classifies from. A host-restart rejection's sentence
carries no legacy marker, so such a card could not be told apart from a
provider's refusal.

The draft table stores the submission's fact next to its reason
(`returned_rejection`, written by the same settlement that sets the reason,
and read back with the reducer's own fact reader), and the card publishes it
as `returnedRejection`. Both are overwritten on every return, so a re-sent
card never keeps an earlier refusal's fact, and a /clear carry inserts a plain
held draft with neither.

Retention moves to queued-message-retention.ts to keep the table module
within max-lines.

* fix(native-chat): fit the queue to main's typed rejections and compaction result

Main (#23026) dropped the disposition's fresh-id retry field, gives a
rejected dispatch a typed sentence plus fact, and types /compact's result.
The queued-draft disposition and the queue tests now use those shapes.

* fix(mobile): a returned queued card reads its typed rejection like a rejected send

The card's label called `dispatchRejectionReasonIsInternal`, which main
removed when rejections became typed facts, so labelling a returned card
threw. A host-restart rejection's reason is also a sentence now, with no
marker for the old string check to recognise.

The label now comes from the shared words a rejected send gets
(`structuredAgentSessionRejectionParts`, 'composer-send'), given the card's
`returnedRejection` fact and falling back to `returnedReason` when a host wrote
none. A Stop withdrawal, read through `dispatchWasWithdrawn` with the fact,
keeps "Held back by Stop — Send to retry"; a provider's words show only where
their audience is the person; a fact kind this build cannot place reads as not
sent rather than trusting the sentence beside it.

* fix(mobile): a returned queued card is worded as the desktop card words it

The card has its own Send, so its words leave out sending again, the way the
desktop card passes retryControl to the shared attempt-failure words. A Stop
withdrawal reads "Stopped before it was sent", the desktop caption.

* fix(mobile): a returned card with a fact this build cannot place shows the host's sentence

The host writes a rejection's reason as a sentence for a person, so when a
newer host's fact kind cannot be placed, that sentence is the best words
available. The card now leaves it to the shared attempt-failure words, as the
desktop card does; an old internal marker still maps to generic words there.

* fix(native-chat): draft bookkeeping can never roll back the journal row it rides

The queued-draft returned transition runs inside every journal append's
transaction. A throw there (a draft table an earlier build created without the
returned_rejection column) rolled back the journal's own rejection row, so a
Stop, a failed start or a provider refusal could not be recorded. The standing
hook now runs in its own savepoint: its failure is logged and rolls back alone,
and the open-time repair re-derives the missed transition from the committed
row. The draft table also gains any missing nullable column at open.

* fix(mobile): a returned card's reason reads whole, and Edit never deletes text it did not copy

A returned card's label was capped at one line, but its reason often reads
only at its end ("... does not support the image type .bmp in a steering
message."). It now wraps; waiting and paused holds stay one line.

Edit copied a card's text into the composer and then deleted the card,
whether or not the copy landed: before the composer mounted, or for an empty
card, the append was a no-op and the delete still ran. The append now reports
whether it copied; Edit stops before Delete when it did not, and focuses the
composer only after a copy. When Edit's Delete loses to the drain, the phone
says "Already sent — your text is still in the composer.", as the desktop
does, so the copy is not sent a second time.

The controller now carries the queued controls as one field, which keeps it
within max-lines.

* fix(mobile): a queued send's replay never vanishes or paints an unretired bubble

A replay answered `withdrawn` is resent under a fresh id only when the
retained id could be released. When the release failed, the answer fell
through as `queued`, which shows nothing, and the text was gone. It now reads
as rejected, so the text returns to the composer.

A replay answered `queued{state:'dispatched'}` was mapped to `accepted`. The
host answers that way only when it cannot find the submission the draft
became, so no echo would retire an optimistic bubble. It now maps to
`queued`, which shows nothing; the transcript or the card owns the text.

* fix(mobile): a queued card hides once its message arrives, and a failed pause says tap Send

A waiting card whose submission has already arrived is hidden, as the desktop
card projection hides it: on a multi-page catch-up the shrunk list rides only
the final page, so the bubble and the card showed together. Returned cards
always show.

The send-failed pause reads "Couldn't send — tap Send to retry".

* fix(native-chat): a draft a Stop or restart took back waits again instead of blocking the queue

Cards A, B and C wait; the turn ends and the drain consumes A, but the agent
has not taken it yet. A Stop then pauses B and C and withdraws A's submission,
which made A a returned card. The user's next send lifted B and C, yet a
returned card blocks everything behind it, so B and C never sent although they
read "sends after your next message". A restart or close before hand-over did
the same.

Nobody failed the user there, so the draft now goes back to waiting at its own
position, under the hold that same event put on the drafts behind it: a Stop's
'stopped', or no stored hold after a restart, whose hold derives from the host
instance. It carries no refusal, and records its spent submission id in
consumed_as, so its next consume (the drain, or Send on the card) mints a fresh
id through the same path a returned card's re-send uses. Provider refusals and
other failures still return the card. The live settlement hook and the
open-time repair share one decision. After a Stop and the user's next turn,
A drains first, then B, then C, one per turn.

* fix(native-chat): Delete and Send on a queued card answer at once during a /compact

A /compact holds the chat's serialized lane for its whole provider call, and
the queued-card Delete and Send ran on that lane, so both hung until the
compaction finished. They now run on the side lane a draft-only send already
uses while a compaction is in flight: Delete completes at once, and Send
reaches its readable "wait for the conversation operation" refusal at once.
The drain stays on the main lane and keeps its command hold, so nothing sends
until the compaction settles.

* fix(native-chat): a re-sent returned card drops the refusal it came back with

Re-consuming a returned card left returned_reason and returned_rejection on the
now-dispatched row, so the row described a refusal that no longer applied. The
consume clears both in the same update that moves the card to dispatched.

* perf(native-chat): the queue gate reads pending prompts without rendering the journal

The prompt check ran on every send admission and drain step, and read
journal.snapshot(), which copies and sorts every item in the chat. It now walks
the reduced items in place with journal.visitItems; the answer is the same,
since the snapshot only sorts those items.

* fix(native-chat): a Stop that fails leaves the queued cards as it found them

Stop holds the waiting cards before it withdraws queued sends and interrupts
the agent. When a later step threw or the Stop was refused, the cards stayed
paused ("sends after your next message") although a failed Stop is meant to
change nothing. A failed Stop now undoes exactly what it added: each card it
held gets back the hold it replaced, a consumed card its withdrawal sent back
to waiting is released, and the user sends it had set aside can again lift the
pause. Holds an earlier Stop or a restart put on the cards stay.

The hold SQL moves to its own module, and the draft store's standalone
transactions share one helper.

* docs(native-chat): confirmed cancellation is no longer a queue rollout prerequisite

Stop withdrawing queued sends with a typed cancellation landed on main with
#23026. The comment gating the queued-messages capability now lists only what
remains: the Codex steer matrix (#21062), the Claude fold receipt, turn-owner
bars, and the desktop and phone clients.

* docs(native-chat): the Claude fold receipt and turn-owner bars have landed; Codex steer and the clients remain

* test(mobile): compare the queued card's Text nodes by name so the test typechecks

* fix(mobile): a draft a Stop requeued stays visible beside its rejected first submission

A Stop that withdraws a consumed draft before delivery puts it back to
waiting under its own id, while its first submission stays in the journal as
rejected. The card filter hid any waiting draft with a same-id submission, and
the transcript hides rejected submissions, so the requeued text could not be
seen, edited, deleted or sent, and later drained unannounced. A rejected
submission no longer hides a card: beside a waiting draft it can only mean a
requeue, since a refusal returns the card instead.

* fix(mobile): a send record storage will not clear never blocks sending that text

A replay answered `withdrawn` is resent under a fresh id only after its
saved record is cleared. While storage kept failing to clear it, every send
of that exact text replayed the old id, was answered `withdrawn`, and came
back rejected, so the text could not be sent until storage recovered. The
resend now goes out once under a fresh id that bypasses the saved record, and
the failed clear is reported through the send's error path.

* fix(mobile): a send record whose draft went out under another id is spent

A send whose answer was lost keeps its record, so the same text replays the
same id. When a Stop requeued that draft and it later drained under a fresh
id, nothing could clear the record: the draft was no longer published, and no
submission carries the old id. Every later send of that text replayed it and
waited for a settlement that would never come.

The host answers that replay with the submission the draft became, under a
different id than the one replayed. That answer now spends the record. Which
send the phone meant stays unconfirmed, as for any retained replay of a live
send.

* fix(native-chat): Send on a queued card during a /compact is refused before it takes a lane

Send-now chose its lane once, at entry. During a /compact it took the side
lane, where it could wait behind a Stop, then run after the compaction had
settled and append a real submission unserialized against the main lane.
While a compaction is in flight, Send-now is now answered with the "wait for
the conversation operation" refusal before entering any lane, and otherwise it
runs on the main lane. Only Delete keeps the side lane, whose compare-and-set
withdrawal is safe on either.

* fix(native-chat): a Stop that fails after reaching the agent keeps the queue paused

A failed Stop undid its queue holds whenever it threw, including after the
interrupt had already gone to the provider (a status-note write failing after
cancelTurn, or after stopping a starting agent). The turn could be stopped
while the cards drained as if no Stop was pressed. The Stop now marks the step
that reaches the provider, and undoes its holds only when it failed before
that. A Stop the agent refused answers ok and keeps its holds; the comment no
longer claims otherwise.

* fix(native-chat): a skipped draft settlement heals on the next drain step, not only at reopen

The draft settlement rides each journal append as bookkeeping, and a failure
there is logged and skipped. Only the open-time repair re-derived it, so a
consumed draft whose submission was rejected stayed dispatched (invisible, and
blocking nothing it should) until the chat reopened. The re-derivation is now
its own function, shared by the open-time repair and the drain: whenever a
dispatched draft's submission is already rejected, the drain step applies the
owed settlement first.

* fix(native-chat): one id is never recorded as a submission twice

A second submission row under an id the journal already holds replaces the
submission with a fresh pending one, so a rejected message could be handed
over again under its own id. Send on a queued card could do exactly that: if
the host died after it consumed the card under the operation's id but before
its answer settled, the rerun consumed again under the same id.

The journal now refuses a submission under an id it already records, so no id
is delivered twice whatever the caller does. And a Send-now rerun that finds
the card consumed under its own operation id answers with that submission
instead of consuming again.

* fix(native-chat): a waiting draft whose first send the agent echoed is withdrawn, never resent

A consumed draft goes back to waiting when its submission is rejected as never
delivered (a Stop's withdrawal, a restart, a close), and then sends again
automatically. That rests on the "never delivered" claim. If the provider then
echoes that message, the first delivery happened, and the automatic resend
would give the agent the same message twice.

The reducer already keeps such an echo apart, since a rejected submission may
not claim it, so the draft store reads it from the appended row itself: a
provider echo of a user message that no live submission claims, matching a
waiting draft whose spent submission is rejected, withdraws that draft the way
a Delete would. The echo-claiming rule is split out of the reducer's aliasing
so both read the same decision, and the per-row draft hook moves beside the
settlement re-derivation.

* feat(native-chat): a submission names the queued draft it hands off

Clients told a queued card's hand-off apart from other sends by comparing the
draft's id with the submission's id. That holds only for a draft's first
hand-off: a re-send, or a draft that goes back to waiting and drains again,
goes out under a fresh id, and the clients showed the card and the sent
message together, or restored text the host still held.

Every submission the host creates by handing off a draft now carries
queuedMessageId, the draft's id. It is written on the submission's journal row
as an optional key (older readers keep it and ignore it), carried by the
reducer, listed in the published submission schema (which otherwise strips
it), and stamped where the row is built from the consume itself, so no
hand-off path can leave it off; a caller naming a different draft is refused.
A direct send names none. The queued-messages capability comment makes the
link part of v1.

* refactor(native-chat): every queued draft goes out under a fresh submission id

A draft's first hand-off reused the draft's own id as the submission id, so
comparing a draft id with a submission id looked right in every first-send test
and failed only on a re-send or a requeued draft. Every hand-off now uses a
fresh id (the drain mints one; Send on a card uses its operation's id), so id
equality is never true and a reader must use the submission's queuedMessageId.

The host gets simpler: queuedMessageNeedsFreshSubmissionId is gone, consumed_as
is set on every dispatched row and cleared when a withdrawal sends the draft
back to waiting (its spent submissions stay findable by their link), the
consume refuses the draft's own id, and the consumedAs ?? messageId fallbacks
collapse. The delivered-echo check finds spent hand-offs by link.

A send this host queued, asked again (a lost answer's replay, or a rerun the
operation ledger no longer covers), answers from its draft and then from the
hand-off that names it, through one function. The rerun path used to be kept
from sending twice only because a submission sat under the send's own id;
with fresh ids that guard is now explicit. A Send-now rerun recognises its own
consume by the link instead of consumed_as.

* fix(mobile): relate a queued draft to its hand-off only through queuedMessageId

The phone matched a draft to the submission it became by comparing ids. The
host now hands off every draft under a fresh submission id and stamps that
submission with `queuedMessageId`, so the ids never match and the link is the
only relation.

- A waiting card hides exactly when a submission names it as its
  `queuedMessageId` and was not rejected; a rejected hand-off is what sent the
  draft back. A direct send sharing the card's id hides nothing.
- A saved send record is spent when its id is a published draft id, or when
  any submission names it as its `queuedMessageId`, in whatever state. This
  clears a lost send that a Stop requeued and the drain later sent under a
  fresh id, from the live submissions stream.
- The replay rule that spent a record when the answer's submission id differed
  from the replayed id is gone. A replay answered by the hand-off reads as
  unconfirmed, as any retained replay of a live send does, and paints no
  optimistic bubble; the stream spends the record once it carries the link.

* fix(native-chat): an echo withdraws a draft only if its rejected hand-off reached the agent

The delivered-echo rule withdrew a waiting draft when a provider echo matched
any rejected hand-off of it, including one a Stop rejected before it was ever
handed over. That hand-off is provably unwritten, so a matching unclaimed echo
is some other message, and the rule silently deleted the card. Only a hand-off
that was handed over and then rejected as never delivered can be disproved by
an echo now.

* fix(mobile): a replay the host answers with its draft's hand-off spends the send record

A retained send replayed after its queued draft was handed off is answered by
the hand-off, which names the replayed id as its `queuedMessageId`. That
answer fell to the retained branch and kept the record, leaving only the live
stream to spend it. The stream can miss the hand-off for good: a reconnect
snapshot covers only the latest page, and submissions ride only with their
items. Every later send of that text then replayed and read "Delivery
unconfirmed" forever.

The answer states the link, so it now spends the record directly, checked
before the rejected and retained branches. Which send the phone meant stays
unconfirmed, as for any retained replay of a live send.

* fix(mobile): a resend past an uncleared send record replays its own id on a retry

When storage could not clear a withdrawn send's record, the phone resent the
text under a fresh id with no record. If that resend's answer was lost and the
user sent the text again, the old record replayed, came back withdrawn, and
the phone minted another fresh id: a second delivery if the first resend had
landed.

The resend's id is now kept in memory for the app run, by operation key, so a
retry replays it; it is forgotten once the host answers it as spent. The
"couldn't update its record" notice now shows only when the resend is known to
have gone out (accepted or queued), never for an unconfirmed one.

* fix(native-chat): a skipped echo withdrawal is re-derived before the draft can send again

The delivered-echo withdrawal rides each journal append as bookkeeping, and a
skipped hook left the draft waiting, so it later sent the same message a
second time. Nothing re-derived it. The draft store now also withdraws, in its
owed-settlement pass, each waiting draft that an echo already in the journal
proves delivered: an unclaimed provider user message (still stored under its
own id), carrying the draft's payload, appended after a hand-off that was
handed over and rejected. The live hook and the re-derivation share one
predicate. The pass runs at open and in the drain step, right before a draft
would send; it reads every item, so it never runs per streamed row.

* fix(native-chat): a rolled-back journal append leaves no draft state cached

The draft store caches its row list by revision. The per-row hook read that
list eagerly inside the append's transaction, after the consume in the same
transaction had already written and bumped the revision, so a failed COMMIT
left the cache showing a hand-off that never happened. The hook now reads the
drafts only once a row holds an unclaimed echo, and any rollback of a journal
append or of its bookkeeping savepoint invalidates the cache, so no other read
inside the transaction can leave it stale either.

* fix(native-chat): a replay of a deleted queued card answers withdrawn, not refused

Once a deleted card's tombstone is pruned, a replay of the send that queued it
found the draft through its last hand-off. When that hand-off had been
rejected (the card came back, and the user then deleted it), the replay
answered with the rejected submission, which clients show as a failed send
with a Retry. Only a withdrawn row is pruned while its last hand-off stands
rejected, so the replay now answers queued, withdrawn.

* test(mobile): a withdrawn replay of a pruned deleted card, position 0, still frees the text

The host now answers a replay of a deleted card whose row was pruned with
queued{state:'withdrawn', position: 0}. The phone reads only the state, so
the next identical send still goes out under a fresh id; the test now uses
that receipt.

* fix(mobile): a bypassed resend's id survives storage recovering, and a malformed answer spends nothing

The remembered id of a resend sent past an uncleared record was consulted
only while the record still could not be cleared. If storage recovered
between that resend's lost answer and the retry, the record cleared, the
normal path minted a fresh id, and the message could be delivered twice. The
retry now hands the remembered id to the new saved record, which replays it,
and memory lets it go. The id is also keyed by the operation key the saved
record matched, not the delivery asked for now, so a capability change in
between cannot mint a fresh id.

A send answer is read as its draft's hand-off only when it carries a
submission, so a malformed answer with neither a submission nor an id no
longer spends the saved record.

* refactor(native-chat): name the queue's pause-lift for what it releases

* refactor(mobile): name a schema-refused replay for what the host decided

* chore(native-chat): one import of the mutation helpers

* refactor(mobile): build the queued-card slot outside the chat view

MobileNativeChatView was over its 400-line limit with the queued cards wired
in. The cards and the composer ref their Edit focuses are now built by
useMobileNativeChatQueuedSlot in the overlay, and the view only places the
cards and hands the ref to the composer.

* test(mobile): say why the overlay test's partial controller is safe

* feat(native-chat): a Stop pauses the whole queue, derived from the journal, with an explicit Resume

After a Stop, each waiting card was held on its own row ('stopped'), lifted
when the host saw, in memory, that a user send made after the Stop had its
turn accepted. The cards read "sends after your next message" one by one,
there was no way to resume the queue without sending something, and the
in-memory record of user sends was lost on a restart or eviction.

The pause is now the queue's, and derived rather than stored as a flag:
- 'stopped': the user's last Stop took effect at a recorded journal position
  and no turn a person asked for has started since. "A person asked for it"
  is the new `origin: 'client'` on the submission row (a send over the client
  send RPC, or a card they sent now); orchestration mail, a restart
  continuation, a host-sent launch prompt and the queue's own drain record
  `host` and never lift it.
- 'restarted': a waiting card was written by another host process and no
  person's turn has started since this conversation opened.
Resume (`agentSession.queuedMessagesResume`) lifts either. Send-now sends one
card; the rest stay paused until that card's turn starts, which is a person's
turn like any other.

The journal's row kinds are closed (an older build truncates a journal at a
row kind it does not know), so the one event the journal cannot carry, where
the Stop took effect, is recorded beside the drafts in `queued_message_pauses`;
everything after it is read from the journal. A Stop records it only once it
takes effect (after withdrawing queued sends, as it reaches the agent), so a
Stop that fails first leaves nothing to undo, and the per-row hold, its undo
and `userSendsAwaitingTurn` are gone. A card keeps a hold of its own only when
its conversion failed ('send_failed').

The pause is published once, as `queuePause` beside `queuedMessages`, on live
frames, catch-up and history. A /clear starts its replacement paused, as after
a Stop, since the carried cards were written for the context it discarded.

* feat(native-chat): a /clear's replacement queue reads paused because of the clear, not an interrupt

The replacement's pause was recorded as 'stopped', which clients show as
"Queue paused because you interrupted" although the user cleared the chat.
It is now its own reason, 'cleared', on queuePause.reason
('stopped' | 'restarted' | 'cleared'). It lifts and resumes exactly like a
Stop's: through Resume, or the user's next turn starting on the replacement.

* feat(mobile): a paused queue shows one header row with Resume, and its cards keep Steer

The host now pauses the whole queue after a Stop, a restart or a /clear, and
publishes that pause once beside the list instead of on each card. The phone
followed the old per-card pause: each card read "Paused — sends after your
next message" and its Steer turned into Send.

- One header row above the cards says why the queue is paused ("Queue paused
  because you interrupted", "... because Orca restarted", "... after you
  cleared the conversation"; an unknown reason reads "Queue paused") and
  offers Resume, which calls agentSession.queuedMessagesResume through the
  same mutation path as Delete, so a refusal reaches the error banner. It
  shows only while there are cards.
- Cards under a paused queue keep Steer, Edit and Delete and read "Queued",
  promising no send time. A card whose own send failed still reads "Couldn't
  send — tap Send to retry" with Send, and returned cards are unchanged.
- Steer's accessibility label is "Submit without interrupting the model".

The pause is kept per session beside the list and changes only on frames that
publish the list. The row lives in the queued-cards component, so the chat
view gains no lines. The queued hook tests' published shapes move to a
fixture module to keep that file within its line limit.

* test(mobile): the paused-queue header words the host's 'cleared' reason

The host now sends queuePause.reason 'cleared' after a /clear carries the
queue over. The header already worded it; the tests now pass it as the
host's typed reason rather than as an unknown one.

* test(mobile): import the subscribe-event type the queued hook test names

* fix(native-chat): a queue pause covers only the cards it paused

A Stop recorded its pause fact even when the queue had no cards, and the fact
outlived the cards it did pause. The published list hid a pause over no cards,
but the drain still treated the queue as paused, so a card typed much later —
during an orchestration-mail turn, or a correction typed before the stopped
turn ended — sat under "paused because you interrupted" with no Stop of its
own.

A Stop now records its pause only if the queue holds a card when the Stop takes
effect (the hand-offs its withdrawal sent back included). The fact is retired
in the same transaction as the Delete, consume or withdrawal that empties the
queue, never from an async publish. A /clear's carry now lands each card with
its 'cleared' pause in one transaction, so a failed insert leaves no pause over
an empty replacement.

* fix(mobile): the paused-queue row keeps Resume's whole target and is announced

- The row no longer pulls itself past the top of the card list with negative
  margins. The list has no padding there, and Android drops touches outside
  a parent, so part of Resume's 44pt target could not be tapped.
- The row is a polite accessibility live region, as the app's other notices
  are, so a screen reader hears that the queue paused.
- A pause published with no reason, or one this build does not know, now
  counts as a pause. The reducer compared reasons only, so it kept a previous
  "not paused" for it; it now reads "Queue paused".
- The card list is keyed by conversation, so a Resume still in flight in one
  chat no longer disables Resume in another.

Tests cover a double tap, a lost Resume answer re-enabling the button, the
row's layout and live region, the conversation keying, a reasonless pause,
and a paused queue's failed-send card and prompt wait.

* perf(native-chat): the queue's pause reads the latest person's turn in O(1)

The pause is derived on every publish, per subscriber, and each derivation
copied and scanned every submission to find a person's accepted turn after the
Stop. The reducer now keeps that fact as it folds rows: the submission row of
the latest accepted turn whose origin is `client`. The Stop's and the
restart's lift both read it directly.

* fix(native-chat): a card handed off after a restart belongs to the process that sent it

A draft's host_instance was only ever the process that first wrote it (or
adopted it while waiting). A returned card from before a restart, sent again
in this process and then withdrawn back to waiting, still carried the old
process, so it raised a 'restarted' pause although no restart happened since
it was sent. Every hand-off (the drain, Send on a card) now stamps the
handing-off process on the draft in the consume's own update.

* fix(native-chat): a queue pause shows only while Resume would send something

After a Stop whose only remaining card was a returned one, or after a restart
with only a card held by its own failed send, the queue published a pause with
a Resume that could send nothing: a returned card waits for the user anyway,
and a held one for its own Send. The pause is now published, recorded by a
Stop, and kept only over a card it can hold back — waiting, with no hold of
its own. The fact is retired in the same transaction as the write that removes
the last such card, a hold or a refusal included.

The publication's dedup also compared only the pause's reason, so a pause
appearing or clearing with no readable reason could read as unchanged; it now
compares presence first.

* fix(mobile): the paused-queue header shows only while Resume would send a card

The host publishes a pause while any card is waiting with no hold of its own,
but it does not exclude cards behind a returned card, which the drain never
sends. With only such cards, or only returned or failed ones, the header
offered a Resume that would do nothing. The header now shows only when a
waiting card with no hold of its own sits ahead of any returned card, as on
desktop.

* fix(native-chat): a queue pause counts only cards Resume would actually send

A waiting card behind a returned one is blocked until the user acts on the
returned card — the drain never sends past it — so a pause over only such
cards still offered a Resume that sent nothing. The rule for "a card Resume
would send" is now one function: waiting, no hold of its own, and not behind a
returned card. The publication, a Stop's record and the fact's retirement all
read it; retirement reads the rows in position order inside the same
transaction as the write that took the last such card.

* fix(native-chat): a returned card that blocks the paused cards hides the pause but keeps it

The last change retired a Stop's pause as soon as a returned card blocked every
paused card. Deleting that returned card then sent the cards behind it at once,
with no Resume — not what the user asked for.

The two rules are now separate. The pause is KEPT (recorded by a Stop, retired
in the same transaction as the write that takes the last one) while any waiting
card with no hold of its own exists, wherever it sits. It is PUBLISHED only
while such a card is not behind a returned one, so the header never offers a
Resume that sends nothing. Deleting the blocking card shows the pause again,
and the cards behind it wait for Resume or the user's next turn.

* fix(native-chat): a Stop pauses a card its withdrawal sent back even when that settlement was skipped

The Stop checked the draft table for a card to pause. When the per-row hook
that settles a withdrawn hand-off was skipped, that card was still
'dispatched', so the Stop recorded no pause; the drain later healed it back to
waiting and sent it, although the user had pressed Stop. Retirement had the
same blind spot and could drop a pause while such a card was owed.

What a pause holds back is now one predicate, judged inside the transaction
that records or retires it: a waiting card with no hold of its own (one SQL
EXISTS), or a dispatched card whose consumed submission was rejected with a
settlement back to waiting (read against the journal's submissions). The Stop
first runs the owed settlement, as the drain does; if that fails, the owed
card still counts, so the pause is recorded rather than skipped. recordPause
now checks inside its own transaction and returns whether it recorded, and any
draft-table write (and the per-row hook, the consume and the open-time
repair) retires a pause that no longer holds anything back.

* test(native-chat): pin the per-row hook's pause retirement; skip the judgement when no pause exists

The retirement test recorded its second pause over a queue with nothing to hold
back, so the recording returned false and the "retired" assertion proved
nothing; ablating the per-row hook's retirement passed every test. The hold
case now asserts the pause was recorded, and a new test has a delivered echo,
through the per-row hook, withdraw the last card a recorded pause holds back.

Retirement runs on every appended journal row, so it now checks the pause row
by key first and judges nothing when no pause is recorded. Two comments were
brought in line with the owed-hand-off rule and rewrapped.

* test(native-chat): match main's append and dispatch shapes in the queue tests

* fix(native-chat): read a compaction's settled submission through the send-result union

* test(mobile): a queued card's Steer groups under the running turn

With turn facts read from the turn record, a draft Steer hands over while a
turn runs is scoped by the host to that turn, under a fresh submission id
that names the card. On the phone it joins that turn's group, gets no bar of
its own, and stays live with it.

* feat(mobile): queued messages sit in one compact box, as on desktop

Each queued message was its own tall card with a "Queued" header and a row of
Steer / Edit / Delete text buttons, and the paused-queue line floated above
them. The phone now draws the desktop's layout: one bordered box whose first
row is the pause line with Resume (only while paused), then one compact row
per message with a leading queue or alert icon, the text (up to two lines,
since a phone has no hover title), a caption only when there is something to
say, and Steer (Send for a returned or failed card), a trash button, and a
"..." menu holding Edit message.

The card model's label becomes a nullable caption with the desktop's rules
(no caption for a plain wait or the queue's pause) plus a needsAttention flag
for the alert icon. Touch targets stay 44pt inside their rows.

* test(mobile): stub the queue's action sheet in the chat view tests

* test(mobile): move the chat view's turn-status wiring tests to their own, type-checked file

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-29 18:43:34 -07:00
Neil cd63c998ab test: retire cases whose predicates never read the varied input (#23965)
* test: retire provider cross-products and mock-delegate tests in store, hooks and mobile transport

Semantic sweep of renderer store/hooks and mobile/src/transport. 8 cases and one
file removed across 8 files; no production file touched.

`full-creation-structured-launch.test.ts` is deleted whole. Its subject,
`beginFullCreationStructuredLaunch`, is a four-line forward to
`beginStructuredAgentSessionProvisionalLaunch` with a fixed argument shape, and
the test mocked exactly that inner call — so the asserted `['begin','reveal','open']`
array was pushed entirely by the mock's own `mockImplementation`, and the second
case returned the mock's `null`. The symbol itself stays; `full-creation-execution.ts:233`
still calls it. The ordering guard against real orchestration code lives in
`lib/worktree-creation-structured-session.test.ts`.

Provider cross-products over paths with no provider branch:
`issue-source-actions.ts:213-240` nulls every field unconditionally, its only
branches being `baseBranchNamesWorkspace`, `name === lastAutoNameRef` and
`noteRef === lastAutoNoteRef` — none provider-dependent. The github/gitlab rows and
the linear/jira block differed only in a selection label, which
`shared/new-workspace/workspace-source.test.ts:107` owns. The github-pr row stays:
it is the only one entering with a non-null `smartGitHubPrStartPointSelectionRef`,
which is the documented reason that reset exists.

Also removed: two discovery cases recombining a single key derivation
(`installed-agent-skill-discovery.ts:207`); a sustained-failure case where only one
write ever occurs, so `mockRejectedValue` and `mockRejectedValueOnce` reach
identical code; a `keeps plain labels when no endpoint is provided` replay of
`isTailscaleEndpoint(undefined)`, owned at `remote-runtime-tailscale-hint.test.ts:44`;
a 1000-host fanout case whose expectation is the output of the helper under test,
with no count-dependent branch.

Kept where the inputs differ even though the assertions repeat: the cellular
escalation trio drives three distinct rpc-client failure paths (connect timeout,
silence after upgrade, a `close()` that never fires `onclose`); the five
connection-log redaction cases map to five distinct regexes in a module with no
test of its own; the `it.each([false, true])` settlement table lands on opposite
sides of `while (dirtyHosts.delete(hostId))`; and the case deleting
`Array.prototype.toSorted` is an engine-compat ratchet, since Hermes lacks it.

* test: retire composer cases whose predicates never read the varied input

Completes the wave-11 sweep of renderer hooks. Four composer files trimmed; no
production file touched.

Each removed case varies something the production path never inspects:
- "treats a slash-containing local branch as reusable" — `resolveComposerBranchReuse`
  never looks at `/`;
- the empty-stack drop-owner case — `at(-1)` has no branch to take;
- "passes a free-form reason straight through" — `compactIpcErrorMessage` is an
  identity on that input, owned by `lib/ipc-error.test.ts`;
- "gives no reason at all when the batch failed for differing reasons" — passes no
  `commonFailure` at all, so its input shape is identical to the case above it;
- "keeps the current request pending until it settles" — asserts `await` semantics
  rather than anything production decides, so no regression can fail it.

Kept deliberately: "does give the shared reason when every path failed the same
way", because removing it would leave `attachment-drop-state.ts:188` — the upload
path's `commonFailure` wiring — with no check at all; the local-path assertion
only covers line 248.
2026-09-29 18:41:55 -07:00
Neil fed1eca486 test: stop restating internal tuning constants, keep the ones that are contracts (#23950)
Removes ~74 assertions of the form `expect(SOME_CONSTANT).toBe(<literal>)` where
the literal is an internal tuning value — a timeout, retry count, debounce
interval, cache TTL, circuit-breaker window, Tailwind class string. Those cannot
fail for any reason a user would notice: they fail only when someone deliberately
changes the number, and then the test is simply updated. They are copies of the
declaration.

The same pattern is NOT junk when the exact value is observable outside this
process, so those were deliberately kept:
- terminal byte contracts: `\r`, `\x03` ETX, Kitty escapes, `\x1b[?1;2c`;
- wire and capability values: `agent.launch.v2`, protocol 3 / min-compatible 2,
  daemon per-feature boundary versions (a daemon survives app updates, so those
  pin what an old field daemon may be trusted with), relay header tokens;
- security invariants: the `127.0.0.1` bind default, an empty iframe `sandbox`;
- values external processes read: exit code 78 (EX_CONFIG) and exit code 3
  (systemd `RestartPreventExitStatus`), `ORCA_AGENT_SESSION_SPAWN_TOKEN`,
  `npx skills …` commands users paste, on-disk journal schema versions,
  the `orca_<hash>` filename prefix the fish sweeper matches;
- third-party names: expo-router's `unstable_settings` / `ErrorBoundary`,
  iOS Safari's 16px zoom threshold.

Where a case asserted a relation rather than a literal — `A < B`, a sum of parts,
a cap compared against a sibling budget — the relation stays and only the literal
went.

Test-only changes: no production file is touched and no test file is deleted.
2026-09-29 17:06:09 -07:00
Brennan Benson 59b746ff3c feat(native-chat): one structured-chat journal database per host, owned by one process (#23613)
* feat(native-chat): one structured-chat journal database per host, owned by one process

Every structured chat on a state directory now lives in one SQLite file,
agent-session-journal.db, opened once by the process holding
agent-session-journal.owner: an empty SQLite file whose held BEGIN EXCLUSIVE is a
kernel byte-range lock, refused while another process holds it and released when
the holder dies.

- Stores own no connection: the per-chat handle, its close contract and the
  close-retry registry are gone; closing a conversation drains its writes, and the
  one connection closes last at teardown.
- The owner lock is taken at runtime start, before orca-runtime.json is written;
  a process that does not own the chats is not published and refuses every
  structured request with journalUnavailable and words that say what to do. It
  retries the lock with backoff and runs the full install once it holds it.
- A journal that will not open fails the host install: every chat says "Unable
  to load this chat." (journalCorrupt), and nothing is renamed, deleted or
  rebuilt. A newer build's database is refused and left byte-identical.
- An append is one INSERT. The listing status is a column, written after the
  rows it describes and keyed by (epoch, sequence).
- A per-chat journal from an earlier build is copied in verbatim (epoch UUID and
  every sequence) on that chat's first open, and its directory is retired only
  after the copy commits.
- auto_vacuum = INCREMENTAL, with freed pages handed back in bounded steps after
  every delete.

* perf(native-chat): key journal rows by block so one chat's rows sit together

Each chat's live epoch owns a block of row ids, block * 2^32 + seq, so a chat's
rows share leaf pages with nobody else's, a replay is one range scan, and
replacing or rewinding a chat deletes one contiguous range. Measured on the
largest real chat (61 MB) beside 19 interleaved peers: 39 ms and 7.5 MB of WAL,
against 214 ms and 102 MB for a (session_id, epoch, seq) key.

- Ids are computed in Number arithmetic, never bitwise. A sequence is refused
  outside [1, 2^32) and a block at 2^21, which keeps every id below 2^53.
- A replace, roll or import allocates a fresh block, moves the chat's pointer,
  and deletes the old block in the same transaction, so no orphan block exists.
- The listing status write moves into its own writer beside the column.

* feat(native-chat): copy a chat's per-chat journal again when an older Orca wrote it after a downgrade

A per-chat journal.db that reappears after its chat was copied in is the newer
history: an older build, run after a downgrade, attached the chat and wrote it.

- journal_imports records the (epoch, tip) each chat was copied from, in the
  copy's own transaction. A file already copied is never copied again, across
  any number of restarts after a failed rename; a file that differs always is.
- Newest writer wins, per chat, with a row saying the chat was continued in an
  older version of Orca. When both builds wrote past the recorded tip under one
  epoch, the copy takes a fresh epoch, so readers reset instead of skipping rows.
- Each copied directory retires to its own .imported-<epoch8>-<ms> name, so a
  second downgrade and re-upgrade never collides with the first.

* test(native-chat): fixture deps match the host journal database shape

Attach-flow and reconcile-attach fixtures stop passing a journal database those inputs do not take, and host and restore fixtures pass the one they now require instead of the removed journal root.

* test(native-chat): state why the runtime-state fixtures' existing casts are safe

* fix(native-chat): start up normally when this process cannot open the chat journal

A process refused the chat journal, because another Orca owns it or because its own journal will not open, failed startup restoration: the window booted in degraded no-save mode and a paired phone could not list any tabs. Startup restoration now treats the refusal structured requests are getting as having no structured host; terminals, tabs and saving go on, structured requests are still refused by the gate, and the install is retried on the next one. Any other install error fails startup as before.

* test(native-chat): name the owner-lock sweep test after the two sweeps it runs

* fix(native-chat): session history and terminal resume work while chats are refused

Session history (listing and preparing a resume) and a terminal typing a resume command only check whether a structured chat owns a provider session. In a process refused the chat journal they failed outright. They now take the refusal chats are getting as having no structured host, the same treatment startup restoration gets, through one shared helper; any other install failure still fails them. Chat requests keep the gate's refusal.

* test(native-chat): the first-work rename's fake journal saves the listing status

* fix(native-chat): open a chat whose per-chat journal file never got its schema

A crash between creating a chat's journal.db and creating its tables left an empty or schema-less file. Each chat used to open that file as an empty chat; the importer instead refused the open as "try again" forever. A file with no journal_sessions table is now read as never written, the same as one with no rows. A file that is not a database, or whose read fails, is still refused.

* fix(native-chat): let the event loop run between chats during startup restore

Opening a chat's journal is synchronous SQLite now that no per-chat directory
is created first, so the restore of every visible chat ran as one main-thread
task. Each chat now waits for a macrotask before it opens.

* fix(native-chat): import a per-chat journal in bounded batches

The one-time copy of an earlier build's per-chat journal ran as one
transaction, which blocked the main thread for 650 ms on the largest chat.
Rows now copy 512 at a time, each batch its own transaction, yielding to the
event loop between batches. The rows go into a block journal_import_blocks
reserves, which no reader follows and no other chat is allocated; the last
batch publishes the chat's pointer, repair marker and import marker together
and releases the reservation. A copy that stops midway leaves only that
block, which the next open clears and copies again. Two opens of one chat
import one after the other.

* fix(native-chat): refuse chats when the owner lock file cannot be opened

A lock file that is not a database, or cannot be opened, made the claim throw
before any refusal was recorded, so startup restoration failed on every
launch. The claim now sits in the same try as the database open and records
the same typed refusal.

* fix(native-chat): finish reclaiming pages a delete frees during a running pass

A reclaim pass ended as soon as the freelist stopped shrinking between steps,
so a second delete that freed more than one step's worth mid-pass ended it
early and left those pages on the freelist. A pass now ends only when a step
itself frees nothing, or the freelist is empty.

* test(native-chat): desktop session history is served while chats are refused

* fix(native-chat): a send to a chat holding a newer Orca's rows says to update

A chat opened read-only because a newer Orca wrote rows to it answered a send
with the generic write failure. It now refuses the way a database a newer Orca
wrote does, with the same reason and words.

* fix(native-chat): keep chat tabs while this process cannot list its chats

A process whose chats another Orca owns, or whose chat journal will not
open, has no structured host. Its session-tabs inventory still answered,
with no chat rows, and the renderer read that as "every chat was closed":
it removed the restored chat tabs and the next session save persisted
their placement away.

The inventory now says `agentSessionsUnverifiable` when the last tab
restore ran with chats on disk but no host to list them. The flag is set
and cleared at the per-client projection point beside the client-hosted
page hold, and the restore is memoised only once a host answered, so a
later lock takeover or journal open republishes the chats and clears it.
The renderer keeps agent-session tabs, and keeps cancellation tombstones,
against an inventory that does not affirm its chat set.

* fix(native-chat): say chats are open in another Orca, with this process's way past it

A process refused because another Orca owns the profile's chats sent the
generic `journalUnavailable` reason, so current desktop and phone surfaces
said "couldn't open this chat's history right now. Try again." — a step
that never helps while the other Orca runs.

The refusal now names its own reason, `journalOwnedElsewhere`, with the
refused process's kind (dev desktop, packaged, orcad) as a fact. Each kind
gets its own step: quit the other Orca, or give this one its own profile
(ORCA_DEV_USER_DATA_PATH) or data folder (ORCA_USER_DATA). The sentences
are added to the shared notice copy, the desktop catalogs in all six
locales, and the boot catalog.

A client that predates the reason reads it as none and keeps the code's
words; an unknown kind reads as the packaged app's step. The `message`
released clients print is unchanged. Which requests refuse does not change.

* fix(native-chat): restore chats on taking ownership, without a list to ask

A refused startup kept its hostless result, so after the owner quit this
process never installed a host, never reconciled restart leases, and kept
telling clients it could not list its chats until a desktop chat request.
Taking the lock now reruns startup restoration once and pushes the chats.

* test(native-chat): a navigation reply says chats are unverifiable while refused

* fix(native-chat): retry a refused owner lock at most every 5 seconds

The lock frees as its holder exits, but a refused process only learns that
on its next retry, and the 30 s cap left a second Orca refusing chats for up
to half a minute after the owner quit. One retry is an open and BEGIN
EXCLUSIVE on an empty file.

* fix(native-chat): install before deciding whether a takeover must republish chats

A list that landed on the refused startup after the lock was taken finished
after the takeover had already checked, so nothing republished. The takeover
now installs first, waits for any restore in flight, and restores only then;
the restore that clears "cannot tell" pushes the frames itself, so a list
that heals the inventory first reaches subscribers too.

* fix(native-chat): no takeover lands a host after the runtime stop

Quitting cancelled a refused claim's retry only at its end, so a retry firing
during the stop's awaits took the lock and installed a host the stop never
tore down, and the lock was then released under an open journal. The stop
now cancels the retry first, keeping the refusal, and repeats its teardown
while an install that began during it (a takeover already under way) is
pending, so no journal connection outlives the lock.

* fix(native-chat): show a thrown refusal in its own words, not its code

A refusal the host throws reaches the client as an RPC error whose message is
the bare code; its reason and facts ride only in the error's data, which no
client read. The chat pane's status line therefore printed
agent_session_journal_unreadable, a send took the bare "not sent" path, and
other writes said the outcome was unconfirmed.

One shared reader, agentSessionThrownRefusal, now reads the refusal from the
error data. A failed history read shows the refusal's read-history words, a
send keeps the refusal behind its Retry exactly as a returned refusal does, and
the other writes (desktop and phone) name the refusal instead of doubting the
outcome. The phone's read failure goes through the same reader.

* fix(native-chat): log a failed journal open once per distinct failure

Every chat request retries a journal open that failed, which is intended, but
each retry also logged the failure with its full stack: a junk database file
logged the same "file is not a database" error 189 times in a minute. The open
now logs a failure only when its code and message differ from the last one
logged, and forgets it once an open succeeds. The retry is unchanged.

The open moves to its own module beside the runtime, which had no room left.

* fix(native-chat): restore lists a chat from its per-chat file and copies it on first use

Startup restore opened every restored chat, and that open copied the chat's
per-chat file into the host database, so the first boot after an upgrade paid
the whole one-time copy before the chat list appeared.

A restore open now reads a chat that is still in its per-chat file straight
from that file, read-only, with the importer's own reader, and closes the file
before moving on. That read drives the listing, the status row and the
restart offer, as it did when every chat had its own file. The copy becomes
owed work on the chat's write queue: it runs before the chat's first write,
and a reader that reaches the chat awaits it. A chat the host already holds,
or that was copied before, still opens through the import and its reimport
rules, and so does a file whose read needs a repair written.

* fix(native-chat): no host stays registered after a stop an install spanned

Each teardown pass clears the registered host before it awaits an install in
flight, and that install registers its host when it finishes. The pass then
tore the host down but left it registered, so a request after the stop was
served by a host whose journal was closed. The stop now clears the slot once
its passes are done.

* fix(native-chat): checkpoint the journal with a full flush on macOS

synchronous = FULL fsyncs each commit, but macOS fsync leaves the drive cache
unflushed, so FULL alone does not survive a power loss there. With
checkpoint_fullfsync, each checkpoint uses F_FULLFSYNC; elsewhere it is a no-op.
The comment that said FULL alone was enough is corrected.

* fix(native-chat): delete a per-chat journal once its copy verifies

An imported chat's per-chat file was kept under an `.imported-*` name, which
doubled the disk its history takes. The copy now reads back from the host
database before it is published: its items, submissions, epoch and tip must
match the file's. Only then does one transaction publish the chat with its
import marker, and the file and its WAL files are deleted, the directory too
when nothing else is in it (a pre-SQLite transcript there is kept).

A copy that does not match is never published: the file stays, the chat is
refused as unreadable ("Unable to load this chat."), and the mismatch is logged
once. A file left behind by a failed delete or a crash matches the marker, so
the next open deletes it rather than copying it again; a file an older build
wrote after a downgrade still differs, and is still copied again.

* fix(native-chat): verify an imported chat a batch at a time

The check that a copied chat reads back as its per-chat file folded both whole,
each in one synchronous task: over half a second on the largest chat. Both
reads now go a batch at a time between turns of the event loop, like the copy
itself, and count rows as well, so a copy that lost a row with no item in it
is caught too.

* fix(native-chat): restore reads a chat's per-chat file a batch at a time

Restore folded a chat still in its per-chat file in one task, so the largest
chat's file held the main thread for about half a second at startup. The fold
now takes the file a batch of rows per turn of the event loop, into the same
fold a replay uses, and nothing reads it before it is done. The file is still
closed before restore moves on.

* fix(native-chat): end a per-chat copy on a turn of its own

A chat's first open ran the copy's last steps (the verified publish and the
per-chat file delete) and the replay of what was copied in one task. The copy
now yields before it returns, so the replay, which every open runs, is a task
of its own.

* fix(native-chat): commit a per-chat copy's batches without an fsync each

Each 512-row batch of a chat's first-use copy committed under synchronous =
FULL, so a large chat paid one fsync per batch, about a quarter of its first
open. The batches now commit under NORMAL, set and restored in the batch's own
task so no other chat's commit runs under it. The publish that makes the copy
visible still commits under FULL, and under WAL that sync makes every earlier
batch durable with it. A crash before it leaves only the unpublished block,
which the next open clears and copies again.

* fix(native-chat): roll back a chat journal transaction whose COMMIT fails

The shared connection's transaction rolled back only when its body threw. A
COMMIT that failed left the transaction open, so every later write, for any
chat, failed with "cannot start a transaction within a transaction", and reads
saw rows that never committed. Under the unsynced copy the failure also tried
to restore the sync level inside the open transaction, which SQLite refuses,
so the caller got that error instead of the COMMIT's.

One transaction helper now covers the body and the COMMIT, rolls back whatever
transaction survives, and rethrows the original error. Schema creation uses it
too. If that ROLLBACK fails as well, the connection is marked stranded: each
later use retries the ROLLBACK, and until one goes through every chat gets the
same "history unavailable, try again" refusal a journal that will not open
gives. The rollback that frees it also restores the FULL sync level.

* fix(native-chat): keep the chat journal connection until its close succeeds

Closing the journal dropped its connection handle before closing it. A close
that failed left the database reporting itself closed with the connection still
open, so the stop that retried the teardown found nothing to close and released
the owner lock over a live connection.

The handle is now dropped only once the close succeeds. A failed close keeps
the runtime pending and the lock held, and the next stop closes that same
connection before it releases the lock.

* fix(native-chat): publish the runtime only once it holds the chat journal lock

When this process could not open the owner lock file at all (a permission
error, or a file that is not a database), the runtime counted that as owning
the chats and wrote orca-runtime.json. That overwrote the real owner's entry,
so the CLI was sent to a process that cannot serve its chats.

A claim that throws is now refused like one another process holds: the runtime
starts but does not publish, the claim's existing retry keeps asking for the
lock, and discovery publishes once the retry takes it. Chats still get the
refusal for the failure itself, and startup restoration reruns on the takeover
the same way it does after another owner quits. A sole process whose lock file
never opens is not found by the CLI until it does.

* fix(native-chat): keep a chat's history when an older build started it over

The first copy deletes a chat's per-chat file, so an older build run after a
downgrade finds no file and starts the chat from nothing. On the re-upgrade that
fresh file was copied in as the newer history, replacing everything the shared
database held for the chat, and then deleted.

A file whose epoch is not the one last copied and that opens with
`session_created` is now kept: neither copied nor deleted, and the chat keeps
the history it has. A file that carried the copied epoch on is still copied
again, as before.

* test(native-chat): pin which chats startup restore copies

Restore copies a chat still in its per-chat file only when restore itself has
to write to it: settling what the last run left open, here a running tool call
or a send handed over and never answered. Every other restored chat stays in
its file until its first use.

* test(native-chat): pin the copy wait on a read that opens a chat restore opened

A read queued behind restore's open of the same chat reaches the conversation
through its own open rather than the listing. It must still wait for the
owed copy, or it reads the chat before its history is in the one database.

* fix(native-chat): record a set-aside per-chat file so no later open reads it

Setting aside a file an older build started over is decided once and kept in
the new `journal_set_aside` table (schema 2, additive), with the file's epoch
and tip as they were. Every later open of the chat skips the file without
opening it, across restarts and after the older build writes more to it:
anything written there grows from that build's own start, never from this
build's history.

The best-effort delete moves beside the per-chat file reader.

* fix(native-chat): set aside any per-chat file at an epoch this build never copied

A chat's per-chat file is deleted once its copy verifies, so a file that
reappears at another epoch was never this build's history, whatever its first
row says: an older build started the chat over, possibly rewinding it after
(`handle_forked`), or rolled the epoch of a file whose delete had failed.
Copying any of them would replace everything the chat holds, so each is set
aside. Only a file still at the copied epoch is copied again (it grew) or
deleted (it did not). The first-row check is gone.

* fix(native-chat): copy a reappearing per-chat file again only while this build has not written past the copy

A per-chat file that an older build carried on under the copied epoch was
copied again even when this build had also written to the chat since the
copy, or had rolled its epoch. The second copy replaced the chat's block,
so what was sent in this build after the copy was gone for good.

Now the file is copied again only when the chat still stands exactly as it
was copied: the same epoch and tip the import marker recorded. Otherwise it
is set aside like any other file that is not this build's history, left on
disk untouched and recorded so no later open reads it. A second copy
therefore never replaces rows this build wrote, keeps the file's own epoch,
and the fresh-epoch rewrite goes away. The row it adds now says the history
includes what the older version recorded, not that anything was replaced.

* test(native-chat): pin that a chat founded here keeps its history, and the v1 schema upgrade

A chat this build founded has a pointer and no import marker, so a per-chat
file an older build later starts for it is set aside. Nothing pinned that
half of the rule: letting such a chat be copied again replaced its history
and every test still passed. A second test pins that a database written at
schema version 1 upgrades in place, gaining the set-aside table and keeping
its import markers.

* test(native-chat): drop a lost copied row by patching the source, not wrapping it

* chore(mobile): restore the mobile lockfile to main's

* fix(native-chat): pass a classified journal refusal through a send or Stop unchanged

* fix(native-chat): refuse a read whose owed copy fails as a failed open does

* test(native-chat): measure only the replace's WAL in the block-key case

Opening the chats starts a free-page pass that waits one event-loop turn,
and the seed never yields one, so that pass was still pending when the
replace committed. It woke during the async stat and reclaimed the pages
the replace freed, adding ~500 KB of WAL whenever the stat lost the race
(Linux CI). Drain that pass before measuring and stub the replace's own.

* test(native-chat): the RPC fixture's status journal can save its listing status

The status feed now hands every projection to the journal, which decides whether it is worth saving.

* test(native-chat): state why the RPC fixture's status journal cast is safe

* fix(native-chat): refuse a per-chat copy whose rows differ from the file, not only its counts

* fix(bench): build the replay benchmark's baseline arm from the base tree and release its handles on failure

* fix(native-chat): retry a failed listing status save on the next read of a cached status

* refactor(native-chat): drop the chat journal owner lock; the process instance lock already guards the profile

The journal carried its own exclusive lock, with a retry loop, an in-process
takeover, lock-gated runtime discovery and a "chats are open in another Orca"
refusal. Every shipped process kind (packaged desktop, serve mode, orcad)
already refuses a second instance on one profile before the journal opens, so
the lock only ever mattered for dev desktops, which the next commit covers at
the process level instead.

The host now opens its one journal connection at install with no lock. What a
sole process whose journal will not open needs stays: the install refusal
recorded for the gate, the no-host startup path, and the unverifiable chat
inventory, now in structured-agent-session-host-refusal.ts. The unreleased
journalOwnedElsewhere reason, its processKind fact and their copy are removed.

* fix(startup): dev desktops take the single-instance lock, and a second one says why it quit

Dev skipped Electron's single-instance lock so parallel `pnpm dev` runs from
several worktrees would not quit silently, but two dev processes on the
default orca-dev profile then write the same stores at once. Dev now takes
the lock like packaged builds: a second launch on the same profile focuses
the first window and exits with code 3, printing one stderr line that names
the taken profile and how to run another copy (ORCA_DEV_USER_DATA_PATH).

Serve mode, the macOS diagnostic bypass and the E2E harness are unchanged:
an E2E launch still skips the lock unless it sets
ORCA_E2E_ENFORCE_SINGLE_INSTANCE_LOCK=1.

* refactor(native-chat): key journal rows by chat, epoch and sequence

Rows in the host's journal database are now addressed by the chat's own
identity, with `(session_id, epoch, seq)` as the primary key, the same
shape each per-chat file already used. The block-keyed layout goes with
everything built on it: the block column and its allocator, the 2^21
block ceiling, and the import's reserved block table.

A first-use copy writes its rows under the file's epoch, which the chat's
pointer does not name until the verified copy publishes it, so no reader
sees a half-copied chat. A try that stopped midway leaves only rows no
pointer names, and the next try deletes them before it copies again.
Replace, rollover and repair delete by (chat, epoch).

This build's history always wins: once a chat was copied or founded here,
any per-chat file that reappears is set aside, and the same-epoch copy
again after a downgrade is removed.

The bounded free-page reclaim after every delete is dropped;
`auto_vacuum = INCREMENTAL` stays at file creation, so a later periodic
reclaim can still be added. Session search keeps its own step.

The schema moves to version 3. Versions 1 and 2 were written only by
unreleased builds of this change and are refused as found, not migrated.

* fix(native-chat): open a chat journal a newer Orca wrote read-only instead of refusing it

After a downgrade, the host's journal database carries a newer user_version. It was refused
outright, so every chat's history disappeared. It now opens on a read-only connection, as the
per-chat journals did: each chat shows what this build can read, from the database or a per-chat
file never copied in, and every write is refused with "Chats were saved by a newer Orca. Update
Orca to keep using them." Nothing is written, copied, repaired or founded, and the file stays
byte-identical. A table the newer schema changed reads as the same read-only refusal, not damage.

* refactor(native-chat): leave the saved listing status to the change that reads it

Nothing in this change reads the per-chat listing status column: it was a stored copy of a fact
the status feed derives, written after every turn end and cleared on every epoch change. The
status_json / status_seq columns, their writer, the saved-status type, the status feed's save and
its retry on a cached projection all go, with their tests. The change that lists chats from a
saved status adds the column back beside its reader.

* fix(native-chat): a chat saved by a newer Orca says to update Orca, not to try again

When a newer Orca wrote the chat journal, this build opens it read-only. A send or a Stop was
refused with the reason `journalUnavailable`, so today's desktop and phone clients chose the
words for an open that can clear: "Orca couldn't open this chat's history right now. Try again."
Retrying never cleared it; only updating Orca does.

The refusal now names its own reason, `journalWrittenByNewerOrca`, whose words are "Chats were
saved by a newer Orca. Update Orca to keep using them." A read refused the same way names it
too. An older client does not know the reason, drops it, and falls back to the code's words
("Orca couldn't read this chat's saved history."), and released clients still print the message.

* fix(native-chat): a chat journal from an unreleased build reads as unusable, not as retryable

A chat journal database stamped with schema 1 or 2 was written only by unreleased development
builds of this change. Opening it threw a plain error, which every chat reported as "Orca couldn't
open this chat's history right now. Try again." Retrying never cleared it.

It now throws a named error that is classified as unusable, so every chat says "Unable to load
this chat." The one log line names the file, says an unreleased development build wrote it, and
says to move it aside. Nothing migrates or renames it.

* docs(native-chat): drop the second-Orca-owns-the-chats case from three comments

The chat-only owner lock is gone, so only a chat journal that will not open leaves a runtime
unable to list its chats.

* docs(native-chat): correct three chat-journal comments the redesign left behind

A per-chat file left without its WAL is set aside, not copied again; nothing runs an incremental
vacuum yet, so the auto_vacuum mode is kept for a later pass; and the idle sweep drops a chat's
in-memory fold, since a chat holds no journal connection.

* refactor(native-chat): stop exporting chat-journal names nothing imports

Each is used only inside its own module now; the teardown's export served a deleted test.

* test(native-chat): name the version-0 test for what it covers, and check every journal table

The test named 'migrates an older user_version forward' covers only a version-0 file that already
has its tables; versions 1 and 2 are refused. The table test now also checks journal_imports and
journal_set_aside.

* fix(startup): a second dev launch's exit line no longer claims it focused a window

The running dev instance may be a background launch or a server, which show no window. The line
now says only that this launch passed its request to that instance.

* fix(native-chat): a failed structured-chat install closes the journal connection it opened

The install opened the chat journal database and closed it only if the record store then failed
to open. A later failure, such as the model catalog wiring or the host constructor, left the
connection open, and the next install opened a second one in the same process. Every failure
after the open now closes it.
2026-09-29 16:42:29 -07:00
Neil 6194a7a1b6 test: drop private-internal and boundary-census tests with behavioral owners (#23941)
Fourth audit wave, cut short by a session restart, so this lands the verified
subset rather than the full batch.

Removes private-predicate cases whose behavior is already covered through the
module's real entry point, and de-exports the seams they reached for. Also drops
three whole files whose every case was a duplicate or a call-shape grep.

The source-grep vein is close to exhausted. One auditor reviewed 15 remaining
flagged files and deleted nothing: what is left is mostly legitimate
architectural ratchets that no type checker and no behavioral test can reach —
AST fences banning `as`/`any` in an RPC operation region, discovered-vs-listed
set equality over subscription sites, count ceilings on unchecked reply readers,
and assertions on generated WebView bundles (no CDN URL, no `</script`
tokenizer escape, parses at the Chrome 74 floor). Those stay.
2026-09-29 16:35:49 -07:00
Brennan Benson 2a001b43e5 fix(mobile): stop a closing phone stream from ending the terminal or chat feed that replaced it (#22939)
* fix(mobile): don't unsubscribe a terminal stream that already ended

* fix(mobile): keep the stream registry under the line cap and expect no error after a streamed end

The streamed `end` now closes a direct stream, so a later reply for that id is unrouted instead of
reaching the listener as an error; the subscription recording test expected the old error. Inline
the error wrapper so the registry fits the 300-line cap, and shorten the comments.

* fix(mobile): reopen a native chat stream the host ended while the screen still shows it

* fix(mobile): only the focused screen takes back a host-ended chat feed, and a reopen keeps loaded history

* fix(mobile): give each native chat subscription its own token, and reopen with the first page

* test(mobile): pin the reopened chat subscribe params without a cast

* fix(mobile): show a host-ended chat feed as an error instead of reopening it

With a token per subscription, the desktop ends a phone chat feed only when the
phone closes it or the socket drops: no client sends the connection-wide chat
unsubscribe, a socket close delivers no end, and every desktop with mobile chat
keys feeds by the client's token. The reopen-with-backoff layer and the
history-keeping reopen merge guarded a trigger that no longer exists, so they
are removed. An end that still reaches the screen settles the feed as an error,
so a dead feed is never shown as live; leaving and re-entering the chat
subscribes again with a fresh token.

* test(mobile): re-record the RPC goldens for the per-subscription chat token

Each nativeChat.subscribe now sends its own token, so the native chat scenarios
expect `claude:session-1:<id>` in the subscribe params, and the four native chat
recordings (native-chat-page-earlier and the session.native-chat-page matrix)
record that token in their subscribe and unsubscribe payloads. Every other
golden moved only its `baseline` header, repinned to the commit recorded from.

* test(mobile): repin RPC recordings to the merge with main and re-record

Merging main moved the product tree the recordings are pinned to, so the
baseline is repinned to the merge commit and the whole corpus re-recorded.
Only the four native chat recordings change content, because each
nativeChat.subscribe now carries a per-subscription token; every other
recording moves only its baseline header.

* test(mobile): repin RPC recordings to the merge with main and re-record

Merging main brought #22762's recordings, #23757's repin and #23720's recorder change, so the
corpus is repinned to the merge commit, the last commit to touch a recorded path, and re-recorded
whole. Against main, 786 recordings move only their baseline header and 4 native chat recordings
change content: their nativeChat.subscribe and nativeChat.unsubscribe carry the per-subscription
token instead of claude:session-1, which also moves their scenarioSha256.
2026-09-29 16:31:47 -07:00
Jinwoo Hong 2d31941286 fix(mobile): the update-mobile wall opens the exact release (#23789)
* fix(mobile): the update-mobile wall opens the exact release

When the in-app checker knows the newest release, the update-mobile wall
offers "Get Orca <version>" and opens that release's page instead of the
generic releases list or a hand-built App Store link. Mounting the wall
runs the single-flight, bounded checker once. Dismissal is ignored here.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): keep the update checker out of the page closure

The wall now takes the offered release as a prop and opens it through
the external-link seam. A shell-only hook runs the checker once per
update-mobile wall and supplies the release; MobileWebShellScreen, which
no page route reaches, wires it. HostProtocolGate is in every page
closure, so it stays unwired pending a ruling.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): platform-split the wall's release offer and throttle its check

The hook is the one module whose behaviour differs by platform: the native
file runs the checker, the .web.ts sibling offers nothing, so both
HostProtocolGate and MobileWebShellScreen wire it the same way and the
page closure never carries the checker. A remount checks only when no
release is known and the last check is absent or over the retry interval
old, since each check is an unauthenticated GitHub call.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): let the checker decide whether the wall's check is due

The hook's own throttle read lastCheckedAt, which moves only on success,
so after a failed lookup every wall remount looked up again, and before
preferences loaded a cold-start wall looked up despite a fresh stored
result. The checker already owns the cadence: checkIfDue waits for the
stored state and checks only past the retry or daily interval.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): the wall reads the known release and triggers nothing

The started checker's own timer, cold-start and foreground paths already
run every due check, so a check requested by the wall could never be due.
The hook now only returns the known release; the checker is back to main.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): the wall reads its release through useWallAppUpdate

The .web.ts sibling alone keeps the page closure clean, so the screen
calls the hook itself and the optional prop, the gate wiring and the
shell wiring go. HostProtocolGate and MobileWebShellScreen are back to
their base content.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-29 19:02:40 -04:00
Jinwoo Hong c634432acd fix(mobile): count opening a cached page generation as use; cache six hosts (#23780)
* fix(mobile): count opening a cached page generation as use; cache six hosts

Eviction order was written only when a download committed, so a host opened
daily but downloaded long ago was evicted first and each revisit cost a full
redownload. An open now refreshes that host's recency inside the store queue.
The ceiling rises to six hosts and the update-failure cap to six hosts' worth.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): stamp page cache recency by order and never clobber it on open

An open stamped wall time, so a backward clock jump made the host in use the
oldest entry and the next download evicted it. Recency now stamps past the
newest entry. An unreadable or torn index read as empty, so every open rewrote
it with one host and demoted the rest; opens now leave it for a commit.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* style(mobile): wrap the page cache index doc comment to 100 columns

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): skip no-op page cache recency writes and stage the index write

Every open rewrote hosts.json even when the host was already the newest, so a
single-host user paid a native write per open. Opens now write only when the
order changes. The index is written to a sibling and renamed over the old one,
so an interrupted write leaves the previous index or none, never a torn file.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): keep page cache recency as an ordered host list

Eviction needs order only, so hosts.json is now the cache keys least recently
used first. That drops the clock, the monotonic stamp, the unreadable-versus-
absent split and the staged write: a torn or old-format file reads as empty,
the same state a missing one gives. Opens and same-build commits share one
touch that writes only when the host is not already last.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-29 19:02:16 -04:00
Neil 70475e0228 test: stop testing private internals through exports no caller needs (#23829)
Third audit wave. The detector looked for production modules exporting three
or more symbols that no production file imports — only tests do. That shape is
the authoring gate's fourth question failing: a test needing a production seam
no caller needs belongs at the real boundary instead.

Most hits were detector false positives and were left alone; the scanner misses
re-export barrels and dynamic imports, so every module was re-verified with rg
before any edit. Where a private predicate's behavior was already covered
through the module's real entry point, the duplicate cases are gone and the
symbol is module-private again. Where it was NOT covered anywhere else, the test
stays — this audit removes tests, it does not author replacements.

Production code deleted where tests were its only callers: the superseded
`filesystem-directory-listing-limit` module, the unused
`format{Hourly,Daily,Adhoc}Version` helpers and their orphaned prerelease
identifiers, the dead `filterByAutomationListSearch*` family superseded by
`matchAutomationListSearchRowKeys`, and the dead
`getAiVaultResumeWorktreeTargetStatus` copy of the live workspace branch.

Also drops two call-shape source greps in `relay-sweep-schedule.test.ts` that
asserted `index.ts` spells `jitteredSweepIntervalMs(30_000)`; the jitter math
has a behavioral owner at the top of the same file. The structural census that
counts role-gated vs total `setInterval(` calls stays — an ungated sweep runs in
every cell, and nothing else can catch that.
2026-09-29 15:53:18 -07:00
Brennan BensonandClaude 29c49aec31 feat(native-chat): hold mid-turn messages in a host-owned queue (#23726)
* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* fix(native-chat): name a chat write by its target, not the owner generation

A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.

Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.

Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.

* fix(native-chat): every journal append reaches the chats that are open

A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.

A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.

* test(native-chat): an epoch replacement reaches the open chat

* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

* perf(native-chat): a publish behind a delivered commit reads nothing

Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.

* test(native-chat): state why the teardown test's fake journal is safe to cast

* docs(native-chat): say mutation admission checks only the writer lease

* docs(native-chat): drop the send rebase from comments that still described it

* fix(native-chat): a message is accepted, then delivered

A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".

A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.

Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.

A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.

Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.

* fix(native-chat): settle queued messages only for the child that ended

A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.

A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.

The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.

* fix(native-chat): an adoption that fails to import keeps the conversation open

The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.

* perf(native-chat): the recovering open reads the journal once

Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.

* fix(native-chat): an attach that fails after indexing its child leaves no child behind

A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.

* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer

The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.

A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.

* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down

The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.

* fix(native-chat): a message rejected while its chat was closed reads as not sent

A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.

* test(orchestration): name why the readiness settlement fakes are cast

* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent

* docs(native-chat): drop the fence from the admission the send effects run behind

* docs(native-chat): give the fence move on release the reason that still holds

* docs(native-chat): stop citing a write fence check in launch and mailbox comments

Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.

* refactor(native-chat): the provider child is its own record

A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.

- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
  patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
  own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
  the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
  dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
  is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
  the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.

* fix(native-chat): the delivery loop alone settles a message its start or child failed

A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.

- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
  reads how it ended: a Stop continues; anything else writes one failure row and rejects every
  queued message with the same words, then stops. A child still starting whose start the adapter
  says did not land fails the same way. The exit, eviction and the settlement retry only settle
  the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
  nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
  failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
  closed, with or without a child, and a start the loop already has in flight is waited for so the
  child it produces is stopped rather than left behind.

* refactor(native-chat): a stopped child ends on the one reading of its stop

The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.

* feat(native-chat): the host says it accepts a send before any agent has it

The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.

* refactor(native-chat): an attach never opens a journal of its own

The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.

* fix(native-chat): a moved fence resends nothing on a host that accepts first

The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.

The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.

* refactor(native-chat): a child's end says whether the user or the host stopped it

The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.

* fix(native-chat): a chat whose only work is a queued message is not offered for resume

A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.

* test(native-chat): type the queued-message fixtures in the resume-offer tests

* fix(native-chat): a start that dies while a message waits on it is that message's failed start

Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.

* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now

* test(native-chat): pin what a failed start settles, and what a resume offer names

A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.

* test(native-chat): the failed-start pins fail on what the message became, not on a timeout

* test(orchestration): the preamble's host stub is typed, not cast

The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.

* fix(native-chat): a Stop that names no turn stops what the conversation has in flight

Between handing a message to the agent and the agent opening its turn, there is no turn id a
client could name, so a Stop in that gap was refused as "already finished" while the agent went
on to answer. A cancel's turn id is now an optional precondition instead of its target: with
none, the host withdraws what is queued and, when the journal still reads working, asks the
adapter to stop whatever the child has in flight. Claude's interrupt is session-scoped, so it
is guarded by fence and acquisition generation rather than a turn identity. Codex interrupts
the turn its latest turn/start answered with until the journal shows one.

A cancel that names its turn behaves exactly as before.

* fix(native-chat): Stop is there from the moment a message is sent

The composer showed Stop only once the agent had opened a turn, so for the second or two after a
send the chat read "thinking" with no way to stop it. Against a host that takes a Stop naming no
turn, Stop now shows whenever the chat reads working (a turn, a queued message, or a handed-over
one still unanswered) or this client still has a message on its way. Pressing it, or Escape,
first drops every outbox entry the journal does not hold yet, so nothing goes out after the
Stop, then sends the conversation-wide cancel. A send already on its way reaches the host ahead
of the cancel, which withdraws it there. Against an older host Stop still needs a running turn.

The unconfirmed-send probe moves into its own hook so the outbox hook stays in budget.

* fix(native-chat): Stop before a turn is gated on its own host capability

A host that accepts sends first (agent-session.accepted-send.v1) can still predate the cancel
that names no turn and would refuse it as invalid, since clients and hosts ship independently.
Hosts that take that cancel now advertise agent-session.conversation-stop.v1, and the renderer
shows Stop before a turn opens, and sends the no-turn cancel, only to a host advertising it.
Every other host keeps a Stop that needs, and names, a running turn.

The host capability probe the accepted-send hook used is generalized so both read one path.

* test(native-chat): a build advertises conversation stop exactly where its cancel may name no turn

* fix(native-chat): a view never restarts a chat whose last start failed

A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.

* test(native-chat): start the child the loop waits on with an attach, not a second view

A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.

* fix(native-chat): settle a gone generation's turn wherever a conversation opens

A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.

* test(native-chat): prove the next child's start settles the turn an earlier child left

The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.

* test(native-chat): count a failed start's rows by row, not by text

Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.

* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget

* test(native-chat): pin the open's and the send's start and row counts, however the view binds

Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.

* fix(native-chat): settle a gone generation's turn at every open but an acquisition's

The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.

* test(native-chat): hold the create's start open until the views bind

The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.

* fix(native-chat): Stop reads the one working rule every session list reads

While Claude retries a rate-limited request it never echoes the message, so no
turn opens: the sidebar read Working from the unanswered send while the composer
showed Send. The chat's working state, the host's session-list status and the
host's no-turn Stop check now call one shared rule instead of three copies.

* test(native-chat): a rate-limit retry pins only that no turn opens, not how its rows are kept

* fix(native-chat): Stop leaves a message waiting on its Retry, and does not show for one

A send that failed holds the queue until the user retries it, and one the host restarted under is
parked the same way. Stop counted both as still on their way, so it showed in an idle chat and
could never go away, and pressing it dropped the failed message along with its Retry.

* test(native-chat): the chat's Stop and a session list read the main agent alike over their own copies

The chat reduces its stream and a list reads the status feed. Driven through the real host for a
rate-limit retry with no turn, a subagent still running after the main turn, and the handed-over
child exiting.

* refactor(mobile): the chat reads the main agent's working state through the shared rule

Behaviour is unchanged: the same two terms, now from the one function the host projection and the
desktop chat read.

* fix(codex): a Stop naming no turn never interrupts an earlier turn

It fell back to the id an earlier turn/start answered with when the latest start went unanswered,
or when the journal showed a compaction Codex had not started, and reported that as stopped.

* fix(native-chat): a Stop naming no turn never says a turn had already finished

When the provider found nothing left to stop, for instance a turn that ended between the host's
check and the interrupt, the chat got "The provider had already finished this turn." for a turn
the Stop never named. It now ends quietly, as a Stop with nothing in flight does.

* fix(native-chat): one Stop the host could not settle no longer refuses every later one

A Stop naming no turn has one operation key per session. When the host could not settle one, it
answered every later Stop under the same id as unknown until the id expired. Once the host says
so, the next press is a new Stop; transport doubt still replays the same id.

* refactor(native-chat): drop the composer's second error formatter

After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.

* test(native-chat): pin the reason on a message rejected while its chat was closed

The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.

* test(native-chat): read Stop operation ids without a cast

* fix(native-chat): a Stop whose answer was lost no longer swallows the next one

A Stop that names no turn has one operation key per chat. When its answer was lost in transit, the
chat kept the id, so every later Stop replayed it; the host answers a replay as already handled, so
for up to a day Stop stopped nothing. The id is now dropped once the call settles, however it
settles. A second press while the first is still on its way still shares its id.

* refactor(native-chat): a Stop naming no target keeps its operation id only for its own call

The chat kept each write's operation id per payload across calls, and dropped it only on some
settle paths. That is right for a write naming what it acts on, but a Stop naming no turn, and a
stop of every background task, share one payload with every later one, so any path that kept the id
made the next Stop replay as already handled and stop nothing. One path was still open: an answer
that arrived after the chat moved to a new fence.

Whether a write names its target is now decided once, before its id is picked. One that names none
keeps its id only while its call is in flight, so a press made meanwhile joins it, and releases it
when the call settles, however it settles. The release runs only while the key still holds that
call's id, so a joined call settling late cannot drop a newer one's. This replaces the per-path
exceptions for a thrown call.

* test(native-chat): read the Stop fences without a cast

* test(native-chat): pin the new id for a named cancel the host could not settle

After the Stop naming no turn moved to a per-call id, the only test of the unknown-refusal release
was gone, and the half that stays, for a cancel naming its turn, could be removed with every test
green.

* fix(native-chat): a Stop pressed after a new message stops it, even while the last Stop is unanswered

A Stop naming no turn shared its operation id with any press made while it was still in flight. The
host runs a chat's writes in order, so a message sent between two presses was accepted after the
first Stop ran, and the second press replayed that Stop as already handled and left the message
running, although the chat had already withdrawn it from the outbox.

A write naming no target now gets a new id on every press and is never kept, so each Stop acts on
whatever is running when the host reaches it. A write naming its target keeps its id exactly as
before. A double press can ask the provider to stop the same turn twice, which it tolerates.

* fix(native-chat): Stop no longer blinks off as Claude opens the turn for a message

Claude's echo of a sent message both answers the send and opens its turn. The echo settled the send
first, so the host published the message as answered one frame before the turn it opened, and for
that frame the chat read nothing running: Stop turned back into Send, and Working blinked off in
every session list, for tens of milliseconds on each turn.

The echo now settles the send after the turn it opens has been emitted, so the running turn is
published first.

* fix(native-chat): a message a Stop withdrew comes back to its sender's composer

A Stop withdraws every message the host holds but has not run, and S also
drops the ones this client had not handed over yet. Either way the message
left the chat and its text survived only in a hidden journal row and the
in-memory ArrowUp history.

The sending client now puts the withdrawn text and images back in that
pane's composer, after whatever is typed there. Withdrawn is read from the
rejection reason through one shared check, which the outbox reconcile now
uses too. The composer is written before the entry leaves storage, so a
failure between the two repeats the text instead of losing it, and an entry
storage no longer holds is never given back again, so a replay, a second
view or a remount restores it once. Only this client's outbox holds the
entry, so other viewers still see the message disappear. A failed Stop
withdraws nothing on the host and gives nothing back.

* fix(native-chat): withdrawn text put back during an IME composition is not lost

While the IME owns the field, the composer ignores a programmatic draft, and
the next composed keystroke wrote the draft without the restored text, after
its outbox entry had already been dropped. The composer now holds text
appended mid-composition, keeps it in the cache after each composed write,
and shows it once the composition settles, the way attachments that land
mid-composition already wait for it.

* test(native-chat): pin that only a withdrawn message comes back to the composer

* test(native-chat): set up the composer's window API for every describe in the composition-race file

* docs(native-chat): note that the withdrawn check reads the legacy reason until a typed category lands

* test(native-chat): pin that text put back mid-composition shows once, even beside a mid-composition clear

* feat(native-chat): host-owned queued-message draft store in the session journal

A queued mid-turn message is a draft row in the session's journal.db,
created idempotently at every writable open with no user_version bump so a
downgrade stays writable. Consume converts one draft into an ordinary
submission inside the journal writer's own transaction (exactly-once), and
a standing writer hook returns a consumed draft only when a committed row
newly settles its current consumed submission to a non-withdrawn rejection
— the same decision the reducer folds rows through. Open-time repair
re-derives returned state behind the stored fact; retention never prunes a
row whose refusal could still return it.

* feat(native-chat): queued-messages wire contract, dark capability, and send classifiers

The send result becomes a union: today's submission arm unchanged, plus a
capability-gated queued arm only clients that sent delivery:'queue-if-active'
ever receive. Whole-list queuedMessages fields ride the subscribe events and
history pages; Stop gains withdrawQueued with the withdrawn bodies in its
result; clear's result carries withdrawn drafts too. Both classifiers treat
queued as accepted/spent. agent-session.queued-messages.v1 is defined but
deliberately NOT advertised: the rollout prerequisites (Claude fold receipt,
integrated Codex steer matrix) are not in this host.

* feat(native-chat): queue a capable mid-turn send as a draft, drain it at turn end, and let Stop and clear return its text

A send carrying delivery:'queue-if-active' while the session owes work — or
behind an actionable backlog — becomes a host-held draft instead of a
submission. A serialized drain woken by journal commits, draft mutations and
conversation opens re-derives its gates from live facts (streamed-event
barrier first, backlog never a gate) and converts the oldest actionable
draft through the exactly-once consume; from that instant today's delivery
pipeline runs unchanged. Stop pauses the withdrawable frontier at the stop
step (a process-level pause set that survives handle eviction and, via the
per-process host instance, restarts), then withdraws it with the text in the
result for capable clients; /clear does the same for the superseded source.
The draft list publishes whole per emit with identity dedup, rides only the
final catch-up page, and attaches to history pages. queuedMessageSend
overrides queue policy only; queuedMessageDelete hands the body back.
Replays for all of it answer from op-stamped tombstone receipts.

* test(native-chat): pin mid-turn queueing against the real host

Accept (working/backlog/text-only/budget/replay), the one-per-settle drain,
returned cards with N1 overtake and the N4 re-send loop, Stop withdraw with
tombstone replays, the process-level pause across evict/reopen, Delete
receipts, /clear returning the withdrawn text, and publication (hydration,
unchanged-cursor insert, same-frame consume, identity dedup).

* test(native-chat): read the queued receipt ids before the wait closures

* chore(native-chat): SAFETY rationales on the sqlite row casts and a cast-free mobile narrowing

* fix(native-chat): queued-draft bookkeeping never costs a publish, an open, a clear or a history read

- Cache the draft list per draft-table revision. The drain re-checks on every
  journal publish, so each streamed delta was running a SELECT and parsing
  every draft body the handle had ever written (tombstones included).
- Open-time repair/prune failures are reported and skipped; they no longer
  fail opening the chat.
- /clear on a source with no drafts answers exactly as before: no empty
  `withdrawnQueued`, no empty write transaction, no extra publish. A draft read
  failure after the committed clear no longer turns it into a refusal.
- History pages read drafts through the same guarded reader as subscribers.
- Publication moves to its own module; the held-draft rule lives with the
  pause state; one pending-prompt check; drop an export nothing calls.
- Tests: restart-held drafts, pre-consume failure pause + Send retry, failed
  open repair, clear with no drafts.

* fix(native-chat): a Stop that withdraws a consumed draft's send gives its text back

A queued draft converted into a submission leaves the sender's outbox, so when
a Stop withdrew that submission before the agent received it, the text had no
holder: the draft stayed `dispatched` forever and nothing restored it.

- The returned-card rule now follows every effective `rejected` settlement of
  a consumed draft's submission, a Stop's withdrawal included, with the
  withdrawal reason stored as the fact (`dispatchWasWithdrawn`). The writer
  hook and the open-time repair share the rule, so no rejected submission can
  leave its draft `dispatched`.
- A capable Stop withdraws the cards it returned itself along with its
  frontier, stamped with its caller-scoped key: the text comes back once in
  `withdrawnQueued` and replays from the tombstone. An old client's Stop
  leaves a returned card.
- Stop's draft steps move to structured-agent-session-queued-stop.ts.
- Tests: Stop between consume and the agent's receipt for both client kinds,
  its replay, a crash after the withdrawal, restart in the window, and the
  repair of a hookless withdrawal.

* perf(native-chat): the queued-draft drain takes no serialized step while the agent works

The drain was woken by every journal publish and, with a draft waiting, queued
a serialized step (streamed-event flush included) per publish, only to find the
session still working. During a streamed turn that is one step per delta,
contending with Stop and every other mutation for the session's queue.

The pre-check now also skips while the session is working. Whatever ends the
work is itself a commit that schedules again, and the step still re-reads every
gate after its flush, so no wake is lost.

- Test: queued sends during a turn take no drain step; settling the turn drains.

* fix(native-chat): a clear withdraws queued text only for a caller that can take it back; paused reasons are markers

An older client running /clear had its source's waiting and returned drafts
withdrawn and their text returned in a `withdrawnQueued` field it does not
read, so the text was lost. Clear now mirrors Stop: `withdrawQueued: true` on
`agentSession.conversationCommand` (strict params, sent only when the
queued-messages capability is advertised) withdraws the drafts and returns
their text once, replaying from the tombstones. Without it the source keeps
its cards: the supersession fence already blocks the drain, and Delete still
hands the text back.

A paused card's reason was host-authored English on the wire. It is now a
typed marker (`send_failed`) the client localizes, like `returnedReason`; a
client treats an unknown marker as a plain pause.

- Tests: an old client's clear leaves the cards and its replay stays
  field-free, then Delete returns the text; a capable clear returns the text
  once and replays it; the paused marker.

* fix(native-chat): a draft pause that commits no journal row still reaches live subscribers

A pause writes no journal row, so it reaches subscribers only on the next
publish. Two pauses had none behind them: the drain's pre-consume failure
(the session is idle by then, so nothing else commits) and an old client's
Stop that interrupted nothing. A live card kept reading as waiting, with no
failure marker, until some unrelated commit arrived.

The drain now publishes after pausing a draft it failed to convert, and an
old client's Stop publishes when it paused a frontier.

- Tests: a failed conversion and an idle old-client Stop each reach a live
  subscriber as a paused card; both fail without the fix.

* fix(native-chat): a failed clear wakes the queued drain, a failed Stop withdrawal still publishes its pause

A conversation command can settle on the record alone (a retried clear that
fails), so drafts held behind its prepared phase waited for an unrelated
journal commit; the command controller now re-derives the drain when any
command finishes. A capable Stop whose withdrawal write failed never
published the pause it set, and a publish failure after a committed
withdrawal (Stop or clear) dropped the bodies from the answer; publishing now
happens outside the withdrawal and can no longer discard its result. Tests
reset the process-level pause set between cases: operation ids repeat per
test, so a shuffled order held later tests' drafts.

* refactor(native-chat): the draft store notifies through the journal's commit listener, the hold is a stored row fact, and one typed gate decides every queue hold

R1: every standalone draft-table transaction that changed rows (insert,
withdraw, hold, open-time repair) fires the journal's own commit listener
after COMMIT, so a draft or hold change publishes and wakes the drain through
the same path a journal row does — no call site can forget. All hand-written
publish/wake plumbing for draft changes is deleted; wakeQueuedDrain survives
only as the record-input wake (a conversation command can settle on the
record alone).

R2: the process-level pause set becomes a hold_reason column on the draft row
(pre-ship, so no migration): holds survive eviction and restart, keep their
send-failed marker across restarts, die with the session's journal, and are
cleared by consume and withdraw in their own UPDATE. The host-instance
derivation stays the one restart mechanism.

R3: one typed structuredQueueHold (blocked | command | prompt | working)
consumed by admission, the drain step and Send-now, with each caller's
override set written beside it. A capable send during a late-result /compact
now queues instead of being refused (PLAN §3.1); the dead prepared-command
branches and the drain's duplicated gate list are gone. prompt outranks
working so Send-now's one override cannot swallow it.

R4: one isUnsettledQueuedMessage predicate for the withdrawable/budget
filters.

Loop 4: a replayed send whose draft was refused answers with the returned
card, never the rejected submission, so the text cannot render twice. Rewind
completion was verified to publish after the record clears (the rewind path's
own publish; the open path's recovery precedes the open snapshot).

* fix(native-chat): a Stop with no drafts writes nothing, and a failed hold still lets a capable Stop withdraw

The stored hold turned Stop's in-memory pause into a draft-table write, so
every Stop (drafts or not, capability advertised or not) opened a BEGIN
IMMEDIATE/COMMIT. An empty hold now returns before the serialized write.

A hold that threw also emptied the frontier, so a capable Stop withdrew only
returned cards and left the waiting drafts unheld to auto-send after the
interrupt. The frontier is read once and survives a failed hold.

* fix(native-chat): a capable Stop with no drafts writes nothing

The empty-hold guard from the previous fix did not reach withdraw, so every
capable Stop still opened a write transaction after the interrupt, and a
closed handle turned its empty answer into a missing field. The draft store
now answers an empty withdraw without a transaction, for every caller.

* refactor(native-chat): Stop and /clear never withdraw queued drafts; no text rides the wire back

Adopt the host-owned-queue model end to end: a Stop holds the waiting
frontier ('stopped') for EVERY client and interrupts — the cards stay
published as paused, Send-now overrides per card, and the pause dies when
the user next starts a turn (an ordinary dispatched send lifts 'stopped'
holds in the same serialized step; 'send_failed' holds still need their
explicit Send). /clear carries the source's unsettled drafts to the
replacement session as born-held rows — identical for every client
version — then tombstones the source. Delete answers with no body: the
card leaving the published list is the outcome.

Removed (never shipped; the capability was dark and unadvertised, so no
wire compatibility is affected): CancelParams.withdrawQueued and its
refine, ConversationCommandParams.withdrawQueued,
CancelResult.withdrawnQueued, ConversationCommandResult.withdrawnQueued,
AgentSessionWithdrawnQueuedMessage, the Delete result body,
settleStopQueuedWithdrawal and the cancel finisher,
withdrawClearedSourceQueuedMessages, replayWithdrawnQueuedMessages, and
cancelPlan's tombstone replay. This also removes the defect where a
withdrawal took every row regardless of which client sent it (a phone
Stop pulled desktop-typed text): nothing moves text anymore, so a Stop
from one client can never relocate another client's drafts.

Hold and carry writes are bookkeeping: a failure is logged and never
gates the interrupt or the clear.

* feat(native-chat): a restart hold lifts like a Stop's, and paused cards say why

The user's next dispatched send lifts every stop-shaped hold in one
UPDATE: stored 'stopped' rows, and restart-held rows (host_instance
mismatch), which are adopted into the running instance — the same fact
the derivation reads, so no second copy of the hold exists. 'send_failed'
still requires its explicit Send. Publication now marks stop/restart
holds with pausedReason 'stopped' (an additive optional value on a dark
capability), so clients can caption them "sends after your next
message" and keep "couldn't send" for 'send_failed'.

* fix(native-chat): only a client's own send lifts a Stop's queue pause

The lift ran for every accepted host send, so orchestration mail, a
restart continuation and a launch prompt released drafts the user had
stopped (and adopted restart-held rows into the running instance). The
client-facing agentSession.send RPC now marks its sends as the user's
own; host-internal senders leave the pause alone. Also drops comments
still describing the withdrawn return-text rule.

* fix(native-chat): a Stop's queue pause lifts when the user's send starts its turn

The pause lifted as soon as the host accepted a user send, so a send the
provider then refused (a failed child start, a refused turn/start) had
already released the stopped drafts into the same failure. The host now
remembers a client's own send, in memory, until the provider answers it:
acceptance lifts the stop-shaped holds, a refusal forgets it with the
holds intact, and a later Stop supersedes it. Nothing is persisted, so a
restart between the send and its turn start leaves the cards held for the
user's next send rather than sending them unasked.

* fix(native-chat): a consumed draft's turn starting lifts a Stop's queue pause

Drafts are only ever a client's own sends, so a drained draft or a
Send-now is a user send for the pause: its submission joins the same
in-memory set a direct send uses, and the provider accepting it lifts the
stop-shaped holds. Before, a message typed while a stopped turn wound
down drained as a draft and left the older stopped cards held, so their
"sends after your next message" caption was false. A refused consumption
lifts nothing, a later Stop still clears the set, and orchestration mail
and restart continuations still never lift.

* fix(native-chat): queue a capable send behind a /compact and re-scope /clear's carried drafts

- A text send with queue-if-active during a /compact in flight is admitted on the
  compact's side lane as a held draft instead of being refused; it may only become
  a draft, so one the gate no longer holds is refused rather than dispatched.
- Drafts /clear carries to the replacement are fingerprinted for the replacement
  session, so the provider's echo folds into the sent bubble.
- The in-memory set of user sends awaiting their turn is capped; sends settling
  unknown no longer grow it without bound.
- Correct the userSend comment: the renderer's launch prompt goes through the
  client RPC and does set it.

* fix(native-chat): a returned queued card carries the typed rejection fact, like a rejected submission

A consumed draft the agent never ran comes back as a returned card. The card
kept only the rejection's sentence, while its submission now also records the
typed fact a client classifies from. A host-restart rejection's sentence
carries no legacy marker, so such a card could not be told apart from a
provider's refusal.

The draft table stores the submission's fact next to its reason
(`returned_rejection`, written by the same settlement that sets the reason,
and read back with the reducer's own fact reader), and the card publishes it
as `returnedRejection`. Both are overwritten on every return, so a re-sent
card never keeps an earlier refusal's fact, and a /clear carry inserts a plain
held draft with neither.

Retention moves to queued-message-retention.ts to keep the table module
within max-lines.

* fix(native-chat): fit the queue to main's typed rejections and compaction result

Main (#23026) dropped the disposition's fresh-id retry field, gives a
rejected dispatch a typed sentence plus fact, and types /compact's result.
The queued-draft disposition and the queue tests now use those shapes.

* fix(native-chat): draft bookkeeping can never roll back the journal row it rides

The queued-draft returned transition runs inside every journal append's
transaction. A throw there (a draft table an earlier build created without the
returned_rejection column) rolled back the journal's own rejection row, so a
Stop, a failed start or a provider refusal could not be recorded. The standing
hook now runs in its own savepoint: its failure is logged and rolls back alone,
and the open-time repair re-derives the missed transition from the committed
row. The draft table also gains any missing nullable column at open.

* fix(native-chat): a draft a Stop or restart took back waits again instead of blocking the queue

Cards A, B and C wait; the turn ends and the drain consumes A, but the agent
has not taken it yet. A Stop then pauses B and C and withdraws A's submission,
which made A a returned card. The user's next send lifted B and C, yet a
returned card blocks everything behind it, so B and C never sent although they
read "sends after your next message". A restart or close before hand-over did
the same.

Nobody failed the user there, so the draft now goes back to waiting at its own
position, under the hold that same event put on the drafts behind it: a Stop's
'stopped', or no stored hold after a restart, whose hold derives from the host
instance. It carries no refusal, and records its spent submission id in
consumed_as, so its next consume (the drain, or Send on the card) mints a fresh
id through the same path a returned card's re-send uses. Provider refusals and
other failures still return the card. The live settlement hook and the
open-time repair share one decision. After a Stop and the user's next turn,
A drains first, then B, then C, one per turn.

* fix(native-chat): Delete and Send on a queued card answer at once during a /compact

A /compact holds the chat's serialized lane for its whole provider call, and
the queued-card Delete and Send ran on that lane, so both hung until the
compaction finished. They now run on the side lane a draft-only send already
uses while a compaction is in flight: Delete completes at once, and Send
reaches its readable "wait for the conversation operation" refusal at once.
The drain stays on the main lane and keeps its command hold, so nothing sends
until the compaction settles.

* fix(native-chat): a re-sent returned card drops the refusal it came back with

Re-consuming a returned card left returned_reason and returned_rejection on the
now-dispatched row, so the row described a refusal that no longer applied. The
consume clears both in the same update that moves the card to dispatched.

* perf(native-chat): the queue gate reads pending prompts without rendering the journal

The prompt check ran on every send admission and drain step, and read
journal.snapshot(), which copies and sorts every item in the chat. It now walks
the reduced items in place with journal.visitItems; the answer is the same,
since the snapshot only sorts those items.

* fix(native-chat): a Stop that fails leaves the queued cards as it found them

Stop holds the waiting cards before it withdraws queued sends and interrupts
the agent. When a later step threw or the Stop was refused, the cards stayed
paused ("sends after your next message") although a failed Stop is meant to
change nothing. A failed Stop now undoes exactly what it added: each card it
held gets back the hold it replaced, a consumed card its withdrawal sent back
to waiting is released, and the user sends it had set aside can again lift the
pause. Holds an earlier Stop or a restart put on the cards stay.

The hold SQL moves to its own module, and the draft store's standalone
transactions share one helper.

* docs(native-chat): confirmed cancellation is no longer a queue rollout prerequisite

Stop withdrawing queued sends with a typed cancellation landed on main with
#23026. The comment gating the queued-messages capability now lists only what
remains: the Codex steer matrix (#21062), the Claude fold receipt, turn-owner
bars, and the desktop and phone clients.

* docs(native-chat): the Claude fold receipt and turn-owner bars have landed; Codex steer and the clients remain

* fix(native-chat): Send on a queued card during a /compact is refused before it takes a lane

Send-now chose its lane once, at entry. During a /compact it took the side
lane, where it could wait behind a Stop, then run after the compaction had
settled and append a real submission unserialized against the main lane.
While a compaction is in flight, Send-now is now answered with the "wait for
the conversation operation" refusal before entering any lane, and otherwise it
runs on the main lane. Only Delete keeps the side lane, whose compare-and-set
withdrawal is safe on either.

* fix(native-chat): a Stop that fails after reaching the agent keeps the queue paused

A failed Stop undid its queue holds whenever it threw, including after the
interrupt had already gone to the provider (a status-note write failing after
cancelTurn, or after stopping a starting agent). The turn could be stopped
while the cards drained as if no Stop was pressed. The Stop now marks the step
that reaches the provider, and undoes its holds only when it failed before
that. A Stop the agent refused answers ok and keeps its holds; the comment no
longer claims otherwise.

* fix(native-chat): a skipped draft settlement heals on the next drain step, not only at reopen

The draft settlement rides each journal append as bookkeeping, and a failure
there is logged and skipped. Only the open-time repair re-derived it, so a
consumed draft whose submission was rejected stayed dispatched (invisible, and
blocking nothing it should) until the chat reopened. The re-derivation is now
its own function, shared by the open-time repair and the drain: whenever a
dispatched draft's submission is already rejected, the drain step applies the
owed settlement first.

* fix(native-chat): one id is never recorded as a submission twice

A second submission row under an id the journal already holds replaces the
submission with a fresh pending one, so a rejected message could be handed
over again under its own id. Send on a queued card could do exactly that: if
the host died after it consumed the card under the operation's id but before
its answer settled, the rerun consumed again under the same id.

The journal now refuses a submission under an id it already records, so no id
is delivered twice whatever the caller does. And a Send-now rerun that finds
the card consumed under its own operation id answers with that submission
instead of consuming again.

* fix(native-chat): a waiting draft whose first send the agent echoed is withdrawn, never resent

A consumed draft goes back to waiting when its submission is rejected as never
delivered (a Stop's withdrawal, a restart, a close), and then sends again
automatically. That rests on the "never delivered" claim. If the provider then
echoes that message, the first delivery happened, and the automatic resend
would give the agent the same message twice.

The reducer already keeps such an echo apart, since a rejected submission may
not claim it, so the draft store reads it from the appended row itself: a
provider echo of a user message that no live submission claims, matching a
waiting draft whose spent submission is rejected, withdraws that draft the way
a Delete would. The echo-claiming rule is split out of the reducer's aliasing
so both read the same decision, and the per-row draft hook moves beside the
settlement re-derivation.

* feat(native-chat): a submission names the queued draft it hands off

Clients told a queued card's hand-off apart from other sends by comparing the
draft's id with the submission's id. That holds only for a draft's first
hand-off: a re-send, or a draft that goes back to waiting and drains again,
goes out under a fresh id, and the clients showed the card and the sent
message together, or restored text the host still held.

Every submission the host creates by handing off a draft now carries
queuedMessageId, the draft's id. It is written on the submission's journal row
as an optional key (older readers keep it and ignore it), carried by the
reducer, listed in the published submission schema (which otherwise strips
it), and stamped where the row is built from the consume itself, so no
hand-off path can leave it off; a caller naming a different draft is refused.
A direct send names none. The queued-messages capability comment makes the
link part of v1.

* refactor(native-chat): every queued draft goes out under a fresh submission id

A draft's first hand-off reused the draft's own id as the submission id, so
comparing a draft id with a submission id looked right in every first-send test
and failed only on a re-send or a requeued draft. Every hand-off now uses a
fresh id (the drain mints one; Send on a card uses its operation's id), so id
equality is never true and a reader must use the submission's queuedMessageId.

The host gets simpler: queuedMessageNeedsFreshSubmissionId is gone, consumed_as
is set on every dispatched row and cleared when a withdrawal sends the draft
back to waiting (its spent submissions stay findable by their link), the
consume refuses the draft's own id, and the consumedAs ?? messageId fallbacks
collapse. The delivered-echo check finds spent hand-offs by link.

A send this host queued, asked again (a lost answer's replay, or a rerun the
operation ledger no longer covers), answers from its draft and then from the
hand-off that names it, through one function. The rerun path used to be kept
from sending twice only because a submission sat under the send's own id;
with fresh ids that guard is now explicit. A Send-now rerun recognises its own
consume by the link instead of consumed_as.

* fix(native-chat): an echo withdraws a draft only if its rejected hand-off reached the agent

The delivered-echo rule withdrew a waiting draft when a provider echo matched
any rejected hand-off of it, including one a Stop rejected before it was ever
handed over. That hand-off is provably unwritten, so a matching unclaimed echo
is some other message, and the rule silently deleted the card. Only a hand-off
that was handed over and then rejected as never delivered can be disproved by
an echo now.

* fix(native-chat): a skipped echo withdrawal is re-derived before the draft can send again

The delivered-echo withdrawal rides each journal append as bookkeeping, and a
skipped hook left the draft waiting, so it later sent the same message a
second time. Nothing re-derived it. The draft store now also withdraws, in its
owed-settlement pass, each waiting draft that an echo already in the journal
proves delivered: an unclaimed provider user message (still stored under its
own id), carrying the draft's payload, appended after a hand-off that was
handed over and rejected. The live hook and the re-derivation share one
predicate. The pass runs at open and in the drain step, right before a draft
would send; it reads every item, so it never runs per streamed row.

* fix(native-chat): a rolled-back journal append leaves no draft state cached

The draft store caches its row list by revision. The per-row hook read that
list eagerly inside the append's transaction, after the consume in the same
transaction had already written and bumped the revision, so a failed COMMIT
left the cache showing a hand-off that never happened. The hook now reads the
drafts only once a row holds an unclaimed echo, and any rollback of a journal
append or of its bookkeeping savepoint invalidates the cache, so no other read
inside the transaction can leave it stale either.

* fix(native-chat): a replay of a deleted queued card answers withdrawn, not refused

Once a deleted card's tombstone is pruned, a replay of the send that queued it
found the draft through its last hand-off. When that hand-off had been
rejected (the card came back, and the user then deleted it), the replay
answered with the rejected submission, which clients show as a failed send
with a Retry. Only a withdrawn row is pruned while its last hand-off stands
rejected, so the replay now answers queued, withdrawn.

* refactor(native-chat): name the queue's pause-lift for what it releases

* chore(native-chat): one import of the mutation helpers

* feat(native-chat): a Stop pauses the whole queue, derived from the journal, with an explicit Resume

After a Stop, each waiting card was held on its own row ('stopped'), lifted
when the host saw, in memory, that a user send made after the Stop had its
turn accepted. The cards read "sends after your next message" one by one,
there was no way to resume the queue without sending something, and the
in-memory record of user sends was lost on a restart or eviction.

The pause is now the queue's, and derived rather than stored as a flag:
- 'stopped': the user's last Stop took effect at a recorded journal position
  and no turn a person asked for has started since. "A person asked for it"
  is the new `origin: 'client'` on the submission row (a send over the client
  send RPC, or a card they sent now); orchestration mail, a restart
  continuation, a host-sent launch prompt and the queue's own drain record
  `host` and never lift it.
- 'restarted': a waiting card was written by another host process and no
  person's turn has started since this conversation opened.
Resume (`agentSession.queuedMessagesResume`) lifts either. Send-now sends one
card; the rest stay paused until that card's turn starts, which is a person's
turn like any other.

The journal's row kinds are closed (an older build truncates a journal at a
row kind it does not know), so the one event the journal cannot carry, where
the Stop took effect, is recorded beside the drafts in `queued_message_pauses`;
everything after it is read from the journal. A Stop records it only once it
takes effect (after withdrawing queued sends, as it reaches the agent), so a
Stop that fails first leaves nothing to undo, and the per-row hold, its undo
and `userSendsAwaitingTurn` are gone. A card keeps a hold of its own only when
its conversion failed ('send_failed').

The pause is published once, as `queuePause` beside `queuedMessages`, on live
frames, catch-up and history. A /clear starts its replacement paused, as after
a Stop, since the carried cards were written for the context it discarded.

* feat(native-chat): a /clear's replacement queue reads paused because of the clear, not an interrupt

The replacement's pause was recorded as 'stopped', which clients show as
"Queue paused because you interrupted" although the user cleared the chat.
It is now its own reason, 'cleared', on queuePause.reason
('stopped' | 'restarted' | 'cleared'). It lifts and resumes exactly like a
Stop's: through Resume, or the user's next turn starting on the replacement.

* fix(native-chat): a queue pause covers only the cards it paused

A Stop recorded its pause fact even when the queue had no cards, and the fact
outlived the cards it did pause. The published list hid a pause over no cards,
but the drain still treated the queue as paused, so a card typed much later —
during an orchestration-mail turn, or a correction typed before the stopped
turn ended — sat under "paused because you interrupted" with no Stop of its
own.

A Stop now records its pause only if the queue holds a card when the Stop takes
effect (the hand-offs its withdrawal sent back included). The fact is retired
in the same transaction as the Delete, consume or withdrawal that empties the
queue, never from an async publish. A /clear's carry now lands each card with
its 'cleared' pause in one transaction, so a failed insert leaves no pause over
an empty replacement.

* perf(native-chat): the queue's pause reads the latest person's turn in O(1)

The pause is derived on every publish, per subscriber, and each derivation
copied and scanned every submission to find a person's accepted turn after the
Stop. The reducer now keeps that fact as it folds rows: the submission row of
the latest accepted turn whose origin is `client`. The Stop's and the
restart's lift both read it directly.

* fix(native-chat): a card handed off after a restart belongs to the process that sent it

A draft's host_instance was only ever the process that first wrote it (or
adopted it while waiting). A returned card from before a restart, sent again
in this process and then withdrawn back to waiting, still carried the old
process, so it raised a 'restarted' pause although no restart happened since
it was sent. Every hand-off (the drain, Send on a card) now stamps the
handing-off process on the draft in the consume's own update.

* fix(native-chat): a queue pause shows only while Resume would send something

After a Stop whose only remaining card was a returned one, or after a restart
with only a card held by its own failed send, the queue published a pause with
a Resume that could send nothing: a returned card waits for the user anyway,
and a held one for its own Send. The pause is now published, recorded by a
Stop, and kept only over a card it can hold back — waiting, with no hold of
its own. The fact is retired in the same transaction as the write that removes
the last such card, a hold or a refusal included.

The publication's dedup also compared only the pause's reason, so a pause
appearing or clearing with no readable reason could read as unchanged; it now
compares presence first.

* fix(native-chat): a queue pause counts only cards Resume would actually send

A waiting card behind a returned one is blocked until the user acts on the
returned card — the drain never sends past it — so a pause over only such
cards still offered a Resume that sent nothing. The rule for "a card Resume
would send" is now one function: waiting, no hold of its own, and not behind a
returned card. The publication, a Stop's record and the fact's retirement all
read it; retirement reads the rows in position order inside the same
transaction as the write that took the last such card.

* fix(native-chat): a returned card that blocks the paused cards hides the pause but keeps it

The last change retired a Stop's pause as soon as a returned card blocked every
paused card. Deleting that returned card then sent the cards behind it at once,
with no Resume — not what the user asked for.

The two rules are now separate. The pause is KEPT (recorded by a Stop, retired
in the same transaction as the write that takes the last one) while any waiting
card with no hold of its own exists, wherever it sits. It is PUBLISHED only
while such a card is not behind a returned one, so the header never offers a
Resume that sends nothing. Deleting the blocking card shows the pause again,
and the cards behind it wait for Resume or the user's next turn.

* fix(native-chat): a Stop pauses a card its withdrawal sent back even when that settlement was skipped

The Stop checked the draft table for a card to pause. When the per-row hook
that settles a withdrawn hand-off was skipped, that card was still
'dispatched', so the Stop recorded no pause; the drain later healed it back to
waiting and sent it, although the user had pressed Stop. Retirement had the
same blind spot and could drop a pause while such a card was owed.

What a pause holds back is now one predicate, judged inside the transaction
that records or retires it: a waiting card with no hold of its own (one SQL
EXISTS), or a dispatched card whose consumed submission was rejected with a
settlement back to waiting (read against the journal's submissions). The Stop
first runs the owed settlement, as the drain does; if that fails, the owed
card still counts, so the pause is recorded rather than skipped. recordPause
now checks inside its own transaction and returns whether it recorded, and any
draft-table write (and the per-row hook, the consume and the open-time
repair) retires a pause that no longer holds anything back.

* test(native-chat): pin the per-row hook's pause retirement; skip the judgement when no pause exists

The retirement test recorded its second pause over a queue with nothing to hold
back, so the recording returned false and the "retired" assertion proved
nothing; ablating the per-row hook's retirement passed every test. The hold
case now asserts the pause was recorded, and a new test has a delivered echo,
through the per-row hook, withdraw the last card a recorded pause holds back.

Retirement runs on every appended journal row, so it now checks the pause row
by key first and judges nothing when no pause is recorded. Two comments were
brought in line with the owed-hand-off rule and rewrapped.

* test(native-chat): match main's append and dispatch shapes in the queue tests

* fix(native-chat): read a compaction's settled submission through the send-result union

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-29 15:22:40 -07:00
Brennan BensonandClaude d60f999f94 fix(native-chat): turn facts come from the turn record, and /compact is a message the chat sends (#23059)
* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* fix(native-chat): name a chat write by its target, not the owner generation

A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.

Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.

Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.

* fix(native-chat): every journal append reaches the chats that are open

A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.

A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.

* test(native-chat): an epoch replacement reaches the open chat

* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

* perf(native-chat): a publish behind a delivered commit reads nothing

Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.

* test(native-chat): state why the teardown test's fake journal is safe to cast

* docs(native-chat): say mutation admission checks only the writer lease

* docs(native-chat): drop the send rebase from comments that still described it

* fix(native-chat): a message is accepted, then delivered

A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".

A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.

Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.

A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.

Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.

* fix(native-chat): settle queued messages only for the child that ended

A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.

A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.

The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.

* fix(native-chat): an adoption that fails to import keeps the conversation open

The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.

* perf(native-chat): the recovering open reads the journal once

Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.

* fix(native-chat): an attach that fails after indexing its child leaves no child behind

A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.

* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer

The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.

A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.

* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down

The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.

* fix(native-chat): a message rejected while its chat was closed reads as not sent

A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.

* test(orchestration): name why the readiness settlement fakes are cast

* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent

* docs(native-chat): drop the fence from the admission the send effects run behind

* docs(native-chat): give the fence move on release the reason that still holds

* docs(native-chat): stop citing a write fence check in launch and mailbox comments

Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.

* refactor(native-chat): the provider child is its own record

A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.

- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
  patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
  own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
  the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
  dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
  is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
  the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.

* fix(native-chat): the delivery loop alone settles a message its start or child failed

A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.

- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
  reads how it ended: a Stop continues; anything else writes one failure row and rejects every
  queued message with the same words, then stops. A child still starting whose start the adapter
  says did not land fails the same way. The exit, eviction and the settlement retry only settle
  the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
  nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
  failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
  closed, with or without a child, and a start the loop already has in flight is waited for so the
  child it produces is stopped rather than left behind.

* refactor(native-chat): a stopped child ends on the one reading of its stop

The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.

* feat(native-chat): the host says it accepts a send before any agent has it

The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.

* refactor(native-chat): an attach never opens a journal of its own

The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.

* fix(native-chat): a moved fence resends nothing on a host that accepts first

The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.

The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.

* refactor(native-chat): a child's end says whether the user or the host stopped it

The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.

* fix(native-chat): a chat whose only work is a queued message is not offered for resume

A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.

* fix(native-chat): the conversation outlives its agent

Opening a chat no longer starts its agent. A conversation is reached through one host
accessor that opens its journal at rest, and a send is what starts the agent, through
the delivery loop. One idle sweep, every five minutes, stops an agent that has been
quiet for thirty minutes and owes no work, then drops an open journal handle that is
only a cache. Its record, tab, status row and readers stay.

- hold and release are no-ops; hold still builds the host for shipped mobile builds.
- The holders, the holds, the release clock and the exit respawn are deleted.
- Options, the model list, the goal and the context meter answer at rest; a model pick
  at rest is recorded as intent for the next start.
- Compact, rewind, clear and goal changes start the agent first. A send does too when
  a rewind is still in doubt after the conversation opens.
- Orchestration routes mail and group addresses on ownership (the record plus the chat
  tab), not on whether the process runs. An open dispatch keeps its worker running.
- The restart continuation is a send; Resume all holds each slot until the message is
  handed over or rejected.
- A read error never replaces a loaded transcript, and shows the host's own words.

* test(native-chat): type the queued-message fixtures in the resume-offer tests

* fix(native-chat): a start that dies while a message waits on it is that message's failed start

Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.

* fix(native-chat): a request that failed reads as failed

A structured chat whose only message the agent's start refused read as a
green finish, and a cancelled structured turn did too: the host published a
verdict only for turn records, and structured rows carried no `interrupted`.

The host projection now reads the session's latest request: its turn's
outcome, or `failure` for a send the agent or its start refused. A send
that was withdrawn, or left undelivered by a restart or a close, fails
nobody and makes nothing listable. The ingest publishes `interrupted` as the
hook lanes do, and every reader decodes the verdict through one accessor, so
a failure reads Failed on the dot, the rollups, history and `worktree ps`,
behaves like a cancellation in every clean-finish policy, and notifies as
"failed".

* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now

* test(native-chat): a verdict change republishes the mobile status projection

* refactor(native-chat): the store's retention trigger keeps its flag compare

A verdict change always moves the completion clock the same check already
reads, so a second verdict compare there caught nothing new.

* test(native-chat): a user message the provider journaled keeps its session listed

* test(native-chat): pin what a failed start settles, and what a resume offer names

A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.

* test(native-chat): the failed-start pins fail on what the message became, not on a timeout

* fix(native-chat): a restart offer ends when the chat's agent starts again

The offer used to end only when the chat's newest user message changed,
because opening a chat started its agent and that start could not be told
apart from real activity. Opening a chat starts nothing now, so the host
reads the fact it already publishes: a chat's status row goes from not
host-owned to host-owned exactly when its agent is started. At that edge the
offer and any failure record for the chat are withdrawn, unless the start is
a resume action's own (its continuation is the oldest undelivered message).

A continuation and a message racing to be first are decided at acceptance:
the continuation is refused, quietly and with nothing filed, when any other
message was accepted since the restart. A failed continuation start leaves
the offer retryable, and each resume action sends its own message id.

Deleted: the newest-user-message comparison, its journal reader, the
continuation filter, and the failure ledger's own "answered by the chat"
check. The marker still carries its message id for one release, so the
previous build can read it.

* fix(runtime): end a transcript stream when its client unsubscribes

Desktop: the IPC subscription controller was dropped as soon as the streaming
handler returned, which for most streams is right after it binds. A later
runtime:unsubscribe then found nothing to abort, so the host kept the subscriber
and derived and sent every publish to a channel no one listened to. The controller
now lives until the renderer unsubscribes, resubscribes the same id, or goes away.

Mobile: disposing an agentSession.subscribe stream now sends agentSession.unsubscribe
with the stream's frame id, so the host ends that subscriber and leaves a sibling
stream on the same socket running. The direct path now passes the frame id the relay
path already passed.

* fix(native-chat): a late provider-session update keeps a failed recovery record failed

A provider-session heartbeat that rewrites a completed recovery record kept
its interrupted flag but dropped the outcome it was copied with, so a live
failed checkpoint read as a clean finish until the next status write.

* test(orchestration): the preamble's host stub is typed, not cast

The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.

* test(native-chat): the terminal-bell check asserts the renamed verdict field

The bell notification test still checked for agentInterrupted, which no
longer exists, so it could not catch a verdict leaking into a bell dispatch.

* fix(native-chat): a failed turn ranks like a completion for attention

Attention readers (completion time, Smart Sort, sticky retention, Cmd+J
Recent) now demote only a turn the user stopped. A failure is news the
user has not seen, so it keeps its completion time, ranks in the Done
class, stays retained after its pane goes away, and a retained failure
reads failed in the worktree rollup instead of done. Clean-finish
policy (hibernation, pane ownership, the value moment) still treats a
failure like a stop.

The retention trigger compares verdicts again: success -> failure no
longer moves the completion clock.

* fix(native-chat): one fact ends a restart offer: the chat moved on since the restart

The offer is live while no other message has been accepted in the chat since the
restart and its agent has not proved a start since. The offer list, the resume's
reservation check and the continuation's acceptance check all read that one fact,
so a message whose start then failed withdraws the offer too, and a stale click
finds nothing to act on.

The fact is read off the conversation's open handle, which the restart closed, so
it is retired durably whenever it may have changed: a message accepted, a start
proven. A close and reopen within the same run therefore cannot bring the offer
back. A continuation rejected before it reached the agent does not count, so a
retry after a failed start still runs.

Deleted: the quit-time gate on withdrawal, which changed nothing because the
withdrawal and the quit's own offer write share one queue; the per-action
"withdrawn" flag and the separate acceptance check it paired with.

* test(native-chat): an older build reads the restart offer this build records

The offer lives in a file the previous release reads after a downgrade. Pin that
against the pinned release's own capsule, and run the lane when the marker or the
capsule changes.

* fix(native-chat): read a restart offer against where the journal stood when it was taken

"Since the restart" was read off the conversation's open handle, which the idle
sweep closes: after a reopen, a message the user had already sent looked older
than the handle and the withdrawn offer came back.

The offer now records the journal position (epoch and sequence) at the moment
it is taken, and a message accepted after that position, or a journal on another
epoch, means the chat moved on. That is derived from the journal, so it holds
across any number of closes and reopens. An older build's offer has no position;
only a start withdraws it. Because the message half is now durable, the offer is
no longer rewritten in the recovery file on every accepted message; a proven
start still writes it, since only the host that saw the start knows of it.

* test(native-chat): wait for the listing's retire write before reading the recovery file

* refactor(native-chat): every journal row states which turn it belongs to

Rows gain a turn scope stated by the write that creates them: the open root
turn, or the conversation. A queued message takes its scope from its handover.
Rows stored before scopes existed are placed on replay by the root turn open
when they were created, so no persisted state is needed for them. Rewind keeps
each retained row's scope and producer, so a subagent's row stays its own.

* fix(native-chat): keep the terminal-backed chat's read error over its local echoes

Messages winning over a read error is right for the structured chat, whose read retries and whose
messages came from the transcript. The terminal-backed view assembles its list from local echoes
too (a launch prompt, a pending send), so a failed read there showed only those bubbles and no
error. Only the structured pane now keeps messages over an error.

* fix(native-chat): a start retries the exit settlement a failed journal write left owed

An agent exit whose journal settlement write failed releases the lease latched until a retry lands.
Reopening the chat used to be that retry; with reveal now only opening the journal, nothing retried
it before the next app launch, and every send was refused. The start the send needs now runs the
retry first, where the attach would.

* fix(native-chat): a failed main agent reads failed while its subagents still work

The verdict is now read from the main agent's own state, not the folded
row: a main agent that is done and failed has a verdict even while its
subagents keep the row working. Without mainAgent (history, worktree ps,
older hosts) the old combined-done rule stands.

Display marks the verdict through agentVerdictDisplayMark: a failure
outranks every combined state on the agent's dot, label, tab badge,
dashboard and activity rows; a stop marks only a done row, so a
successful or stopped main agent with live subagents still reads
working. Subagent rows keep their own state. The worktree card, terminal
tab and Cmd+J rollups share one pane fold and rank a pending question,
then failed, then working, monitoring, interrupted and done.

worktree ps publishes the main agent's outcome on a working row, and the
mobile mirror reads it. The store's change check, the paired-client
mirror's equality and its epoch now see a verdict change on a working
row, which otherwise moves no state or clock and left the worktree card
reading working. Clean-finish policy is unchanged: a working row is never
hibernated and has no completion time.

* perf(native-chat): answer the owner check without opening the chat

Worktree activation calls agentSession.handoffStatus for every chat tab in the worktree, and the
answer comes from the session record alone. Reaching it through the accessor opened each resting
chat's journal (a full read, the crash-boundary write and a restored status publish), then kept it
open for the idle window. It now checks the record and the adapter's support, as before this series,
and opens nothing.

* fix(native-chat): a read waiting on the session lock opens nothing once quit began

The accessor checked for quit before queueing the open, so a read queued behind a session task ran
its open after teardown had begun and indexed a journal no teardown step would close. The check now
runs at the open itself.

* test(native-chat): pin stated turn scopes, the upcast of unscoped rows, and rewind attribution

* fix(native-chat): /compact is a message the chat sends, run as a turn of its own

The conversation command RPC now accepts /compact into the queue like any
send and answers once it is handed over. The delivery loop opens the command's
own turn, starts the provider on it, and waits for the provider's end off the
session's queue, so messages typed meanwhile are held and delivered after it,
even when it fails. It settles by re-reading the journal: a child that died
meanwhile already wrote the verdict. Stop ends the command at once. The 180 s
completion window, the unconfirmed row and the recovery of an older build's
compaction record are gone; that record no longer gates anything. On Codex the
provider turn the command opens is claimed into the command's turn.

* fix(native-chat): read a failed resume's chat before calling it retryable

Whether a failed resume is retryable is the offer's own rule: the chat has not moved on since the
restart, read from its journal. The failure list read it only for a chat already open, so once the
idle sweep closed a chat the user had moved on in, its failure showed Retry again, and the click did
nothing. The list now opens the failed chats first, as the offer list does.

* test(native-chat): type the provider event sink the settlement test reaches for

* docs(native-chat): the worktree ps outcome comment no longer claims old hosts send it

The field is new: an old host sends no outcome at all, so a reader falls
back to interrupted. The removed clause said old hosts send it on done
rows, which never shipped.

* fix(native-chat): say the structured read keeps trying only where it does

The structured pane's "Orca keeps trying to load it" line never showed: the view state filled in an
untranslated fallback whenever the read error had no text, and the empty state prefers any message.
The view state now leaves the message out, so the structured pane shows that line and the
terminal-backed pane its own translated one. Mobile's structured lane does not resubscribe after an
error frame, so it no longer makes the claim.

* fix(native-chat): rows group under the turn their record names, not the one above them

Each row's turn is the turn its stated scope names, anchored on the entry
that opened it, or on the turn itself when the provider opened it unasked.
So /compact groups its own rows and the previous turn is untouched, a message
typed into a running turn joins it, and a provider-resumed turn folds under
its own Worked-for. A row reporting how a turn ended, an error or the
compaction separator, never folds. Desktop and mobile read the same keys; a
host that states no scope keeps today's positional grouping.

* test(native-chat): await the send's settlement instead of polling for the start

The at-rest send tests polled for the provider start with vi.waitFor's one-second default, which a
loaded machine outran. They now await the host's own settlement of the message.

* docs(native-chat): the status-store listing rule names provider-journaled user messages

* fix(native-chat): a restart offer resumes any time after the quit, and knows its own continuations

The continuation's message id was dated by the quit, and the ledger refuses a new id dated more than
a day back, so Resume or Retry a day after quitting was always refused (on main too). It is now
dated by the resume action.

Telling a rejected continuation from the user's own message read the operation ledger, whose rows
expire after about a day; after that a failed resume stopped being retryable. The offer now
records the continuation each action sends on its own capsule entry, bounded to the newest 16, so
the ids end with the offer. The ledger read is deleted.

* fix(native-chat): a /compact is not a request the sidebar, notifications or restart resume report

The sidebar's prompt, preview, verdict and instant, the turn-completion feed,
and the restart-resume marker read past a conversation command and its turn to
the last real request, so a /compact neither notifies nor re-dates the row,
and a command in flight is never offered as work to resume. An older client
shown a command's turn in the legacy form names the session's own agent.

* fix(orchestration): route no mail to a structured worker its orchestration released

A structured worker is routed on ownership, and a resting worker's lease is released, so ownership
held while its chat tab stayed listed. A worker the coordinator abandoned and then released, found
at rest by the release, therefore still took peer mail and @worktree: broadcasts, and each one
restarted its agent. Routing now also reads the orchestration's own resource row: once it is
released, direct mail, group addressing and worker-show's addressable answer drop the worker, as
they would a terminal worker whose terminal closed. The chat tab stays, and nothing new is stored.

* fix(native-chat): a failed retry names the user's prompt, not Orca's continuation

A resume's continuation is written to the chat before its start, so after a failed attempt the chat's
newest user message is that rejected continuation. A second failure then showed Orca's own restart
text as the chat's prompt. A retry now keeps the prompt its first failure named.

* test(native-chat): pin what a conversation command's admission refuses at rest and at handover

* test(native-chat): tests merged from the base state which turn their rows belong to

* fix(native-chat): a refused send notifies failed through the completion feed

The host's completion feed followed only the newest turn, so a send the
agent or its start refused, which creates no turn, read Failed on its row
but sent no notification. The feed now follows the session's latest
request, read from the projection the status feed already makes for the
commit: a turn keeps its id, a refused send is named by its journal item
key. It announces only while the session is idle, as the row reports a
verdict, so queued sends refused one commit at a time notify once, and a
withdrawn send falls back to a request already announced.

* fix(orchestration): read the released row optionally, as the authority does

worker-show's observation called the row lookup directly, which a runtime double without it threw on
and failed the structured tab-retirement release.

* chore(native-chat): one import per module and no unexplained casts in the turn-scope changes

* test(claude): pin which turn a Claude row joins, including a subagent's after the turn ends

* fix(native-chat): the status bar drops a restart offer the chat moved on from

The renderer re-read the host's restart offer only when a failed chat showed activity, so after a
message withdrew a pending offer the host answered no chats while the status bar kept counting one,
and clicking it opened nothing. The same watch now covers pending offers: a status change in an
offered chat asks the host again, once.

* fix(native-chat): a refused steer is read from the turn its handover named

The latest-request reader decided whether a refused send had joined a running turn by comparing
host clocks: its handover time against the previous turn's end. The handover row now states the
turn it delivered into, so the reader reads that instead and the clock comparison goes. A journal
written before handover rows stated a turn is scoped on replay from the turn open when each row
was written, which can differ from the clock reading only when a send and a turn's end share a
millisecond.

* fix(mobile): the native-chat controller contract carries the turn journal

The controller and overlay already pass nativeChatTurnJournal, but the
contract type never declared it, so mobile failed to typecheck.

* fix(native-chat): the live turn is the running turn, not the newest user row

A turn the provider opened on its own (a background wake, a resumed turn)
anchors on its own record, but the list still treated the newest user row
as the live turn. While such a turn ran, the settled user turn before it
lost its duration and the running turn's own rows were drawn as settled,
so its tool calls lost their live state.

nativeChatTurnMembership now answers both questions from the turn record:
each row's turn, and the live turn (the running root turn's anchor, else
the newest user row, which is also all an unscoped host has). Desktop and
mobile key liveness, the timing clock and the live status's row on it.

* test(native-chat): a turn the provider opened keeps its own clock

Pins that the local turn clock follows the live turn, so a wake after a
settled turn does not restart that turn's clock when no host durations
are recorded.

* fix(native-chat): a running turn no message opened draws its status on no row

Its live status belongs to the transcript-tail indicator alone. Once it
settles, its duration draws at its first row as before; a running turn a
message opened still draws on that message.

* fix(native-chat): every copy of a row carries the main agent's own status

History entries, sleep records and `worktree ps` rows carried a flattened
top-level `outcome`, copied under different gates and without the main agent's
clock. They now carry `mainAgent` (state, outcome, stateStartedAt), the type
the live row already persists and sends, and every copy site takes it with
`interrupted` through one function, `agentVerdictFields`.

- The accessor reads `mainAgent` then the legacy flag; the mobile mirror
  matches it line for line.
- Sleep records admit `mainAgent` with `normalizeMainAgentStatusField`, so a
  malformed value drops the field, never the record.
- Mobile dates a main agent that failed under live subagents by its own clock,
  as desktop does, and its row equality compares `mainAgent`.
- The activity feed reads a history entry's own `mainAgent` instead of
  rebuilding one; the sync key and history equality compare it.

* test(native-chat): pin the worktree ps verdict across host and phone versions

Pairs the real v1.4.212 host and phone row reader with this build: an old phone
reads a new host's rows by `interrupted`, a new phone reads an old host's rows
(no `mainAgent`) the same way, and a new phone reads a failure under live
subagents as Failed, dated by `mainAgent.stateStartedAt`. The release checkout
now carries the phone's self-contained row reader, and the lane runs when the
`worktree ps` row producers change.

* test(mobile): name the parity table's row for its role

* test(native-chat): a roster of idle or finished children does not keep an agent awake

The sweep reads owed background work through the shared child-work liveness that upstream's
release clock adopted; a child that went idle or finished is not work the agent still owes.

* fix(native-chat): a request that settles while the user is asked something notifies once

The completion edge waited for an idle session, and a pending prompt (including a
subagent's approval) is not idle. Structured chat has no other attention producer,
so a main turn that finished while a subagent waited on the user sent nothing
until the prompt was answered.

The edge now waits only on owed work (a running turn or an unanswered send), which
the projection reports even beneath a pending prompt. A request that settles with
a prompt pending announces once; the renderer words it "needs input" from the
host status mirror's `attention`, and answering the prompt keeps the same request
identity, so it does not announce again. The wire shape is unchanged.

* fix(orchestration): a task dispatched into a resting structured worker keeps it running

The sweep's open-dispatch check read only the worker-start dispatch that owns the worker's terminal
resource, so a task later dispatched to the same worker (orchestration dispatch --to, which writes a
dispatch with no worker row) did not count: after thirty quiet minutes the worker was stopped while
that task was open, and its coordinator read exited. Any unsettled dispatch addressed to the worker's
process incarnation now counts, derived from the existing rows.

* fix(native-chat): a command's wait ends when its child does

The delivery loop waited for a /compact only on the adapter's compaction
tracker, which learns of the child's end only on some exit paths: a Codex
exit or close, and a Claude close, never reach it. The wait then never
ended, so nothing queued behind the command was delivered again, Stop had
no child to answer through, and the tracker's leftover entry refused the
next /compact.

Every way a child ends passes endProviderChild, so the host now offers a
per-child end signal there. The loop races the tracker against it (the
dead-generation settlement has already written the command's verdict),
and on that end asks every adapter to release the command, so a later
command runs and no later provider turn is claimed into the dead one.
The adapters' own exit-time releases were unreachable (Codex) or covered
one path of several (Claude), and are removed.

The Codex RPC test harness moves to its own module so the exit can be
driven through the real adapter's connection callback.

* fix(native-chat): keep refusing sends during a command on an older host

An older host's controller still refuses a send while a conversation
command runs, so dropping the client's block turned every message typed
during /compact into a 'not sent' row with Retry there. The block stays
for hosts that do not run the command as a send-path turn, and goes only
for those that do.

The signal is one the client already holds: a host that runs /compact on
the send path states a turn scope on every journal row it writes, the
same fact turn membership uses to tell it from an older host. Both now
read it from one predicate. On an empty conversation, or one whose rows
all predate the upgrade, the signal is absent until the command's own
entry streams in, so that brief window keeps the old local refusal; no
capability or wire field is added.

* docs(native-chat): comments stop describing the hold this PR removed

Eight comments still justified orderings and teardown choices by a viewer or dispatch hold that
pinned the provider child. Nothing holds any more; the orderings stand for the binding's redrive
subscription and parked mail, and a chat's agent runs from a send until the idle sweep rests it.
Comment-only.

* fix(native-chat): the completion says when the user is being asked

A request that settles while a prompt waits on the user was worded "needs input"
from the renderer's status-feed mirror. Remote clients receive the status and
completion streams over separate sockets, so they can arrive in either order and
the wording could be wrong both ways.

The host already knows at emit time, so the completion now carries an optional
`awaitingUser: true` in that case and omits it otherwise. The renderer words the
notification from that field alone and no longer reads the status mirror. Old
clients ignore the field and word by outcome; old hosts never send it.

* fix(native-chat): a restart offer keeps the start its own continuation made

Whose start ended an offer was decided at read time, from whether the offer's continuation was
still the queued message. Once the provider refused that continuation, the child it had started
read as someone else's start, so the offer ended and its failure showed no Retry. The delivery
loop now records which queued message a start is for on the in-memory child, and the child's end
carries it; the offer counts a start as its own when that message is one of its continuations.

* fix(native-chat): a rewound turn still names the message that opened it

A Codex rewind rebuilds the epoch without submissions, so each sent message survives only under
its provider key. The kept turn records still named the submission key, so each turn anchored on
itself and its rows grouped apart from the message that opened it. The rewind now renames the
turn's opener along with the message.

* fix(native-chat): Stop ends only the command it names

Stop on a command turn abandoned whatever compaction the session had pending, so a late Stop for
an earlier /compact cancelled the one running now. The tracker now ends a command only when the
Stop names its turn, and the cancel reply reports whether it did.

* fix(native-chat): an agent gets a full idle window after its owed work ends

The sweep measured quiet only from the last journal row, so once a subagent, command, monitor or
dispatch that had outlived the window ended, the agent was stopped at the next tick. A child can
read done before the lead's wake-up turn writes anything, and stopping in that gap loses the
wake-up. The sweep now counts owed work it observes as activity, which gives the agent the full
window afterwards, as the release clock it replaced did.

* test(claude): the options-read fixture runs a live child

The fixture marked its conversation running with a hasProviderChild field the
session type does not have, so the read took the at-rest path and refused a
session with no record. It now carries a child, which is what the read checks.

* test(native-chat): host tests reach its collaborators through a typed seam

The rest-test rig and three test files read the host's private members with
Reflect.get and cast the result. The host now exposes one test-only accessor,
collaboratorsForTests(), and the subscribers class a subscriberCountForTests()
beside its existing retainedActivityCountForTests(), so the tests are checked
against the real types and the casts are gone.

* fix(worktree-status): a departed agent's failure yields to live work on the worktree card

A retained failed agent has no expiry, so ranking it with a live failure pinned the card to Failed over other panes' live work. It now ranks below working, monitoring and permission, and above every finished outcome.

* refactor(orchestration): one owner answers a structured worker's custody

Routing, group addressing, worker-show and the idle sweep each composed their own reading of
whether orchestration still holds a structured worker, so each new obligation or retirement state
had to be added to every reader. structured-worker-custody now derives both answers from the
worker-terminal list state coordinators see in worker-list: addressable is owned and not released,
and owed work is an active custody or an unsettled task dispatched to the same incarnation. The
owner's state is read through the remote dispatch attachment too, as the terminal transfer lookup
already does. Behaviour is unchanged; a settled worker awaiting its coordinator still rests.

* refactor(orchestration): owed work is an open dispatch on the worker's incarnation

A supervised worker's own dispatch context stays open exactly while the worker is active, so the
separate active-custody branch only repeated it. Owed work is now one fact, which also states the
policy that a worker awaiting its coordinator's decision may rest, and both custody decisions are
written once at the top of the module.

* docs(agent-status): a departed agent's failure ranks below live work on the worktree card

* fix(native-chat): a restart offer knows its continuations by a tag in their id

The offer recorded each continuation id in a list on its capsule entry, capped at 16, and a running
action's id in memory. Both could disagree with the journal: past the cap an old rejected
continuation read as the chat moving on, and a crash during a retry restored the failure's older
entry, which lacked the retry's id. Each continuation id now carries a tag derived from the offer
(its teardown and chat), then the action's own part, so any continuation of this offer, queued or
rejected, is recognised from the journal row and the marker alone. The persisted list, its cap and
the in-memory action map are deleted; the agent-start withdrawal keeps an offer whose own
continuation the start was for, read against the stored marker.

* test(runtime): the legacy-worker reveal test judges its stale snapshot inside the wait

The tui-idle probe reads through readTerminal, which now awaits the structured
worker check before the PTY read, so the probe's snapshot request starts a
microtask later. vi.waitFor missed it on its first check and polled again at
50 ms, the same moment the wait's own 50 ms timeout fired. The stale snapshot
then resolved after the wait had already timed out, so the test passed without
judging it, and the rejection landed before any handler was attached. Vitest
reported that as an unhandled error and failed the shard.

Polling every 1 ms sees the request within a few ms, so the snapshot is judged
while the wait is still pending.

* fix(native-chat): a message held behind /compact is drawn where it was handed over

A message typed while /compact runs was drawn above the compaction's result, between
itself and its own answer. The reducer kept every item at the sequence and timestamp of
the row that created it, and a queued message is created at acceptance, long before the
command it waits behind writes its result. The phone orders by that sequence and the
desktop by that timestamp, so both put the message first.

A queued message now takes its position from its handover row, the same row that already
states its turn scope. Everything the agent did before the handover, a command it waited
behind included, draws above it. This holds for every held message, not only /compact's,
and needs no client change: every client, older builds included, reads the position the
host publishes. A live batch already carries the item when its dispatch row lands, and
history pages cut the reduced timeline by sequence, so paging stays contiguous.

* fix(native-chat): a phone's send during /compact answers without waiting out the compaction

A client that predates accepted-send replies, which is every phone build, has its send
reply held until the host hands the message over. A message sent during /compact is not
handed over until the compaction ends, so the phone's 15 s request timeout fired first
and showed the message as unconfirmed.

That wait now also ends once the message is queued behind a running command. This is
read from the journal's running turn and needs no new state. Every other wait still
ends at the handover: behind a starting child or an ordinary turn, and for restart
resume, the command front door and orchestration, which keep the plain handover point.

* perf(native-chat): a rewind places provider items with one pass over the merged rows

A Codex rewind gives each provider item the old epoch never held the turn record for its
provider turn. It found that record by scanning every merged row, restoring each row's
body, once per provider item. That is quadratic, and it runs on the host's main thread
up to the journal's 10,000-row cap, twice per rewind. A rewind record written before
rows carried their scope holds no scope for any provider item, so it paid the full cost.

The merge now indexes turn records by provider turn id once, keeping the first match as
the scan did, and each provider item looks its record up.

* fix(native-chat): a view never restarts a chat whose last start failed

A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.

* fix(native-chat): a message waiting behind /compact is drawn after it until it is sent

A message sent while /compact runs is placed where it was handed over. It was still
drawn where it was accepted until then. /compact writes its result one step before the
handover, so for that step the waiting message sat above the compaction's separator.

A message the host accepted but has not handed over is not part of the conversation
yet, so both clients now draw it after everything the agent has done. The shared
projection moves it to the end, which is the order the phone draws. The desktop ranks
it with the other not-yet-sent rows, after the streaming preview. At handover it takes
its place from its handover row, which is also after the separator, so it never
appears above the compaction it waited for.

* fix(native-chat): the idle sweep reads owed work every tick

Owed work counted as activity, but the sweep read it only once the idle window had elapsed, so it
refreshed the clock at most once a window. Work that ended just before the next read left the
agent to be stopped at that read, moments after the work ended, which is the gap the refresh was
meant to cover. The sweep now reads owed work on every tick for a started agent, so the window
always runs from the last tick that saw work owed.

* fix(native-chat): a continuation handed to the agent stays sent

The offer read its own continuation as not reaching the agent while its dispatch was pending, which
also covered one already handed over and still unanswered. When the wait for that answer ended first,
the failure it filed read as retryable, and a retry sent a second continuation to an agent that may
have acted on the first. Only a continuation still queued, or rejected, is now read as unsent.

* test(native-chat): start the child the loop waits on with an attach, not a second view

A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.

* fix(native-chat): settle a gone generation's turn wherever a conversation opens

A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.

* test(native-chat): prove the next child's start settles the turn an earlier child left

The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.

* test(native-chat): count a failed start's rows by row, not by text

Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.

* test(cross-version): load the phone row readers without mobile's toolchain

Vite transforms a file against its nearest tsconfig, and mobile/tsconfig.json
extends expo/tsconfig.base.json, which the root-only cross-version lane never
installs. The worktree ps verdict suite imported the current phone row reader
from mobile/ directly, so CI failed with TSConfckParseError before any test ran.

The harness now imports a copy of the working-tree reader placed under the
checkout cache, where the root tsconfig applies, as it already does for the
release checkout's copy. Both readers are still the real files.

* test(cross-version): keep the checkout path-guard message and justify the copy import's cast

* fix(native-chat): a command ends only by its own provider answer or its child's end

Stop no longer settles a conversation command. It interrupts it like any turn,
and when the provider cannot take that (Codex has not opened the command's turn
yet, or Claude refuses the interrupt) it stops the child, whose dead-generation
settlement writes the verdict.

The pending command now lives on the provider child's own session instead of an
adapter-wide map keyed by session, so it dies with the child and nothing has to
release it. Claude's /compact is sent under a uuid the slot records, and only a
root result naming that input (or naming none) ends it; its outcome is read with
the ordinary result reading, so a stopped /compact is a cancellation.

* fix(native-chat): a command's settle answers its message before ending its turn

The two writes are not one batch. Writing the message's answer first means a
crash between them leaves a running command turn, which the stale-turn sweep
already settles, instead of an ended turn whose message reads as in flight
forever. The settle now writes only while the command turn is still running.

* fix(native-chat): "Worked for" counts from the handover, not the send

A message held behind /compact, or behind a cold start, used to count the wait
as the agent's work, although its row is drawn at the handover. Every handed-over
submission's turn, the command's own included, now starts at the handover row's
instant, falling back to the send time for a host that recorded none.

* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget

* test(native-chat): the interrupted create's own retry continues again

The merge of main's lease-latch fix replaced that test's retry of the interrupted create, under its
own operation id, with a fresh start whose result nothing read. That fresh start passes with the
released-reservation continuation deleted, so the case the fix exists for went untested. The retry
and its assertion are main's again.

* docs(native-chat): three comments that still had views starting agents

A start with nothing queued now comes from a command, goal change or rewind; an interrupted compaction
left alone would refuse every send, so no agent would ever start to finish it; and a current host
raises the unattached read refusal only once quit began, with the attach window belonging to an older
host.

* test(native-chat): pin the open's and the send's start and row counts, however the view binds

Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.

* fix(native-chat): a second Stop on a command ends its child; one compaction verdict for every provider

A Stop's note now names itself in its key, so a later Stop on a command still
running reads, from the journal, that the provider was already asked and never
answered, and stops the child instead of interrupting again. Nothing is held in
memory for it.

Adds the rule both translators will read a compaction's end by: only a
compaction the provider reported is a success; none after Orca's interrupt is a
cancellation; anything else is a failure. A real Claude capture, pinned as a
fixture, is why: a stopped /compact ends in the same success result as a
finished one.

* test(native-chat): a reader's open settles the turn a failed exit settlement left running

An exit whose settlement write failed leaves its turn running in the open journal. PR 1's open now
settles it, and this pins the two reads that reach it here: a reader reopening a chat the idle
sweep closed, and a read that opens the chat before the restart restore reaches it.

* test(native-chat): the view-start test's starting window outlasts two subscriptions on a loaded runner

A subscription reads the conversation before it returns, so under load the two views took longer
than the create child's 300 ms start, which then exited before the test checked that it had not.
The child now takes a second to fail.

* fix(native-chat): settle a gone generation's turn at every open but an acquisition's

The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.

* test(native-chat): hold the create's start open until the views bind

The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.

* refactor(native-chat): the provider's translator ends a command's turn; the loop holds no command state

A conversation command is now a turn of the provider child's own journal
pipeline. The adapter-wide tracker, its promise and the loop's settle step are
gone.

- Codex: the translator claims the provider turn that carries the command, scopes
  its rows to the command's turn, and writes the command's end in the same batch
  that settles that turn. Codex's own compaction marker is the success row.
- Claude: the command's turn is the translator's open turn until the result that
  answers the /compact input ends it. The command's own frames, such as the
  continuation summary, its echo and "Compaction canceled.", draw nothing.
- Both read the end with the one compaction rule: success needs the provider's
  report of the compaction; none after Orca's interrupt is a cancellation.
- The message resolves at the provider's receipt, as any send does: the Codex
  ack, or the Claude slash-command waiter on its result. The host writes a
  command's end only when the provider never took it.
- The delivery loop stops while a command's turn runs, and every journal commit
  re-wakes it through the session's serialize, so an end that lands while a step
  decides to stop is never lost. A child that ends first is settled with it.

* test(native-chat): pin a command's end to real /compact frames and to each path it threads

The captured /compact frames drive the Claude translator's command turn: a
finished compaction ends as a success with only the separator drawn; a stopped
one ends as a cancellation with no failure row, and the next send answers in its
own turn; a result naming another input ends nothing. The command's end is
checked at each point the ordinary result path threads through: the reopen latch
after a failure, the settling of a child still working, the context facts the
result reports, and the provider's own error row.

On the host: a message held behind a command is handed over when the command
ends just as the loop stops for it, a refused command settles as a failure and
the loop moves on, and a Claude child that exits mid-command settles the command
and hands what waited to a fresh child.

* test(native-chat): tests merged from the base state which turn their rows belong to

* refactor(native-chat): drop the child-end waiter nothing waits on

A command no longer waits for its child here: its turn ends from the provider's frames or from
that child's settlement, and the delivery loop is woken by the commit. The waiter and its test
were left from the earlier shape.

* fix(native-chat): a command holds the queue only while its child runs it

The delivery loop stopped whenever the journal showed a command's turn running. When the
command's child ended and its settlement could not be written, that turn stayed running with
no child to end it, and the loop's gate kept it from ever starting the next child, which is
what settles a gone generation's leftovers. Every later send was held for good, and Stop had
no child to end.

The gate now holds only while the conversation has a child: with none, the command belongs to
a gone generation, and the loop's start settles it like any turn a dead child left running.

* fix(native-chat): a Claude /compact succeeds only on its compaction boundary

The command's evidence counted Claude's `compact_result: 'success'` status as the compaction
done. That status comes before the boundary that replaces the history, so a Stop landing
between the two read as a finished compaction even though no boundary was ever written. Only
the boundary now counts, as the rule for both providers states; the capture's finished
compaction carries one, so it still reads as a success.

* fix(native-chat): a Claude child's exit says why the turn it ended stopped

When a Claude child exited mid-/compact, the command showed "Worked for 0s" and no reason. The
child's translator ends its open turn the moment the exit is reported, stamped with the exit's
instant, so by the time the exit settlement ran nothing was running. The settlement recognises a
turn the exit already ended by that same instant, but the Claude lifecycle event dropped it on the
way to the host, which then used its own clock, matched nothing, and wrote no row. When the clocks
did agree, the row was scoped to the running turn, of which there was none, so it landed outside
the turn it explained.

The exit's instant now reaches the host, and the exit row belongs to the turn the exit ended:
still running, or ended by the translator at that instant.

* fix(native-chat): a message waiting behind /compact draws below its live activity

A message sent while /compact runs waits on the host until the command ends. Both clients moved
it to the end of the transcript rows, but the running turn's live activity line ("Compacting the
conversation") draws after every row, so the waiting message sat between the command and its own
live status.

A row that is queued, and not what the live turn is for, now draws after that live activity: on
desktop outside the transcript window, below the activity line; on the phone in the list footer,
below the live status. A message whose own start is pending still draws above the activity that
start reports.

* fix(native-chat): only a running command holds a message below its live activity

A message is accepted, then handed over a moment later, and in between it reads as waiting. Every
message waiting behind a live turn drew below that turn's activity line, so an ordinary message
sent while the agent was working crossed below "Thinking" and jumped back up once it was handed
over, on desktop and phone. Only a conversation command's turn holds the queue on the host.

A message now waits below the live activity only while the running turn is one a command opened,
read from the entry that opened it. The phone test also typechecks, which the mobile test ratchet
requires.

* test(codex): the claim test names its notification params as a record

* test(native-chat): a read that reaches a crashed chat before the startup reconcile settles its turn

On desktop the chat on screen at relaunch reads before startup reconciles the leases, while the
dead process's lease still reads live. The open settles the turn it left running anyway, and the
restore that follows finds it settled.

* refactor(native-chat): drop the composer's second error formatter

After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.

* test(native-chat): pin the reason on a message rejected while its chat was closed

The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.

* docs(native-chat): drop the removed dispatch hold from six comments

A worker's session no longer takes a dispatch hold, and no release clock
rests a chat by visibility; the agent-launch comments, the abandon test,
the teardown test and the refusal census still said so.

* test(native-chat): rest the owner-status chat through the idle sweep, not a hold

The activation-gate test from #22808 put its chat at rest by holding and
releasing it, and passed the release-clock grace. This branch deleted both,
so the case threw before it reached its assertions. It now moves the host's
clock past the idle window and lets the sweep stop the agent and close the
conversation, then asserts the same owner answer and activation gate.

* fix(native-chat): show the structured pane's retrying line when a read fails

The read transport always hands the pane the host's words, so the error
state's "Orca keeps trying to load it" line, which showed only when there
were none, was never seen: the pane showed the host's text twice, as its
subtitle and on the status line under it. The structured pane now always
says its read keeps retrying, and the host's text stays on the status line.
The terminal-backed chat is unchanged.

* fix(native-chat): a send the provider never received after a restart has no verdict

Restart reconciliation rejects a crash-stranded send that is absent from a
trustworthy provider history with reason 'not_delivered'. Nobody failed that
send, but the verdict allowlist did not name it, so after a crash the chat
read Failed, was listed, and could notify "failed". Give the reason a shared
constant (persisted value unchanged), add it to the no-verdict set, and treat
it as an internal marker so the Retry row no longer shows the raw string.

* fix(native-chat): a failed Codex compaction's late completion writes no turn of its own

Codex ends a failed turn with an error and then still completes it as failed.
The error settled the compaction and released its claim on the provider turn,
so the completion read that turn as an ordinary one and wrote a stray record.
The claim now lasts until the completion, which adds nothing to a command the
error already ended.

* test(native-chat): the mid-command exit case resumes its next child as a real one does

The case's fake started every child as a newly created thread with the same generation. The
store refuses a created link once the conversation has a thread, so the next child's start
failed and wrote its own error row, which landed before or after the case read the journal.
The next child now resumes the thread under its own generation, and the case reads the
journal once the waiting message is delivered, which also proves the loop moved on.

* test(native-chat): wait for a send's background start before the refusal oracle removes its store

An accepted send wakes the delivery loop, which starts the agent in the background. The oracle's teardown disposed the loop but did not wait for that start, so its lease write could create a temp file in the store directory while the directory was being removed, failing the test with ENOTEMPTY about one run in four. The teardown now drains tracked starts before it closes the journals.

* fix(native-chat): a start a message waited on gets one failure row, the delivery loop's

When a queued message's start failed, two writers could report it under the same row: the delivery loop, when the adapter settled the start without proving it, and the exit settlement, when the child's exit landed. The last one won, so the chat's row could name a different cause than the one the message was rejected with, or be written twice.

The exit settlement now writes the start's row only when no message is queued and the loop has not already recorded that start. A start for a command, goal change or rewind, with nothing queued, still gets its row from the exit.

* fix(native-chat): a /compact whose start failed says to run /compact again

The failure-words context named only /clear as a command to retry, so a
/compact whose agent failed to start read "Send your message to try again."
on its row, its rejected message and the command reply. The context now
carries any conversation command; the host derives it from the oldest
message still waiting on the provider, which is the one a failed start
fails first, and the /compact reply names it directly.

* fix(native-chat): a Codex /compact ends only on its turn's completion, below Codex's own error row

Since only turn/completed ends a Codex turn, Codex's turn-ending `error` is a row
inside the still-open command turn, and the failed completion that follows it is
the command's end: completed, outcome failure, at the completion's receipt time.
The command's own "Compaction failed" row was written on that completion too, so a
failed /compact read its reason twice.

The command turn now notes when Codex's turn-ending error for the turn it carries
was written as a row, and its end then adds no second row. A retried stream error
ends nothing and is not counted. The flag that let the error end the command and
kept the claim until the completion is gone with the error-driven end.

A test replays the captured failed compaction from the real app-server through a
claimed command turn.

* test(native-chat): main's crash-turn test states its row's turn, and a dead /compact settles on its recorded exit

Two tests the main merge brought together:
- The crash-turn test from #23456 writes a turn record through the event sink
  without options; every row here states its turn scope, and a turn record's is
  the thread.
- The /compact whose exit settlement could not be written no longer stays running
  until the next start: main now settles an open chat from the exit it recorded, so
  the command reads interrupted before the next message, which is then delivered.

* test(native-chat): main's new journal tests state each row's turn

The crash-turn, stale-turn and sink-queue tests main added wrote rows without a
turn scope, which every item write now states. Rows written inside a running
turn name that turn; the sink-queue batch and a send handed over with no live
turn name the thread.

* fix(native-chat): draw a queued turn's message after the earlier turn's rows

A message sent while A runs is written to the journal when it is sent.
When the provider queues it (Claude answers it after A), A's remaining
rows - its last tool run and its answer - are written after that
message, and the message's own turn opens only after them. Grouping put
those rows in A's turn, but the transcript still drew them in journal
order, below B's bubble and bar, where A's answer read as B's reply. This
is the residual #23671 left open.

A message that opened a turn now draws after the earlier turns' rows the
journal wrote after it, just before its own turn's rows
(nativeChatTurnDrawOrder, returned by nativeChatTurnMembership as
drawOrder). Desktop and mobile both draw in that order. A steer, and a
message that has opened no turn yet, stay where they were written. It
applies on hosts that state turn scopes and, through journal order, on
older ones.

* test(native-chat): run #23026's Stop tests against #23059's command turns

Two of #23026's tests call APIs #23059 changed, and failed after the
merge:

- codex-structured-conversation-stop: a compaction now goes through
  adapter.compact with the command run the host wrote (#23059), not a
  bare turn id, and answers with the provider's receipt. With the command
  claimed, a Stop that names no turn while the compaction's provider turn
  has not opened still interrupts nothing.
- main-agent-working-agreement: a provider row states its turn scope
  (#23059's appendItem contract); the retry and subagent rows are
  conversation-scoped.

* fix(native-chat): typecheck main's Stop and restore-grouping code against #23059

A Stop's compaction interrupt reads the narrowed requested turn, and the
restore-grouping test states whether each row reports its turn's outcome.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-29 13:17:03 -07:00
Neil 31012aeb09 test: remove assertion-free probes, copied inventories and export-shape checks (#23816)
Second audit wave, targeting three more junk patterns:

- assertion-free cases that run code and assert nothing, so they pass no
  matter what the code does;
- inventory literals re-typed from a production declaration, where the only
  way the assertion can fail is someone editing one of the two copies;
- export key-set and export-shape loops (`typeof x === 'function'` over every
  export) that restate what TypeScript already enforces.

Yield is much smaller than wave 1 on purpose: the assertion-free scanner has
a high false-positive rate, because many flagged blocks assert through a
shared helper or their oracle is "this must not throw". Those were kept.

`mobileWebCheckArgs` in `config/scripts/run-mobile-web-app-checks.mjs` is
de-exported — after the inventory comparison went away, nothing outside the
module read it.
2026-09-29 02:21:47 -07:00
Brennan BensonandClaude c49388cd33 fix(native-chat): Stop is there from the moment a message is sent (#23026)
* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* fix(native-chat): name a chat write by its target, not the owner generation

A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.

Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.

Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.

* fix(native-chat): every journal append reaches the chats that are open

A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.

A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.

* test(native-chat): an epoch replacement reaches the open chat

* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

* perf(native-chat): a publish behind a delivered commit reads nothing

Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.

* test(native-chat): state why the teardown test's fake journal is safe to cast

* docs(native-chat): say mutation admission checks only the writer lease

* docs(native-chat): drop the send rebase from comments that still described it

* fix(native-chat): a message is accepted, then delivered

A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".

A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.

Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.

A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.

Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.

* fix(native-chat): settle queued messages only for the child that ended

A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.

A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.

The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.

* fix(native-chat): an adoption that fails to import keeps the conversation open

The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.

* perf(native-chat): the recovering open reads the journal once

Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.

* fix(native-chat): an attach that fails after indexing its child leaves no child behind

A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.

* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer

The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.

A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.

* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down

The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.

* fix(native-chat): a message rejected while its chat was closed reads as not sent

A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.

* test(orchestration): name why the readiness settlement fakes are cast

* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent

* docs(native-chat): drop the fence from the admission the send effects run behind

* docs(native-chat): give the fence move on release the reason that still holds

* docs(native-chat): stop citing a write fence check in launch and mailbox comments

Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.

* refactor(native-chat): the provider child is its own record

A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.

- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
  patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
  own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
  the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
  dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
  is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
  the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.

* fix(native-chat): the delivery loop alone settles a message its start or child failed

A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.

- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
  reads how it ended: a Stop continues; anything else writes one failure row and rejects every
  queued message with the same words, then stops. A child still starting whose start the adapter
  says did not land fails the same way. The exit, eviction and the settlement retry only settle
  the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
  nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
  failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
  closed, with or without a child, and a start the loop already has in flight is waited for so the
  child it produces is stopped rather than left behind.

* refactor(native-chat): a stopped child ends on the one reading of its stop

The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.

* feat(native-chat): the host says it accepts a send before any agent has it

The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.

* refactor(native-chat): an attach never opens a journal of its own

The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.

* fix(native-chat): a moved fence resends nothing on a host that accepts first

The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.

The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.

* refactor(native-chat): a child's end says whether the user or the host stopped it

The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.

* fix(native-chat): a chat whose only work is a queued message is not offered for resume

A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.

* test(native-chat): type the queued-message fixtures in the resume-offer tests

* fix(native-chat): a start that dies while a message waits on it is that message's failed start

Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.

* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now

* test(native-chat): pin what a failed start settles, and what a resume offer names

A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.

* test(native-chat): the failed-start pins fail on what the message became, not on a timeout

* test(orchestration): the preamble's host stub is typed, not cast

The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.

* fix(native-chat): a Stop that names no turn stops what the conversation has in flight

Between handing a message to the agent and the agent opening its turn, there is no turn id a
client could name, so a Stop in that gap was refused as "already finished" while the agent went
on to answer. A cancel's turn id is now an optional precondition instead of its target: with
none, the host withdraws what is queued and, when the journal still reads working, asks the
adapter to stop whatever the child has in flight. Claude's interrupt is session-scoped, so it
is guarded by fence and acquisition generation rather than a turn identity. Codex interrupts
the turn its latest turn/start answered with until the journal shows one.

A cancel that names its turn behaves exactly as before.

* fix(native-chat): Stop is there from the moment a message is sent

The composer showed Stop only once the agent had opened a turn, so for the second or two after a
send the chat read "thinking" with no way to stop it. Against a host that takes a Stop naming no
turn, Stop now shows whenever the chat reads working (a turn, a queued message, or a handed-over
one still unanswered) or this client still has a message on its way. Pressing it, or Escape,
first drops every outbox entry the journal does not hold yet, so nothing goes out after the
Stop, then sends the conversation-wide cancel. A send already on its way reaches the host ahead
of the cancel, which withdraws it there. Against an older host Stop still needs a running turn.

The unconfirmed-send probe moves into its own hook so the outbox hook stays in budget.

* fix(native-chat): Stop before a turn is gated on its own host capability

A host that accepts sends first (agent-session.accepted-send.v1) can still predate the cancel
that names no turn and would refuse it as invalid, since clients and hosts ship independently.
Hosts that take that cancel now advertise agent-session.conversation-stop.v1, and the renderer
shows Stop before a turn opens, and sends the no-turn cancel, only to a host advertising it.
Every other host keeps a Stop that needs, and names, a running turn.

The host capability probe the accepted-send hook used is generalized so both read one path.

* test(native-chat): a build advertises conversation stop exactly where its cancel may name no turn

* fix(native-chat): a view never restarts a chat whose last start failed

A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.

* test(native-chat): start the child the loop waits on with an attach, not a second view

A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.

* fix(native-chat): settle a gone generation's turn wherever a conversation opens

A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.

* test(native-chat): prove the next child's start settles the turn an earlier child left

The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.

* test(native-chat): count a failed start's rows by row, not by text

Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.

* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget

* test(native-chat): pin the open's and the send's start and row counts, however the view binds

Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.

* fix(native-chat): settle a gone generation's turn at every open but an acquisition's

The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.

* test(native-chat): hold the create's start open until the views bind

The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.

* fix(native-chat): Stop reads the one working rule every session list reads

While Claude retries a rate-limited request it never echoes the message, so no
turn opens: the sidebar read Working from the unanswered send while the composer
showed Send. The chat's working state, the host's session-list status and the
host's no-turn Stop check now call one shared rule instead of three copies.

* test(native-chat): a rate-limit retry pins only that no turn opens, not how its rows are kept

* fix(native-chat): Stop leaves a message waiting on its Retry, and does not show for one

A send that failed holds the queue until the user retries it, and one the host restarted under is
parked the same way. Stop counted both as still on their way, so it showed in an idle chat and
could never go away, and pressing it dropped the failed message along with its Retry.

* test(native-chat): the chat's Stop and a session list read the main agent alike over their own copies

The chat reduces its stream and a list reads the status feed. Driven through the real host for a
rate-limit retry with no turn, a subagent still running after the main turn, and the handed-over
child exiting.

* refactor(mobile): the chat reads the main agent's working state through the shared rule

Behaviour is unchanged: the same two terms, now from the one function the host projection and the
desktop chat read.

* fix(codex): a Stop naming no turn never interrupts an earlier turn

It fell back to the id an earlier turn/start answered with when the latest start went unanswered,
or when the journal showed a compaction Codex had not started, and reported that as stopped.

* fix(native-chat): a Stop naming no turn never says a turn had already finished

When the provider found nothing left to stop, for instance a turn that ended between the host's
check and the interrupt, the chat got "The provider had already finished this turn." for a turn
the Stop never named. It now ends quietly, as a Stop with nothing in flight does.

* fix(native-chat): one Stop the host could not settle no longer refuses every later one

A Stop naming no turn has one operation key per session. When the host could not settle one, it
answered every later Stop under the same id as unknown until the id expired. Once the host says
so, the next press is a new Stop; transport doubt still replays the same id.

* refactor(native-chat): drop the composer's second error formatter

After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.

* test(native-chat): pin the reason on a message rejected while its chat was closed

The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.

* test(native-chat): read Stop operation ids without a cast

* fix(native-chat): a Stop whose answer was lost no longer swallows the next one

A Stop that names no turn has one operation key per chat. When its answer was lost in transit, the
chat kept the id, so every later Stop replayed it; the host answers a replay as already handled, so
for up to a day Stop stopped nothing. The id is now dropped once the call settles, however it
settles. A second press while the first is still on its way still shares its id.

* refactor(native-chat): a Stop naming no target keeps its operation id only for its own call

The chat kept each write's operation id per payload across calls, and dropped it only on some
settle paths. That is right for a write naming what it acts on, but a Stop naming no turn, and a
stop of every background task, share one payload with every later one, so any path that kept the id
made the next Stop replay as already handled and stop nothing. One path was still open: an answer
that arrived after the chat moved to a new fence.

Whether a write names its target is now decided once, before its id is picked. One that names none
keeps its id only while its call is in flight, so a press made meanwhile joins it, and releases it
when the call settles, however it settles. The release runs only while the key still holds that
call's id, so a joined call settling late cannot drop a newer one's. This replaces the per-path
exceptions for a thrown call.

* test(native-chat): read the Stop fences without a cast

* test(native-chat): pin the new id for a named cancel the host could not settle

After the Stop naming no turn moved to a per-call id, the only test of the unknown-refusal release
was gone, and the half that stays, for a cancel naming its turn, could be removed with every test
green.

* fix(native-chat): a Stop pressed after a new message stops it, even while the last Stop is unanswered

A Stop naming no turn shared its operation id with any press made while it was still in flight. The
host runs a chat's writes in order, so a message sent between two presses was accepted after the
first Stop ran, and the second press replayed that Stop as already handled and left the message
running, although the chat had already withdrawn it from the outbox.

A write naming no target now gets a new id on every press and is never kept, so each Stop acts on
whatever is running when the host reaches it. A write naming its target keeps its id exactly as
before. A double press can ask the provider to stop the same turn twice, which it tolerates.

* fix(native-chat): Stop no longer blinks off as Claude opens the turn for a message

Claude's echo of a sent message both answers the send and opens its turn. The echo settled the send
first, so the host published the message as answered one frame before the turn it opened, and for
that frame the chat read nothing running: Stop turned back into Send, and Working blinked off in
every session list, for tens of milliseconds on each turn.

The echo now settles the send after the turn it opens has been emitted, so the running turn is
published first.

* fix(native-chat): a message a Stop withdrew comes back to its sender's composer

A Stop withdraws every message the host holds but has not run, and S also
drops the ones this client had not handed over yet. Either way the message
left the chat and its text survived only in a hidden journal row and the
in-memory ArrowUp history.

The sending client now puts the withdrawn text and images back in that
pane's composer, after whatever is typed there. Withdrawn is read from the
rejection reason through one shared check, which the outbox reconcile now
uses too. The composer is written before the entry leaves storage, so a
failure between the two repeats the text instead of losing it, and an entry
storage no longer holds is never given back again, so a replay, a second
view or a remount restores it once. Only this client's outbox holds the
entry, so other viewers still see the message disappear. A failed Stop
withdraws nothing on the host and gives nothing back.

* fix(native-chat): withdrawn text put back during an IME composition is not lost

While the IME owns the field, the composer ignores a programmatic draft, and
the next composed keystroke wrote the draft without the restored text, after
its outbox entry had already been dropped. The composer now holds text
appended mid-composition, keeps it in the cache after each composed write,
and shows it once the composition settles, the way attachments that land
mid-composition already wait for it.

* test(native-chat): pin that only a withdrawn message comes back to the composer

* test(native-chat): set up the composer's window API for every describe in the composition-race file

* docs(native-chat): note that the withdrawn check reads the legacy reason until a typed category lands

* test(native-chat): pin that text put back mid-composition shows once, even beside a mid-composition clear

* fix(native-chat): land a late settlement from a streamed turn's end after that turn's rows

A settlement that says a streamed turn ended waits for the session's event
sink to drain before writing its dispatch row. The journal reducer still
refuses to overwrite an accepted or rejected send.

* fix(codex): settle a send from the end of the turn Codex answered it into

The turn/start answer names the turn that holds a send. The send's echo
entry now keeps that binding, in memory only. If the bound turn is
interrupted without echoing the send, the send is withdrawn: Codex clears a
turn's pending input on interrupt, so the model never saw it. If the turn
fails first, the send is rejected in Codex's words. A completed turn settles
nothing, since Codex records pending input when it finishes and the echo is
still due. An answer read after its turn already ended is settled by that
end. The echo is still the acceptance and carries the item key.

* test(codex): a send settles from the end of the turn Codex answered it into

The fake Codex keeps 0.157's turn bookkeeping, and can deliver the turn/start
answer after turn/started or after turn/completed. The tests cover:
- a Stop before any echo withdraws the send, and the working rule reads idle;
- a steered follow-up is withdrawn when the turn is interrupted;
- a failed turn rejects the send in Codex's words;
- a completed turn leaves the send to its echo;
- a normal echo and a late echo;
- two steered sends in one turn;
- an answer read after the turn ended;
- a timed-out answer;
- child-thread turns;
- how a binding dies.

* refactor(native-chat): drop the stream flush before a late turn-end settlement

Nothing reads the order of a dispatch row against the turn's terminal row:
the reducer keeps a settled send terminal and the working state is derived
from both. The echo acceptance on the same path never waited either, and the
wait could drop the settlement on a failed sink barrier.

* fix(codex): settle a failed turn's sends at its end, not at its error

Codex keeps a failed turn's pending input and records it after the error
frame, before turn/completed. Settling at the error rejected a steered
follow-up the model had in fact received, so a Retry would send it twice.

* docs(codex): say a completed turn echoes what it took before it ends

Codex records a completed turn's pending input before `turn/completed`, so
a bound send that turn never echoed is left for recovery, not awaiting an
echo. The comments and one test title said the echo was still due.

* test(codex): settle a send whose answer is read after its turn failed or completed

A failed turn that ended before the answer rejects the send in Codex's words,
once; a completed one leaves it admitted and still armed for its echo.

* refactor(codex): read a failed turn's reason with the typed thread-fact reader

* fix(native-chat): a Stop that names no turn and ends nothing says why

The host sends a Stop naming no turn to the agent only while the chat reads
working. When the agent ended nothing, the Stop wrote no row, so it looked
ignored. It now writes one: Stop could not reach the agent, in the agent's
own words when it refused the interrupt.

* fix(codex): hold a cold send until Codex opens its turn, and stop the turn it opened

Codex answers turn/start before it opens the turn, and refuses an interrupt
until then. A Stop queued behind a cold send ran in that gap, named the
answered turn, and was refused. The send's handover now lasts until Codex
opens that turn, or provably will not: the turn ended, the primary thread
stopped running, or the child ended, bounded by the turn/start deadline.
A steered send, whose turn is already open, does not wait.

A Stop naming no turn now interrupts the turn the journal shows, else the
primary-thread turn Codex reported started and not yet ended. The per-start
answered id is gone: it was never cleared at a turn's end.

* fix(native-chat): give back a send already on its way only when the host withdraws it

Stop took the in-flight send out of the outbox and put its text back in the
composer at once. The send still landed ahead of the Stop, so the chat read
working for a moment before the host withdrew it, and a send the agent had
already taken came back as well. The in-flight send now stays until the host
answers it, and the withdrawn-message restore gives it back from that answer.

* refactor(native-chat): read a Stop's withdrawal through the rejection classifier

dispatchWasWithdrawn matched the legacy reason string. It now asks the
classifier, which reads the typed fact first and keeps that string only as its
own fallback, so a withdrawal written as a fact with a sentence is still given
back to the composer.

* fix(native-chat): a Stop Codex took but Orca could not confirm is not reported as reaching nothing

When Codex acknowledged the interrupt but Orca could not verify the turn's
processes ended, a Stop naming no turn wrote "it had no turn running to stop",
though the turn then ended as interrupted. The adapter now says the Stop was
taken but unconfirmed, and the row says Cancellation was not confirmed.

* fix(codex): a held cold send never delays closing the chat or quitting

A cold send's handover waits for Codex to open its turn, and that wait sat
in the session queue ahead of the close and quit eviction, so either could
wait out the 30 s request deadline. The host now releases those waits before
it queues a close or starts quit teardown, and the adapter releases them when
it is asked to close the child and on every exit, including one whose end
publication is backpressured.

* fix(native-chat): a refused Stop says the agent declined, and names it

"Stop could not reach the agent" was wrong: the agent was reached and
declined. The row now reads "Codex had no turn running to stop." or
"Codex didn't stop: <Codex's words>.", naming the chat's agent.

* fix(codex): a Stop waits for the turn Codex answered to open; the send no longer does

Codex answers turn/start before it opens the turn and refuses turn/interrupt
until then, so a Stop naming no turn in that window was lost. The send's
handover used to wait for the turn to open, and a close or quit needed its own
release to get past that wait.

Now only the Stop waits. A Stop naming no turn, finding no journal turn and no
open one, reads the turn Codex answered the latest pending send into and has
neither opened nor ended, and waits for it: bounded at 5 s, under the quit
eviction budget, and ended by that turn opening or ending, the thread going
idle or failing, or the child exiting. If the turn opened it is stopped;
otherwise the Stop reports that Codex had no turn running. The send returns at
Codex's answer, so a close or quit with no Stop pending is never delayed, and
the release plumbing through the adapter, router, close and quit is gone.

* fix(codex): the Stop's wait reads a stopped thread from the thread-facts reader main kept

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-29 01:39:49 -07:00
Neil 6e1b7e7fa3 test: remove junk tests that assert source text instead of behavior (#23815)
Deletes 101 test files and trims 112 more, all matching documented junk
patterns: exact source/import/string greps, copied inventories and export
lists, duplicate invocations of a contract another test already owns,
typeof-shape checks TypeScript already enforces, and self-comparisons.

The largest group read a production `.ts` file and asserted on its text —
for example a TaskPage test that required the source to contain
`selectedRepos.find((r) => r.id === newIssueRepoId) ?? selectedRepos[0] ?? null`.
Any behavior-preserving rename broke it; no behavior change ever did.

Production-side follow-through: exports that only these tests imported are
de-exported or deleted, stale comments pointing at removed censuses are
dropped, and the reliability-gate registry, `cloud/package.json` test lists,
and orphaned source-reading helpers are updated so nothing references a
deleted file.

Two files kept their real coverage and lost only the census scaffolding:
`agent-status-producer-census.test.ts` now drives all five producers end to
end instead of grepping the source tree, and `config-toml-trust-stale-writes`
replaces an export-list parity check.
2026-09-29 01:21:53 -07:00
Brennan Benson 357a2fed08 fix(native-chat): group chat rows by the turn that produced them (#23671)
* fix(native-chat): keep a turn's bar on the prompt that opened it

A message sent while a structured turn runs appears in the transcript at
once, so "the newest user message" is not the running turn's owner. The
live "Working for" bar moved to the mid-turn message and counted from the
earlier prompt's start, and a send queued behind the running turn counted
its wait twice: once in the previous turn and again from its own send.

Derive both from the host's turn records in one ordered pass:
- The running turn's bar belongs to the user message its lifecycle row
  names (resolved exactly as settled timing resolves it). A message sent
  mid-turn gets no bar until its own turn opens; a send folded into the
  running turn never gets one. Surfaces fall back to the latest user
  message only when the host names no opener.
- A turn counts from its send, but never before the previous turn in the
  journal ended (its recorded end, else its row's last host revision),
  capped at the turn's own start. The same origin feeds the live counter
  and the settled duration.

Desktop and mobile share the derivation; no wire, host, or storage change.

* feat(native-chat): derive each transcript row's owning turn from the journal

Rows between a turn record and the next belong to that record's turn, so a
message the provider folds into a running turn no longer captures the rows
produced after it. A turn whose opener the host cannot name in the loaded
window anchors to its own record instead of a bystander prompt, and shared
nativeChatRowTurnKeys keeps positional preceding-user grouping for anything
the host does not attribute (older hosts stay pixel-identical).

* fix(native-chat): fold and time transcript rows by their owning turn on desktop

A settled turn's bar now folds every row the turn produced, including rows
after a mid-turn send; the steered bubble stays visible and carries no bar. A
provider-opened turn renders its bar above its first row instead of borrowing
the newest prompt, and row liveness follows the owning turn rather than the
newest user message, so a running turn's rows stay live while a send waits.

* fix(mobile): group phone transcript rows and bars by their owning turn

Same shared derivation as desktop: the opener's bar owns every row of its
turn across a mid-turn send, a steered bubble never grows a bar, a
provider-opened turn's bar sits above its first row, and a running turn's
tool rows stay live while a newer message waits behind it.

* fix(mobile): declare the turn ownership map on the chat controller contract

* fix(native-chat): one diff rollup per turn, and no wake-turn clock on a later prompt

A turn's rows are no longer contiguous once rows are grouped by owning turn: a
prompt sent before the running turn's last row lands among its rows. The diff
rollup was drawn at every run boundary, so such a turn showed its rollup twice
and the later prompt's rollup appeared under its bubble. It now renders once,
under the turn's last row.

On the phone, a turn keyed to its own record is not a user message, so when it
ended its clock was treated as a replaced optimistic echo and handed to the
newest prompt - a message sent during a wake turn got a bar with the wake turn's
duration. Host-attributed turn keys now count as live turn keys.

* fix(native-chat): keep a Codex turn's bar on its send until the echo lands

Codex reports turn/started before it runs hooks and prewarm and before it
echoes the send, so for that gap the turn names a provider key no alias
resolves yet. Anchoring it to its own record left the running turn with no
bar at all; treat the send still in flight ahead of the record as its opener.

* fix(codex): restore each turn's record ahead of that turn's items

Rows are grouped by the nearest turn record before them in journal order,
which holds on the live path because a turn's record is written when the turn
opens. Full-history restore (the fallback for Codex app-servers that reject
excludeTurns) wrote each turn's items first and its record after, so every
restored turn's rows were credited to the previous turn and turn 1's answer
folded away.

The restore now appends the record before the turn's items. The record itself
is unchanged: same identity, state, outcome, opener key, endpoints and
duration, one append each. The restore-order test now expects the record
first, because that order is what keeps grouping correct; its old order was
incidental, not a contract any reader relied on. Readers that key turns by id
or opener key, or that scan for the newest record (all restored records are
settled), read the same result in either order.

Journals already imported in the old order stay as written; no migration.
2026-09-29 00:55:45 -07:00
Brennan Benson 50a8ef18e4 fix(native-chat): keep a turn's bar on the prompt that opened it (#23573)
A message sent while a structured turn runs appears in the transcript at
once, so "the newest user message" is not the running turn's owner. The
live "Working for" bar moved to the mid-turn message and counted from the
earlier prompt's start, and a send queued behind the running turn counted
its wait twice: once in the previous turn and again from its own send.

Derive both from the host's turn records in one ordered pass:
- The running turn's bar belongs to the user message its lifecycle row
  names (resolved exactly as settled timing resolves it). A message sent
  mid-turn gets no bar until its own turn opens; a send folded into the
  running turn never gets one. Surfaces fall back to the latest user
  message only when the host names no opener.
- A turn counts from its send, but never before the previous turn in the
  journal ended (its recorded end, else its row's last host revision),
  capped at the turn's own start. The same origin feeds the live counter
  and the settled duration.

Desktop and mobile share the derivation; no wire, host, or storage change.
2026-09-28 21:22:22 -07:00
Brennan Benson c5fc0c6f26 fix(ci): keep a squash-merged RPC recording pin reachable through its pull request (#23720)
* fix(ci): keep a squash-merged RPC recording pin reachable through its pull request

Main's "RPC recording pin" check has been red since #22762: that branch pinned
the recording corpus to its own commit 03995ae, and the squash-merge left that
commit out of main's history. Every behaviour-change squash did the same, and
each needed a hand-made repin PR to clear it (#23565, #23535, #23046 and more).

The guard now accepts a pin that is either in this history or in the head of the
pull request whose squash wrote it into the manifest. It finds that pull request
from the `(#n)` subject of the commit that added the pin and fetches
`refs/pull/<n>/head`, which GitHub keeps after the branch is deleted. The
reproduce step uses the same lookup, so it can still check the pinned tree out.

* fix(ci): give the recording pin lookup room to walk a blobless clone

In CI's blobless clone, `git log -S` fetches the manifest's blobs one commit at a
time, a few seconds each. Under the 30 s process default the walk was killed after
a handful of manifest commits, which main's history already exceeds (up to 7
manifest commits between a pin landing and the next pin change), and the guard
then failed with an empty "Could not find the commit that pinned" error. The
lookup and the pull request fetch now carry explicit budgets and say when they
timed out.

The not-an-ancestor instruction now names the pull request whose head was
checked, or says the commit that pinned it names none.

Adds the two merge-preview shapes the guard runs on: a branch opened after a
squash resolves main's pin through the squash's pull request, and a branch whose
rebase dropped its own pinned commit fails on its pull request instead of on main.

* fix(mobile): tell a missing recording pin apart from product drift

After a squash the pinned commit can live only in its pull request's head, so a
clone that never fetched it makes `git diff --quiet <baseline>` exit 128. The
recorder reported that as "Product sources or lockfile differ from the pinned
main baseline", which sends the developer to repin a tree that may match. It now
prints git's error and the command that fetches the pin.

* fix(ci): ask GitHub which pull request holds a squash-dropped recording pin

The recording pin guard found the pull request that keeps a squash-dropped
pin by walking main's first-parent history for the commit that wrote the pin
into the manifest and reading "(#n)" off its subject. A merger who edits the
squash title loses the number, and the push to main turns red anyway. That
already happened on main: of the 22 squashes that left a pin outside main's
history, #21674's title had no "(#n)".

The guard now asks GitHub for the pull requests associated with the pinned
commit (GET /repos/{owner}/{repo}/commits/{sha}/pulls) and, for each in turn,
fetches refs/pull/<n>/head and accepts only when git proves the pin is an
ancestor of that head. GitHub only nominates candidates, so a wrong answer can
fail the guard but never pass it. The endpoint named the right pull request
for all 22 historical cases, #21674 included, and names none for commits a
force-push orphaned.

This removes the pickaxe walk, its 600 s budget and its lazy blob fetches in
a blobless clone, the first-parent subtlety, and the subject regex. A revert
that restores an older pull-request-only pin now resolves too, because the
lookup is by the pin itself rather than by the commit that last wrote it.

CI passes the job token to both guard steps and grants the job
pull-requests: read. Local runs work without a token on this public repo and
send GITHUB_TOKEN or GH_TOKEN when set. A failed lookup throws with the HTTP
status, and names the rate limit when an unauthenticated call is refused.
2026-09-28 21:17:07 -07:00
Brennan Benson e899809ff8 fix(mobile): unsubscribe session tabs by request on the direct connection (#22943)
* fix(mobile): unsubscribe session tabs by request on the direct connection

* fix(mobile): hold a direct session tabs unsubscribe until the first snapshot

The desktop registers a tab-list stream only as it emits the first snapshot. A direct
unsubscribe sent before that found nothing, and with per-request addressing no later
worktree-wide sweep collects the late stream, so it kept its desktop listener until the
socket closed. Hold the unsubscribe until the snapshot arrives, as the relay connection does.

* test(mobile): cover a held session tabs unsubscribe whose subscribe fails

* fix(mobile): keep the session tabs hold within the registry line budget after merging main

Move the pre-snapshot hold into the session tabs stream module, note that only older hosts need it, and give the unsubscribe test the real registration version now that a worktree-wide unsubscribe spares later streams.
2026-09-28 21:15:39 -07:00
Jinwoo Hong a134d1259e feat(mobile): tell users when a newer app binary is installable (#23755)
* feat(mobile): tell users when a newer app binary is installable

With OTA page updates, store releases get rare and users stop looking.
The shell now asks the channel that installed it. Android sideload
reads GitHub's mobile-android-v* tag refs and proves the release has
an APK. iOS reads the App Store lookup. A home card above Desktops,
dismissible per version, and Settings rows surface the result. The
check runs on the desktop updater's cadence: cold start, foreground
once 24 h have passed, and a 1 h retry after a failure.

The releases atom feed was not used because it lists only the 10
newest releases, which are all desktop builds, so it never carries a
mobile tag.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): parse update replies with zod schemas

The anti-slop gate refuses Reflect.get on dynamic input. The GitHub
refs, the release, the App Store lookup and the stored update record
are now parsed into named schemas before they are read.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): say why the Android update source reads tag refs

Record why the Android source reads tag refs. The releases atom feed
and /releases?per_page=100 are both newest-first windows that desktop
releases fill. Either would silently report "current" once a run of
desktop builds pushes the newest mobile release out.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): load update state once and apply review rulings

Every check and dismissal now awaits one shared store load. This
replaces the merge that guessed whether a check had landed during the
load. A manual check that fails while the store loads therefore keeps
its 1 h retry instead of re-checking at once.

- checking is derived from the in-flight check.
- start() uses a per-start flag, so a StrictMode double start applies
  one load.
- A check that finishes after stop() writes nothing.
- A corrupt stored update record loses only itself.
- Tag refs are parsed with a single schema.
- The runtime wiring is folded into one file, and the card moves to
  home/.
- The recorded App Store fixture is oxfmt-formatted, with the same
  parsed value.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): keep the update timer armed across a stop and restart

A check that stayed in flight across stop and restart returned
'failed' without rescheduling. The restart skipped arming because a
check was in flight, which left a live checker with no timer until
the next foreground. The stop counter is removed. schedule() already
arms nothing while no start is active, and saving a real result after
a stop is harmless.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* chore: retrigger CI after the RPC recording repin (#23757) landed on main

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): trim the update checker and Settings rows

- The load sets prefs and the due time only. start() re-arms the
  schedule after it.
- The Settings result hides through one effect keyed on the result.
- onUpdate receives the URL.
- The version row is bound once.
- The retry and timeout constants are no longer exported.
- The unused AppUpdateChecker type is deleted.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): run the update check when its timer fires

The armed timer is the due time. Re-checking the wall clock when the
timer fired meant a clock stepped back skipped the check and re-armed
nothing. The due-time guard now applies only on foreground.

The binary version still comes from expoConfig.version. SDK 55
removed Constants.nativeAppVersion, so the no-expo-updates invariant
is now named in the comment.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): recover from a future check time and use Apple's page URL

If the device clock was ahead when a check ran and was corrected
later, the stored check time is in the future. Cold starts then armed
a timer for the whole skew, and foreground never came due. The stored
state now reads as never checked in that case.

The iOS link is the lookup's trackViewUrl instead of a URL built from
trackId.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* style(mobile): fit the future-check-time comment in the print width

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-28 22:28:50 -04:00
Neil ccdb324b63 Add CodeBuddy as a built-in coding agent (#23740)
* feat(agents): integrate CodeBuddy launch, status and session history

* docs: record CodeBuddy lifecycle verification

* fix(codebuddy): backfill scoped history and negotiate remote resume

* test(cli): include CodeBuddy in known search agents
2026-09-28 18:11:25 -07:00
Brennan Benson e8e144bf3c fix(native-chat): a subagent's words are presented as that subagent's, never the parent's (#23605)
* fix(native-chat): a subagent's words are presented as that subagent's, never the parent's

The journal already names the agent that produced every row, but the transcript
projection dropped it, so a subagent's prose rendered as the parent's reply, its
tool calls folded into the parent's runs, and a settled turn could fold down to
a subagent's words as its only visible answer.

The transcript message now keeps the row's producer. The fold keeps each agent's
calls in that agent's own run, a turn's answer is the session's own agent's last
prose, and a subagent's row names the subagent on desktop, mobile and a worker's
transcript text.

* test(native-chat): give the window fixture's slot the attribution field it now carries

* fix(mobile): read the subagent label the row is given, and pin the caption

* fix(native-chat): keep interleaved agents in order and each agent's own run live

Review follow-ups:
- the fold is main's adjacency fold plus one condition: a row never folds into
  another agent's run, so an agent's later call stays below its subagent's work
  instead of jumping back into its earlier row
- each agent has its own live frontier, so a parent still inside its spawn call
  reads as running while its subagent works below it
- mobile names no one on a row whose only content is hidden behind its settled turn
- a pending question from a subagent keeps its producer
- worker reads serve only the producing agent's id, bounded like the roster key
  that names it, and drop the provenance fields
- the single-message worker formatter is private, so no caller can drop names
2026-09-28 15:28:01 -07:00
Brennan Benson 29c7d5d983 fix(mobile): start AI-button agents through agent.launch, never a bare shell (#22762)
* fix(mobile): start AI-button agents through agent.launch, never a bare shell

"Fix checks with AI", "Resolve conflicts with AI", commit-failure recovery and
diff review's "New Agent Session" created a terminal with no agent and typed the
multi-line prompt into the shell, so each line ran as a shell command.

They now call agent.launchReplay into the existing workspace with the prompt;
the host picks chat or terminal from the user's default and delivers the prompt.
Hosts without the launch capabilities get the buttons disabled with update copy.

The agent comes from the desktop's own resolution (moved to src/shared). The
replay loop and capability read are shared with the workspace-create launch.

* test(mobile): repin bridged-parity tallies for the AI-button launch goldens

The corpus goes from 787 to 790 goldens: five shell-path goldens are removed and
eight agent.launch ones added; one lands in identical and two in
result-absent-settlement.

* test(mobile): re-record goldens for AI-button launches through agent.launch

Repinned baseline to 514ab7f868 and re-recorded all goldens. Against the branch
point: 781 header-only (baseline on all; adapterSha256 on the 48 goldens whose
adapter module changed; scenarioSha256 on 3), one body moved
(pr-triage-launch: createTerminal + terminal.send becomes agent.launchReplay),
eight added (the new launch outcomes and their reply matrices) and five deleted
(the shell-path scenarios and their matrices).

* fix(mobile): show review notes' agent launch progress and failures, once

"New Agent Session" left the sheet open with no progress for the whole launch
(up to a minute while a terminal agent readies), so a second tap started a
second agent, and a launch that never started or could not be confirmed
rejected an unobserved promise, showing nothing. The sheet now closes on tap,
the review screen says "Starting an agent...", one launch runs at a time, and
every outcome lands in the review screen's status line.

Marking notes sent now reads the screen state when the launch settles, so a
note written during the wait is not dropped by the whole-list save.

* test(mobile): re-record goldens for review notes' launch outcome on the review screen

Repins the corpus to 6f3018576d. One golden body moves:
review-create-agent-refused now fulfils with "Workspace not found" in the
review screen's status line and the sheet closed, where it previously
rejected an unobserved promise and left the status line empty. The other
789 goldens move only their baseline header.

* test(agent-status): drop the retired PR-triage terminal send from the identity inventory

The phone's AI buttons no longer create a terminal and send the prompt into it
(`createTerminalAndSendPrompt` is gone); the host's agent launch delivers it.
There is no terminal action consumer left in that file to pin.

* fix(runtime): publish saved source-control launch recipes to paired clients

settings.get is an allowlist and omitted sourceControlAi, so the phone never
saw an agent saved globally for "Fix checks", "Resolve conflicts" or commit
recovery and always fell back to the default agent. The host now publishes
the launch actions' recipes (agent, prompt template, agent args), normalized
so legacy saved defaults are already migrated. A new optional reply field:
older clients ignore it, and a client talking to an older host sees none and
keeps using the default agent.

* fix(mobile): ask to update Orca only when the host answered without agent launch

An unread or failed status read settles with no capabilities, which the AI
buttons read as an old host and showed "Update Orca on your computer". The
update copy now needs a status the host actually returned; an unread one keeps
the buttons disabled without blaming the desktop's version.

* fix(mobile): send an AI button's saved agent arguments with its launch

The desktop's direct launches for "Fix checks", "Resolve conflicts" and commit
recovery pass the action's saved agent arguments to agent.launch; the phone
honoured the saved agent but dropped its arguments. It now sends them the same
way: absent when none are saved, so the host keeps the user's configured
defaults. A host that predates the field ignores it.

* test(mobile): repin the RPC recording corpus after the launch recipe and availability fixes

Repins to fe85e0346f. All 790 goldens move only their baseline header: no
scenario saves agent arguments or reads an unreadable status, so no recorded
behaviour changes.

* fix(mobile): say the host status is unreadable instead of nothing when it is

With the update copy now reserved for a host that answered without agent
launch, an unread status left the AI buttons disabled with no explanation.
They now say "Could not read this host's status. Go back and reopen it.", the
words the mobile web shell already uses for the same failure; leaving the host
re-reads its status.

* fix(mobile): wrap an AI button's prompt in the action's saved template

The desktop renders every source-control launch's prompt through the action's
saved template (buildSourceControlRecoveryAgentCommandInput); the phone sent
its built-in prompt as is. Now that the host publishes the recipes, the phone
renders through the same shared function, refuses an empty result as the
desktop does, and offers the rendered text when it could not be delivered.
Review notes have no recipe and are unchanged.

* test(mobile): re-record goldens for the templated AI-button prompt

Repins to 7fd1555d20. One golden body moves: pr-triage-prompt-not-delivered
now carries the prompt as sent (rendered through the action's template) on its
prompt-not-sent result, which is what Copy prompt offers. The other 789
goldens move only their baseline header.

* fix(mobile): re-read a host status that failed while the connection stayed up

A status.get that timed out or was cut over settled the host's gates closed
and was never asked again until the connection state changed, so the phone's
AI buttons stayed disabled behind "Could not read this host's status" on a
link that was working. The gate still settles closed at once, so a failed
read never holds the host screen, but it now re-asks in the background with
the same backoff the runtime capability probe uses, and opens once a status
lands. A reply this app cannot decode is not re-asked.

* test(mobile): repin the RPC recording corpus after the host status re-read

The status gate change moves no recorded behavior: every golden's body is
unchanged and only its baseline header moves to the new pin.

* fix(mobile): show a launch's host warning as a note, not an error

A launch that went ahead can carry a host warning (a structured chat ignores saved agent
arguments, including the '' a template-only save writes). The AI buttons rendered it in the red
error line beside a success haptic. The notice now carries it separately as secondary text, and
review notes keep saying they were sent. Commit recovery also takes the synchronous in-flight lock
the PR triage buttons use, so two taps before a re-render start one agent.

* chore(mobile): record the host status re-read timer for React Doctor

The status re-read arms one retry timer from inside its read and clears it in the effect's
cleanup. React Doctor reports that self-rescheduling shape even in its minimal form, which failed
both changed-lines gates. Suppressed the same way as the session startup timers.

* fix(mobile): say review notes are waiting for the desktop instead of doing nothing

With no live connection, New Agent Session threw from a handler whose promise the sheet drops, so
the tap did nothing visible while the button stayed enabled (proven capabilities survive a drop).
It now closes the sheet and shows "Waiting for desktop..." as the other AI buttons do.

* fix(mobile): stop sending an AI button's saved agent arguments

Whether saved arguments apply depends on the route and shell the host settles after the request
(a chat ignores them and warns; malformed ones fail after admission), and the desktop sends them
only when they apply. The phone cannot know that, so it now leaves them out and the agent's default
arguments apply, as before this series. The saved agent and prompt template still apply.

* fix(mobile): mark review notes sent through the latest save

The sent marks after an agent launch went through the save callback captured at tap time, whose
rollback restores the screen from that moment, so a failed save could drop notes written during
the launch. It now uses the latest render's save, as it already did for the screen state.

* refactor(mobile): own the host status re-read outside the effect

The re-read loop lived inside the effect body, so React Doctor could not see its cleanup and
needed an inline suppression plus a config allowlist entry. The loop is now a plain function that
returns its stop handle, and the effect returns that handle, the same shape every caller of the
runtime capability probe uses. Both suppressions are removed; behaviour is unchanged.

* test(mobile): record the host descriptor from a background status re-read

Pins that the status read records the host descriptor when a re-read succeeds after a failed first
read, not only on the first answer.

* fix(mobile): show a PR AI launch notice only under the button that launched it

Fix checks and Resolve conflicts shared one error, warning and undelivered prompt, so a Fix checks
launch whose prompt was not sent also offered "Copy prompt" under Resolve conflicts, copying the
fix-checks prompt. Notices are now kept per button. The host availability notice stays under each
disabled button, since it explains why that button cannot be tapped.

* fix(mobile): say the host status is being retried instead of asking to reopen it

The host status gate now re-reads a failed status in the background, so "Go back and reopen it"
asked the user for a step that is no longer needed. The review sheet hint uses the same words.
The mobile web shell keeps its own copy.

* refactor(mobile): run the host status gate on the shared status probe

The gate had its own copy of the status probe's retry loop (same delays, same cutover and backoff
split, same stop on an undecodable status). The probe now takes an optional callback for each
failed attempt, which the gate uses to settle closed on the first failure, and the duplicate loop
and its now-unused reader are removed. Existing probe callers are unchanged.

* test(mobile): repin the RPC recording corpus after merging main

Re-records every golden against the merge commit and drops the three goldens whose
scenarios this branch removed, which the merge had restored from main.

* test(mobile): re-record the RPC goldens on the merge with main

Conflicted goldens were seeded from main and re-recorded against the merged
tree; every value either side recorded survives except main's terminal.send in
the PR triage launch, which this branch removes. Drops three goldens main still
had for scenarios this branch deleted.

* feat(mobile): confirm an AI button's agent started, naming the workspace

Fix checks, Resolve conflicts and commit recovery now show "Agent started in
<workspace>" under the button once the host started the agent with its prompt,
so a tap is no longer silent. The workspace label comes from the Source Control
panel and falls back to the branch.

* test(mobile): repin the recording baseline to the success-confirmation commit (header-only)

* fix(mobile): name the workspace in the diff review's AI-button confirmation

The diff review screen mounted the PR sidebar without a workspace label, so
"Agent started in ..." under Fix checks and Resolve conflicts named the branch
while the screen header named the workspace. The sidebar now requires the label
so no screen can drop it, and the diff review passes the one its header shows.

* test(mobile): re-record the RPC goldens on the merge with main

Repins baseline to the merge commit, the last commit to touch a fenced
path. Against this branch before the merge, only header fields move:
baseline on every golden, and adapterSha256 on the 14 review-action goldens
whose adapter main now drives through the review sheet state. No recorded
body changed.

* test(mobile): re-record the RPC goldens on the merge with main

Repins baseline to the merge commit, the last commit to touch a fenced
path. Against this branch before the merge only the baseline header moves,
on every golden; no recorded body changed.

* test(mobile): give the send-sheet stacking test the review controller's host status inputs

The merge with main brought in #22951's stacking test, which builds the review
controller without the host capability and status inputs this branch made
required, so the mobile test typecheck ratchet failed.

* test(mobile): re-record the RPC goldens on the merge with main

Repins baseline to the merge commit 03995ae29d, the last commit to touch a
fenced path. Against this branch before the merge only headers move: baseline
on every golden, and adapterSha256 on the 14 goldens recorded through the
terminal adapter main changed in #23080. No recorded body changed, and the
merged corpus differs from main exactly as this branch did before.
2026-09-28 14:28:09 -07:00
Jinwoo Hong fe34acda3b fix(mobile): truncate oversize markdown reads instead of failing them (#23676)
* fix(mobile): truncate oversize markdown reads instead of failing them

The desktop bridge refused markdown over a private 512 KiB cap with
file_too_large, which reached the phone as a generic runtime_error that
the reader discarded, so a 632 KB file showed "Couldn't load markdown".
Reads now return a UTF-8-boundary prefix under one shared 2 MiB budget,
marked truncated with the full byteLength and read-only. The phone shows
the truncation like file tabs do and maps refusal codes to real copy, so
an older desktop's refusal reads "File too large for mobile preview".

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): pin that a truncated markdown read is never editable

The read budget sits above the edit budget, so every truncated document is
already read-only as file_too_large. Pin that ordering so a future budget
change cannot make a prefix editable, and name the constant as the markdown
preview budget, separate from the file preview's own.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): refuse saves from a truncated markdown read

A phone holding a truncated prefix must not write it back; the 256 KiB
save guard refuses it as file_too_large before any version check. The
shared budget test shrinks to the ordering it pins and names the bridge
test as the behavioural pin.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): tighten markdown truncation types and measure once

The read truncation measures the document once. The disk fallback drops
its truncated-only read-only text, which the status line never showed,
and both truncation fields are optional there. A markdown doc's flag is
only ever true, and the schema comment names the hook, not line numbers.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): drop a stale disk-fallback comment

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* chore: retrigger CI after #23675 landed on main

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-28 16:58:15 -04:00
Jinwoo Hong 0034ede120 fix(mobile): left-align every line of the desktop host card (#23673)
* fix(mobile): left-align every line of the desktop host card

StatusDot carried its own marginRight on top of each row's spacing, so the
host card's status text sat 14 px in and the worktree line was hand-indented
to match. Move the dot spacing to the row gap in every consumer and drop the
worktree-line indent.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-pin tasks style parity hash for the title-row gap

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-28 15:31:40 -04:00
8b410b4893 feat: add first-class Qoder CLI support (#23581)
feat: add first-class Qoder CLI support

Integrate Qoder launch, identity, canonical hook status, trust and resume.
Verify with captured Qoder 1.1.64 transcripts and hidden Electron sidebar checks.

Builds on and cross-reviews #7502, #8611, #9655, #12910, #13311 and #15291.

Co-authored-by: dalveytech-vincent <vincent@dalveytech.com>
Co-authored-by: Eridanus117 <45489268+Eridanus117@users.noreply.github.com>
Co-authored-by: xingqingzzp-gif <xingqingzzp-gif@users.noreply.github.com>
Co-authored-by: jyang2004 <jyang2004@users.noreply.github.com>
Co-authored-by: yunqian <yunqian@alibaba-inc.com>
Co-authored-by: huzhening.hzn <huzhening.hzn@alibaba-inc.com>
2026-09-28 02:59:50 -07:00
400e4e7957 feat(agents): add Freebuff launch and sidebar status support (#23567)
Add Freebuff launch support and execution-host status reporting for the sidebar, including running, question, blocked, and settled states. Validate against captured CLI transcripts and real rendered sidebar evidence.

Cross-referenced community implementations #17065, #20839, and the Freebuff portion of #18790. Preserve their agent/catalog/mobile/documentation coverage and add canonical status publication and regression tests.

Co-authored-by: Harkaran Brar <18134082+harkaranbrar7@users.noreply.github.com>
Co-authored-by: Prarambha369 <98906077+Prarambha369@users.noreply.github.com>
Co-authored-by: Lesley Murfin <260182349+LesleyMurfin@users.noreply.github.com>
2026-09-28 02:32:41 -07:00
Brennan BensonandClaude a68d62911e fix(native-chat): the host writes chat failures for a person, with a typed fact beside them (#23116)
* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* fix(native-chat): name a chat write by its target, not the owner generation

A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.

Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.

Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.

* fix(native-chat): every journal append reaches the chats that are open

A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.

A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.

* test(native-chat): an epoch replacement reaches the open chat

* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

* perf(native-chat): a publish behind a delivered commit reads nothing

Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.

* test(native-chat): state why the teardown test's fake journal is safe to cast

* docs(native-chat): say mutation admission checks only the writer lease

* docs(native-chat): drop the send rebase from comments that still described it

* fix(native-chat): a message is accepted, then delivered

A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".

A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.

Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.

A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.

Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.

* fix(native-chat): settle queued messages only for the child that ended

A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.

A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.

The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.

* fix(native-chat): an adoption that fails to import keeps the conversation open

The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.

* perf(native-chat): the recovering open reads the journal once

Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.

* fix(native-chat): an attach that fails after indexing its child leaves no child behind

A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.

* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer

The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.

A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.

* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down

The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.

* fix(native-chat): a message rejected while its chat was closed reads as not sent

A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.

* test(orchestration): name why the readiness settlement fakes are cast

* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent

* docs(native-chat): drop the fence from the admission the send effects run behind

* docs(native-chat): give the fence move on release the reason that still holds

* docs(native-chat): stop citing a write fence check in launch and mailbox comments

Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.

* refactor(native-chat): the provider child is its own record

A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.

- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
  patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
  own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
  the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
  dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
  is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
  the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.

* fix(native-chat): the delivery loop alone settles a message its start or child failed

A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.

- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
  reads how it ended: a Stop continues; anything else writes one failure row and rejects every
  queued message with the same words, then stops. A child still starting whose start the adapter
  says did not land fails the same way. The exit, eviction and the settlement retry only settle
  the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
  nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
  failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
  closed, with or without a child, and a start the loop already has in flight is waited for so the
  child it produces is stopped rather than left behind.

* refactor(native-chat): a stopped child ends on the one reading of its stop

The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.

* feat(native-chat): the host says it accepts a send before any agent has it

The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.

* refactor(native-chat): an attach never opens a journal of its own

The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.

* fix(native-chat): a moved fence resends nothing on a host that accepts first

The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.

The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.

* refactor(native-chat): a child's end says whether the user or the host stopped it

The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.

* fix(native-chat): a chat whose only work is a queued message is not offered for resume

A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.

* test(native-chat): type the queued-message fixtures in the resume-offer tests

* fix(native-chat): a start that dies while a message waits on it is that message's failed start

Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.

* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now

* test(native-chat): pin what a failed start settles, and what a resume offer names

A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.

* test(native-chat): the failed-start pins fail on what the message became, not on a timeout

* test(orchestration): the preamble's host stub is typed, not cast

The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.

* feat(native-chat): a typed failure fact beside every failure sentence

Adds the shared vocabulary the host writes a failure with: a closed failure kind, a
provider diagnostic that says who it is for (a person, or a log), and a refusal cause
beside the refusal code. Status rows gain an optional failure fact and rejected
submissions an optional rejection fact; the dispatch row carries it, the reducer reads
it field by field, and the projection forwards it. Older rows and older readers are
untouched: every field is optional and the schemas stay open.

* fix(native-chat): durable failure rows and rejection reasons are written for a person

Every host writer that records a failure now writes a sentence for a person beside a
typed fact, instead of embedding a refusal's message, an exception or a composed exit
string. A provider's own words travel as a separate diagnostic from the places Orca
composes them - the Claude and Codex exit stderr (a log), Codex's JSON-RPC message,
Claude's compact_error and Codex's turn error (for a person) - and are never inferred
from a string afterwards. Not signed in and oversized history are typed at the
adapter that detects them.

Covers start and restart failures, the delivery loop, dispatch rejections (content,
queue-full, write failures, provider refusals), cancel and answer confirmation rows,
compaction, the rewind fallback, and not_delivered, which released clients printed
as it was. Two leaks close on the way: a settlement retry no longer writes Orca's
probe evidence into the exit row, and an attach or journal-sink failure is recorded
as Orca's fault rather than as the provider stopping. The legacy rejection markers
and the reasons on sends in doubt stay byte-identical.

* feat(native-chat): refusals name their cause, and a failed start is worded in one place

A refusal now carries an optional cause beside its code: one closed enum of the situations
a chat write can meet, set at every emitter a structured-chat write reaches. Returned
refusals build it with refuse(code, cause, message). Store and host paths that raised a
bare Error(code) now throw AgentSessionRefusalError, whose message is still the code and
which has no code property; the RPC error mapper handles it before any other passthrough,
keeps today's wire code and message byte-identical, and adds { refusal: { code, cause } }
to the error's data. The hold throws it, and restart-resume files the cause beside the
unchanged reason. The operation ledger stores the cause beside the code, so a replay names
the same situation as the first answer. The store fallback copy picks its words and cause
by situation, so a stale replay or a moved lease no longer reads as a latched owner.

Every failed start is worded by structuredAgentSessionStartFailure(cause, context), which
returns the row sentence and the typed fact together; the delivery loop, the exit
settlement and the dispatch that met a starting child all call it. Provider diagnostics
are capped at the lease record's 512 characters wherever a fact is built.

* refactor(native-chat): one reader of why a submission was rejected

classifyDispatchRejection(submission) returns { category, verdict, kind? }. It reads the
typed rejection fact when the row carries one this build can place, and the legacy
markers otherwise - all six, including not_delivered, which released clients printed as
it was. The verdict is null only for a withdrawal, a host restart and a closed chat;
write failures, a full queue and not-delivered stay failures. It replaces
dispatchRejectionReasonIsInternal and every string comparison against the markers: the
outbox reconcile, the rejection notice, and the send disposition, where a replayed
Stop-withdrawn send no longer surfaces as a failed send.

The journal reducer's echo-aliasing guard reads the narrow isWriteFailureSubmission,
which matches the typed kind or the legacy prefix in any dispatch state, so its behaviour
on legacy unknown rows is unchanged.

* test(native-chat): pin provider diagnostics where they are composed

The Claude exit status and stderr, the Codex stderr tail and Codex's JSON-RPC message are
each checked at the place Orca composes its own error around them, so the typed detail
is proven to come from the provider's value and never from Orca's wording.

* fix(native-chat): an attachment Orca refuses says which limit it broke

The content check's refusals (20 images, 5 MB per image, 20 MB in total, supported types) are
written for a person, but the rejection writer replaced them all with one generic sentence.
Each refusal now carries its own sentence, in MB rather than bytes, and the writer records it
beside kind attachmentInvalid. Only an attachment that could not be read keeps the generic
sentence.

* fix(native-chat): the chat tab table refuses with a typed cause

Showing a chat tab refused with bare Error('agent_session_conflict') and
Error('agent_session_identity_required'), the only chat-reachable refusals still thrown without a
cause (opening a chat from history can reach the second when the chat is removed mid-open). Both
now throw the typed refusal; wire code and message are unchanged.

* fix(native-chat): a compaction Codex refuses up front keeps Codex's words

When Codex refused thread/compact/start, the adapter passed on only Orca's wrapped error text and
dropped Codex's own message, so the chat's row read just "Compaction failed." The refusal now
carries Codex's message as the failure detail, as a compaction that fails later already did.

* fix(native-chat): an unreadable chat record no longer promises an update fixes it

The recordUnreadable copy said a newer Orca saved the chat, but the store marks a record
unreadable for damage and key mismatches too, where updating does nothing. The sentence now says
Orca can't read it and gives both next steps.

* fix(native-chat): a restart or a close leaves released clients a sentence, not a marker

host_restarted_before_delivery and provider_closed_before_delivery are not in the markers released
desktop and mobile builds hide, so they printed raw on every rejected message a restart or a chat
close left. Neither marker has shipped. New rows carry a sentence plus kind hostRestarted or
chatClosed, as not_delivered already did; the classifier still reads both markers, and the
verdict for both stays no-failure.

* fix(native-chat): log why an attachment could not be read

The rejected message now says only that the attachment couldn't be read, so the error that said
why (a missing file, a permission, or an unexpected throw) went nowhere. It is logged instead. The
content rejection moves beside the content check that owns its errors.

* refactor(native-chat): drop an unused thrown-refusal cause reader

It had no callers, and its comment claimed it looked through wrappers, which it did not.

* fix(native-chat): a compaction the provider never confirmed is recorded as unconfirmed, not failed

* fix(native-chat): an empty or non-user message is not recorded as a bad attachment

* fix(native-chat): an undelivered preamble's error ends in one period

* refactor(native-chat): a compaction ends as compacted, failed, or unconfirmed, never an unlabelled error

* refactor(native-chat): a failure's sentence is written only from its fact

A writer could choose its sentence and its kind separately, so five writers
put hand-written words beside a fact that said something else. One shared
constructor, agentSessionFailureWords(fact, { surface, agentName }), now
makes both, and the journal types refuse anything else: a status row, a
rejected message or a conversation command that carries a fact must carry
the sentence that constructor branded. Persisted rows keep their shape.

- The sentence table and the restart table move to src/shared, as the English
  default a client copy table can reuse.
- A rejection's legacy markers come from the same constructor. A write
  failure is now the bare `provider_write_failed` marker, which released
  clients already hide; its error goes to the log.
- An image Orca refuses carries which check it failed (and the limit) in the
  fact instead of a sentence; an empty message gets its own kind, and a
  non-user message is Orca's fault.
- The exit row, the interrupted compaction, the /clear failure and the rewind
  placeholder no longer carry their own words beside a fact: the exit row
  says what its fact says, and the placeholder carries no fact.

* fix(native-chat): a start that failed without an observed exit no longer blames the provider

Every untyped start error was recorded as "The provider stopped before it
finished starting.", so an Orca fault, a failed spawn or a close that ended
a start was blamed on the provider. Only an exit the adapter observed says
so now: the Claude adapter marks the error it saw the child exit with, and
anything else is a new `startFailed` kind, "<Agent> couldn't start.", keeping
the provider's diagnostic when the error carried one. A child gone with no
end observed, and a /clear whose new conversation was refused, read the
same way.

* fix(native-chat): a view never restarts a chat whose last start failed

A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.

* fix(native-chat): a chat a terminal agent still holds says to quit that agent

The merged base words a restart a terminal agent's claim refused with the
refusal's own message, which names the process. That message is Orca's text
and never reaches a durable row here, so the row read only "<Agent>
couldn't restart." and lost the one step that frees the chat. The restart
sentence now derives it from the refusal's cause: a claimConflicted refusal
adds "This chat is still open in a terminal agent. Quit that agent to
continue the chat here." The live refusal still names the process.

The restart-resume ledger test now sets up a claim the base still refuses:
a terminal owner that is proven running.

* refactor(native-chat): keep the changed files inside their lint limits

The legacy-marker lookup is a table, not a non-exhaustive switch; the
preamble tests read the error without a cast; and the Codex history refusal
goes through a named constructor so its file stays under the line limit.

* test(native-chat): start the child the loop waits on with an attach, not a second view

A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.

* fix(native-chat): settle a gone generation's turn wherever a conversation opens

A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.

* test(native-chat): prove the next child's start settles the turn an earlier child left

The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.

* test(native-chat): count a failed start's rows by row, not by text

Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.

* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget

* refactor(native-chat): refusals carry details keyed by their code

A refusal named its situation with one flat cause list shared by every
code, so nothing stopped a site pairing a code with a situation that code
never means, and the loose fields released clients read (fence, revision,
resolution, verdict, rewind reason) were written by hand at each emitter.

A refusal is now one variant per code with optional details: a reason that
code lists plus that code's own facts. refuse(code, details, message)
rejects a reason the code does not list at compile time, and it is the one
place the loose top-level fields are copied from details, so released
clients read exactly what they read before. A site that cannot name its
situation uses refuseUnclassified, which carries facts but no reason, the
same as an older host; there is no catch-all reason.

Thrown refusals put { refusal: { code, details } } in the RPC error's data
(wire code and message unchanged). The operation ledger, the restart-resume
record and a restart failure's embedded refusal keep details beside the
code and read them back against it; a row an unreleased build wrote with a
cause parses and reads as naming none.

* test(native-chat): pin the open's and the send's start and row counts, however the view binds

Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.

* test(native-chat): a restart-resume failure keeps its refusal details

The recovery capsule reads a failure's details back against the refusal
code in its reason: facts the code does not list and a reason another code
owns are dropped, and a record an unreleased build wrote with a cause
still parses, naming nothing.

* test(native-chat): import the failure words once in the provider child test

The merge left two imports of the same modules.

* fix(native-chat): settle a gone generation's turn at every open but an acquisition's

The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.

* test(native-chat): hold the create's start open until the views bind

The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.

* fix(native-chat): a start the provider refused or Orca broke no longer says the provider stopped

A failed start whose cleanup proved the child gone was typed as
`providerStartFailed` whatever failed it: Codex refusing to resume a thread,
a timeout, or Orca's own store fault. The chat then read "The provider
stopped before it finished starting.", which was untrue, and the provider's
own words were dropped. A refused restart also inferred the same from an
`exited` verdict, which only says nothing runs now.

Only an exit the adapter observed names that situation now; anything else
is refused with no reason, keeping its verdict, and the chat reads
"<Agent> couldn't restart.". The provider's words travel host-side from
where the acquisition failed into the start-failure fact's detail, never
onto the refusal, and the sentence does not quote them. What failed is
logged once where the start failed.

* fix(native-chat): a refused /clear start keeps the situation it named

The replacement start that /clear makes built its own start-failure fact,
so a refusal that named its situation, such as not being signed in or a
history too large to restore, was recorded as a bare "couldn't start". It
now takes its fact from the same start-failure function as every other
start, as a new session that failed to start.

* fix(native-chat): the journal schema and comments describe a refusal's details, not its cause

The persisted failure fact's schema still described `refusal.cause`, which
this branch replaced with `details`. It now describes `details` as an
optional open object; a row an earlier build wrote with a `cause` still
parses. A comment and three test descriptions that still named the cause
now name the details.

* fix(native-chat): a Claude child that exits while being acquired still reads as the provider stopping

Now that only an exit the adapter observed says the provider stopped, the
exit the Claude adapter saw during acquisition has to be marked where it is
seen, as the exit after acquisition already is. Without the mark, a Claude
CLI that exited at spawn read "Claude couldn't restart." instead of "The
provider stopped before it finished starting."

* fix(native-chat): word a failed chat start's refusal as its start failure

A chat whose agent failed to start answered the create with the raw error: the launch strip read "Chat could not be started. claude stream-json exited (code 1): claude: not signed in", and the ledger replayed the same text. The first answer and the replay now carry the sentence the chat's start-failure row reads as ("The provider stopped before it finished starting.", "Claude couldn't start.", the not-signed-in and history-too-large sentences), and the raw error goes to the log. A store refusal's code and the unproven-exit marker are unchanged.

* fix(native-chat): show Claude's API retries as one sentence row

While Claude retried a refused request (a 429, say), the chat gained one red row per attempt reading "rate_limit", with the raw retry frame behind Details. Each retry run now writes one warning row that later attempts revise in place: "Claude is rate-limited and retrying." for a rate limit (error `rate_limit` or status 429), and "Claude hit a temporary problem and is retrying." otherwise. The row carries a `providerRetrying` fact with the provider's error type and status, and the frame as a log detail capped at 512 characters.

* fix(native-chat): tell the user to run /clear again when its new conversation can't start

When /clear's replacement conversation failed to start, the result told the user to "send your message again", which would go into the old conversation. The failure words now take the command the start was for, so a failed /clear reads "Codex is not signed in for the selected account. Sign in, then run /clear again.", "Codex couldn't start. Run /clear again." or "The provider stopped before it finished starting. Run /clear again." A message send keeps its wording.

* test(native-chat): expect a failed Claude create to be refused in a sentence

The runtime suites asserted the CLI's stderr reached the create refusal; it now goes to the log and the refusal reads as the chat's start failure.

* fix(native-chat): name the agent that stopped starting and say how to retry a failed start

A start the provider ended now reads "Claude stopped before it finished starting." (or "The agent ..." when the chat's agent is unknown) instead of naming "the provider". A start or restart that failed with a chat left to retry now ends in "Send your message to try again.", or "Run /clear again." for /clear; released clients print only this sentence. A message rejected at dispatch because its child died while starting names the same agent as the start's row.

* refactor(native-chat): drop the composer's second error formatter

After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.

* test(native-chat): pin the reason on a message rejected while its chat was closed

The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.

* fix(native-chat): answer a create refused before spawn in the words its replay reads

A create that failed before any process started threw Orca's own error text as its first answer,
while its replay from the operation ledger read the generic start sentence. The two refusals a
person can act on, a launch whose Anthropic sign-in variables override the managed Claude account
and a Claude account switch in progress, are now typed where they are thrown and worded by the
shared failure constructor, so the first answer and the replay say the same thing. Every other
pre-spawn failure reads the generic start sentence, with its own text in the log. The first answer
keeps its wire code; a message that is itself a code is unchanged.

* fix(native-chat): say how to retry after an agent stopped before it finished starting

"<Agent> stopped before it finished starting." gave no next step outside /clear, unlike every other
failed start. It now ends "Send your message to try again.", the same step a start or restart that
could not run gives; after /clear it still says "Run /clear again."

* fix(native-chat): say how to reach a Claude chat when a WSL Claude account blocks it

A Claude chat that restarts while a Claude account is added in WSL and no Windows Claude account is selected was refused before spawn with Orca's own text as its first answer, and the generic start sentence on replay. The refusal is now typed where the account gate throws, and both answers read the same sentence: choose or add a Windows Claude account in Claude Accounts settings, then send the message again. Account settings that cannot be read name no situation and keep the generic sentence.

* fix(native-chat): say the reason a host names for a refused chat write

A refused Stop, answer, setting, goal, command or queued message now reads the refusal's reason as
well as its code. A code stands for several situations, so the code alone could only say what did
not happen; with the reason, the notice says why and, where the person has a step to take, what it
is: "The agent is still responding. The command didn't run. Wait for the agent to finish
responding, or stop it." The phone uses the same words.

The notice table keys on code, then reason, then the kind of write. Every reason of every code has
an entry, so a reason the host adds does not compile until it has words; a reason whose honest
words are its code's keeps the code's row. A refusal with no reason, or one this build does not
know, reads exactly as before, which is what an older host gets. A start that failed reuses the
failure row's own sentence rather than a second one.

A queued message keeps the reason, the rewind reason and the owner's verdict with its saved
failure, and a rejected one keeps the host's typed fact without its provider detail. Nothing that
moves with the owner or comes from the provider is saved; entries saved before this load as they
were. The saved failure moves to its own module beside the words chosen from it.

* fix(native-chat): retry a refused send under a new id only once its agent is proven gone

A send refused because Orca could not tell who owns the chat keeps its operation id, since the
first attempt may still land. When the refusal also says the agent process has exited, nothing can
run that attempt, so the message moves to its Retry row under a new id instead of holding the queue.

The owner's verdict is read as a floor: a saved `exited` is final, and any other saved verdict never
changes the id on its own. Only a verdict re-derived from the current lease can raise it to
`exited`, and nothing lowers a saved `exited`. Today no host sends a verdict on a send refusal, so
nothing a person sees changes; the rule is in place for the saved verdict a reload reads back.

* fix(native-chat): say a /clear that never finished did not finish, instead of that it cleared the chat

A send into a chat whose /clear started but never committed was refused as if the conversation had been cleared: "This conversation has been cleared. Your message was not sent. Open the current conversation to continue." The clear never finished, so its new conversation may not exist and there is nothing to open. That refusal now has its own reason and reads "The last /clear didn't finish. Your message was not sent. Start a new chat to continue.", which is the only way on today. Its message for released clients says the same: "The last /clear didn't finish. Start a new chat to continue." Only a committed /clear still says the conversation was cleared.

* fix(native-chat): a send the provider never received after a restart has no verdict

Restart reconciliation rejects a crash-stranded send the provider's history
proves it never received. Nobody failed that send, but the verdict table
treated it as a failure. Each rejection kind now has its verdict in one
exhaustive table, so a new kind does not compile until its verdict is chosen;
no verdict for a withdrawal, a host restart, a chat close, or this lost send.

A rejection whose kind this build cannot place, such as one a newer host
added, now reads as undelivered with no verdict instead of falling back to the
reason beside it: all it proves is that the message did not happen. The host
keeps such a fact's kind when it reads the row back, rather than dropping it
and letting the reason decide.

Only kinds that can be why a message was not sent may reject one, by type:
compaction, cancel/answer confirmation and provider-retry kinds stay on status
rows. The one dispatch-row builder takes its input from the type that makes a
rejected row carry its fact.

* fix(native-chat): a send refused after its agent exited keeps its place in the queue

When Orca cannot tell who owns a chat but the refusal says the agent process
has exited, the next attempt may use a new id, since nothing can run the old
one. It no longer marks the message as rejected: nothing recorded it, so it
still holds the head of the queue, and later messages wait behind it instead
of being sent ahead of it.

* fix(native-chat): stop reading a provider's words from an error that contains itself

A cleanup that aggregates errors restarted the depth count for each one, so an
aggregate error that contains itself recursed until the host ran out of stack.
One depth bound now covers both the cause chain and the aggregated errors.

* fix(native-chat): say a chat whose history can't be read can't continue, and to start a new one

A read of a chat's history is refused with `agent_session_journal_unreadable` only when the chat's journal file is corrupt or not a database, which no retry can change. The notice table had no words for a read at all, so a pane had nothing to show but the raw code. Reading a chat's history is now its own request kind, `read-history`, and that refusal on it reads "This chat's history couldn't be read, so it can't continue here. Start a new chat to continue.", whether the host names the reason or raises the bare code. A write refused under the same code keeps its words, because its cause is any failed open, which can clear. Any other read refusal says only "This chat's history couldn't be loaded." `agentSessionReadHistoryRefusalParts(code, details)` gives a pane those words from a read error.

The notice sentences move to their own module so the table stays within its size limit.

* fix(native-chat): decide what every way a child ends means for queued messages in one table

What a child's end means for the messages queued behind it was an if-chain: a user's Stop was checked in one place, a host stop in another, and every other cause, including one added later, fell through to "the provider exited". It is now one table over every end cause, so a new cause does not compile until someone says whether it fails what is queued and how. A user's Stop still fails nothing, a host stop is still Orca's fault, and an exit, a failed attach or an eviction still carry the failure the end recorded. Nothing a person sees changes.

* test(native-chat): a close that stops the child and then fails rejects what is queued by how the child ended

When a chat closes, stops its agent, and then fails a later step, the chat stays open with its messages still queued. The delivery loop then rejects them by the way the child ended: an eviction during startup reads as a failed start, otherwise as the provider having stopped, and either counts as a failure. The cases are rows over the end cause, so another way a chat closes is one more row.

* test(native-chat): type the refusal a persisted-schema test admits

* fix(native-chat): say whether a chat's history is damaged or just couldn't open

A write refused because the chat's journal would not open named one reason, `journalUnreadable`, for every failed open, and a read of the history took that same reason as final. So the words depended on what was asked, not on what happened: a busy or permission-denied open could tell a person to start a new chat, and a damaged one could read as something that clears.

The host now decides at the refusal which it was. `journalCorrupt` is set only when SQLite itself reports the journal damaged or not a database (SQLITE_CORRUPT or SQLITE_NOTADB, extended codes included, read from the driver's result code and never from message text, through any `cause` chain). Every other failed open is `journalUnavailable`. A corrupt history reads "This chat's history couldn't be read, so it can't continue here. Start a new chat to continue.", with "Your message was not sent." before the step on a send. One that couldn't open reads "Orca couldn't open this chat's history right now. Try again.", likewise on a send. A host that names no reason gets "Orca couldn't read this chat's saved history.", which promises neither, because damage can't be proven from the code alone.

The refusal's message, which released clients print for a send, is now that person sentence instead of the open error's own text; the error is logged instead. `journalUnreadable` is replaced outright: no released build wrote it, and a stored one reads as a refusal with no reason.

* docs(native-chat): say why a history that couldn't open names its retry step

* fix(native-chat): a failed start's row keeps the words its rejected messages carry

When the delivery loop settles a failed start before the child's exit is published, it writes the
start's error row and rejects every queued message with the adapter's startup answer. The exit
settlement then rewrote the same row from the exit event, so the row could say one thing while the
rejected messages, which are terminal, said another. The exit now leaves a start's row it finds
already written.

* fix(native-chat): a Claude start Orca itself failed no longer says Claude stopped

Every error that ended a Claude session was marked as an exit the adapter
observed, including a start Orca failed while the CLI was still running: a
saved option whose restore lost its answer, an init frame naming another
session, or a journal write fault. Those read "Claude stopped before it
finished starting." although Claude never stopped on its own. The mark now
stays where the child's exit is seen (the connection's exit callback), so
those starts read "Claude couldn't start." with any diagnostic beside it,
and a real exit before the start lands still says Claude stopped.

* fix(native-chat): name the agent that stopped, and blame Orca for its own closes

A chat that lost its agent mid-response said "The provider stopped…", and a
started Claude session that Orca itself closed after a journal fault said the
same, as if Claude had exited on its own.

The exit row and the rejected-message reason now name the chat's agent ("Claude
stopped while this response was in progress…", "Codex stopped before this
message was sent."), or "The agent" when the name is unknown; the stale-state
settlement now passes the agent name too. After a Claude start has landed, the
ended event reports providerExited only when the child's own exit was observed;
any other close is Orca's fault and reads as one.

* test(native-chat): pin the sidebar verdict to the rejection classifier for every kind and legacy marker

* fix(native-chat): blame Orca, not Codex, when Orca closes the Codex child

A Codex chat that Orca itself closed (a journal sink that could not take a
frame, or a forced close) said "Codex stopped while this response was in
progress", as if Codex had exited on its own.

Orca's own close path now reports hostFault. providerExited is left to the
app-server connection's exit callback, which the connection withholds while Orca
is closing the child, so it only ever reports the child's own exit.

* fix(native-chat): a chat whose history is damaged reads "Unable to load this chat."

* fix(native-chat): route the conversation-outlives-agent writers through the typed refusals

Three writers that arrived with the merge wrote refusals the old way:

- An operation that starts the agent itself, such as a goal change, turned any error the start
  threw into a refusal whose message was Orca's own error text, which released clients print.
  It now logs the error and says only that the agent couldn't restart.
- An option picked while the chat is at rest, for a key the provider would not accept, is refused
  with the rejected-option reason, like the same pick on a running agent.
- A restart continuation whose agent was refused a start filed the rejected message's sentence as
  the failure's reason. It files the refusal's code with its details again, which is what the
  restart-failure guidance keys on.

* fix(native-chat): a start Orca stopped because it never finished reads as that

The idle sweep now stops an agent whose start never finished and rejects the messages waiting on
it. The chat read "Orca ran into a problem, so this didn't go through. Try again." for that,
because every host stop was worded as Orca's own fault. It now reads "Codex never finished
starting, so Orca stopped it." in the chat's row and on each rejected message, carried as its own
failure kind so newer clients can tell it apart. The message counts as failed, like any start that
did not land.

* test(native-chat): pin the merged close and host-stop rows to their typed facts

The merge left two expectations on the old words: the close tests looked for the marker a close
used to write, and the host-stop test for the host-fault sentence. A close now writes "The chat
closed before this message was sent." with its fact, and a host stop the hostStopped sentence the
constructor gives, whatever reason the stop carried. Also folds the conversation command's two
imports from send preparation into one.

* fix(native-chat): a read of a chat this host cannot open says why

Reads now reach a chat through one accessor, which refused a missing record and a provider this
host does not run as a bare code with nothing beside it. Revealing the same chat already names
those reasons, so a client could tell "this chat no longer exists" and "update Orca" apart there
but not on the history or subscribe read that follows. The accessor now throws the same typed
refusals. The wire code and message are unchanged; the reason rides only in the error's data,
which released clients ignore.

* test(native-chat): a Claude retrying past the idle window keeps its conversation open

Every api_retry frame publishes the journal, and that publish is the activity the idle sweep reads, so a retry run revised into one row still renews the clock on each attempt.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-28 01:53:11 -07:00
Jinwoo Hong ba858ee446 feat(mobile): the keyboard covers the page like a native screen, and the shell says its height (#23110)
The shell no longer shortens the WebView for the keyboard; it publishes the keyboard height like the safe-area insets, so native's keyboard lift, refit hold and dismiss key run on the page unchanged. Keyboard and inset arithmetic read the shell's OS through a host-os seam. One page-version floor (manifest pageVersion, shell floor 1) replaces per-feature accept negotiation; a page below the floor gets the existing update wall, a desktop with no bundle keeps native screens. iOS shell drops the form accessory bar and its own keyboard observers. Native session screens untouched.
2026-09-28 04:07:35 -04:00
Jinwoo Hong 2077956254 fix(mobile): size a terminal's first subscribe from the document's reported cell box (#23080)
* fix(mobile): size a terminal's first subscribe from the document's reported cell box

#22960 sent phone dims on a terminal's first subscribe by opening a throwaway
empty terminal (init 80x24 ""), awaiting its ready and measuring, behind a
per-document first-subscribe mark whose lifetime was tied to web-ready. That
cost a second xterm/WebGL instance and ~150 ms per open, plus lifecycle state.

The document now measures the cell box without a terminal (xterm 6's
CharSizeService strategy, rounded as the renderer rounds it) for every
text-size preset and reports it with its viewport in web-ready; a table,
because the text scale only reaches the document after that notify. Each
init's ready reports the box xterm actually laid out, which replaces the
probe's entry. The controller answers fitDimensions/measureFitDimensions from
that table and the view's layout with no message; without a table it asks the
document as before.

The session seeds an unmeasured viewport synchronously in subscribeToTerminal,
so the first subscribe carries dims by construction. Deleted: the empty init,
its awaitReady gate, deferFirstSubscribeUntilViewportMeasured and the
subscribedDocuments mark. The fit pass is unchanged and still covers a
document that reports no cell box.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): read the reported cell box through in-narrowing, not Reflect.get

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): correct the probe's cell-box guess from the box xterm lays out

The web-ready probe is a guess: building the WebGL addon creates no context,
so a context that fails on load lands on the DOM renderer, whose width is not
snapped and depends on the column count. Before, a ready box that differed was
only logged; the first subscribe had carried the wrong column count, the host
echoed it, the fit pass saw the viewport equal to the host's dims, and the grid
stayed slightly shrunk. The store also kept the WebGL width after a context loss.

The document now reports the box xterm laid out whenever it changes (from
onRender, which covers a renderer swap and a DPR change that
onDimensionsChange does not fire for, and at ready). The store replaces the
guess; when that changes the current text size's entry, the view calls
onCellBoxChange with xterm's grid and the session re-fits, running the bounded
fit pass if the dims moved (one resubscribe). Equal boxes do nothing.

The RN layout box now survives a document reload; the document's own viewport
only stands in until the view reports a layout (on the page, web-ready arrives
first). The mismatch console.log is gone, and the probe's rounding names the
xterm version it copies.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): correct each cell-box guess at most once, so a DOM renderer cannot loop

On the DOM renderer the cell width is the rounded canvas width divided by the
column count, so every re-init at new cols reported a new box. Each one counted
as a correction, a floor over floats could flip the fit between two sizes, and
each flip landed converged, which reset the resubscribe budget: an unbounded
series of full-snapshot resubscribes.

Only the first laid-out box for a guessed text size may be a correction; later
reports still update the store, so fits stay truthful, but never resubscribe on
their own. The fit's floor gains a 1e-6 epsilon so floating-point error at an
exact boundary cannot flip a column or row. New document tests pin the render
report after a renderer swap and the report at ready for a paused renderer.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): make xterm the only terminal cell measurer

The document builds its real terminal before web-ready, at the app's
text scale, and reports the box xterm laid out; the first init reuses
that terminal. The page-side prediction, the per-scale guess table and
the once-per-document correction are gone. The app remembers the box
per text scale for its lifetime, so a later open at a known scale
subscribes with phone dims at once. A box that changes at the same grid
(renderer swap, pixel ratio) refits the open terminal in place.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): keep commands queued before the terminal WebView first loads

A subscribe sized from the stored cell box can queue init before the
native WebView reports its first load start, which cleared the queue
and left the terminal blank. Only a reload now drops queued commands.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): re-init a document that lacks the subscription's init, and fit one frame width

- Web-ready now says whether the document holds the terminal's latest init
  (a reload before the first ready drops a queued one); the session
  resubscribes any initialized terminal whose document lacks it.
- One grid fit, shared by the app and the document, fed the unrounded frame
  width React Native laid out; it keeps exact fits whole at fractional pixel
  ratios. The document's viewport-width fits are gone.
- The page builds every document at the scale the view mounted with, as the
  native WebView does.
- A new document's first cell box is compared against the grid the
  subscribe fitted from the stored box.
- The terminal built before ready stays hidden until its first init.
- The cell-box census matches glyph-measurement techniques, not names;
  the store's unused clear() is gone.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): build the terminal before ready only for the view shown at mount

A session mounts one terminal view per tab, and each built xterm and a
WebGL context before ready: 20 tabs made 20 contexts at load, past the
~16 a page (or Android's shared WebView renderer) holds, and native logged
32 context losses. Only the view shown when it mounts builds early now;
the rest build at their first init as before. Deferring the WebGL addon
instead would change the reported box: the DOM renderer lays out 7.8x15
where WebGL lays out 7.667x15 at the same font.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): write a WebView document's start values into its page, not an injected script

Android ran the pre-content injected script after the document's own in
1 of 22 documents on the emulator; that document started with no text
scale or shown flag and built a terminal it should not have. The values
now sit in the page ahead of the document script, one source object per
start pair so a render never reloads the WebView.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): the pre-ready terminal measures and reports while hidden

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): measure only the laid-out frame, and refit on a new grid, not a new width

- A measure needs both of the frame's dimensions from React Native; the
  document's viewport-height fallback is gone, and before the first layout
  the handle answers no fit without asking the document.
- A frame width change that still fits the PTY's grid from the stored box is
  a no-op, so sub-pixel layout jitter no longer re-measures. The width ref is
  written in that effect rather than during render (react-doctor).
- One "last grid" ref: the last reported grid, or the one a subscribe fitted
  from the stored box.
- The page render rig measures through the frame it laid out, as the session
  does, and lets the replay's fit settle before its resize-refit witness.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): let only the current terminal document's ready flush

A reload kept the WebView and its onMessage, so the old document's late
web-ready flushed the queue into the reloading view and the new document
got a second init. Each document now gets its own view (keyed on a
generation the controller owns), every notify carries the generation of
the view that received it, and a web-ready from a replaced document
flushes nothing and stamps nothing.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): drop every notify from a replaced terminal document

One rule at the receive boundary: a notify from any generation but the
current one is dropped, whatever its type, not only web-ready.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): make fitDimensions a pure question; name each generation counter

- fitDimensions no longer records the grid. A width change to a new grid
  asked it first, so the DOM renderer's report of that grid's box read as
  "same grid, new box" and refit again. Only the first-subscribe seed
  (seedFitDimensions) records the grid the document's first report is
  checked against.
- viewGeneration counts the views, readyGeneration counts web-readies.
- replaceDocument no longer resets the load flag; the load-start reset
  stays as the guard for a view that reloads itself.
- The name-based lifecycle census is replaced by a behavioural test.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): typecheck the handle mocks, drop the unused cell-box get

- The two handle mocks carry both fitDimensions and seedFitDimensions,
  and the fake-timer acts return nothing, so the three test files check
  under tsconfig.test.json again.
- terminalCellBoxes.get had no product caller; the store's tests assert
  through fit.
- The load-start comment says what the controller does now.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): hold the grid the document has, ignore a replaced view's load start, dispose a failed pre-ready terminal

- The document reports a new grid even with an unchanged box, so an
  in-place reflow on WebGL is held before a later renderer swap at that
  grid; the swap then refits. The app's apply paths do not hold the grid
  themselves: the DOM renderer's box follows cols, and a grid held on
  apply would read its own box as a renderer change and loop. One
  writer (holdGrid) holds the seeded or reported grid.
- A load start from a view a replacement unmounted is ignored, as its
  notifies already are.
- A terminal whose open throws before ready is disposed, not only
  unreferenced.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): ignore every native event from a replaced terminal view

One wrapper binds each WebView lifecycle event (load start, error, HTTP
error, render process gone, content process terminated) to the view's
generation, so a replaced view's late event cannot reset, replace or
put an error over the current document.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): a DOM seed refits once on its first report, not on the refit's own

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): subscribe a terminal only after its document is ready

The document still builds its terminal before ready and reports the cell
box xterm laid out in web-ready; the app now subscribes after that ready
and fits from that box, so nothing is sent to a document before it is
ready. Everything that made a pre-ready subscribe safe goes: the
app-lifetime box store, the seed fit, the per-document view generations
and their event filtering, the init tracker and the hasInit resubscribe.
The native view reloads in place again and web-ready keeps main's reload
rule. Boxes are kept per view; the grid a document last reported still
guards the in-place refit against the DOM renderer's cols-dependent box.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): hold one reported cell box and the grid the subscribe fitted

The controller keeps only the box the current document last reported,
not a per-text-size store: the document re-reports on a scale change.
The subscribe after ready fits from that box and holds the grid it
fitted, so the DOM renderer's first report at that grid (a new box)
refits once in place and converges; refit and apply paths hold nothing.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): fit only a ready box at the app's scale; forget a reloaded document's box and grid

A reload keeps the document's mount scale, so a ready after a text-size
change reports a box at the old scale; that box no longer sizes the first
subscribe, which then takes the no-box path. A readiness reset drops the
old document's box and held grid, so the new document's first DOM report
at the same grid does not refit.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): give the terminal document its frame at init, and fit text scale over it only

A subscribe sized from the ready box sends no measure, so the document
had no frame when the text size changed and reported the pre-refit row
pitch. The app's init now carries the frame it laid out, in the fields a
measure uses; the router takes it from either. The text-scale fit reads
only that frame, with no viewport fallback.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): say why a frameless text-scale change skips the resize

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): one cell box per terminal notify, not an array

web-ready and cell-metrics carry `cellBox: {fontScale, cellWidth,
cellHeight} | null`; the document's `laidOutCellBox` returns one or
null and the parser validates one object. The text-scale match moves
from web-ready into `handle.fitDimensions`, the one place a box is fitted.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): fit terminals in the app from the reported box; drop the measure round trip

The app already holds the box the document reported, so the refit and
the fit pass await the init's ready and call `handle.fitDimensions`
instead of posting `measure` and waiting on `measure-result`. The
document's measure, its retries, and the measure promise and timeout go.
The document still resizes locally on a text-size change, so every grid
the app sends (init, resize, reflow) carries the laid-out frame it was
fitted to. `holdSubscribedGrid` replaces `subscribeFitDimensions`, so
the only fits are `fitDimensionsFromCell` and `handle.fitDimensions`.
The render rig reads its fit from the ready box. The recorder adapter
mounts the new handle with the same recorded effects; the goldens it
mounts move on their adapterSha256 header only.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): keep the terminal frame in one ref, and notify a new width imperatively

The session held the frame in a height ref, a width ref, a width state
and the refit's own width ref. It now holds one `terminalFrameRef`
({width, height} | null until the first layout; a hidden 0x0 layout
keeps the last box). onLayout notifies a new width imperatively, as it
does height, and the refit's notify skips a width whose fit is the grid
the PTY has. `terminal-frame-width-refit.ts`, the width state and its
effect go. The subscribe's layout gate reads "no frame yet" directly.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): subscribe a held-back terminal on the frame's first layout only

`handleTerminalFrameLayout` ran on every onLayout; it now runs once, when
the frame first has a size. Later layouts only notify a new width.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): size the first subscribe inline in subscribeToTerminal

`sizeTerminalViewportFromCellBox` wrapped five lines in a 37-line
module; the subscribe now fits the ready box against the frame, holds
that grid and records the diagnostic itself. The helper's tests fold
into the subscription tests, which move to the subscription's name.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): drop the unreachable font-size guard on the reported cell box

xterm 6.1.0-beta.303 updates the render service's cell box in the same
task that sets `options.fontSize`: CharSizeService.measure fires
onCharSizeChange, and RenderService.handleCharSizeChanged runs the
renderer's `_updateDimensions` (DomRenderer.ts:359, WebglRenderer.ts:229).
`term.onRender` fires from RenderService._renderRows after the rows
are drawn (RenderService.ts:213, CoreBrowserTerminal.ts:538), and the
document writes its text scale and the font size in one task
(text-scaling.ts applyTextScale, terminal-init.ts init). So no report
can read a box between the font and the scale; the guard and its test
go. A new test pins the real order: no report when the font is set,
the new box at the new scale on the next render.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): one start seam, no source cache, the reported box as an object

- `useState` already pins each view's WebView source at mount (a new
  test re-renders at another text scale and gets the same object), so
  the module-level `webViewSources` Map goes.
- `initialTextScale` and `buildsTerminalBeforeReady` become one
  `start(): { textScale, shown }` seam.
- `reportedCellBox` holds the last reported box and grid, not a string key.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the RPC recordings to this branch and re-record

The terminal refit now fits in the app from the reported box and reads
one frame ref, so the recorder's terminal adapter mounts the new handle
(`awaitReady` + `fitDimensions`) and options (`terminalFrameRef`),
keeping its recorded effects. `baseline` is repinned to 21954dbd2f, the
last commit to touch a fenced path, and every golden is re-recorded.
Proof by class against HEAD: 787 header-only, 0 body moved, 0 added,
0 deleted. Header keys moved: `baseline` on all 787, and `adapterSha256`
on the 14 goldens `terminal-mount-adapters.ts` mounts (query-reply 3,
accessory-raw-send 4, takeover-report 4, viewport-refit 3). No recorded
traffic moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): hold the reported cell box and its grid in one ref

The controller kept the box in `cellBoxRef`, the grid in a string
`lastGridRef` and wrote it through `terminal-held-grid.ts`. One
`heldRef` now holds `{ cellBox, grid }`, as the document's own
`reportedCellBox` does: web-ready writes the box, every cell-metrics
report writes both, `holdSubscribedGrid` writes the grid, and a
readiness reset clears it. Same write points, so the one-refit bound
holds; the DOM-loop and refit-once tests pass unchanged.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the RPC recordings to the hold-rule commit and re-record

H (f00bebba48) touched a fenced path after the last repin, so
`baseline` moves to it and every golden is re-recorded. Against the
corpus before this branch's refreshes (21954dbd2f): 787 header-only,
0 body moved, 0 added, 0 deleted; `baseline` on all 787 and
`adapterSha256` on the 14 goldens `terminal-mount-adapters.ts` mounts.
Against the previous refresh: `baseline` only. No recorded traffic moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the RPC recordings to the main merge and re-record

The merge (b3b1b0def2) is the last commit to touch a fenced path, so
`baseline` moves to it and every golden is re-recorded. Against
97b5bb2b9a: 787 header-only, 0 body moved, 0 added, 0 deleted;
`baseline` on all 787, and `adapterSha256` on the 14
session.diff-review-actions goldens whose adapter #22951 edited. Against
origin/main: 787 header-only, 0 body moved/added/deleted; `baseline` on
all 787 and `adapterSha256` on this branch's 14 terminal goldens.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): the terminal document holds the grid and decides each refit

The document already kept the last reported box and grid; the app kept
a mirror of both to decide the refit. Now the document decides: its
`cell-box` notify carries `{ cellBox, refit }`, sent only when the box
changes, with `refit` a box that changed at a kept grid. web-ready
records the pre-ready terminal's box at its 80x24 grid, and the first
init that reuses that terminal holds the init's grid, so the DOM
renderer's first report refits once, as the subscribe's hold did. A
re-init no longer clears the record, so a new renderer at the same grid
still refits. The app keeps one `cellBoxRef` and `holdSubscribedGrid`,
`heldRef` and the grid on the notify go. The one-refit, DOM-loop and
renderer-swap tests move to the document with the same scenarios.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): one init options object, and a frame on every grid

`init` takes `{ cols, rows, data, preserveScroll, oscLinks, frame }`
instead of six positionals, and `init`, `resize` and `reflow` (handle
and messages) require `frame: TerminalFrame | null`. The refit's reflow
check reads `!dims` alone, and the controller's test file is named for
the `cell-box` notify it now covers.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): one notifyTerminalFrame for the frame's layout

The frame's onLayout made four calls and held the classification
itself. It now calls `notifyTerminalFrame({ width, height })`, and the
session's terminal-webview hook keeps the one frame ref, notifies the
height, subscribes the document held back for the first layout, and
notifies a later width change. `handleTerminalFrameLayout` is named for
what it does: `subscribeIntendedActiveTerminal`. The layout tests move
to that hook.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the RPC recordings to the round-8 head and re-record

85d421963c is the last commit to touch a fenced path. Against
a676c1b65a: 787 header-only, 0 body moved/added/deleted, `baseline`
only. Against origin/main: 787 header-only, 0 body moved/added/deleted;
`baseline` on all 787 and `adapterSha256` on this branch's 14 terminal
goldens. No recorded traffic moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): name the init option initialData, as the message does

The init option `data` becomes `initialData`, the message field's name,
so the controller passes it through unrenamed. The `preserveScroll` why
stays on the message type only, and the document test's title names the
three grids that carry the frame.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the RPC recordings to the round-9 head and re-record

486566c82b is the last commit to touch a fenced path. Against
3371c39715: 787 header-only, 0 body moved/added/deleted, `baseline`
only. Against origin/main: 787 header-only, 0 body moved/added/deleted;
`baseline` on all 787 and `adapterSha256` on this branch's 14 terminal
goldens. No recorded traffic moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* ci: rerun checks against main with #23560 landed

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-28 03:26:49 -04:00
Brennan BensonandClaude 7a24d3d335 fix(native-chat): the conversation outlives its agent (#22835)
* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* fix(native-chat): name a chat write by its target, not the owner generation

A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.

Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.

Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.

* fix(native-chat): every journal append reaches the chats that are open

A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.

A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.

* test(native-chat): an epoch replacement reaches the open chat

* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

* perf(native-chat): a publish behind a delivered commit reads nothing

Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.

* test(native-chat): state why the teardown test's fake journal is safe to cast

* docs(native-chat): say mutation admission checks only the writer lease

* docs(native-chat): drop the send rebase from comments that still described it

* fix(native-chat): a message is accepted, then delivered

A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".

A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.

Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.

A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.

Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.

* fix(native-chat): settle queued messages only for the child that ended

A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.

A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.

The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.

* fix(native-chat): an adoption that fails to import keeps the conversation open

The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.

* perf(native-chat): the recovering open reads the journal once

Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.

* fix(native-chat): an attach that fails after indexing its child leaves no child behind

A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.

* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer

The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.

A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.

* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down

The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.

* fix(native-chat): a message rejected while its chat was closed reads as not sent

A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.

* test(orchestration): name why the readiness settlement fakes are cast

* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent

* docs(native-chat): drop the fence from the admission the send effects run behind

* docs(native-chat): give the fence move on release the reason that still holds

* docs(native-chat): stop citing a write fence check in launch and mailbox comments

Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.

* refactor(native-chat): the provider child is its own record

A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.

- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
  patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
  own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
  the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
  dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
  is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
  the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.

* fix(native-chat): the delivery loop alone settles a message its start or child failed

A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.

- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
  reads how it ended: a Stop continues; anything else writes one failure row and rejects every
  queued message with the same words, then stops. A child still starting whose start the adapter
  says did not land fails the same way. The exit, eviction and the settlement retry only settle
  the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
  nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
  failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
  closed, with or without a child, and a start the loop already has in flight is waited for so the
  child it produces is stopped rather than left behind.

* refactor(native-chat): a stopped child ends on the one reading of its stop

The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.

* feat(native-chat): the host says it accepts a send before any agent has it

The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.

* refactor(native-chat): an attach never opens a journal of its own

The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.

* fix(native-chat): a moved fence resends nothing on a host that accepts first

The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.

The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.

* refactor(native-chat): a child's end says whether the user or the host stopped it

The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.

* fix(native-chat): a chat whose only work is a queued message is not offered for resume

A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.

* fix(native-chat): the conversation outlives its agent

Opening a chat no longer starts its agent. A conversation is reached through one host
accessor that opens its journal at rest, and a send is what starts the agent, through
the delivery loop. One idle sweep, every five minutes, stops an agent that has been
quiet for thirty minutes and owes no work, then drops an open journal handle that is
only a cache. Its record, tab, status row and readers stay.

- hold and release are no-ops; hold still builds the host for shipped mobile builds.
- The holders, the holds, the release clock and the exit respawn are deleted.
- Options, the model list, the goal and the context meter answer at rest; a model pick
  at rest is recorded as intent for the next start.
- Compact, rewind, clear and goal changes start the agent first. A send does too when
  a rewind is still in doubt after the conversation opens.
- Orchestration routes mail and group addresses on ownership (the record plus the chat
  tab), not on whether the process runs. An open dispatch keeps its worker running.
- The restart continuation is a send; Resume all holds each slot until the message is
  handed over or rejected.
- A read error never replaces a loaded transcript, and shows the host's own words.

* test(native-chat): type the queued-message fixtures in the resume-offer tests

* fix(native-chat): a start that dies while a message waits on it is that message's failed start

Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.

* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now

* test(native-chat): pin what a failed start settles, and what a resume offer names

A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.

* test(native-chat): the failed-start pins fail on what the message became, not on a timeout

* fix(native-chat): a restart offer ends when the chat's agent starts again

The offer used to end only when the chat's newest user message changed,
because opening a chat started its agent and that start could not be told
apart from real activity. Opening a chat starts nothing now, so the host
reads the fact it already publishes: a chat's status row goes from not
host-owned to host-owned exactly when its agent is started. At that edge the
offer and any failure record for the chat are withdrawn, unless the start is
a resume action's own (its continuation is the oldest undelivered message).

A continuation and a message racing to be first are decided at acceptance:
the continuation is refused, quietly and with nothing filed, when any other
message was accepted since the restart. A failed continuation start leaves
the offer retryable, and each resume action sends its own message id.

Deleted: the newest-user-message comparison, its journal reader, the
continuation filter, and the failure ledger's own "answered by the chat"
check. The marker still carries its message id for one release, so the
previous build can read it.

* fix(runtime): end a transcript stream when its client unsubscribes

Desktop: the IPC subscription controller was dropped as soon as the streaming
handler returned, which for most streams is right after it binds. A later
runtime:unsubscribe then found nothing to abort, so the host kept the subscriber
and derived and sent every publish to a channel no one listened to. The controller
now lives until the renderer unsubscribes, resubscribes the same id, or goes away.

Mobile: disposing an agentSession.subscribe stream now sends agentSession.unsubscribe
with the stream's frame id, so the host ends that subscriber and leaves a sibling
stream on the same socket running. The direct path now passes the frame id the relay
path already passed.

* test(orchestration): the preamble's host stub is typed, not cast

The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.

* fix(native-chat): one fact ends a restart offer: the chat moved on since the restart

The offer is live while no other message has been accepted in the chat since the
restart and its agent has not proved a start since. The offer list, the resume's
reservation check and the continuation's acceptance check all read that one fact,
so a message whose start then failed withdraws the offer too, and a stale click
finds nothing to act on.

The fact is read off the conversation's open handle, which the restart closed, so
it is retired durably whenever it may have changed: a message accepted, a start
proven. A close and reopen within the same run therefore cannot bring the offer
back. A continuation rejected before it reached the agent does not count, so a
retry after a failed start still runs.

Deleted: the quit-time gate on withdrawal, which changed nothing because the
withdrawal and the quit's own offer write share one queue; the per-action
"withdrawn" flag and the separate acceptance check it paired with.

* test(native-chat): an older build reads the restart offer this build records

The offer lives in a file the previous release reads after a downgrade. Pin that
against the pinned release's own capsule, and run the lane when the marker or the
capsule changes.

* fix(native-chat): read a restart offer against where the journal stood when it was taken

"Since the restart" was read off the conversation's open handle, which the idle
sweep closes: after a reopen, a message the user had already sent looked older
than the handle and the withdrawn offer came back.

The offer now records the journal position (epoch and sequence) at the moment
it is taken, and a message accepted after that position, or a journal on another
epoch, means the chat moved on. That is derived from the journal, so it holds
across any number of closes and reopens. An older build's offer has no position;
only a start withdraws it. Because the message half is now durable, the offer is
no longer rewritten in the recovery file on every accepted message; a proven
start still writes it, since only the host that saw the start knows of it.

* test(native-chat): wait for the listing's retire write before reading the recovery file

* fix(native-chat): keep the terminal-backed chat's read error over its local echoes

Messages winning over a read error is right for the structured chat, whose read retries and whose
messages came from the transcript. The terminal-backed view assembles its list from local echoes
too (a launch prompt, a pending send), so a failed read there showed only those bubbles and no
error. Only the structured pane now keeps messages over an error.

* fix(native-chat): a start retries the exit settlement a failed journal write left owed

An agent exit whose journal settlement write failed releases the lease latched until a retry lands.
Reopening the chat used to be that retry; with reveal now only opening the journal, nothing retried
it before the next app launch, and every send was refused. The start the send needs now runs the
retry first, where the attach would.

* perf(native-chat): answer the owner check without opening the chat

Worktree activation calls agentSession.handoffStatus for every chat tab in the worktree, and the
answer comes from the session record alone. Reaching it through the accessor opened each resting
chat's journal (a full read, the crash-boundary write and a restored status publish), then kept it
open for the idle window. It now checks the record and the adapter's support, as before this series,
and opens nothing.

* fix(native-chat): a read waiting on the session lock opens nothing once quit began

The accessor checked for quit before queueing the open, so a read queued behind a session task ran
its open after teardown had begun and indexed a journal no teardown step would close. The check now
runs at the open itself.

* fix(native-chat): read a failed resume's chat before calling it retryable

Whether a failed resume is retryable is the offer's own rule: the chat has not moved on since the
restart, read from its journal. The failure list read it only for a chat already open, so once the
idle sweep closed a chat the user had moved on in, its failure showed Retry again, and the click did
nothing. The list now opens the failed chats first, as the offer list does.

* test(native-chat): type the provider event sink the settlement test reaches for

* fix(native-chat): say the structured read keeps trying only where it does

The structured pane's "Orca keeps trying to load it" line never showed: the view state filled in an
untranslated fallback whenever the read error had no text, and the empty state prefers any message.
The view state now leaves the message out, so the structured pane shows that line and the
terminal-backed pane its own translated one. Mobile's structured lane does not resubscribe after an
error frame, so it no longer makes the claim.

* test(native-chat): await the send's settlement instead of polling for the start

The at-rest send tests polled for the provider start with vi.waitFor's one-second default, which a
loaded machine outran. They now await the host's own settlement of the message.

* fix(native-chat): a restart offer resumes any time after the quit, and knows its own continuations

The continuation's message id was dated by the quit, and the ledger refuses a new id dated more than
a day back, so Resume or Retry a day after quitting was always refused (on main too). It is now
dated by the resume action.

Telling a rejected continuation from the user's own message read the operation ledger, whose rows
expire after about a day; after that a failed resume stopped being retryable. The offer now
records the continuation each action sends on its own capsule entry, bounded to the newest 16, so
the ids end with the offer. The ledger read is deleted.

* fix(orchestration): route no mail to a structured worker its orchestration released

A structured worker is routed on ownership, and a resting worker's lease is released, so ownership
held while its chat tab stayed listed. A worker the coordinator abandoned and then released, found
at rest by the release, therefore still took peer mail and @worktree: broadcasts, and each one
restarted its agent. Routing now also reads the orchestration's own resource row: once it is
released, direct mail, group addressing and worker-show's addressable answer drop the worker, as
they would a terminal worker whose terminal closed. The chat tab stays, and nothing new is stored.

* fix(native-chat): a failed retry names the user's prompt, not Orca's continuation

A resume's continuation is written to the chat before its start, so after a failed attempt the chat's
newest user message is that rejected continuation. A second failure then showed Orca's own restart
text as the chat's prompt. A retry now keeps the prompt its first failure named.

* fix(orchestration): read the released row optionally, as the authority does

worker-show's observation called the row lookup directly, which a runtime double without it threw on
and failed the structured tab-retirement release.

* fix(native-chat): the status bar drops a restart offer the chat moved on from

The renderer re-read the host's restart offer only when a failed chat showed activity, so after a
message withdrew a pending offer the host answered no chats while the status bar kept counting one,
and clicking it opened nothing. The same watch now covers pending offers: a status change in an
offered chat asks the host again, once.

* test(native-chat): a roster of idle or finished children does not keep an agent awake

The sweep reads owed background work through the shared child-work liveness that upstream's
release clock adopted; a child that went idle or finished is not work the agent still owes.

* fix(orchestration): a task dispatched into a resting structured worker keeps it running

The sweep's open-dispatch check read only the worker-start dispatch that owns the worker's terminal
resource, so a task later dispatched to the same worker (orchestration dispatch --to, which writes a
dispatch with no worker row) did not count: after thirty quiet minutes the worker was stopped while
that task was open, and its coordinator read exited. Any unsettled dispatch addressed to the worker's
process incarnation now counts, derived from the existing rows.

* docs(native-chat): comments stop describing the hold this PR removed

Eight comments still justified orderings and teardown choices by a viewer or dispatch hold that
pinned the provider child. Nothing holds any more; the orderings stand for the binding's redrive
subscription and parked mail, and a chat's agent runs from a send until the idle sweep rests it.
Comment-only.

* fix(native-chat): a restart offer keeps the start its own continuation made

Whose start ended an offer was decided at read time, from whether the offer's continuation was
still the queued message. Once the provider refused that continuation, the child it had started
read as someone else's start, so the offer ended and its failure showed no Retry. The delivery
loop now records which queued message a start is for on the in-memory child, and the child's end
carries it; the offer counts a start as its own when that message is one of its continuations.

* fix(native-chat): an agent gets a full idle window after its owed work ends

The sweep measured quiet only from the last journal row, so once a subagent, command, monitor or
dispatch that had outlived the window ended, the agent was stopped at the next tick. A child can
read done before the lead's wake-up turn writes anything, and stopping in that gap loses the
wake-up. The sweep now counts owed work it observes as activity, which gives the agent the full
window afterwards, as the release clock it replaced did.

* test(claude): the options-read fixture runs a live child

The fixture marked its conversation running with a hasProviderChild field the
session type does not have, so the read took the at-rest path and refused a
session with no record. It now carries a child, which is what the read checks.

* test(native-chat): host tests reach its collaborators through a typed seam

The rest-test rig and three test files read the host's private members with
Reflect.get and cast the result. The host now exposes one test-only accessor,
collaboratorsForTests(), and the subscribers class a subscriberCountForTests()
beside its existing retainedActivityCountForTests(), so the tests are checked
against the real types and the casts are gone.

* refactor(orchestration): one owner answers a structured worker's custody

Routing, group addressing, worker-show and the idle sweep each composed their own reading of
whether orchestration still holds a structured worker, so each new obligation or retirement state
had to be added to every reader. structured-worker-custody now derives both answers from the
worker-terminal list state coordinators see in worker-list: addressable is owned and not released,
and owed work is an active custody or an unsettled task dispatched to the same incarnation. The
owner's state is read through the remote dispatch attachment too, as the terminal transfer lookup
already does. Behaviour is unchanged; a settled worker awaiting its coordinator still rests.

* refactor(orchestration): owed work is an open dispatch on the worker's incarnation

A supervised worker's own dispatch context stays open exactly while the worker is active, so the
separate active-custody branch only repeated it. Owed work is now one fact, which also states the
policy that a worker awaiting its coordinator's decision may rest, and both custody decisions are
written once at the top of the module.

* fix(native-chat): a restart offer knows its continuations by a tag in their id

The offer recorded each continuation id in a list on its capsule entry, capped at 16, and a running
action's id in memory. Both could disagree with the journal: past the cap an old rejected
continuation read as the chat moving on, and a crash during a retry restored the failure's older
entry, which lacked the retry's id. Each continuation id now carries a tag derived from the offer
(its teardown and chat), then the action's own part, so any continuation of this offer, queued or
rejected, is recognised from the journal row and the marker alone. The persisted list, its cap and
the in-memory action map are deleted; the agent-start withdrawal keeps an offer whose own
continuation the start was for, read against the stored marker.

* test(runtime): the legacy-worker reveal test judges its stale snapshot inside the wait

The tui-idle probe reads through readTerminal, which now awaits the structured
worker check before the PTY read, so the probe's snapshot request starts a
microtask later. vi.waitFor missed it on its first check and polled again at
50 ms, the same moment the wait's own 50 ms timeout fired. The stale snapshot
then resolved after the wait had already timed out, so the test passed without
judging it, and the rejection landed before any handler was attached. Vitest
reported that as an unhandled error and failed the shard.

Polling every 1 ms sees the request within a few ms, so the snapshot is judged
while the wait is still pending.

* fix(native-chat): a view never restarts a chat whose last start failed

A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.

* fix(native-chat): the idle sweep reads owed work every tick

Owed work counted as activity, but the sweep read it only once the idle window had elapsed, so it
refreshed the clock at most once a window. Work that ended just before the next read left the
agent to be stopped at that read, moments after the work ended, which is the gap the refresh was
meant to cover. The sweep now reads owed work on every tick for a started agent, so the window
always runs from the last tick that saw work owed.

* fix(native-chat): a continuation handed to the agent stays sent

The offer read its own continuation as not reaching the agent while its dispatch was pending, which
also covered one already handed over and still unanswered. When the wait for that answer ended first,
the failure it filed read as retryable, and a retry sent a second continuation to an agent that may
have acted on the first. Only a continuation still queued, or rejected, is now read as unsent.

* test(native-chat): start the child the loop waits on with an attach, not a second view

A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.

* fix(native-chat): settle a gone generation's turn wherever a conversation opens

A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.

* test(native-chat): prove the next child's start settles the turn an earlier child left

The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.

* test(native-chat): count a failed start's rows by row, not by text

Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.

* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget

* test(native-chat): the interrupted create's own retry continues again

The merge of main's lease-latch fix replaced that test's retry of the interrupted create, under its
own operation id, with a fresh start whose result nothing read. That fresh start passes with the
released-reservation continuation deleted, so the case the fix exists for went untested. The retry
and its assertion are main's again.

* docs(native-chat): three comments that still had views starting agents

A start with nothing queued now comes from a command, goal change or rewind; an interrupted compaction
left alone would refuse every send, so no agent would ever start to finish it; and a current host
raises the unattached read refusal only once quit began, with the attach window belonging to an older
host.

* test(native-chat): pin the open's and the send's start and row counts, however the view binds

Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.

* test(native-chat): a reader's open settles the turn a failed exit settlement left running

An exit whose settlement write failed leaves its turn running in the open journal. PR 1's open now
settles it, and this pins the two reads that reach it here: a reader reopening a chat the idle
sweep closed, and a read that opens the chat before the restart restore reaches it.

* test(native-chat): the view-start test's starting window outlasts two subscriptions on a loaded runner

A subscription reads the conversation before it returns, so under load the two views took longer
than the create child's 300 ms start, which then exited before the test checked that it had not.
The child now takes a second to fail.

* fix(native-chat): settle a gone generation's turn at every open but an acquisition's

The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.

* test(native-chat): hold the create's start open until the views bind

The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.

* test(native-chat): a read that reaches a crashed chat before the startup reconcile settles its turn

On desktop the chat on screen at relaunch reads before startup reconciles the leases, while the
dead process's lease still reads live. The open settles the turn it left running anyway, and the
restore that follows finds it settled.

* refactor(native-chat): drop the composer's second error formatter

After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.

* test(native-chat): pin the reason on a message rejected while its chat was closed

The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.

* docs(native-chat): drop the removed dispatch hold from six comments

A worker's session no longer takes a dispatch hold, and no release clock
rests a chat by visibility; the agent-launch comments, the abandon test,
the teardown test and the refusal census still said so.

* test(native-chat): rest the owner-status chat through the idle sweep, not a hold

The activation-gate test from #22808 put its chat at rest by holding and
releasing it, and passed the release-clock grace. This branch deleted both,
so the case threw before it reached its assertions. It now moves the host's
clock past the idle window and lets the sweep stop the agent and close the
conversation, then asserts the same owner answer and activation gate.

* fix(native-chat): show the structured pane's retrying line when a read fails

The read transport always hands the pane the host's words, so the error
state's "Orca keeps trying to load it" line, which showed only when there
were none, was never seen: the pane showed the host's text twice, as its
subtitle and on the status line under it. The structured pane now always
says its read keeps retrying, and the host's text stays on the status line.
The terminal-backed chat is unchanged.

* test(native-chat): wait for a send's background start before the refusal oracle removes its store

An accepted send wakes the delivery loop, which starts the agent in the background. The oracle's teardown disposed the loop but did not wait for that start, so its lease write could create a temp file in the store directory while the directory was being removed, failing the test with ENOTEMPTY about one run in four. The teardown now drains tracked starts before it closes the journals.

* fix(native-chat): a start a message waited on gets one failure row, the delivery loop's

When a queued message's start failed, two writers could report it under the same row: the delivery loop, when the adapter settled the start without proving it, and the exit settlement, when the child's exit landed. The last one won, so the chat's row could name a different cause than the one the message was rejected with, or be written twice.

The exit settlement now writes the start's row only when no message is queued and the loop has not already recorded that start. A start for a command, goal change or rewind, with nothing queued, still gets its row from the exit.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-27 23:46:14 -07:00
Brennan Benson 56e691344f fix(native-chat): show the "Working for" bar while a turn runs (#23537)
* fix(native-chat): show the "Working for" bar while a turn runs

The turn bar under the user's message only rendered once a turn settled, so a
running turn had no bar, only the spinner line at the tail. The running turn
now draws the same bar with a live clock, and it settles in place to "Worked
for". The tail line keeps its spinner but drops the clock, so the clock is said
once: activity text, else "Thinking", else "Working…". Mobile gets the same split.

* fix(native-chat): keep the turn bar mounted through the settle

When a turn ends without a host-recorded duration (a send folded into a
running turn, or a host that stamps no start), the local duration is stamped
one pass after the working flag drops. For that pass the turn status resolved
to nothing, so the live "Working for" bar unmounted and remounted as "Worked
for" - invisible on desktop (layout effect) but a painted blink on mobile,
where the stamp runs in a passive effect.

The shared selector now keeps the just-ended turn's running status until its
duration lands, unless the host says it never saw the end, and mobile reads
the active turn's row from that selector the way desktop does.
2026-09-27 23:39:43 -07:00
Neil 45f3512a33 feat(agents): add first-class DeepSeek Harness (dsh) support (#22468)
* feat(agents): add first-class DeepSeek Harness (dsh) support

Register DSH as a supervised Orca agent: catalog entry and detection for its
dsh-tui profile, status/question hooks through DeepSeek's own Claude-Code hook
bridge, composer-ready prompt delivery, session resume, headless Source Control
AI, and title identity that no longer collides with Gemini's.

* fix(dsh): reach Orca through DSH's credential scrub and stop reading its title as Gemini

DSH runs command hooks through its own shell executor, which drops every env var whose
name contains KEY, TOKEN, SECRET or PASSWORD — taking ORCA_PANE_KEY and
ORCA_AGENT_LAUNCH_TOKEN with it, so every hook exited without posting. Mirror both onto
scrub-safe aliases at spawn and restore them at the top of the DSH hook script.

Its title collided too: DSH rests on the same glyph Gemini works on, so a resting DSH
pane was relabelled Gemini CLI and reported working forever. Defer both the Gemini
classifier and the title status detector on DSH's whale, in the base module both copies
of that classifier read.

* test(mobile): repin the session-route closure for the DSH agent icon

* fix(dsh): address review — never splice user rows, cover remote panes, keep the diff off argv

- findManagedDshPatchRegion paired an orphan start marker with a later block's end, so a
  truncated write made install/remove delete the user's own rows. Pair each end with the
  nearest preceding start; regression test fails without the fix.
- The relay PTY env builder never applied the scrub-safe aliases, so remote DSH status
  silently never appeared even with the remote hook installed.
- Source Control AI sent the whole diff on argv; send it over stdin with DSH's '-' marker.
- dsh-tui/dst already chose the interactive profile, so a workspace folder named 'web' or
  'plugin' no longer marks a live agent pane non-interactive.
- Isolate USERPROFILE as well as HOME so a Windows run cannot edit the real home.
- Drop the duplicate README badge and revert an incidental doc reformat.

* refactor(dsh): share the managed-hooks reader and tighten the new modules

Reuse before reimplementing: readManagedDshHookEvents was a near-verbatim copy of Muse's,
with byte-identical private helpers. Both now call one readManagedHookEventsFromJson.

Also: one readTextOrAbsent instead of two spellings of the same read (dropping an
existsSync TOCTOU), one status() builder instead of four inline literals, rmSync(force)
instead of exists-then-unlink, and a redundant empty-string guard before JSON.parse.
The patch-file transforms lose their index juggling for a predicate plus a filter.

* fix(dsh): refuse a flow-style patch file, keep its mode, and stop the relay inheriting a pane

- applyManagedDshPatch matched only an exact `[]`, so `[] # keep empty` or a non-empty
  flow sequence got a block entry appended after it — invalid YAML that would leave DSH
  unable to load the user's own patch layer either. It now strips the token from an empty
  sequence (keeping a trailing comment) and returns null for a non-empty one; install
  reports that and changes nothing.
- The patch rewrite dropped an owner-only file to the umask default (CWE-732); pass
  preserveMode.
- The relay PTY env never dropped inherited pane identity the way the local and daemon
  builders do, so a spawn that specified none could inherit the relay's own and every
  agent's hook would report against that pane.

* fix(dsh): keep the flow-style refusal in every status read, and scope the mode test to POSIX

A refused patch file carries no managed region, so getStatus() fell through to a bare
not_installed with detail null — the actionable 'rewrite it as a block sequence' message
only ever reached the one-shot install() return. Export the predicate and check it first,
behind one shared message constant.

The owner-only mode assertion cannot hold on Windows, where chmod only toggles the
read-only attribute and mode & 0o777 reads 0o666 for any writable file.

* docs(readme): restore the DeepSeek Harness badge lost in the rebase

* test(mobile): repin the session-route closure to the measured 4221

Measured, not derived: 4220 without the DSH icon entry, 4221 with it. Two of the three
modules above main's 4218 pin are not this change's — they arrived with the mobile work
after #22570 and were never repinned; the changelog records that split explicitly.

* fix(dsh): settle tui-idle on the agent's own hook, so supervised workers see it ready

Reported by a tester on the adhoc build: `terminal wait --for tui-idle` ran to its 90s
timeout against an already-ready DSH composer, so a supervised worker never sees the agent
as ready.

Every existing tier reads the title, and DSH deliberately carries no title status: its rest
prefix is Gemini's working glyph, so the detector reports none. A fresh first-party `done`
is better evidence than any title anyway — it is the agent's own account of its own turn,
and normalizeDshEvent drops subagent events, so it is the lead's. Scoped to DSH: for agents
whose hooks report child turns, a mid-turn `done` is the #6011 class this file prevents.

* test(daemon): record the DSH transcript's true-colour I2 divergences

Adding the dsh-tui capture to __fixtures__ enrolled it in the serialize replay sweep, where
it reports 10 I2 divergences and failed the unlisted-transcript default of 0.

Every one is the same shape — visible-grid row=0, a 24-bit background the round trip does
not restore to default — which is DSH's whale intro painting whole rows of true colour.
Verified as an upstream limitation rather than a regression by replaying against the
previous build (build-serialize-addon-at-ref.mjs --ref origin/main): I1 and I3 both hold.

* fix(dsh): return the new tui-idle verdict from the first-party done lane

Main refactored isTuiIdleSatisfied into evaluateTuiIdle, which returns a verdict rather
than a boolean. The DSH lane still returned `true`; it is tier-1 positive evidence, so it
returns READY_STRONG like the title/body lane above it. Re-verified the regression test
still fails without the lane.

* test(relay): pin the scrub-safe pane-identity aliases on the relay spawn path

The relay builds a remote pane's env itself, so the alias mirroring there had no
test: removing the call left every suite green while remote DSH status silently
vanished. Both cases fail without it.

* docs(dsh): point the hook service at the integration reference

The reference doc had no inbound link from anywhere in the repo.
2026-09-27 22:44:18 -07:00
Brennan BensonandClaude 85067494a1 fix(native-chat): a request that failed reads as failed (#22944)
* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* fix(native-chat): name a chat write by its target, not the owner generation

A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.

Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.

Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.

* fix(native-chat): every journal append reaches the chats that are open

A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.

A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.

* test(native-chat): an epoch replacement reaches the open chat

* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

* perf(native-chat): a publish behind a delivered commit reads nothing

Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.

* test(native-chat): state why the teardown test's fake journal is safe to cast

* docs(native-chat): say mutation admission checks only the writer lease

* docs(native-chat): drop the send rebase from comments that still described it

* fix(native-chat): a message is accepted, then delivered

A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".

A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.

Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.

A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.

Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.

* fix(native-chat): settle queued messages only for the child that ended

A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.

A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.

The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.

* fix(native-chat): an adoption that fails to import keeps the conversation open

The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.

* perf(native-chat): the recovering open reads the journal once

Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.

* fix(native-chat): an attach that fails after indexing its child leaves no child behind

A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.

* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer

The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.

A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.

* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down

The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.

* fix(native-chat): a message rejected while its chat was closed reads as not sent

A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.

* test(orchestration): name why the readiness settlement fakes are cast

* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent

* docs(native-chat): drop the fence from the admission the send effects run behind

* docs(native-chat): give the fence move on release the reason that still holds

* docs(native-chat): stop citing a write fence check in launch and mailbox comments

Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.

* refactor(native-chat): the provider child is its own record

A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.

- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
  patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
  own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
  the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
  dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
  is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
  the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.

* fix(native-chat): the delivery loop alone settles a message its start or child failed

A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.

- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
  reads how it ended: a Stop continues; anything else writes one failure row and rejects every
  queued message with the same words, then stops. A child still starting whose start the adapter
  says did not land fails the same way. The exit, eviction and the settlement retry only settle
  the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
  nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
  failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
  closed, with or without a child, and a start the loop already has in flight is waited for so the
  child it produces is stopped rather than left behind.

* refactor(native-chat): a stopped child ends on the one reading of its stop

The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.

* feat(native-chat): the host says it accepts a send before any agent has it

The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.

* refactor(native-chat): an attach never opens a journal of its own

The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.

* fix(native-chat): a moved fence resends nothing on a host that accepts first

The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.

The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.

* refactor(native-chat): a child's end says whether the user or the host stopped it

The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.

* fix(native-chat): a chat whose only work is a queued message is not offered for resume

A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.

* test(native-chat): type the queued-message fixtures in the resume-offer tests

* fix(native-chat): a start that dies while a message waits on it is that message's failed start

Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.

* fix(native-chat): a request that failed reads as failed

A structured chat whose only message the agent's start refused read as a
green finish, and a cancelled structured turn did too: the host published a
verdict only for turn records, and structured rows carried no `interrupted`.

The host projection now reads the session's latest request: its turn's
outcome, or `failure` for a send the agent or its start refused. A send
that was withdrawn, or left undelivered by a restart or a close, fails
nobody and makes nothing listable. The ingest publishes `interrupted` as the
hook lanes do, and every reader decodes the verdict through one accessor, so
a failure reads Failed on the dot, the rollups, history and `worktree ps`,
behaves like a cancellation in every clean-finish policy, and notifies as
"failed".

* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now

* test(native-chat): a verdict change republishes the mobile status projection

* refactor(native-chat): the store's retention trigger keeps its flag compare

A verdict change always moves the completion clock the same check already
reads, so a second verdict compare there caught nothing new.

* test(native-chat): a user message the provider journaled keeps its session listed

* test(native-chat): pin what a failed start settles, and what a resume offer names

A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.

* test(native-chat): the failed-start pins fail on what the message became, not on a timeout

* fix(native-chat): a late provider-session update keeps a failed recovery record failed

A provider-session heartbeat that rewrites a completed recovery record kept
its interrupted flag but dropped the outcome it was copied with, so a live
failed checkpoint read as a clean finish until the next status write.

* test(orchestration): the preamble's host stub is typed, not cast

The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.

* test(native-chat): the terminal-bell check asserts the renamed verdict field

The bell notification test still checked for agentInterrupted, which no
longer exists, so it could not catch a verdict leaking into a bell dispatch.

* fix(native-chat): a failed turn ranks like a completion for attention

Attention readers (completion time, Smart Sort, sticky retention, Cmd+J
Recent) now demote only a turn the user stopped. A failure is news the
user has not seen, so it keeps its completion time, ranks in the Done
class, stays retained after its pane goes away, and a retained failure
reads failed in the worktree rollup instead of done. Clean-finish
policy (hibernation, pane ownership, the value moment) still treats a
failure like a stop.

The retention trigger compares verdicts again: success -> failure no
longer moves the completion clock.

* fix(native-chat): a failed main agent reads failed while its subagents still work

The verdict is now read from the main agent's own state, not the folded
row: a main agent that is done and failed has a verdict even while its
subagents keep the row working. Without mainAgent (history, worktree ps,
older hosts) the old combined-done rule stands.

Display marks the verdict through agentVerdictDisplayMark: a failure
outranks every combined state on the agent's dot, label, tab badge,
dashboard and activity rows; a stop marks only a done row, so a
successful or stopped main agent with live subagents still reads
working. Subagent rows keep their own state. The worktree card, terminal
tab and Cmd+J rollups share one pane fold and rank a pending question,
then failed, then working, monitoring, interrupted and done.

worktree ps publishes the main agent's outcome on a working row, and the
mobile mirror reads it. The store's change check, the paired-client
mirror's equality and its epoch now see a verdict change on a working
row, which otherwise moves no state or clock and left the worktree card
reading working. Clean-finish policy is unchanged: a working row is never
hibernated and has no completion time.

* docs(native-chat): the worktree ps outcome comment no longer claims old hosts send it

The field is new: an old host sends no outcome at all, so a reader falls
back to interrupted. The removed clause said old hosts send it on done
rows, which never shipped.

* docs(native-chat): the status-store listing rule names provider-journaled user messages

* fix(native-chat): a refused send notifies failed through the completion feed

The host's completion feed followed only the newest turn, so a send the
agent or its start refused, which creates no turn, read Failed on its row
but sent no notification. The feed now follows the session's latest
request, read from the projection the status feed already makes for the
commit: a turn keeps its id, a refused send is named by its journal item
key. It announces only while the session is idle, as the row reports a
verdict, so queued sends refused one commit at a time notify once, and a
withdrawn send falls back to a request already announced.

* fix(native-chat): every copy of a row carries the main agent's own status

History entries, sleep records and `worktree ps` rows carried a flattened
top-level `outcome`, copied under different gates and without the main agent's
clock. They now carry `mainAgent` (state, outcome, stateStartedAt), the type
the live row already persists and sends, and every copy site takes it with
`interrupted` through one function, `agentVerdictFields`.

- The accessor reads `mainAgent` then the legacy flag; the mobile mirror
  matches it line for line.
- Sleep records admit `mainAgent` with `normalizeMainAgentStatusField`, so a
  malformed value drops the field, never the record.
- Mobile dates a main agent that failed under live subagents by its own clock,
  as desktop does, and its row equality compares `mainAgent`.
- The activity feed reads a history entry's own `mainAgent` instead of
  rebuilding one; the sync key and history equality compare it.

* test(native-chat): pin the worktree ps verdict across host and phone versions

Pairs the real v1.4.212 host and phone row reader with this build: an old phone
reads a new host's rows by `interrupted`, a new phone reads an old host's rows
(no `mainAgent`) the same way, and a new phone reads a failure under live
subagents as Failed, dated by `mainAgent.stateStartedAt`. The release checkout
now carries the phone's self-contained row reader, and the lane runs when the
`worktree ps` row producers change.

* test(mobile): name the parity table's row for its role

* fix(native-chat): a request that settles while the user is asked something notifies once

The completion edge waited for an idle session, and a pending prompt (including a
subagent's approval) is not idle. Structured chat has no other attention producer,
so a main turn that finished while a subagent waited on the user sent nothing
until the prompt was answered.

The edge now waits only on owed work (a running turn or an unanswered send), which
the projection reports even beneath a pending prompt. A request that settles with
a prompt pending announces once; the renderer words it "needs input" from the
host status mirror's `attention`, and answering the prompt keeps the same request
identity, so it does not announce again. The wire shape is unchanged.

* fix(native-chat): the completion says when the user is being asked

A request that settles while a prompt waits on the user was worded "needs input"
from the renderer's status-feed mirror. Remote clients receive the status and
completion streams over separate sockets, so they can arrive in either order and
the wording could be wrong both ways.

The host already knows at emit time, so the completion now carries an optional
`awaitingUser: true` in that case and omits it otherwise. The renderer words the
notification from that field alone and no longer reads the status mirror. Old
clients ignore the field and word by outcome; old hosts never send it.

* fix(worktree-status): a departed agent's failure yields to live work on the worktree card

A retained failed agent has no expiry, so ranking it with a live failure pinned the card to Failed over other panes' live work. It now ranks below working, monitoring and permission, and above every finished outcome.

* docs(agent-status): a departed agent's failure ranks below live work on the worktree card

* fix(native-chat): a view never restarts a chat whose last start failed

A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.

* test(native-chat): start the child the loop waits on with an attach, not a second view

A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.

* fix(native-chat): settle a gone generation's turn wherever a conversation opens

A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.

* test(native-chat): prove the next child's start settles the turn an earlier child left

The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.

* test(native-chat): count a failed start's rows by row, not by text

Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.

* test(cross-version): load the phone row readers without mobile's toolchain

Vite transforms a file against its nearest tsconfig, and mobile/tsconfig.json
extends expo/tsconfig.base.json, which the root-only cross-version lane never
installs. The worktree ps verdict suite imported the current phone row reader
from mobile/ directly, so CI failed with TSConfckParseError before any test ran.

The harness now imports a copy of the working-tree reader placed under the
checkout cache, where the root tsconfig applies, as it already does for the
release checkout's copy. Both readers are still the real files.

* test(cross-version): keep the checkout path-guard message and justify the copy import's cast

* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget

* test(native-chat): pin the open's and the send's start and row counts, however the view binds

Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.

* fix(native-chat): settle a gone generation's turn at every open but an acquisition's

The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.

* test(native-chat): hold the create's start open until the views bind

The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.

* refactor(native-chat): drop the composer's second error formatter

After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.

* test(native-chat): pin the reason on a message rejected while its chat was closed

The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.

* fix(native-chat): a send the provider never received after a restart has no verdict

Restart reconciliation rejects a crash-stranded send that is absent from a
trustworthy provider history with reason 'not_delivered'. Nobody failed that
send, but the verdict allowlist did not name it, so after a crash the chat
read Failed, was listed, and could notify "failed". Give the reason a shared
constant (persisted value unchanged), add it to the no-verdict set, and treat
it as an internal marker so the Retry row no longer shows the raw string.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-27 22:23:49 -07:00