Commit Graph
39 Commits
Author SHA1 Message Date
Brennan BensonandMerge Sim 7faa9f7cd3 fix(native-chat): reasoning rows with a real open/finished state, and readable Claude thinking (#19221)
* fix(native-chat): render structured reasoning as collapsible messages

* fix(native-chat): align expanded reasoning with summary

* fix(native-chat): place reasoning chevron after summary

* De-emphasize reasoning headlines with observed timing labels

* Exclude later turn work from observed reasoning duration

* fix(native-chat): give reasoning rows a host-owned open/closed lifecycle

A reasoning row now says whether its block is still streaming (`state`)
and, when the host saw it end, when (`completedAt`). The row's own
start is its first-write time, so "Thought for N s" is measured on the
execution host instead of by whichever window happened to be watching.

Rows open with their first non-empty text and close on every path that
ends them: the block's final frame, a new message in the same stream,
every Claude turn end through one hook on the open turn, Codex
item/completed, turn and session settlement, active-item eviction, and
both host sweeps for a dead generation. Closing writes are lifecycle
writes so backpressure cannot leave a row open. Rows without the field
(older hosts, older journals) never read as live.

The reasoning row's headline reads the row's lifecycle and its own
turn's liveness: Thinking while open in a running turn, then Thought
for N s, Thought when no span was seen, Reasoning when the host kept no
lifecycle.

* feat(native-chat): ask Claude for readable thinking summaries

Under Orca's launch the Claude CLI streams thinking blocks with empty
text, so no reasoning row ever had anything to show. Pass only
`--thinking-display summarized`: it fills thinking blocks with the
API's summaries without turning thinking on, so a user who disabled
thinking keeps it off.

Orca runs the user's own binary, and a CLI older than the flag exits on
it before the session starts. The launch probes the binary it is about
to run, overlapping the rest of launch resolution, and passes the flag
only when that probe has already answered with 2.1.94 or newer. A slow
or failed probe never delays a launch and is never remembered; a
successful one is kept per binary until the binary changes.

* test(native-chat): type the superseding send in the reasoning lifecycle test

* fix(native-chat): stop reading an ended reasoning row as thinking

The spinner line infers "Thinking" from the newest root row being a
reasoning message. With summaries on, a closed reasoning row stays the
newest row while Claude streams a tool's input, so the line read
Thinking for the whole Write. A reasoning row now counts only while its
own state is running; a row from a host that keeps no state reads as
before.

* fix(native-chat): measure reasoning from the block's start, not its first text

A row opens with its first summary text, which trails the thinking
block's start by seconds, and the journal stamps a row with its queued
append time. So "Thought for N s" read 1 s for blocks that ran 4.65 s
and 12.95 s. The block registry now records the host time at each
block's content_block_start (else its first delta), and every write of
that block's row carries it as the row's observed time; the close keeps
the frame's own receipt time. Codex reasoning rows likewise carry their
item/started time.

* perf(native-chat): stop rewriting the reasoning row for every thinking token

Claude sends a thinking_tokens frame after every thinking delta. Every
frame the stream path did not consume forced the streamed text out
first, bypassing the checkpoint widening and the coalescing window, so
a single thinking block was rewritten and republished once per token:
212 full-row writes and 258 KB for one captured block. A frame that can
write no row (token tallies, stream deltas no registry carries, pings)
no longer forces that flush; every frame that can write still does.
The same block now takes 14 writes.

* fix(native-chat): time Codex reasoning by receipt and close evicted rows

Codex summary text lands 13-50 ms before item/completed, inside one
coalescing window, so a reasoning row's first write can be its
completion. Item boundaries are now stamped with their host receipt
time the way turn boundaries already were, so a retried or buffered
delivery keeps it, and the item translator falls back to the host
clock rather than Date.now(). The row carries its item/started time on
every first write, including the completion, so its span is
item/started to item/completed.

An evicted active item is now closed from the text streamed so far,
like both settle paths, and the eviction runs before the incoming item
is tracked: tracking first let the stream bound drop the evictee's text
before the eviction could close its row, stranding it running.

completedAt now has one meaning everywhere: the host time the message
was seen to end, or the end of the turn or stream that cut it off;
absent only when no end was seen live.

* refactor(native-chat): keep the Codex streaming body translation pure

The streaming translation preserves a reasoning body as-is again; the
stream writer, which is what knows the item has not completed, stamps
it running.

* test(native-chat): keep a re-collapsed reasoning row collapsed through a revision

* fix(native-chat): estimate a collapsed reasoning row as its trigger

A reasoning row renders collapsed, as one small button, but its height
was estimated from its full text: a 4,129-character summary reserved
about 950 px for a 24 px row, so long chats jumped as rows were
measured. It is now estimated at the trigger's height; opening the row
remeasures it.

* fix(native-chat): probe the CLI the launch will run, and learn from a refusal

The thinking-display gate probed `claude --version` with Orca's own cwd
and env, while the launch spawns with the workspace's cwd and the shell
env. Behind a version manager's shim those can pick different CLIs, so
the probe could approve a CLI the launch never ran, and an older CLI
exits on the unknown flag before the session starts. The probe also
only counted if it had already finished when resolution did, so a
first launch, or the first after a CLI update, usually went without
the flag.

The probe now runs with the launch's own resolved cwd and env, is keyed
by the binary and the workspace, and a launch waits up to 200 ms for it
(about 3x the probe's measured p95) before going without the flag. Only
answers are kept, so a slow or failed probe is asked again next launch.
A child that exits with commander's "unknown option '--thinking-display'"
marks that binary in that workspace so the next launch skips the flag;
that one start fails exactly as any CLI startup failure does today.

* test(native-chat): type the thinking-display probe mock with both of its parameters

* fix(native-chat): recheck the account switch after the probe, last as before

Moving the invocation ahead of the probe, the transcript check and the
permission mode put its account-switch recheck before those awaits, so a
switch that began during them launched unchecked. The invocation is the
last await again. The probe gets its own env from the same sources the
launch uses, the inherited env and the overlay with the CLI's runtime on
PATH, built by the same code, minus every credential: asking a CLI its
version needs none.

* fix(native-chat): wait up to 1.5 s for a cold probe, and remember every outcome

200 ms only covered a warm CLI; a cold disk, a node install or an
antivirus scan exceeds it, and that chat's child then ran its whole
life without summaries. A launch now waits up to 1.5 s, once per binary
per workspace, measured from when that binary's probe began, so a later
launch never waits again on a probe already past it. The probe gets its
own 10 s kill timeout, and every outcome, including no version printed,
a failure or a kill, is kept for the binary's life, so a probe that
hangs costs one launch rather than every one. A refusal seen while a
probe still runs wins over its late answer.

* fix(native-chat): leave Thinking to the activity line while reasoning runs

A reasoning row still being written drew a pulsing "Thinking…" header
right under the turn's activity line, which already says Thinking: two
live indicators for one fact. In a running turn an open reasoning row
now draws nothing and reserves no height; it appears when it closes, as
"Thought for N s". A row from a host that keeps no state, a closed row,
and a row left open by a turn that ended draw as before. The row's
Thinking headline is gone with its catalog key.

* fix(native-chat): end every unfinished Codex item through one rule

A reasoning completion with no text of its own left the row its stream
wrote running for good: the completion translated to nothing and the
item left the active set, so no settle could reach it. It now closes
from the text streamed so far.

Settlement, eviction and that completion now build an unfinished item's
row through one choice (the streamed text when there is any, else the
item as it started) and end it through one rule. An evicted file change
with streamed tool output no longer keeps that output as its patch; it
reads as interrupted, as a settled one does. A completion whose start
was never recorded claims no span, so it reads "Thought".

* fix(native-chat): catalogue Claude's stream keep-alive as benign

An uncatalogued `ping` stream frame classified as substantive, so the
fallback wrote a visible "claude · message:stream_event:ping" row, and
since such frames no longer force streamed text out first, a ping
inside a coalescing window landed above the open reasoning row. A ping
is now benign: it writes no row.

* fix(native-chat): bound finding the binary by the probe budget, and keep it LRU

Resolving the command's real path and its mtime was awaited before the
budgeted wait, so a slow filesystem could hold a launch indefinitely;
it now counts against the same budget, and running out caches nothing.
The cache is least recently used rather than first written, and holds
32 binary-and-workspace entries rather than 16.

* feat(mobile): collapse reasoning rows the way desktop does

With summaries on, every Claude turn now carries reasoning text, and the
phone drew all of it inline, dimmed, between the prompt and the answer.
Mobile now draws a reasoning row as desktop does: collapsed to "Thought
for N s", "Thought" or "Reasoning", its text mounted only once opened,
and nothing at all while the row is still being written in the live
turn or has no text. The headline and the visibility rule live in one
shared module both clients read, so they cannot drift.

* fix(native-chat): let a failed CLI probe heal instead of latching

A probe killed at its timeout, failing to spawn under a loaded boot, or
printing no version was cached as "no flag" for the binary's life, so
that workspace never got summaries again in that run. Only a version
(either side of the floor) or the CLI's own refusal is kept for good
now; a probe that gave no version is kept for 10 minutes, so a hung CLI
still costs one wait per stretch and a boot-time failure heals.

* docs(native-chat): say exactly what the version probe's env leaves out

* fix(native-chat): record a Codex item's start whatever its first frame carried

The start was recorded only for an item/started that wrote no row, so a
reasoning item that started with text lost it and its completion
claimed no span. Every tracked item/started now records its receipt
time, and the started write carries it too.

* fix(mobile): label the reasoning toggle and give it a full touch target

The toggle now tells a screen reader what it is, "Reasoning: Thought
for 3s", as desktop's prefix does, and reaches a 44 pt target. The
shared English copy stays private to the module that formats it.

* refactor(native-chat): build reasoning rows from one provider-neutral helper

Claude, Codex and the terminal sweeps each built the reasoning row body and
its running/ended stamp themselves. They now share journal-reasoning-row:
blank text journals no row, text is bounded the same way, and an end carries
completedAt only when the host saw it.

* refactor(codex): move the active journal item type into the contracts file

codex-unfinished-item-body imported the type from the settlement module,
which imports values from it.

* test(claude): read the launch PATH the way Windows spells it

* test(claude): compare the probe's PATH to the launch's without Orca's CLI dir

When the CLI's directory also holds node (Linux CI's /usr/local/bin), the
runtime pairing puts that directory first, ahead of the Orca CLI directory the
launch adds, so the launch PATH no longer ends with the probe's. Both still
resolve the same claude and shims. The test now checks that exactly, for a CLI
with and without a sibling node.

* feat(native-chat): lead the reasoning row with a brain glyph in the tool-row column

* fix(native-chat): route the reasoning glyph through the shared icon names, keep its chevron findable, and match it on mobile

* fix(mobile): keep the long-press actions sheet on reasoning rows for Android

* fix(native-chat): forward every exit argument through the thinking-display connection wrapper

* refactor(codex): keep the receipt-timed notification methods with the event they stamp

* feat(native-chat): read an open reasoning block through the one live "Thinking" line

While the agent's open reasoning block has text, the turn's live activity line is its
disclosure: collapsed by default, expandable to the live text (capped and scrollable), and
the block's row draws nothing meanwhile. When the block ends, its row appears in place,
open if the reader opened it live, because the line and the row read one disclosure key.
Which block the line discloses is derived from the line's own render condition, so a row
is never hidden while nothing on screen shows it; any other open block (a subagent's, or
one a prompt pushed off the line) draws as "Reasoning". Desktop and mobile alike; no host
or wire change.

* fix(native-chat): a slot kept for its turn bar or diff rollup no longer draws its message

The transcript row drew the message of every message slot, while the slot builder pushes a slot
for a row it declined to draw whenever that row also carries its turn's bar or diff rollup. So the
open reasoning block the live line discloses still drew as a "Reasoning" row when it was a
provider-opened turn's first row or the last row of a turn that changed files, and one click
opened both. The builder's decision now travels on the slot (`drawsMessage`) and the row draws
only the bar and rollup when it is false; the row-level `folded` guard it made redundant is gone.

* refactor(native-chat): draw the live line from one shared value, with one live region

Desktop and mobile now render the live activity line from one pure function,
`nativeChatLiveLine`: whether it draws, what it says, and the open reasoning block it
discloses with the text it has so far. The open block's row is hidden from that same value,
so desktop no longer restates the line's render condition beside it, and the lines no longer
re-derive the text.

The line keeps one element, and so one live region, through every state; only its trigger
and body come and go, so a screen reader hears "Thinking" and the label after it. On mobile
the live text gets the finished row's Android long press (copy or select through the message
actions sheet), the line's touch target is the row's 44 pt, and its label and body sit in
the finished row's column so nothing moves when the row takes over.

* fix(mobile): keep the live text's actions sheet on the block it was opened for

On Android the sheet opened by a long press on the live reasoning was a flag gated on a live
block: it vanished when the block ended, mid Select text, and the stale flag reopened it
unprompted on the next block. The sheet now holds the message it was opened for.

* test(native-chat): the reasoning body owns its tone, live and once landed

Rendered QA on a pre-merge build showed the live line's open reasoning in full foreground and the
landed row's in muted text, so it dimmed as the block landed: the body set no colour of its own and
inherited one from wherever it was mounted. Since the main merge (017ad743fa) the body carries the
chat's faint tone itself; this pins that, and that the line and the row draw the same body.

---------

Co-authored-by: Merge Sim <sim@local>
2026-10-06 07:19:24 -07:00
Brennan Benson 9905765e3b feat(native-chat): open structured chat's wire and stored records to registered agents, behind a negotiated capability (#25159)
* refactor(native-chat): keep the provider resume handle opaque to shared code

Shared structured-chat code parsed each provider's resume handle: Claude's
session id and branch leaf, Codex's thread id, through a 'claude' | 'codex'
union every new agent had to widen. The in-memory handle is now
{ transport, agent, nativeId, providerData? }: shared readers use nativeId,
lease and handle-chain checks compare transport and agent, and only the
Claude adapter reads its leaf (providerData).

Stored and wire forms are unchanged for Claude and Codex. One encoding
module writes their typed shapes and decodes both those and the neutral
shape a new transport uses, which an older build refuses as unreadable
rather than reading as Codex. Key and root strings, which fork seeds,
superseded creations and resume offers persist, stay byte-identical.

The journal's own handle type becomes the journal-row and attach-wire
encoding of the same handle, and the journal identity carries the
neutral handle (null before the provider proves one).

No user-visible change.

* fix(native-chat): derive journal-row provider handles from the journal identity

The journal row converter now takes the identity every caller already holds,
so a row's handle has one obvious constructor. Tests that wrote the in-memory
handle straight into journal rows now build it through that converter, and the
processless Claude fixture names a not-yet-proved handle as null.

* fix(native-chat): refuse a stored provider handle written in both forms

A typed Claude or Codex handle that also carries the neutral form's
transport, agent, native id or provider data named two identities; it
was read as Claude or Codex and the next write dropped the other one.
Such a row now stays unreadable and is set aside untouched.

* refactor(native-chat): route structured agents through registered definitions

The structured-chat shared layer named Claude and Codex in a closed router
type, in option reads, in frame classification, and in renderer label and
validation branches, so adding an agent meant editing shared types. Each
agent now has a definition (capabilities, resting option rules) owned by its
own module; the router is a registry of definition + adapter pairs built at
runtime setup, and shared code reads the definition instead of the name.

No behavior change for Claude or Codex; no wire or stored shape change.

* refactor(native-chat): name the structured agent list once in host types

The host types, conversation commands, provider-session ownership and the
runtime's create path spelled the structured agent list inline; they now use
the one shared type the record and wire already validate against.

* fix(native-chat): narrow the record before reading its agent's option rules

* refactor(native-chat): make the router's registrations the only agent definition lookup

The at-rest option reads looked definitions up in a second, separately populated map, so a
registered definition could route and declare capabilities while the static map decided which
options a chat at rest accepted. The router now answers `definition(agent)` from its
registrations, the at-rest readers take it through the host's adapter, and the built-in
definitions are listed only where the runtime composes the router.

* feat(native-chat): open the structured-chat wire and stored records to registered agents

A host's structured agents are the ones its runtime registered. Records, RPC
params, persisted tabs and the model catalog accept any registered agent instead
of naming Claude and Codex; each agent's definition declares the transport its
handles live in and the variable its account home pins. A new runtime
capability, agent-session.structured.registered-agents.v1, advertises that a
host accepts and lists its agents (agentSession.agents, with each agent's
capability record), and the host withholds any other agent's tabs and restart
offers from clients that do not advertise it.

* test(native-chat): cover registered agents on the wire, in storage and across versions

* refactor(native-chat): let the record store decide which agents' tabs exist

* test(native-chat): declare the pilot test agent's storage

* test(native-chat): read the old build's saved tabs through a parsed shape

* fix(native-chat): act on restart offers only for agents the calling client can show

A paired client too old to show an agent's chat was listed only the offers it could show, but
dismissing or continuing all reached every offer on the host, and a named continuation answered
with the host's whole remaining inventory. The client's audience now goes to the host with every
restart operation: only offers it sees are reserved, dismissed or returned. Without an audience
(this host's own process, or a client that shows every agent) nothing changes.

* test(native-chat): read an agent-registering baseline's storage on its own terms

The registered-agents downgrade test assumed its baseline release predates registered agents: it
expected the saved-tab parser to erase an unknown agent and called the record reader without the
agents list. Once a release with this change becomes the baseline, both break. The expectations now
follow what the baseline host advertises, and an agent-registering baseline is handed its own
Claude and Codex storage.

* refactor(native-chat): derive record-store admission from the runtime's agent registrations

Which agents a stored record may name and which agents the router drives came from two lists in
the runtime, so a newly registered agent could be routed while its records were set aside. One
list of registrations now holds each agent's definition and the factory for its adapter: the
store's admitted agents are derived from it before the store opens, and the adapters are built
from it once it has.

* fix(native-chat): hand the exit drain a promise for every registered agent

* fix(native-chat): let the adoption conflict check read any agent's ownership

Ownership rows name any registered agent since the stored records opened to them; the adoption
check compares by agent, so it takes the same open id. Only Claude and Codex still adopt.

* test(native-chat): use opaque handle in queued rejection fixture

* test(native-chat): share one Codex journal identity in the integration suite

Main grew the suite to the 800-line limit; the opaque-handle import pushed it
over. The two tests built the same identity inline.

* refactor(agent-session): name the handle's adapter state resumeCursor

Rename the neutral provider handle's providerData to resumeCursor before any
row persists the neutral form: it is an adapter-owned resume position (Claude's
transcript leaf), never identity. Claude/Codex stored and wire bytes are
unchanged; their typed shapes never carried the field.

State the stored-form contract (a handle's field set is closed; later per-link
data goes on the chain link, which every build preserves) and pin it with a
record round-trip test. Document that transport records the id space the
native id was minted in, which can differ from the agent's current transport.

* refactor(agent-session): one required agent registry; declarations admit what they claim

A StructuredAgentRegistry, built once from the {definition, adapter}
registrations, is now a required host dependency and the router routes with
it. The adapter interface loses its optional router-only capabilities?() and
definition?(); every reader (options at rest, thread goal, rewind, the
/compact handover) asks the registry. A live session still narrows rewind
through the adapter, and the host combines declared and narrowed in one
helper. The registry refuses a registration that declares compact, a thread
goal or rewind without the adapter method behind it.

/compact is admitted by the declared capability, so an agent that declares
compact:false gets the commandRefused fact instead of a thrown error.

Also: the cut-turn notice names the agent from the catalog, create-support
builds the account home through agentSessionAccountHome, and the turn
status text older clients read names the session's agent instead of
defaulting to Codex (Claude/Codex text unchanged).

* refactor(agent-session): the router applies the declared rewind itself

The router holds the registry, so its rewindSupport answers the owner's
declared rewind narrowed by the adapter, in one place no reader can bypass.
The host-side combining helper is gone; readers ask the adapter they hold.

* test(agent-session): register the agents the merged-in tests now need

The adoption replay test builds its host over this build's registered agents
instead of an adapter definition method, and the stored-form test passes the
stored agents the record guard now requires.

* chore(agent-session): the registry is the one lookup; drop the test-only empty capability record

* fix(agent-session): keep "dismiss all" restart offers dismissed for the desktop

The desktop renderer calls the restart methods as a paired runtime client
without the registered-agents capability, so it was handed a Claude/Codex
audience and its "dismiss all" took the scoped path: no dismissal fence and
no clearing of unwritten teardown witnesses, so a late teardown write could
bring a dismissed offer back.

The audience is now derived against this host's registry: a client that can
show every agent the host registered gets none, and dismisses exactly as
before (one fence for every offer). A client that truly cannot show some
agent still gets a final dismissal for what it sees: each cleared session is
fenced on its own (bounded, superseded by a dismiss-all fence) and only the
witnesses of agents it sees are dropped.

One predicate now answers which agents a client renders, for both tab
projection and restart offers; the Claude capability rule is folded in.

* fix(agent-session): a changed agent definition never hides that agent's chats

A record was readable only if every handle used the transport its agent's
definition declares today, and its account variable was declared by some
registered agent. A later build that drives the same agent over another
protocol, or under a renamed variable, would have set every existing chat of
that agent aside: no tab, no history, no error.

Readability now asks only that the record agree with itself: its agent is
registered, every handle shares one namespace owned by that agent, and the
account variable is a well-formed name. Whether this build can drive it (the
chain's transport is the one the agent speaks, and the account variable is
that agent's own, since it becomes the child's environment) is decided in
performAttach, the one admission every start of every agent passes through,
and refused there as hostUnsupported. A record pinning another agent's
variable is refused the same way rather than passing because some agent
declares it.

* refactor(agent-session): each agent's registration says where it runs and which account it pins

createSupport, create and the model catalog's account read still chose a
location rule and an account resolver by name (Claude, Codex, else none),
so a new agent would have needed a third branch beside the registration
list that already decides storage, routing and publication.

Each runtime registration now carries supportsLocation(location) and
resolveAccountHomePath({ launchEnv, location, purpose, workspacePath }),
and Claude's and Codex's rules move into their entries unchanged: a read
still syncs no home and starts no bridge, a Codex launch still trusts the
workspace first, and WSL locations resolve as before. The list exists
before the host is built, so none of these reads installs the host, and an
agent the list does not hold is answered no without opening the journal.

* test(agent-session): drop a duplicate registry key; keep fences out of the capsule state

A spread input already carries the registry, so the explicit key was
overwritten (TS2783). The session-fence helper now returns only the fence,
not the capsule state it was given.

* fix(agent-session): a scoped dismiss-all persists no per-session fence

The per-session dismissal fence added a forever-persisted capsule field
with an arbitrary cap that served no caller: only the local desktop calls
the restart methods, and on a host whose agents it all shows it gets no
audience and takes the unscoped, fenced dismissal. A scoped dismiss-all
(a caller that cannot show some registered agent) keeps that agent's
offers, clears only its audience's unwritten witnesses, and writes no
fence; a late write from this process is serialized behind it. The field
never left this unreleased branch.

* fix(agent-session): refuse an attach whose agent is not the session's own

The start check judges the record's agent while the router starts the adapter of params.agent, and the attach wire let the two differ. Admission now refuses agent !== provider as requestMalformed, so the agent checked is the agent started. Every host-built attach already sets them equal.

* fix(agent-session): offer to start a chat only when the start would accept it

The restart offer, a failure's retryable flag and the pre-send check asked only whether the adapter runs the chat's location, while the start also refuses a record this build cannot drive; such a chat was offered Resume and its Retry failed forever. One host predicate, hostCanStartRecord (location support and agentDrivesSession), now answers all four; adapterSupportsRecord stays the reader gate, and an undrivable chat's offer is kept, not retired.

* fix(agent-session): the model-catalog probe uses a record's account only if this build drives it

A record-scoped catalog read started the agent's lister under the record's account path whatever variable the record pinned. It now uses that path only when agentDrivesSession holds, and otherwise resolves the account as for a read with no record.

* refactor(agent-session): the record store admits agent ids; comments say where transport is checked

The store only ever asked whether an agent is registered, so its admission list is now the registered ids; the definition is the one home of an agent's transport and account variable. Comments that said the record store checks transport, and one that cited a nonexistent function, now point at agentDrivesSession and the attach admission. The isPersistedAgentSessionRecord(value, agents) call shape and the test fixture the cross-version probe imports are kept.

* docs(agent-session): the record store admits the registered agents' ids

* test(native-chat): dismiss-all through a remote client's audience keeps a newer Orca's offer

Uses an audience production sends (one that cannot show every agent), per review.

* chore(native-chat): keep the record store and recovery capsule under max-lines after the main merge

* test(native-chat): build the Stop-opened-turn test's identity and turn context the current way

The test (#25056) landed before the opaque provider handle (#24991), so main still built
the old {kind, threadId} handle; the turn context also needs this branch's agent registry.

* test(ratchet): require src/main/provider-process now that it has landed

* test(native-chat): keep main's opaque-handle import in the Stop-opened-turn test

Main's #25706 and this branch both added the import at different lines; the merge kept both.

* Keep saved chats readable without provider registration

* Keep stored-record compatibility checks independent of registration

* Keep saved providers in restart client audiences

* test(native-chat): type reveal fixtures without assertions

* test(wire): expose known agents in structured host fixture

* Supply startability dependency in the new Codex catalog fixture
2026-10-06 00:41:14 -07:00
Brennan Benson 3a03441580 refactor(native-chat): structured agents declare their capabilities instead of shared code naming Claude and Codex (#25076)
* refactor(native-chat): keep the provider resume handle opaque to shared code

Shared structured-chat code parsed each provider's resume handle: Claude's
session id and branch leaf, Codex's thread id, through a 'claude' | 'codex'
union every new agent had to widen. The in-memory handle is now
{ transport, agent, nativeId, providerData? }: shared readers use nativeId,
lease and handle-chain checks compare transport and agent, and only the
Claude adapter reads its leaf (providerData).

Stored and wire forms are unchanged for Claude and Codex. One encoding
module writes their typed shapes and decodes both those and the neutral
shape a new transport uses, which an older build refuses as unreadable
rather than reading as Codex. Key and root strings, which fork seeds,
superseded creations and resume offers persist, stay byte-identical.

The journal's own handle type becomes the journal-row and attach-wire
encoding of the same handle, and the journal identity carries the
neutral handle (null before the provider proves one).

No user-visible change.

* fix(native-chat): derive journal-row provider handles from the journal identity

The journal row converter now takes the identity every caller already holds,
so a row's handle has one obvious constructor. Tests that wrote the in-memory
handle straight into journal rows now build it through that converter, and the
processless Claude fixture names a not-yet-proved handle as null.

* fix(native-chat): refuse a stored provider handle written in both forms

A typed Claude or Codex handle that also carries the neutral form's
transport, agent, native id or provider data named two identities; it
was read as Claude or Codex and the next write dropped the other one.
Such a row now stays unreadable and is set aside untouched.

* refactor(native-chat): route structured agents through registered definitions

The structured-chat shared layer named Claude and Codex in a closed router
type, in option reads, in frame classification, and in renderer label and
validation branches, so adding an agent meant editing shared types. Each
agent now has a definition (capabilities, resting option rules) owned by its
own module; the router is a registry of definition + adapter pairs built at
runtime setup, and shared code reads the definition instead of the name.

No behavior change for Claude or Codex; no wire or stored shape change.

* refactor(native-chat): name the structured agent list once in host types

The host types, conversation commands, provider-session ownership and the
runtime's create path spelled the structured agent list inline; they now use
the one shared type the record and wire already validate against.

* fix(native-chat): narrow the record before reading its agent's option rules

* refactor(native-chat): make the router's registrations the only agent definition lookup

The at-rest option reads looked definitions up in a second, separately populated map, so a
registered definition could route and declare capabilities while the static map decided which
options a chat at rest accepted. The router now answers `definition(agent)` from its
registrations, the at-rest readers take it through the host's adapter, and the built-in
definitions are listed only where the runtime composes the router.

* test(native-chat): use opaque handle in queued rejection fixture

* test(native-chat): share one Codex journal identity in the integration suite

Main grew the suite to the 800-line limit; the opaque-handle import pushed it
over. The two tests built the same identity inline.

* refactor(agent-session): name the handle's adapter state resumeCursor

Rename the neutral provider handle's providerData to resumeCursor before any
row persists the neutral form: it is an adapter-owned resume position (Claude's
transcript leaf), never identity. Claude/Codex stored and wire bytes are
unchanged; their typed shapes never carried the field.

State the stored-form contract (a handle's field set is closed; later per-link
data goes on the chain link, which every build preserves) and pin it with a
record round-trip test. Document that transport records the id space the
native id was minted in, which can differ from the agent's current transport.

* refactor(agent-session): one required agent registry; declarations admit what they claim

A StructuredAgentRegistry, built once from the {definition, adapter}
registrations, is now a required host dependency and the router routes with
it. The adapter interface loses its optional router-only capabilities?() and
definition?(); every reader (options at rest, thread goal, rewind, the
/compact handover) asks the registry. A live session still narrows rewind
through the adapter, and the host combines declared and narrowed in one
helper. The registry refuses a registration that declares compact, a thread
goal or rewind without the adapter method behind it.

/compact is admitted by the declared capability, so an agent that declares
compact:false gets the commandRefused fact instead of a thrown error.

Also: the cut-turn notice names the agent from the catalog, create-support
builds the account home through agentSessionAccountHome, and the turn
status text older clients read names the session's agent instead of
defaulting to Codex (Claude/Codex text unchanged).

* refactor(agent-session): the router applies the declared rewind itself

The router holds the registry, so its rewindSupport answers the owner's
declared rewind narrowed by the adapter, in one place no reader can bypass.
The host-side combining helper is gone; readers ask the adapter they hold.

* chore(agent-session): the registry is the one lookup; drop the test-only empty capability record

* test(native-chat): build the Stop-opened-turn test's identity and turn context the current way

The test (#25056) landed before the opaque provider handle (#24991), so main still built
the old {kind, threadId} handle; the turn context also needs this branch's agent registry.

* test(ratchet): require src/main/provider-process now that it has landed

* test(native-chat): keep main's opaque-handle import in the Stop-opened-turn test

Main's #25706 and this branch both added the import at different lines; the merge kept both.
2026-10-05 21:09:26 -07:00
Brennan Benson 57fedeed79 fix(native-chat): the agent's exit ends its record, and an unconfirmed stop is joined instead of held (#24862)
* fix(native-chat): a child's root exit is reported even during its close, and bookkeeping after it never reads as unproven

- Both connections report the root process's exit once, with `expected` set when a close had
  begun. A close that came back unproven and whose root exits later is finished by the adapter,
  and its end reaches the host like any other.
- A Claude close whose resume-point write fails after the exit was proven, and a Codex close whose
  terminal row is refused, now end the session and report the failure, instead of keeping a dead
  child indexed as if its exit were unproven.
- A Codex close whose forced tree kill can't prove the descendants gone but saw the root exit
  reports the descendants and counts the root exit.
- Every child exit with an identity, expected or not, is forwarded to the host.

* fix(native-chat): the exit ends the child's record; an unfinished stop is the child's own close, which everyone joins

- The host keeps no stored "stop still owed" record any more. A stop begins the child's close
  (`child.close`), which lives on the child and ends with it. A second Stop, the idle reaper,
  quit, a send and an option/answer/goal/rewind all join that close instead of retrying a
  separate obligation.
- A caller waits on the close only as long as the step deadline; the close itself is never
  abandoned. A proof that lands after every caller stopped waiting reaches the host as the
  adapter's report of that exit, which ends the record through the same handler.
- Once the exit is proven, draining, settling, the lease release and the adapter's
  acknowledgement are each attempted and reported on failure; none keeps the child on record.
  A start, and the handle's close, write a release that failed from this host's proof of that
  exit, so a failed write never refuses a send.
- A start that meets a close still unverifiable is refused with `previousExitUnverifiable`, so the
  queued message is rejected with a send-again reason; nothing is held and nothing starts beside
  the old process.
- The idle sweep goes back to idle reaping only.
- Removes #24333's retry entry points, the wait row and its hold rule, the ask/failure cursors on
  the stored record, and the stop's own wake.

Tests replace the #24333 unproven-stop test: a send joining an unproven close and an in-flight
one, a late proof past the caller's bound, a root exiting after its close gave up, a proven exit
whose resume-point write and lease release both failed, an unverifiable close rejecting the send
and refusing an option change, a surviving descendant, quit and the idle reaper; and Codex's
unverifiable, late-exit and joined-close cases.

* fix(native-chat): a message refused because the old process's exit is unverifiable says so, and to send again

The start failure for a refusal with reason `previousExitUnverifiable` reads "Orca couldn't
confirm Claude's previous process ended. Send your message to try again." instead of "Claude
couldn't restart." The status-row kind and the refusal reason stay in the shared lists for rows
and hosts that still carry them; the catalogs keep one sentence for both.

* fix(native-chat): a close's verdict is the root's exit alone, and what follows it is logged

- A Claude close resolves as soon as the root's exit is proven: the session ends and its `ended`
  report goes out then. Saving the resume point runs afterwards and a failure is logged, so a slow
  or hung write never reads as an unproven exit or keeps a dead child on record.
- A root that exits after its close came back unproven finishes that close through the same path
  as any close, so the session's child work is published as ended (background tasks and subagents
  no longer stay shown running for a dead agent), and a failure there is logged.
- Codex logs a refused final row, and reports a root exit whose forced tree kill could not prove
  the rest of the tree gone the way Claude does, so the host logs it and blocks nothing.
- Both adapters take the host's logger for this bookkeeping.

* fix(native-chat): one handler ends every child's exit, and a join waits on the adapter's own close

- One exit handler (`structured-agent-session-child-exit`) ends a child's record for an exit
  expected or not. `expected` only changes what the chat is told: the stop's cause, its end at the
  stop's ask, the settlement id, and no crash outcome row. The lease release keeps the exit's
  evidence; the handoff guard, lifecycle barrier, sink release and adapter acknowledgement apply to
  both. A Claude journal-sink failure ends in the same step as its stop, as Orca's own fault.
- Joining a close is asking the adapter, whose close is memoized while it runs and bounded by its
  own kill escalation; the host keeps no attempt of its own and no 10 s caller bound. An ask after
  a close came back unproven runs the stop again.
- A close's end is stamped where its stop was asked for (a repeated ask moves it), so the closed
  chat and failed start checks order a message accepted meanwhile after it.
- A start refused because the old exit is unverifiable rejects what was queued in the same step.
- The end of a close the host asked for no longer waits on the cross-session recovery chain.
- The kill no longer waits for the stop event's write; the journal writes rows in order.

* fix(native-chat): an exit's lease release lands whatever the length of its reason

A crash's reason can carry kilobytes of the provider's stderr, and a lease whose death detail is
over 512 characters fails the store's own check. The exit handler cut it, but the release a start
or the chat handle's close re-derives did not, so after a crash whose own release failed every
message was refused as not resumable until restart. The record's builder now cuts the detail to
the record's bound, so no writer can hand it one too long.

* fix(claude): a proven close waits at most 2 s for the output it already wrote

Once the root's exit is proven, the close still waited for the SDK's output reader to end. Something
outside the process tree that holds the output open would keep that close, and every send, Stop
and quit joining it, waiting with no bound. The wait is now bounded; past it the close resolves as
proven and the open output is logged.

* fix(codex): an exit reported inside Orca's close keeps the reason Orca closed it for

The connection reports the app-server's exit inside the close that ends it, so that report ended
every Codex close and replaced the close's own reason (for example, a provider frame that could not
be recorded) with the connection's stderr text in the ended record and the lease's exit evidence.
The session now records Orca's close with its reason, and the exit it ends keeps that reason. The
test connection reports its exit inside close the way the real one does.

* fix(native-chat): quit stops delivery before it drains exit recovery

Every exit now wakes delivery, and teardown drained exit recovery before it stopped delivery, so an
exit settled in that window could start a fresh agent that teardown then killed. Teardown stops
delivery first; queued messages wait for the next launch.

* docs(native-chat): the unverifiable-exit refusal no longer names a caller's wait

The caller's bounded wait was removed; the comment describes the close as it is now.

* fix(native-chat): a stop whose kill did not take is logged, and the next ask kills again

When a close's kill leaves the agent's root running, the host now logs it. Tests pin what a later
ask does: each connection runs its whole stop again (Codex sends SIGKILL a second time), refuses
input meanwhile, and proves the exit once the kill takes.

* fix(native-chat): a start refused over the old process says Orca couldn't stop it

The host reaches an unverifiable verdict only after its own kill left the agent's root running, on
the machine that runs the agent, so the sentence now says that: "Orca couldn't stop {agent}'s
previous process." The refusal reason, failure kind and wire shapes are unchanged. The host test
also checks the failed kill is logged.

* docs(native-chat): an unverifiable close verdict is a root that survived the kill

The host's close runs where the agent runs, so lost contact never yields this verdict; the comment no longer says it does.

* fix(native-chat): a kill that did not take is reported once, by whoever met it

The log added at the close fired beside a Stop's own failure report for the same event. A stop
still reports it through its failure; a send or option change refused over it now logs it at the
refusal, the only place it is otherwise invisible.

* test(native-chat): a second Stop joins a close the first could not prove and retries its kill

* fix(native-chat): say a start refused beside an unstopped process plainly

The rejection now reads "Couldn't stop {{agent}} from before. Send your message again to try once more."
This kind has its own send-again step; every other failure keeps "Send your message to try again."
2026-10-05 12:14:47 -07:00
Brennan Benson c6cfcc034e refactor(native-chat): structured chat failures always reach the diagnostics log (#24312)
* refactor(native-chat): give the structured chat host one required logger

The structured chat runtime took an optional onError callback that the
desktop never passed, so a late dispatch settlement, an unanswered-dispatch
release, a journal event-sink write and a provider lifecycle delivery that
failed were dropped with no trace. Other host failures went to scattered
console.warn calls, which reach nothing in a packaged desktop build.

The runtime and host now take one required logger (warn/error with a scope
and fields). The production logger writes each entry as a failed span to
<userData>/logs/main.trace.ndjson, which the diagnostic bundle collects, and
to the console (stderr under a supervised headless host). The runtime and the
host wrap it so a logger that throws never fails what it reports, and the
install refuses without one. Sites that deliberately kept a recovery-capsule
error out of the log still log no error object.

* refactor(native-chat): hand the chat host's collaborators the logger, and give orcad its trace file

The delivery loop, idle sweep, queued-message drain, lease renewer, event
sink, conversation map and provider start/exit settlement each took an
internal error callback that the host mapped onto the logger. They now take
the logger itself and log under their own scope. The event sink keeps one
onFailed hook, which decides whether to stop the provider, not whether to
report. The dead-generation settlement returns its failure so each caller
logs it under its own scope.

orcad now installs the desktop's local trace sink under its own data root, so
a headless host's chat failures reach <data-root>/logs/main.trace.ndjson as
well as stderr.

Also passes the logger in the test fixtures the first commit missed, which
tc:node caught.

* fix(native-chat): keep repeated chat failures from flooding the trace file, and record their causes

- The production structured-chat logger writes a repeated failure (same level, scope, session,
  message and error text) once per 5 minutes, carrying how many repeats it swallowed; the
  tracked set is capped at 256.
- Trace entries now carry the error's code (and SQLite errcode) and up to three causes by name and
  message.
- A chat read whose conversation will not open is logged through the host's logger
  (open-for-read), and so are the runtime's chat-tab bookkeeping failures that already hold the
  host.
- orcad writes its own orcad.trace.ndjson, closes it after every quit handler, and flushes it on
  process exit; a trace file that cannot be opened leaves tracing off instead of stopping the app
  or orcad.
- Tests: the desktop wiring test proves the logger reaches the trace sink, and the privacy tests
  read every level the logger received.

* fix(native-chat): log a created chat's tab-publication and launch-prompt failures through the host's logger

* fix(native-chat): key a repeated chat failure on everything its entry writes

The repeat suppression keyed on the message and the error's text, so two refusals with the same
code but different causes, a plain error and a refusal of one code, or two object-valued errors
shared a key and the second was swallowed for five minutes. The key is now the entry's whole
written content (fields, code, errcode, refusal reason, cause chain, a stable rendering of a
non-error value) plus the error's name and message; a refusal's reason is also written.

* test(native-chat): pin that an error's name keeps two repeated failures apart

* test(native-chat): build the refusal in the repeat-key test as the wire does

* fix(native-chat): read an error's code and a refusal's reason by narrowing, not Reflect.get
2026-10-01 11:21:50 -07:00
Brennan Benson 0b79720c2e feat(native-chat): the chat strip and the sidebar read the host's child records (#22614)
* feat(native-chat): publish the host's child records to the status summary and the chat strip

The status summary and the background-task channel now read a session's child
records from the host's canonical store, through the status sink its row
landed in, and derive the legacy task and subagent shapes from the same views.
The parent row folds its child-work liveness from those records at ingest,
not from the summary's task list. The adapters no longer push their task DTO
to clients: the onBackgroundTasksChanged path is gone, and a child-work ingest
is what republishes both the summary and the strip. Finished children stay
listed until the session's own next turn starts. A reader that predates child
views never receives a roster whose rows are all settled.

* feat(native-chat): the chat strip reads the host's child records with its parent's verdict

The strip's roster now renders from the child views its channel carries, and
passes the verdict the session's own status row gives its children, built the
way the sidebar builds it (the row's freshness and the status feed's
observation). So one child reads the same in the strip and the sidebar, live,
after the transport drops, and once the row goes stale. A roster of finished
children stays shown until the next turn but no longer animates the monitoring
indicator or blocks conversation commands.

Tests: an end-to-end run on a host with no renderer (a real hook server as the
status sink) shows the summary and the strip channel carrying the same records
at every step, the parent row folded from them, retention, and an older
client's task list holding live work only; a wired renderer test shows both
surfaces agree when live, lost and stale.

* test(native-chat): a newer host's view degrades, an older reader keeps its live roster, a finished roster holds nothing open

- The view decoder ignores unknown keys, degrades unknown kinds, states,
  outcomes and memberships, and drops only rows it cannot identify.
- At the RPC boundary a reader that predates child views gets no strip for a
  roster of finished children and never the views themselves; a stop-only
  reader keeps rows whose host offers no targeted stop.
- The strip shows finished children without reading them as live work.
- The row keeps its child list's identity when a summary repeats it.

* refactor(native-chat): the summary's task list is the legacy projection's live rows, unfiltered

* test(native-chat): type the switch tests' mocks instead of asserting them

* test: remote clients advertise reading child views

* docs(agent-status): the structured row folds the store's child records

* fix(agent-status): keep the view reader's header from reading as a value import to the renderer boundary

The renderer node-builtin boundary test scans raw text, so a header comment
that said "imports" ahead of the import block turned the type-only import of
agent-status-child-work into a value edge that reaches node:crypto.

* refactor(native-chat): the status summary's broadcast equality gets its own module

The status feed crossed the file-size limit once the summary gained the main agent's turn
outcome beside the child views. Which summary changes reach every session list now lives in
structured-agent-session-status-summary-equality.ts.

* fix(native-chat): command admission reads the strip's child records

A conversation command was refused on the provider tracker's own roster
while the strip read the host's child records, so a drift between the two
rule sets could refuse /clear with a stop instruction the strip had no
button for. Admission now reads the same records through the same read as
the strip, uses the strip's liveness fold, and asks for a stop only when
the strip renders one. The adapter contract no longer exposes the tracker
roster, so no host decision can read it.

Also records when the legacy child shapes die, every earlier death of a
settled child, and the display-precision invariant behind the summary's
clock tolerance.

* refactor(native-chat): command admission takes only what it reads of a turn

* fix(native-chat): the session list drops a session's children when the store does

A session's end no longer removes its child records: a child still running
settles with an outcome nobody reported, and a finished one stays listed.
Records now leave only at the session's own next turn, at the cap on settled
records, or when the host lets go of the session and its row leaves the store.

The summary kept after the host lets go used to strip its children on close,
a rule of its own. It now re-reads them from the store when the row leaves,
through the same read every live summary uses, so the session list and the
chat strip list the same children at each step, including a forget with no
close. Closing only revokes ownership, as before the child records existed.

* test(native-chat): write the Codex frame script's parent row out step by step

Once every surface reads the child records, the provider tracker's roster is
no oracle: it and the records read the same child executions, so agreeing
with it cannot catch a defect in either. Each frame now states the child
liveness and the parent row it must fold to.

* fix(native-chat): the idle sweep and the restart snapshot read the host's child records

The idle sweep (keep an agent running while its subagents or commands run) and
the restart-resume snapshot (what a chat was doing when Orca stopped it) both
read the provider tracker's roster through the adapter interface, which no
longer carries it. Both now take the host's one child-record read, the same one
the status summary, the chat strip and command admission use.

The snapshot's working test also folded that roster through the shared fold's
old `backgroundTasks` input, which the fold no longer reads, so a settled lead
whose subagent was still running would have been offered nothing. It now hands
the fold the records.

* test(native-chat): the child-record tests follow the merged command lifecycle

A command is a live child record from its start and is removed, not settled,
when it stops, whatever Codex tagged it. The end-to-end switch now shows the
child's `npm test` as a live row beside its dev server, and both are gone once
they exit; only the finished subagent stays listed until the next turn. Letting
go of the session is its tab closing, since a closed conversation whose tab
remains keeps its row.

Command admission's finished row is a subagent, the one kind that settles, and
the failed-verdict row test admits its live subagent as a host record, the only
thing the row folds.

* refactor(native-chat): the status feed's journal projection cache gets its own module

The status feed crossed the file-size limit once the child records joined the
agent-start signal and the completion feed's status read. The per-journal
projection, cached per commit, now lives in
structured-agent-session-status-journal-projection.ts.

* test(native-chat): the admission test's compaction resolves with a real outcome

Main's compaction result is a tagged outcome; the host-level admission test
resolved its mock compaction with an empty object.

* feat(native-chat): the sidebar lists running subagents; the strip, running then the newest finished

The host keeps every child record; what each surface lists is picked from them on every read, so
nothing is stored twice. The status summary, which every session list reads, now carries only
running children (and a finished one whose shell still runs, which reads monitoring): a finished
or failed subagent leaves the sidebar and stays in the chat's strip. The strip lists every running
child, then the newest finished ones, 100 rows in all; more than 100 running all show.

This matches common practice: sidebars show live subagents, and finished ones stay in the chat's
panel, newest first. No wire field is added. An older client reads fewer rows: its legacy task
lists were already live-only in the summary, and the strip's settled tasks come from the same
bounded roster.

* fix(sidebar): one rule for what the worktree sidebar lists: running children, from every source

`worktreeSidebarListsChild` is the one definition: a child that runs, counting a finished one
whose own shell still runs (it reads monitoring). The sidebar's row builder applies it to every
child source it reads, a terminal agent's hook roster and a chat session's records alike, and the
host's status summary applies the same predicate, so the sidebar's payload stays small. The chat's
strip keeps finished children, newest first.

A terminal agent's hook roster already drops a child on its own stop, so nothing changes there:
a teammate between turns and a child gone quiet still run, and still show. The selection module
moves to `agent-child-work-listing.ts`, since it now covers every source, not only chat sessions.

* test(native-chat): the switch test passes the startup child key main's status bar takes

* fix(native-chat): a finished child stays until the user's next send, not a turn Claude opens on its own

Claude wakes the agent on its own when a background task ends, and that wake is a
new root turn. Keying retention on the newest root turn retired every finished
child about two seconds after a background agent or shell finished, so its
outcome never showed in the strip.

Retention now keys on the user's newest send the provider accepted (a message, a
steer or a command), with the journal epoch so a rewind still retires. A turn the
provider opens itself and a subagent's turn carry no send. Replayed captured wake
orders through the real adapter, hook server and status feed.

* fix(native-chat): a background Stop reaches the tasks the child records show

The strip draws a row's Stop, and /clear, /compact and rewind wait for background
work, from the host's child records, but the Claude adapter still resolved which
tasks a Stop reached from its own tracker's roster, and refused to stop at all
once that roster was empty. A task the records kept live after the roster dropped
it showed a Stop that sent nothing and blocked those commands until the chat tab
closed.

The host now resolves the provider ids a Stop sends from the records (the same
per-row rule the strip and admission use; every such row for stop-all), and the
adapter stops exactly those, with no tracker guard. An acknowledged stop ends the
record: a running task sends its own stopped frame first, and the CLI answers
success with no frame for a task it no longer knows. A refused stop leaves the
record live. No production code reads the tracker's roster any more.

* fix(native-chat): one rule for a finished child that still owns live work, at any depth

The listing kept a finished child whose work ran through any depth of ownership,
but retention at the user's next turn protected only the direct owner, so a
finished agent whose finished subagent still ran a shell was removed and that
subagent jumped to the top level. Both now read settledOwnersOfLiveWork.

* fix(native-chat): an older client sees a Codex child's shell as it did before views

Clients that predate child views read a flat task roster derived from the views.
It listed a Codex child agent's shell as an extra row beside the running agent,
then as a bare command once the agent finished. The derivation now hides a
running agent's commands and names a finished agent's as "<agent> — <command>",
as the Codex tracker did; the label rule moves to a shared module both use.

* fix(native-chat): the chat decodes a roster's child rows once, as the frame arrives

The client reducer compared raw wire rows, so a row shaped by a newer host could
throw there, and the strip decoded a new array on every render, which defeated its
grouped-rows memo while a turn streamed. Rows are now decoded where the frame
enters the reducer, an unchanged roster keeps its identity, and the strip's parent
context is rebuilt only when one of its values changes.

* fix(native-chat): the strip channel forgets a closed conversation's roster

It kept the last roster fingerprint of every conversation for the host's lifetime.
The conversation map now tells observers when one leaves it. Also corrects the
summary's children comment: it carries running children only.

* docs(native-chat): rewrap the retention comment

* fix(native-chat): a task's own ending replaces a Stop's, and a child finished after the user wrote stays

Two lifecycle gaps from the round-1 fixes.

A Stop acknowledged ahead of the task's own ending relabelled it. The SDK hands
Orca a control answer as soon as it reads it and queues other frames, so when a
task finished just as the user pressed Stop, the acknowledgement arrived before
the task's completion the CLI wrote first, and the task read "Stopped" with its
result lost. An acknowledged Stop now ends a record provisionally (outcome basis
`stop-acknowledged`); the task's own terminal frame replaces it, and nothing
replaces an ending the task reported itself.

A finished child still vanished with no new action from the user when the send
the provider took was written before the child finished: a steer Claude takes at
its next boundary, or a queued draft handed over at the end of the turn. The
user's next turn now carries when they acted (the send's written time, or the
draft's queued time, both on the host clock), and only children that finished at
or before it retire; a later one stays until the user's next send. A rewind
still retires every finished child.

Also: a Stop-all keeps stopping the remaining tasks after one request fails, then
reports the failure.

* fix(native-chat): the strip keeps one empty list for a roster that omits one

A roster with only running or only finished rows made a new empty array on every
render, so the strip regrouped its rows each time. One shared empty list keeps
its memo.

* fix(native-chat): a strip row whose owner the 100-row budget cut renders under the main agent

The budget can keep a finished child and cut its finished owner. The child still
named that owner, so it rendered nowhere. The selection now clears an owner it
did not keep, as the view contract says for an owner outside the projection.

* fix(native-chat): a stop-all that times out stops asking, and "no children" is sent once

A stop-all kept asking after a request timed out, so a Claude CLI that stopped
answering control requests cost one full deadline per task, while the chat's
sends, Stop and /compact waited behind it. A timeout now ends the loop; other
request failures still let the remaining tasks be stopped.

The chat strip channel never remembered that it had sent "no children", so
every change in a chat with none re-sent that frame to each subscriber. It now
remembers it, and forgets only when the conversation closes.

Also renames agent-child-work-stop.ts to agent-child-work-stop-targets.ts, which
says what it answers: the provider ids a background Stop reaches.

* fix(native-chat): the status feed reads a journal snapshot with no submissions, and e2e tests use the child-work reader

CI on dbd2439cd4 was red in three places:
- first-work-branch-rename and the agentSession.subscribeStatus RPC test feed the status feed a
  journal whose snapshot lists items only. The projection reads the user's newest accepted send
  from `snapshot.submissions`; it now tolerates their absence, as the status projection beside
  it already did.
- the cross-version downgrade test still passed `backgroundTasks` to the teardown's working
  marker, which now takes `childWork`.
- an e2e unit test still gave the status feed the removed `readBackgroundTasks` dependency
  (harmless at run time, a type error in the tests/ project).

* fix(native-chat): the chat strip lists running children only, by the sidebar's rule, and hides when none runs

A finished subagent's result is already in the transcript ("Ran N subagents ·
completed"), so the strip is for work that runs. It now lists exactly what the
sidebar lists, by one predicate (a running child, or a finished one whose own
shell still runs, which reads monitoring), and the host sends no roster once none
runs, so the strip hides.

Gone with it: the 100-row budget and the running-then-newest-finished
selection, the re-homing of a child whose owner the budget cut, and the RPC
gate's rule for a roster of finished rows only (no such roster exists now).
Older clients still get their derived task list, running work only.

Finished records still stay in the host's store until the user's next accepted
message: they refuse a late frame of their run, let a task's own ending replace
an acknowledged Stop's, and keep a running shell's owner. Dating that retention
by when the user wrote the message only kept finished rows visible longer, so it
is removed.

* fix(native-chat): the strip shows running work only from any host, and hides after a released session's last child

- A new app paired with an older host no longer shows that host's finished task
  rows: the strip lists running work only, whatever host sent it, and hides when
  an older host's roster has only finished rows left.
- A test for the path that hides the strip when a session's last running child
  settles after the provider let go of the session (Claude's release path): the
  channel sends `null` though no provider answers for the session any more.
- A test comment still described the strip keeping finished children.
2026-09-30 18:23:24 -07:00
Brennan Benson 5cda0f4508 refactor(native-chat): keep agent-session records in the chat journal database (#24006)
* fix(native-chat): report a failed startup chat reconcile instead of failing app startup

At startup the chat host re-checks every saved chat's lease and writes the
result to agent-sessions.json. If that write failed (the file lock gave up,
the file could not be written, or the file was written by a newer Orca and
is read-only here), reconcileRestartLeases rejected, the startup IPC call
rejected, and the renderer fell into its degraded "Session restore failed.
Changes won't be saved until restart" mode.

The reconcile is bookkeeping: a lease left unreconciled grants no writer,
and every attach, send and read of a chat reconciles its own lease again.
So the startup reconcile now reports its failure through a new optional
host dependency, onStartupReconcileFailure, and resolves. The runtime
routes it to its onError sink under the scope
structured-agent-session-startup-reconcile, or logs it when no sink is
installed (the desktop installs none).

* fix(native-chat): read restored chats without waiting on lease bookkeeping

With native chat on and a chat tab open at quit, the renderer's startup
also awaits the chat tab restore (session.tabs.listAll). That restore
re-ran the lease reconcile before reading each chat and rethrew its store
failure, then recorded each restored tab as visible through a store
transaction that throws on a held lock or a read-only store. Either one
failed the restore, so startup still fell into "Session restore failed".

Reading a chat grants no writer, so the reconcile startup and the restore
run is now a reader's: createReaderReconcile never throws, answers whether
every lease is settled (recovery is resolved only then; the journal opens
either way), and reports each distinct failure once until a reconcile
settles. Attach and agent start keep the strict reconcile. The restore's
tab republish logs a failed visibility write and still publishes the tab,
since a client drops every unpublished chat tab; user-driven publishes
still refuse.

The host dependency is renamed onLeaseReconcileFailure (scope
structured-agent-session-lease-reconcile), since it now also reports for
reads.

* fix(native-chat): keep every record-store write off the startup chat read path

Round-2 review found two more writes on the startup chat restore that
could still fail it and put the app into "Session restore failed":
republishing a /clear replacement recorded its tab visibility strictly,
and resolving a chat's recovery rethrew its store error. The restore
also paid one lock wait per tab and per batch of chats while the lock
stayed held.

The restore now derives tabs from state it already holds:
- publishStructuredAgentSessionTab splits into the strict write and
  projectStructuredAgentSessionTab, which only updates the runtime's
  snapshot. The restore and /clear replacements only project: a saved
  tab index already lists every restored chat, and a /clear moves the
  tab in the same write that commits it. visibilityWriteMayFail is gone.
- Chats a legacy profile restores that the index does not list are
  recorded in one best-effort transaction (store.showSessionTabs), so a
  failure leaves the index absent to seed again rather than partial.
- The read restore's recovery resolution is caught and reported through
  onLeaseReconcileFailure, deduplicated with the reconcile's reports.
- Once lease bookkeeping fails in a restore pass, the rest of that pass
  skips it, so a held lock costs one wait for the startup reconcile and
  one for the restore, however many chats are open.

User actions (create, reveal, attach, send, the /clear commit) keep
their strict writes.

* test: open, seed and read the agent-session record store through one harness

Tests that open the durable agent-session record store, seed it, or read
back what it persisted now go through agent-session-record-store-test-harness.ts
instead of calling AgentSessionRecordStore.open or touching agent-sessions.json
themselves. A later change that moves the store into the chat database then
changes the harness instead of every test. No production code changes.

Tests whose subject is the JSON file itself (its .bak recovery, salvage,
schema versions, permissions, and what older builds read back) keep reading
and writing the file directly; the storage move rewrites or deletes them.

* fix(native-chat): start each restore pass from one lease check and stop its bookkeeping at the first failure

The restore now runs one reader lease check for the pass and lets each chat
re-check and resolve recovery only while the pass is still settled. The
first refusal or failed write clears it for the rest of the pass, and every
chat is still opened for reading. With another process holding the lock,
startup waits on it once in prepare and once in the restore, however many
chats are open; a legacy profile waits once more for its tab-index seed.

* docs(native-chat): correct restore comments and a test name to match the final design

* test: address the record-store harness by the host's state directory

The harness took the store's own folder, so each caller picked one
(join(root, 'store'), or 'agent-sessions' where a test read the store the
runtime owns). A later change that moves the store into the state
directory's journal database could not tell those apart, and would have
had to edit every caller again.

Every harness function now takes the state directory, the one the test's
journal database and recovery capsule already live in, and keeps the
store in the same subfolder the runtime uses. Callers pass that directory;
store-only tests pass their temp directory unchanged. Format tests that
share a directory with harness calls take the file path from
testAgentSessionStoreFilePath.

The folder name moves from a private constant in the runtime to
AGENT_SESSION_STORE_DIR_NAME beside the store's file name, so the harness
shares it without importing the runtime. Its value and every path built
from it are unchanged.

* refactor(native-chat): keep agent-session records in the chat journal database

The record store's records, operation ledger, retired claim keys and chat tab
index become tables in agent-session-journal.db (user_version 4). The version-4
migration copies agent-sessions.json in its own transaction and never writes,
renames or deletes that file or its .bak. Each store write is one journal
transaction over exactly the rows it changed, checked with the load rules; the
file lock, the external-change refresh and its hash, the .bak rotation, salvage
and the hot-path recovery fence are gone from the store.

* wip: importer tests

* test(native-chat): cover the records migration, the import, row writes and read-only records

* docs(native-chat): retire comments that describe the records file as the live store

* test(native-chat): drop the record-store harness's leftover file path and type the import fixture

* test(native-chat): let the host harness cleanup wait out a recovery-offer read's lock

* fix(native-chat): let Stop reach the agent when its ledger row cannot be written

Stop's operation-ledger row now shares the database with the chat history, so
damage, a full disk or a stranded transaction on that write refused the Stop
before the interrupt. A cancel plan now takes its decision from the committed
ledger in memory, runs without settling, and warns that the row was skipped.
Other mutations answer proven damage with the typed "Unable to load this chat."
refusal instead of the raw SQLite error.

* fix(native-chat): answer whether a profile holds chats from the database's rows

Every host install creates agent-session-journal.db, chats or not, and the
version probe created it too, so its mere existence made every profile that
ever installed the host wait on host install and reconcile at startup. The
check now opens the database read-only and looks for a record or tab row,
lets the records file answer while its import is still owed, and counts an
unreadable database as present. The version probe no longer creates the file.

* fix(native-chat): open a chat from history when its tab index cannot be written

Over records a newer Orca wrote, every write is refused, so opening a closed
chat from Agent Session History failed on the tab-visibility write and the chat
read as unreachable. Like closing a tab, opening one now reports a failed
restore-index write and still publishes the tab.

* fix(native-chat): keep the records import owed when the backup read fails transiently

A torn records file whose .bak could not be read (EACCES, EIO) was reported as
unusable, so the migration completed with nothing copied and never retried.
A non-ENOENT read failure of either copy now carries its cause, which the
importer classifies as a read that can clear.

* test(native-chat): pin that an unreadable records file never falls back to its backup

* fix(native-chat): restore imported chats' tabs when the records file had no tab index

A chat created while the import was owed recorded a tab index holding only
itself. When the file it later imported had no index, that index still read
as recorded, so the imported chats' tabs never came back. The import now
clears the recorded marker in that case, and restore falls back to the
profile's tabs.

* refactor(native-chat): drop the unused in-transaction store write

Nothing called it, and it bypassed the write queue and the read-only refusal.

* docs(native-chat): say that an unusable records file is left untouched but never re-imported

* refactor(native-chat): keep the provider handle chain check as main has it

The chain-validation refactor has no measured need in this change.

* docs(native-chat): retire lease-renewer comments that describe the records file as the live store

* fix(native-chat): keep a throwing failure sink from failing the startup chat read

The lease bookkeeping failure reporter called the host's failure sink
directly, so a sink that threw turned a reported, recoverable store failure
back into a rejected startup reconcile or read restore. The reporter now
catches a sink throw and logs both the original failure and the sink error
with console.warn.

* test(native-chat): wait for a replaced host's restart-offer writes before cleanup

A restart test replaces the host without tearing the old one down, so the old
host's fire-and-forget restart-offer withdrawal could still hold the recovery
capsule's lock directory when cleanup removed the test directory (ENOTEMPTY).
The harness now hands hosts a capsule that tracks running operations and waits
for them before removing the directory, replacing the rm retries.

* docs(native-chat): retire the abandon helper's note that the store re-creates its directory

* fix(native-chat): restore a chat opened while the import was owed beside the profile's chats

When the imported records file had no tab index, restore fell back to the
profile's saved tabs, which never list a Claude chat, and the seed then
rewrote the tab table without the chat opened while the import was owed.
The tab rows that chat left are now loaded as unrecorded, restore takes them
together with the profile's chats, and the seed keeps their tab ids.

* test(native-chat): pin that a create whose tab index write fails still opens the chat

* docs(native-chat): say why restore puts chats opened while the import was owed first

* test(native-chat): replace a ledger row rather than change it in place in the Send-now rerun test

The record store freezes published rows in tests, so setting a row's outcome
in place threw; the test now swaps in a changed copy, as its sibling cases do.
2026-09-30 14:45:32 -07:00
Brennan Benson 007b7c0d32 fix(claude): end a message Claude started but never confirmed, and keep Claude running while it holds one (#23898)
* fix(claude): settle a queued send the CLI withdrew from its own cancelled frame

Claude reports each uuid-stamped command's lifecycle (queued, started,
completed, cancelled). A send it withdraws from its queue gets `cancelled`
before the interrupt or cancel_async_message answer, so a lost or failed
answer no longer leaves that send pending: it settles as withdrawn, with the
same reason and words as the receipt path.

A command the CLI already started also ends `cancelled` when its turn is
interrupted or fails, so `cancelled` after `started` is not a withdrawal;
an echoed send has left the waiter lists and is never reached.

Tests replay real 2.1.280 captures, scrubbed.

* fix(claude): release a doubted send when the CLI reports its session idle

A Claude send whose write ended in doubt is recorded `unknown`, and a live
`unknown` reads as work still owed, so the chat showed Working until the
child exited. Claude sends `session_state_changed idle` only once its whole
queue has drained, so it can no longer be holding that send. The runtime now
routes that report to the host's existing release, the same one Codex's
thread-stopped report uses; it retires `unknown` only, never `pending`.

* fix(claude): keep a command's started mark when a redelivery re-emits queued; fixtures name msg_lifecycle_v1

* fix(claude): settle every terminal lifecycle state of a send the CLI never echoed

A send the CLI started, then cancelled before any echo, stayed pending: it may
already be in the conversation, so it is released as doubt (unknown, recovered),
never withdrawn and never re-sent. A late echo still accepts it.

The 2.1.280 schema has two more terminal states. `discarded` (the CLI ended its
session with the send still queued) settles as not delivered; `refused`
(declined before it queued) settles as not accepted by the provider. After
`started`, either one is doubt, as `cancelled` is.

The late-settlement path gains an `unknown` outcome, which the host records as
released doubt.

* fix(claude): release a send the CLI took but left unanswered when it goes idle

`session_state_changed idle` comes only once the CLI's queue has drained, so a
send it took that is still unanswered there got no echo and never will: a turn
that throws can leave `started` with no terminal state. Idle releases it as
doubt.

What proves the CLI took a send is its lifecycle frame. On a CLI that reports no
lifecycle, it is the send's place on stdin: one whose write finished before an
interrupt went out was read before the interrupt was, so the first idle after
that interrupt releases it too. A send armed ahead of the interrupt but written
after it is left alone, since the CLI may still run it.

* fix(native-chat): keep the idle sweep off a Claude child that holds a send

A Claude retrying a rate-limited request has taken the send but echoes nothing,
so no turn row exists yet and the sweep rested the child after the idle window,
turning the send into doubt. The adapter now reports whether the CLI holds a
send (lifecycle `queued` or `started`, not yet echoed or ended), derived from
the live waiters, and owed work counts it.

Nothing is stored: every held send leaves the live set on its echo, its
terminal lifecycle state, the CLI's idle, or the child's exit, so the hold ends
with the send.

* fix(claude): count only a started send at idle and as a held send

2.1.280's end-of-turn cleanup can report idle before it re-reads its queue, so
a send read in that window goes queued, idle, started. Releasing every taken
send at idle doubted that live send and dropped Working. Only a `started` send
is released at idle or keeps the child from the idle sweep; a `queued` one ends
by starting and echoing, by a terminal lifecycle frame, or with the child.

The stdin-order path for CLIs without lifecycle frames is removed: a doubted
send retired there disables content matching on CLIs that mint their own echo
ids, and no Orca failure called for it. Those CLIs keep the earlier behaviour.

Comments that said only a failed write or child exit ends a waiter, or that
idle comes only once the queue has drained, now say what ends one.

* docs(claude): say only what the CLI's lifecycle frames and idle actually prove

* fix(claude): hold the idle sweep while Claude has a send queued, not only started

The sweep rested a child whose CLI had queued a follow-up behind a turn, dropping
the send it had already taken. The hold now spans the CLI reporting it took the
send until its echo, a terminal lifecycle state, or the child's exit. The idle
release still covers only started sends: 2.1.280 can report idle before it
re-reads its queue.

* refactor(native-chat): give provider-proven late dispatch settlement its own module

* test(claude): pin a steer a Stop interrupts after it started as doubt, not withdrawn
2026-09-30 14:11:41 -07:00
Brennan Benson ba39c6d5fd fix(native-chat): a failed startup chat-lease save no longer puts the app into "Session restore failed" (#23964)
* fix(native-chat): report a failed startup chat reconcile instead of failing app startup

At startup the chat host re-checks every saved chat's lease and writes the
result to agent-sessions.json. If that write failed (the file lock gave up,
the file could not be written, or the file was written by a newer Orca and
is read-only here), reconcileRestartLeases rejected, the startup IPC call
rejected, and the renderer fell into its degraded "Session restore failed.
Changes won't be saved until restart" mode.

The reconcile is bookkeeping: a lease left unreconciled grants no writer,
and every attach, send and read of a chat reconciles its own lease again.
So the startup reconcile now reports its failure through a new optional
host dependency, onStartupReconcileFailure, and resolves. The runtime
routes it to its onError sink under the scope
structured-agent-session-startup-reconcile, or logs it when no sink is
installed (the desktop installs none).

* fix(native-chat): read restored chats without waiting on lease bookkeeping

With native chat on and a chat tab open at quit, the renderer's startup
also awaits the chat tab restore (session.tabs.listAll). That restore
re-ran the lease reconcile before reading each chat and rethrew its store
failure, then recorded each restored tab as visible through a store
transaction that throws on a held lock or a read-only store. Either one
failed the restore, so startup still fell into "Session restore failed".

Reading a chat grants no writer, so the reconcile startup and the restore
run is now a reader's: createReaderReconcile never throws, answers whether
every lease is settled (recovery is resolved only then; the journal opens
either way), and reports each distinct failure once until a reconcile
settles. Attach and agent start keep the strict reconcile. The restore's
tab republish logs a failed visibility write and still publishes the tab,
since a client drops every unpublished chat tab; user-driven publishes
still refuse.

The host dependency is renamed onLeaseReconcileFailure (scope
structured-agent-session-lease-reconcile), since it now also reports for
reads.

* fix(native-chat): keep every record-store write off the startup chat read path

Round-2 review found two more writes on the startup chat restore that
could still fail it and put the app into "Session restore failed":
republishing a /clear replacement recorded its tab visibility strictly,
and resolving a chat's recovery rethrew its store error. The restore
also paid one lock wait per tab and per batch of chats while the lock
stayed held.

The restore now derives tabs from state it already holds:
- publishStructuredAgentSessionTab splits into the strict write and
  projectStructuredAgentSessionTab, which only updates the runtime's
  snapshot. The restore and /clear replacements only project: a saved
  tab index already lists every restored chat, and a /clear moves the
  tab in the same write that commits it. visibilityWriteMayFail is gone.
- Chats a legacy profile restores that the index does not list are
  recorded in one best-effort transaction (store.showSessionTabs), so a
  failure leaves the index absent to seed again rather than partial.
- The read restore's recovery resolution is caught and reported through
  onLeaseReconcileFailure, deduplicated with the reconcile's reports.
- Once lease bookkeeping fails in a restore pass, the rest of that pass
  skips it, so a held lock costs one wait for the startup reconcile and
  one for the restore, however many chats are open.

User actions (create, reveal, attach, send, the /clear commit) keep
their strict writes.

* fix(native-chat): start each restore pass from one lease check and stop its bookkeeping at the first failure

The restore now runs one reader lease check for the pass and lets each chat
re-check and resolve recovery only while the pass is still settled. The
first refusal or failed write clears it for the rest of the pass, and every
chat is still opened for reading. With another process holding the lock,
startup waits on it once in prepare and once in the restore, however many
chats are open; a legacy profile waits once more for its tab-index seed.

* docs(native-chat): correct restore comments and a test name to match the final design

* fix(native-chat): keep a throwing failure sink from failing the startup chat read

The lease bookkeeping failure reporter called the host's failure sink
directly, so a sink that threw turned a reported, recoverable store failure
back into a rejected startup reconcile or read restore. The reporter now
catches a sink throw and logs both the original failure and the sink error
with console.warn.
2026-09-30 11:45:05 -07:00
Brennan Benson 75040eba5a test: open, seed and read the agent-session record store through one test harness (#23986)
* test: open, seed and read the agent-session record store through one harness

Tests that open the durable agent-session record store, seed it, or read
back what it persisted now go through agent-session-record-store-test-harness.ts
instead of calling AgentSessionRecordStore.open or touching agent-sessions.json
themselves. A later change that moves the store into the chat database then
changes the harness instead of every test. No production code changes.

Tests whose subject is the JSON file itself (its .bak recovery, salvage,
schema versions, permissions, and what older builds read back) keep reading
and writing the file directly; the storage move rewrites or deletes them.

* test: address the record-store harness by the host's state directory

The harness took the store's own folder, so each caller picked one
(join(root, 'store'), or 'agent-sessions' where a test read the store the
runtime owns). A later change that moves the store into the state
directory's journal database could not tell those apart, and would have
had to edit every caller again.

Every harness function now takes the state directory, the one the test's
journal database and recovery capsule already live in, and keeps the
store in the same subfolder the runtime uses. Callers pass that directory;
store-only tests pass their temp directory unchanged. Format tests that
share a directory with harness calls take the file path from
testAgentSessionStoreFilePath.

The folder name moves from a private constant in the runtime to
AGENT_SESSION_STORE_DIR_NAME beside the store's file name, so the harness
shares it without importing the runtime. Its value and every path built
from it are unchanged.
2026-09-29 23:38:43 -07:00
Brennan Benson 59b746ff3c feat(native-chat): one structured-chat journal database per host, owned by one process (#23613)
* feat(native-chat): one structured-chat journal database per host, owned by one process

Every structured chat on a state directory now lives in one SQLite file,
agent-session-journal.db, opened once by the process holding
agent-session-journal.owner: an empty SQLite file whose held BEGIN EXCLUSIVE is a
kernel byte-range lock, refused while another process holds it and released when
the holder dies.

- Stores own no connection: the per-chat handle, its close contract and the
  close-retry registry are gone; closing a conversation drains its writes, and the
  one connection closes last at teardown.
- The owner lock is taken at runtime start, before orca-runtime.json is written;
  a process that does not own the chats is not published and refuses every
  structured request with journalUnavailable and words that say what to do. It
  retries the lock with backoff and runs the full install once it holds it.
- A journal that will not open fails the host install: every chat says "Unable
  to load this chat." (journalCorrupt), and nothing is renamed, deleted or
  rebuilt. A newer build's database is refused and left byte-identical.
- An append is one INSERT. The listing status is a column, written after the
  rows it describes and keyed by (epoch, sequence).
- A per-chat journal from an earlier build is copied in verbatim (epoch UUID and
  every sequence) on that chat's first open, and its directory is retired only
  after the copy commits.
- auto_vacuum = INCREMENTAL, with freed pages handed back in bounded steps after
  every delete.

* perf(native-chat): key journal rows by block so one chat's rows sit together

Each chat's live epoch owns a block of row ids, block * 2^32 + seq, so a chat's
rows share leaf pages with nobody else's, a replay is one range scan, and
replacing or rewinding a chat deletes one contiguous range. Measured on the
largest real chat (61 MB) beside 19 interleaved peers: 39 ms and 7.5 MB of WAL,
against 214 ms and 102 MB for a (session_id, epoch, seq) key.

- Ids are computed in Number arithmetic, never bitwise. A sequence is refused
  outside [1, 2^32) and a block at 2^21, which keeps every id below 2^53.
- A replace, roll or import allocates a fresh block, moves the chat's pointer,
  and deletes the old block in the same transaction, so no orphan block exists.
- The listing status write moves into its own writer beside the column.

* feat(native-chat): copy a chat's per-chat journal again when an older Orca wrote it after a downgrade

A per-chat journal.db that reappears after its chat was copied in is the newer
history: an older build, run after a downgrade, attached the chat and wrote it.

- journal_imports records the (epoch, tip) each chat was copied from, in the
  copy's own transaction. A file already copied is never copied again, across
  any number of restarts after a failed rename; a file that differs always is.
- Newest writer wins, per chat, with a row saying the chat was continued in an
  older version of Orca. When both builds wrote past the recorded tip under one
  epoch, the copy takes a fresh epoch, so readers reset instead of skipping rows.
- Each copied directory retires to its own .imported-<epoch8>-<ms> name, so a
  second downgrade and re-upgrade never collides with the first.

* test(native-chat): fixture deps match the host journal database shape

Attach-flow and reconcile-attach fixtures stop passing a journal database those inputs do not take, and host and restore fixtures pass the one they now require instead of the removed journal root.

* test(native-chat): state why the runtime-state fixtures' existing casts are safe

* fix(native-chat): start up normally when this process cannot open the chat journal

A process refused the chat journal, because another Orca owns it or because its own journal will not open, failed startup restoration: the window booted in degraded no-save mode and a paired phone could not list any tabs. Startup restoration now treats the refusal structured requests are getting as having no structured host; terminals, tabs and saving go on, structured requests are still refused by the gate, and the install is retried on the next one. Any other install error fails startup as before.

* test(native-chat): name the owner-lock sweep test after the two sweeps it runs

* fix(native-chat): session history and terminal resume work while chats are refused

Session history (listing and preparing a resume) and a terminal typing a resume command only check whether a structured chat owns a provider session. In a process refused the chat journal they failed outright. They now take the refusal chats are getting as having no structured host, the same treatment startup restoration gets, through one shared helper; any other install failure still fails them. Chat requests keep the gate's refusal.

* test(native-chat): the first-work rename's fake journal saves the listing status

* fix(native-chat): open a chat whose per-chat journal file never got its schema

A crash between creating a chat's journal.db and creating its tables left an empty or schema-less file. Each chat used to open that file as an empty chat; the importer instead refused the open as "try again" forever. A file with no journal_sessions table is now read as never written, the same as one with no rows. A file that is not a database, or whose read fails, is still refused.

* fix(native-chat): let the event loop run between chats during startup restore

Opening a chat's journal is synchronous SQLite now that no per-chat directory
is created first, so the restore of every visible chat ran as one main-thread
task. Each chat now waits for a macrotask before it opens.

* fix(native-chat): import a per-chat journal in bounded batches

The one-time copy of an earlier build's per-chat journal ran as one
transaction, which blocked the main thread for 650 ms on the largest chat.
Rows now copy 512 at a time, each batch its own transaction, yielding to the
event loop between batches. The rows go into a block journal_import_blocks
reserves, which no reader follows and no other chat is allocated; the last
batch publishes the chat's pointer, repair marker and import marker together
and releases the reservation. A copy that stops midway leaves only that
block, which the next open clears and copies again. Two opens of one chat
import one after the other.

* fix(native-chat): refuse chats when the owner lock file cannot be opened

A lock file that is not a database, or cannot be opened, made the claim throw
before any refusal was recorded, so startup restoration failed on every
launch. The claim now sits in the same try as the database open and records
the same typed refusal.

* fix(native-chat): finish reclaiming pages a delete frees during a running pass

A reclaim pass ended as soon as the freelist stopped shrinking between steps,
so a second delete that freed more than one step's worth mid-pass ended it
early and left those pages on the freelist. A pass now ends only when a step
itself frees nothing, or the freelist is empty.

* test(native-chat): desktop session history is served while chats are refused

* fix(native-chat): a send to a chat holding a newer Orca's rows says to update

A chat opened read-only because a newer Orca wrote rows to it answered a send
with the generic write failure. It now refuses the way a database a newer Orca
wrote does, with the same reason and words.

* fix(native-chat): keep chat tabs while this process cannot list its chats

A process whose chats another Orca owns, or whose chat journal will not
open, has no structured host. Its session-tabs inventory still answered,
with no chat rows, and the renderer read that as "every chat was closed":
it removed the restored chat tabs and the next session save persisted
their placement away.

The inventory now says `agentSessionsUnverifiable` when the last tab
restore ran with chats on disk but no host to list them. The flag is set
and cleared at the per-client projection point beside the client-hosted
page hold, and the restore is memoised only once a host answered, so a
later lock takeover or journal open republishes the chats and clears it.
The renderer keeps agent-session tabs, and keeps cancellation tombstones,
against an inventory that does not affirm its chat set.

* fix(native-chat): say chats are open in another Orca, with this process's way past it

A process refused because another Orca owns the profile's chats sent the
generic `journalUnavailable` reason, so current desktop and phone surfaces
said "couldn't open this chat's history right now. Try again." — a step
that never helps while the other Orca runs.

The refusal now names its own reason, `journalOwnedElsewhere`, with the
refused process's kind (dev desktop, packaged, orcad) as a fact. Each kind
gets its own step: quit the other Orca, or give this one its own profile
(ORCA_DEV_USER_DATA_PATH) or data folder (ORCA_USER_DATA). The sentences
are added to the shared notice copy, the desktop catalogs in all six
locales, and the boot catalog.

A client that predates the reason reads it as none and keeps the code's
words; an unknown kind reads as the packaged app's step. The `message`
released clients print is unchanged. Which requests refuse does not change.

* fix(native-chat): restore chats on taking ownership, without a list to ask

A refused startup kept its hostless result, so after the owner quit this
process never installed a host, never reconciled restart leases, and kept
telling clients it could not list its chats until a desktop chat request.
Taking the lock now reruns startup restoration once and pushes the chats.

* test(native-chat): a navigation reply says chats are unverifiable while refused

* fix(native-chat): retry a refused owner lock at most every 5 seconds

The lock frees as its holder exits, but a refused process only learns that
on its next retry, and the 30 s cap left a second Orca refusing chats for up
to half a minute after the owner quit. One retry is an open and BEGIN
EXCLUSIVE on an empty file.

* fix(native-chat): install before deciding whether a takeover must republish chats

A list that landed on the refused startup after the lock was taken finished
after the takeover had already checked, so nothing republished. The takeover
now installs first, waits for any restore in flight, and restores only then;
the restore that clears "cannot tell" pushes the frames itself, so a list
that heals the inventory first reaches subscribers too.

* fix(native-chat): no takeover lands a host after the runtime stop

Quitting cancelled a refused claim's retry only at its end, so a retry firing
during the stop's awaits took the lock and installed a host the stop never
tore down, and the lock was then released under an open journal. The stop
now cancels the retry first, keeping the refusal, and repeats its teardown
while an install that began during it (a takeover already under way) is
pending, so no journal connection outlives the lock.

* fix(native-chat): show a thrown refusal in its own words, not its code

A refusal the host throws reaches the client as an RPC error whose message is
the bare code; its reason and facts ride only in the error's data, which no
client read. The chat pane's status line therefore printed
agent_session_journal_unreadable, a send took the bare "not sent" path, and
other writes said the outcome was unconfirmed.

One shared reader, agentSessionThrownRefusal, now reads the refusal from the
error data. A failed history read shows the refusal's read-history words, a
send keeps the refusal behind its Retry exactly as a returned refusal does, and
the other writes (desktop and phone) name the refusal instead of doubting the
outcome. The phone's read failure goes through the same reader.

* fix(native-chat): log a failed journal open once per distinct failure

Every chat request retries a journal open that failed, which is intended, but
each retry also logged the failure with its full stack: a junk database file
logged the same "file is not a database" error 189 times in a minute. The open
now logs a failure only when its code and message differ from the last one
logged, and forgets it once an open succeeds. The retry is unchanged.

The open moves to its own module beside the runtime, which had no room left.

* fix(native-chat): restore lists a chat from its per-chat file and copies it on first use

Startup restore opened every restored chat, and that open copied the chat's
per-chat file into the host database, so the first boot after an upgrade paid
the whole one-time copy before the chat list appeared.

A restore open now reads a chat that is still in its per-chat file straight
from that file, read-only, with the importer's own reader, and closes the file
before moving on. That read drives the listing, the status row and the
restart offer, as it did when every chat had its own file. The copy becomes
owed work on the chat's write queue: it runs before the chat's first write,
and a reader that reaches the chat awaits it. A chat the host already holds,
or that was copied before, still opens through the import and its reimport
rules, and so does a file whose read needs a repair written.

* fix(native-chat): no host stays registered after a stop an install spanned

Each teardown pass clears the registered host before it awaits an install in
flight, and that install registers its host when it finishes. The pass then
tore the host down but left it registered, so a request after the stop was
served by a host whose journal was closed. The stop now clears the slot once
its passes are done.

* fix(native-chat): checkpoint the journal with a full flush on macOS

synchronous = FULL fsyncs each commit, but macOS fsync leaves the drive cache
unflushed, so FULL alone does not survive a power loss there. With
checkpoint_fullfsync, each checkpoint uses F_FULLFSYNC; elsewhere it is a no-op.
The comment that said FULL alone was enough is corrected.

* fix(native-chat): delete a per-chat journal once its copy verifies

An imported chat's per-chat file was kept under an `.imported-*` name, which
doubled the disk its history takes. The copy now reads back from the host
database before it is published: its items, submissions, epoch and tip must
match the file's. Only then does one transaction publish the chat with its
import marker, and the file and its WAL files are deleted, the directory too
when nothing else is in it (a pre-SQLite transcript there is kept).

A copy that does not match is never published: the file stays, the chat is
refused as unreadable ("Unable to load this chat."), and the mismatch is logged
once. A file left behind by a failed delete or a crash matches the marker, so
the next open deletes it rather than copying it again; a file an older build
wrote after a downgrade still differs, and is still copied again.

* fix(native-chat): verify an imported chat a batch at a time

The check that a copied chat reads back as its per-chat file folded both whole,
each in one synchronous task: over half a second on the largest chat. Both
reads now go a batch at a time between turns of the event loop, like the copy
itself, and count rows as well, so a copy that lost a row with no item in it
is caught too.

* fix(native-chat): restore reads a chat's per-chat file a batch at a time

Restore folded a chat still in its per-chat file in one task, so the largest
chat's file held the main thread for about half a second at startup. The fold
now takes the file a batch of rows per turn of the event loop, into the same
fold a replay uses, and nothing reads it before it is done. The file is still
closed before restore moves on.

* fix(native-chat): end a per-chat copy on a turn of its own

A chat's first open ran the copy's last steps (the verified publish and the
per-chat file delete) and the replay of what was copied in one task. The copy
now yields before it returns, so the replay, which every open runs, is a task
of its own.

* fix(native-chat): commit a per-chat copy's batches without an fsync each

Each 512-row batch of a chat's first-use copy committed under synchronous =
FULL, so a large chat paid one fsync per batch, about a quarter of its first
open. The batches now commit under NORMAL, set and restored in the batch's own
task so no other chat's commit runs under it. The publish that makes the copy
visible still commits under FULL, and under WAL that sync makes every earlier
batch durable with it. A crash before it leaves only the unpublished block,
which the next open clears and copies again.

* fix(native-chat): roll back a chat journal transaction whose COMMIT fails

The shared connection's transaction rolled back only when its body threw. A
COMMIT that failed left the transaction open, so every later write, for any
chat, failed with "cannot start a transaction within a transaction", and reads
saw rows that never committed. Under the unsynced copy the failure also tried
to restore the sync level inside the open transaction, which SQLite refuses,
so the caller got that error instead of the COMMIT's.

One transaction helper now covers the body and the COMMIT, rolls back whatever
transaction survives, and rethrows the original error. Schema creation uses it
too. If that ROLLBACK fails as well, the connection is marked stranded: each
later use retries the ROLLBACK, and until one goes through every chat gets the
same "history unavailable, try again" refusal a journal that will not open
gives. The rollback that frees it also restores the FULL sync level.

* fix(native-chat): keep the chat journal connection until its close succeeds

Closing the journal dropped its connection handle before closing it. A close
that failed left the database reporting itself closed with the connection still
open, so the stop that retried the teardown found nothing to close and released
the owner lock over a live connection.

The handle is now dropped only once the close succeeds. A failed close keeps
the runtime pending and the lock held, and the next stop closes that same
connection before it releases the lock.

* fix(native-chat): publish the runtime only once it holds the chat journal lock

When this process could not open the owner lock file at all (a permission
error, or a file that is not a database), the runtime counted that as owning
the chats and wrote orca-runtime.json. That overwrote the real owner's entry,
so the CLI was sent to a process that cannot serve its chats.

A claim that throws is now refused like one another process holds: the runtime
starts but does not publish, the claim's existing retry keeps asking for the
lock, and discovery publishes once the retry takes it. Chats still get the
refusal for the failure itself, and startup restoration reruns on the takeover
the same way it does after another owner quits. A sole process whose lock file
never opens is not found by the CLI until it does.

* fix(native-chat): keep a chat's history when an older build started it over

The first copy deletes a chat's per-chat file, so an older build run after a
downgrade finds no file and starts the chat from nothing. On the re-upgrade that
fresh file was copied in as the newer history, replacing everything the shared
database held for the chat, and then deleted.

A file whose epoch is not the one last copied and that opens with
`session_created` is now kept: neither copied nor deleted, and the chat keeps
the history it has. A file that carried the copied epoch on is still copied
again, as before.

* test(native-chat): pin which chats startup restore copies

Restore copies a chat still in its per-chat file only when restore itself has
to write to it: settling what the last run left open, here a running tool call
or a send handed over and never answered. Every other restored chat stays in
its file until its first use.

* test(native-chat): pin the copy wait on a read that opens a chat restore opened

A read queued behind restore's open of the same chat reaches the conversation
through its own open rather than the listing. It must still wait for the
owed copy, or it reads the chat before its history is in the one database.

* fix(native-chat): record a set-aside per-chat file so no later open reads it

Setting aside a file an older build started over is decided once and kept in
the new `journal_set_aside` table (schema 2, additive), with the file's epoch
and tip as they were. Every later open of the chat skips the file without
opening it, across restarts and after the older build writes more to it:
anything written there grows from that build's own start, never from this
build's history.

The best-effort delete moves beside the per-chat file reader.

* fix(native-chat): set aside any per-chat file at an epoch this build never copied

A chat's per-chat file is deleted once its copy verifies, so a file that
reappears at another epoch was never this build's history, whatever its first
row says: an older build started the chat over, possibly rewinding it after
(`handle_forked`), or rolled the epoch of a file whose delete had failed.
Copying any of them would replace everything the chat holds, so each is set
aside. Only a file still at the copied epoch is copied again (it grew) or
deleted (it did not). The first-row check is gone.

* fix(native-chat): copy a reappearing per-chat file again only while this build has not written past the copy

A per-chat file that an older build carried on under the copied epoch was
copied again even when this build had also written to the chat since the
copy, or had rolled its epoch. The second copy replaced the chat's block,
so what was sent in this build after the copy was gone for good.

Now the file is copied again only when the chat still stands exactly as it
was copied: the same epoch and tip the import marker recorded. Otherwise it
is set aside like any other file that is not this build's history, left on
disk untouched and recorded so no later open reads it. A second copy
therefore never replaces rows this build wrote, keeps the file's own epoch,
and the fresh-epoch rewrite goes away. The row it adds now says the history
includes what the older version recorded, not that anything was replaced.

* test(native-chat): pin that a chat founded here keeps its history, and the v1 schema upgrade

A chat this build founded has a pointer and no import marker, so a per-chat
file an older build later starts for it is set aside. Nothing pinned that
half of the rule: letting such a chat be copied again replaced its history
and every test still passed. A second test pins that a database written at
schema version 1 upgrades in place, gaining the set-aside table and keeping
its import markers.

* test(native-chat): drop a lost copied row by patching the source, not wrapping it

* chore(mobile): restore the mobile lockfile to main's

* fix(native-chat): pass a classified journal refusal through a send or Stop unchanged

* fix(native-chat): refuse a read whose owed copy fails as a failed open does

* test(native-chat): measure only the replace's WAL in the block-key case

Opening the chats starts a free-page pass that waits one event-loop turn,
and the seed never yields one, so that pass was still pending when the
replace committed. It woke during the async stat and reclaimed the pages
the replace freed, adding ~500 KB of WAL whenever the stat lost the race
(Linux CI). Drain that pass before measuring and stub the replace's own.

* test(native-chat): the RPC fixture's status journal can save its listing status

The status feed now hands every projection to the journal, which decides whether it is worth saving.

* test(native-chat): state why the RPC fixture's status journal cast is safe

* fix(native-chat): refuse a per-chat copy whose rows differ from the file, not only its counts

* fix(bench): build the replay benchmark's baseline arm from the base tree and release its handles on failure

* fix(native-chat): retry a failed listing status save on the next read of a cached status

* refactor(native-chat): drop the chat journal owner lock; the process instance lock already guards the profile

The journal carried its own exclusive lock, with a retry loop, an in-process
takeover, lock-gated runtime discovery and a "chats are open in another Orca"
refusal. Every shipped process kind (packaged desktop, serve mode, orcad)
already refuses a second instance on one profile before the journal opens, so
the lock only ever mattered for dev desktops, which the next commit covers at
the process level instead.

The host now opens its one journal connection at install with no lock. What a
sole process whose journal will not open needs stays: the install refusal
recorded for the gate, the no-host startup path, and the unverifiable chat
inventory, now in structured-agent-session-host-refusal.ts. The unreleased
journalOwnedElsewhere reason, its processKind fact and their copy are removed.

* fix(startup): dev desktops take the single-instance lock, and a second one says why it quit

Dev skipped Electron's single-instance lock so parallel `pnpm dev` runs from
several worktrees would not quit silently, but two dev processes on the
default orca-dev profile then write the same stores at once. Dev now takes
the lock like packaged builds: a second launch on the same profile focuses
the first window and exits with code 3, printing one stderr line that names
the taken profile and how to run another copy (ORCA_DEV_USER_DATA_PATH).

Serve mode, the macOS diagnostic bypass and the E2E harness are unchanged:
an E2E launch still skips the lock unless it sets
ORCA_E2E_ENFORCE_SINGLE_INSTANCE_LOCK=1.

* refactor(native-chat): key journal rows by chat, epoch and sequence

Rows in the host's journal database are now addressed by the chat's own
identity, with `(session_id, epoch, seq)` as the primary key, the same
shape each per-chat file already used. The block-keyed layout goes with
everything built on it: the block column and its allocator, the 2^21
block ceiling, and the import's reserved block table.

A first-use copy writes its rows under the file's epoch, which the chat's
pointer does not name until the verified copy publishes it, so no reader
sees a half-copied chat. A try that stopped midway leaves only rows no
pointer names, and the next try deletes them before it copies again.
Replace, rollover and repair delete by (chat, epoch).

This build's history always wins: once a chat was copied or founded here,
any per-chat file that reappears is set aside, and the same-epoch copy
again after a downgrade is removed.

The bounded free-page reclaim after every delete is dropped;
`auto_vacuum = INCREMENTAL` stays at file creation, so a later periodic
reclaim can still be added. Session search keeps its own step.

The schema moves to version 3. Versions 1 and 2 were written only by
unreleased builds of this change and are refused as found, not migrated.

* fix(native-chat): open a chat journal a newer Orca wrote read-only instead of refusing it

After a downgrade, the host's journal database carries a newer user_version. It was refused
outright, so every chat's history disappeared. It now opens on a read-only connection, as the
per-chat journals did: each chat shows what this build can read, from the database or a per-chat
file never copied in, and every write is refused with "Chats were saved by a newer Orca. Update
Orca to keep using them." Nothing is written, copied, repaired or founded, and the file stays
byte-identical. A table the newer schema changed reads as the same read-only refusal, not damage.

* refactor(native-chat): leave the saved listing status to the change that reads it

Nothing in this change reads the per-chat listing status column: it was a stored copy of a fact
the status feed derives, written after every turn end and cleared on every epoch change. The
status_json / status_seq columns, their writer, the saved-status type, the status feed's save and
its retry on a cached projection all go, with their tests. The change that lists chats from a
saved status adds the column back beside its reader.

* fix(native-chat): a chat saved by a newer Orca says to update Orca, not to try again

When a newer Orca wrote the chat journal, this build opens it read-only. A send or a Stop was
refused with the reason `journalUnavailable`, so today's desktop and phone clients chose the
words for an open that can clear: "Orca couldn't open this chat's history right now. Try again."
Retrying never cleared it; only updating Orca does.

The refusal now names its own reason, `journalWrittenByNewerOrca`, whose words are "Chats were
saved by a newer Orca. Update Orca to keep using them." A read refused the same way names it
too. An older client does not know the reason, drops it, and falls back to the code's words
("Orca couldn't read this chat's saved history."), and released clients still print the message.

* fix(native-chat): a chat journal from an unreleased build reads as unusable, not as retryable

A chat journal database stamped with schema 1 or 2 was written only by unreleased development
builds of this change. Opening it threw a plain error, which every chat reported as "Orca couldn't
open this chat's history right now. Try again." Retrying never cleared it.

It now throws a named error that is classified as unusable, so every chat says "Unable to load
this chat." The one log line names the file, says an unreleased development build wrote it, and
says to move it aside. Nothing migrates or renames it.

* docs(native-chat): drop the second-Orca-owns-the-chats case from three comments

The chat-only owner lock is gone, so only a chat journal that will not open leaves a runtime
unable to list its chats.

* docs(native-chat): correct three chat-journal comments the redesign left behind

A per-chat file left without its WAL is set aside, not copied again; nothing runs an incremental
vacuum yet, so the auto_vacuum mode is kept for a later pass; and the idle sweep drops a chat's
in-memory fold, since a chat holds no journal connection.

* refactor(native-chat): stop exporting chat-journal names nothing imports

Each is used only inside its own module now; the teardown's export served a deleted test.

* test(native-chat): name the version-0 test for what it covers, and check every journal table

The test named 'migrates an older user_version forward' covers only a version-0 file that already
has its tables; versions 1 and 2 are refused. The table test now also checks journal_imports and
journal_set_aside.

* fix(startup): a second dev launch's exit line no longer claims it focused a window

The running dev instance may be a background launch or a server, which show no window. The line
now says only that this launch passed its request to that instance.

* fix(native-chat): a failed structured-chat install closes the journal connection it opened

The install opened the chat journal database and closed it only if the record store then failed
to open. A later failure, such as the model catalog wiring or the host constructor, left the
connection open, and the next install opened a second one in the same process. Every failure
after the open now closes it.
2026-09-29 16:42:29 -07:00
Brennan Benson 18327d9665 fix(claude): a queued message Claude withdrew is settled from Claude's own cancelled event (#23862)
* fix(claude): settle a queued send the CLI withdrew from its own cancelled frame

Claude reports each uuid-stamped command's lifecycle (queued, started,
completed, cancelled). A send it withdraws from its queue gets `cancelled`
before the interrupt or cancel_async_message answer, so a lost or failed
answer no longer leaves that send pending: it settles as withdrawn, with the
same reason and words as the receipt path.

A command the CLI already started also ends `cancelled` when its turn is
interrupted or fails, so `cancelled` after `started` is not a withdrawal;
an echoed send has left the waiter lists and is never reached.

Tests replay real 2.1.280 captures, scrubbed.

* fix(claude): release a doubted send when the CLI reports its session idle

A Claude send whose write ended in doubt is recorded `unknown`, and a live
`unknown` reads as work still owed, so the chat showed Working until the
child exited. Claude sends `session_state_changed idle` only once its whole
queue has drained, so it can no longer be holding that send. The runtime now
routes that report to the host's existing release, the same one Codex's
thread-stopped report uses; it retires `unknown` only, never `pending`.

* fix(claude): keep a command's started mark when a redelivery re-emits queued; fixtures name msg_lifecycle_v1
2026-09-29 13:18:53 -07:00
Brennan Benson c5330d0d52 fix(native-chat): stop killing processes that only inherited a chat's spawn tag (#23460)
* fix(native-chat): stop signalling processes that only inherited a spawn token

A spawn token is an environment variable, so every descendant of a provider child
carries it. The Linux-only startup scan treated any carrier no lease claimed as a lost
provider child and sent it SIGTERM, which also hit editors, tmux servers and nested
Orca processes the agent had started. Remove that scan's killing consumer; the token
scan stays for the reservation probe, and recorded owners are still stopped by
identity during recovery.

* fix(codex): remove the token-scan kill path from app-server teardown

Every descendant inherits the spawn token, so killing each pid that carries it can
reach processes the agent started that are not the provider. Production never
injected this path; teardown always uses the process-group and descendant-snapshot
proof. Drop it, its deps, and the now-unused spawn-token argument.
2026-09-28 15:25:24 -07:00
Brennan Benson 153d3fd3fa feat(native-chat): Codex sessions write their subagents into the host status store (#22553)
* refactor(native-chat): the Codex acquire names its turn-boundary methods as a set

Behavior-neutral: the same two methods stamp receipt time. Keeps the file
under the size limit once the child-work sink lands.

* feat(native-chat): Codex sessions write their subagents into the host status store

A Codex child thread and each persistent command become host child records,
fed through the same delivery, ingest and reducer the Claude lane uses. The
child's own turn decides it: turn start is live, turn completion settles it
with the outcome Codex reports, and a follow-up turn reopens the same record
as a new run. Its open tool call, last message, usage and waiting-on-user flag
come from its own thread's frames. A parent turn ending settles nothing.

* fix(native-chat): close a Codex child's tool call by its item id alone

A completion frame need not restate the tool it ran, so reading the tool name
before closing left the call open and the record naming a finished tool.

* test(native-chat): pin the Codex child-work evidence and every hop to the host's records

Child turn start/end/follow-up, open tool call, last message, usage, waiting,
the persistent command a child owns and its monitoring display, a primary
turn end settling nothing, and session end. End to end through the real
adapter: evidence after the journal and the legacy republish, and the parent
state the records imply equals today's at every frame of a scripted session.
Through the production runtime: a Codex session's child work reaches the
status sink under its own address, and a provider exit ends it there.

* test(native-chat): a Codex child's new run never inherits the last run's open call

* test(native-chat): a Codex session with no child-work sink holds no evidence

* test(native-chat): deliver a Codex child's announcement twice, as Codex does, before counting edges

* refactor(native-chat): hand the Codex producer's pending edge over directly

* fix(native-chat): name every Codex turn state in the outcome map; type the runtime test's fake opener

* fix(native-chat): a Codex child's turn ends on the error that ends it, or on its thread closing

Codex can end a child's turn with no turn/completed: an error it will not
retry is that turn's own end (the verdict the transcript already settles the
same turn on), and a closed thread ran its last turn. The executions, the one
owner of child turn state, now end the turn on both, so the strip drops the
child and its record settles (failed, or unknown for a close) together,
instead of reading working for the life of the session. A systemError status
is not an ending: Codex raises it for errors that leave the turn running.

A child fact whose frame names no turn now belongs to the turn the child is
running, instead of counting for every run.

* test(native-chat): a Codex child's turn ending by fatal error or thread close settles strip and record together

* test(native-chat): the Codex parity script reads a waiting child through the shared fold's waiting arm

* test(native-chat): a Codex child row's journal attempt is its record's generation

The journal numbers a Codex child's runs by the turns it observed on the
child's thread; the host record numbers them by the runs its evidence
opened. Both are keyed by the child's own turn id, so they must agree run for
run, including when Codex reports the child's first turn before the spawn
that announces it.

* test(native-chat): a Codex session's end settles its live children and keeps the ended ones

The host no longer erases a session's children when its provider goes away: a
child still running settles with an outcome nobody reported, and a child that
had already ended keeps what it said. The producer tests now expect exactly
that, from the close path and from an unexpected exit.

* fix(native-chat): a Codex subagent's shell is its open tool until the process exits

Codex runs every agent shell through unified exec, so every subagent shell
arrives with the source the persistent-command tracker keys on. The producer
skipped those items, so a working subagent never named its shell, and an
approved command (started on the approval path, completed from unified exec)
stayed its open tool until the turn ended. The tracker still records the
process separately, so a command that outlives the turn reads as monitoring.

* fix(native-chat): a Codex shell becomes a subagent's own work only once it outlives its turn

Codex runs every agent shell through unified exec and never says when one is
left running, so the producer turned every shell, even a millisecond `rg`, into
a command record the moment it started. Each settled into the session's pool
of 32 settled records, so a busy turn evicted a finished subagent's record
(its outcome row would vanish) and listed dozens of finished shells beside it.

A command now becomes a record at the first turn boundary of the thread that
launched it while its process still runs: until then it is the agent's open
call. A shell that exits within its turn never becomes a record.

* refactor(native-chat): child records keep every settled child and can be removed outright

Settled child records now stay until the host drops the session's row; the
32-record trim is gone. A producer can say work stopped with nothing to
report, and its record (and the handles it answered to) goes instead of
settling. Evidence stays host-internal: the producer and the store share
one process.

* fix(native-chat): a Codex command is live work from its start until its process stops

The command tracker is now the one owner of a Codex command's lifetime. It
admits every command whatever `source` Codex tags it with (the approval
path starts one as `agent`), and ends it when its process exits, when its
thread closes (Codex stops the processes first, so no exit ever arrives),
or when the session ends. The producer mirrors that one-to-one: a live
record from the start, removed when the command stops, never settled.

This removes the turn-boundary rule: a command that was only recorded at
its turn's end left the parent reading done for one publish when the main
agent's turn ended with a shell still running. The parity script now
checks the parent at every journal write, not only at frame end.

* fix(native-chat): a Codex command whose approval its turn abandoned never ran

Codex starts an approval's command item before it asks, and when the turn
ends with the question unanswered (the user stops at the approval), it drops
the question and never completes the item. The command tracker admitted that
start as a running process, so the strip kept a phantom command row and the
session row read working until the session ended.

The prompt registry, which owns which approvals are still unanswered, reports
the command approvals a turn ended without; the tracker ends those commands
with the frame that ended the turn. An answered approval keeps its command.

* test(native-chat): start the Codex child-work runtime test without the removed hold

Main no longer has host.hold: creating the session starts its child, and
nothing a viewer does keeps it running. The test attaches and asserts the
one child that attach started, then drives it as before.
2026-09-28 14:56:48 -07:00
Brennan BensonandClaude 7a24d3d335 fix(native-chat): the conversation outlives its agent (#22835)
* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* fix(native-chat): name a chat write by its target, not the owner generation

A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.

Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.

Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.

* fix(native-chat): every journal append reaches the chats that are open

A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.

A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.

* test(native-chat): an epoch replacement reaches the open chat

* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

* perf(native-chat): a publish behind a delivered commit reads nothing

Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.

* test(native-chat): state why the teardown test's fake journal is safe to cast

* docs(native-chat): say mutation admission checks only the writer lease

* docs(native-chat): drop the send rebase from comments that still described it

* fix(native-chat): a message is accepted, then delivered

A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".

A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.

Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.

A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.

Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.

* fix(native-chat): settle queued messages only for the child that ended

A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.

A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.

The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.

* fix(native-chat): an adoption that fails to import keeps the conversation open

The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.

* perf(native-chat): the recovering open reads the journal once

Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.

* fix(native-chat): an attach that fails after indexing its child leaves no child behind

A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.

* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer

The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.

A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.

* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down

The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.

* fix(native-chat): a message rejected while its chat was closed reads as not sent

A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.

* test(orchestration): name why the readiness settlement fakes are cast

* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent

* docs(native-chat): drop the fence from the admission the send effects run behind

* docs(native-chat): give the fence move on release the reason that still holds

* docs(native-chat): stop citing a write fence check in launch and mailbox comments

Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.

* refactor(native-chat): the provider child is its own record

A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.

- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
  patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
  own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
  the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
  dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
  is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
  the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.

* fix(native-chat): the delivery loop alone settles a message its start or child failed

A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.

- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
  reads how it ended: a Stop continues; anything else writes one failure row and rejects every
  queued message with the same words, then stops. A child still starting whose start the adapter
  says did not land fails the same way. The exit, eviction and the settlement retry only settle
  the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
  nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
  failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
  closed, with or without a child, and a start the loop already has in flight is waited for so the
  child it produces is stopped rather than left behind.

* refactor(native-chat): a stopped child ends on the one reading of its stop

The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.

* feat(native-chat): the host says it accepts a send before any agent has it

The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.

* refactor(native-chat): an attach never opens a journal of its own

The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.

* fix(native-chat): a moved fence resends nothing on a host that accepts first

The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.

The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.

* refactor(native-chat): a child's end says whether the user or the host stopped it

The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.

* fix(native-chat): a chat whose only work is a queued message is not offered for resume

A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.

* fix(native-chat): the conversation outlives its agent

Opening a chat no longer starts its agent. A conversation is reached through one host
accessor that opens its journal at rest, and a send is what starts the agent, through
the delivery loop. One idle sweep, every five minutes, stops an agent that has been
quiet for thirty minutes and owes no work, then drops an open journal handle that is
only a cache. Its record, tab, status row and readers stay.

- hold and release are no-ops; hold still builds the host for shipped mobile builds.
- The holders, the holds, the release clock and the exit respawn are deleted.
- Options, the model list, the goal and the context meter answer at rest; a model pick
  at rest is recorded as intent for the next start.
- Compact, rewind, clear and goal changes start the agent first. A send does too when
  a rewind is still in doubt after the conversation opens.
- Orchestration routes mail and group addresses on ownership (the record plus the chat
  tab), not on whether the process runs. An open dispatch keeps its worker running.
- The restart continuation is a send; Resume all holds each slot until the message is
  handed over or rejected.
- A read error never replaces a loaded transcript, and shows the host's own words.

* test(native-chat): type the queued-message fixtures in the resume-offer tests

* fix(native-chat): a start that dies while a message waits on it is that message's failed start

Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.

* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now

* test(native-chat): pin what a failed start settles, and what a resume offer names

A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.

* test(native-chat): the failed-start pins fail on what the message became, not on a timeout

* fix(native-chat): a restart offer ends when the chat's agent starts again

The offer used to end only when the chat's newest user message changed,
because opening a chat started its agent and that start could not be told
apart from real activity. Opening a chat starts nothing now, so the host
reads the fact it already publishes: a chat's status row goes from not
host-owned to host-owned exactly when its agent is started. At that edge the
offer and any failure record for the chat are withdrawn, unless the start is
a resume action's own (its continuation is the oldest undelivered message).

A continuation and a message racing to be first are decided at acceptance:
the continuation is refused, quietly and with nothing filed, when any other
message was accepted since the restart. A failed continuation start leaves
the offer retryable, and each resume action sends its own message id.

Deleted: the newest-user-message comparison, its journal reader, the
continuation filter, and the failure ledger's own "answered by the chat"
check. The marker still carries its message id for one release, so the
previous build can read it.

* fix(runtime): end a transcript stream when its client unsubscribes

Desktop: the IPC subscription controller was dropped as soon as the streaming
handler returned, which for most streams is right after it binds. A later
runtime:unsubscribe then found nothing to abort, so the host kept the subscriber
and derived and sent every publish to a channel no one listened to. The controller
now lives until the renderer unsubscribes, resubscribes the same id, or goes away.

Mobile: disposing an agentSession.subscribe stream now sends agentSession.unsubscribe
with the stream's frame id, so the host ends that subscriber and leaves a sibling
stream on the same socket running. The direct path now passes the frame id the relay
path already passed.

* test(orchestration): the preamble's host stub is typed, not cast

The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.

* fix(native-chat): one fact ends a restart offer: the chat moved on since the restart

The offer is live while no other message has been accepted in the chat since the
restart and its agent has not proved a start since. The offer list, the resume's
reservation check and the continuation's acceptance check all read that one fact,
so a message whose start then failed withdraws the offer too, and a stale click
finds nothing to act on.

The fact is read off the conversation's open handle, which the restart closed, so
it is retired durably whenever it may have changed: a message accepted, a start
proven. A close and reopen within the same run therefore cannot bring the offer
back. A continuation rejected before it reached the agent does not count, so a
retry after a failed start still runs.

Deleted: the quit-time gate on withdrawal, which changed nothing because the
withdrawal and the quit's own offer write share one queue; the per-action
"withdrawn" flag and the separate acceptance check it paired with.

* test(native-chat): an older build reads the restart offer this build records

The offer lives in a file the previous release reads after a downgrade. Pin that
against the pinned release's own capsule, and run the lane when the marker or the
capsule changes.

* fix(native-chat): read a restart offer against where the journal stood when it was taken

"Since the restart" was read off the conversation's open handle, which the idle
sweep closes: after a reopen, a message the user had already sent looked older
than the handle and the withdrawn offer came back.

The offer now records the journal position (epoch and sequence) at the moment
it is taken, and a message accepted after that position, or a journal on another
epoch, means the chat moved on. That is derived from the journal, so it holds
across any number of closes and reopens. An older build's offer has no position;
only a start withdraws it. Because the message half is now durable, the offer is
no longer rewritten in the recovery file on every accepted message; a proven
start still writes it, since only the host that saw the start knows of it.

* test(native-chat): wait for the listing's retire write before reading the recovery file

* fix(native-chat): keep the terminal-backed chat's read error over its local echoes

Messages winning over a read error is right for the structured chat, whose read retries and whose
messages came from the transcript. The terminal-backed view assembles its list from local echoes
too (a launch prompt, a pending send), so a failed read there showed only those bubbles and no
error. Only the structured pane now keeps messages over an error.

* fix(native-chat): a start retries the exit settlement a failed journal write left owed

An agent exit whose journal settlement write failed releases the lease latched until a retry lands.
Reopening the chat used to be that retry; with reveal now only opening the journal, nothing retried
it before the next app launch, and every send was refused. The start the send needs now runs the
retry first, where the attach would.

* perf(native-chat): answer the owner check without opening the chat

Worktree activation calls agentSession.handoffStatus for every chat tab in the worktree, and the
answer comes from the session record alone. Reaching it through the accessor opened each resting
chat's journal (a full read, the crash-boundary write and a restored status publish), then kept it
open for the idle window. It now checks the record and the adapter's support, as before this series,
and opens nothing.

* fix(native-chat): a read waiting on the session lock opens nothing once quit began

The accessor checked for quit before queueing the open, so a read queued behind a session task ran
its open after teardown had begun and indexed a journal no teardown step would close. The check now
runs at the open itself.

* fix(native-chat): read a failed resume's chat before calling it retryable

Whether a failed resume is retryable is the offer's own rule: the chat has not moved on since the
restart, read from its journal. The failure list read it only for a chat already open, so once the
idle sweep closed a chat the user had moved on in, its failure showed Retry again, and the click did
nothing. The list now opens the failed chats first, as the offer list does.

* test(native-chat): type the provider event sink the settlement test reaches for

* fix(native-chat): say the structured read keeps trying only where it does

The structured pane's "Orca keeps trying to load it" line never showed: the view state filled in an
untranslated fallback whenever the read error had no text, and the empty state prefers any message.
The view state now leaves the message out, so the structured pane shows that line and the
terminal-backed pane its own translated one. Mobile's structured lane does not resubscribe after an
error frame, so it no longer makes the claim.

* test(native-chat): await the send's settlement instead of polling for the start

The at-rest send tests polled for the provider start with vi.waitFor's one-second default, which a
loaded machine outran. They now await the host's own settlement of the message.

* fix(native-chat): a restart offer resumes any time after the quit, and knows its own continuations

The continuation's message id was dated by the quit, and the ledger refuses a new id dated more than
a day back, so Resume or Retry a day after quitting was always refused (on main too). It is now
dated by the resume action.

Telling a rejected continuation from the user's own message read the operation ledger, whose rows
expire after about a day; after that a failed resume stopped being retryable. The offer now
records the continuation each action sends on its own capsule entry, bounded to the newest 16, so
the ids end with the offer. The ledger read is deleted.

* fix(orchestration): route no mail to a structured worker its orchestration released

A structured worker is routed on ownership, and a resting worker's lease is released, so ownership
held while its chat tab stayed listed. A worker the coordinator abandoned and then released, found
at rest by the release, therefore still took peer mail and @worktree: broadcasts, and each one
restarted its agent. Routing now also reads the orchestration's own resource row: once it is
released, direct mail, group addressing and worker-show's addressable answer drop the worker, as
they would a terminal worker whose terminal closed. The chat tab stays, and nothing new is stored.

* fix(native-chat): a failed retry names the user's prompt, not Orca's continuation

A resume's continuation is written to the chat before its start, so after a failed attempt the chat's
newest user message is that rejected continuation. A second failure then showed Orca's own restart
text as the chat's prompt. A retry now keeps the prompt its first failure named.

* fix(orchestration): read the released row optionally, as the authority does

worker-show's observation called the row lookup directly, which a runtime double without it threw on
and failed the structured tab-retirement release.

* fix(native-chat): the status bar drops a restart offer the chat moved on from

The renderer re-read the host's restart offer only when a failed chat showed activity, so after a
message withdrew a pending offer the host answered no chats while the status bar kept counting one,
and clicking it opened nothing. The same watch now covers pending offers: a status change in an
offered chat asks the host again, once.

* test(native-chat): a roster of idle or finished children does not keep an agent awake

The sweep reads owed background work through the shared child-work liveness that upstream's
release clock adopted; a child that went idle or finished is not work the agent still owes.

* fix(orchestration): a task dispatched into a resting structured worker keeps it running

The sweep's open-dispatch check read only the worker-start dispatch that owns the worker's terminal
resource, so a task later dispatched to the same worker (orchestration dispatch --to, which writes a
dispatch with no worker row) did not count: after thirty quiet minutes the worker was stopped while
that task was open, and its coordinator read exited. Any unsettled dispatch addressed to the worker's
process incarnation now counts, derived from the existing rows.

* docs(native-chat): comments stop describing the hold this PR removed

Eight comments still justified orderings and teardown choices by a viewer or dispatch hold that
pinned the provider child. Nothing holds any more; the orderings stand for the binding's redrive
subscription and parked mail, and a chat's agent runs from a send until the idle sweep rests it.
Comment-only.

* fix(native-chat): a restart offer keeps the start its own continuation made

Whose start ended an offer was decided at read time, from whether the offer's continuation was
still the queued message. Once the provider refused that continuation, the child it had started
read as someone else's start, so the offer ended and its failure showed no Retry. The delivery
loop now records which queued message a start is for on the in-memory child, and the child's end
carries it; the offer counts a start as its own when that message is one of its continuations.

* fix(native-chat): an agent gets a full idle window after its owed work ends

The sweep measured quiet only from the last journal row, so once a subagent, command, monitor or
dispatch that had outlived the window ended, the agent was stopped at the next tick. A child can
read done before the lead's wake-up turn writes anything, and stopping in that gap loses the
wake-up. The sweep now counts owed work it observes as activity, which gives the agent the full
window afterwards, as the release clock it replaced did.

* test(claude): the options-read fixture runs a live child

The fixture marked its conversation running with a hasProviderChild field the
session type does not have, so the read took the at-rest path and refused a
session with no record. It now carries a child, which is what the read checks.

* test(native-chat): host tests reach its collaborators through a typed seam

The rest-test rig and three test files read the host's private members with
Reflect.get and cast the result. The host now exposes one test-only accessor,
collaboratorsForTests(), and the subscribers class a subscriberCountForTests()
beside its existing retainedActivityCountForTests(), so the tests are checked
against the real types and the casts are gone.

* refactor(orchestration): one owner answers a structured worker's custody

Routing, group addressing, worker-show and the idle sweep each composed their own reading of
whether orchestration still holds a structured worker, so each new obligation or retirement state
had to be added to every reader. structured-worker-custody now derives both answers from the
worker-terminal list state coordinators see in worker-list: addressable is owned and not released,
and owed work is an active custody or an unsettled task dispatched to the same incarnation. The
owner's state is read through the remote dispatch attachment too, as the terminal transfer lookup
already does. Behaviour is unchanged; a settled worker awaiting its coordinator still rests.

* refactor(orchestration): owed work is an open dispatch on the worker's incarnation

A supervised worker's own dispatch context stays open exactly while the worker is active, so the
separate active-custody branch only repeated it. Owed work is now one fact, which also states the
policy that a worker awaiting its coordinator's decision may rest, and both custody decisions are
written once at the top of the module.

* fix(native-chat): a restart offer knows its continuations by a tag in their id

The offer recorded each continuation id in a list on its capsule entry, capped at 16, and a running
action's id in memory. Both could disagree with the journal: past the cap an old rejected
continuation read as the chat moving on, and a crash during a retry restored the failure's older
entry, which lacked the retry's id. Each continuation id now carries a tag derived from the offer
(its teardown and chat), then the action's own part, so any continuation of this offer, queued or
rejected, is recognised from the journal row and the marker alone. The persisted list, its cap and
the in-memory action map are deleted; the agent-start withdrawal keeps an offer whose own
continuation the start was for, read against the stored marker.

* test(runtime): the legacy-worker reveal test judges its stale snapshot inside the wait

The tui-idle probe reads through readTerminal, which now awaits the structured
worker check before the PTY read, so the probe's snapshot request starts a
microtask later. vi.waitFor missed it on its first check and polled again at
50 ms, the same moment the wait's own 50 ms timeout fired. The stale snapshot
then resolved after the wait had already timed out, so the test passed without
judging it, and the rejection landed before any handler was attached. Vitest
reported that as an unhandled error and failed the shard.

Polling every 1 ms sees the request within a few ms, so the snapshot is judged
while the wait is still pending.

* fix(native-chat): a view never restarts a chat whose last start failed

A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.

* fix(native-chat): the idle sweep reads owed work every tick

Owed work counted as activity, but the sweep read it only once the idle window had elapsed, so it
refreshed the clock at most once a window. Work that ended just before the next read left the
agent to be stopped at that read, moments after the work ended, which is the gap the refresh was
meant to cover. The sweep now reads owed work on every tick for a started agent, so the window
always runs from the last tick that saw work owed.

* fix(native-chat): a continuation handed to the agent stays sent

The offer read its own continuation as not reaching the agent while its dispatch was pending, which
also covered one already handed over and still unanswered. When the wait for that answer ended first,
the failure it filed read as retryable, and a retry sent a second continuation to an agent that may
have acted on the first. Only a continuation still queued, or rejected, is now read as unsent.

* test(native-chat): start the child the loop waits on with an attach, not a second view

A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.

* fix(native-chat): settle a gone generation's turn wherever a conversation opens

A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.

* test(native-chat): prove the next child's start settles the turn an earlier child left

The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.

* test(native-chat): count a failed start's rows by row, not by text

Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.

* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget

* test(native-chat): the interrupted create's own retry continues again

The merge of main's lease-latch fix replaced that test's retry of the interrupted create, under its
own operation id, with a fresh start whose result nothing read. That fresh start passes with the
released-reservation continuation deleted, so the case the fix exists for went untested. The retry
and its assertion are main's again.

* docs(native-chat): three comments that still had views starting agents

A start with nothing queued now comes from a command, goal change or rewind; an interrupted compaction
left alone would refuse every send, so no agent would ever start to finish it; and a current host
raises the unattached read refusal only once quit began, with the attach window belonging to an older
host.

* test(native-chat): pin the open's and the send's start and row counts, however the view binds

Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.

* test(native-chat): a reader's open settles the turn a failed exit settlement left running

An exit whose settlement write failed leaves its turn running in the open journal. PR 1's open now
settles it, and this pins the two reads that reach it here: a reader reopening a chat the idle
sweep closed, and a read that opens the chat before the restart restore reaches it.

* test(native-chat): the view-start test's starting window outlasts two subscriptions on a loaded runner

A subscription reads the conversation before it returns, so under load the two views took longer
than the create child's 300 ms start, which then exited before the test checked that it had not.
The child now takes a second to fail.

* fix(native-chat): settle a gone generation's turn at every open but an acquisition's

The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.

* test(native-chat): hold the create's start open until the views bind

The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.

* test(native-chat): a read that reaches a crashed chat before the startup reconcile settles its turn

On desktop the chat on screen at relaunch reads before startup reconciles the leases, while the
dead process's lease still reads live. The open settles the turn it left running anyway, and the
restore that follows finds it settled.

* refactor(native-chat): drop the composer's second error formatter

After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.

* test(native-chat): pin the reason on a message rejected while its chat was closed

The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.

* docs(native-chat): drop the removed dispatch hold from six comments

A worker's session no longer takes a dispatch hold, and no release clock
rests a chat by visibility; the agent-launch comments, the abandon test,
the teardown test and the refusal census still said so.

* test(native-chat): rest the owner-status chat through the idle sweep, not a hold

The activation-gate test from #22808 put its chat at rest by holding and
releasing it, and passed the release-clock grace. This branch deleted both,
so the case threw before it reached its assertions. It now moves the host's
clock past the idle window and lets the sweep stop the agent and close the
conversation, then asserts the same owner answer and activation gate.

* fix(native-chat): show the structured pane's retrying line when a read fails

The read transport always hands the pane the host's words, so the error
state's "Orca keeps trying to load it" line, which showed only when there
were none, was never seen: the pane showed the host's text twice, as its
subtitle and on the status line under it. The structured pane now always
says its read keeps retrying, and the host's text stays on the status line.
The terminal-backed chat is unchanged.

* test(native-chat): wait for a send's background start before the refusal oracle removes its store

An accepted send wakes the delivery loop, which starts the agent in the background. The oracle's teardown disposed the loop but did not wait for that start, so its lease write could create a temp file in the store directory while the directory was being removed, failing the test with ENOTEMPTY about one run in four. The teardown now drains tracked starts before it closes the journals.

* fix(native-chat): a start a message waited on gets one failure row, the delivery loop's

When a queued message's start failed, two writers could report it under the same row: the delivery loop, when the adapter settled the start without proving it, and the exit settlement, when the child's exit landed. The last one won, so the chat's row could name a different cause than the one the message was rejected with, or be written twice.

The exit settlement now writes the start's row only when no message is queued and the loop has not already recorded that start. A start for a command, goal change or rewind, with nothing queued, still gets its row from the exit.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-27 23:46:14 -07:00
Brennan Benson acf8e679ea feat(native-chat): Claude sessions write their subagents into the host status store (#22536)
* refactor(native-chat): the host hands out client delivery's status subscriptions as they are

subscribeStatus and subscribeTurnCompletions wrapped client delivery's bound
methods in forwarding lambdas; they are now the same members, the way
waitForSendSettlement already is. The host is at its size limit, and the
next channel it hands out needs the line.

* feat(native-chat): Claude sessions write their subagents into the host status store

The Claude background-task tracker queues child-work evidence at each decision it
already makes (start, update, progress, terminal frame, roster replacement, turn
end, session end), plus the two facts its legacy row ignores: a foreground child's
progress and a foreground spawn call's result. The adapter drains that evidence
after the journal handled the frame and the parent row was republished, and the
host folds it into one record per child in its canonical store.

Nothing reads the records yet; the strip and sidebar keep their current sources.

* test(native-chat): pin the Claude child-work evidence and the host reduction of it

* test(native-chat): prove every hop from a Claude frame to the host's child record

The adapter delivers evidence after the frame's journal rows and the parent's
republished row; the frame script keeps the parent state today reads while the
records add outcome and activity; the runtime hands the evidence to the status sink
under the session's own address; both entry points wire the sink to the ingest.

* test(native-chat): read an optional task list as optional in the producer script

* test(native-chat): an address whose publish threw carries no child work

* test(agent-status): a foreign record differs from ours by producer alone

* feat(native-chat): a foreground Claude child's own tool call is what its record says it is doing

A child's tool traffic reaches the parent stream only for a foreground child. Read
after the journal handled the frame, the child's newest call still awaiting a result
becomes its open operation, previewed the way a hook-reported row previews its own
tool; the result closes it. The open call is derived from the journal's own
bookkeeping, not held a second time.

* fix(native-chat): a Claude child restarted under a new spawn call keeps reporting to its record

A task that ended and starts again stays hidden from the legacy row until a roster
lists it, so the tracker held no run for it: the new run's progress reached nothing
and a foreground re-run's own spawn result settled nothing. The run is now held
beside the live map, where the legacy row never reads it, until a roster hands it
back or it ends. A parity test pins the record's run count to the journal roster's
attempt on a new spawn call, the one event both count.

* refactor(native-chat): the Claude child-tool queries and translator contract get their own homes

The translator's child-tool queries move into claude-child-tool-queries.ts and its contract
type into claude-journal-translator-contract.ts. Brings the translator back under the size
limit.

* refactor(native-chat): Claude child evidence carries only its own edge's facts

Admission now keeps what a child's record already knows: labels, model,
owner, residency, the last message within an invocation, and a token count
that never shrinks. The evidence side copied all of those forward itself, a
second owner of the same rule. It now sends only what this edge observed,
and a task's token count comes from the frame that reported it.

* refactor(native-chat): Claude child evidence hands admission its raw labels

Admission now folds provider text to one line and drops a malformed fact
instead of refusing the record, so the evidence side no longer folds labels
itself. The description keeps admission's longer bound.

* fix(agent-status): admission alone decides a settled child's second ending

The reconciliation returned before admission whenever a record had already
settled with a definite outcome. That dropped the evidence an `unknown` ending
carries (its last message and tokens), which admission's refine-only rule keeps,
so that rule never ran for the structured producers.

The latch goes. Admission keeps the definite outcome, lands the late evidence,
and refuses a conflicting definite ending as `stale-invocation`, which the host
ingest already counts as the fence doing its job, not a fault.

Pinned through the real Claude producer and the host's own ingest.

* perf(agent-status): keep child records off the status hot paths

Child records made every store write and every status notification scale with the
whole store. Each Claude child progress frame cost about 2 ms with 5 chats holding
~200 child records (about 14 ms at ~1,400), and every status change on any lane
re-parsed every child record just to list parent rows.

- The store derives each frozen record's key once instead of re-parsing it on every
  mutation's validation and every alias lookup.
- Settled history is trimmed only when a batch settles something.
- Parent listing and the structured row's revision stamp read the parents and the
  revision directly instead of building a full snapshot.

A progress frame now costs about 0.3 ms at the same size, and listing parent rows no
longer depends on how many child records the store holds.

* fix(native-chat): an errored Claude spawn result no longer decides how the child ended

Interrupting a foreground Claude agent while it runs a tool delivers the spawn call's errored
result before the child's own killed/stopped frames. The spawn result settled the record
`failed` first, and admission then refused the later `cancelled` as a conflicting ending, so an
interrupted child read as a failure.

An errored spawn result now settles the child `unknown`; the child's own terminal frame refines
it to `cancelled` or `failed`. A successful spawn result still settles `succeeded`. The test
replays both frame orders the real CLI produced when interrupted.

* test(native-chat): pin a Claude foreground child's real finishing order

The real CLI ends a foreground agent with its own completed update, then a notification
carrying the final summary and usage, and only then the spawn call's result. Existing tests
modeled the spawn result arriving first, so nothing checked that the notification's summary
and tokens still land on a record the update already settled.

* perf(agent-status): a store write costs what it touches, not the whole store

With child records on the host, every mutation copied all five store maps and re-validated
every record, and reads scanned every child and alias. A parent status publish cost about
10 ms with 4,000 child records in the store, and a child update about 13 ms.

- A mutation writes into drafts over the committed maps and lands in place; a refused one
  is dropped with nothing to undo. The drafts keep the exact map order a copy would have.
- Only what a mutation touched is re-validated: touched parents, children, aliases, facts
  and tombstones, plus every alias of a touched child and whatever a removed parent owned.
  The full validation stays for snapshot restore.
- The snapshot byte budget is a running total instead of a re-measure.
- Children by parent, facts by parent, aliases by child, aliases by identity and retired
  aliases are indexed, so reads return stored records without scanning or re-parsing.
- The memoized alias identity and tombstone-key checks are gone: indexes derive them once.

A parent publish now costs about 0.015 ms and a child update about 0.06 ms at 40, 1,000 and
4,000 children alike. A seeded fuzz holds the store to the copy-and-validate-everything
path decision for decision, snapshot for snapshot and read for read, and a replica fed the
envelopes ends identical.

* fix(native-chat): a Claude child ends only on its own terminal frame

The child records were fed from the legacy background-task tracker's display decisions, so
they inherited rules that are not truth: a turn ending swept foreground children, a roster
omitting a background child settled it, a foreground spawn call's result ended the child,
and a new background start after any roster produced no record. Captured from the real CLI,
an agent moved to the background keeps its own shell running for 40 s after the parent's
turn ends, and that shell was settled `unknown` at the parent's `result`. Replayed with the
spawn result ahead of the roster, the same agent settled as a false success and was then
revived as a spurious second run.

A new decoder reads the task frames directly. `task_started` opens a child (a start for an
ended task id is a restart, the way messaging a finished agent resumes it), progress and a
live `task_updated` update it, and a terminal `task_updated` or `task_notification` ends it.
Rosters, turn ends and spawn results say nothing about a child. Every child in every capture
gets its own terminal frame, so no evidenced ending is lost. The notification's `tool_use_id`
names the run that ended (captured on a resumed agent's second run), so an ending from a run
that is already over no longer ends the current one; a run id the record never saw still
ends it, so nothing strands.

The tracker, its settled-task retention and the frame readers are back to exactly what main
has: the aggregate-roster split and the restart holding map are deleted, and the legacy row is
unchanged by construction.

* fix(agent-status): a session's end settles its live children instead of erasing them

When a structured session ended, the reducer removed every child record it held, finished
or not, so a reader could no longer tell how the session's work had ended. Now a child still
live when its session ends settles `unknown` (nothing reported how it ended), and a child that
had already ended keeps its outcome. The records still die with their parent: closing or
releasing the session drops the parent row, and the store drops its children with it. A
child's own outcome arriving after the session ended still refines the `unknown`.

The `inventory` and `turn-ended` edges, and the rules that settled children on a roster
omission or at a turn boundary, are deleted: no producer sends them any more. A restart is
now its own flag on a live edge, which is what a producer reports when a finished child
starts again under the same run handle.

* test(native-chat): replay the real Claude CLI's frame orders into a real host

Scrubbed cuts of five Claude CLI 2.1.280 stream-json captures (ids, paths and prompts replaced,
frame order and relative clock kept), replayed through the adapter into a hook server:

- an agent moved to the background keeps its own shell live past the parent's turn, and the
  shell settles at its own notification's time;
- the same capture with the spawn result ahead of the move ends the agent once, from its own
  notification, with no second run;
- a roster that omits a background child without its own ending leaves it live;
- a session that ends settles what still runs `unknown` and keeps every record;
- messaging a finished background agent opens its second run, which ends from its own frame;
- interrupts in both captured orders end `cancelled`, and a finished foreground agent keeps
  its summary and usage.

* test(agent-status): hold the store's running indexes and byte total to a rebuild

The copying-store fuzz never reaches the snapshot byte budget, so a drift in
the running byte total (or any index the public reads do not surface) passed
it. After every fuzzed step, including refusals, compare every index with one
rebuilt from the committed maps.
2026-09-25 12:50:50 -07:00
Brennan BensonandClaude d443320af2 refactor(native-chat): remove the unused terminal handoff (#22783)
* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-25 10:17:36 -07:00
Brennan Benson e0144a9eb6 fix(native-chat): show the Codex and Claude model picker the moment a chat opens (#22756)
* fix(native-chat): show the Codex and Claude model picker the moment a chat opens

A new structured chat showed no model picker until its session had been
created, spawned, initialized and had answered a model listing — and the
picker then listed the models a second time. Codex's listing often goes to
the network, so the picker took 0.6-2 s to appear.

- Keep a host-owned model catalog per agent and account home, persisted on
  success only and refreshed in the background once it ages out. A new
  read-only agentSession.modelCatalog RPC answers from it without a live
  session; sessions reuse it instead of listing again.
- Render the picker while the launch is still provisional, showing the saved
  default. A pick made before the session exists is held and applied once it
  publishes; only the host's acceptance saves it as the default.
- Mark the model and effort set in the user's Codex config as the listing's
  default (config/read), so the first frame names what the chat will run.
- Resolve the account a record-less read would use without running launch
  preparation, which writes and syncs account state.

* fix(codex): disable plugins in the model catalog probe app-server

* fix(native-chat): read the host model catalog only for panes on this machine

* fix(native-chat): name a pre-report model only for a chat this view launched

* test(native-chat): pin the launch latch across publish

* test(native-chat): pin a held pick reaching the host before the first send

* fix(claude): pin the catalog probe's config dir by the session spawn's rule

* fix(native-chat): read the host model catalog only for a visible chat

* fix(native-chat): rewrite the model catalog file only when a listing changes

* test(native-chat): type-check the first-send order fixture

* refactor(native-chat): keep the structured options hook under the line cap

* fix(native-chat): send the first turn only after every pick held during launch settles

* fix(claude): name no default effort from the catalog probe listing

* fix(native-chat): name no listed default model for a chat resumed from history

* refactor(native-chat): let the launch own picks made before it publishes

A pick made while a chat launches had no fence to go to, so the pane held it
and flushed it after publish; every other sender (the outbox, the launch
prompt) then needed its own gate to wait for that flush. The launch now keeps
those picks in its own state, applies them against the create receipt's fence
before it counts as published, and every sender follows publish by
construction. The pane flush, the outbox gate and the module-wide held-pick
registry are gone.

The launch also snapshots the saved selection its create seeds when the intent
is built, so a pick in another chat no longer relabels one still launching, and
a pick the host refuses is reported the way a refused mid-session pick is.

* fix(native-chat): name no default model a workspace's own config can replace

The catalog's default is the account's, read without a working directory, but
a chat runs in its worktree, where a project config (Codex's .codex/config.toml
between the project root and the worktree, or a Claude .claude settings file
that sets a model) picks the model instead. The picker named the account
default there while the chat ran the project's model.

A new chat's catalog read now names its worktree. The host checks that
workspace for such config (existence only for Codex, the model key for
Claude) and, when any is present or the workspace is not a local directory,
serves the listing with no default, so the picker names nothing until the chat
reports its model.

* fix(native-chat): name the listed default model before the report only for Codex

* fix(claude): let an option pick made while Claude starts wait for it instead of being refused

* fix(codex): name no listed default when the configured model is not in the listing

* fix(native-chat): show the picker as unavailable until a published chat attaches

* fix(codex): keep the catalog probe's listing when config/read stalls

* fix(native-chat): write the pending model catalog save before quit

* chore: drop an unrelated lockfile rewrite

* fix(native-chat): rename the catalog store's listing parameter off the global fetch name

* chore: drop an unrelated lockfile rewrite

* fix(native-chat): name the model Claude will run before its first turn

* fix(native-chat): keep Claude's pre-turn applied effort out of the saved session options
2026-09-24 23:04:07 -07:00
Brennan BensonandClaude 6ae6ed08bb fix(claude): open structured chat without a startup deadline, and make Retry start fresh (#22364)
* fix(claude): open structured chat without a startup deadline, and make Retry start fresh

Publish the Claude session as soon as its process is spawned instead of racing
initialize against a fixed 10s deadline. Prompts sent before startup lands are
held and written in order once it does. An exit or sign-in failure before startup
ends the session with the reason and the CLI's stderr.

A create that failed because the process provably exited now carries
ownerVerdict 'exited', so the client marks the launch failed and Retry mints a
new operation instead of replaying the stored failure.

* fix(native-chat): sending into a chat that failed to start restarts it

* fix(native-chat): a send with no live owner restarts it once

A provider child that timed out or exited hands its lease back, and every
later send was refused agent_session_ownership_unknown. Clients read that
code as "not admitted yet" and resend forever, while only a surface hold
could make a new child, once per mount, with its failure swallowed.

The send now routes to a live owner, otherwise restarts one from the
persisted resume state where resume eligibility allows it (single-flight
per session), otherwise refuses with the new settled
agent_session_owner_unrecoverable. Unverifiable, reserved and handed-off
leases are left alone. The desktop hold now logs its failure.

* test(native-chat): pin the unrecoverable refusal as settled in the outbox

* test(native-chat): pin the release clock after a send restarts an unheld owner

* test: read the sent operation id without a cast

* fix(native-chat): type the send-recovery record lookup as the store returns it

* fix(native-chat): a send ensures its owner before admission, and an unheld owner idles for 30 minutes

* fix(native-chat): a create that throws releases its event sink

A child that dies between spawn and journal attach can still write through
the host's event sink, which attach unbound in onAcquiring and never re-bound
because onAttached never ran. The orchestration released that sink only when
performAttach returned a refusal; a thrown failure (the root-exit path) kept
the sink cached with its queued write, so the next attach's drain barrier and
runtime shutdown's flush waited forever.

Also pins the publish-on-root-exit clause for a start that never proved:
deleting it reddened nothing before.

* fix(native-chat): a resend the journal answers restarts nothing, and a send joining a restart rebases from the fence it replaced

* fix(native-chat): the host learns a Claude start positively, and persists only proven options

A publish-first create used to read the session's options before Claude had
answered initialize. With startup pending that read fell back to the built-in
catalog's default, so `record.options.model` was persisted as `sonnet` for
every user whose CLI default is something else; an owner handoff or a reopen
then replayed `set_model('sonnet')` and silently switched their model.

The adapter now reports `started` once startup facts are applied and saved
options restored. The host keeps a `providerChildPhase` on the session it
owns: a starting child hands over nothing but the saved options as intent,
and the `started` event re-reads the options as fact and persists them through
the same record write a user's option change takes. The status summary carries
`hostExecutionPhase` (optional, wire-safe), and the chat pane says the agent is
still starting instead of showing nothing.

A child whose exit already reached the adapter before acquire returns is no
longer handed over as live; the create fails with the CLI's diagnostic.

* fix(native-chat): a hold and a send that find the owner gone share one restart, and a send the ledger already holds restarts nothing

* fix(native-chat): a failed create answers one refusal shape, stamped once at the boundary

A create whose Claude process was seen to exit answered twice in two shapes:
the first call threw a generic runtime error, and only the replay of the same
operation carried the `ownerVerdict: 'exited'` refusal that lets a client
retry under a new operation. Three sites stamped the verdict and the store
failure path stamped nothing.

The first-hand root exit is now returned as the refusal on the first call,
with the provider's own diagnostic as its message. The verdict is stamped in
one place, at the boundary of the attach, from the durable row the operation
settled to, so every refusal shape answers the same fact and no site can
forget it. The per-site stamps are gone.

* fix(native-chat): a send into a session whose child ended restarts it before admission

A session that published and then lost its Claude child before startup (not
signed in, for one) keeps a released lease and a chat the user can still type
into. The send was refused as ownership-unknown, the outbox parked it as
pending admission, and nothing ever restarted the child: the message sat
there until the user closed and reopened the tab.

A send reaching a session with no provider child now runs the same resume a
surface's first hold runs, before the write is admitted. The resume reserves
a new fence, so that send is answered stale with the published fence and the
client's outbox re-drives under it, as after any fence change. A resume that
fails is not this send's answer; admission reports the lease as it stands.

* chore: restore pnpm-lock.yaml to origin/main (local pnpm rewrote it)

* test(native-chat): pin the pre-handover exit as a failed acquire; stub the status feed in the delivery test

An exit the adapter observes before acquire returns now fails the acquire
with the CLI's diagnostic instead of handing over a dead child; the
published-then-ended path stays pinned by the slow-init startup case. The
delivery test renders the pane, which now activates the host status feed.

* test(native-chat): a same-ID re-hold over the wire joins the one resume, and a replay reopen goes on the idle clock

* test(native-chat): a re-hold that joins a failing resume proves one resume ran

* fix(native-chat): a create whose child was proven gone answers the refusal on the first call

The previous change answered a first-hand root exit as the exited refusal on the
first call, but the common failed start never took that path: when the close
ladder proves the whole tree dead the acquisition error is a plain one, the
store-failure classifier rethrows it, and the client still saw a runtime error
first and the refusal only on replay.

The cleanup that proves the child gone now names such a failure
`AgentSessionAcquisitionExitProvenError`, carrying the provider's diagnostic,
unless it already names its own verdict (a refusal, a typed exit proof, a host
store code). The attach answers both proven-exit kinds as the refusal its replay
gives. How a failed acquisition settles and how it is first answered now live
beside the verdict stamp, in the failed-create module.

* test(native-chat): pin the outbox re-driving a stale-refused send under the resumed fence

A send into a session whose child ended is answered stale once the host has
restarted the child. The outbox keeps that operation queued and blocked, and the
fence change the resume publishes re-drives the same operation under the new
fence; the host admits it.

* fix(native-chat): a child restarted for a send nobody holds is still released

The restart a send runs for a childless session takes no holder, on the premise
that the sending surface already holds one. A one-shot writer holds nothing, so
the child it restarted had no release clock and lived until the app quit. The
write resume now arms the clock when no holder is present, as the first-hold
resume already does. The send-after-failed-start cases also pin that the stale
answer's operation is admitted when re-sent under the new fence, and that two
racing sends restart the child once.

* test(native-chat): pin the picked Claude model across a resume whose child starts on its own default

The started event re-reads and persists what the child reports. A resumed child
answers initialize with its CLI default before the saved pick is restored over
it; the record must hold the pick while starting and after started.

* Revert "fix(native-chat): a child restarted for a send nobody holds is still released"

This reverts commit e52c4a6f08.

* Revert "test(native-chat): pin the outbox re-driving a stale-refused send under the resumed fence"

This reverts commit 136a39deb0.

* Revert "fix(native-chat): a send into a session whose child ended restarts it before admission"

This reverts commit 39234e44bf.

* refactor(native-chat): make ensure-owner a step of the serialized send

A send that found the owner gone restarted it OUTSIDE the host's per-session
serialize, through a single-flight resume map shared with the surface hold, then
rebased its fence by heuristic. The attach body is now callable from inside
`serialize` (`attachStructuredAgentSessionUnderSerialize`), and every restart
runs there: a hold, a send's ensure-owner step, provider-exit recovery and the
rewind owner replacement take turns on one queue, so the first to run attaches
and the next finds its child. The single-flight map and `isResuming` are gone.

Admission is two-phase for a send: the ledger's answer comes first and places
nothing; a send it will admit gives the session an owner, and only then are the
row placed and the lease and fence checked. A send it will replay into a closed
session makes the journal readable and spawns nothing. The session entry
records the released fence the child replaced (`resumedFromFence`), so a writer
current as of that owner is admitted at the new fence by bookkeeping, whether it
ran the restart or arrived behind it.

The resume reads its record only after this host has reconciled it and exited
any recovery stage a failed attempt latched, so a hold behind a failed attempt
makes its own attempt against the lease as it now stands.

* fix(native-chat): a Claude start proving itself no longer waits on the CLI

The host handles a Claude child's `started` on the recovery chain every
session's unexpected-exit handling shares, under that session's serialized
step. It then asked the CLI for the model list and settings again, so one slow
CLI held every other session's exit recovery, and its own close, behind up to
two request timeouts.

The adapter already holds those answers when startup proves: the settings read
at startup, the restore's confirmations, and the initialize result the SDK
answers the model list from. `started` now carries that snapshot, and the host
turns it into one record write without any provider I/O.

* test(native-chat): a hold reads its lease only after this host has reconciled it and exited a latched recovery stage

* fix(claude): a chat whose first start failed resumes as the same conversation

A Claude start that dies before initialize writes no transcript, so the next
start launches the chain head's provider id fresh instead of `--resume`. The
launch flag that chose that mode also chose the provider-handle link's origin,
so the fresh launch published a second `created` link onto a chain that
already had a head. The store refused it, the healthy child was closed, and
every later reopen, hold or send spawned and killed another Claude.

The launch now carries the two facts separately: `resumesTranscript` (launch
mode, from whether Claude wrote a transcript) and `continuesChain` (lineage,
from the record's chain head). The link origin reads lineage; rewind and the
Fast opt-in carry-over read launch mode.

* refactor(native-chat): a resume answers with a typed refusal the send classifies

`resumeHeldStructuredAgentSession` and the holds' `ensureProviderChild` answer
`{ ok: true } | { ok: false, refusal }` instead of throwing the refusal code.
The refusal is the attach's own, with its message and, when the failed attach
proved its child gone, its `ownerVerdict`. An attach that settles a failed
acquisition in the ledger and then rethrows the cause is read back off that row,
so a durably failed restart is a refusal and only an unrecorded error is a fault.

The send classifies the refusal through a `Record` over every wire code — a new
code does not compile until it is placed — into transient (the lease is someone
else's to settle; the send runs as the lease stands) or terminal. A terminal
one answers `agent_session_owner_unrecoverable` carrying the cause, forwards the
verdict, and writes the same status row into the chat that a start that failed
leaves, so the user sees why after the error strip is gone. Nothing about the
failure is remembered; a Retry is a fresh attempt. A fault thrown by the restart
itself is reported and the send runs as the lease stands, since bookkeeping
never gates a user's action.

`hold()` still raises the refusal code for its RPC caller.

* fix(native-chat): a child's event sink belongs to the attach attempt that spawned it

The runtime kept one event sink per session id and handed it to every attach.
An attach that acquired a new child unbound that sink first, so when the
acquire then failed its dead child's queued frames stayed in the cached,
unbound sink. The earlier guard only discarded it when no session entry was
left, which a resume of a still-indexed session never satisfies: the next
attach's drain and shutdown's flush waited on it forever. A TUI-to-native
handoff acquire had the same shape.

Each acquiring attempt now mints its own sink. Only a successful attach (or a
proven handoff owner) adopts it as the session's, closing the one it
replaces; any other exit closes it with whatever its child queued. A re-attach
to a live child keeps the sink that child already writes through. A sink that
is not the session's own can no longer force the session's provider down.

The native handoff acquisition moves to its own module, which keeps the
handoff file under its line budget.

* test(native-chat): pin that only the adopted child's event sink still takes writes

Closing a failed attempt's sink and closing the sink a resume replaces were both
unpinned: removing either left every suite green, because neither sink is in the
map that drains and flushes read. The resume test now asserts the failed
attempt's sink and the exited generation's sink refuse writes, and the adopted
one accepts them; deleting either close reddens its own assertion.

* perf(native-chat): the chat reads only the host's startup phase from the status feed

The chat took the whole status summary to read one field, so every status change
for its session (prompt, update time, background tasks) re-rendered the chat
view. It now subscribes with the phase itself as the snapshot, so it re-renders
only when the phase changes.

* fix(native-chat): the startup-phase hook answers a phase or null, never undefined

* fix(native-chat): every restart is counted from the moment it is asked for, and a handoff clears the restart fence

Provider-exit recovery now restarts through the holds' `ensureProviderChild`
like a hold and a send do, so a child whose only surface left while the attach
ran goes on the idle clock instead of living until quit. A hold's resume and a
client attach are tracked as in flight from enqueue, not from their turn on the
queue, so a quit's drain waits for one queued behind a close before it decides
what to evict. A handoff back to native moves the fence in place and now clears
`resumedFromFence`: only a restart may rebase a writer. The failed-restart
status row is keyed by the send's operation id, not the clock, so a resend of
the same id that fails again adds no second row.

* fix(native-chat): a Claude start no longer waits behind another session's exit recovery

The runtime delivered every Claude lifecycle event on the single chain
exit recovery uses so teardown can drain it. That chain orders nothing
across sessions, and an exit recovery on it can run a full reacquisition,
so one chat's `started` waited on an unrelated chat's respawn and kept
its 'still starting' line up. `started` now takes only its own session's
serialized step, is queued the moment it is emitted (ahead of any later
exit of that child), and is tracked in a set the same teardown drain waits
on.

* fix(native-chat): a Claude create that dies at spawn is refused with the CLI's own diagnostic

A CLI that exited before its acquisition handed the child over was
refused with 'claude stream-json for session … exited while being
acquired', or with an unreadable start time, and the stderr the exit
carried (for example 'not signed in') appeared nowhere. The acquisition
now keeps the error its connection ended with and answers with it at
both sites; the generic message is only a fallback when none exists.

* test(native-chat): pin that a reopened Claude chat dying before initialize says why

A resumed start is published at spawn, so a CLI that exits before it
answers initialize fails a chat the user is looking at. Pin that the
open chat is sent the 'stopped before it finished starting' row with the
CLI's diagnostic even when the child's tree cannot be proven gone, and
that a message held for that start is refused rather than left in doubt.

* fix(native-chat): Stop while a Claude start drains its held prompts withdraws the rest

Stop withdrew held prompts only while startup was pending. Once startup
landed and the gate began writing them one by one, a Stop interrupted the
CLI and the prompts still waiting were written straight after it. Stop
now withdraws whatever the gate still holds in both states; the drain
takes each prompt off the queue immediately before writing it, so a
withdrawn prompt can never be written. The one already written still
gets the interrupt.

* fix(native-chat): the release clock keeps a session that still owes a sent message

When the last surface stops holding a session, the release clock evicts
it after the grace unless a turn is running. A message sent while Claude
is still starting is held, not running, so switching away from that chat
for the grace evicted the session and refused a message the user had
already sent. The clock now asks whether the session owes work: a running
turn, or a submission the provider has not taken yet (still pending in
the journal). Both are read from the journal; nothing new is stored. A
starting session that owes nothing is still released, and an explicit
close still ends everything.

* fix(native-chat): only a start that holds a sent message keeps a released session

The release clock kept any session with a pending submission. A Codex
send is admitted and stays pending until its echo, which may never come,
and only an eviction retires it, so such a session was never released
while the app ran. A pending send now keeps the session only while its
child is still starting, which is when the send is held for that start.
Pins that a ready session with an unechoed send is evicted, and that a
Claude chat whose turn finished is released after the grace.

* fix(native-chat): a send waits for the owner it met to prove its start before it is admitted

A Claude child is published before the CLI has answered initialize, so a send admitted right
behind a restart — or right behind the first start — was dispatched into a child that could die
milliseconds later, and learned of the death only as a delivery nobody could confirm. The terminal
refusal the send was written to give was unreachable on the real adapter for exactly the failure
it was written for.

The send's serialized step now admits nothing against a `starting` child. It registers for the
child's startup verdict and returns having placed nothing; the send waits off the session's queue
(the `started` and `ended` settlements run on it) and admits again once the child is `ready`, or is
refused `agent_session_owner_unrecoverable` with the child's own exit reason when it exits first.
The exit settlement writes the one status row, decided by the host's own phase rather than only
the provider's flag. A close, an eviction or a replacement answers the wait too, and quit releases
whatever is left; there is no timer. One spawn per user action holds across re-entries.

* fix(native-chat): restart the release grace when a start writes its held prompts

A prompt held while Claude starts is written when the start lands, but
its turn opens only when Claude echoes it. The release clock stopped
counting it once the child read ready, so a tick landing in that gap
stopped the child before it ran the user's first message. The start
landing now restarts a pending release's full grace, the same grace a
message sent to a ready chat gets before it is released.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): a starting child owns the send; the adapter holds the message for its start

A send that meets a child still proving its start is admitted against it, as it was before the
off-queue startup wait: the adapter holds the message until startup lands and rejects it with the
child's own diagnostic when the child dies first, the exit settlement writes that cause into the
chat, and the release clock keeps a starting session that holds a sent message. The startup watch,
the off-queue wait loop and their teardown phase are gone; the exit settlement still reads a start
that failed off the host's own phase when the provider omits the flag.

The scripted-CLI test now pins that contract end to end: a restart a send asked for whose CLI dies
at initialize leaves the message rejected with the diagnostic, one row naming it, and the fence
moved by two; a healthy CLI is restarted once and written to; a send during the first start is
held and written once initialize answers, or rejected with the diagnostic when the CLI dies.

* fix(native-chat): a failed send restart says why, and offers a new chat only when nothing can restart it

The refusal a send gets when the host cannot restart the chat's agent is renamed
agent_session_owner_restart_failed and now reads "<Agent> couldn't restart: <reason>." with the
restart's own cause. "Start a new chat to continue." is added only when the resume was refused
because this host has no record to restart from or cannot run the one it has. Any other failure,
such as a CLI that is not signed in, leaves the chat retryable: the outbox stops auto-retrying, and
a manual Retry or a new send tries the restart again, since a refusal before admission leaves no
ledger row.

* fix(native-chat): a Claude start skips an option write the CLI never answers instead of faulting at the request deadline

* test(native-chat): wait for the recovery's reserved lease, not the released one it replaces at once

* fix(native-chat): a send whose restarted child dies before starting is rejected with the child's diagnostic

A child that never proved its start has accepted nothing: input is written
only after it initializes. A send admitted against such a child, whose
dispatch then found no session, settled unknown, twice, and the outbox took
Retry away. It now settles rejected with the child's own diagnostic, both
when the dispatch throws and when the exit settles the sends it left
unanswered, so the chat says why and offers Retry. A proven child's
unanswered sends stay in doubt, as before.

* test(native-chat): expect a send held for a start that never proved itself to settle rejected

* fix(claude): write a prompt held after startup already drained, instead of stranding it pending

* test(native-chat): pin that a send to a child that died before starting is answered rejected

* test(native-chat): leave the cast exit-session fixtures as they were, since a proven exit never rejects

* fix(claude): a saved option the CLI never answered stays saved instead of being replaced by the CLI's value

A start skips an option write the CLI does not answer within the request
deadline, and then persisted what the CLI reported in its place, so a slow
answer silently replaced the user's saved model or dropped their saved
permission mode. Silence is not a refusal: the unanswered option is now
recorded apart from a rejected one, the live child keeps running on the
CLI's value, and the saved choice stays on the record for the next start to
retry. An option the CLI rejects is still dropped as before.

* fix(native-chat): a rejected send opens no turn, so the row naming why it failed is not folded away

A send whose restarted child died before starting is rejected, and the
exit writes a row naming the cause. The chat's local clock had watched the
send go pending and stop, so it gave the message "Worked for 0s"; that
settled a turn that never ran, and the fold hid every non-prose row after
the message behind it, including the one naming the cause. The row only
appeared when a later send moved the turn anchor, which read as two rows
for one Retry. The host's journal already says the send was rejected; it
now answers that such a message opened no turn, which outranks the local
clock on desktop and mobile alike. A rejected send whose journal does
record a turn keeps its duration.

* fix(native-chat): a send whose restart died starting leaves the same row as any start that died

One failed attempt already leaves one row, but which row depended on when
the child died. A child that died after the send was admitted left "The
provider stopped before it finished starting: <cause>."; one that died
before the send was admitted left "Claude couldn't restart: <cause>." So the
same failure read two ways from one Retry to the next. When the refused
restart proved its child exited, the send now writes the startup-failure row
itself, as its comment always said it did. The refusal under the composer
still says the restart failed; a restart that failed for a reason other than
a child exiting keeps its own wording.

* fix(native-chat): a send rejected because the agent never started names the cause under the composer

When the child a send was admitted against died before starting, the host
rejected the send with the child's diagnostic behind the internal transport
marker. The client rightly hides that marker's detail, so the red line read
"Couldn't reach the agent" while the cause sat in the record. A startup
death is not a failed write: the host now words that rejection the way the
chat row does, "The provider stopped before it finished starting: <cause>.",
at every site that rejects for it. Desktop and mobile show a reason in words
verbatim already, and older clients do too, so no client change is needed.
Real write failures keep the marker and the generic copy.

* fix(claude): a saved option the CLI never answered survives a later change to a different option

The saved choice a start could not apply was kept on the record, but the next
option the user set persisted only what the child had applied, so changing the
permission mode or effort, or clearing the chat, silently dropped the saved
model. The adapter now reports which saved options are still unanswered, every
option write keeps those saved values, and a write the child accepts for that
option retires it.

* fix(claude): a send that meets a child whose exit already settled names that exit's cause

When the child a send was admitted against died starting and its exit
finished settling before the send reached it, the send was rejected with
"no live claude stream-json session for <id>", now shown under the composer
as the cause. The adapter keeps a settled exit's diagnostic until the chat is
acquired or closed again, so that send names what the CLI said. A refused
restart whose child died at spawn or while its start time was read is pinned
to leave one row in the words any failed start uses.

* test(native-chat): pin the words an exit settlement rejects a never-started send with

The startup gate and the dispatch reject a send first in every existing
scenario, so the exit settlement's own rejection had no test of its wording.

* fix(claude): derive which saved options are still unanswered from what the child applied

A write that lands already puts its option in the session's applied set, so
the unanswered list is that list minus what has since been applied, rather
than a second copy every option write must remember to edit. Session
fixtures built without the new set no longer throw on an ordinary write.

* fix(native-chat): a cleared chat starts from a saved choice the child never answered

Clearing a chat seeded the replacement from the values the child reported,
so a saved model or effort whose restore write the CLI never answered was
replaced by the CLI's own value in the new chat, even though the retired
record kept it. The replacement now keeps those saved values too, and its
start retries them.

* fix(claude): closing a chat forgets its exit's diagnostic even when the exit settles during the close

The diagnostic was dropped when the close began, but closing over an exit
that was still settling finishes that settlement, which kept it again, so a
closed or deleted chat held it until its next acquire. It is now dropped once
the close finishes. Pins that an acquire and a close each retire it.

* refactor(claude): keep a saved option the CLI never answered as the wanted value, not a list beside it

A restore cleared the session's wanted options and added back only the writes
the CLI answered, so an unanswered one lost the user's value and every later
writer had to be told to put it back: the start report, each option change and
/clear each carried a list of unanswered keys. The restore now keeps the saved
value as wanted and unconfirmed, so what the session reports and persists
already carries it, and the list, its adapter method and the started-event
field are gone. A refused option is still dropped.

/clear now starts the replacement from the record's options instead of reading
the child's live values, which can be a model the CLI fell back to.

* test(claude): wait for the start to finish before changing the saved model

The record holds the saved model from creation, so waiting for it returned at
once and the option write could reach Claude while it was still starting,
which refuses it. Wait for the effort the finished start reports instead.

* fix(i18n): translate the still-starting chat notice

The notice that a structured chat is still starting was only in English.

* test(claude): pin the failed acquisition's own reading-control release

The merge re-pointed this test at a child that exits after publish, where the
exit path also releases the binding, so it passed with the acquisition's release
deleted. A child that exits before publish leaves only that release. Also drops
the create 'init' phase, which lost its last producer when rewind stopped
proving before publish.

* refactor(claude): move unexpected-exit handling into the exit lifecycle module

The adapter crossed the 300-line limit once main's context-usage change
landed beside this branch's growth. The two methods that turn a Claude
process exit into an ended event now live next to the existing exit
helpers; behavior is unchanged.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-24 15:23:15 -07:00
Brennan Benson 60bd1dfdea feat(native-chat): one shell-environment setting for every structured chat (#22387)
* feat(native-chat): one shell-environment setting for every structured chat

Structured Codex chats started from the login-shell environment, while
structured Claude chats started from Orca's own process environment, so a
variable exported in .zshrc reached one and not the other. Both now start
from the same base, chosen by a new setting:

- on (default): the whole login-shell environment, as a terminal gets
- off: Orca's environment plus PATH, locale, SSH_AUTH_SOCK, and the
  variable names the user lists

The setting is re-read each time a chat starts or resumes. It is shown
only when Chat UI, the Chat UI default view, and structured native chat
are all on. Terminal-backed chat is unchanged.

* fix(native-chat): normalize the shell-environment settings when a profile loads

A hand-edited settings file could store the variable list as something other
than an array, and the structured runtime called `.filter` on it per launch, so
a malformed value failed every structured chat create and resume, and the
settings pane render. Normalize both keys where the profile loads, the same way
the other array settings are, through one shared normalizer the runtime policy
also uses. Also pin that an uncommitted name draft survives an unrelated
settings re-render.

* fix(native-chat): keep the pinned account as the only source of a structured chat's Claude home

The session record owns which Claude home a structured chat uses, and the
acquisition pin (claudeConfigDirEnvPatch) is the only emitter of
CLAUDE_CONFIG_DIR, compared against what the child would otherwise inherit.
With the login-shell snapshot as the inherited base, a CLAUDE_CONFIG_DIR
exported only in a shell rc flipped that comparison and produced an explicit
pin to the CLI default home, which moves the CLI off its default Keychain item.

Drop the inherited CLAUDE_CONFIG_DIR in the Claude launch resolver before the
pin runs, as Codex already does for an inherited CODEX_HOME. A configured
per-agent overlay still passes through, since the record already honors it.

* fix(native-chat): drop Orca's own CLAUDE_CONFIG_DIR from a structured Claude child too

The process spawner merges Orca's process env under the launch env, so a
CLAUDE_CONFIG_DIR exported to Orca itself reached the child around the launch
resolver's drop and unseen by the account pin. One helper now strips it from
both inherited bases. Also declare the two shell-environment settings on the
runtime store contract and add the six new strings to every locale catalog.

* feat(native-chat): add shell variables one at a time with a removable list

* fix(native-chat): return focus to the name input after removing a shell variable

* fix(native-chat): use a neutral placeholder for the shell variable input

The empty input showed a grey HTTPS_PROXY as its placeholder, which reads as a
saved value, especially right after that exact entry is removed from the list.
Use "Variable name" instead, in every locale catalog.
2026-09-23 14:29:20 -07:00
Brennan Benson dc8cf30554 fix(native-chat): end a structured turn when the agent reports it failed (#22047)
* fix(native-chat): end a structured turn when the provider reports it failed (#22044)

A turn reads as working while its durable turn row says `running`, and only two
events could write a terminal row: the provider's turn-completed notification and
the provider process going away. A provider error that ends a turn is neither, so
the row stayed `running` with nothing re-deriving it, and the chat counted
"Working for N" for the life of the session.

Codex reports such a failure as an `error` notification naming the turn it ended,
with `willRetry` distinguishing it from a stream error it is about to retry. That
frame now settles the turn it names. Claude's CLI reports the same through its
session-state frame, whose `idle` arm the SDK documents as the authoritative
turn-over signal; that now settles the open turn too.

Codex's `thread/status/changed` deliberately settles no open turn: the app server
clears `running` on every error, including ones it reports as not affecting turn
status, so a turn still open there is still running. What it does settle is a send
whose dispatch was never answered — a timed-out dispatch is recorded as unverified
delivery, reads as work still owed, and nothing in a live session retired it.
Retiring it never makes the send re-deliverable.

Splits the codex notification translator so the file stays inside its line budget.

* fix(codex): defer idle dispatch release until turn end

* fix(claude): enable session state lifecycle events
2026-09-21 12:44:13 -07:00
Brennan Benson 434365d2de Offer to reconnect native chats that were working when Orca restarted (#21096)
* feat(native-chat): resume structured chats that were working at restart

Teardown records a marker for every session this host was genuinely running a
turn for, derived from the LIVE runtime rather than a persisted status row, so
a stale `running` row left by an older crash can never trigger a resume. On the
next launch a modal lists exactly which chats would resume and resumes them via
native continuation (Claude resume/resumeSessionAt, Codex thread id) — never by
re-sending the prompt, which is what makes an agent redo finished work.

A session resumes only when all of these hold: a teardown marker exists and has
not expired, the record's lease is released and reconciled, a provider resume
cursor exists and still matches the marker, the journal's own turn record names
the same turn, and the marker has not already been spent. Markers are consumed
before the resume is submitted, so a crash mid-resume cannot double-fire, and an
admission gate refuses a second concurrent resume for one session. Resumes are
staggered three at a time rather than spawning every provider at once.

The modal's "Don't ask again" checkbox writes the nativeChatResumeWorkOnRestart
setting, which Settings can turn back off; automatic mode runs the identical
predicate and staggering and reports what it did. Declining consumes the markers
so the prompt cannot return every launch — nothing is lost, because opening a
chat still re-acquires it at the same cursor.

* fix(native-chat): compare handle ROOT and turn state when offering a resume

Four defects QA found in the restart-resume offer, fixed together because the
first two interact: shipping the root fix without the state fix would convert a
silent no-op into actively offering finished chats.

1. Claude was never offered (0/4). The marker recorded agentSessionProviderHandleKey,
   which embeds Claude's leaf uuid — a branch cursor. The adapter's own close path
   appends a `resumed` link with an advanced leaf during the SAME teardown, so the
   marker went stale seconds after it was written and the drift guard refused every
   Claude session forever. Record and compare agentSessionProviderHandleRoot instead:
   the root is the part a resume must preserve, and changing it is a fork, which is
   exactly what this guard is for. Codex is unaffected (its thread id is the whole
   key) but uses the root too, so the rule is uniform.

2. The predicate compared turn IDENTITY but discarded turn STATE, so a `completed`
   turn satisfied it as readily as an interrupted one. Eviction rewrites `running`
   to `interrupted` and never to `completed`, so the state is what separates work
   that was cut off from work that finished. Require `interrupted` or `unverifiable`.

3. A chat blocked on a pending approval or question was marked as working, because
   the teardown reader accepted any `running` turn while the product's own projection
   calls that state `attention`. Teardown now defers to that projection: an agent
   waiting on the USER is not interrupted work.

4. "Resume all" could silently no-op. The modal fetched candidates at mount; by click
   time the chat's own pane may have bound and taken the hold, moving the lease to
   `live` so the predicate dropped it and the call returned no results, leaving the
   dialog open behind a dead button. Re-derive at click time and settle an
   already-live session as resumed — it is running, which is what the user asked for.

Test fakes now model the Claude close path that advances the leaf, which is why no
unit test could previously exhibit defect 1. Ablation covers all eleven guards.

* fix(native-chat): gate the already-live settlement on the full resume predicate

Two follow-ups from re-QA, both cases of a rule stated by intent rather than by
discriminator.

1. The already-live path bypassed the predicate. "Resume all" sends no session
   ids, so the fallback's target set was every marker, and it was gated only on
   the session having a live provider child. A chat the predicate had refused --
   a completed turn, say -- whose pane happened to own the lease was therefore
   settled as `already_live` and had its marker spent, inflating the "Resumed N"
   count with chats that were never eligible. No provider spawned and no tokens
   were spent, but a marker the predicate rejected must never be consumed.

   The resumable set now takes an explicit `leaseState`. The already-live path
   derives a second set with ONLY the released-lease clause relaxed, and settles
   a session just when it is in that set. Every other clause still applies.

2. The `attention` rule was one-sided. Teardown refuses to mint a marker for a
   chat blocked on the user, but the set predicate had no equivalent, so a marker
   arriving by any other route was offered once eviction rewrote its turn to
   `interrupted` -- the same asymmetry the completed-turn case had.

   Gated on projectStructuredAgentSessionStatus === 'attention'. That projection
   tests for a pending approval or question BEFORE it looks at turn state, so it
   still reports `attention` after the turn is settled, which makes it the durable
   signal and keeps one source of truth with teardown.

Ablation now covers thirteen guards, including one for each of the above.

* fix(native-chat): capture awaits-user on the marker instead of re-deriving it

The awaits-user clause could never fire. It asked the live projection for
`attention`, which needs a prompt whose resolution is still `pending` -- but
teardown CANCELS that prompt a few phases after it writes the marker. By the next
launch the evidence is gone, for precisely the sessions the clause was written
for. QA measured the injection still being offered and then resumed.

This is the same shape as the leaf-drift bug: state read after teardown is not the
state that justified the marker. The discriminator, now applied across the whole
predicate:

  - a fact teardown itself destroys or mutates must be CAPTURED on the marker
    while it is still true;
  - a fact that evolves on its own must be RE-DERIVED at read time, never
    snapshotted.

So `awaitsUser` is now recorded at teardown and the predicate reads the recorded
value. Teardown still declines to mint a marker for such a session, so the
recorded flag is the second line rather than the only one.

Audit of every other clause against the same test:

  - turn id (captured) -- teardown rewrites turn STATE but never the id. Correct.
  - provider handle root (captured) -- the close path appends a resumed link, and
    appendAgentSessionProviderHandleLink refuses one that changes the root, so the
    root is invariant under exactly the mutation that broke the key. Correct.
  - turn state (re-derived) -- DELIBERATE exception, stated here rather than left
    implicit: we are not reading the state that justified the marker, we are
    reading teardown's receipt that it settled the turn. A turn still `running`
    means eviction never finished, and we refuse. Correct, and intentionally so.
  - lease reconciled / released / handoff stage (re-derived) -- these answer a
    different, launch-time question: may this host take the lease NOW. The
    teardown-time value would be meaningless, and `unreconciled` is cleared by
    this launch's own reconciliation. Correct.
  - adapter support, marker TTL, marker consumption (re-derived) -- all evolve
    independently of teardown. Correct.

Only awaitsUser was on the wrong side.

* fix(native-chat): drop the unreachable awaits-user marker flag

The captured flag was dead code. `awaitsUser` could only be true when the
projected status was `attention`, and `attention` hits the `continue` above the
push -- so every marker teardown can ever write carries `false` (QA measured
22 of 22 across two real teardowns). The predicate clause reading it was
unreachable by any production path.

A flag that is structurally always false is worse than no flag: it reads as a
safeguard, so the next person to touch this trusts it. The asymmetry it was
added to close was only ever reachable by fault injection, because teardown is
the sole writer of markers and already refuses attention sessions.

Removing it also drops an upgrade discontinuity: as a required field it made a
marker written by the previous build fail validation and be silently discarded,
costing a resume offer on precisely the upgrade where the user was mid-turn.
Markers predating the providerHandleRoot rename still will not parse, but those
carry a leaf-sensitive key the predicate would refuse anyway, so nothing usable
is lost.

In its place the teardown gate now states that `status !== 'working'` is the
SINGLE gate for awaiting-user sessions, why a predicate-side mirror would be
unreachable, and why it could not even re-derive the fact -- so the reasoning is
inherited rather than rediscovered.

Ablation is back to twelve guards; every other clause is unchanged.

* fix(native-chat): say reconnect, not resume, and show each offer's age

Two changes, both independent of the parked continuation decision.

1. The copy claimed something QA disproved. "Resuming continues each agent where
   it left off" is false: reconnection restores the session at the point it
   stopped, with full context and without re-sending the prompt, but the
   interrupted reply does not continue on its own. The toast's "Resumed N chats"
   implied work had restarted.

   Audited every user-facing string against the rule that none may claim work
   continues or that a reply resumes -- which caught more than the three strings
   the fix started from. The title, the row button, "Resume all", "Resuming...",
   the not-now hint ("picks it up where it left off"), the checkbox and its hint
   ("resume on their own"), the list's aria-label and the Settings row all made
   the same claim. The user-facing verb is now reconnect throughout; the body and
   update variant state outright that the interrupted reply will not continue.
   en.json synced, runtime boot catalog regenerated.

   If we later decide to send a continuation instruction, this is one commit to
   change back. Shipping text we know to be false was the worse option.

2. Rows now show each offer's age. The TTL is 24 hours and a stale offer looked
   identical to a fresh one. The marker already carried `recordedAt`, so this is
   a render change plus one field on the renderer's candidate type, formatted
   with the existing formatUiRelativeTime helper rather than a new one.

   The clock is stamped once when the list arrives rather than read during render:
   ages then stay stable across re-renders, and the render stays pure, which the
   react(purity) rule requires.

Guards, predicate and RPC are untouched; ablation still covers twelve.

* feat(native-chat): show the workspace name on each reconnect row

A row read `codex · folder:8f3a1c22-… · 8 hours ago`. Recognising which chats
would reconnect is the entire point of the list, and at twenty rows a UUID
identifies nothing.

No RPC or host change was needed: the renderer can already resolve this id.
Resolved the way automation dispatch resolves the same id space
(resolveAutomationDispatchWorkspace) -- a folder workspace by its full
`folder:<uuid>` key via getKnownWorktreeById, a git worktree by its bare
`repoId::path` id via allWorktrees. Both return a Worktree, whose displayName is
a required field, and DetectedWorktree extends Worktree so either shape answers.

Falls back to the id when nothing resolves, which is what the row showed before
and also covers the window before the worktree store has hydrated.

The lookup lives in a per-row subcomponent because a hook cannot run inside
`map`, and its selector returns a primitive string so repeated selector runs
cannot churn referential equality.

* feat(native-chat): group the reconnect modal by worktree and add opt-in continuation

Grouping. Rows are now grouped under a worktree heading with the repo glyph and
an agent count, using the sidebar's own collapse mechanics. Only presentational
pieces are reused -- RepoIconGlyph, CompactAgentExpansion, AgentIcon and
formatShortTimeAgo. The sidebar's agent row cannot be: worktree-card-compact-agent-row
imports DashboardAgentRow, the dashboard's own type, so both surfaces render one
live-agent model requiring a pane, tab and status entry. Every chat offered here
is by definition stopped, so supplying that would mean inventing live state.

Two things I had assumed were reusable and were not:

  - DashboardHostBadge returns null unless hostKind is ssh or remote. Structured
    chat is local-only, so it would always render nothing. The host line is
    omitted rather than faked; the badge is the right element to add if and when
    structured chat gains remote support.
  - No state dot. Every AgentDotState misleads here: idle and unverifiable both
    presuppose a live pane, interrupted renders red like an error, done green,
    working a spinner. A missing dot beats one saying these agents are running.

One worktree renders flat with no heading -- a name, count and chevron around a
single group says nothing the dialog has not already said.

The age column now uses formatShortTimeAgo for sidebar consistency. It takes
(timestamp, now) and subtracts internally rather than taking a delta, so the call
is (recordedAt, listedAt); passing the old delta would have rendered plausible
nonsense. The clock is still stamped once into state, so ages stay stable and the
render stays pure.

Continuation. A secondary "Reconnect and continue" action sends one message, from
a single shared constant, identical for both providers. Reconnect is unchanged and
still sends nothing. An info popover quotes the literal message read from that
same constant, so what is shown cannot drift from what is sent.

Ablation now covers fourteen guards. Two are new: continuation only follows a
reconnect that actually happened, and -- inversely -- a send injected into the
reconnect path must turn the test red, since "don't ask again" rests on reconnect
never sending.

* feat(native-chat): say terminal sessions kept running, and clear the quality gate

The modal lists stopped chats with no way to tell that CLI agents are fine, and
the true state of the world is counterintuitive: the terminal sessions survived
the restart and the chats did not. One line now says so, next to the heading
where it frames the list rather than as a footnote at the bottom.

Wording follows the app's own vocabulary rather than inventing a term: the
catalog settles on "terminal sessions" (terminalSessionCount, "Terminal sessions
are grouped by workspace", "No terminal sessions yet"), and UpdateCard already
reassures with "Your terminal sessions won't be interrupted during the update" in
the same text-xs text-muted-foreground treatment. "kept running" rather than
"were restored" -- nothing reconnected them, they never stopped, and the line
says nothing about why.

Also clears check:code-quality:changed, which I had not been running -- oxlint
alone covers neither the design-system nor the casting audit, so 18 findings had
accumulated across the branch.

  - design system (4): Button spacing hand-rolled as gap-1/px-2 is just size="xs";
    PopoverContent and DialogTitle own their typography and spacing, so the
    text-xs moved to the popover's own children and the title's icon gap moved to
    a plain wrapper.
  - casting (14): production code loses its assertions outright via Reflect.get,
    the idiom already used in managed-hook-detection-commands and
    worktree-name-retirement. The marker validator reads each field through
    Reflect.get and now checks recordedAt is a number rather than asserting it;
    the store-file parse uses the existing `file` shape instead of a second
    assertion; the runner narrows the admission error's owner with typeof.
    Test fixtures keep their assertions behind the line-specific SAFETY:
    rationale the repo mandates for exactly this case.

One trap worth recording: the audit reports an assertion at the line its
EXPRESSION OPENS, not where `as` appears, so a disable-next-line above the
closing brace of a multi-line literal is inert and silently changes nothing.

Guards unchanged; ablation re-proved 14/14 at this head.

* fix(native-chat): give the reconnect row's provider icon an accessible name

Every row rendered the provider as a bare AgentIcon, whose svg carries no
aria-label, title or alt. With a Claude chat and a Codex chat in one worktree the
two rows were identical to any non-visual consumer, and the dialog offered
several identically-named "Reconnect" buttons with nothing to tell them apart.

A regression from 233e37b2bd, where the row read `${agent} · ${workspace} · …` as
text. Moving the workspace name into the group heading was right; dropping the
provider to an unlabelled glyph is what lost the information.

AgentIcon takes no label prop, so the icon is wrapped the way
NativeChatSupportedAgents already names it: a span with role="img" and an
aria-label from formatAgentTypeLabel, the same labeller the sidebar and dashboard
rows use.

The per-row button also names its agent now ("Reconnect Claude chat"). The
identical buttons were half the reported harm, and an accessible name that opens
with the visible word keeps WCAG 2.5.3 satisfied. Say so if you would rather ship
only the icon label -- it is one attribute and one catalog key to drop.

Age code untouched, as asked: formatShortTimeAgo still takes (timestamp, now) and
is still called with (recordedAt, listedAt).

* fix(native-chat): scope resume markers to one launch and report the real dispatch

Three defects in the restart-resume path, all of which could resume a session
that was not genuinely working or claim one was continued when it was not.

Launch scoping. A durable marker with a 24h TTL is a write-ahead latch: a
teardown write that failed or timed out, or a store restored from its backup,
left a previous generation's marker actionable, and automatic reconnect would
have acted on it silently. Markers now carry the id of the launch that wrote
them, and only the launch immediately after may claim them. The launch id lives
in its own file with no backup mechanism, so it cannot roll back in step with
the markers it is proving adjacency for. Startup claims the previous launch's
markers into launch-scoped memory and deletes every durable copy in the same
step, so the durable fact dies at claim time rather than at use time. Both
halves fail closed: an unprovable predecessor and a clear that throws each
claim nothing.

Dispatch states. The send layer answers ok as soon as Orca owns the message;
the provider's own answer lives in the submission. Continuation read only the
envelope, so a rejected turn/start was reported as continued and stamped the
journal saying the agent had been asked to carry on. All four states are now
preserved, and only an accepted dispatch appends the attribution note.

Claude pre-echo sends. Claude cannot write a running turn until the SDK echoes
the user message back, which is seconds on a real journal, so a turn-id-only
marker dropped exactly the sessions that were working hardest. A send that has
not become a turn now carries its own identity, and the launch-side predicate
asks the journal about that submission's dispatch state instead.

* fix(native-chat): follow an accepted send to its turn, and settle before judging

Two defects found in QA, both reproduced twice.

Follow the submission forward. The launch-side predicate accepted a
submission-shaped marker only while its dispatch was pending or unknown, but the
window in which work is submission-shaped is precisely the window in which the
dispatch is about to be accepted: the send settles during teardown and the turn
it opened is then cut off as interrupted. Judgement was frozen at the moment the
marker was written, so the predicate refused the very sessions this was built
for and fired only when the send never reached the provider. An accepted
submission is now followed to the turn it opened -- matched through the user
item key a turn names and a submission is aliased by -- and that turn is judged
by the existing turn rule. Accepted alone still proves nothing: without the link,
or with a turn that completed, this refuses as before.

Settle before judging. A send resolves as soon as Orca owns the message, while
its dispatch is still pending; that is the ordinary successful path. Reading the
dispatch off the send result therefore reported every delivered continuation as
pending and never wrote the attribution note. The outcome is now decided on the
settled submission, through the host's existing settlement waiter, with the send
result as fallback when nothing settles in time.

The failed-note path no longer swallows its error. It stays best effort -- a
journal that refuses the note must not turn a delivered continuation into a
failure -- but the failure is reported through the host's error sink instead of
being discarded, so it cannot regress unseen again.

The surface's send is typed against the wire result rather than a hand-written
subset, which is what let a test assert a shape the host never returns. Binding
the surface to the host moves into its own file: the host was one line under the
line cap, and the bindings carry decisions that belong beside their consumer.

* feat(native-chat): show the reconnect offer the way the worktree sidebar does

The offer is a list of workspaces, so it should read like the one users already
know. Rows are now three tiers -- repo or project, then workspace, then the agent
sessions inside it -- and each agent carries a checkbox rather than its own
button, checked by default, with the footer acting on whatever is ticked.

Reused rather than rebuilt. The host chip is the sidebar's own: its markup lived
inline in the card's meta row, so it moves to a shared component both surfaces
render, and the label comes from getHostContextLabel, which is where "Local Mac"
has always come from. The repo glyph is RepoIconGlyph; a group with no repo uses
the FolderTree the sidebar's own project-group metadata uses. The agent row
reuses AgentIcon, the agent-type label helpers, formatShortTimeAgo and the same
model treatment.

Two things could NOT be reused, and both are deliberate. The sidebar's
CompactAgentRow needs a live pane, tab and status entry, and every chat here is
stopped by definition. And the sidebar has no git-worktree-vs-folder glyph
resolver at all -- both kinds render the same card, and the difference people
read is its status lane choosing GitBranch when a workspace has branch identity;
that single precedent is what the workspace glyph follows.

The model, the execution host and the workspace kind now travel with each
offered chat. All three are read off the durable record the predicate already
holds -- the model through the same normalizer the status feed uses -- so the
glyph is never inferred from a name and no new data source appears. They are
optional on the wire, so an older host still renders a row.

Selection changes which ELIGIBLE chats are acted on, never what is eligible. Ids
are seeded from the host's own answer and intersected back against it before any
call, and the host re-derives the predicate regardless of what it is sent.
Continuing still requires an explicit click, and the automatic path still calls
the reconnect method, which contains no send.

The badge's treatment becomes a variant instead of a pile of overrides, which is
what the design-system gate asks for once the markup is somewhere it can see it.

* fix(native-chat): title a folder workspace group with its project name

A folder workspace's synthetic worktree borrows the `repoId` slot to name the
project group it belongs to, so that field is NEVER null. The reconnect offer
read a non-null `repoId` as proof of a git repo, looked it up in the repos list,
found nothing, and rendered the raw `folder-workspace:<uuid>` string as the group
header. The project glyph written for the no-repo case was unreachable for the
one workspace kind it was meant for, and the string fallback behind it was dead
for the same reason.

The project group name was available all along and the sidebar already titles
these with it, which is what this list is meant to mirror.

Recognising the id now lives beside the code that mints it, so the two cannot
drift: there was no such helper, only forward constructions of the same prefix in
five places. The header choice itself moved into a pure resolver, so the branch
that was wrong is now the branch under test.

The dead fallback string is gone, along with its catalog entries.

* fix(native-chat): offer an accepted send the provider never opened a turn for

QA: a chat that was genuinely working was silently dropped from the offer. The
discriminator was how far the send had progressed -- it was the last chat
prompted before quitting, reachable by quitting a second or two after sending.

Mechanism, reproduced against the predicate. The marker was written while the
send was still pending, so it is submission-shaped. During teardown the dispatch
then settled to `accepted`, which took it out of the pending/unknown branch and
into the follow-forward branch. But the provider died before writing a turn row
for that send, so there was no turn to follow forward TO, and the branch demanded
a proved link before it would answer. Both the no-turn-at-all case and the
newest-turn-belongs-to-an-earlier-exchange case therefore refused.

An accepted send that never became a turn cannot be finished work, because
finishing writes a turn row. The marked send is also the newest work in the
session, so any turn it opened would be the newest turn.

That makes the link unnecessary to prove for a safe answer. When the newest turn
is interrupted or unverifiable the two readings agree: if the row really is this
send's under a key we failed to match, it was cut off; if it belongs to an
earlier exchange, this send opened no turn at all. Either way the work was
interrupted. A journal with no turn row at all is the same case with nothing to
disagree about.

The readings only diverge on a `completed` row, where an unmatched one might be
this very send's finished turn under a key we did not recognise. That stays
refused. Ambiguity resolves to no, because resuming finished work is the one
outcome never worth risking.

* fix: write the grouping separators as escapes so the files stay text

Five separators in the reconnect-offer redesign were written as raw NUL bytes
instead of the `\0` escape. The runtime strings were correct and the app behaved,
but git classifies a file containing a NUL as binary -- so the two central files
of that redesign rendered as "Binary file not shown" in review, and `rg` skipped
them silently, returning no matches rather than an error.

The escape produces the identical string, so the NUL separator is kept: the
previous separator was a space, and a workspace id containing one would corrupt
the join/split pair this grouping depends on.

Nothing could have caught this. Typecheck, lint, the quality gate, the
localization verifiers and the full suite all passed throughout, because none of
them look at file encoding. So this adds a check that does, wired into the
pre-commit hook where it costs nothing and catches the next one at the moment it
is written.

Two files already on main carry a raw NUL for the same reason -- one a template
separator, one a deliberately tricky test alphabet whose neighbours are all
written as escapes. They are grandfathered rather than fixed here, since they
belong to their own change, and the gate fails if the list ever grows or goes
stale.

* fix: parse markers into a domain type, and declare the four restart methods

Two CI failures, both ours.

Static analysis. `Reflect.get` was adopted to clear the casting audit, and the
anti-slop rule forbids it -- the two gates disagree, and the rule text says what
both want: parse dynamic input into a named type once, then read typed fields off
it. Markers re-enter from a file this process may not have written and decide
whether an agent is handed a provider child, so they now go through a single zod
parse. Unknown keys still pass, and a malformed marker is still dropped rather
than thrown, so a bad entry cannot make a user's sessions unreadable. The launch
stamp is parsed the same way, the resume-admission refusal becomes a named error
carrying a typed `owner` instead of a bag assigned onto `new Error`, and the test
harness gets a named journal type instead of reaching into `unknown`.

Cross-version wire. The four restart methods are added to the manifest rather
than the count being bumped, so the suite now exercises them in both skews. They
are bare additions, not capability-negotiated: an unknown RPC method answers
`method_not_found`, which is explicit and visible during negotiation, unlike a
stream opcode that is dropped in silence. The whole `agentSession.*` surface
already sits behind its runtime capability, so an old client is told it does not
exist and never reaches a host method.

The stub's spies stay a flat map because callers iterate it asserting each entry
is a spy that did not run; a composer reassembles the member the host really
exposes. The manifest and its params builders move to their own module, which is
what keeps the suite under its line cap as the surface grows.

* Prevent duplicate restart continuation and release reconnect holds

* fix: recheck interrupted work when admitting restart continuation

* fix(native-chat): invalidate restart offers after newer user work

* Consume restart recovery offers from an isolated advisory capsule

* Refuse completed restart work and report recovery outcomes

* fix(native-chat): honor queued completion and uncertain restart delivery

* fix(native-chat): preserve restart refusal and teardown evidence

* fix(native-chat): rederive recovery evidence before continuing

* Validate restart continuation at provider dispatch

* fix(native-chat): finish restart refusal and attribution delivery

* fix(native-chat): keep recovery teardown errors out of logs

* fix(native-chat): validate restart continuation at provider dispatch

* Revalidate restart continuation when Claude dequeues input

Check continuation authority after the SDK input queue wait and arm replay correlation only after authorization. Preserve typed pre-dispatch refusal, ordinary send behavior, and cleanup when the provider exits or capacity fills during authorization.

* Deduplicate settlement test import

* Keep merge update scoped to restart recovery

* Polish continuation popover spacing
2026-09-17 12:59:03 -07:00
Brennan Benson 7f5141ae2d Make the Agent Permissions toggle apply to Codex chat (#20977)
* fix(structured-chat): deliver the permission posture through each transport's own contract

Codex posture moves off app-server argv onto typed `thread/start` and
`thread/resume` params. Manual states `on-request` / `workspace-write`
explicitly instead of omitting the fields, which app-server resolved through the
mirrored config.toml — a Manual thread on a home carrying
`approval_policy = "never"` never prompted.

Claude keeps its owned `--dangerously-skip-permissions` flag through SDK
`extraArgs`; the SDK's typed bypass option emits a newer allow flag that older
user-installed binaries reject.

Posture is re-derived from current settings on every session acquisition.

* fix(structured-chat): parse permission arguments as argv

* fix(structured-chat): keep permission policy authoritative
2026-09-16 18:21:52 -07:00
Brennan Benson c702e77bc7 Stop reading the terminal arguments field on the structured chat route (#20944)
* fix(native-chat): stop reading the terminal arguments field on the structured chat route

Setting Claude's Arguments to "--dangerously-skip-permissions --model Opus" made
every new Claude tab open in the old terminal-backed chat instead of the new
structured one, with nothing on screen to explain why. Removing "--model Opus"
fixed it.

The cause was a whole-string comparison: the configured arguments were checked
against a single blessed value per agent, so any added token at all — including
one the agent supports — stopped the string matching and the launch was demoted.

Structured chat does not run the interactive CLI. It drives Claude through the
Agent SDK and Codex through app-server, and those take narrower option sets that
are versioned separately from the CLI's, so one free-text field cannot have a
guaranteed meaning for all three. The structured route now reads only what it can
actually honour: a replaced launch command, or a launch that names its own working
directory. Terminal launches still apply the field exactly as before.

Permission posture no longer travels as a raw flag. It is derived from the
resolved launch arguments, which is the same fact a terminal launch acts on and
which falls back to the default Orca ships when the field was never touched, so
bypass stays on by default and Manual is still honoured. Claude gets the SDK's
typed permissionMode and allowDangerouslySkipPermissions at query start; Codex
gets its bypass flag placed before the app-server subcommand. Both are re-derived
per acquisition beside the auth policy and environment overlay rather than stored
in the session record, so nothing can disagree with the setting.

Codex also loses the --profile, --add-dir and -c passthrough that reached
app-server through that field. Only the permission posture comes back.

* test(native-chat): pin routing authority on the narrowed feasibility input

The routing-authority pin still named the old bundled blocker and built its
"customized" fixture out of the arguments field, which is no longer a feasibility
input. Both are now the launch command, and arguments and environment are
customized on both passes of the loop, so the flag handed to the shared resolver
tracks the command alone — a caller that resumed reading either one fails here.

No case is dropped and no assertion is relaxed: the blocker list is still
exhaustive and every caller must still honour a refusal from the shared resolver.
2026-09-15 23:38:04 -07:00
Brennan BensonandMerge Sim f55b7ba680 fix(native-chat): cancel pending prompts precisely (#20601)
* fix(native-chat): hide activity while awaiting input

* fix(native-chat): keep approval turns cancellable

* test(native-chat): satisfy split PR quality gate

* fix(native-chat): catalog approval cancellation label

* fix(native-chat): include approval cancellation runtime label

* fix(codex): settle prompts when cancelled turns complete

* fix(codex): settle prompt registry fallbacks

* test(native-chat): cover pending interaction fallbacks

* test(native-chat): split prompt state coverage

* test(native-chat): keep prompt state isolated

* fix(native-chat): bound prompt turn backfill

* refactor(codex): centralize prompt registry bounds

* fix(native-chat): cancel pending prompts precisely

* fix(native-chat): consolidate capability imports

* fix(native-chat): harden precise prompt cancellation

* fix claude cancellation teardown races

* retry claude prompt lifecycle admission

* bound claude prompt cancellation retry work

* fix(codex): bound prompt turn identity on registration

* fix(native-chat): route rejected late dispatch settlements

* fix(codex): retain exact cancellable prompt turn ids

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-14 14:59:03 -07:00
Brennan BensonandMerge Sim 955051ded0 fix(codex): settle a structured send on admission, and stop minting a colliding identity (#20138)
* fix(codex): settle a structured send on admission, and stop minting a colliding identity

Two sends could be written into the journal under one durable identity.

Codex coalesces a mid-turn `turn/start` into the running turn rather than
refusing it -- measured against real `codex app-server` builds 0.147.0,
0.150.1 and 0.153.4, none of which refuse and none of which fire a second
`turn/started`. The dispatch path read the turn id from the turn/start
response and stamped every accepted send `ordinal: 0`. Since a coalesced
send gets the running turn's id back, two submissions persisted the same
`providerItemId`. That string is durable, and it is the key a restore uses
to match a submission against provider history, so the second message's real
history row matched nothing and rendered as an extra bubble on replay.

On 0.147.0 it is worse than a collision: the coalesced response returns a
turn id that never starts and never completes, so the persisted key named a
turn absent from history and NEITHER message could match.

Identity is now minted from the echoed user message at `identityFor` -- the
single point that mints the journal row's own identity -- so the settled key
is by construction the one replay computes, rather than a parallel
calculation that can drift.

Dispatch returns `admitted` when the transport write completes; identity
settles on the echo through a channel that did not previously exist for
Codex. Waiters are keyed by client message id instead of being shifted off
the front of an array by arrival order, and they are cleared on session
close and child exit -- previously a timeout was the only thing that ever
ended one.

`TURN_ID_WAIT_MS` is deleted. It was never reachable on any build measured:
`readCodexTurnId` returns non-null on all three, so the 10s wait never
fired. The comment justifying it claimed older builds acknowledge before the
id exists, which no tested build does.

Three comments asserting Codex answers a mid-turn send with `turn already
running` are corrected. Their only backing was a test fixture inventing that
error string. The correction is factual only -- every changed line in
`src/main/runtime/orchestration/` is a comment, and mid-turn delivery is
still refused for both providers. Whether that policy is right is a separate
question; it was resting on a false premise.

Known gap, stated rather than implied: this prevents new collisions and does
not repair journals already written with a colliding or phantom key. Those
conversations keep duplicating on restore. Repairing them means re-matching
persisted submissions against provider history and rewriting
`providerItemId` -- which is what `journal-submission-reconciler.ts` is
written for, and it still has no production caller.

* test(codex): drop the synchronous-accept contract and the colliding `:0` from the integration fakes

Three tests in the structured-session integration suites encoded the dispatch
contract this branch replaces, and two of them pinned the defect it fixes.

They asserted `agentSession.send` answers `dispatchState: 'accepted'` carrying
`providerItemId: codex:<thread>:<turn>:0` at send time. That ordinal was never
observed; it was stamped on every accepted send, which is exactly the collision
this branch removes -- a send coalesced into a running turn is answered with the
running turn's id, so two submissions persisted one durable key.

The visible failure was a 30s timeout rather than a failed assertion. The fake
client advertised no `agent-session.pending-send-result.v1`, and without it the
host holds the reply until the send settles: a shim for clients too old to
render a pending bubble. The fake provider then echoed the user message with no
`clientId`, so nothing could correlate that echo back to the submission, and the
wait ran to its own 30s ceiling. Real Codex sends `clientId` on that echo, and
the fake now does too, which is what makes it a model of the provider rather
than a sketch of one.

The identity assertion is kept rather than dropped. Each send now asserts
`pending` with no identity at admission, then asserts the submission settles
`accepted` at `codex:<thread>:<turn>:0` once the echo lands. Same ordinal, but
earned from `identityFor` on the echo -- the key a replay recomputes -- instead
of guessed from the turn/start response. Ablated: removing `clientId` from the
two echoes leaves both submissions `pending` and fails both assertions, so the
assertion is load-bearing and not satisfied by something incidental.

Both suites' client fixtures now advertise the capability set the desktop
renderer sends in `src/main/ipc/runtime.ts`, which is what these suites mean by
a client. The older-client settlement wait keeps its own coverage in
`src/main/runtime/rpc/methods/structured-agent-session.test.ts`.

`structured-agent-session-runtime-exit.test.ts` asserts `pending` for the same
reason; it drives the host directly, so it never took the compatibility path,
and what proves delivery there is still the turn the reacquired provider starts.

The replay suite's "without dispatching it twice" property is untouched: one
`turn/start` call, one replayed ledger row.

* fix(codex): preserve unsettled dispatch correlations

* test(codex): type the dispatch fixtures instead of asserting over them

main's new casting gate (#20367 base) flags type assertions on changed
lines. Replace them with checked types: the recording sink already
satisfies its interface, both CodexSession fixtures are now annotated and
carry real collaborators, the settlement assertion compares whole
identities, and the integration helper reads submissions through the
host's public journalSnapshot instead of its private session map.

* fix(test): merge the duplicate doubt-reasons import the merge left behind

Both sides added an import from journal-dispatch-doubt-reasons and the
merge kept both statements, which the whole-repo native plugin gate
refuses under --deny-warnings.

* test(codex): a Fast mode turn is admitted, not accepted

#20506 landed its Fast mode tests against the dispatch contract this
branch replaces: a Codex send now returns admitted and settles its
identity on the provider echo. The tier assertions the test exists for
are untouched.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-14 13:37:13 -07:00
Brennan Benson 539d4d1f32 fix(native-chat): resume structured chats cleanly after restart (#20509)
* fix(native-chat): retire provider ownership on restart

* fix(native-chat): stop showing a restart eviction as a provider death

Restarting Orca turned a resumable structured chat into a user-visible
`Provider exited: recorded pid absent on host`. Quit never released the durable
lease, so restart probed the recorded pid, adjudicated the session evicted, and
wrote a synthetic status row against a chat that was perfectly resumable.

The fix is the missing teardown phase plus the missing fence check: quit now
evicts every provider child this host owns — stopping it, settling its journal
and handing the lease back — and the release compare-and-swaps on the fence it
expected. Restart then finds a released lease and reopens the chat silently.

What the user sees is decided by the typed death evidence rather than the shape
of a settlement id: only an `exit-observed` death writes copy, and that copy now
carries its cause so an auth failure and an OOM kill do not read alike. The
reassuring wording stays. Historical synthetic rows are filtered out of the
render projection, which needs no schema change and leaves every real
provider-exit row alone.

Also:
- Bound the new eviction phase well below the quit deadline; a quit that dies
  mid-eviction leaves the lease unreleased, which is the original bug.
- Scope the interruption verdict to work that was mid-response. A provider that
  died while waiting on an approval interrupted nothing.
- Keep host bookkeeping in step with the adapter: the provider-child flag clears
  when the child is proven stopped, not seven steps later.
- Drop the router's duplicate shutdown gate and acquisition drain — both
  adapters already own theirs — and latch the router closed so a late acquire
  cannot fan a session back out to closed adapters.
- Attach the real cause to the settlement failure a quit reports, and remove a
  recovery-ticket field that was hardcoded at its only construction site.

* fix(native-chat): scope the legacy status filter to the copy it retires

The read-time filter hid every status row carrying a `restart-eviction:`
identity. That identity is still minted, so a genuine provider death settled
under it would have been dropped from every rendered page. Match the retired
`Provider exited` copy as well, so only the legacy rows are hidden.

Three smaller corrections alongside it:

- The settlement retry path now applies the same unfinished-work check the
  live exit path uses, so a provider that died waiting on an approval no
  longer gets told a response was in progress.
- Bound the exit reason before composing the outcome copy, so a stderr dump
  in the reason cannot push the "you can continue" sentence past the row's
  byte cap.
- Correct the teardown comment: tail rows are protected by eviction's own
  per-session ordering, and `closeAll` is a backstop for children eviction
  never took, including one whose eviction was refused.

* fix(native-chat): retire legacy status rows at the projection source

The read-time filter that hides the retired `Provider exited …` rows ran on the
way OUT of the page builder, after the paging math had already measured the
unfiltered timeline. A backward window landing entirely on those rows returned
an empty page that still reported `hasOlder: true` with a null `window.oldest`,
so the renderer's backfill loop re-asked from the same anchor forever. Its only
no-progress guard compares `window.oldest?.sequence` to the anchor, and
`undefined === n` never breaks. The live subscription opens behind that loop, so
the transcript never finished loading either.

Filter where items ENTER the page pipeline instead: the reduced snapshot gets
one renderable timeline, the forward path gets one renderable batch, and the
window bound, effective limit, `hasOlder`, `window.oldest` and `nextCursor` are
all computed over that single array. A window with nothing left behind it now
reports end-of-history.

Also restore the eviction retry contract. Clearing `hasProviderChild` as soon as
the adapter proves the child gone is honest, but it is a different fact from the
wind-down this host still owes. A retry after a step aborted between the two was
reading "no child here" and skipping both the dead-generation settlement and the
lease release the aborted attempt had promised to repeat. The obligation is now
tracked separately and cleared only by a release that actually landed.

And rename the filter to the copy it retires: it drops only rows carrying the
retired `Provider exited` text, not restart-eviction status rows in general.

* fix(native-chat): read the wind-down a close owes from the live child

An eviction recorded "nothing owed" whenever it ran over a session with no
provider child of its own, and the retry then read that record in preference
to the child in front of it. A session suspended to an agent terminal is
exactly that shape, and the trip back to native re-acquires into the SAME
session object rather than replacing it, so the next close skipped both the
dead-generation settlement and the lease release — leaving the record claiming
a live owner this host had just stopped, and a pending send unsettled.

The obligation is now derived the way the quit sweep already derived it, from
one shared predicate: a live child always owes a wind-down, and a remembered
`false` only carries the obligation forward, never cancels it.

Also drops a memoization in the history page that could never hit. Its key was
the snapshot's items array, which the reducer rebuilds on every `snapshot()`
call, so each backward page allocated a fresh key; the one reader that does
share a snapshot across pages reads forward and never calls it. The comment
claimed a multi-page read filtered once, which was not true of either path.

Tests: the handoff round trip that strands the lease, and the quit sweep
picking up an eviction whose close retry never came.

* chore(native-chat): scope three helpers to their file and pin the teardown order

retryUnexpectedExitSettlement, hasUnfinishedStructuredAgentSessionWork and
isRetiredProviderExitStatusItem each have no consumer outside the file that
defines them, so they no longer advertise an external contract.

The quit-path phase list documents its order as load-bearing, but nothing
asserted it. Pin the phase names so evict-owned-sessions cannot drift out of
its slot between drain-attaches and flush-event-sinks.

* fix(native-chat): stop the router reporting a stop it never observed

`closeAll` cleared the route table and set one boolean, after which that boolean
was the only surviving evidence about any session. Two call sites then spent it:
`releaseAcquisition` and the stop path each turned a route-lookup MISS into
reported success. Eviction reads a `true` from the stop path as proof the
provider child is gone and releases the durable lease on it, so a session the
router never routed could have its lease handed back on the strength of "I have
no record, but everything is closed."

Loss of contact is not evidence of process death. The fix keeps the evidence
instead of the inference: adapter shutdown only resolves once every child is
proven stopped, so `closeAll` now marks each routed session `stopped` rather
than forgetting it. A routed session still answers `true` from its own retained
proof; a session with no route answers `false`, which leaves it indexed for a
real retry. `releaseAcquisition` drops its short-circuit and asks the adapters,
which answer from their own session maps.

The acquire-side latch is unchanged: once closed, the router stays closed and
refuses new work.

Behaviour that changed: a post-`closeAll` stop for a session the router never
routed, or one the host already acknowledged as released, now reports unproven
instead of proven. That matches what the same call already answered before
`closeAll`, and no real flow reaches it — quit evicts every owned session before
`closeAll` runs, and eviction only asks the adapter for sessions whose provider
child this host acquired through the router.

* test(native-chat): ratchet the retired provider-exit copy out of production

The retirement filter hides a status row on two facts: a restart-eviction item
id and copy that opens with the retired prefix. The identity half is still
minted today, so the filter cannot tell a new producer's row from the legacy row
it exists to hide — any future writer of that copy would be dropped from every
transcript with no trace. Until now that safety property lived only in a doc
comment.

Scan the shipped tree for a string literal that OPENS with the retired prefix,
which is exactly what the filter's `startsWith` reads. Comments are stripped
first, so prose about the retirement is not a producer, and the filter's own
constant is exempt. Tests are excluded: writing the copy is how the filter is
exercised.

* revert(native-chat): drop the read-time retired provider-exit filter

Fix forward instead. The lifecycle change in this branch stops any new
`Provider exited: <reason>` row from being written; rows a previous build
already persisted stay in those transcripts and age out with them. A
permanent read-time filter for a cosmetic, shrinking set was not worth its
maintenance cost, and its paging seam was the only place a backward window
could land entirely on hidden rows.

Removes the filter module and its test, restores agent-session-history-page.ts
to its pre-branch form, and drops the tests that only existed to prove the
filter did not over-match or wedge the backfill loop.

The copy ratchet stays and now carries the whole guarantee: with no filter in
front of it, any production writer that resurrects the retired prefix reaches
the user's transcript directly.
2026-09-14 00:23:20 -07:00
Brennan BensonandMerge Sim ebb1acfa37 refactor(agent-status): publish structured sessions into the hook server store (#19683)
* refactor(agent-status): publish structured sessions into the hook server store

Structured (native chat) sessions have no PTY and no hook script, so their
status never reached the hook server's store; #19217 gave `worktree ps` its
own adapter over the structured feed instead. The feed now writes every
projection into that store through a status sink the runtime wires, drops
the row when the host closes the session, and `worktree ps` reads the one
snapshot like every other agent.

Rows carry a `structuredHost` marker and the journal clock; they are never
persisted to last-status.json, and the main process does not forward them
to the renderer yet, whose feed bridge still owns them until it is retired.

Design and the two follow-ups: docs/reference/agent-status-store.md.

* chore: drop stray @pnpm/exe lockfile entry

An unrelated local pnpm run added @pnpm/exe as a packageManagerDependency
with no package.json change, so CI's --frozen-lockfile install failed
before any job ran.

* docs(agent-status): describe the step that actually landed

The design record claimed PR 1 deletes RuntimeAgentRowStore, drops the
retained-versus-hook reconciliation, stamps terminalHandle on OSC rows, and
tags rows with a source field of 'structured-host'. None of that is true of
the shipped code: the retained store and its reconciliation are still in
place, and the row field is structuredHost: 'held' | 'owned'.

AGENTS.md points every future contributor here before they touch agent
status, so split the roadmap into the 1a that landed and the 1b that has not,
and name the fields the code actually writes.

* fix(agent-status): pair session removal with the status-row forget

A session dropped from the host's map without an explicit forget left its row
in the store forever: `structuredHostOwned` bypasses the staleness check, so a
failed re-attach (the Claude rewind path reaches one) stranded a permanently
working agent in `worktree ps` and on mobile with no UI able to clear it.
Deletion and forget are now one operation both callers route through.

* fix(agent-status): give orcad the store worktree ps reads from

`orcad` constructed its runtime with neither `getAgentStatusSnapshot` nor
`structuredAgentStatusSink`, so once `worktree ps` sourced rows only from that
snapshot the headless host published nowhere and listed nothing. The hook
server's store is a module singleton whose import tree never reaches Electron,
and its file paths come from `start()`, which orcad never calls.

* fix(agent-status): drop a structured row without a renderer clear

`dropStructuredStatus` went through `clearPaneState`, which fans a pane clear
out to the renderer for a pane key the renderer's own feed bridge still writes
- so 'exactly one writer per pane key' held for writes and not for deletes.
`dropStatusEntry` routes through the status-drop tap instead, and skips the
resume-identity remnant: a structured session has no pane to resume into, and
every null-status publish would otherwise re-mint one.

* test(agent-status): pin both half-migration structured-row filters

Neither the `agentStatus:getSnapshot` filter nor the main-window listener's had
a single assertion, so deleting either — the first step of PR 2 — was green
everywhere. Also covers the perf skip and the drop's lack of a renderer clear.

* docs(agent-status): correct three statements this PR made false

The sink JSDoc claimed only tests construct a host without one; `orcad` did.
The doc argued a structured row needs no tab mirror 'because headless serve has
no renderer', reasoning about exactly the topology the wiring had not reached.
The deleted runtime adapter's warning that the pane key must be the DERIVED one
- never a bearer handle or minted worker key - was lost with it.

* test(agent-status): declare orcad in the hook-row producer census

Wiring the hook store into the orcad runtime added a production site that
hands hook rows to a consumer, which the census ratchet pins deliberately.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-10 11:27:27 -07:00
Brennan BensonandMerge Sim 5868fdc9e3 feat(native-chat): report Codex background tasks in the chat strip (#19346)
* feat(native-chat): report Codex background tasks in the chat strip

The background-tasks strip works for Claude only; a structured Codex
session shows nothing in it. Feed it from the Codex app-server stream.

The strip stands for work that OUTLIVED a turn, which is what the
monitoring header, Claude's foreground suppression, and the conversation
command gate all already assume. Codex has no `is_backgrounded` flag, so
that fact is derived from the turn boundary: a `subAgentActivity` child or
a primary-thread `commandExecution` becomes visible once the turn it
belongs to completes and it is still unsettled.

`turn/completed` only reveals a task here, never settles one — measured on
`codex app-server` 0.153.4, a spawn_agent child reported `completed` 95.8s
after its parent turn ended. Only a child's own activity kind settles it.

Codex exposes no honest stop: `turn/interrupt` on a child ends its turn
without emitting a terminal activity item and leaves its shell running. So
the state carries a new optional `supportsStopAll: false`, the strip hides
a control that could not act, and the blocked-command message asks the user
to wait rather than to press a button that does not exist.

* refactor(codex): move session teardown out of the structured adapter

Merging main crossed the 300-line cap on
`codex-structured-session-adapter.ts`: the rewind backend (#19235) and this
branch's close-time strip clear both landed in it. The four close paths move
verbatim into `codex-structured-session-teardown.ts`, where they funnel
through one `settled` helper instead of repeating the notification-retry and
background-task cleanup at each call site. No ratchet bump.

Also normalize a background task's description once at receipt rather than on
every projection; the roster is re-projected on each observed frame.

* fix(codex): drop the shell row the journal already settles

A `commandExecution` still `inProgress` when its turn ends was reported as a
`command` task. But `settleCodexJournalTurn` writes exactly those items to the
journal as `state: 'failed'` on `turn/completed` and forgets them, so the strip
row would have claimed a shell was still running at the same instant Orca
recorded that it was not — two surfaces contradicting each other about the same
process.

A subagent is the opposite case and stays: the roster pointedly does not sweep
at a turn boundary, because children measurably outlive it. That leaves the
producer making exactly one claim — these spawn_agent children are still live
after their turn — which the durable roster row corroborates.

* fix(native-chat): track Codex background execution lifetimes

* fix(native-chat): keep running tool groups from claiming completion

* Fix runtime catalog and capability expectation

* fix(codex): keep a child's name on the command row that outlives it

A child agent's commands stay hidden behind its agent row while the child
works. Once the child's turn settles with a command still running, that
command surfaces as its own row labelled from the raw command string, so
'long_probe' became "/bin/zsh -lc 'ping -c 300 127.0.0.1 > /dev/null'"
at the moment that row was the only remaining signal for the work.

Qualify a child's command row with the child's label. Resolved on read,
so a label registered after the command still lands, and bounded by the
existing description cap so admission accounting stays valid. Primary-
thread commands are left unqualified: they have no child to name.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-09 00:13:20 -07:00
Brennan BensonandMerge Sim b7b6ea3942 fix(native-chat): auto-rename the workspace on a structured chat's first turn (#19138)
* fix(native-chat): auto-rename the workspace on a structured chat's first turn

Structured native chat (Claude and Codex) never reached the first-work
workspace rename. The orchestrator has a single production caller, the
agent-hook server listener, and structured sessions never set
ORCA_PANE_KEY, so no hook event could ever be attributed to one. The
renderer knew this and suppressed pendingFirstAgentMessageRename for
structured launches at three sites, which also closed the gate the
folder-workspace title rename depends on.

The host's status feed already computes the exact edge: status 'working'
with a latestPrompt normalized the same way the hook payload is, and a
workspaceId that IS the worktree id. Publish that projection to the host,
thread it out to the runtime, and hand it to the same orchestrator the
hook path uses.

Re-projections of state the host already knew (restore, an arriving
subscriber) are flagged as replays and map to the orchestrator's existing
isReplay gate, so a host restart cannot rename off a stale journal.

One host and one journal serve both providers, so this covers Claude and
Codex together.

Verified in a live Electron instance, worktrees created through the real
composer and prompts sent through the real chat composer:
  Codex  langouste -> retry-helper-exponential-backoff
  Claude prowfish  -> parse-csv-headers

* fix(native-chat): preserve first-work rename across runtime and queued turns

* fix(native-chat): skip branch rename for folder projects

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-06 17:38:06 -07:00
Brennan BensonandMerge Sim ad4dc353f3 fix(native-chat): settle a structured send the provider proves it received after the ack window (#19140)
* fix(native-chat): settle a structured send the provider proves it received after the ack window

A send waits a bounded window for the provider to echo the message it was given.
On timeout the dispatch resolves `unknown`. The echo that arrives later IS matched
— `recoverLateIdentity` uses it to repair the session's turn identity — but nothing
tells the journal, and `unknown` is terminal there. The submission stays unknown for
the life of the session.

Two consequences, both reachable on any ordinary session:

- The composer renders "Message delivery is unconfirmed." with a Retry, forever,
  for a message that was delivered and answered.
- Retry redispatches, because the host only replays a recorded outcome unless
  `retryUnknown` is set, which that button is the only thing that sets. So the
  banner is a duplicate delivery armed and waiting for a click — and a user who
  believes the banner and resends is doing exactly that by hand.

Every send made while a turn is already running takes this path: the provider does
not echo a queued message until the running turn ends, which is far past the 10s
ack window. Sends made while idle are unaffected, which is why this reads as
intermittent.

Carry the `clientMessageId` on the dispatch waiter and settle the journal
submission `accepted` when the late echo proves delivery. Deliberately unfenced
against the dispatch sequence: that fence decides which turn owns the identity,
while delivery is settled either way. Already-terminal rows are untouched.

* fix(native-chat): persist late dispatch receipts before session close

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-06 16:41:09 -07:00
Neil af82126058 fix(native-chat): give the Claude exit barrier a handle on unpublished exits (#18826)
A first-hand Claude exit is not published where it is observed. `handleExit`
re-enters the close ladder and persists the transcript cursor before it emits
`ended`, and only that emission reaches the runtime's recovery chain. So the
runtime's `waitForRecovery` — whose whole job is to drain an in-flight recovery
before teardown stops children — returns immediately for an exit that is still
climbing the ladder, and nothing outside the adapter can tell an observed exit
from a published one.

The integration test for fenced host reconciliation had no handle on that
barrier, so it bounded-polled the lease for 100ms instead. Measured under 16x
local concurrency, publication alone takes 77-204ms: 19/24 runs failed.

Retain the ladder-then-settle tail on the exit record and expose
`drainObservedExits`, fold it into `waitForRecovery`, and export the barrier so
a caller that needs the settled lease can await it. Codex publishes inside its
own exit callback and needs nothing. The test now awaits the barrier: 0/24
under the same load, and it fails on an idle machine without the drain.
2026-09-05 03:50:19 -07:00
Brennan BensonandMerge Sim e89deb63c9 Show Claude background task status in Native Chat (#18757)
* feat(native-chat): show Claude background task status

* fix(native-chat): carry background task fence forward

* fix(claude): bound background task stop requests

* Show running Claude background task details

* Harden Claude background task status updates

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-05 00:21:42 -07:00
a65332a8bd feat(claude): move structured native chat onto the Claude Agent SDK and enable it on macOS and Linux (#18560)
* Join structured attach teardown through journal bind

* fix: restore structured chat parity

* feat: add Claude structured session adapter

* fix: harden Claude structured adapter

* fix: close Claude adapter edge cases

* fix: start Claude init deadline after launch

* feat: wire Claude structured sessions

* fix: harden Claude structured runtime

* fix: fence Claude structured compatibility

* fix: preserve Claude free-text prompt answers

* fix: decode addressed Claude prompt text

* feat: enable Claude structured chat on mobile

* fix(mobile): keep structured chat provider-aware

* fix(mobile): negotiate Claude structured tabs

* fix: keep scoped RPC tests native-free

* fix: secure mobile structured image delivery

* fix: close structured session data-loss gaps

* fix: prove real Claude structured startup

* fix: consume pre-spawn proof before retry

* feat(native-chat): add desktop structured sessions

* fix(native-chat): satisfy structured session cleanup gates

* fix(native-chat): keep structured renders pure

* fix(native-chat): open composer pickers upward

* fix(native-chat): use existing view for structured sessions

* fix: harden structured desktop status projection

* fix: close structured desktop lifecycle gaps

* fix: fence structured AI Vault resumes

* fix: fence structured AI Vault resumes

* fix: preserve structured tabs during activation

* feat: toggle structured sessions between chat and TUI

* fix: harden structured session handoffs

* fix: bind structured TUI before rollout proof

* fix: complete structured chat round trips

* fix: align structured TUI return readiness

* fix(native-chat): make reverse handoff transactional

* Add Claude structured TUI handoff seams

* fix(native-chat): clear sticky handoff recovery

* fix(native-chat): complete mobile reverse after TUI exit

* fix(native-chat): keep TUI transcripts readable

* fix(native-chat): recover TUI transcript gaps

* fix(native-chat): recover claimed TUI owners

* fix(native-chat): retain cold TUI proof authority

* fix(native-chat): preserve Claude handoff authority

* fix(native-chat): recover TUI transcripts read-only

* fix(native-chat): harden Claude handoff recovery

* fix(native-chat): serialize structured handoff recovery

* fix(native-chat): close handoff admission races

* fix(native-chat): validate pinned launch environment

* fix(native-chat): revalidate restored and retried owners

* fix(native-chat): gate restart recovery publications

* fix(i18n): catalog Claude session controls

* fix(native-chat): wait for structured TUI process proof

* fix(native-chat): queue stale idle TUI handoffs

* fix(native-chat): route structured Codex options directly

* fix(native-chat): persist structured session options

* fix(native-chat): hydrate resumed structured options

* fix(native-chat): preserve options across structured handoffs

* fix(native-chat): replay pending option mutations

* fix(native-chat): rotate settled handoff operations

* fix(native-chat): rotate refused send operations

* test(native-chat): derive refusal retry state from host

* test(native-chat): give the host-oracle matrix test an explicit timeout

* fix(native-chat): keep Claude option controls idle

* fix mobile structured first-send hydration race

* fix(native-chat): preserve handoff launch authority

* fix(native-chat): harden shared handoff recovery

* fix(native-chat): serialize structured handoff recovery

* fix(native-chat): close handoff admission races

* fix(native-chat): validate pinned launch environment

* fix(native-chat): revalidate restored and retried owners

* fix(native-chat): gate restart recovery publications

* fix(i18n): catalog structured session recovery control

* fix(native-chat): wait for structured TUI process proof

* fix(native-chat): queue stale idle TUI handoffs

* fix(native-chat): keep structured recovery provider-neutral

* fix(native-chat): drop local terminal topology from structured sync

* fix structured outbox and tab restore races

* fix(native-chat): preserve Claude question groups

* fix structured provider visibility and request handling

* fix structured session TUI handoff recovery

* fix reverse structured session handoff

* fix(native-chat): recover Claude outbox and resume state

* chore(mobile): preserve the working-tree lockfile state before the main merge

Carries the pre-existing uncommitted mobile/pnpm-lock.yaml modification into history so the
main merge cannot overwrite it. Verified benign pnpm drift (babel 7.29.7->7.29.8 transitives
plus deprecation metadata); drops no patchedDependencies (the mobile lockfile declares none).

* test(native-chat): drop orphaned Claude handoff-auth test left by the main merge

'pins Claude handoff auth through the terminal provider boundary' is absent from main and its
production counterpart preserveClaudeAuthEnv no longer exists outside this test - orphaned residue
of the terminal/native handoff work this PR excludes by scope.

Removed rather than repaired: the failure was a renamed field (providerHome -> providerRoot), and
renaming it would have carried out-of-scope handoff code into the merge. Body preserved as evidence
and logged in CLAUDE-STRUCTURED-DISPOSITION-TABLE.md.

* Fix mobile structured turn state

* fix Claude structured session blockers

* fix claude structured lane blockers

* fix Claude acquisition exit proof

* fix(claude): route stream-json launch through process wrapper

* fix(claude): gate structured chat support

* Fix Claude structured launch gating

* fix(claude): split session acquisition and prune mobile scope

* test(claude): align structured session fixtures

* fix(agent-session): preserve handoff launch arguments

* fix(claude): open journals through the factory after origin/main split

The journal opener moved to journal-store-factory on main; retarget the
Claude structured tests that still imported the old path.

* fix(claude): resolve Claude structured launch args, auth, and win32 proof

The origin/main merge re-expressed the lane's Claude wiring onto main's split
orca-runtime facade and dropped three wires past green typecheck and lint.

- resolveLaunchArgs discarded its provider parameter, so structured Claude
  sessions were launched with Codex app-server flags; Claude exits on
  --dangerously-bypass-approvals-and-sandbox, and a Codex arg-parse throw
  could block Claude session creation outright.
- resolveClaudeLaunchEnv was no longer supplied, so the launch resolver fell
  back to the whole process env as configuredEnv and
  buildClaudeChildProcessEnv re-applied every auth var it had just stripped.
  The resolver now merges the Claude overlay onto a strip-applied copy of the
  inherited env, which also keeps PATH intact for withCliRuntimeOnPath.
- The windowsProcessStartTimeAvailable producer was gone while the contract
  field and both consumers survived, so the renderer gate fail-closed and
  structured native chat was unreachable on every win32 host.

Separately, structured Claude pinned CLAUDE_CONFIG_DIR unconditionally. An
explicit pin makes the CLI abandon the macOS Keychain even when it names the
CLI's own default, so a default claude.ai account could not authenticate where
the legacy Claude terminal could. Pin only a home the CLI would not resolve on
its own, matching ClaudeRuntimePathResolver, and compare against the env the
child would otherwise inherit so a diverging overlay cannot outrank the
record's account home.

Also await the now-async revealNativeSession in its regression test, and set
the native status before revealing so a rejecting reveal cannot leave a
session released but never marked native.

Claude-Session: https://claude.ai/code/session_013UqKCRB6k5e8UaYhXUHeWY

* fix(claude): scrub case-insensitive Windows auth env

* fix(native-chat): settle handoff outcome-write failures instead of leaking them

A store write failure while recording a handoff outcome escaped the flow
runner's catch handler, so the client never received the failure and the
flow surfaced as an unhandled rejection (seen as an intermittent
agent_session_store_corrupt error in the proven-dead-retry suite, whose
teardown raced the flow's trailing outcome write). Record the failed
outcome best-effort, and drain the coordinator before that test's
teardown removes the store root.

Claude-Session: https://claude.ai/code/session_011aXkcHyeiRJuezupQdjZaM

* fix(native-chat): make the structured close-failure toast provider-neutral

The structuredSessionCloseFailed toast fires for any structured session,
but its copy said 'Codex chat', so a Claude structured session that fails
to close showed the wrong provider name. The launch-failure toast is only
reachable behind the agent === 'codex' gate, so its copy stays as is.

Claude-Session: https://claude.ai/code/session_013ugSpCx4AWkySaJb69BQax

* fix(native-chat): wire structured handoff proof recovery

* fix(native-chat): wire structured handoff proof recovery

* fix(native-chat): correct the structured chat opt-in copy

The one `experimentalStructuredNativeChat` toggle gates both providers —
`useStructuredAgentSessionCreate` runs `canUseStructuredNativeChat` for
`'claude'` as well as `'codex'` — but its description named only Codex.

Its scope line also said Windows keeps using terminal chat, while the gate
refuses win32 only until the host proves it can read a process start time.
`structured-native-chat-availability.test.ts` already pins that Windows is
allowed once the proof is cached, so the two contradicted each other.

Claude-Session: https://claude.ai/code/session_01RJFsidQWmKYFmeoUuVu4Tp

* test(claude): pin @anthropic-ai/claude-agent-sdk 0.3.251 contracts against a scripted CLI

PR 1 of the SDK migration: dependency + test-only harness, no product wiring.

- Pin @anthropic-ai/claude-agent-sdk to exactly 0.3.251 — not the newest
  release — because 0.3.251 (published 2026-08-28) clears the repo's 3-day
  minimumReleaseAge supply-chain gate with no exclusion, while the newest
  release was minutes old and would have required excluding a brand-new
  publish from the exact control built to catch brand-new malicious
  publishes. Every contract this design depends on was verified identical
  on 0.3.251: the full option surface, no pid on SpawnedProcess (custom
  spawner stays mandatory), env defaulting to process.env when omitted, and
  --replay-user-messages appearing only via extraArgs.
- Exclude all eight bundled CLI platform binaries via
  ignoredOptionalDependencies. The setting lives in pnpm-workspace.yaml
  because pnpm 12 no longer reads the package.json "pnpm" field (it warns
  and ignores it; verified by install ablation). Excluding the binaries is
  what makes Orca's pathToClaudeCodeExecutable override mandatory rather
  than merely preferred. Note: pnpm 12.0.0 honors the ignore list when
  reconciling an existing lockfile but not on fresh resolution of a new
  dependency, so the lockfile's SDK entry was pinned surgically; both
  'pnpm install' and 'pnpm install --frozen-lockfile' verify clean and
  stable against the committed lockfile.
- Contract-pin suite drives the real SDK against a scripted fake CLI and pins:
  unknown type/field/content-block pass-through (and keep_alive interception),
  spawner env fidelity plus the omitted-env process.env inheritance sharp edge,
  extraArgs producing --replay-user-messages, argument parity for every
  CLAUDE_STRUCTURED_BASE_ARGS entry plus --session-id/--resume/
  --resume-session-at, canUseTool wire request_id stability and abort on
  control_cancel_request, one spawn per query, pathToClaudeCodeExecutable
  honored by the default spawner, the exact SDK version, and the eight platform
  binaries staying uninstalled.

Claude-Session: https://claude.ai/code/session_01FGCRfYUnb4hbvfTAHGtJKQ

* feat(claude): drive the structured transport through the agent SDK

Replaces the hand-rolled `claude -p --input-format stream-json` transport with
@anthropic-ai/claude-agent-sdk 0.3.251, keeping the existing connection
interface for this commit so the acquisition path changes minimally. The
control-plane rewrite is a separate change.

Orca still supplies the process. `spawnClaudeCodeProcess` routes through
`spawnProcess`, retains the child and its pid — the triple the durable lease
adjudicates on — drains stderr so exit errors keep their tail, and hands `.cmd`
shims to Orca's Windows argument encoder rather than the SDK's plain spawn.
`close()` keeps Orca's own bounded tree-kill and exit deadline, so it still
resolves true only after an observed exit.

Launch resolution emits an SDK options object instead of argv; durable
`launchArgs` translate to a typed option where one exists and to `extraArgs`
otherwise, refusing a token neither can carry rather than dropping it. The
child env is always passed explicitly — omitting it would let the SDK inherit
`process.env` and reintroduce the ambient `ANTHROPIC_*` leak. The stdout line
parser is deleted; the SDK owns framing, and unknown frames still reach the
translator verbatim.

Claude-Session: https://claude.ai/code/session_01JMhFjh9HEnkcJ5YTfCdgD3

* fix(claude): settle the frame the SDK pulled but never wrote

The SDK's input pump is `for await (frame of prompt) { await transport.write(frame) }`.
When that write rejects — the child dies between Orca's liveness guard and the
write — the for-await ends abruptly and calls the generator's `return()`, so the
code after `yield` never runs. The frame was already shift()ed out of `queued`,
so the later `fail()` from the exit path could not reach it and `send()` never
settled: `dispatchClaudeTurn` awaits that send before it can return `unknown`,
wedging the caller and the durable outbox. The pre-SDK transport rejected on the
stdin write callback instead.

Retain the in-flight entry and settle it from the generator's cleanup, and let
fail() reach it too for the pump that never resumes at all.

Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum

* fix(claude): keep the agent SDK behind the structured-Claude boundary

The ordinary OrcaRuntimeService graph statically reaches the Claude adapter and
so the transport module, whose first line imported @anthropic-ai/claude-agent-sdk.
The SDK is evaluated whenever the regular runtime loads, before any structured
Claude session is chosen: it sets process.env.NoDefaultCurrentDirectoryInExePath,
changing Windows executable resolution for later subprocesses, and a missing or
incompatible install would break normal runtime startup — for a user who never
leaves the terminal/TUI path.

Defer the SDK to the connection, memoized so it loads once per process, and add
the import-graph ratchet: a walk from the Electron main entry that fails on any
static import of the package, plus a clean-fork check that loading the runtime
leaves the Windows search variable untouched and a child-process pin that the
side effect is still real.

Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum

* fix(claude): answer list_models so the picker stops serving the seed

sendControlRequest had no list_models case, so every request hit the default
reject; readClaudeStructuredSessionOptions swallows that with .catch(() => null)
and falls back to the static catalog. Every structured session therefore served a
hardcoded model list with no per-model effort levels, no resolvedModel and no
default detection, and nothing surfaced the failure. The pre-SDK transport got the
live catalog from the CLI.

Route it through the SDK's supportedModels(), wrapped in the { models } envelope
the existing parser reads.

Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum

* fix(claude): reap the child's descendants before killing it

The forced step of the exit ladder went through the Codex helper, which spawns
`pkill -KILL -P <pid>` and SIGKILLs the parent in the same tick: the parent
usually dies first, the descendants reparent to pid 1, and `-P` matches nothing.
An MCP or launcher descendant of a stubborn Claude child was left running. The
test named for that requirement declined to assert it and killed the survivor by
hand instead, so it could not fail for the thing it was named after.

Route the Claude reap through Orca's existing sweep, which snapshots descendants
while their parent link still exists and signals them before the root goes, and
on Windows uses the identity-gated `taskkill /T /F`. The test now asserts the
descendant is dead; the manual kill stays only as a failure-safe. close() still
returns true only on an observed exit.

Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum

* fix(native-chat): merge the duplicated handoff type import

CI's static-analysis lint (`oxlint --config
config/oxlint-code-quality-native-plugins.json src config tests mobile
--deny-warnings`) exits 1 on the two separate `import type` statements from the
same module.

Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum

* fix(claude): answer a permission callback whose signal already aborted

settleFrom registered the abort listener and then delivered the request. A
callback that arrives already aborted never fires that event, so the promise
stayed pending behind a durable prompt with no cancel path. Check the signal
first, emit the cancel, and resolve the SDK's null sentinel without registering.

Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum

* test(claude): wait for the child to record the frame, not just for its report

The scripted CLI writes its report at startup, so `until(readReport)` returned a
report with no user messages whenever the child had not yet read the line. The
assertion then failed under parallel load. Poll for the frame instead of for the
file.

Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum

* fix(claude): coalesce partial deltas onto one assistant item and stop painting result frames

Under --include-partial-messages every stream_event frame carries its own
uuid, and the final assistant frame for a block carries yet another; only
message.id ties them. The translator keyed each delta by its frame uuid, so a
reply painted as one bubble per delta chunk followed by a complete duplicate
under the final frame's uuid. The block's first stream frame now mints the
claude:(sessionId, uuid) identity, deltas coalesce onto it through the shared
60ms seam, and the final frame reconciles onto that same item.

Known SDK bookkeeping no longer reaches the provider-fallback row: result
subtypes are catalogued and settled by the turn lifecycle, an empty thinking
block (redacted thinking) is a modeled kind, a string-content user replay is a
text block, and an empty user frame paints nothing. An unmodeled result
subtype or content kind still lands on the bounded fallback row.

Claude-Session: https://claude.ai/code/session_01GaP5HpYQbvy2hYehVhwfEW

* fix(claude): prove descendant exit at the close boundary instead of on an unref'd timer

close() reported proven=true as soon as the direct child exited while the
descendant sweep's SIGKILL sat on an unref'd 2 s timer, so a SIGTERM-resistant
MCP server outlived the lease release. The reaper now composes the same shared
primitives the Codex structured provider uses: snapshot, verified bounded
descendant termination on POSIX, taskkill /T /F on Windows. The proof is false
whenever descendants outlive the deadline, a retried close re-verifies the
retained snapshot rather than trusting the dead root, and the raw pipe child no
longer goes through the PTY job sweep it never owned a job for.

Measured on macOS: a killed child of a SIGSTOPped parent stays a matching zombie
row in ps, so the root is killed while verification runs rather than stopped
first as the Codex non-group path does.

Claude-Session: https://claude.ai/code/session_0161QFm3KVRNJKfdzWVGVNWk

* feat(claude): replace the hand-rolled control plane with the SDK's native surface

PR 3 of the Claude structured SDK migration removes the wire-frame scaffolding
PR 2 kept, so Orca drives the SDK's typed control surface directly.

Inbound permissions move from a rebuilt control_request dispatch to the SDK's
canUseTool / onUserDialog callbacks. The prompt registry now carries the
callback's own resolver: a decodable can_use_tool becomes a durable prompt whose
answer settles the callback; a malformed one is denied without registering; the
SDK's abort signal (fired on control_cancel_request, which the SDK matches and
dedups itself) forgets the prompt and settles it null, and a late answer after
abort finds no prompt and is refused. Closing settles every in-flight callback so
no promise dangles. The claude-agent-sdk-control-bridge that rebuilt the wire
frame is deleted.

Outbound control maps to Query methods: interrupt() for cancel, setModel /
setPermissionMode / applyFlagSettings for options, supportedModels for the model
list, initializationResult() for init proof, each under Orca's own request
deadline and error classification. Cancel is interrupt-receipt aware: a CLI
advertising interrupt_cancel_queued_v1 gets cancel_queued in one round trip,
otherwise the receipt's still_queued uuids are swept with cancel_async_message so
a cancelled turn cannot spawn a later unexpected turn; older CLIs resolve no
receipt. Init keeps the 10s deadline and the unauthenticated-startup guidance.

Every behavior is failing-first and ablation-proven; the toggle-off import
boundary and the accepted loss of unknown-control visibility rows are unchanged.

Claude-Session: https://claude.ai/code/session_01Pqjduxt5G4rr9aYvtp7rNm

* fix(claude): arm the descendant snapshot before stdin closes and make the tree verdict unproven by default

A healthy Claude root leaves within the graceful window, and the close ladder
only snapshotted descendants when the root was still alive after that window.
So the common close never looked at the tree: `treeExited` stayed null,
`!== false` passed it, and close() reported a proven exit with an MCP child
still running. A root that died before the walk made the snapshot vacuous too.

The proof is now unproven by default. The reaper holds one verdict in Orca's
vocabulary (exited / live / unverifiable), assigned in exactly one place from
the bounded verification, and close() returns true only on `exited`. The
snapshot is armed before stdin closes, while the root can still be walked, and
is verified after the root exits; a root that left before any snapshot could
be armed stays unverifiable rather than vouching for descendants it never
showed us. The shared verifier gains the three-way verdict behind its boolean
face, and the connection reports the root and tree verdicts separately along
with the child's exit status.

One verification per close attempt: the retried close re-verifies, so the
intra-attempt re-reap is gone from the teardown budget.

Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9

* fix(claude): verify the Windows tree after taskkill instead of trusting that it ran

`terminateWindowsProcessTree` resolves from taskkill's callback whatever the
error says, so a timeout, an access denial, a recycled root and a surviving
descendant all looked identical to the reaper — which then returned a proven
exit unconditionally. close() reported true and the lease was released with an
MCP descendant potentially still live.

The Windows branch now snapshots the root's descendants while it is alive and,
after taskkill, polls a fresh process table to a bounded deadline: a row still
matching by pid AND creation time is `live`, an unreadable table is
`unverifiable`, and only a table with no match is `exited`. Creation time is
the PID-reuse guard the POSIX path gets from ps lstart, so a descendant that
denied a creation-time query is omitted rather than signalled on a bare pid.
A root already observed exited is never taskkilled: `/T /F` on a recycled pid
would take an unrelated tree down with it.

The captured tree is tagged by platform so neither verifier can be handed the
other's rows.

Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9

* fix(claude): release a reservation on a first-hand root exit instead of latching it into manual recovery

Making close() strict about the descendant tree exposed a second defect at the
same boundary. A create-time acquisition has no ownerProcess until publication,
so an unproven cleanup mapped to handoffStage `manual-recovery`, and
adjudication then refuses every later attach with agent_session_ownership_unknown.
A user who was merely signed out, or whose --resume the CLI rejected, wedged the
session id permanently.

Each question now answers from its own evidence. close() is unchanged and stays
strict about the tree. Separately, the lease is keyed on the root's pid and
start time, so when Orca's own child handle observed that root exit and no
descendant snapshot was ever admissible, the reservation is released and the
CLI's exit code and stderr reach the user. A descendant observed still alive,
or a root Orca never saw leave, stays unproven and keeps the reservation.

The settlement records only what was observed: the released lease says the
provider process exited and its descendants were not verifiable, rather than
reusing the wording that claims cleanup proved no child remains.

Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9

* fix(claude): surface an API error a result frame reports instead of settling the turn on it

The SDK models an API failure as a SUCCESS-subtype result whose `result` string
is the user-facing error text, with no assistant frame behind it. The translator
suppressed every catalogued result subtype as turn bookkeeping, so that turn
tombstoned its lifecycle and showed the user a completed, empty reply with no
sign anything had failed.

Suppression is now by meaning. A result reporting a failure routes to the
bounded provider-error surface, leading with the provider's own sentence and
keeping the raw frame behind the row's disclosure; ordinary successful results
stay off the timeline as before. A turn the user aborted also stays suppressed:
its interrupt frame already says so, and its execution diagnostic would only be
noise on every stop.

Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9

* fix(claude): drop the stream state of turns that never received their final frame

Every streamed delta recorded its block's identity, latest text and checkpoint
length. Only the final assistant frame removed them, so an interrupted turn left
its whole accumulated reply reachable until the session was disposed, and a long
session with repeated interruptions grew those maps without bound. The partial
text was already journaled by the flush that precedes settlement, so the live
copy was pure retention.

That state now lives in its own module, named for what it does — grow a streamed
block's journal row between its deltas and its final frame — and turn settlement
drops every block still awaiting a final. The translator reports how many remain,
which is the invariant: a settled turn leaves none.

Also makes a timed-out process-table read retryable while the root is still
alive. A loaded host can miss the table's one-second deadline, and latching that
as "no descendants" both lost the descendant sweep and, on a busy machine, made
the close ladder report unproven for a tree it never actually looked at. Only
the root's death still makes a missing snapshot final.

Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9

* perf(claude): capture the Windows descendant tree from one process-table read

The capture walked the descendant tree and then read the table again for the
creation times the walk's projection drops. Each read is bounded in seconds and
both run inside the close ladder's budget, so the second one cost the worst-case
teardown three seconds for data the first read already held.

The walk is now exported from the module that owns it and runs over rows the
caller has already read, which is also what lets the snapshot keep the
PID-reuse guard the projection cannot carry.

Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9

* fix(pty): spend the descendant verification window instead of surrendering on one slow table read

The verification abandoned the whole check the first time a process-table read
missed its own one-second deadline, with seconds of its window still unspent.
On a loaded host that reported a tree unverifiable without ever having looked at
it, which the Claude close ladder then turned into an unproven close and a
retried teardown. It also made the descendant-exit tests flake under a parallel
suite run, for the same reason and with the same honest-but-premature verdict.

A read that missed its deadline is now simply not an answer: the loop waits and
reads again until its own deadline, and only a window that ends without a
readable table reports unverifiable. This can only turn a premature verdict into
one backed by evidence; it never manufactures a proof.

Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9

* fix(claude): never let a later failed look collapse an observed live descendant into unverifiable

The reaper's single assignment site latched only 'exited', so a second reap
whose table reads all missed their deadline overwrote an earlier completed
verification's 'live' with 'unverifiable'. The acquisition release gate
discriminates on exactly that pair, so a root exit after such a decay released
the lease over a descendant that had been observed alive. The latch is now
monotone in trust order: exited is final, and live is only ever raised to exited.

Claude-Session: https://claude.ai/code/session_01HfdhsvSJucLw4cTZxzg2CP

* fix(claude): never prove a Windows tree gone while a descendant denied identification

The Windows snapshot dropped rows that denied the creation-time query, and an
emptied snapshot was judged exited without any table read: a descendant Orca was
refused information about was treated as one that had left. The snapshot now
counts the unidentified rows it saw, and verification caps its verdict at
unverifiable while any exist. Nothing is ever signalled on a bare pid, as before.

Claude-Session: https://claude.ai/code/session_01HfdhsvSJucLw4cTZxzg2CP

* fix(claude): classify cleanup after a first-hand exit as a root exit instead of a proven tree

When the CLI died between a successful acquire and the host's commit or proof
of the lease, handleExit had already removed the session, so releaseAcquisition
found nothing and reported true. The attach flow then settled exit-proven with
deathEvidence claiming cleanup proved no provider child remains, though the
tree was never verified. The adapter now keeps the exit that removed a
published session until the session is acquired again; acquisition cleanup runs
that connection's close ladder and classifies its verdict exactly as a
start-time failure would be, so the record reads root-exit-observed. The wire
helper keeps that typed classification and its provider diagnostic instead of
wrapping it as unproven, and the router gives up its owner even when the
release throws.

Claude-Session: https://claude.ai/code/session_01HfdhsvSJucLw4cTZxzg2CP

* fix(claude): integrate SDK teardown and picker lifecycle fixes

* fix(claude): preserve resume leaf and settle processless spawns

* fix(claude): reacquire from persisted resume leaf

* fix(native-chat): restore Claude grouped question handling

* fix(claude): persist only resumable transcript leaves

* fix(claude): recover structured session exits safely

* fix(claude): close remaining structured session P1s

* fix(claude): harden transcript branch proof

* Remove superseded root fix reports

* fix(windows): restore indexed descendant row walk

* fix(router): forward force-close lifecycle

* fix(claude): fence stale turn cancellations

* fix(claude): fence cancellation after unknown dispatch

* fix(claude): fence replay and option recovery races

* fix(claude): block replay fallback after waiter eviction

* fix(claude): fence evicted slash results

* fix(claude): fence ambiguous results and restore options safely

* fix(claude): scrub SDK child env and localize pending launch

* fix(claude): pin transcript roots and exit recovery proofs

* fix(claude): retain unproven SDK exits

* fix(claude): settle retained exit before reacquire

* fix(claude): resume from settled retained cursor

* chore: remove tracked review artifact

* fix: harden Claude SDK transport session cleanup

* fix: close Claude sessions safely

* fix(claude): close races with fresh child snapshots

* fix(claude): fail closed on recycled child identities

* fix(claude): gate root cleanup on process identity

* fix(claude): fence same-second root identity reuse

* fix(claude): restore the root SIGKILL fallback the identity gate took away

The direct root kill goes through the handle Node owns, not through a pid:
libuv drops that handle in the same turn it reaps, so the signal either
reaches the process Orca spawned or reaches nothing at all. Gating it on a
process-table probe therefore bought no safety and cost the tree its only
fallback whenever the probe declined -- a first capture landing in the fork's
own second, a recycled descendant pid voiding the snapshot, or a process table
that could not be read on either platform.

Identity verification stays where a bare pid is genuinely addressed: Windows
`taskkill /T /F`, and the descendant sweep's own revalidation before it signals.

Also stops a declined root probe from collapsing an observed `live` or `exited`
descendant verdict into `unverifiable`, and stops a successful taskkill from
reporting `unverifiable` because a later probe found the root correctly dead.

* docs(claude): rewrap the root-kill ordering comment

* Match the Claude structured launch to the terminal path's managed-account auth rules

The SDK path stripped ambient Anthropic auth unconditionally, let an explicit
agentDefaultEnv override beat a pinned managed account, and had no account-switch
guard. Reuse the terminal preflight's own predicate and messages so both transports
strip, refuse, and report identically, and cover the CLI transcript location that
mobile native chat depends on.

* Reach the Claude structured chat lane from the desktop UI

The main process has had a complete, correctly gated Claude Agent SDK lane for
a while, but no renderer ever asked for it: the launch route accepted only
`codex`, and the create path was typed `agent: 'codex'` end to end.

Widen both to the structured provider union that already exists
(`AgentSessionHandleProvider`), and generalize the codex-named create path
instead of adding a Claude twin beside it. The pending-launch registry is now
keyed by agent as well as workspace — a shared key handed a second caller the
first agent's intent, so a Claude and a Codex launch in one worktree collided.

Windows, per agent. Codex's client-side win32 refusal is deliberate and settled
elsewhere, so it stays exactly as it was. Claude's answer is no longer guessed
from the client's platform: a structured session fences its provider child on
that child's process start time, and only the executing host knows whether it
can read one. `agentSession.createSupport` already answers precisely that, per
agent, and had no renderer caller — so the Claude create path asks it before
creating and turns a "no", or a probe it cannot get answered, into the
definitive refusal the launch fallback already handles. Fail closed either way.

That refusal mapping also closes a real gap: the host reports an unsupported
location by throwing `structured_agent_session_unsupported`, which reaches the
client as a transport rejection rather than a refusal envelope, so
`StructuredAgentSessionCreateRefusalError` never fired. The launch would retry
the create, strand itself in `visibilityUnknown`, run no legacy fallback, and
show an error toast.

Close a fail-open hole while Claude and win32 become reachable: `create` with a
client-supplied location, and `ensure`, both skip the worktree-resolving support
check. They now ask the executing host the same question directly, so a host
that cannot fence a provider child no longer creates one on a client's say-so.

Also deletes `structured-agent-session-provider-routing.ts`, a duplicate of
`structured-agent-session-provider-support.ts` with no importers.

WSL, SSH and paired hosts, floating workspaces, draft prompt delivery, explicit
TUI customization and initial session options all keep refusing; folder
workspaces keep working.

* P1-1: make the structured Claude auth policy required and testable

The optional dep plus a {stripAuthEnv:false} fallback meant a dropped wiring
under-stripped silently. Required at all three hops, asserted at install time for
the @ts-nocheck caller, and the settings-to-policy mapping is now a named tested
function.

* P2-3: mobile's default Claude transcript root must follow CLAUDE_CONFIG_DIR

session-file-resolver's default ignored the variable the pinned account home
follows, so a CLAUDE_CONFIG_DIR launch wrote one tree and mobile read another. The
Task-4 test now resolves with no root override (mobile's own call) and checks the
answer against the root the CLI itself reports, instead of mirroring the code under
test's own expression.

* P2-1/P2-2/P3: close the teardown window, join the live-auth gate, align the refusal

P2-1: a switch beginning inside the acquire teardown left a dead chat and no
replacement. Past that point the launch waits the swap out and refuses only if it
never settles; the entry guard still refuses outright, because nothing is torn down
there yet.
P2-2: structured children now hold the same OAuth-refresh gate a Claude PTY does,
so a managed refresh cannot rotate the token out from under a live turn.
P3: the refusal now matches the strip it guards (case-folded on win32, presence not
truthiness), and the dead structured-to-TUI builder states its auth policy instead
of silently signing a system-auth user out.

* Make the live-auth gate tests independent of sibling connection teardown order

* Do not offer structured Claude under a WSL-only managed account

Structured Claude launches against the ambient Claude config, which the account
service keeps in sync with the selected HOST account. A WSL-bound managed
account lives inside the distro and is never synced there, so on Windows a
structured session would authenticate as whatever the ambient identity happens
to be while the UI names the WSL account — the user is told one identity and
given another.

That was unreachable only because nothing offered structured Claude on win32.
Enabling it makes it reachable, so gate it here rather than patching the auth
layer: refuse the structured path when the active managed Claude account is
WSL-bound, and let the terminal-backed path — which resolves the account per
runtime — handle that account shape.

The answer rides the agentSession.createSupport seam the renderer already
consumes, so no new capability and no renderer knowledge of account internals.
A create the host declines becomes the definitive refusal the launch fallback
already turns into a legacy native chat tab, with no error toast.

Unknown answers refuse. An install with no managed accounts claims no identity
and is fine, but an active selection that cannot be resolved — or account state
that cannot be read at all — is not evidence that the ambient identity is right.

Claude only. Codex resolves its account through a different path and its
createSupport answer is untouched, as is every Codex routing decision.

* Read the structured Claude account gate through the auth policy's accessor

The gate resolved the active account from the account-service snapshot's
runtime map; the auth policy resolves it with
getSelectedClaudeAccountIdForTarget(settings, { runtime: 'host' }). Those are
two sources and two resolution rules, and they disagree on a legacy settings
blob that carries the selection only in the flat activeClaudeManagedAccountId:
the accessor falls through to it, a direct read of the runtime map does not. The
gate would then refuse a launch the policy would have run under host-1 — and in
the mirror case a session could be admitted under a policy computed from a
different account than the gate approved.

Read the same settings through the same accessor so agreement is structural
rather than coincidental, and drop the controller accessor that existed only to
reach the snapshot.

No behaviour change for any state both already agreed on; Codex is untouched.

* Round-3 review fixes: N-1 empty-value regression, N-2 gate leak window, N-4 lost history

N-1: my presence-based conflict predicate refused a terminal launch that works
today. 'ANTHROPIC_API_KEY=' is how a user blanks a variable and the settings
pipeline preserves that empty value; an empty override cannot beat the pinned
account and the strip removes the name anyway. Back to truthiness for the value,
keeping the win32 case folding.
N-2: enter the live-auth gate only after the exit/close handlers that release it,
so no throw in between can leave an entry nothing reconciles.
N-4: the Claude transcript resolver searches config-dir-then-default and de-dupes,
matching the Codex sibling in the same file, so adopting CLAUDE_CONFIG_DIR no
longer hides history written before it.

* Run the managed-account gate on every Claude acquisition, not just create

createSupport gates the create path, but a session's account state can change
while it lives. A reacquire after an unexpected child exit re-resolves the
launch and re-derives auth, with nothing re-checking the gate — so a session
created while supported could come back up in the refused shape. With the strip
predicate keyed on there being an active non-WSL account, the WSL-only user's
normalized steady state (accounts exist, none active) does not strip, and that
reacquire reaches the child with ambient auth while the UI names the account.

Gate at resolveLaunch, the one choke point every acquisition passes through,
refusing with the pre-spawn error the caller already handles. Same predicate as
create-time, now sharing one settings reader so the two cannot drift.

Claude only; Codex resolves its account on a different path and is untouched.

The runtime class that wires this does not typecheck its own `this` calls — a
missing hookup compiles clean — so the wiring is pinned behaviourally rather
than trusted to the compiler.

* Move the structured Claude gate out of the @ts-nocheck runtime files

Both call sites of the managed-account gate sat in files whose first line is
`// @ts-nocheck`, so neither was typechecked: three arguments to a one-argument
function plus an undeclared identifier compiled clean. New auth-identity
decision logic had no compiler behind it.

Move the verdict into a checked module that takes the two facts the runtime
owns — the adapter's answer and a settings getter — and decides. The runtime
class now only forwards. Move the gate reader's construction into the checked
installer too, so the nocheck file passes a plain settings closure and never
names a gate symbol.

Every reference to the gate predicate and its reader now lives in a checked
file, so the ablation that used to pass silently is a compile error at both the
create-support and reacquire sites.

Removing the file-level @ts-nocheck is a separate, larger job and is not
attempted here.

* Derive the gate test's auth policy from the settings under test

A hardcoded stripAuthEnv asserts a gate/policy pairing production cannot
produce, and false additionally lets launch.env inherit the runner's real
process.env. Derive via claudeStructuredAuthPolicyForSettings instead: the
gate settings type is the same Pick the policy takes, and both resolve the
account through getSelectedClaudeAccountIdForTarget.

* Pin the absent-vs-empty distinction in the managed-account gate

An empty claudeManagedAccounts array is a real answer: the user has no managed
accounts, nothing claims an identity, and the ambient path is legitimate. A
readable settings object with no such field is settings we failed to parse —
the same unknown as unreadable — so it refuses.

The two are one character apart in the code and the difference is invisible
without the reasoning, so record it at the branch and pin both sides. The test
fails under the obvious "consistency fix" of treating a missing field as empty.

* fix(claude): keep command queue bookkeeping out of the transcript

Claude Code 2.1.258 emits a `command_lifecycle` frame for every uuid-stamped
command it starts, completes or cancels. The frame carries a command uuid and a
state and no content, and the CLI keeps it out of its own transcript -- but it
is absent from the SDK's SDKMessage union and so from Orca's frame catalogue,
where an uncatalogued kind defaults to a substantive row. Every structured turn
therefore painted raw JSON rows into the user-visible transcript.

Catalogue it and disposition it as status chrome. The unknown-kind default stays
`timeline-substantive`: a kind we have never seen is likelier to carry content
than to be chrome, and a visible row we can catalogue later beats content we
silently dropped. A lifecycle state that reads as a failure still surfaces,
because the payload error check in `classifyProviderFrame` outranks the
catalogue.

* fix(claude): let a re-walked descendant become eligible for the forced sweep

A descendant first observed by a capture inside its own birth second could never
be SIGKILLed: `ps lstart` is second-resolution, so that capture cannot rule out
a pid recycled later in the same second, and the merge pinned each retained row
to the boundary of the walk that first saw it. SIGTERM-resistant children forked
in that window were signalled and then never escalated -- they survived close,
quit and restart, reparented to init, and had to be killed by hand.

Advancing that boundary on any later capture would be unsound: a later capture
matching pid, pgid and start-second is exactly what an impostor would also show.
But a capture is not a match -- it is a fresh ppid walk from a root Node pins
through its own handle, so a row it re-derives is proved ours at that instant
without appealing to its start time. Chain the fence from there instead, and
take that walk at the close boundary while the root certainly still lives: the
root may leave inside the grace window, and the post-timeout refresh never runs.

A row absent from the later walk still keeps its earlier boundary, and a row no
walk has ever re-derived in a later second is still never escalated.

* Treat an absent managed-account list as empty, not as unreadable

An empty claudeManagedAccounts array and a missing one are the same answer:
this user has no managed Claude accounts, so nothing claims an identity and
ambient auth is the truth. Refusing on absence strands any profile that simply
never wrote the key, and it disagrees with the auth policy, whose own predicate
takes `(accounts ?? [])` for exactly this reason.

Only settings that cannot be READ stay unknown, and those still refuse — as do
a WSL-bound active account and a selection naming an account the list does not
explain.

The earlier reasoning treated a missing field as settings we failed to parse.
That conflated "not present" with "not readable"; only the second is unknown.

* Support structured Claude when accounts are registered but none is selected

Registered-but-deselected Claude accounts were refused, which is behaviourally
identical to having no accounts at all: the auth policy does not strip, ambient
auth is the truth, and the UI names no host identity. A user who deselected
their accounts silently got legacy chat with nothing explaining why.

Nothing selected for the host runtime is two states the settings cannot tell
apart after the fact, because pruneInvalidClaudeRuntimeSelection empties the
host slot and persists null in the second one:

  honest deselection      -> ambient auth, UI names nothing   -> SUPPORTED
  the WSL-only steady state -> ambient auth, UI names the WSL account -> REFUSED

The presence of any WSL-bound account in the list decides. Simplifying this to
"none active -> supported" re-opens the auth-identity misrepresentation, so the
tests fail loudly on exactly that: five of them, across the unit rule and the
createSupport path.

* Stop treating an unanswerable create-support probe as a refusal

A worktree is not resolvable for a beat after createWorktree resolves, so a
probe fired immediately after creation fails the RPC with selector_not_found
instead of answering. The catch collapsed that into `supported = false`, so the
composer refused and quietly built a terminal session — the gate never said no,
it was never asked successfully. Elapsed time was the only input that decided
whether a Claude launch went structured.

"Could not answer" and "answered no" are different states and only the second
is a verdict. Retry while the host cannot yet resolve the selector, with a
bounded backoff that covers the measured window with margin, and keep refusing
on the first ask for everything else. Fail-closed is unchanged: a probe that
still cannot be answered when the budget is spent refuses.

The retry is narrowed with the shared error-code matcher, which classifies a
token that transports re-wrap into a longer message without matching prose that
merely mentions it.

Codex never probes, so this race has never been able to refuse a Codex launch —
the race itself is identical for it. Recorded at the early return, because
whoever gives Codex a probe inherits the bug.

* fix(claude): fence the forced sweep on re-derivation, not on lstart's second

A descendant forked in the same wall-clock second as every walk that sees it was
signalled with SIGTERM and then never escalated, so a SIGTERM-resistant child
survived tab close, app quit and a full relaunch. Two children of one parent
96ms apart across a second boundary took opposite paths. The leak predates this
branch: it reproduces with the change reverted.

`ps lstart` has one-second resolution, so a walk landing inside a row's birth
second can never rule out a pid recycled later in that same second. But a walk
is not a match: a ppid walk only reaches what the root actually parents, and the
root is pinned by Node's own handle, so a row the walk re-derived is ours
whatever second it was born in -- a stranger would have to have been forked into
our tree, and then it is not a stranger. Fence the escalation on that.

Rows a merge retained from an earlier walk are not re-derived and still answer
to the start-time fence, which remains correct for them.

Scoped to callers that revalidate identity before signalling, which is the
Claude close path. Codex teardown reaches this same verifier and is unchanged;
the argument holds there too, but widening it is its own deliberate change.

Also reverts two changes from the previous attempt at this leak. Advancing the
capture boundary on a later walk is inert once the sweep fences on re-derivation
-- both key on the same set of rows, so the new term short-circuits for exactly
the rows whose boundary it advanced. The extra ladder refresh was a duplicate
full process-table read: close() already awaits tree.refresh() immediately
before proveClaudeChildExit, on the only path that reaches it.

Known property: the kill lands roughly a grace window after the walk that proved
membership, so a pid recycled inside that gap could in principle be signalled.
It is bounded -- matchingSnapshotRows already requires the live row to carry the
same start-second and pgid, so an impostor must be born in the remainder of that
one second, land on that exact pid, and sit in the same process group, and it
has already received the unfenced SIGTERM from the same loop.

* Run the Claude structured integration suite as a runtime client

The suite exercises agentSession.* for Claude, not the mobile surface: nothing
in it asserts anything mobile-specific and its sibling integration suites use
'runtime'. Mobile now additionally requires the experimental structured-chat
setting, which structured-agent-session.test.ts pins in both states, so the
stale 'mobile' fixture was claiming coverage it never had.

* fix(claude): report effort from get_settings, which is the only frame that has it

The composer's Effort pill rendered blank in every structured session. This is
not a missing source: the publication reads `effortLevel` off the `system/init`
frame, and that frame has never carried an effort of any kind, while the correct
value is already fetched at acquisition and thrown away on the auth diagnostic.
Verified two ways -- a live get_settings probe against Claude Code 2.1.258, and
the shipped binary's own init frame construction, which lists `model` and no
effort. So `reportedOptions.effort` was always empty, the options reader dropped
the key, and the pill had no value. Model survived only because
`currentModelId()` has a fallback chain.

The get_settings call acquisition already makes reports the session's current
effort as `effective.effortLevel`; pass that into the publication instead.
Selecting an effort already worked, so this is the arrival value only.

The legacy PTY path is unaffected and must not be "fixed" to match: it reads its
effort by parsing the startup banner (`CLAUDE_MODEL_EFFORT` in
src/renderer/src/components/native-chat/claude-terminal-session-options.ts),
which is why it shows a value where the structured path does not.

Also removes the fixture that hid this: the fake init frame invented
`effortLevel: 'high'`, a field the CLI does not send, which is why every gate
stayed green over a value that is always empty in production. The fixture's
get_settings now returns the real {applied, effective, sources} shape instead of
a bare `{env: {}}`, so the two adapter tests that asserted an effort keep
asserting it through the path production actually uses.

The reader returns null rather than defaulting: an effort nothing measured would
repeat the fixture's mistake, and a blank pill is the honest degradation if the
provider ever renames the key.

* fix(claude): only record an effort the child confirms it adopted

apply_flag_settings answers `success` for an effort it then ignores. Measured
against Claude Code 2.1.258: applying `bogus-effort-xyz` returns
subtype "success" with no error while `applied.effort` stays at its previous
value, and a valid `low` moves it. The option write treated the absence of a
throw as adoption and recorded the requested value unconditionally, so Orca
would show and persist an effort the child was not using, with nothing anywhere
reporting a problem.

Read the effort back after applying it, through the same reader the arrival
value uses, and reject when the child reports a different one. A readback that
could not be taken is not evidence of a refusal -- the apply itself succeeded --
so it still records; only a readback that disagrees rejects.

Not reachable from today's picker, which offers catalog values only, but the
CLI's effort catalog is server-delivered and has changed before, so a retired id
would otherwise become a pill confidently displaying a setting that never took.

* test(claude): assert the effort contract against the real binary

The blank pill survived every gate because the only tests that touched it were
fixture-backed, and the fixture invented the field. A test that pins the shape
we read cannot catch the provider renaming the key, which is the failure mode
that produced this defect.

Asserts both halves against a live authenticated CLI: that no frame it publishes
carries an effort at all, and that the session's current effort arrives through
get_settings. Which frame proves the session varies by host -- this machine
proves it with a SessionStart hook rather than a system/init frame -- so the
negative half asserts over every published frame rather than picking one.

Skips with the rest of the file when no authenticated CLI is present.

* fix(claude): stop the synthesised content-part kinds leaking into the transcript

Sending an image put a bare `claude · message:user:content:image` row between
the user's bubble and the answer. Two causes, and only the second is a family.

An image part counted as modelled only when `source.type === 'url'`, but
claudeDispatchMessageContent sends a local attachment as a base64 source and the
CLI replays that shape back, so every attached image was classified unmodelled.
Accept the base64 and file sources Orca itself sends.

The family is the real defect. `message:<role>:content:<type>` kinds are
synthesised at runtime from whatever `part.type` arrives, so unlike the
top-level frame catalogue they can never be enumerated ahead of time -- the
`?? 'timeline-substantive'` default then prints the synthesised name at a user
who cannot act on it. That default is right for top-level frames, where
"substantive" means show the frame; here it meant show our own vocabulary, which
drops the content AND leaks the opcode.

So an unrenderable part now renders a sentence saying exactly that, with the
kind and payload still on the row's disclosure. A part that carries its own
readable sentence keeps it -- the placeholder is a fallback, not an override.

An unknown future part type is therefore visible, never silently dropped and
never printed as a kind: the same principle as the effort readback, which
records only what the provider confirms.

* Declare agentSession.requestHandoff on the cross-version wire surface

The manifest is a ratchet for cross-version reachability, so the method is
declared with real HandoffParams rather than counted. requestHandoff is
capability-gated through requireStructuredHost and has no client caller, so
declaring it is the whole of the change.

Also model two host capabilities the harness omitted: the stub host's
supportsCreate, and the fake adapter's, without which adapterSupportsCreate
falls through to a supportsLocation the fake also lacks. Every ensure was
refused for the harness's silence rather than for its location.

* Gate structured Claude session tabs on the client capability that names them

The Claude structured lane deleted the projection's `agent !== 'codex'`
filter and added CLAUDE_STRUCTURED_AGENT_SESSION_RUNTIME_CAPABILITY in the
same commit, but never wired the constant to anything. Paired clients then
received agent-session tabs for Claude, which no shipped client renders --
mobile's resolveMobileNativeChat returns null for every agent but codex, so
the row listed and selected into a pane with neither chat nor terminal.

Restore the filter behind the declared capability instead of the bare agent
name. No client advertises it yet, so this matches main's behaviour today
and becomes a negotiation a future client can opt into.

* Confirm the structured Claude model against the model the CLI reports

set_model answers success for any string, including a model it cannot
resolve — the failure only surfaces when the turn runs — and get_settings
reports the settings-file model, not the session's. The init frame that
opens each turn is the only channel carrying the adopted model, so keep
the session's reported model current from it instead of reading it once
at acquisition.

Also stop rejecting an effort the readback cannot represent: max is
session-scoped and excluded from the persisted effortLevel, so a readback
reporting the level underneath it is an absence of evidence, not a refusal.

* Clear the session-option hedge when the provider confirms the value

The pill claimed every option was unconfirmed for the life of the session:
the renderer recorded each write as dispatched and nothing ever moved it,
so a model the CLI had already reported back still read as unconfirmed.

Carry the provider's own confirmation to the surface. Main reports which
option ids the provider named rather than merely accepted, and the client
re-reads options as a turn changes, because the frame that opens a turn is
where the adopted model arrives. A value the provider has not reported
stays hedged, including an effort whose readback could not be taken.

The confirmed list is optional on the wire: a host that predates it sends
nothing and the client keeps hedging, which is the behaviour it had.

* Keep the model report current across an acquisition fence bump

* Show the picked session-option value and let the provider report correct it

The pill showed a "not confirmed" second tooltip line for any value we had sent
but not yet seen reported back. Nothing acts on it, and for the PTY lane it was
permanent — that transport has no report channel. The pill now shows the picked
value immediately and the provider's per-turn report corrects it when the two
disagree; a newer local write still outranks a report that precedes it.

`dispatched` stays as a provenance member rather than collapsing into `applied`:
it is produced independently by the PTY lane, and it is where the `confirmed`
wire field lands, which would otherwise be unobservable.

Effort keeps its readback and its rejection path. That matters more now, not
less: with the hedge gone the rejection is the only user-visible failure signal
on this surface, so a spurious one would be the loudest bug here. Skipping the
readback for an effort the settings response structurally cannot echo is what
prevents it — the response carries the persisted level, so reading it back for a
session-scoped value would report the level underneath and fail a valid write.

* Hedge a session-option value only when the terminal transport sent it

Both lanes emit `dispatched`, so it could never say which one produced a value.
The descriptor now carries the transport that built it, set once in the shared
snapshot builder from a parameter that is required rather than defaulted — the
builder is the only place a descriptor is constructed, so a new producer has to
name its lane or fail to compile.

The structured lane confirms every value from the provider's own per-turn report,
which makes the hedge transient noise there. The terminal lane can only learn an
outcome by parsing the screen back, and only for Claude: every other agent's
`dispatched` value stays unconfirmed for the life of the session, so the line is
the only signal that we sent something we never saw land.

* Refuse an effort the session's model advertises no control for

* Refuse tab mutations on a Claude row the client never negotiated

The branch added a case asserting a client advertising only
agent-session.structured.v1 may mutate a claude row. That is the same
ungated behaviour the projection gate removes, encoded a second time —
mutation authorization reads the projection, so hiding the row refuses
the write. Assert that contract instead, and add the positive case for a
client that does negotiate Claude rows.

* Resolve the Claude session's current model in one place so the effort guard and the pill agree

* Record an effort the child did not adopt instead of refusing the write

apply_flag_settings answers success for an effort it then ignores, so the
readback exists to detect that. Refusing on it made the detection a veto,
and a veto is only correct if the readback can never be wrong about which
model is current -- which it was, twice. The pre-flight guard already
refuses a level the model advertises no control for, so the veto guarded a
door that is now locked upstream.

Keep the detection, drop the refusal: a disagreement records the child's
own answer and omits the option from confirmed, so main stops vouching for
a value the provider rejected without blocking the user's write.

* Stop a slow whole-machine ps from being read as an absent process

`ps -axo ...command=` pays a per-pid argv read: measured 1.15s for 1,948
processes (0.03s without `command=`), and CPU contention stretched the same
capture to 6.0s. Two budgets sized for a cheap look then misreport a readable
machine.

The reader's 3s ceiling killed 6 of 20 consecutive captures at load 27, so
every consumer answered "unverifiable" about a table it could read. Raise it
to 15s, and stamp the capture instant at ps START so `capturedAgeMs` is the
upper bound its contract promises -- a 6s capture used to report itself as
freshly taken, understating staleness against a 5s kill gate. The TTL keys on
completion so a slow capture still coalesces instead of forking ps per caller.

`readStructuredTuiProcessIdentity` then spent its whole 5s wait inside one
capture and concluded "no exact child" after a single look taken before the
child existed (observed landing at ~3.5s). Absence needs a look that did not
race the spawn, so require two captures before the deadline can end the loop.

Both surfaced by the real-binary Claude TUI resume test, which failed ~1 in 5
under load; 14/14 now, 8 of those runs containing a capture the old 3s budget
would have killed.

* Let the desktop renderer negotiate Claude structured tabs

The paired-client gate hides agent-session rows an agent the client cannot
render. The desktop renderer's own IPC dispatches as clientKind 'runtime'
advertising only agent-session.structured.v1, so the gate hid Claude rows
from the surface this feature ships on. It renders them; it should say so.

* Stop a slow process table from silently blinding every freshness gate

Stamping `capturedAgeMs` at ps START made the number honest, and honest broke
both consumers that read it. `ps -axo ...command=` measured 2.5-9.0s on an idle
2,002-process laptop and 4.0-18.6s at load 46, so the age it now reports lands
past every budget: `planRelayPtySweep` refuses the stop as "too old", and the
renderer's `admitRemoteForegroundEvidence` refuses the record outright. That
second one is the expensive half and was outside the diff -- a refusal bumps
`consecutiveInspectionErrors`, the poll scheduler backs off to its 10s floor,
and agent-completion detection stops for the pane. The subsystem went blind on
exactly the loaded hosts the honest stamp was meant to serve.

The evidence-publishing read now gives up at 1,200ms instead of waiting out
`PS_TIMEOUT_MS`. It is one budget for one question: these consumers ask whether
an observation describes NOW, and past this it does not -- a late answer is
refused by the age gate anyway, having first blocked a polled path for the whole
capture, so a prompt `unverifiable` is both the truthful verdict and the cheap
one. Both relay call sites already produce it from a rejection, and an admitted
`unverifiable` costs a poll where a refusal costs the cadence. Identity proof
keeps the full 15s through `getFreshProcessTableSnapshot`, because it asks
whether a process EXISTS and must never read slow as absent. The budget bounds
the wait, never the capture: the reader coalesces, so an abandoned wait leaves
its capture running to fill the cache rather than forking a second whole-machine
`ps` on the host that can least afford one.

1,200ms is bracketed rather than picked. The floor is the capture's own cost --
`command=` measured 1.15s for 1,948 processes on an idle host, and a budget
under that answers `unverifiable` about a machine nobody is straining. The
ceiling is the consumer's: 2,000ms, less the 500ms a TTL-shared capture may
already have aged, leaves 1,500ms, and transit takes the rest.

That ceiling only fits once the capture stops being charged twice. `ps` runs
inside the RPC round trip, so its duration is already in `receiveDelay`, and
`capturedAgeMs` is that same duration on the host's clock; summing them halved
the budget this gate grants a host from ~2.0s of `ps` to ~1.0s, which is why a
1.2s capture arriving at 1.3s read as 2.5s old and was refused. Admission now
takes the larger of the two. The sweep's gate keeps its sum, which is correct
there: `evidenceAgeSinceListingMs` is stamped after the listing ARRIVES, so it
measures planning time and overlaps nothing.

A stated limit rather than an assumed one: 15s is not proven sufficient for
identity proof. The same capture reached 18.6s at load 46, so that path can
still time out and answer "no exact child" about a host it simply could not read
in time. Narrowing it needs a cheaper question than a whole-machine argv read,
not a larger number.

The one test guarding this field could not fail. `beginPtyHandlerTest` installs
fake timers, so `Date.now()` is frozen, the real reader reports exactly +0, and
`0 <= 500` held identically for a hardcoded zero, for completion-stamping and
for start-stamping -- while the real reader on that host returns thousands of
ms. It now drives a measured age in and asserts the handler publishes it rather
than restamping; that the reader MEASURES it correctly stays pinned separately,
against a controllable clock. Both consumers get boundary coverage either side,
and each new gate was ablated red before it went green.

* Keep the compatibility fields off the capture the budget just abandoned

inspectProcess falls back to processHasChildren and listProcesses to
getForegroundProcessName, and both read the same TTL-shared capture with
no budget of their own. On a slow host they joined the in-flight capture
the budgeted evidence read had just given up on, so the call still blocked
for the full 6-18s and the budget bought nothing -- once for inspectProcess
and once per managed PTY for listProcesses.

Use the degraded answers those helpers already give for an unreadable
table, reached promptly. pty.hasChildProcesses keeps its unbudgeted fresh
probe: it is a one-shot destructive gate that can afford to wait.

---------

Co-authored-by: Merge Sim <merge-sim@local>
Co-authored-by: Merge Sim <sim@local>
2026-09-04 15:55:20 -07:00
Brennan BensonandMerge Sim f4c2821167 refactor(agent-session-journal): move the session journal onto SQLite (#18652)
* refactor(agent-session-journal): move the session journal onto SQLite

The agent-session journal kept its state in three hand-rolled file formats: an
append-only `log.jsonl` with torn-tail repair, a `snapshot.json` holding folded
state plus a retained tail, and byte-quarantine files for anything unreadable.
This replaces all of it with one SQLite database per session — `journal.db`
beside the existing `blobs/` store — using the in-house adapter and the
open/pragma/migrate/harden pattern the orchestration database already follows.

Two tables: `journal_rows` (the append-only log, keyed by
`(session_id, epoch, seq)`) and `journal_sessions` (the derived projection,
upserted in the SAME transaction as every insert). Rows stay JSON in one
column, so the row schema, the version upcast chain, and the reducer survive
byte for byte — `journal-reducer.test.ts` and four other suites pass unchanged
and are the regression proof.

Deleted: `journal-log-file.ts`, `journal-compaction.ts`,
`journal-corruption-quarantine.ts`, and the public `compact()` /
`compactionBoundary` / `autoCompact` members, none of which had a non-test
caller.

Existing `log.jsonl` / `snapshot.json` journals are deliberately abandoned. No
importer: a session created on the old path stops working, which is acceptable
because the feature is off by default.

## The physical quota is repriced, because SQLite does not charge like a file

The 256 MiB per-session bound is unchanged, but the arithmetic under it could
not survive: SQLite grows the database in pages and the WAL in frames, and the
checkpoint that copies the WAL forward holds the same pages in both files at
once, so a transaction's peak is about twice its content. Admission now charges
the candidate transaction's own measured page cost, validated against a sweep
that runs as a regression test (`journal-database-space.test.ts`) rather than
derived from reasoning about the allocator.

Four things are load-bearing rather than tuning, each measured:

- `auto_vacuum = INCREMENTAL` must be set BEFORE `journal_mode = WAL`. Set it
  after and it is ignored with no error, reclamation silently becomes a no-op,
  and the file never shrinks again. Both halves are asserted.
- `wal_autocheckpoint = 0` plus an explicit `wal_checkpoint(TRUNCATE)` at the
  end of every write path, so the one moment the same pages live in two files is
  a moment the charge accounts for.
- Reclamation runs in bounded chunks. A single unbounded `incremental_vacuum`
  took a 252 MB directory to 504 MB — the reclamation added to defend the bound
  would have breached it. `PRAGMA incremental_vacuum(N)` also frees exactly one
  page unless it is stepped to completion, which no size assertion catches, so
  the freed page count is asserted directly.
- A blocked checkpoint leaves the WAL on disk together with the database growth
  it already copied, so admission charges that deferred copy explicitly. The
  term is zero whenever the last checkpoint succeeded, so the uncontended path
  admits and refuses an identical set.

The epoch discard is `DELETE FROM journal_rows` with no WHERE clause, which
takes SQLite's truncate optimization: measured at ~0.26% of the database in WAL
bytes where the `WHERE session_id = ?` form rewrote every emptied leaf at up to
99%. One database per session is what makes the unqualified form correct.

An open, empty journal costs 57,344 bytes before a single row exists, so a
configured quota below `JOURNAL_MIN_SESSION_BYTES` now fails loudly at open with
the existing `journal_bound_exceeded` instead of as a run of identical append
failures. No production caller configures one; the affected surface is test
fixtures, rescaled to the smallest value that restores what each case proves.

## One deliberate behaviour change

Compaction was the only mechanism that shed bytes inside an epoch, and the write
path called it precisely so an append at the bound was not refused. The
SQLite-shaped replacement — a bounded prefix delete — cannot be used: with the
snapshot gone the surviving rows ARE the state, so dropping the oldest of them
loses the oldest transcript silently at the next reopen. So no row is ever shed
inside an epoch, and a session whose row bytes alone reach the bound now refuses
every append where it previously compacted and continued. A loud typed refusal
beats silent data loss.

What still sheds is unreferenced BLOB bytes — the dominant and unbounded byte
source — on the same write-path hook. The escape from the hard stop is the fold
that already exists, `replaceEpochItems`, which now actually returns bytes to
the filesystem instead of leaving them on the freelist.

The prune's protected set is a union of live reducer digests AND the candidate
row's own digests, including those cited only by a nested lifecycle-batch
mutation. Content addressing never rewrites a digest already on disk, so
protecting live state alone deletes the blob the append is about to cite — a
dangling reference that surfaces one reopen later as an empty expansion on an
item the user can see. `journal-store-blob-budget.test.ts` pins it, and it goes
red when the set is narrowed back.

## Handle ownership

A file handle used to be opened and closed per append; a SQLite handle is held
for the session's lifetime. Every path that can open a connection now has one
owner: the open function owns its raw connection until it returns, the store
owns its retained one and releases it in a new `close()`, and every other
connection is closed by the call that opened it. The attach, recovery,
eviction, map-overwrite and host-teardown paths close what they drop, and host
teardown is failure-complete — the sink-barrier flush throws by design, so a
trailing close statement would be skipped on exactly the path that leaks.

`close()` has a stated contract: admission at enqueue and permanent, the close
step on the same queue past that gate, one shared in-flight attempt, fulfilment
terminal, and the release last and deliberately unguarded so a retry re-enters
it. Guarding the release would skip it on retry, guaranteeing a permanent leak
in exactly the case where it did not release.

`journal_closed` joins the error union for a write after `close()`; no file
outside the directory references any of these codes.

* fix(agent-session-journal): make a COMMIT final, stop repairs deleting valid rows, and keep rejected closes retryable

Six review findings on the SQLite journal migration.

1. A successful COMMIT is now the point of no return. The ordinary append,
   the epoch roll and the epoch replacement each adopt the committed row or
   epoch BEFORE any post-commit filesystem work; checkpoint, reclaim, blob
   prune and directory measurement run through `runJournalPostCommit`, which
   is best-effort by design and falls back to the transaction's own charge as
   a conservative footprint. Previously a post-COMMIT scan failure rejected a
   durable append and the next one reused its sequence, and a failed epoch
   housekeeping step left the store writing into a prefix already deleted.

2. Corruption repair preserves instead of destroying. A rejected suffix is
   copied into a new `journal_quarantine` table and removed from the live
   epoch in ONE transaction per chunk, charged against the session bound
   before a byte is written; a journal that cannot afford the copy refuses to
   open rather than falling back to deletion. The repair state is exposed as
   `journal.repair` and the rows are readable through
   `recoverQuarantinedRows()`, so Orca-owned submission, receipt and
   lifecycle identity survives a gap or a malformed row.

3. The physical charge covers the B-tree key payload. `session_id` and
   `epoch` are stored in both tables and both primary-key indexes and appear
   nowhere in `row_json`, so the journal boundary now bounds them and
   `journalTxnPhysicalCost` charges those bounds plus the projection upsert.
   The charge sweep runs the exact production transaction at maximum admitted
   key sizes.

4. A rejected `close()` no longer orphans its handle. Callers hand the
   journal to `agentSessionJournalCloseRetries` instead of swallowing the
   rejection, the attach map replacement is ABORTED when the previous
   journal will not close, host teardown retries what the registry holds, and
   a failed runtime teardown is retained so the next stop is a real retry.

5. `journalWalBytes()` returns zero only for ENOENT and propagates every
   other stat error, so admission and reclamation fail closed.

6. The WAL contention test closes the writer before removing its temp root
   and asserts the directory is removable once handles close.

Regression coverage: post-commit divergence (4), corruption repair (5),
key bounds (5), WAL stat (8), close retry (5), plus a runtime stop-retry
case. Each fix was ablated on this head and the matching tests go red.

* fix(agent-session-journal): anchor replay at sequence 1, make quarantine append-only, and charge it in bytes

Three ways the corruption quarantine still lost rows it was written to keep.

Replay validated contiguity from the first row that HAPPENED to remain, so an
epoch missing only its sequence-1 row declared the leftovers contiguous and set
no `truncateFrom`. The load was still corrupt, so recovery imported provider
history and `replaceEpochItems` deleted every live row — including Orca-minted
submission, receipt and lifecycle identity that no transcript can reconstruct,
and that nothing had quarantined. Replay now anchors at sequence 1, so a missing
epoch row rejects the whole surviving range before any replacement runs.

`journal_quarantine` was keyed on `(session_id, epoch, seq)` and copied with
`INSERT OR REPLACE`. A repair frees the sequences it removed and the live epoch
reuses them, so a second repair in the same epoch silently deleted what the
first preserved. The table is now keyed on a surrogate `quarantine_id`, the copy
is a plain append, and `(epoch, seq)` is metadata; existing v1 databases are
rekeyed in the migration that already bumps `user_version`.

The admission charge read `length(row_json)`, which counts CHARACTERS for a TEXT
value where `journalTxnPhysicalCost` expects physical UTF-8 bytes. A multibyte
suffix was charged at up to a third of what it writes, which defeats the
pre-write physical bound — over a megabyte on a maximum-size lifecycle batch.

* fix(agent-session-journal): keep a repaired epoch anchored and stop the v1 quarantine migration doubling the file

Replay validated numeric contiguity from sequence 1 but never that sequence 1
IS the epoch row. When the anchor was missing the repair set aside every
surviving row, and if provider-history import then failed — a transcript that
is temporarily gone is enough — the journal reopened as a clean, row-less
epoch: an ordinary append took sequence 1, replay accepted it, read-restore
published it as history, and automatic recovery never ran again while the
user's real messages sat in quarantine.

Replay now rejects an unanchored prefix, the open publishes an
`unreconcilable_prefix` anchor for an epoch its repair emptied, and that anchor
keeps reporting corrupt — so provider history is retried on every attach —
until the timeline is rebuilt or the session writes content of its own. A
repair also discloses rows it set aside when no line was unreadable at all,
which is the case that removes the most.

The v1 quarantine rekey copied every legacy row into the new table inside one
transaction and dropped the old one. A quarantine holds whole rejected rows: a
single 8 MiB row nearly doubled the database past the physical bound the open
had already checked, the dropped pages only reached the freelist, and the next
open refused the session it had just migrated. The v1 table is renamed and
frozen instead, and reads take both generations. Table creation also moves
inside the migration transaction, so a crash can no longer leave a v2-shaped
database still reporting version 0 for an older build to write into.

* fix(agent-session-journal): stop an empty provider transcript retiring the repair marker

A transcript that exists but decodes to zero messages was imported as a
success: the import published an empty `legacy_import` replacement that
deleted the `unreconcilable_prefix` anchor and its disclosure, so the next
probe read the session as clean and every later attach skipped provider
recovery while the user's rows sat in quarantine for good.

The import now leaves the epoch untouched when nothing decodes, reporting
`replaced: false`, and recovery treats that like a transcript it could not
read — the marker stands and a later attach with real history rebuilds the
timeline.

* style(agent-session-journal): merge the duplicate journal-database-space import

* refactor(agent-session-journal): drop quarantine, byte bound, blob spill and rate limit

Match what comparable implementations do: the journal is an unbounded
append-only SQLite log with no side tables and no admission control.

Corruption: the rejected suffix is DELETED rather than copied into a
quarantine table. The load still reports `corrupt` and recovery still
rebuilds the epoch from provider history, so the observable outcome is
unchanged — only the preservation half is gone. The schema is back to one
version with two tables; no v1 database exists outside unmerged commits of
this branch, so the rekey migration and the two-generation read path go with
it. Sequence-1 epoch anchoring and the empty-provider-transcript retry are
kept: both are about the corrupt signal being correct.

Size: no `maxSessionBytes`, so no page-cost arithmetic, reclaim band,
incremental vacuum, lifecycle byte reservations or `journal_bound_exceeded`.
`auto_vacuum` and `wal_autocheckpoint = 0` existed only to make a
transaction's physical cost predictable for that charge; with the charge gone
SQLite's default checkpointing is what the journal wants, and the explicit
pre-close checkpoint is redundant with the one `db.close()` performs. WAL,
`synchronous = FULL` and `busy_timeout` stay.

Payloads: an oversized body is truncated at the existing inline cap with the
existing marker and the remainder is discarded, bounded at the translation
layer that already calls these helpers. The truncation point and message do
not change; the content-addressed blob directory and all digest tracking do.

Rate: no `maxAppendsPerWindow` and no `journal_rate_exceeded`.

`JournalPayloadLimits` is now just the inline cap.

* fix(agent-session-journal): mark a partial repair pending and bound multi-block tool input

A repair that keeps its prefix had nothing durable to show for the suffix it
deleted: a sequence gap costs no malformed row, so no disclosure is appended,
and the surviving rows keep their epoch anchor. The next probe read a
contiguous anchored prefix, called it clean, and the deleted stretch of
timeline was never asked for again — silent loss, with the deletion already
committed. The deletion now writes a `journal_repairs` marker in the SAME
transaction, and replay keeps reporting corrupt while it stands. It retires
under exactly the rule the emptied-epoch anchor takes: a fresh epoch carries
the rebuild, or the session writes content of its own past the sequence the
repair left free. The repair's own disclosure is not that content.

Legacy import bounded a tool call's input only when it was the message's sole
block; the multi-block path returned `tool-call` unchanged, so a mixed message
from Claude, Grok or an omp execution cell persisted the whole input despite
`inlineHeadBytes`. `boundBlock` now routes it through `boundToolInput`.

Also drops canonical comments describing quarantine, snapshot files, blob
storage and blob compaction — none of which exist any more.

* fix(agent-session-wire): stop awaiting the synchronous journal probe

loadJournal runs on a sync-database connection and returns JournalLoad | null, so both wire call sites were awaiting a non-Promise. The type-aware code-quality gate flags it; the native gate does not.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-04 15:09:23 -07:00
Brennan BensonandMerge Sim 872bd51d47 fix(native-chat): reland large structured command results (#17720)
* fix(native-chat): preserve large structured command results (#17707)

* fix(native-chat): preserve large structured command results

* chore: place native chat validation artifacts under docs

* chore: drop stale root package config

* fix(native-chat): enforce rebuilt lifecycle append slots

---------

Co-authored-by: Merge Sim <sim@local>

* chore: omit native-chat reland planning docs

* fix(native-chat): remove journal store import cycle

* fix(native-chat): keep journal factory acyclic

---------

Co-authored-by: Merge Sim <sim@local>
2026-08-31 13:12:19 -07:00
Brennan Benson 894ed75abb Revert "fix(native-chat): preserve large structured command results (#17707)" (#17719)
This reverts commit 5fe37729ea.
2026-08-31 12:34:49 -07:00
Brennan BensonandMerge Sim 5fe37729ea fix(native-chat): preserve large structured command results (#17707)
* fix(native-chat): preserve large structured command results

* chore: place native chat validation artifacts under docs

* chore: drop stale root package config

* fix(native-chat): enforce rebuilt lifecycle append slots

---------

Co-authored-by: Merge Sim <sim@local>
2026-08-31 12:34:03 -07:00
Brennan Benson fd9125ea8c feat(native-chat): Codex structured native chat restructure (#16729)
* feat(native-chat): port structured Codex sessions from restructure-recovery

Rebuilds the desktop structured native-chat implementation from
brennanb2025/native-chat-restructure-recovery (tip 4e31c08db3) on top of
current main as a single commit, scoped to the local Codex path.

Ported:
- Structured agent-session core: durable record store + single-writer lease,
  canonical journal, agent-session wire host/attach/eviction/subscribers,
  `agentSession.*` RPC surface (registered via ALL_RPC_METHODS; host-side
  mobile allowlist included for wire compat), pty write gate, transcript
  additions, and the Codex app-server adapter/launch resolution.
- Renderer: NativeChatStructuredSession view/composer stack, structured
  launch path with the single-flight guard, local structured session tabs
  sync, activation gate + structured inventory (read-only
  `agentSession.handoffStatus` probe), agent-session tabs in the tab strip,
  AI-vault structured session activation, and the settings pane with the
  parent Experimental Chat UI toggle plus the nested "Use updated structured
  native chat" toggle. New sessions require both flags, agent codex, no
  prompt, and a local non-WSL, non-Windows-host execution host
  (structured-native-chat-availability).
- Fixes 72c013cea6 (verified Codex launch recovery), 8ddbaf5e3d (defer
  native terminal view switching affordances), and 4e31c08db3 (release the
  launch gate after a visibility retry) with their regression tests,
  including the third-launch-after-retry guard case.
- Cross-version agent-session wire test + CI lane, packaging entries
  (proper-lockfile, agent-tooling asar excludes), and the wire-compat doc
  section.

Deliberately not ported: mobile/ changes, the Claude structured runtime
(only the claude-transcript-branch-proof and claude-structured-owner-identity
leaf modules remain, backing the kept TUI-recovery arms), the terminal↔chat
adoption/handoff flow (`agentSession.adoptTerminal`/`requestHandoff`, the
handoff request engine, TUI adoption machinery, orca-runtime adoption
methods), renderer switching affordances and their dead leftovers, the
hook/subagent-status refactor cluster, and unrelated branch changes. The
crash-during-acquisition recovery path (restart handoff adjudication,
restore/reverse re-acquire, lease schema handoff keys) is kept because every
plain direct launch depends on it; a trimmed handoff coordinator exposes
only status/restore/close.

Branch edits that targeted files main has since split (ipc/pty.ts,
worktrees.ts, rpc/methods/terminal.ts, useIpcEvents, pty-connection,
store/slices/terminals.ts, runtime-types, web preload) were re-applied to
the split modules, preserving main's newer logic (Windows CIM fallback,
browser tab close rework, cold-restore resume flow, dispatcher threading).

Known seam: the mobile clipboard image-provenance CONSUMER gate ships
(agentSession.send refuses unproven mobile image refs with
agent_session_image_untrusted) but the producer hunk in
rpc/methods/clipboard.ts stays with the unported mobile cluster, so mobile
image sends into structured chat fail closed until that side ports.

* fix(native-chat): trust only authenticated local image uploads

* fix(build): preserve Windows process-tree patch application

* test(windows): include process creation time in addon fixture

* fix(build): run windows-process-tree node-gyp from the physical package dir

gyp expands the node-addon-api dependency by probing node, whose cwd
resolves to the package's physical directory in the store, so the emitted
target is a store-relative ../../../../node-addon-api@... hop. gyp then
resolves that hop against the rebuild cwd; from the node_modules
symlink/junction it escapes the store and configure fails with
"node_addon_api.gyp not found" (run 32999886072).

Rebuild from realpath(package dir) so both bases agree, matching how the
package manager itself runs native install scripts. The regression test
replays gyp's expansion+resolution against the planned cwd and fails
without the fix.

* fix(native-chat): keep chat tabs visible through terminal closes and empty-worktree launches

Two proven blockers in the native Codex tab contract:

closeTerminalTab pre-empted the canonical unified close. With one terminal
left it deactivated the worktree on a terminal/editor/browser-only check,
blanking a workspace that still held a renderable agent-session tab; with
two or more it pre-picked a successor from terminal entities only,
re-stamping the group active before closeUnifiedTab's MRU/neighbor repair
could land on the chat tab. Successor choice now defers to the unified
contract whenever the terminal has a unified row, and deactivation is
gated on the unified renderable count (matching leaveWorktreeIfEmpty),
with the legacy pre-pick kept only for terminals without a unified row.

A structured session created on an empty worktree was published into the
host's headless group while preserveLocalLayout froze the local layout,
leaving the tab in store but permanently off screen. A preserveLocalLayout
owner now always takes client-owned placement — repairing a rendered
leaf whose group record is missing, or materializing a rendered group on a
truly empty worktree — and applies the client-derived layout repair while
still rejecting host-authored layout.

Regression tests drive the real store through closeTerminalTab (git
worktree and folder workspace) and the real snapshot applier for the
empty-worktree adoption states; all fail without the fixes.

* fix(native-chat): close stale turns and retry rejected sends

* fix(native-chat): retire hosted rows on structured tab activation

* fix(native-chat): preserve rpc defaults across main merge

* chore: format remote wire compatibility guide

* test(native-chat): cover retry after unconfirmed send

* fix(native-chat): reload outbox on session switch

* docs(settings): disclose structured chat platform limits

* fix(native-chat): await Codex launch-home preparation

* fix(codex): align child-process allowlist with async trust bridge

* test(identity): update inventory for tab surface refactor

* fix(windows): preserve process-tree CRLF patch sources

* fix(native-chat): anchor an unmatched chat echo where it was sent (#16117)

* fix(native-chat): anchor an unmatched chat echo where it was sent

The reported symptom was old user messages replaying below every new turn, so the
conversation read as scrambled. The cause was not that the echo failed to match a
transcript row. Claude consumes a mid-turn send through a `queued_command`
attachment and writes no `type:"user"` record for it, so some echoes can never
match, and no amount of matching will change that. The cause was WHERE an
unmatched echo rendered: buildMobileNativeChatTransientData appended every pending
item after the entire transcript, so it re-read below each turn that landed
afterwards.

Render each echo directly after the transcript row it was sent against, using the
baseline the send already captures. An unmatched echo is then at worst a duplicate
in the right position rather than a scrambled one, and it stays visible. Echoes
sharing an anchor keep send order; a send with no baseline, or one whose anchor
folding dropped, still falls back to the tail.

Deliberately NOT fixed by deleting the echo. Inferring from send ordering that an
echo can never match, then removing it, loses the user's own text for a message
the agent did receive, and it cannot fire in the common case anyway - measured
drain groups are 1,017 of size 1 against 55 larger. It also escalates an existing
gap: the count pass has no baseline-tail guard, unlike the glue pass, while
`messages` is a 40-row window that head-trims, resets on reconnect and grows at
the front on loadEarlier, so a false landing there would license deleting a
DIFFERENT outstanding message.

That count-pass gap is real and left for a separate change; anchoring makes its
worst case a duplicate in place rather than a scrambled conversation.

* fix(native-chat): preserve folded echo anchors

* fix(native-chat): preserve forward-folded echo anchors

* fix(native-chat): keep leading folded echoes in place

* fix(workspace-cleanup): show git status for every row (#16690)

* fix(native-chat): refuse structured chat on every Windows execution path

canUseStructuredNativeChat only refused win32 when a project runtime
resolved, so folder-workspace keys (and other keys with no project
runtime) failed open into structured chat on Windows. Fail closed on
win32 unconditionally after the host check, matching the settings copy:
local macOS/Linux only; Windows/WSL/SSH stay on terminal chat.

* fix(native-chat): restore runtime refusals behind the win32 gate

506d375de3 replaced the project-runtime checks with a bare platform test,
so a WSL or repair-required runtime resolution would no longer refuse
structured chat off-win32. Keep the unconditional win32 refusal and
re-run the runtime resolution after it, so the gate does not depend on
the resolver's own platform guard. Tests inject WSL and repair-required
resolutions on darwin/linux and fail against the regressed gate.

* fix structured session journal durability

* fix structured tab active pointer after restart

* fix(native-chat): await optional lease renewal callbacks

* refactor(skills): extract install error messages

* fix(agent-session): harden recovery ownership

* fix(native-chat): retain panes across tab activation

* fix(native-chat): address round-one review findings

* test(native-chat): align integration coverage after main merge

* fix(native-chat): harden round-two reliability

* fix(native-chat): harden round-three reliability

* fix(native-chat): close round-four recovery gaps

* fix(native-chat): separate bounded journal key forms

* fix(native-chat): reset outbox error in render on session switch

The switch effect adjusted error state after the sessionId prop changed,
tripping react-doctor's no-adjust-state-on-prop-change on the changed-code
gate and flashing the old session's banner for a frame. Reset it with the
render-time previous-value guard instead.

* fix(native-chat): invalidate stale outbox settlements

* test(native-chat): restore settled-error session-switch regression

a6e2379bd1 replaced this test with the in-flight settlement race test,
leaving the render-time error reset unpinned: deleting the reset block
still passed the whole native-chat suite. Keep both scenarios pinned;
they are distinct (settled error clears on switch vs stale settlement
invalidated in the commit-to-passive window).

* test(wire): make release checkouts race safe

* test(wire): pin cross-process checkout single-flight and importer specifier contract

* test(wire): harden release checkout lifecycle

* fix(build): drop CR-byte residue from windows-process-tree patch

The two trailing CR bytes on the patch's deletion lines are a proven
no-op: pnpm hashes patches CRLF-normalized (both forms hash to the
lockfile's 946ffb2b) and materializes this package without applying the
patch in either form, so the load-bearing build edits come solely from
applyWindowsProcessTreeBuildFixes() (#16947), which handles both source
EOL forms. Restore byte-identity with main and repin the contract test
to the post-#16947 reality: LF-only patch bytes plus lockfile hash sync.

* fix(native-chat): skip empty startup recovery
2026-08-28 16:45:58 -07:00