Files
orca/src/shared
Brennan BensonandClaude 6ae6ed08bb fix(claude): open structured chat without a startup deadline, and make Retry start fresh (#22364)
* fix(claude): open structured chat without a startup deadline, and make Retry start fresh

Publish the Claude session as soon as its process is spawned instead of racing
initialize against a fixed 10s deadline. Prompts sent before startup lands are
held and written in order once it does. An exit or sign-in failure before startup
ends the session with the reason and the CLI's stderr.

A create that failed because the process provably exited now carries
ownerVerdict 'exited', so the client marks the launch failed and Retry mints a
new operation instead of replaying the stored failure.

* fix(native-chat): sending into a chat that failed to start restarts it

* fix(native-chat): a send with no live owner restarts it once

A provider child that timed out or exited hands its lease back, and every
later send was refused agent_session_ownership_unknown. Clients read that
code as "not admitted yet" and resend forever, while only a surface hold
could make a new child, once per mount, with its failure swallowed.

The send now routes to a live owner, otherwise restarts one from the
persisted resume state where resume eligibility allows it (single-flight
per session), otherwise refuses with the new settled
agent_session_owner_unrecoverable. Unverifiable, reserved and handed-off
leases are left alone. The desktop hold now logs its failure.

* test(native-chat): pin the unrecoverable refusal as settled in the outbox

* test(native-chat): pin the release clock after a send restarts an unheld owner

* test: read the sent operation id without a cast

* fix(native-chat): type the send-recovery record lookup as the store returns it

* fix(native-chat): a send ensures its owner before admission, and an unheld owner idles for 30 minutes

* fix(native-chat): a create that throws releases its event sink

A child that dies between spawn and journal attach can still write through
the host's event sink, which attach unbound in onAcquiring and never re-bound
because onAttached never ran. The orchestration released that sink only when
performAttach returned a refusal; a thrown failure (the root-exit path) kept
the sink cached with its queued write, so the next attach's drain barrier and
runtime shutdown's flush waited forever.

Also pins the publish-on-root-exit clause for a start that never proved:
deleting it reddened nothing before.

* fix(native-chat): a resend the journal answers restarts nothing, and a send joining a restart rebases from the fence it replaced

* fix(native-chat): the host learns a Claude start positively, and persists only proven options

A publish-first create used to read the session's options before Claude had
answered initialize. With startup pending that read fell back to the built-in
catalog's default, so `record.options.model` was persisted as `sonnet` for
every user whose CLI default is something else; an owner handoff or a reopen
then replayed `set_model('sonnet')` and silently switched their model.

The adapter now reports `started` once startup facts are applied and saved
options restored. The host keeps a `providerChildPhase` on the session it
owns: a starting child hands over nothing but the saved options as intent,
and the `started` event re-reads the options as fact and persists them through
the same record write a user's option change takes. The status summary carries
`hostExecutionPhase` (optional, wire-safe), and the chat pane says the agent is
still starting instead of showing nothing.

A child whose exit already reached the adapter before acquire returns is no
longer handed over as live; the create fails with the CLI's diagnostic.

* fix(native-chat): a hold and a send that find the owner gone share one restart, and a send the ledger already holds restarts nothing

* fix(native-chat): a failed create answers one refusal shape, stamped once at the boundary

A create whose Claude process was seen to exit answered twice in two shapes:
the first call threw a generic runtime error, and only the replay of the same
operation carried the `ownerVerdict: 'exited'` refusal that lets a client
retry under a new operation. Three sites stamped the verdict and the store
failure path stamped nothing.

The first-hand root exit is now returned as the refusal on the first call,
with the provider's own diagnostic as its message. The verdict is stamped in
one place, at the boundary of the attach, from the durable row the operation
settled to, so every refusal shape answers the same fact and no site can
forget it. The per-site stamps are gone.

* fix(native-chat): a send into a session whose child ended restarts it before admission

A session that published and then lost its Claude child before startup (not
signed in, for one) keeps a released lease and a chat the user can still type
into. The send was refused as ownership-unknown, the outbox parked it as
pending admission, and nothing ever restarted the child: the message sat
there until the user closed and reopened the tab.

A send reaching a session with no provider child now runs the same resume a
surface's first hold runs, before the write is admitted. The resume reserves
a new fence, so that send is answered stale with the published fence and the
client's outbox re-drives under it, as after any fence change. A resume that
fails is not this send's answer; admission reports the lease as it stands.

* chore: restore pnpm-lock.yaml to origin/main (local pnpm rewrote it)

* test(native-chat): pin the pre-handover exit as a failed acquire; stub the status feed in the delivery test

An exit the adapter observes before acquire returns now fails the acquire
with the CLI's diagnostic instead of handing over a dead child; the
published-then-ended path stays pinned by the slow-init startup case. The
delivery test renders the pane, which now activates the host status feed.

* test(native-chat): a same-ID re-hold over the wire joins the one resume, and a replay reopen goes on the idle clock

* test(native-chat): a re-hold that joins a failing resume proves one resume ran

* fix(native-chat): a create whose child was proven gone answers the refusal on the first call

The previous change answered a first-hand root exit as the exited refusal on the
first call, but the common failed start never took that path: when the close
ladder proves the whole tree dead the acquisition error is a plain one, the
store-failure classifier rethrows it, and the client still saw a runtime error
first and the refusal only on replay.

The cleanup that proves the child gone now names such a failure
`AgentSessionAcquisitionExitProvenError`, carrying the provider's diagnostic,
unless it already names its own verdict (a refusal, a typed exit proof, a host
store code). The attach answers both proven-exit kinds as the refusal its replay
gives. How a failed acquisition settles and how it is first answered now live
beside the verdict stamp, in the failed-create module.

* test(native-chat): pin the outbox re-driving a stale-refused send under the resumed fence

A send into a session whose child ended is answered stale once the host has
restarted the child. The outbox keeps that operation queued and blocked, and the
fence change the resume publishes re-drives the same operation under the new
fence; the host admits it.

* fix(native-chat): a child restarted for a send nobody holds is still released

The restart a send runs for a childless session takes no holder, on the premise
that the sending surface already holds one. A one-shot writer holds nothing, so
the child it restarted had no release clock and lived until the app quit. The
write resume now arms the clock when no holder is present, as the first-hold
resume already does. The send-after-failed-start cases also pin that the stale
answer's operation is admitted when re-sent under the new fence, and that two
racing sends restart the child once.

* test(native-chat): pin the picked Claude model across a resume whose child starts on its own default

The started event re-reads and persists what the child reports. A resumed child
answers initialize with its CLI default before the saved pick is restored over
it; the record must hold the pick while starting and after started.

* Revert "fix(native-chat): a child restarted for a send nobody holds is still released"

This reverts commit e52c4a6f08.

* Revert "test(native-chat): pin the outbox re-driving a stale-refused send under the resumed fence"

This reverts commit 136a39deb0.

* Revert "fix(native-chat): a send into a session whose child ended restarts it before admission"

This reverts commit 39234e44bf.

* refactor(native-chat): make ensure-owner a step of the serialized send

A send that found the owner gone restarted it OUTSIDE the host's per-session
serialize, through a single-flight resume map shared with the surface hold, then
rebased its fence by heuristic. The attach body is now callable from inside
`serialize` (`attachStructuredAgentSessionUnderSerialize`), and every restart
runs there: a hold, a send's ensure-owner step, provider-exit recovery and the
rewind owner replacement take turns on one queue, so the first to run attaches
and the next finds its child. The single-flight map and `isResuming` are gone.

Admission is two-phase for a send: the ledger's answer comes first and places
nothing; a send it will admit gives the session an owner, and only then are the
row placed and the lease and fence checked. A send it will replay into a closed
session makes the journal readable and spawns nothing. The session entry
records the released fence the child replaced (`resumedFromFence`), so a writer
current as of that owner is admitted at the new fence by bookkeeping, whether it
ran the restart or arrived behind it.

The resume reads its record only after this host has reconciled it and exited
any recovery stage a failed attempt latched, so a hold behind a failed attempt
makes its own attempt against the lease as it now stands.

* fix(native-chat): a Claude start proving itself no longer waits on the CLI

The host handles a Claude child's `started` on the recovery chain every
session's unexpected-exit handling shares, under that session's serialized
step. It then asked the CLI for the model list and settings again, so one slow
CLI held every other session's exit recovery, and its own close, behind up to
two request timeouts.

The adapter already holds those answers when startup proves: the settings read
at startup, the restore's confirmations, and the initialize result the SDK
answers the model list from. `started` now carries that snapshot, and the host
turns it into one record write without any provider I/O.

* test(native-chat): a hold reads its lease only after this host has reconciled it and exited a latched recovery stage

* fix(claude): a chat whose first start failed resumes as the same conversation

A Claude start that dies before initialize writes no transcript, so the next
start launches the chain head's provider id fresh instead of `--resume`. The
launch flag that chose that mode also chose the provider-handle link's origin,
so the fresh launch published a second `created` link onto a chain that
already had a head. The store refused it, the healthy child was closed, and
every later reopen, hold or send spawned and killed another Claude.

The launch now carries the two facts separately: `resumesTranscript` (launch
mode, from whether Claude wrote a transcript) and `continuesChain` (lineage,
from the record's chain head). The link origin reads lineage; rewind and the
Fast opt-in carry-over read launch mode.

* refactor(native-chat): a resume answers with a typed refusal the send classifies

`resumeHeldStructuredAgentSession` and the holds' `ensureProviderChild` answer
`{ ok: true } | { ok: false, refusal }` instead of throwing the refusal code.
The refusal is the attach's own, with its message and, when the failed attach
proved its child gone, its `ownerVerdict`. An attach that settles a failed
acquisition in the ledger and then rethrows the cause is read back off that row,
so a durably failed restart is a refusal and only an unrecorded error is a fault.

The send classifies the refusal through a `Record` over every wire code — a new
code does not compile until it is placed — into transient (the lease is someone
else's to settle; the send runs as the lease stands) or terminal. A terminal
one answers `agent_session_owner_unrecoverable` carrying the cause, forwards the
verdict, and writes the same status row into the chat that a start that failed
leaves, so the user sees why after the error strip is gone. Nothing about the
failure is remembered; a Retry is a fresh attempt. A fault thrown by the restart
itself is reported and the send runs as the lease stands, since bookkeeping
never gates a user's action.

`hold()` still raises the refusal code for its RPC caller.

* fix(native-chat): a child's event sink belongs to the attach attempt that spawned it

The runtime kept one event sink per session id and handed it to every attach.
An attach that acquired a new child unbound that sink first, so when the
acquire then failed its dead child's queued frames stayed in the cached,
unbound sink. The earlier guard only discarded it when no session entry was
left, which a resume of a still-indexed session never satisfies: the next
attach's drain and shutdown's flush waited on it forever. A TUI-to-native
handoff acquire had the same shape.

Each acquiring attempt now mints its own sink. Only a successful attach (or a
proven handoff owner) adopts it as the session's, closing the one it
replaces; any other exit closes it with whatever its child queued. A re-attach
to a live child keeps the sink that child already writes through. A sink that
is not the session's own can no longer force the session's provider down.

The native handoff acquisition moves to its own module, which keeps the
handoff file under its line budget.

* test(native-chat): pin that only the adopted child's event sink still takes writes

Closing a failed attempt's sink and closing the sink a resume replaces were both
unpinned: removing either left every suite green, because neither sink is in the
map that drains and flushes read. The resume test now asserts the failed
attempt's sink and the exited generation's sink refuse writes, and the adopted
one accepts them; deleting either close reddens its own assertion.

* perf(native-chat): the chat reads only the host's startup phase from the status feed

The chat took the whole status summary to read one field, so every status change
for its session (prompt, update time, background tasks) re-rendered the chat
view. It now subscribes with the phase itself as the snapshot, so it re-renders
only when the phase changes.

* fix(native-chat): the startup-phase hook answers a phase or null, never undefined

* fix(native-chat): every restart is counted from the moment it is asked for, and a handoff clears the restart fence

Provider-exit recovery now restarts through the holds' `ensureProviderChild`
like a hold and a send do, so a child whose only surface left while the attach
ran goes on the idle clock instead of living until quit. A hold's resume and a
client attach are tracked as in flight from enqueue, not from their turn on the
queue, so a quit's drain waits for one queued behind a close before it decides
what to evict. A handoff back to native moves the fence in place and now clears
`resumedFromFence`: only a restart may rebase a writer. The failed-restart
status row is keyed by the send's operation id, not the clock, so a resend of
the same id that fails again adds no second row.

* fix(native-chat): a Claude start no longer waits behind another session's exit recovery

The runtime delivered every Claude lifecycle event on the single chain
exit recovery uses so teardown can drain it. That chain orders nothing
across sessions, and an exit recovery on it can run a full reacquisition,
so one chat's `started` waited on an unrelated chat's respawn and kept
its 'still starting' line up. `started` now takes only its own session's
serialized step, is queued the moment it is emitted (ahead of any later
exit of that child), and is tracked in a set the same teardown drain waits
on.

* fix(native-chat): a Claude create that dies at spawn is refused with the CLI's own diagnostic

A CLI that exited before its acquisition handed the child over was
refused with 'claude stream-json for session … exited while being
acquired', or with an unreadable start time, and the stderr the exit
carried (for example 'not signed in') appeared nowhere. The acquisition
now keeps the error its connection ended with and answers with it at
both sites; the generic message is only a fallback when none exists.

* test(native-chat): pin that a reopened Claude chat dying before initialize says why

A resumed start is published at spawn, so a CLI that exits before it
answers initialize fails a chat the user is looking at. Pin that the
open chat is sent the 'stopped before it finished starting' row with the
CLI's diagnostic even when the child's tree cannot be proven gone, and
that a message held for that start is refused rather than left in doubt.

* fix(native-chat): Stop while a Claude start drains its held prompts withdraws the rest

Stop withdrew held prompts only while startup was pending. Once startup
landed and the gate began writing them one by one, a Stop interrupted the
CLI and the prompts still waiting were written straight after it. Stop
now withdraws whatever the gate still holds in both states; the drain
takes each prompt off the queue immediately before writing it, so a
withdrawn prompt can never be written. The one already written still
gets the interrupt.

* fix(native-chat): the release clock keeps a session that still owes a sent message

When the last surface stops holding a session, the release clock evicts
it after the grace unless a turn is running. A message sent while Claude
is still starting is held, not running, so switching away from that chat
for the grace evicted the session and refused a message the user had
already sent. The clock now asks whether the session owes work: a running
turn, or a submission the provider has not taken yet (still pending in
the journal). Both are read from the journal; nothing new is stored. A
starting session that owes nothing is still released, and an explicit
close still ends everything.

* fix(native-chat): only a start that holds a sent message keeps a released session

The release clock kept any session with a pending submission. A Codex
send is admitted and stays pending until its echo, which may never come,
and only an eviction retires it, so such a session was never released
while the app ran. A pending send now keeps the session only while its
child is still starting, which is when the send is held for that start.
Pins that a ready session with an unechoed send is evicted, and that a
Claude chat whose turn finished is released after the grace.

* fix(native-chat): a send waits for the owner it met to prove its start before it is admitted

A Claude child is published before the CLI has answered initialize, so a send admitted right
behind a restart — or right behind the first start — was dispatched into a child that could die
milliseconds later, and learned of the death only as a delivery nobody could confirm. The terminal
refusal the send was written to give was unreachable on the real adapter for exactly the failure
it was written for.

The send's serialized step now admits nothing against a `starting` child. It registers for the
child's startup verdict and returns having placed nothing; the send waits off the session's queue
(the `started` and `ended` settlements run on it) and admits again once the child is `ready`, or is
refused `agent_session_owner_unrecoverable` with the child's own exit reason when it exits first.
The exit settlement writes the one status row, decided by the host's own phase rather than only
the provider's flag. A close, an eviction or a replacement answers the wait too, and quit releases
whatever is left; there is no timer. One spawn per user action holds across re-entries.

* fix(native-chat): restart the release grace when a start writes its held prompts

A prompt held while Claude starts is written when the start lands, but
its turn opens only when Claude echoes it. The release clock stopped
counting it once the child read ready, so a tick landing in that gap
stopped the child before it ran the user's first message. The start
landing now restarts a pending release's full grace, the same grace a
message sent to a ready chat gets before it is released.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): a starting child owns the send; the adapter holds the message for its start

A send that meets a child still proving its start is admitted against it, as it was before the
off-queue startup wait: the adapter holds the message until startup lands and rejects it with the
child's own diagnostic when the child dies first, the exit settlement writes that cause into the
chat, and the release clock keeps a starting session that holds a sent message. The startup watch,
the off-queue wait loop and their teardown phase are gone; the exit settlement still reads a start
that failed off the host's own phase when the provider omits the flag.

The scripted-CLI test now pins that contract end to end: a restart a send asked for whose CLI dies
at initialize leaves the message rejected with the diagnostic, one row naming it, and the fence
moved by two; a healthy CLI is restarted once and written to; a send during the first start is
held and written once initialize answers, or rejected with the diagnostic when the CLI dies.

* fix(native-chat): a failed send restart says why, and offers a new chat only when nothing can restart it

The refusal a send gets when the host cannot restart the chat's agent is renamed
agent_session_owner_restart_failed and now reads "<Agent> couldn't restart: <reason>." with the
restart's own cause. "Start a new chat to continue." is added only when the resume was refused
because this host has no record to restart from or cannot run the one it has. Any other failure,
such as a CLI that is not signed in, leaves the chat retryable: the outbox stops auto-retrying, and
a manual Retry or a new send tries the restart again, since a refusal before admission leaves no
ledger row.

* fix(native-chat): a Claude start skips an option write the CLI never answers instead of faulting at the request deadline

* test(native-chat): wait for the recovery's reserved lease, not the released one it replaces at once

* fix(native-chat): a send whose restarted child dies before starting is rejected with the child's diagnostic

A child that never proved its start has accepted nothing: input is written
only after it initializes. A send admitted against such a child, whose
dispatch then found no session, settled unknown, twice, and the outbox took
Retry away. It now settles rejected with the child's own diagnostic, both
when the dispatch throws and when the exit settles the sends it left
unanswered, so the chat says why and offers Retry. A proven child's
unanswered sends stay in doubt, as before.

* test(native-chat): expect a send held for a start that never proved itself to settle rejected

* fix(claude): write a prompt held after startup already drained, instead of stranding it pending

* test(native-chat): pin that a send to a child that died before starting is answered rejected

* test(native-chat): leave the cast exit-session fixtures as they were, since a proven exit never rejects

* fix(claude): a saved option the CLI never answered stays saved instead of being replaced by the CLI's value

A start skips an option write the CLI does not answer within the request
deadline, and then persisted what the CLI reported in its place, so a slow
answer silently replaced the user's saved model or dropped their saved
permission mode. Silence is not a refusal: the unanswered option is now
recorded apart from a rejected one, the live child keeps running on the
CLI's value, and the saved choice stays on the record for the next start to
retry. An option the CLI rejects is still dropped as before.

* fix(native-chat): a rejected send opens no turn, so the row naming why it failed is not folded away

A send whose restarted child died before starting is rejected, and the
exit writes a row naming the cause. The chat's local clock had watched the
send go pending and stop, so it gave the message "Worked for 0s"; that
settled a turn that never ran, and the fold hid every non-prose row after
the message behind it, including the one naming the cause. The row only
appeared when a later send moved the turn anchor, which read as two rows
for one Retry. The host's journal already says the send was rejected; it
now answers that such a message opened no turn, which outranks the local
clock on desktop and mobile alike. A rejected send whose journal does
record a turn keeps its duration.

* fix(native-chat): a send whose restart died starting leaves the same row as any start that died

One failed attempt already leaves one row, but which row depended on when
the child died. A child that died after the send was admitted left "The
provider stopped before it finished starting: <cause>."; one that died
before the send was admitted left "Claude couldn't restart: <cause>." So the
same failure read two ways from one Retry to the next. When the refused
restart proved its child exited, the send now writes the startup-failure row
itself, as its comment always said it did. The refusal under the composer
still says the restart failed; a restart that failed for a reason other than
a child exiting keeps its own wording.

* fix(native-chat): a send rejected because the agent never started names the cause under the composer

When the child a send was admitted against died before starting, the host
rejected the send with the child's diagnostic behind the internal transport
marker. The client rightly hides that marker's detail, so the red line read
"Couldn't reach the agent" while the cause sat in the record. A startup
death is not a failed write: the host now words that rejection the way the
chat row does, "The provider stopped before it finished starting: <cause>.",
at every site that rejects for it. Desktop and mobile show a reason in words
verbatim already, and older clients do too, so no client change is needed.
Real write failures keep the marker and the generic copy.

* fix(claude): a saved option the CLI never answered survives a later change to a different option

The saved choice a start could not apply was kept on the record, but the next
option the user set persisted only what the child had applied, so changing the
permission mode or effort, or clearing the chat, silently dropped the saved
model. The adapter now reports which saved options are still unanswered, every
option write keeps those saved values, and a write the child accepts for that
option retires it.

* fix(claude): a send that meets a child whose exit already settled names that exit's cause

When the child a send was admitted against died starting and its exit
finished settling before the send reached it, the send was rejected with
"no live claude stream-json session for <id>", now shown under the composer
as the cause. The adapter keeps a settled exit's diagnostic until the chat is
acquired or closed again, so that send names what the CLI said. A refused
restart whose child died at spawn or while its start time was read is pinned
to leave one row in the words any failed start uses.

* test(native-chat): pin the words an exit settlement rejects a never-started send with

The startup gate and the dispatch reject a send first in every existing
scenario, so the exit settlement's own rejection had no test of its wording.

* fix(claude): derive which saved options are still unanswered from what the child applied

A write that lands already puts its option in the session's applied set, so
the unanswered list is that list minus what has since been applied, rather
than a second copy every option write must remember to edit. Session
fixtures built without the new set no longer throw on an ordinary write.

* fix(native-chat): a cleared chat starts from a saved choice the child never answered

Clearing a chat seeded the replacement from the values the child reported,
so a saved model or effort whose restore write the CLI never answered was
replaced by the CLI's own value in the new chat, even though the retired
record kept it. The replacement now keeps those saved values too, and its
start retries them.

* fix(claude): closing a chat forgets its exit's diagnostic even when the exit settles during the close

The diagnostic was dropped when the close began, but closing over an exit
that was still settling finishes that settlement, which kept it again, so a
closed or deleted chat held it until its next acquire. It is now dropped once
the close finishes. Pins that an acquire and a close each retire it.

* refactor(claude): keep a saved option the CLI never answered as the wanted value, not a list beside it

A restore cleared the session's wanted options and added back only the writes
the CLI answered, so an unanswered one lost the user's value and every later
writer had to be told to put it back: the start report, each option change and
/clear each carried a list of unanswered keys. The restore now keeps the saved
value as wanted and unconfirmed, so what the session reports and persists
already carries it, and the list, its adapter method and the started-event
field are gone. A refused option is still dropped.

/clear now starts the replacement from the record's options instead of reading
the child's live values, which can be a model the CLI fell back to.

* test(claude): wait for the start to finish before changing the saved model

The record holds the saved model from creation, so waiting for it returned at
once and the option write could reach Claude while it was still starting,
which refuses it. Wait for the effort the finished start reports instead.

* fix(i18n): translate the still-starting chat notice

The notice that a structured chat is still starting was only in English.

* test(claude): pin the failed acquisition's own reading-control release

The merge re-pointed this test at a child that exits after publish, where the
exit path also releases the binding, so it passed with the acquisition's release
deleted. A child that exits before publish leaves only that release. Also drops
the create 'init' phase, which lost its last producer when rewind stopped
proving before publish.

* refactor(claude): move unexpected-exit handling into the exit lifecycle module

The adapter crossed the 300-line limit once main's context-usage change
landed beside this branch's growth. The two methods that turn a Claude
process exit into an ended event now live next to the existing exit
helpers; behavior is unchanged.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-24 15:23:15 -07:00
..
2026-09-03 17:32:59 -07:00