Commit Graph
12954 Commits
Author SHA1 Message Date
Brennan Benson acea59c6d8 refactor(native-chat): remove the records-file and per-chat journal imports (#26038)
* refactor(native-chat): remove the records-file and per-chat journal imports

Every native chat user is on a build that already moved chat records and
history into the app-wide database, so the one-time copies are dead code:
the agent-sessions.json import and its owed-copy flag, the per-chat
journal.db import and the write-queue hold behind it, and the pre-SQLite
log.jsonl notice.

* refactor(native-chat): remove what only the deleted imports used

- JournalHostDatabase.unsyncedTransaction and the synchronous-pragma reset
  that undid it; JOURNAL_SYNCHRONOUS is now module-local. The stranded
  rollback test drives transaction() instead.
- deleteUnpublishedJournalRows and its SQL, plus its test.
- boundJournalStatusText.
- Stale comments naming the removed first-use copy (queued messages,
  send refusal example, JOURNAL_SYNCHRONOUS doc).
- Write-queue and readInOrder comments: a write also lags when issued
  behind one still waiting in line.
- Resolved-append ordering test binds the sink from inside a running read,
  so a resolver that read at handover now fails it.

* test(native-chat): run resolved-append in the SQLite runtime project
2026-10-06 23:03:32 -07:00
Brennan Benson 6bc529176b fix(orchestration): word the worker brief for a chat worker (#26036)
* fix(orchestration): word the worker brief for a chat worker

* test(orchestration): prove the chat brief wording through the real send paths

* test(orchestration): read the redispatch paragraphs from the preamble module

* test(orchestration): normalise the chat/terminal wording in the mode-opacity check
2026-10-06 23:01:13 -07:00
Brennan Benson d0b0f13b74 fix(native-chat): hold a queued message until the turn ahead opens (follow-up to the Stop-event plan, fixes STA-9348) (#25217)
* test(native-chat): a Stop over a card sent now into the running turn keeps it paused

Red on main: Codex's turn end withdraws the steered hand-off and the queue
sends the card again as a new host turn, with no pause recorded.

* fix(native-chat): a Stop's queue pause holds a card whose hand-off is still unanswered

A card sent now into the running turn was still pending when Stop judged the
pause, so nothing was recorded; the interrupt then withdrew the hand-off, the
card went back to waiting unpaused, and the queue sent it again as a new turn.
The pause now counts a hand-off that may still return to waiting, judged with
the appended row applied, so a withdrawal lands under the pause and an
acceptance retires it in that same write. Codex and Claude both hit it.

* test(native-chat): the Claude re-send case fails on its diff, inside the test's budget

* test(native-chat): a pause held by an unanswered hand-off ends on every path that ends it

The provider's answer, the provider dying, the chat closing, a restart, a
withdrawal still owed at open, and a /clear (refused until the hand-off ends,
then carrying every waiting card paused 'cleared'); each ends with the queue
sending again.

* fix(native-chat): narrow the pause's settled hand-off, and assert the queued receipt's card

* refactor(native-chat): derive the queue's pause from Stop and Resume journal rows

Stop now appends one journal row where it takes effect, before the interrupt,
whatever the queue holds; Resume appends its own. The pause is a pure function
of the fold: the latest Stop with no later Resume and no later accepted turn a
person asked for. A /clear's carried cards name their source, which is the
replacement's 'cleared' pause. Host-origin turns never lift either.

One predicate decides which cards a pause holds; by default every waiting card
without a hold of its own, including one queued after the Stop. The drain's
consume re-judges it inside its own transaction.

The rows are tombstones of an id no item takes, carrying the mark: a released
host reads an unknown row kind as corruption and truncates the journal there.

Deletes the stored pause (recordPause, the retire hook on every appended row,
the settle-before-record step, mayReturnToWaiting and its row overlay) and the
tests that only proved it retires. The queued_message_pauses table stays in the
schema, unread and unwritten, for downgrade safety.

* fix(native-chat): a card queued after a Stop sends normally, never ahead of held ones

A Stop's pause now holds only the cards queued before its row, plus a steer it
withdrew, which returns to its own place. Each card records the journal
position it was queued at, and the one hold rule compares that with the Stop
row. A card queued after the Stop is a new instruction: it sends as usual, but
the drain still stops at the first held card, so it never overtakes them.
/clear's pause holds the cards it carried. Holding every card again is a
one-line switch in that rule.

* fix(native-chat): the queue's own send re-checks the no-overtake rule in its transaction

The drain's pick and its consume now read one function, nextSendableQueuedCard,
so a Stop row that lands between them holds a newer card behind an older held
one exactly as the pick would. Notes why Stop and Resume ride a tombstone row.

* fix(native-chat): stop creating the unused queue pause table

The queue's pause is derived from journal rows, so nothing reads or writes
queued_message_pauses. It was still created on every open "for downgrade
safety", but an older build creates it itself when it opens the database, so
the table only sat empty in every new database. The tests now pin that no
pause table exists.

* fix(native-chat): a Stop's pause never hides the restart pause

A Stop holds only the cards queued before it. The pause derivation still
returned the Stop alone whenever it was in force, so the restart pause was
never considered: a card queued after the Stop, written by a host process
that has since exited, sent by itself after Orca restarted, with no pause
header and no Resume. A /clear pause that held nothing could hide it the
same way.

Every pause in force is now derived. A card is held if any of them holds
it, and it names the first that does. The drain's pick, the consume
transaction's re-check and the published header all read that one rule;
the header names the pause holding the first card Resume would send.

* test(native-chat): pin the Stop's no-resend, lift and held-card rules

- The Claude and Codex Stop-withdraws-a-steer tests checked "not sent
  again" at one instant, before a queue ignoring the pause re-sends. They
  now wait for the stopped turn to end and re-check after a quiet window.
- The deleted-card test read a card queued after the Stop, which sends
  whether or not a person's turn lifts it; it now reads the Stop's pause
  before and after that turn.
- Unit cases pin that a Stop holds a card with no recorded position and one
  queued before a rewind.

* refactor(native-chat): a Stop writes one Stop event with its reason, turn and caller

The Stop row that paused the queue becomes the general Stop event
{ reason, turnId?, at, caller? }, whose reason is the host's existing stop
cause. It still rides a tombstone of a host-only id (a released host deletes
the journal from the first unknown row kind), and Resume keeps its own marker
on its own id. Only a person's Stop (reason user-stop) pauses the queue.

* test(native-chat): a rewind keeps a lifted /clear pause lifted and restates the same Stop event

* test(native-chat): pin that Stop and Resume rows never reach apps or count as history

* test(native-chat): only a person's Stop event pauses the queue

* test(native-chat): pin that a Stop's event precedes the interrupt and the at-start stop

Through the real host: the event names the turn and who asked and is in the
journal when the interrupt reaches the agent; at an agent still starting it is
there before the start is ended and holds a card queued before it; an idle Stop
writes one only when it withdrew a send; and the queue's claim re-judges a
pause that landed after its pick.

* test(native-chat): a card held at a starting agent is checked before the Stop's timing

Also says precisely what the claim's in-transaction pause check defends
against: the Stop and the drain share one serialized lane.

* test(native-chat): a released build keeps and folds a journal holding Stop events

Replays this build's rows from the released build's own journal database: every
row is kept, the history after the Stop still folds, and an older client is sent
only removed ids no item uses.

* style(native-chat): format the Stop event changes

* test(native-chat): type the released build's exports through one checked helper

* fix(native-chat): the Stop/Resume row guard narrows to those tombstones only

* test(native-chat): run the Stop-event downgrade test in CI, and cover a writable downgrade

The Stop-event downgrade test ran in no CI lane: unit shards exclude the
cross-version folder, and the cross-version lane runs a fixed file list that
did not name it. It is now on that list.

Its only case replayed the rows into a release's own fresh database, because
that release cannot open the current host database. A second case opens the
journal this build wrote with a main build that shares the database: it opens
writable, keeps every row, appends, and this build then reopens it with the
person's Stop still pausing the queue.

* fix(native-chat): a Stop that stops nothing new writes no Stop event

A Stop reaching a running agent wrote a Stop event on every press. Two
presses before the first interrupt landed wrote two events, so a card
queued between them counted as before the latest Stop and was held,
though a card queued after a Stop should send normally. A Stop naming a
turn that had already ended, as a phone sends late, also wrote an event
for a turn it never stopped.

It now writes one only when it withdrew a queued send, or stops something
no event records yet: not a turn the journal no longer runs, and not the
live turn a Stop still in force already names, unless a card was handed
over into it since, which this Stop's interrupt sends back and must hold.
The interrupt and the "already finished" note are unchanged. A Stop at a
starting agent still always writes.

* test(native-chat): pin that a later host, eviction or close Stop never lifts a person's Stop

* chore(native-chat): put each Stop-row doc on its own declaration, and say only user-stop is journaled

* fix(native-chat): any later Stop event ends a person's Stop pause

A person's Stop paused the queue until their next accepted turn or Resume,
and a later Stop of another reason (the host stopping the agent, an
eviction, a close) was ignored. Now the pause is the latest Stop event's:
a later Stop of any reason ends a person's pause, and only a person's Stop
pauses. The fold keeps the latest Stop event whatever its reason.

An eviction of a resting chat writes no Stop event (a Stop that stops
nothing writes nothing), so it cannot release held cards; a test pins that
no event means no lift.

* fix(native-chat): a second Stop press is a repeat even when the first came before the turn showed

A Stop pressed before the agent's turn shows in the journal (before
Claude's echo, or before Codex opens the turn) records no turn. A second
press once the turn showed compared that missing turn with the live one,
wrote a second Stop event, and held a card queued between the presses.

A repeat is now judged by what was sent since the Stop in force: with
nothing sent after it (a refused send aside), a Stop that named no turn,
or named the live one, is repeated and writes nothing. Anything sent since
and not refused, including a send whose fate is unknown, makes the new
press write, since its interrupt may send that card back to waiting.

Tests: the two-press case across the turn showing; a steer between the
presses settled unknown; and a Stop naming a turn that ended while the next
card is sent but shows no turn yet, which writes and holds that card. The
fold test that claimed an eviction path is renamed.

* fix(native-chat): the queue's pause ignores a Stop or Resume row holding a value no build writes

A Stop or Resume row's value is read from disk with no shape check, and
the pause fold stored whatever it found. A stored `stopEvent: null` would
then throw on every pause check for that chat: the queue's pick, its
send, and every queue update to clients. No build writes such a row, so
this is hardening.

The fold now reads a Stop only when it is an object with a string reason
and a finite time, and a Resume only when it is `true`. Anything else is
ignored: it pauses nothing and ends nothing. The row is still not treated
as malformed, which could cut the history short.

* fix(native-chat): a Stop still reads as yours after Orca restarts before the turn ends

Every stop that ends work now writes the Stop's event before it ends the child: a
person's close of the chat, an eviction (worktree teardown, orchestration stop, tab
cleanup) and the idle sweep's stop of a start that never landed. A stop that ends
nothing writes nothing, and quit writes none: its resume marker records why.

The turn-end write reads the latest Stop event where every turn row is built, so the
adapter's settle, the host's fallback and the relaunch's settle all agree: a turn a
person's Stop or close named, ending with no verdict of its own after that Stop, ends
as their cancellation. A relaunch's probe-bounded end is no earlier than a Stop that
found the turn running. When the provider refuses the interrupt and the turn runs on,
a refusal row answers the Stop, so a later crash still reads Failed; pressing Stop
again after a refusal is a new Stop.

* refactor(native-chat): a stop no longer carries its cause; the turn's end reads the Stop event

The cause of a stop was threaded in memory from each entry through the host's stop
step, the adapter router and each adapter's close onto the `ended` it settled with,
and Claude kept a per-turn copy of a Stop it sent. All of that is gone: adapters
settle a turn they cut as interrupted with no verdict, the host's fallback does the
same, and the one rule where a turn row is built (`turnEndAfterStop`) reads the
journal's latest Stop event to say whether it was a person's.

- `closeSession` / `disposeSession` take no cause; `ended` has no `stopCause`.
- Claude reads an error result after a person's Stop as their cancellation from the
  journal's Stop event (through the event sink), not from a per-turn slot, and a
  refused interrupt is the host's refusal row, not `withdrawTurnStop`.
- An owed wind-down keeps no cause: its retry's fallback reads the Stop event.
- The mutation context's Stop passes no cause: its step already wrote the event, and
  the delivery loop's child-end reason is read back from it.
- A Stop pressed before its turn showed applies to the turn that opens under it,
  unless a send a person made since was accepted.

* test(native-chat): a turn a later send opened is no Stop's that named no turn

* test(native-chat): the restart test's death proof carries its detail

* refactor(native-chat): a refused Stop leaves no record; a Stop only ever ends the turn it names

The stop-refused mark is gone: its tombstone kind, its fold, the clock-keyed match that tied it to
a Stop, and the exception that let a second press after a refusal write a new Stop. A Stop that
stops nothing writes nothing. A Codex refusal names a turn that is no longer its active one, and
the Stop names that turn, so the turn running instead never reads as the person's by its id alone.

* fix(native-chat): a Stop pressed before any turn showed stops only the turn opened next

A Stop that named no turn read as the person's cancellation for every later turn that opened
after it, until a send a person made was accepted. The queue's drain, orchestration mail and a
restart continuation send as the host, so a turn they opened long after, cut by a crash, read
"Interrupted" as if the person had stopped it. The Stop now applies only to the first turn
opened after it.

* fix(native-chat): an older Claude's error end after a Stop pressed before its echo reads Interrupted

Claude CLIs before 2.1.91 end an interrupted turn with an error result that names no reason. The
translator judged whether a person's Stop explained it by its own copy of the Stop rule, which
ignored a Stop that named no turn, so a Stop pressed before Claude echoed the send read "Failed".
The translator now writes such an end as interrupted with no verdict and no error row whenever a
person's Stop may name the turn, and the journal's one rule decides as it writes the end.

* fix(native-chat): a person's Stop and /clear each name why they end the agent

The host's mutation path ended the agent with one "recorded" ending for every caller, which read
back the reason of whatever Stop event the journal held last, however old. /clear writes no Stop
event, so its end took an unrelated earlier reason. Each caller now names its own: the chat's Stop
`user-stop`, whose event its own step wrote, and /clear `user-close`, the user replacing this chat.

* fix(native-chat): a host stop judges whether it ends work after the provider's rows land

A close, eviction or host stop decided whether it ended a running turn from the journal as it
stood, while the provider's own rows (the turn its echo opened) could still be in the session's
event sink. A close landing in that gap wrote no Stop event, so the turn it cut read as news. It
now reads after the sink drains, as a person's Stop does, through the same check; a drain that
fails or takes over a second reads working.

* fix(native-chat): a Claude Stop naming a turn that just ended still marks the follow-up it cuts

A phone names the turn it last saw. When that turn had ended and a follow-up was still unechoed,
Claude's Stop interrupted the follow-up and ended the child, but the Stop's event named the ended
turn, so the follow-up's turn the child's end cut read "Failed" under "Cancellation requested.".
A Stop that ends the provider's session ends whatever is in flight, so its event now names the
live turn or none, and a Stop that names none binds the turn opened next. Codex keeps naming only
the turn the Stop names.

The Claude Stop turn-end tests move to their own file, since the session-ending Stop suite is at
its line budget.

* fix(native-chat): the idle sweep reads working by the same rule as a stop's event

The sweep judged a chat resting while a send whose reply was lost was still unanswered, but the
stop's event writer counts that send as work. So the sweep evicted it and wrote an evict event,
which ends a person's Stop pause and let the cards behind it drain on their own. The sweep's owed
work now reads the main agent working the way every session list and the event writer do.

* test(native-chat): an aborted eviction's injected drain failure lands on the eviction's own drain

A host stop now drains the session's sink once to judge whether it ends work, so the tests that
fail the eviction's drain-published step skip that first drain.

* fix(native-chat): the idle sweep's rest writes no Stop event; it evicts a send that never echoes

The previous commit made the sweep count an unanswered send as owed work, which pins a chat whose
admitted send Codex never echoes forever, and the sweep exists to retire exactly that. That rule
returns. The sweep stops only an agent it judged resting, so its eviction now writes no Stop
event, whatever send it retires: a person's Stop pause holds through it.

* fix(native-chat): stopping a start that carries no send writes no Stop event

A host stop, eviction or close of a starting child wrote a Stop event whatever the start carried.
A start with a send already reads working, so the clause only mattered for a start with none,
which ends no turn and no send: its event only lifted a person's Stop pause and bumped the idle
clock, which is why the idle sweep had been changed to close the conversation in the same pass.
The clause goes and the sweep is #24072's again. The child's end still reads host-stop, as before.

* test(native-chat): a Stop's pause across a restart is tested with a restart that writes no event

The rig's restart closes the chat with an eviction, which now writes a Stop event when work runs
and so ends a person's Stop pause. "A Stop never hides a restart's pause" then passed with no Stop
pause left to hide anything. Those tests, and the pause-lift test whose dropped assertion returns,
restart as a process that dies with no close, which like a quit writes no Stop event, and assert
that both the Stop's and the restart's pauses are in force first.

* fix(native-chat): a host stop of a turn a person's Stop is still ending keeps that Stop's reason

An eviction or host stop that landed while a person's Stop or close was already ending the same
turn wrote a newer Stop event, and the turn's end reads only the latest, so the person's Stop of
that turn read as news. A host reason now writes nothing while a person's Stop still decides what
runs: the live turn it names or bound, or, with none, the turn a send opens next. The person's
own close still writes. The E2 tests now open and end the stopped send's own turn, as Codex does,
so the mail turn after it is not the turnless Stop's.

* fix(native-chat): an older Claude's error on a later turn keeps its error text after a Stop

The translator left an error result that names no reason to the journal's Stop rule whenever a
person's Stop named the turn or none, but the rule binds a Stop naming no turn only to the turn
opened next. So a real error on a later turn read "Failed" with its error text dropped. The
translator now asks the journal's rule itself (`personStopDecidesTurn`, the one core
`turnEndAfterStop` and a host stop's in-force check share), so the two cannot disagree.

* fix(native-chat): a Stop of a start that never landed binds no later turn, whatever sent it

A person's Stop pressed while the agent starts names no turn, and the send it stopped is
cancelled before it opens one. The Stop then bound the next turn anything opened (orchestration
mail, a restart continuation, the queue's drain, all of which send as the host), so a host
eviction of that turn wrote nothing and its crash or close read as the person's cancellation. A
Stop that named no turn now binds only a turn no send journaled after it opened: any send since,
of any origin and not refused, opens its own. The E2 test's mail send is accepted as Codex
accepts it, instead of opening the stopped send's own turn first.

* test(native-chat): a rewind's restated turnless Stop binds no turn opened after the rewind

A Codex rewind restates a person's Stop still in force after the turns it keeps, at a new
sequence, so by sequence alone it would bind the next turn opened after the rewind. A send
journaled after the restated row voids that binding (the previous commit), which this pins.

* fix(native-chat): a relaunch settles a person's stopped turn with no "stopped while in progress" row

After a restart, a turn a person's Stop ended reads "Interrupted after N" with the muted mark, but
the relaunch still added the error row saying the provider stopped mid-response, which a live Stop
never writes. The settle now skips that row when every turn it interrupts is the person's Stop's
by the journal's one rule; a crash nobody stopped keeps it.

* test(native-chat): the unexpected-exit settle's journal fake answers whether a person's Stop decides a turn

* fix(native-chat): a host stop whose sink drain fails reads the journal as it stands

A host stop drains the session's sink before judging whether it ends work, and a failed or slow
drain read as working. So an eviction of an agent at rest wrote a Stop event that ended nothing,
which lifts a person's Stop pause, and a close wrote a person's event naming no turn. The drain is
now best effort: the stop goes ahead either way and only its record is at stake, so a failed or
slow drain leaves the journal's read as it stands. A person's Stop keeps its own rule.

* fix(native-chat): a Stop that named no turn applies only to a turn a send it stopped opened

A person's Stop pressed before any turn showed names no turn. It bound the first turn opened
after it, then (5ead1f6bcc) any turn opened by no later send, so a turn the host started for
a card the Stop held, or for orchestration mail, read as the person's cancellation, and a host
eviction of it wrote no Stop event when its send had been abandoned by the close first.

The rule is now the concept itself: a Stop naming no turn applies to a turn opened by a send it
stopped, one already handed to the agent at the Stop's position. Nothing new is stored. The turn's
row names the send that opened it (Codex: the submission's key; Claude: the echo, which the journal
aliases to the submission), and a handed-over send's item sits at its handover, so the target set
is derived from the journal. A card the Stop held is handed over after it, so it is no target; a
Stop of a start whose send never opens a turn binds nothing; a rewind keeps no submissions, so a
restated Stop binds no turn opened after it. With no turn running, a host stop defers to the
person's Stop only while every unanswered send is one it stopped. Claude's translator, which asks
before its echo row lands, passes the send its echo acknowledged.

* fix(native-chat): a host stop whose sink drain runs long reads the agent working; a failed one reads the journal

A drain past its bound may still hold the turn's row, while the echo's acceptance has already
landed, so the journal as it stands read nothing running: a person's close of that turn wrote no
Stop event and the turn read as news. The two drain outcomes now differ: one that failed has
nothing more to deliver, so the journal's read holds (as before); one still running reads working.

* test(native-chat): a host stop with no turn running defers only while every unanswered send is the Stop's

The branch had no test. An eviction with only the stopped send unanswered writes nothing; one
with a send made after the Stop still unanswered writes its event.

* fix(native-chat): a slow sink drain reads working only while an accepted send's turn row is due

The previous commit read every drain past its bound as working, so a host eviction or stop of an
agent at rest during a sink backlog wrote a Stop event that ended nothing and lifted a person's
Stop pause. A slow drain now reads working only when the latest send the agent accepted has opened
no turn the journal holds, the race it was for; otherwise the journal's read holds.

* test(native-chat): the host-stop control keeps the stopped send unanswered beside the later one

With both unanswered, the host writes only because not every unanswered send is the Stop's; a rule
that deferred when any one was would pass the old control.

* fix(native-chat): a steer is no send owed a turn when a slow drain judges a host stop

A slow drain reads working when the latest accepted send has opened no turn yet. A Codex steer or
a Claude fold is accepted into the running turn and never opens one, so a chat at rest whose last
send was a steer still read working, and an eviction lifted a person's Stop pause. Sends delivered
into a running turn, whose item carries that turn's scope, are skipped.

* feat(native-chat): a chat reads Stopping from the person's Stop until the work it stopped ends

The host derives it on each journal publish from the Stop's event, the live turn and the Stop's
own answer, and publishes it as an optional field on the session status and the main agent's row.
Clients present it: the chat's tail line and Stop control, the sidebar row, worktree ps and the
phone's row. The chat and the phone also read their own Stop press until its request answers.

* test(native-chat): pin Stopping on the phone and across mixed versions

* test: give touched fake journals and mocks their SAFETY notes

* test(native-chat): a turn waiting on the person reads attention, never Stopping

* test(native-chat): the sidebar row follows Stopping when it is the only field that moved

* test(mobile): the phone reads Stopping from its own Stop until the request answers

* test(mobile): type the held Stop request instead of casting it

* chore: keep the base lockfile (a local pnpm run rewrote it)

* fix(native-chat): narrow the Stop note's optional failure; type the phone test's reply

* test(native-chat): type the Stop test envelope's fields narrowly

* fix(native-chat): a Stop the agent declined, or whose child end failed, says so while the turn runs on

A Stop naming the turn that still runs, refused by the agent, now writes the
Stop's refused fact instead of 'already finished'. A session-ending Stop whose
child end fails while the work runs on revises its note to unconfirmed.

* fix(native-chat): Stopping holds while any press of the Stop took

A repeat press refused after an earlier press took no longer clears Stopping.
Exit early when the Stop named a turn that is not the live one.

* fix(native-chat): keep Stop enabled while the host says Stopping

Only this client's own Stop request in flight disables Stop and Esc. A repeat
Stop is how a stop the provider took but never answered escalates.

* fix(sidebar): every agent row says Stopping in place of its tool line

The dashboard row, which the sidebar's non-compact mode also draws, read the
last tool line while a person's Stop ended the turn. It now shares the compact
row's rule.

* fix(mobile): the worktree list sees Stopping change on its own

A snapshot whose only change was the host dropping Stopping compared equal and
was thrown away, leaving the row on Stopping.

* refactor(native-chat): fold the status feed in src/shared for both clients

The snapshot merge and the contact-loss strip move out of the renderer feed so
the phone folds the same stream the same way.

* feat(mobile): let phones read the structured session status stream

agentSession.subscribeStatus joins the mobile allowlist. The agent-session
methods move to their own file, which the at-cap allowlist spreads in, and the
allowlist test reads the Set instead of parsing the source.

* feat(mobile): the phone chat reads Stopping from the host, like the desktop

One status stream per client, opened on a host that advertises the status feed;
a refusal to phones reads as no feed. The chat reads Stopping from the host or
its own press, and holds Stop only while its own request is in flight.

* test: give the new fakes checked types or a SAFETY reason

* revert(native-chat): drop the refused-named-turn rewrite of a Stop's note

Codex can send its refusal before the turn's end frames, so reading the turn as
still live after a flush races; the Codex Stop that ends nothing is handled by
ending the process instead. The base's 'already finished' note and its test
expectation return. The Claude wind-down failure revise stays.

* fix(mobile): release the status stream when the host ends it; refusals last one connection

The feed now drops the handle of a stream the host ended or refused, so the
logical client never replays it on a later session. A refusal to phones holds
for one connection, so a host updated while the phone stays paired is asked
again.

* perf(native-chat): read the live turn's opener from its record when deriving Stopping

After a Stop that named no turn, every later commit walked and copied the whole
journal to find the live turn's record. The derivation now reads that record
off the rendered snapshot's tail and decides with the same rule.

* refactor(mobile): move the method-unavailable check into transport

The status feed imported it from the Files tab's fallback. No behaviour change.

* test(native-chat): a send after a Stop reads Working before its turn opens

Pins the derivation's running-only read of the newest turn: the stopped turn,
already ended, must not keep the next send on Stopping.

* fix(sidebar): the compact row leads with Stopping so a narrow sidebar keeps it whole

At the default width the row read 'Codex Chat - Stoppin…': the model and time
keep their room and the line truncates from the end. Stopping now leads the line
the way monitoring already does, so the chat name is what gets cut. Also pins
that the turn bar keeps its running clock while the tail line says Stopping.

* test(native-chat): the retry of a close whose exit was unproven writes no second Stop event

The idle sweep finishes a stop left owed with that stop's own cause. It is the same stop, so its
event stands alone and the child's end keeps the cause, for a person's close and an eviction.

* fix(native-chat): read and write a Stop's answer by the turn its event records

Stop notes are now one row per turn, keyed by the turn the Stop's event
records. Performing a Stop and deriving Stopping share that key. A refused or
unconfirmed answer never overwrites one that took at the same key; only a
session-ending Stop's failed wind-down downgrades it, and that step now revises
the note the Stop actually wrote, carried on the wind-down. Stopping reads the
turn's note whenever it was first written, plus newer notes no other turn owns.
Test fixtures gain the host logger and the phone's quietRepeatedStop.

* refactor(native-chat): read a Stop's target once for its event and its note

The chat's Stop now reads what it is aimed at (the named turn and whether the
Stop ends the provider session) once, and both its event and its note's key
derive their turn from that one value through the same rule. Adds the host test
for a Stop naming an ended turn on a provider whose Stop ends the session.

* fix(native-chat): the host never steers a message into a turn a Stop is ending

A queued card's Send-now, or a send made while a person's Stop ends the turn,
went to the agent as a steer into that turn. The delivery loop now holds any
waiting message while the host's own Stopping reading holds, and sends it as
its own turn once the turn ends. The Stopping reader also stops at the Stop's
position and looks the turn's note up by key, instead of walking the whole
journal.

* feat(native-chat): while Stopping, the composer says a message runs after the stop

Desktop and phone: the composer placeholder reads "Queue a message to run
after the stop" while the chat reads Stopping, and a queued card's Steer (and
the desktop's steer shortcut) is held. New key translated in all 6 catalogs.

* refactor(native-chat): the chat pane's Stop controls live in their own module

The pane went over its line limit once merged with main. Its Stopping reading,
the press that holds Stop, and the steer and placeholder it hands the composer
move to native-chat-structured-stop-controls.ts.

* fix(native-chat): hold a send at its handover, reading the feed's own Stopping

The hold was checked when the delivery step started, but the handover runs in a
later step after waiting on the agent's start, so a Stop landing in between let
a new send steer into the stopping turn. The check now runs at the handover.
It reads the status feed's projection for the commit instead of rendering the
journal again, so holding a send adds no journal read of its own.

* refactor(native-chat): one display status decides Stopping on every surface

agentStopDisplayStatus combines whether the agent works, the host's flag and
this client's own press. The chat pane, sidebar and dashboard rows, and the
phone's chat all read it, instead of each combining the flags.

* fix(native-chat): Stopping holds until the stopped turn ends, whatever the Stop's answer

A Stop the agent declined, or whose end went unconfirmed, used to drop the chat
back to Working. It now stays on Stopping until the turn ends, and Stop stays
enabled so a repeat press escalates. The Stop's answer is no longer read for
Stopping, so its note goes back to the key the base gives it (the restore of
queued-stop.ts and the removed key test landed in the previous commit). The
note still keeps a press that took over a later refusal, and a failed process
end still says the Stop went unconfirmed.

* fix(native-chat): a Stop binds only the turn it actually stopped

A Stop pressed before any turn showed used to claim, at end-write time,
whatever turn the stopped send later opened, even when the Stop stopped
nothing. A turn that then died on its own read as "Interrupted" (your
cancellation) instead of "Failed".

Now a person's Stop that named no turn binds, in memory only, every turn
that ends while the Stop settles, and afterwards only the turn its
interrupt took. The settle ends a still-running stopped turn once. A
relaunch finds nothing in memory, so an unsettled turnless Stop binds no
turn. A Codex Stop whose answered turn does not open within its wait, or
whose send's answer was lost, now answers refused, so the host ends the
child and the turn can never run.

* fix(native-chat): keep the person's queue pause and close binding after a Stop settles

A host stop or eviction with no turn running now defers to a person's
Stop while its queue pause still holds with nothing sent since, read from
rows, so a held card is not handed off on reopen after a Stop that did
nothing or whose kill failed.

A person's close that named no turn opens a settle around its child's
end, so a turn that end cuts reads as theirs.

A press opens its settle only when the latest Stop event is its own or
the one in force it repeats: a late Stop, a card's interrupt or a lost
event row reopens no earlier Stop.

A Codex Stop that cannot reach a turn still able to open says the Stop is
unconfirmed rather than that no turn ran, and a second Stop still reaches
a turn an earlier wait left unopened.

Also drops the unused openedBy plumbing and the unreachable "a written
cancellation stays one" rule, and pins a relaunch after a named Stop.

* fix(native-chat): keep a failed Stop's turn display-only, and settle edges off the commit path

A Stop that failed marks the turn it could not stop for "Stopping…" only
(JournalStopSettle.failedOn): no turn-end rule reads it, so that turn's
own end with no verdict reads as a failure, not the person's.

A settle edge writes no row, so it no longer goes through the journal's
commit listener, which also delivers history, counts as activity for the
idle sweep and schedules the queue drain. A narrow settle-edge hook
republishes the status row and wakes the steer hold's handover, and
nothing else.

* fix(native-chat): a Codex Stop agrees on both presses when a turn is still owed, and pin the close's settle

A Codex Stop that waited for a turn Codex answered a send into now answers
"may still open" whenever that turn neither opened nor ended and its send
is still owed, however the wait ended (it ran out, or the thread went
idle). Before, a first press after an idle thread said no turn was
running and kept Codex, while an identical second press ended it.

Adds a test that a person's close the conversation outlives (as /clear
does) closes its settle, so a later turn that ends on its own reads as a
failure.

* fix(native-chat): a Stop that failed before its turn showed still reads Stopping through that turn

A Stop that failed with no turn open marked nothing, so the chat dropped
to Working and the turn that then opened never read "Stopping…". The
display-only mark now also covers that case: the first turn that opens
after the Stop failed, provided no message was handed to the agent in
between. No turn-end rule reads the mark, so that turn's own end with no
verdict still reads as a failure.

* test(native-chat): a Stop whose event row failed binds no turn to an earlier Stop

With one ordered journal writer the Stop's event is in the fold when its
write returns, so the press reads whether it owns the latest Stop from the
fold instead of awaiting the write. Pins the case the read must refuse.

* test(native-chat): name the settle, not a stream drain, in the Stop's own-end test

* test(native-chat): a Codex Stop answered before Codex ends the turn reads interrupted throughout

Codex answers an interrupt it took before it sends turn/completed (interrupted):
on TurnAborted the app-server answers pending interrupts, then ends the turn, on
one channel. The test fake did the reverse. It now answers first and ends the
turn on a later read, and the tests that read the turn's end right after a Stop
wait for it.

New end-to-end test through the shipped host, journal and Codex adapter: with
the real order, every end row of the stopped turn reads interrupted by the Stop
(named, unnamed, and a Stop pressed while turn/start was in flight). Breaking the
settle window turns the in-flight case red: the Stop's own end row then has no
verdict, which reads as failed until Codex's end lands.

* test(native-chat): Stopping ends with a Codex turn whose interrupt is answered before its end

With Codex's real order (the interrupt's answer, then turn/completed interrupted),
the status shows Stopping while the Stop settles, drops it once the turn ends, and
never carries a verdict other than the person's cancellation.

* fix(native-chat): a Codex Stop interrupts a turn Codex answered but has not opened at once

A Stop that named no turn, made after Codex answered a send but before the turn
opened, used to wait up to 5 s for the turn to open before interrupting, and
ended the Codex process when it didn't. The stated reason, that Codex refuses an
interrupt until it opens the turn, holds only part of the time: with no turn
active, Codex takes an interrupt once its thread runs (turn_interrupt_inner),
and refuses it with -32600 "no active turn to interrupt" before that or once
the turn has ended.

The Stop now sends the interrupt at once. Only on that refusal, while the turn
has neither opened nor ended, does it wait for the turn to open (bounded at
5 s) and send it once more. A turn that ended meanwhile was nothing to stop. One
that never opens, or that an earlier wait already gave up on, fails the Stop,
and the host ends the child as before. Sends still wait for the turn to open
before steering into it.

The test fake models Codex taking an interrupt once the thread runs (run()).

* fix(mobile): name how the phone's status stream is released in the subscription inventory

Main made each inventory entry state its release; the status feed's stream is
released from its subscribe params, as the session event stream is.

* fix(native-chat): every Codex Stop waits for an answered turn to start, as the first did

A Stop whose interrupt Codex refused as finding no active turn skipped the wait
when an earlier wait, a Stop's or a send's, had already given up on that turn.
Every press now waits its own bound and retries once if the turn starts, so a
turn that opens during a later press is still stopped. Both presses still reach
the same verdict when it never starts.

* fix(codex): never steer a turn whose interrupt Codex answered

Codex answers an interrupt as the turn aborts, before it sends that turn's
turn/completed. In that gap the adapter still counted the turn as running, so a
message handed over right after a Stop settled (the Stop's own end row already
reads the turn ended) went out as turn/steer, which Codex refused with -32600
"no active turn to steer", and only then as turn/start. The adapter now marks a
turn whose interrupt Codex answered as aborted until its turn/completed, never
steers into it, and starts the message's own turn directly. Steering a turn
that is genuinely running is unchanged.

* fix(native-chat): a message sent while Stopping is queued as a card, whatever the setting

While the chat reads Stopping (the host's flag or this client's own Stop in
flight) there is no turn left to steer into, so the desktop asks the host to
queue the send even with the queueing setting off, and it is never drawn as a
bubble inside the turn being stopped. The phone already queued every send on a
capable host; a test now pins that it does so while Stopping.

* fix(native-chat): a message queued while Stopping is a card at once, not after the stop

A person's Stop holds the session's lane until Codex answers its interrupt, and
a send was admitted only behind it. By then the turn read ended, so a send that
asked to be queued went out plain: no card for the whole of Stopping, then a
bubble and a new turn.

While the host reads that a person's Stop is ending the work, a text send that
asks to be queued is admitted without waiting for the lane: the same ledger and
lease admission, and a plan that only writes the card through the journal's
ordered writer. The card runs when the stop lands; the Stop's pause holds only
cards queued before it. Anything else, including a Stop that settled by the
time the send runs, takes the lane as before.

The Codex test fake now drops the active turn when it takes an interrupt, as
Codex does before it answers, so a turn/start after the answer opens a new turn.

* fix(native-chat): a Codex Stop that ends the child before any turn opened withdraws its send

A Stop on a Codex turn that was answered but never opened ends the Codex
process. That end settled the send as in doubt (unknown, recovered), and the
client's outbox holds every later send behind a send in doubt until the person
presses Retry, which re-sends the very message they stopped. The chat looked
stuck.

Codex records a prompt only once its turn has started, so a send whose turn
never opened never ran. When the Stop's refusal says so (turnMayOpen), the child
end now settles the unanswered sends as withdrawn, the verdict Codex's own
interrupted-turn end already gives an unechoed send. The flag rides on the owed
wind-down, so a retry after a failed child end withdraws them too. Claude's
child end still leaves its unanswered send in doubt.

* fix(native-chat): derive the withdrawal of a Codex send whose turn never opened

Replaces the flag the Stop carried to the child's end, and its copy on the owed
wind-down, with a reading of the journal at the settlement that lands. A Codex
child's unanswered send is withdrawn when a person's Stop is in force since it
was sent and no turn row ran, or was written, after it; any other end (a turn
that opened, a host's close, a crash, Claude) still leaves it in doubt. A
retried wind-down reads the same rows, so it withdraws the same sends.

* fix(native-chat): a card queued while Stopping runs past the cards the Stop holds

A card queued before a person's Stop waits under its pause until Resume. One
queued after it, as a message sent while Stopping now is, was stuck behind them
too, since the queue never reorders. Such a card was asked for after the Stop,
so it runs when the stop lands, past the cards held only by that Stop's pause;
a returned card and the restart and /clear pauses still hold everything behind
them.

Also: the host's own Stopping reading is gated on working, as the published
flag is, so a failed Stop's mark never reads Stopping on an idle session; a send
that falls back to the lane re-reads the conversation's journal there; and the
end-to-end test asserts the queue's pause rather than a per-card field.

* fix(native-chat): a paused queue labels only the cards it holds

Since a card queued after a person's Stop runs past the cards the Stop holds,
labelling every card "paused" while the queue's pause is published misreads that
card. The host now marks each card its pause holds (heldByPause, a new optional
field), and the desktop and phone label only those. An older host marks none,
so a client keeps today's labels; an older client ignores the field.

Adds a test of the desktop's own send through the real outbox: while the chat
reads Stopping, the request asks the host to queue it and no bubble is drawn,
on a host that advertises the queue.

* revert(native-chat): defer the per-card queue pause label to the queue's rollout

The heldByPause field and its labels are visible only where the host
advertises the queued-messages capability, which shipped hosts do not yet do.
Deferred to that rollout; the real-outbox send test stays.

* fix(native-chat): while Stopping, say and show what a send does where the queue is dark

Shipped hosts do not advertise the queued-messages capability, so a message
sent while Stopping goes out plain: the host holds it until the stopped turn
ends and then runs it as its own turn. The composer still said "Queue a message
to run after the stop", and the message was drawn inside the turn being
stopped.

Now the placeholder reads "Send a message to run after the stop" where the host
does not queue sends, and "Queue a message…" only where it does (desktop and
phone, all six catalogs). A send this client made that the host has not
recorded yet is drawn after the Stopping line while the chat reads Stopping, as
a message held behind a running command already is; once the host hands it
over it opens its own turn. Client presentation only.

* fix(native-chat): keep a send in doubt when a turn was open for it

The derived withdrawal read a turn as open for a send only if it still ran or
was written after the send. A send steered into a running Codex turn whose
interrupt failed met neither once the adapter's end settled that turn ahead of
the host's settle, so it read withdrawn, though Codex drains a steer into the
running turn and may hold it. A turn that ended after the send was handed over
was open for it too: such a send stays in doubt, as before.

Pins that case, and that a send made after the Stop, to a child that then dies
before its turn opens, stays in doubt.

* fix(native-chat): restore the per-card queue pause label

Kept after all: a paused queue labels only the cards it holds (heldByPause),
which is visible only where the host advertises the queued-messages capability.

* fix(native-chat): draw only a send made while Stopping after the Stopping line

Every send the host had not recorded yet was drawn after the Stopping line,
including one made just before the Stop, which the host steers into the turn;
it then jumped up into that turn once recorded. The outbox now marks a send
made while the chat reads Stopping, and only those wait after the line.

* fix(native-chat): withdraw a Codex send by whether it started its own turn, not by timing

Whether a turn was open for a send was read from end times: a turn that ended
after the send's handover counted. A send made while a Stop ended the turn is
handed over once that turn reads ended, yet Codex's own end for it can arrive
later, so such a send whose own turn never opened read in doubt again, and the
chat's queue held behind it.

The handover already records where the send went: its message joins the turn
running then (a steer) or belongs to no turn (it starts its own). Only a send
that started its own turn, with none opened since, is withdrawn; one that
joined a running turn, or has no recorded place, stays in doubt.

The Codex test fake takes an answered interrupt as Codex does, dropping the
turn before its end arrives.

* refactor(native-chat): move queue-while-stopping to its own follow-up

The queued-messages capability is off on every shipped host (#21062), so the
parts of this PR that act only when it is on move to a follow-up stacked on
this one: admitting a queued card while a Stop holds the session's lane, a card
queued after a Stop running past the cards it holds, the per-card pause mark,
and asking the host to queue a send made while Stopping. This PR keeps the
Stopping state, the host's steer hold, the rule that never steers a turn whose
interrupt Codex answered, and what a send while Stopping looks like where the
queue is off.

* fix(native-chat): leave no Stop row when the Stop took back a send that never ran

A Stop on a Codex send whose turn never opened ends the child, and the child's end
takes the send back into the composer. The Stop still wrote "Cancellation
requested." at the conversation level, so with the send gone it sat under the
previous finished turn and read as if that turn had been stopped. A Stop that found
no turn running and whose child end took back every send it found now writes no
row; a Stop of a running turn, or one that leaves a send in doubt, still does.

* fix(native-chat): count a send whose answer was lost when a Stop takes it back

The no-row rule counted only pending sends, but the child's end also takes back a send this process left in doubt when Codex's turn/start answer was lost. That case still wrote "Cancellation requested." under the previous turn. Both now read one predicate, so they cannot drift apart.

* test(native-chat): pin which sends a Stop's child end can take back

A send an earlier process left in doubt is never withdrawn and never holds the row back, and a queued card's send is never counted.

* fix(native-chat): hold the next handover while the send ahead opens its turn

The delivery loop handed queued messages to the provider back to back. Codex steers a second send into the first one's turn once it opens, and Claude folds it into the running cycle, yet its handover row was written before that turn existed, so it read as belonging to no turn and was drawn ahead of the reply. The loop now hands the next message over only once the send ahead has opened its turn, settled, or been stopped, all read from the journal; each is a commit, which wakes the loop again. The message then goes in scoped to the opened turn.

* fix(native-chat): doubt a Codex send whose turn never opened once Codex goes idle

A Codex before 0.148 fails a turn before opening it with only an error. The send stayed pending, and the hold behind it waited on it. When Codex reports its thread not running with no turn open, a send answered into a turn it never opened or ended now settles as doubt with the existing idle reason, which releases the hold. No clock: a slow but healthy turn start is never doubted.

* fix(native-chat): end an unopened Codex turn on its final error, not at idle, and drop the Stop release

Codex publishes the thread idle ahead of turn/completed, and between an aborted turn and the next picked one, so releasing at idle doubted sends in turns Codex did open or was about to. A final error naming a turn Codex never opened (its only end before 0.148) now ends that turn as failed instead, settling its send as a turn/completed failure would. The test fake publishes idle before turn/completed, as Codex does. A Stop no longer releases the hold: the seven rig tests that needed it now settle the stopped send the way the provider does.

* docs(native-chat): put each turn-end settlement comment on its own function

* test(native-chat): echo the first send before the next in the restarted-child test

The test's child admitted the first message and never answered it, which no live provider does; the next send was then held behind a turn still opening. The child now echoes the first message, so the test still checks that the next send restarts nothing.

* fix(native-chat): draw a message held behind an opening turn after that turn's live status

While the send ahead was still opening its turn, the live turn was taken to be the newest user message, the held one, so its live status drew under the held message and the held message drew above it. While a send is opening, its turn is now the live one, and messages sent after it, queued or not yet recorded, wait behind the live turn as a message held behind /compact or a Stop does. The host's hold and the client's drawing read one shared predicate.

* fix(native-chat): keep held messages waiting while Stopping, and leave Retry-only sends in place

While Stopping, a message typed then returned early from the waiting rule, so a message held behind the opening turn moved back above the Stopping line. The rule now waits the union of both. A send only the user's Retry sends again is not held by the host, so it no longer waits behind an opening turn; the projection marks it.

* fix(native-chat): keep a turn on the send that opened it when Codex echoes a steer first

A message held behind an opening turn is steered in moments after that turn opens, and
Codex can echo the steer before the send that opened the turn. Both carry the turn's one
provider key, so the steer was read as the turn's opener: for that moment the live status
moved under it and restarted its clock. While the opener is still in flight ahead of the
turn record, a steer into that turn no longer takes it.

* fix(native-chat): never anchor an opening turn on a message still queued above its send

Two messages queued behind /compact, or sent while Stopping, are accepted above the first
send's handover, and the hold makes the first send's turn record land before the second is
handed over. The turn fell back to the first unechoed send ahead of its record, which was
the queued one, so its live status moved under it. A send still queued is not in flight.

The phone keeps no outbox, so its frame tests now draw only recorded rows, through the
phone's own fold rather than the desktop's transcript order.

* fix(native-chat): read the host's Stopping beside main's startup phase

Main now reads only the startup phase from the status feed and no longer publishes which
child is starting. The chat reads the host's Stopping from its own hook beside it, and the
Stopping bridge test mocks the execution-host lookup main's owner resolution now calls.

* test(orchestration): open the working send's turn before a `now` send joins it

The rig's working send was handed over with no turn record, so the running turn the test
names never existed; the hold rightly kept the `now` send until that turn opened. The setup
now opens it as a provider does. The assertions are unchanged.

* test(native-chat): pin the phone's own echo of an accepted send behind an opening turn

The phone keeps no outbox, but it does show its echo of a send the host accepted until that
send's row arrives, and the shared waiting rule moves that echo behind a turn still opening.
The phone test now builds its list as the phone's view does, echoes included, and covers it.

* test(native-chat): read the outbox reconcile from where main moved it

* test(native-chat): pass the projection's options after main's rejected-in-place rows

* fix(native-chat): a retried message no longer waits behind a later Stop

Retry dropped the Stop it had outlived but kept the mark that it was sent while
a Stop was ending a turn, so a retried message waited behind whatever later,
unrelated turn a Stop was ending. Retry is a new send: drop that mark too.

* fix(native-chat): word a send after a Stop by whether this send will queue

The 'queue a message to run after the stop' placeholder read the host's
queue capability alone. A send queues only when the host queues and this
send asks it to: the queue setting is on and no pending prompt blocks the
queue. Desktop and phone now word the placeholder from that same decision
their send uses.

* test(native-chat): one test per case for the words of a send after a Stop

* refactor(native-chat): the dictation hook owns the composer's dictation state

Keeps NativeChatComposer within its line limit after the Stop props and
main's /context answer both landed in it.

* refactor(native-chat): name the dictation hook for what it owns now

* fix(native-chat): a Stop's note says it took once a joined close proves the exit

A session-ending Stop whose child's end failed revises its note to
'unconfirmed'. Since main's #24862, the next Stop joins that close rather
than stopping again, and wrote no note, so a close that then proved the
exit left 'unconfirmed' under a turn that ended. The close now carries
the note it settles, and its proven exit revises it to 'Cancellation
requested.'. A join that fails again leaves it unconfirmed.

* fix(native-chat): a proven turn end says a Stop's unconfirmed note took

Replaces the note carried on the child's close. Every settlement that
ends turns interrupted on a proven exit, live or after a crash, also
revises an unconfirmed Stop note on those turns to 'Cancellation
requested.', found by the note's turn scope, in the same batch. A note a
Stop wrote before its turn showed is re-keyed onto the running turn when
it becomes unconfirmed, in one batch, so that end finds it. Known limit:
with no turn open yet, the note keeps its key and no turn's end revises
it.

* test(native-chat): a Stop's unconfirmed note says it took when the agent exits on its own

* test(native-chat): a message that joins a Stop's unproven Claude close goes to the resumed child, the next waits for its echo

Covers the case main's #24862 retired with its unproven-stop test: the
first message after an unproven close joins it and reaches the resumed
child, and a second waits for that message's turn to open.

* test(native-chat): name the Codex handle as main's opaque handle does

* test: restore the provider handle import the main merge dropped

* fix: derive Stop note wording from interrupted turns

* test: name the raw replay case for what it covers

* refactor(native-chat): move the waiting-slot split into its own hook

* test: give the opening-send hold host the agent registry main now requires

* test: follow main's chat font-size rename in the stopping tests

* test: follow main's chat font-size change in the opening-send test

* test: follow main's single live-line value in the Stopping tests

* test: give the android live-line fixtures the stopping field

* chore: keep the session host under its line limit after the main merge

* chore: keep the composer test and the phone chat view under their line limits after the main merge

The Stop control now disables itself while Stopping, so the composer passes
the flag through and its test file stays as main has it. The phone chat
header's Stop moves to its own component.

* perf(native-chat): read a Stop note's fields before parsing its key on every snapshot

Every snapshot projects each item through the Stop-note read, and each new
snapshot rebuilds the index of Stop notes by turn. Both parsed every item's
key first; they now check the row's kind and turn scope (and, for the
projection, its unconfirmed-stop failure) before the parse. Every Stop note
is a status row, so what each finds is unchanged.

* fix(native-chat): a Stop ends the hold on the next message, even when the stopped send's turn never opened

* test: a Stop's note ends the hold on the next message

* fix(native-chat): only a Stop that took ends the hold; a refused or unconfirmed one leaves the turn opening

* fix: import the moved Stop-note helpers where the host module uses them; type the test's failure note

* test(native-chat): pin that a taken Stop's note lands after the turn the provider opened before answering it

* test(native-chat): give the opening-send hold test's runtime the launch arguments main now requires

* fix(native-chat): a message a Stop takes back while the turn ahead opens stays after that turn, with its stop row; move the rows that wait behind the live turn into their own module

* test(native-chat): run the two opening-send hold tests, which open a real journal database, in the Node runtime project

* test(native-chat): with a send still opening its turn, a starting child gets only the first message; a failed start still rejects both in its words
2026-10-06 22:57:51 -07:00
Brennan Benson 988f87e743 fix(native-chat): move the attachment failure words to their own module so lint passes on main (#26079)
The failure-words table grew past the 300-line limit when two changes landed back to back,
which fails Lint on main and makes every PR's static-analysis job skip typecheck.
2026-10-06 22:49:50 -07:00
Jinwoo HongandClaude 358e4f92e0 fix(claude): share account history on Windows with junctions and a same-drive hardlink (#26067)
* fix(claude): share account history on Windows with junctions and a same-volume hardlink (STA-3698)

Windows skipped history sharing for Orca-managed Claude account folders on
the mistaken premise that it needs symlink privilege. Session folders now
use directory junctions and history.jsonl a same-volume hardlink, neither
of which needs elevation. A different volume or a failed link keeps that
account's prompt history private and reports it. A small link record tells
a replaced shared file apart from the account's own copy so prompts the
user removed are never merged back. Renames retry on Windows file locks.

* fix(claude-accounts): refuse a cross-volume Windows share before draining set-aside prompt history

A pending set-aside copy was appended to the shared history.jsonl before the
same-volume check refused the share, leaking the account's prompts into it.

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-07 01:43:00 -04:00
64bb9373da Claude account profiles: dormant WSL guest setup (Step 3 of 4) (#24384)
* feat(claude): add dormant profile setup and history sharing

* fix(claude): make profile setup one gated, typed, fail-safe entry

Review round 1 of the dormant profile setup found that the pieces could
be called without their safety checks, that one failed write or an
unreadable bookkeeping file could silently stop sharing for good, and
that Windows prompt history could bring back history the user cleared.

- One entry, provisionClaudeAccountProfile: the profile gate (namespace,
  no linked components, outside ~/.claude and ~/.config/claude, and an
  ownership marker beside the home naming the account and target) runs
  first and refuses before creating anything; then history sharing,
  config provisioning, and the hook install after the settings merge.
  Results come back per surface with closed warning codes instead of
  message text.
- The sharing ledger is keyed by surface name, records a value only
  after its write succeeded, and an unreadable ledger starts empty and
  is rewritten instead of blocking every surface.
- The profile state file goes through the same locked writer as folder
  trust (Claude's <file>.lock plus the in-process queue), generalized as
  updateClaudeGlobalConfig. Onboarding and trust are still applied when
  the personal state file is unreadable.
- WSL descriptors build guest POSIX paths; the state-file path style
  follows the injected platform.
- Orca's managed statusLine has one owner in a profile: the settings
  merge never shares it, a user's own statusLine is shared over it, and
  the profile installer follows the default home's slot so a default
  opt-out reaches every profile. remove() takes the same destination;
  the remote installer cannot accept one.
- Prompt history compares file identity (bigint dev+ino) on every
  platform, never drains the shared file into itself, drains retained
  copies in generation order, never reuses a stale cursor, and on Windows
  keeps a replaced default's old copy aside instead of replaying it.
  Directory merges keep going past a failed entry.

* fix(claude): share the user's own hooks and keep merged history whole

A user's own Claude hooks in ~/.claude (notifications, formatters) did
not run under a managed account, because the whole hooks key stayed
private. They are now shared like any other settings key: Orca's own
hook entries and its managed statusLine are stripped from both the
personal value and the profile's current value before the per-key
ledger comparison, so they never travel through the merge and never make
the key look user-owned. Orca entries already in the profile are kept on
write, and the profile hook installer adds them on top as before.

Prompt history: merged bytes that lack a final newline are terminated,
so Claude's next record no longer fuses onto the last merged line. When
a CLI rewrote the profile's history file (old records plus new), only
the lines past the part it shares with the default history are added,
instead of the whole file again.

* fix(claude): close review round 2 gaps in profile setup

Hooks and statusLine sharing:
- When ~/.claude holds only Orca's hook entries, the user's shared hooks
  now read as an empty value instead of a missing key. Removing the
  user's last own hook in ~/.claude therefore reaches profiles that
  never edited it, and deleting the only shared hook inside a profile
  stays deleted.
- A custom statusLine Orca shared, and the profile never edited, goes
  away when the default home drops it. When a shared custom line
  replaced Orca's line in a profile, the profile's statusline marker is
  dropped so Orca's line comes back once the default returns to it; a
  profile that opted out stays opted out. No other key gains deletion.
- install/remove/getStatus with a profile directory refuse when it is
  the default home, or its settings.json resolves to the default one,
  instead of editing System Default's hooks and opt-out state.
- The profile statusline rule reads the default settings under the
  userHome passed to the setup entry, not os.homedir().

Profile state and ownership:
- A malformed `projects` value skips only folder trust (new warning
  code trust-refused); onboarding and shared keys still apply.
- The ownership marker stores only host-local facts (account, runtime,
  distro). The execution host id is the caller's view of the host, so
  it stays in the in-memory descriptor and is not compared.

Prompt history interruption paths:
- With no cursor yet, a retained copy starts past the bytes it shares
  with the default history, so an interrupted share no longer replays
  the whole history.
- A retained name for the shared file itself is removed with its cursor
  instead of lingering until a later scrub makes it look new.
- The Windows link record is read three-state: unreadable stops the
  share instead of reading as "no link". If the record cannot be
  written after linking, the fresh link is undone.
- An unreadable retained copy is reported and no longer blocks linking.

* build(cli): list the new Claude hook modules in the CLI project

hook-service.ts and hook-settings.ts are compiled into the packaged CLI
project, which lists every file explicitly. The statusline policy and
profile destination modules they now import were missing, so the CLI
typecheck failed with TS6307. The CLI still loads hook-service through
the existing managed-agent-hook-controls build entry, which bundles
both modules; neither imports electron.

* fix(claude): close review round 3 regressions in profile setup

- A profile whose hooks hold only Orca's entries and that sharing never
  recorded is no longer treated as a user edit, so the user's first own
  hook in ~/.claude reaches it (for example when the profile was set up
  before ~/.claude had any hooks).
- A retained prompt-history file is removed as a second name for the
  shared file only when the default history does not itself link to it;
  otherwise it holds the only copy and is kept.
- Default-home checks compare file identity: the profile hook
  destination check uses device and inode, and the profile/default
  separation check resolves on-disk case, so a case-only alias of
  ~/.claude is refused on case-insensitive filesystems.
- A test pins that an unreadable leftover session tree no longer blocks
  linking.

* fix(claude): let shared keys leave a profile when ~/.claude drops them

QA found that removing a setting from ~/.claude never reached a managed
account: deleting the whole `hooks` block left the user's hook running
there. Only statusLine followed the default away.

Every shared key now follows the same rule through the existing per-key
ledger: when a key disappears from ~/.claude/settings.json (or
mcpServers/theme from the personal state file), it is removed from the
profile if the profile still holds exactly what Orca last shared. A
value changed inside the account is kept. Keys Orca never shared,
including denylisted ones, are never touched. Deleting the whole hooks
block removes the user's shared hooks and keeps Orca's own entries. A
missing source counts as empty; an unreadable source removes nothing.

* feat(claude): add dormant profile routing and account consumers

* fix(claude): drop the dormant profile selection RPC; clients negotiate by capability

Restores the inline mobile allowlist so its source-scan guard sees every
accounts.* method again, and the generated params catalog to generator order.

* fix(claude): guard the claude shell function and honour a hand-exported config dir

The function is defined only in a routed pane where claude is a real
executable (the codex function's guard), re-reads the pointer only while
CLAUDE_CONFIG_DIR is unset or still Orca's injected twin, accepts Git Bash
drive paths, and starts on its own line after the fish/PowerShell codex text.

* fix(claude): spawn-time profile env, total account listing, setup at lifecycle triggers

Round-1 review fixes for the dormant profile routing:
- Panes get the selected profile's CLAUDE_CONFIG_DIR plus an Orca twin at
  spawn, so nested shells and scripts inherit the account; System Default
  injects nothing and its home is the inherited CLAUDE_CONFIG_DIR.
- An absent routing owner is System Default, never a throw; AI Vault and
  session-search scans receive profile roots from their parent, and the
  capability is advertised only where an owner is installed.
- Account listing never throws: per-account readiness, a stale pointer is
  republished in the background and reported on the snapshot.
- Profiles are set up at select and startup; a launch only sets up one that
  never was, and a worker fault on a prepared profile is a warning. The
  Claude version probe is cached per binary identity.
- Pre-trust goes through the existing deadline- and realpath-guarded writer
  against the launch env's profile config.
- Skill discovery keeps a caller's Claude root and a broken Claude selection
  no longer fails other providers.
- The durable record carries a provider-neutral launchAccountHome, read
  through one helper by the launch fallback and the model catalog.

* test(claude): pin the version-probe cache, launch-account record and temp-home readers

* test(claude): pin dormant bash rc text alongside fish and PowerShell

* test(claude): read the fish launch init without a nullable index

* fix(claude): withdraw the profile pointer when a selection cannot be published

A pointer left naming the previous account would launch it silently; a
missing pointer makes the claude function refuse visibly. A newer selection
that raced the failed one keeps its pointer.

* fix(claude): read the fish profile pointer with read -z for fish older than 3.4

Shell tests skip system config and abort unless claude resolves to the fake.

* fix(claude): only the newest publish withdraws the pointer; total config dir lookup

- An overtaken publish that fails leaves the newer selection's pointer.
- The runtime config dir falls back to the legacy home for an unresolvable
  account or a WSL target, so skill roots never fail for other providers.
- WSL guest reader roots merge verbatim, never realpathed on this thread.
- History readers include ~/.claude, where step-1 setup pools profile history.
- System Default ignores a config dir an outer Orca injected (twin-marked).

* fix(claude): System Default launches and probes use the structured create resolver

A Claude agent-env CLAUDE_CONFIG_DIR the create path stored is now the home
the launch pins and the model probe accepts.

* test(claude): type the System Default launch record as an agent-session record

* fix(claude): install profile hook scripts under the setup job's home

A worker thread's os.homedir() ignores its own env, so the hook and
statusline scripts now go under the home the job names. The worker test pins
the process HOME to a sentinel, refuses to run unless the worker sees it, and
asserts nothing lands there.

* test(claude): skip shell cases whose shell the runner lacks

* fix(claude): remove env vars in the PowerShell claude function instead of setting null

On .NET 9+ (pwsh 7.5+) SetEnvironmentVariable with $null creates an empty
variable, so stripped auth vars reached claude as empty strings and the
restore left CLAUDE_CONFIG_DIR empty in the user's session.

* feat(claude): add dormant WSL guest profile setup

* fix(claude): open WSL panes without guest calls and coalesce same-profile publishes

A WSL pane now gets the same non-throwing, guest-free spawn env as a host
pane; only select, startup and Claude launches publish into the guest.
Overlapping publishes of one target share the newest publish while the
selection still names the same profile, instead of failing as superseded.
Publish issues name their WSL distro and drop out when the target is no
longer routed. A late inspect from an older selection no longer replaces
the newer one's verification, a failed guest request evicts the cached
guest, and readiness is derived per account from the guest's owned homes.

* fix(claude): roll back only the target whose selection failed

With profiles, a failed select or remove republishes just its own target
instead of running startup over every WSL distro, and a rollback failure is
logged instead of replacing the error that caused the rollback.

* fix(claude): scan WSL profile history only in running distros

Vault and usage scans pass Claude profile roots through the same
running-distro filter as every other WSL root, so a stopped distro's UNC
paths are never walked.

* fix(wsl): ship the Claude profile helper only in the WSL bundle dir

The helper only ever runs inside WSL from the desktop, so it moves out of
the SSH relay artifacts (no upload, no relay version change) into
out/relay/wsl beside the other WSL-only guest bundles. The three WSL bundle
resolvers share one candidate list.

* fix(wsl): refuse old glibc before downloading, and keep the shared download per caller

The pinned Node runtime needs glibc 2.28, so a distro below the floor is
refused before any download with a message naming both versions, as SSH
hosts are. The shared download again owns its own deadline and each caller
waits on its own signal, and the OpenCode reader keeps its architecture
error text.

* fix(claude): bound each WSL guest operation and run the helper through the WSL runner

A cached guest no longer carries its 180 s preparation deadline into later
requests. The helper runs through runWslProcess (stdin payload, WSL_UTF8),
the distro is confirmed running once per preparation and once per request,
a failed `claude --version` probe continues with an unknown version like
native setup, the helper resolves from the WSL bundle dir, and the guest
entry decodes stdin once so split UTF-8 survives.

* test(claude): cover WSL profile pre-trust routing and its deadline

* refactor(claude): drop WSL refresh cleanup that the failed publish's withdraw already does

* fix(claude): catch rollback failures only when profiles route the selection

With the gate off, select and remove surface the rollback error exactly as
before; only profile routing logs it and keeps the original error.

* fix(claude): give every WSL pane a guest-relative Claude profile pointer

WSL panes now always carry `~/.local/share/orca/claude-profiles/selected-wsl`,
which the bash/zsh and fish claude functions expand against the guest $HOME
at each invocation, so a pane opened before Orca has met the distro still
follows the selected account instead of falling back to ~/.claude. Absolute
pointers are untouched, PowerShell is unchanged, and a missing pointer file or
profile still refuses visibly. CLAUDE_CONFIG_DIR is set at spawn only when the
selection resolves without a guest call.

* test(claude): assert a missing guest-relative pointer refuses with a visible message

* test(claude): type the WSL runner mock in the transport test

* fix(claude): route only WSL distros that hold an Orca account, and re-derive their publish

A WSL distro is routed only while host settings hold an Orca Claude account
for it, decided from settings with no guest call. An unrouted distro behaves
as before profiles: its panes get no pointer or profile env, and a Claude
launch is System Default with no guest prepare. A distro that loses its last
account has its pointer withdrawn best-effort so older panes stop launching
the removed account.

A routed distro without a current publish (for example stopped at startup)
gets one non-blocking background publish from its next pane spawn, coalesced
per target; its failure stays that distro's issue and a later success clears
it. A late setup result from an older publish no longer replaces the newer
selection's verification. The owner contract moves to its own module so the
routing service stays under the size limit.

* fix(claude): read WSL profile history in native chat and adoption only in running distros

Native chat resolves Claude transcripts from host roots first and reads WSL
profile roots only after a miss, filtered to running distros like Codex's WSL
homes. Structured adoption candidates go through the same filter.

* fix(claude): target registration rollbacks and keep their errors in profile mode

A failed add or re-authentication rolls back only the account's own target.
With profiles, a failed re-authentication rollback is logged instead of
replacing the original error; with the gate off both behave as before.

* fix(claude): spell the guest pointer location once and keep set -u safe

The guest helper, the withdraw script and the pane pointer all derive from
one home-relative constant, and the posix claude function reads ${HOME:-}
so `set -u` with HOME unset refuses cleanly instead of aborting.

* test(claude): cover the IPC preflight and daemon WSLENV paths for WSL profile env

The renderer preflight is tested for wsl.exe and Windows shells with a \\wsl$
cwd (which always launch wsl.exe) and with the gate off, the daemon launch
plan imports the pointer and profile home without a WSLENV flag, and the
Windows launch test uses the guest-relative pointer production sends.

* fix(wsl): report why the guest runtime failed, with download context and trimmed stderr

The install's promote output is classified with the SSH classifier, so a
self-test failure shows the exit code and the loader's words (for example a
missing libstdc++ on Alpine) and a security-software change is named. A failed
runtime download says it was Orca's Node runtime for WSL, while a checksum
mismatch keeps its own text. Guest stderr is trimmed before it reaches a
refusal message.

* test(claude): pin that pointer retirement never runs for host targets or with the gate off

* test(claude): give the routed WSL preflight fixture its required authMethod

* fix(claude): let the pane-triggered WSL publish repair a distro stopped at startup

"Distro not running" is now a typed refusal: it never withdraws the pointer
(the distro's last pointer cannot be stale, and a withdraw racing the boot
could delete a valid one) and never records a distro issue. The background
publish a pane fires now waits a few seconds for the pane's own spawn to boot
the distro, probing three times, and is dropped silently and re-armed if the
distro stays down. It joins any publish already in flight for that target
instead of preparing the guest a second time. Per-target generations and
pointer-write ordering move to ClaudeProfilePointerQueue so the routing
service stays under the size limit.

* fix(claude): remove the last selected WSL account without a guest publish

With profiles, removal writes the account list and the selection in one
update, so a distro losing its last account is already unrouted when it syncs
and its pointer is retired best-effort. Removal no longer needs the distro to
be running or able to run Orca's runtime. The gate-off order is unchanged.

* fix(claude): keep native chat's legacy Claude roots first and unfiltered

Only roots added by WSL profiles are read after a miss and filtered to
running distros; a host CLAUDE_CONFIG_DIR on a \\wsl$ share is searched first
and unfiltered, as before profiles.

* test(claude): cover stopped-at-startup repair, launch join and last-account removal end to end

* test(claude): assert no running probe before the pane has had a turn to boot the distro

* fix(claude): let user-initiated profile work boot an idle-stopped WSL distro

WSL distros idle-stop on their own, and the legacy path boots them with its
spawn or \\wsl$ write. With profiles on, a Claude launch, a select, a remove,
a failed-change rollback and the retire after removing a distro's last
account now skip the running pre-check and let their first bounded guest
command (`wsl -d <distro> --exec ...` through runWslProcess) boot the
distro. They refuse only if that command fails, with wsl.exe's own reason,
for example a distro that does not exist. Startup, the pane-triggered repair
and the history readers keep the running pre-check and its typed refusal, so
background work never boots a distro. With the gate off nothing changes.

* fix(claude): let startup join a launch or select already publishing a WSL distro

Startup no longer overtakes a user's in-flight publish for the same target,
so a launch that is booting an idle-stopped distro is not handed startup's
"not running" refusal.

* fix(claude): remove accounts of a WSL distro that no longer exists, and name the helper once

wsl.exe's own failures (exit 0xFFFFFFFF, empty stderr, the diagnostic and its
WSL_E_* code on stdout) are now read by one shared reader used by the git
runner and the WSL profile transport, so profile refusals show wsl.exe's
message. WSL_E_DISTRO_NOT_FOUND becomes ClaudeProfileHostMissingError: with
profiles, removing an account from a distro that no longer exists keeps the
removal and logs a warning, while select and launch still refuse visibly.
The helper's file name is defined once in shared/relay-artifacts.ts and used
by the relay build and the transport.

* fix(claude): give plain fish tabs the claude function through the codex hand-off

Main now gives a plain fish tab Orca's codex function through a vendor_conf.d
snippet instead of a -C init. The claude function only rode the -C path, so a
plain fish tab would not re-read the account selection per invocation once
profiles are on. Define it at the first prompt beside codex; it stays empty
while the profile gate is off.

* fix(claude): share personal rules, themes, workflows and keybindings into account profiles

A managed account launches Claude with its own config folder, so user-level
rules/, custom themes/ (which a shared `custom:<slug>` theme points at),
personal workflows/ and keybindings.json silently stopped applying. Link the
three directories like skills and commands, and copy keybindings.json with the
same edit-preserving ledger as CLAUDE.md. routines/ stays unshared: routines
belong to the claude.ai account and the folder holds per-run state.

* test(claude): wait for the running child to read its account before switching

The test switched the selection after a fixed 20 ms, so under load the backgrounded claude
had not yet read the pointer and picked up the new account. The stand-in now marks when it has
started, and the test waits for that mark (bounded) before switching.

* fix(claude): accept WSL setup warnings for every shared Claude file

The guest reply schema listed CLAUDE.md by name, so a warning about the newly shared
keybindings.json would have rejected the whole reply. It now takes the shared-file list
from provisioning, like the shared folders.

* fix(claude): import the personal CLAUDE.md into account profiles instead of copying it

Claude also loads ~/.claude/CLAUDE.md as a parent folder's memory for any project under home,
so a copied account CLAUDE.md made every such session read the user's instructions twice
(checked live with Claude 2.1.288). An @~/.claude/CLAUDE.md import resolves to the same real
file, which Claude loads once from home, from projects under home and from folders outside it.

* refactor(claude): simplify account profile setup toward the prior art

- Windows keeps each account's history private; drop the hardlink, link
  record and conflict-copy machinery that only Windows reached.
- Share hooks and statusLine as ordinary settings keys: Orca writes the
  same entries into every folder, so the installer finds them present.
  Drops the Orca-entry carve-out, the per-profile statusline follow
  logic and its marker.
- Unreadable ledger is just an empty ledger.
- Share from the user's own CLAUDE_CONFIG_DIR when they set one (marked
  so Orca's injected value is never mistaken for it), and refuse a
  profile at or around it.
- Pin the one canonical profile path spelling in a test.

* refactor(claude): route launches through one account router, superset-shaped

Replace the routing service, owner interface, setup worker thread, reader-root
merging, persisted launch account and capability string with one
ClaudeProfileRouter: the pointer is written first and setup runs best-effort
after it (superset's order); a missing pointer means System default.

The claude shell function re-reads the pointer on every launch, refuses only
a selected account whose folder is missing, and prints a note when the user's
own CLAUDE_CONFIG_DIR overrides the selected account in that terminal.

Still dormant: claudeProfileRoutingEnabled() is false.

* test(claude): type router test settings instead of casting

* fix(claude): run account setup on a worker thread, never Electron main

publish() writes the pointer and starts setup in the background, so neither
startup nor an account switch blocks on a history merge. Each setup runs in a
one-shot worker (the profile-state backup worker's pattern); one setup per
account at a time, reused by later requests. A launch waits only for a folder
that was never set up, and refuses with a clear message if that setup fails.

* fix(claude): do not await the synchronous pointer publish

* refactor(claude): route WSL distros through a small guest router on the Step 2 shape

Replaces the WSL owner/transport/guest-inspect stack with ClaudeWslProfileRouter:
publish writes the guest pointer with one sh command and kicks Step 1's setup
best-effort; prepareLaunch checks the folder over the distro share and waits only
for a never-set-up folder; preparation returns main's WSL shape, so trust, rate
limits and readers need no new code. Setup runs as Linux in the guest on Orca's
pinned Node via a bundled helper (argv in, exit code out), without hooks.

Restores OpenCode's WSL runtime prep, git's wsl-host-failure, wsl-runner,
workspace trust, readers and account selection/registration to Step 2.
Names the guest pointer per Orca build so dev and packaged never share it.

* test(claude): give the routing launch test the merged resolver deps and handle shape

* test(claude): type the WSL routing mock's original() without an inline import()

* fix(claude-accounts): dedupe merged prompt history, drop drained copies, link setup folders by path

- Prompt-history drain appends only lines the shared file lacks, so a purge never re-adds lines.
- A set-aside history copy whose saved offset reaches its end is deleted on the next run.
- Setup folders link to the default home's own entry, not its resolved target.
- The profile gate and folder creation run once, in provisionClaudeAccountProfile.
- installHooks receives only configDir; drop a duplicate test key that fails CI.

* fix(claude-accounts): refuse a routed resume whose transcript is in another account; zsh claude function; setup timeout

- With account routing, a chat resume checks its transcript is in the launch folder; a missing one
  with a stored leaf refuses with historyInOtherAccount instead of starting fresh.
- The launch folder of a selected account comes from prepareLaunch(); the resolver stays for System default.
- zsh panes get the claude function like bash, fish and PowerShell (empty while routing is off).
- The setup worker is terminated after 60 s so a later launch can retry.
- Document that the setup marker means setup started, not finished.

* fix(claude-accounts): write the WSL account pointer before a launch returns; one relay bundle candidate list

- prepareLaunch awaits writePointer, so a missing or stale guest pointer cannot run another account.
- Startup's WSL republish runs inside serializeMutation, like rollback.
- relayBundleCandidates takes 'wsl'; the hook relay, browser relay and Claude helper use it, and
  wsl-relay-bundle-dirs.ts is gone.
- One setup-marker path helper for host and WSL; the guest pointer path is home-relative and only
  the pane value carries '~/'; drop a no-op esbuild external.

* fix(claude-accounts): refuse a routed resume only when the transcript is found in another folder

A transcript found in no known folder keeps the old stored-leaf resume.

* fix(claude-accounts): a WSL launch writes the pointer for the selection current at write time; bound the pointer read

A selection made while a launch waited on setup was overwritten by the launch's stale account.
A hung \\wsl.localhost read no longer stalls startup's serialized publish.

* fix(claude-accounts): record installed hooks as Orca-shared; skip symlink tests on Windows

After Orca installs its hooks into an account, record the account's hooks in
the settings ledger so a later run can still bring the user's own hooks in.
Tests that create real symlinks now skip on Windows.

* fix(claude-accounts): trim the which-account file in the PowerShell claude function

Co-Authored-By: Claude <noreply@anthropic.com>

* test(claude-accounts): spell the user's own config folder as an absolute path on every platform

Co-Authored-By: Claude <noreply@anthropic.com>

* test(claude): skip the POSIX-only WSL profile test on Windows

A WSL profile's data root is a POSIX path, so building one from a Windows
temp dir fails the absolute-path check there.

---------

Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
2026-10-07 01:25:41 -04:00
Brennan Benson 2f49377425 feat(native-chat): Grok as a structured chat over the Agent Client Protocol (#25225)
* Leave a stopped turn's running tools to the agent's own end

When another writer settles the open turn (a person's Stop), the assembler
now only stops that turn's text and cancels its pending prompts. Running tool
calls stay the agent's: a progress update or completion it reports after the
Stop lands as reported, and whatever is still running settles at the agent's
turn end for that turn, the next turn's open, or the session's end.

An agent's end for an earlier turn while a newer one is open no longer clears
the open turn's activity line or ends its anonymous reply. An unnamed end right
after a Stop ends the stopped turn instead of being dropped. The test rig's
restart no longer writes the dead assembler's window text, matching dispose.

* Pin that a stopped turn's running tools hold budget until the agent's end

* Type the stopped turn's tool progress update as a tool body

* List every event the assembler hands to the decision step

The type-aware lint requires an exhaustive switch with no default case.
Also retitle a Stop test to say what it asserts.

* End a running call as its turn's journal row ends

A call still running when its turn ends takes the state of that turn's
row: a row another writer settled first (a person's Stop) stands, so its
calls read interrupted whatever the provider's later end reports. The
no-ending path that settled calls from the Stop row is gone, since a Stop
now leaves running calls to the provider. Adds the two Spanish strings.

* Say why a Grok turn failed, and keep task rows in Grok's own words

A failed Grok turn ended with no reason on screen: the translator dropped
every copy of Grok's message. The failed turn now gets one status row in
Orca's existing "provider did not accept this message" words with Grok's
reason, read from whichever copy arrives first (the given-up retry, the
turn's end, the prompt's completion notice, or the prompt's error answer);
later copies only fill a reason the row still lacks.

A running background command no longer reads "Background task <id>
started": a task's summary is mapped only once it has settled. A monitor
stays a monitor when the agent reads its output: a frame that names no
kind keeps the known one, and a "[monitor" command is a monitor.

A prompt's turn is marked started, so a late frame for an ended prompt
neither reopens it nor becomes the active turn. A tool's turn is held in
one place at a time.

* Read a monitor from Grok's exact output prefix

* Word a failed Grok turn in Grok's own text, not as a refused message

A turn that started and then failed was told "The provider did not accept
this message", Orca's sentence for a message refused before its turn. The
row now reads as a Codex turn-ending error does: an error status row with the
provider's own words. With no words, the dialect names the failure ("Grok
ended this turn with an error." / "Grok usage limit reached."), else the
agent's display name does.

* Settle a stopped turn's running call as its turn row ended after a restart too

The restart sweep ended every running call by the death evidence alone, so after
a person's Stop with no proof the child died the call read failed under a turn
that read interrupted. The sweep and the live dead-generation settlement now ask
the same rule the assembler does: a call in a turn already settled ends as that
row ended; only a turn still running leaves its calls to the evidence.

* Keep the dead-generation settlement under the line cap

* Register the ACP schema verify step in the PR preflight phase test

* refactor(agent-session): one required agent registry; declarations admit what they claim

A StructuredAgentRegistry, built once from the {definition, adapter}
registrations, is now a required host dependency and the router routes with
it. The adapter interface loses its optional router-only capabilities?() and
definition?(); every reader (options at rest, thread goal, rewind, the
/compact handover) asks the registry. A live session still narrows rewind
through the adapter, and the host combines declared and narrowed in one
helper. The registry refuses a registration that declares compact, a thread
goal or rewind without the adapter method behind it.

/compact is admitted by the declared capability, so an agent that declares
compact:false gets the commandRefused fact instead of a thrown error.

Also: the cut-turn notice names the agent from the catalog, create-support
builds the account home through agentSessionAccountHome, and the turn
status text older clients read names the session's agent instead of
defaulting to Codex (Claude/Codex text unchanged).

* Read ACP permissions, session events and prompt errors through the protocol client's own types

The translator now reads a permission request with the client's lenient reader, a session update
with its session-event reader, and takes only the agent's own error answer as a failed prompt's
reason, so an Orca-side error never reads as the provider's words. Tests cover protocol values
newer than this build.

* refactor(agent-session): the router applies the declared rewind itself

The router holds the registry, so its rewindSupport answers the owner's
declared rewind narrowed by the adapter, in one place no reader can bypass.
The host-side combining helper is gone; readers ask the adapter they hold.

* test(agent-session): register the agents the merged-in tests now need

The adoption replay test builds its host over this build's registered agents
instead of an adapter definition method, and the stored-form test passes the
stored agents the record guard now requires.

* chore(agent-session): the registry is the one lookup; drop the test-only empty capability record

* fix(agent-session): keep "dismiss all" restart offers dismissed for the desktop

The desktop renderer calls the restart methods as a paired runtime client
without the registered-agents capability, so it was handed a Claude/Codex
audience and its "dismiss all" took the scoped path: no dismissal fence and
no clearing of unwritten teardown witnesses, so a late teardown write could
bring a dismissed offer back.

The audience is now derived against this host's registry: a client that can
show every agent the host registered gets none, and dismisses exactly as
before (one fence for every offer). A client that truly cannot show some
agent still gets a final dismissal for what it sees: each cleared session is
fenced on its own (bounded, superseded by a dismiss-all fence) and only the
witnesses of agents it sees are dropped.

One predicate now answers which agents a client renders, for both tab
projection and restart offers; the Claude capability rule is folded in.

* fix(agent-session): a changed agent definition never hides that agent's chats

A record was readable only if every handle used the transport its agent's
definition declares today, and its account variable was declared by some
registered agent. A later build that drives the same agent over another
protocol, or under a renamed variable, would have set every existing chat of
that agent aside: no tab, no history, no error.

Readability now asks only that the record agree with itself: its agent is
registered, every handle shares one namespace owned by that agent, and the
account variable is a well-formed name. Whether this build can drive it (the
chain's transport is the one the agent speaks, and the account variable is
that agent's own, since it becomes the child's environment) is decided in
performAttach, the one admission every start of every agent passes through,
and refused there as hostUnsupported. A record pinning another agent's
variable is refused the same way rather than passing because some agent
declares it.

* refactor(agent-session): each agent's registration says where it runs and which account it pins

createSupport, create and the model catalog's account read still chose a
location rule and an account resolver by name (Claude, Codex, else none),
so a new agent would have needed a third branch beside the registration
list that already decides storage, routing and publication.

Each runtime registration now carries supportsLocation(location) and
resolveAccountHomePath({ launchEnv, location, purpose, workspacePath }),
and Claude's and Codex's rules move into their entries unchanged: a read
still syncs no home and starts no bridge, a Codex launch still trusts the
workspace first, and WSL locations resolve as before. The list exists
before the host is built, so none of these reads installs the host, and an
agent the list does not hold is answered no without opening the journal.

* test(agent-session): drop a duplicate registry key; keep fences out of the capsule state

A spread input already carries the registry, so the explicit key was
overwritten (TS2783). The session-fence helper now returns only the fence,
not the capsule state it was given.

* fix(agent-session): a scoped dismiss-all persists no per-session fence

The per-session dismissal fence added a forever-persisted capsule field
with an arbitrary cap that served no caller: only the local desktop calls
the restart methods, and on a host whose agents it all shows it gets no
audience and takes the unscoped, fenced dismissal. A scoped dismiss-all
(a caller that cannot show some registered agent) keeps that agent's
offers, clears only its audience's unwritten witnesses, and writes no
fence; a late write from this process is serialized behind it. The field
never left this unreleased branch.

* fix(agent-session): refuse an attach whose agent is not the session's own

The start check judges the record's agent while the router starts the adapter of params.agent, and the attach wire let the two differ. Admission now refuses agent !== provider as requestMalformed, so the agent checked is the agent started. Every host-built attach already sets them equal.

* fix(agent-session): offer to start a chat only when the start would accept it

The restart offer, a failure's retryable flag and the pre-send check asked only whether the adapter runs the chat's location, while the start also refuses a record this build cannot drive; such a chat was offered Resume and its Retry failed forever. One host predicate, hostCanStartRecord (location support and agentDrivesSession), now answers all four; adapterSupportsRecord stays the reader gate, and an undrivable chat's offer is kept, not retired.

* fix(agent-session): the model-catalog probe uses a record's account only if this build drives it

A record-scoped catalog read started the agent's lister under the record's account path whatever variable the record pinned. It now uses that path only when agentDrivesSession holds, and otherwise resolves the account as for a read with no record.

* refactor(agent-session): the record store admits agent ids; comments say where transport is checked

The store only ever asked whether an agent is registered, so its admission list is now the registered ids; the definition is the one home of an agent's transport and account variable. Comments that said the record store checks transport, and one that cited a nonexistent function, now point at agentDrivesSession and the attach admission. The isPersistedAgentSessionRecord(value, agents) call shape and the test fixture the cross-version probe imports are kept.

* docs(agent-session): the record store admits the registered agents' ids

* refactor(native-chat): Grok's registration declares where it runs; ACP no longer borrows Codex's location rule

The rule a self-supervised agent child runs under (this machine, no WSL, Windows only with process
start-time proof) is its own module that Codex and the ACP adapter both use. Grok's registration
takes the full account-home resolver signature, and D3's tests build hosts with the agent registry.

* fix(native-chat): Grok follows the ACP runtime's request contract and the managed process's close

A request the agent or a Stop cancels is answered with the agent's own cancelled reply by the code that
owns it (the runtime no longer answers a silent handler), so a Stop needs no separate decline pass. A
permission answer still being saved when the agent stopped waiting is reported unconfirmed, since the
protocol already answered it cancelled. Cancelling the agent's own turn is the plain cancel. Request
rows are matched under their generation-scoped ids. A refusal's reason comes from the dialect's wording
path. The child drops its own stderr tail and close policy for the managed process's, and a close
whose process tree was not proven gone is reported as the adapter contract asks.

* fix(native-chat): a Grok chat Orca already holds resumes without writing what Grok replays

A chat with a saved Grok session reattaches with session/resume where the agent offers it, else
session/load. Either way the call runs inside the translator's load window, so what Grok sends while
it reattaches (its saved exchange, a task the dead process left running, ended by the restart) opens
no turn and writes no row; only context usage reads on. A reply an Orca or Grok crash cut short is no
longer completed from Grok's saved history: it reads like a Claude or Codex chat's, with the existing
notice. The attach window also closes after a failed attach, and a created session that session/resume
reports missing is replaced like one session/load reports missing.

The replay reconciliation is removed: the lane no longer reads the journal, and D3's replayed-input
grammar test and completed-turn check in the assembler go with it.

* refactor(native-chat): a failed Grok reattach needs no window close of its own; its lane is replaced

* test(native-chat): D3's merged tests use the shipped declarations and the launch options main requires

* fix(native-chat): typecheck fallout of the base merges; any agent's empty chat is reusable

Main's idle-empty-chat lookup and launch join now take any registered agent, as the rest of the
launch path does. The refusal check moved into the prompt turns and the prompt-block conversion beside
the turns that send it, keeping both files in their line limit.

* fix(native-chat): a Grok Stop ends the process once Grok settles its turn; the next send resumes

Grok's session/cancel ends only the running turn: work it already moved to the background keeps
running and can begin a turn of its own after the person pressed Stop. Stop is now a session
boundary, as it is for Claude: the cancel answers open requests and lets Grok end the turn, the host
waits a bounded grace for that, then ends the process; the next send relaunches and resumes.
The adapter's own bounded close of a turn Grok began is gone. Its named-turn check stays: the host
ends the session unless the provider declines a Stop naming a turn that has since ended.

* test(native-chat): a Grok Stop ends the process only after Grok answered the cancel

* fix(native-chat): Steer on a Grok card cancels the running prompt, then sends it

A send that reached Grok while a prompt ran was held in the adapter until that turn ended: Steer
on a queued card took the card out of the host's editable queue and meant 'send after this turn'.
It now cancels the running prompt (session/cancel; the session stays) and sends as the next prompt
once Grok answers the cancel, as the common pattern does; a steer behind another cancels it in
turn, so the last one runs. The adapter holds a send only while that cancel lands, so its general
held-send queue and its holdsDispatch report are gone (every send it holds has its turn open in
the journal). An older client's mid-turn send takes the same path. capabilities.steering is
unchanged and still unread.

* refactor(native-chat): a close or Stop cancels a start through the acquire's own abort signal

The host owns the acquire it runs, so it now owns its cancellation: each attach's acquire gets an
AbortSignal, aborted from outside the session's queue by a close and by a Stop admitted now (the
same admission rule as before). The optional abandonStart adapter hook, the router's fan-out to
every adapter and the ACP adapter's session-keyed start map are gone; the ACP adapter keeps an
unkeyed set of starts only so quit can prove their children gone, and keeps a failed start's
unproven child until its exit is proven.
The hook also let a later close ask that child again. The host now does that from state it holds:
a close of a chat with no live child whose record still names an owner process with no death
evidence asks the adapter to release it. The answer is not recorded as proof (the lease probe
does that), so an owner pid an earlier Orca left is never killed or marked gone. Claude and Codex
ignore the signal and hold no such child; their release is a no-op (tested).

* fix(native-chat): a Grok crash that closes stdout before its exit still ends with Grok's last words

On macOS and Linux the agent's stdout ends before its exit is observed, with or without the
supervisor's EOF forwarding, so the connection's loss closed the journal first and its error text
became the session's ended reason, dropping Grok's stderr. The reason is now read at the proven
exit: the agent's last words when it left any, else why the connection closed. The failure already
carried them. Comments that assumed the exit comes first, that early frames past the cap refuse the
start, and that dispatch re-checks image support are corrected.

* fix(native-chat): nothing Grok sends while a held chat reattaches is written, marked as replay or not

The reattach window relied on the dialect's replay verdict, and Grok's frames read as live unless
they carry isReplay, so an unmarked chat frame during session/resume opened a turn that never
ended. D3 now marks every frame inside the window as replay before the translator reads it, so the
translator keeps only context usage whatever the agent marked; options and commands are still
adopted. The translator's load semantics are unchanged.

* test(native-chat): a Stop after a resume finds no turn an unmarked old reply opened

* test(native-chat): a resumed Grok chat keeps its last context reading; the resume refreshes only the window

* test(native-chat): a Grok background task a Stop ended reads as stopped reporting

* refactor(native-chat): quit's stop of each start answers through one promise kind

* fix(native-chat): quit aborts every start the host has in flight before draining attaches

A Grok that never answered its handshake held quit until the start's own 60 s bound, past the
20 s quit deadline. The host's teardown now aborts each in-flight acquire (and any the drain
still begins), so the adapter's own quit controller and its map of starts are gone: a start
has one canceller, the host's signal.

* fix(native-chat): a Grok start's abort stops reaching its child once the start has returned

The listener stayed on the host's signal until the attach finished committing, so a Close in that
window killed the now-live child behind the host's back and it read as Grok crashing. The start
now detaches it when it ends; a later Close goes through the session's own stop.

* fix(native-chat): a close or Stop during any attach phase stops the start before it launches

The attach began its abort controller only after reconciling leases, resolving recovery and
probing the previous owner, so a close or admitted Stop in those phases reached nothing and Grok
launched anyway. The controller now begins first, and the acquisition checks it before asking the
adapter to start.

* test(native-chat): a close during the attach's owner probe asks no adapter to start

Also renames the close test after the hook it no longer exercises.

* test(native-chat): a close's re-ask closes a Claude or Codex child a failed cleanup left

The re-ask is not a no-op for them: when the adapter still holds the child its cleanup could not
prove gone, the close stops it again as a requested close, and Claude persists the handle of the
conversation it ran so the next send resumes it. Corrects the tests' and comment's wording; the
close awaits the re-ask, bounded by each adapter's kill ladder.

* fix(native-chat): Steer during a turn Grok began itself cancels it and sends once it ends

A send while Grok ran a turn of its own (a background task waking it) went straight to Grok, which
queued it behind that turn where Orca could no longer withdraw it, while Stop treated the same turn
as the running reply. The send now waits as a steer, the turn is cancelled once, and the message
goes when the turn ends; a Stop withdraws it and an exit rejects it as never sent.

* test(native-chat): a steer whose cancel Grok never answers ends Grok and is rejected as never sent

Pins the bounded steer cancel kept from the runtime: past the bound the connection closes, the
running reply reads unverifiable, Grok's end reads as its exit, and the waiting steer is rejected
as never sent.

* fix(native-chat): a Grok crash stays a crash when a stop lands before its exit is proven

After the connection broke and the close could not prove Grok's exit, any later stop Orca asked
for (the next start, a Stop, a Close) marked the child as closed by Orca, so the crash read as a
requested close and Grok's last words were dropped; a send meanwhile was recorded unconfirmed.
The connection loss now decides the cause, and a send on that session is rejected as never sent.

* test(native-chat): fixtures this PR's registered Grok and desktop capability made stale

CI's unit shards failed on tests outside the PR's own lists. Each encodes something this PR changes
on purpose: Grok is now a registered agent (the seam test's unregistered agent is now Cursor); the
desktop now advertises registered agents (the restart-offer tests' older client drops that
capability explicitly); the attach context carries the start's abort controllers (the forget-status
double gains them); and the ACP real-host test rig sends to the host directly (listed beside the
other real-host rig in the send ratchet).

* fix(native-chat): a start quit stops is not the queued message's start failure

With quit now aborting a start it would have waited for, the delivery step recorded the aborted
start as the message's failure ("couldn't restart"). After quit has stopped delivery, the step
leaves the message to quit, which settles it as a close does ("The chat closed before this message
was sent."). The test that pinned quit waiting for that start and stopping its child now pins that
nothing is launched behind quit.

* fix(native-chat): a message sent after a Stop or close aborted a start gets its own start

A start the host aborts (an admitted Stop, a close, or quit) returned its refusal to the delivery
loop, which then rejected whatever was queued at that moment with "couldn't restart", including a
message the user sent after the Stop. The attach now reports that the host aborted it, and the loop
re-derives from the journal instead: what the Stop or close withdrew is already settled, a message
accepted since gets a start of its own, and quit's next step stops the loop. This replaces the
quit-only carve-out with the same rule for every abort and every agent.

* test(native-chat): the message sent after an aborted start is answered, so no settlement outlives the test

* fix(native-chat): a Grok model pick Grok never answers no longer holds Stop or Close

The pick runs on the session's queue. It now registers in the host's out-of-queue
abort registry beside a start, so a close, an admitted Stop or quit abandons it, and
the ACP adapter bounds it at 30 s like Claude and Codex. A late answer is still adopted.

* fix(agent-launch): a phone's launch opens a terminal for an agent whose chat it cannot show

agent.launch now reads the caller's capabilities by the rule tabs and restart offers
use (clientRendersStructuredAgent). A phone without registered-agents.v1 gets Grok as
a terminal again, as on main; the host's own callers and desktop clients are unchanged.

* fix(acp): strip every agent hook variable from the ACP child, from the shared list

ACP_CHILD_ENV_TO_DELETE was a second copy of the hook runtime keys that missed
ORCA_AGENT_HOOK_TRANSPORT; it now spreads AGENT_HOOK_RUNTIME_ENV_KEYS beside the pane
identity keys.

* refactor(native-chat): the mutation context carries the provider-wait registry itself

Keeps the host file within its line limit; one field instead of two closures over it.

* fix(agent-launch): agent.launch.v2 still vouches for Claude and Codex chats

The caller rule from the previous commit also turned Claude and Codex into terminals
for a client advertising only agent.launch.v2, whose contract says it opens a chat
(mobile retry-authority tests). Only an agent beyond those two now needs the client to
read it (clientRendersStructuredAgent); the test fixtures go back to what they were.

* refactor(native-chat): drop saved-history adoption from the timeline assembler

The common pattern discards the history a provider replays while loading a
session, so the assembler has no use for an input.history event.

* refactor(acp): drop session/load history adoption from the translator

The common pattern discards the history an agent replays during session/load,
keeping only what it says about the context window. Remove the adoption path
(acp-history-adoption.ts, the adopt option, and the historical background-task
liveness rewrite it fed) so load replay is always dropped except usage.

* refactor(native-chat): a pending input is only Orca's send now

Review follow-up to the adoption removal: drop the comment naming the
provider's saved message, and make requestedAt required since every pending
input comes from input.accepted.

* test(acp): keep the task-result status table on live frames

Review follow-up to the adoption removal: the result-status mapping was only
tested through adopted history, so run the same table on live frames, and
cover an unmarked task notice during a load being dropped.

* test(acp): a frame helper for a shell command Grok is running

* fix(acp): a Grok crash settles through the host's provider-exit batch, scoped to the turn it ended

A Grok crash ended the journal unverifiable before the adapter reported the exit, so the host's
provider-exit settlement found no running turn and wrote nothing: the adapter's failure (with
Grok's last words) never reached the journal, and a later stale-session pass wrote a bare,
thread-scoped cut-short row, so the partial reply was not folded as Claude's and Codex's are.

At a proven exit the ACP lane now ends its running turn interrupted at the exit instant, as the
host's exit contract expects of a child's own translator (Codex's does the same). When Grok's
stdout closed first (every POSIX crash), the turn is unverifiable only until the exit is proven:
the host's provider-exit settlement now takes the exit as proof naming the child's fence and
revises what that child left unverifiable in the same batch, with the turn-scoped row and the
adapter's failure. Claude and Codex write no unverifiable turn of a live child except a command
whose hand-off is in doubt; that turn is now revised at the exit instead of at the next open.

* test(acp): a crash seen first leaves the host no Grok turn to revise

* refactor(native-chat): what a gone generation left unfinished gets its own module

The settlement file passed 300 lines with the exit-proof revision. The unfinished-work reads
(capture, interrupted-by-the-exit, in-progress) are their own concept and move out unchanged,
apart from the exit proof they now take.

* refactor(native-chat): a watched exit revises what its child left unverifiable without reading Stop marks

An exit's own instant is the turn's end, so the revision needs only each row's fence: the
settlement's journal type gains itemFence alone, and the host test fakes say so.

* test(native-chat): drop the duplicate itemFence on the fake that already had one

* test(claude, codex): an exit whose stdout ended first still reports as it always did

The provider supervisor now ends Orca's stdout when the agent's ends, so on every crash EOF
arrives before the exit is seen. Claude's and Codex's connections report nothing at EOF and
report the exit, with its usual reason, once it is seen.

* fix(acp): reopen a chat with session/load, as the common pattern does

An agent that offers both now reloads its session instead of resuming it; the
reattach window still discards what it replays except context usage.

* fix(acp): drop the 60 s handshake bound; an abort fails the start's waits at once

Neither common design bounds an ACP handshake: Close, Stop and quit end a start
that never answers. The abort now also closes the connection, as a kill there
does, so the start settles even before the child's exit is proven. The
host-stopped start refusal only this bound produced goes with it; the idle
sweep keeps its words.

* fix(acp): a Stop naming an ended turn follows Claude's rule

It still stops nothing while another turn is live, but in the gap before a
follow-up's turn opens, which no client can name, it now stops what is in
flight and the session ends, as a Claude Stop does.

* fix(native-chat): a close no longer re-asks a failed start's unproven child

Neither common design retries that stop at Close, and Orca's Claude contract
re-asks only at the next start and at quit. The ACP adapter keeps the child
until its exit is proven and asks it again there, as Claude does.

* fix(acp): a message sent during a turn the agent began itself goes at once

Both common designs send it straight to the agent with no cancel; only Orca's
own running prompt is steered (cancelled, then re-prompted).

* test(native-chat): dismiss-all through a remote client's audience keeps a newer Orca's offer

Uses an audience production sends (one that cannot show every agent), per review.

* fix(acp): launch Grok as `grok agent stdio`, without the update and leader flags

The common pattern passes neither --no-auto-update, --no-leader nor
GROK_DISABLE_AUTOUPDATER; full access still adds --always-approve.

* fix(acp): an agent that ends its stdout, or answers unreadably, is not a lost connection

As in the common pattern, only a broken stdin (or Orca's own close) ends the
agent; one that closed its output but can still be written to stays until a
Stop, a close or its exit. The provider supervisor goes back to its base
content, so Claude and Codex no longer get the forwarded stdout end either.

* fix(native-chat): a person's close joining a failed one still binds the turn its child end cuts

On main every close of the chat writes its own Stop and settle. Here a later close joins the
first and writes no row, and the first's settle closed when its kill failed, so a turn that opened
in between and was cut by the next close read as failed. A person's close joining a person's close
whose Stop opened a settle now reopens that settle until its attempt is done.

Tests: a close whose kill failed still closes its settle; a turn opened between a failed close and
the next reads as the person's cancellation (each fails without its half of the fix).

* test(native-chat): Grok opens as a chat only behind the structured-chat setting

agent.launch and orchestration worker-start read the same setting as the
renderer route; pin both states for Grok on each. The setting's description no
longer names only Codex and Claude, in every catalog.

* docs(acp): generic ACP comments say what holds for every agent, not Grok

Stop ends the session for every ACP agent, as in the common pattern; the
adoption hook comment goes (adoption is not planned); a failed start's child is
retried at the next start or quit.

* test(claude, codex): type the EOF-before-exit test's streams; the supervisor no longer forwards EOF

The Claude test wrote to the child's stdout and stderr through their Readable
type, which the node typecheck rejects; it now holds its own PassThrough
streams. The comments no longer credit the reverted supervisor change.

* feat(acp): a steer's cancel asks once and never ends the agent

The runtime had one cancel: send session/cancel, wait at most 10 s for Orca's prompt to settle,
then close the connection, which ends the agent. A steer used it too, so a slow agent lost its
process just because the person added a message. requestSteerCancel() now sends session/cancel
once per prompt, cancels the agent's open requests and answers later permissions cancelled, and
never bounds or closes: the prompt's own reply ends it and the steer's prompt follows. cancel()
stays the Stop: bounded, then close. A Stop after a steer still bounds and closes. Both cancel
paths move into acp-prompt-cancel.ts over one cancel channel.

* chore(native-chat): keep the record store and recovery capsule under max-lines after the main merge

* fix(acp): a repeated steer shares the cancel in flight; say what the caller owns

Per review: a second steer before the first write lands returns that write instead of resolving
early. The steer's JSDoc says the wait for the prompt's reply is unbounded and that a prompt that
fails instead must not take the steer until the caller rebuilds the session; the Stop's says a
prompt that settles in time leaves the agent for the Stop's owner to end. The steer test now gives
the runtime a handler that would allow: the open permission's signal aborts and the late one never
reaches it.

* fix(native-chat): drop the stopDelivery the A3 merge doubled

* fix(acp): a steer's cancel asks Grok once and never ends it

A steer now uses D1's notify-only cancel. Two messages sent during a reply Grok began itself
cut that reply, as the common pattern does, and then both run; before, the queued first
message could not answer the bounded cancel and Orca ended Grok although Grok answered.
A Stop keeps the bounded cancel and its 4 s grace.

* fix(acp): a permission Grok asks with no prompt of Orca's running is declined

During a turn Grok began itself nobody asked it to act, so the request is answered
cancelled at once instead of opening a card that waits, as the common pattern does.

* fix(acp): a Grok that dies while starting is reported with its own last words

A dying process's stdout ends before its exit is seen, so the start failed as a closed
connection and Grok's stderr was lost. A start whose connection closed now waits, bounded by
the Stop grace (or a Close/Stop), for the exit before it is told.

* test: a Stop after a steer sends its own cancel; drop the import the A3 merge doubled

* test(native-chat): main's Stop-note test builds its turn context with the agent registry

* test(claude): say why the close test's fake child cast is safe

* test(native-chat): build the Stop-opened-turn test's identity and turn context the current way

The test (#25056) landed before the opaque provider handle (#24991), so main still built
the old {kind, threadId} handle; the turn context also needs this branch's agent registry.

* test(native-chat): build the Stop-opened-turn test's identity with the opaque handle

The test (#25056) landed before the opaque provider handle (#24991), so main still built
the old {kind, threadId} handle.

* test(ratchet): require src/main/provider-process now that it has landed

* test(native-chat): keep main's opaque-handle import in the Stop-opened-turn test

Main's #25706 and this branch both added the import at different lines; the merge kept both.

* test(native-chat): keep main's opaque-handle import in the Stop-opened-turn test

Main's #25706 made the same fix as this branch at a different line; the merge kept both imports.

* feat(native-chat): record a fresh provider conversation that replaced one the agent could not restore

A chat whose saved conversation the agent cannot reopen can now continue in a
fresh one: the handle chain records the new conversation as a creation that
replaces the lost one (which, why, and when), keeping every earlier link.

Rows keep a shape older builds read: the stored chain starts at the latest
replacement and carries the earlier links inside it.

* test(native-chat): build this stack's journal identities with main's opaque provider handle

Main's #24991 replaced the {kind, ...} handle with {transport, agent, nativeId}; three test
files from this stack still wrote the old shape. Same lines the downstream ACP branch uses.

* docs(acp): every reattach drops the agent's replay, not only for a chat the journal holds

* refactor(native-chat): store a replaced conversation flat; refuse it where older builds read the row

Older builds only read Claude and Codex records, so the nested stored form
protected rows no replacement can reach while adding a cap mismatch after a
downgrade. Store the chain as held, refuse a replacement in a Claude or Codex
chain until one has a stored shape older builds read, and refuse a supersession
key on a replacement that names no creation in the chain.

* refactor(native-chat): read hosts' structured agents from the app-shell services

Main grew the startup hydration hook to its line limit; the host agents sync is an app-lifetime subscription like the structured session tabs sync beside it, so it moves there.

* Use current provider handles in transition tests

* Use current provider handles in timeline fixtures

* test(native-chat): prove replacement rows survive downgrade and re-upgrade

* Require the ACP directory in the runtime import check

* test(ratchet): require src/main/acp now that this PR lands it

* feat(acp): a saved session the agent cannot reopen continues in a new one, with one warning row

When session/load (or session/resume) of a saved ACP session fails, the chat starts a new session and records it as a creation that replaces the lost one (#25747's 'replaces' link), and writes one warning row that the agent no longer remembers the earlier messages. A created session the agent reports missing is still superseded silently; a signed-out agent or a start that is over (Close, Stop, a lost agent) still fails the start.

* chore(acp): rewrap the acquire header comment

* test(acp): a start closed while the agent reopens fails without opening or announcing a new session

* Let ACP connections own their supervised agent process

* Preserve ACP cleanup evidence and isolate exit observers

* Expose ACP cleanup observations and type the permission fixture

* refactor(native-chat): the registered-agents capability lives in its own module

Main's growth put protocol-version.ts one counted line over its 300-line limit once the capability
was added; like main's other per-feature capabilities, it now has its own module, and importers read
it from there.

* refactor(acp): one connection owns the Grok process and its protocol

D3 now opens each ACP agent through createAcpAgentConnection (ACP-ALIGN #25810): one object spawns the
process on the execution host, owns its stdio and protocol, and reports its proven exit. It is built and
tracked before the handshake, so a start's abort (Close, Stop, quit) still reaches it, and a failed start
keeps that same connection for the next close to retry rather than spawning another process.

Deleted: the spawnAcpStructuredChild wrapper and its test, the raw-stream runtime assembly, the caller's
exit -> runtime.close wiring, the stdout-EOF heuristic (the connection no longer treats stdout EOF as
exit), and the 10 s steer/Stop cancel bound with requestSteerCancel. Reader control maps to
pauseReading/resumeReading; a close is connection.close after the host's existing 4 s Stop grace.

The adapter owns what the protocol no longer does: one session/cancel per running prompt however many
steers arrive (cleared with that send's settlement, retried after a failed write), and a Stop or steer
answers every open agent request the person has not already answered with the agent's own cancelled
reply. An answer already being saved when the Stop lands is sent.

Tests: blocked cancel write never holds Stop's grace, two quick steers send one cancel, a failed cancel
write is retried, a real process exiting while a child holds its stdout ends the session, and the
existing start-abort, retention, crash, connection-loss and reload-failure suites on the new rig.

* fix(acp): Grok signs in on its own machine with its API key or cached sign-in

When Grok reports that it needs authentication, Orca now names a sign-in method on the machine Grok runs
on, read from the same environment Grok was launched with: xai.api_key when XAI_API_KEY is set there and
Grok offers that method, else cached_token when Grok offers it, else none and the chat keeps the existing
not-signed-in refusal. The rule lives in Grok's launch spec; the adapter applies any agent's rule for new
and reopened sessions through the protocol client's caller-named method (authenticate, then retry once).
No new sign-in UI; interactive methods are never chosen.

* fix(acp): the adapter decides which of Grok's requests reach the person

The turn owner now admits every agent request, permission or question, from its own turn state: a
request reaches the person only while Orca's prompt runs and no steer or Stop is cutting it short (a
question may also come from a turn Grok began itself, until a Stop). Anything else gets the agent's
own cancelled reply and opens no card, so a question arriving after Stop or during a steer never
appears. A steer, like a Stop, withdraws the requests already open; an answer already being saved is
still sent. The protocol client's abort-on-cancel path is no longer used: after the connection
change its request signal aborts only when the connection closes.

* fix(acp): a plan Grok proposes shows as a plan, with no approval card

When Grok leaves plan mode it asks the client to approve its plan (x.ai/exit_plan_mode). Orca showed a
blocking 'Approve plan / Request changes' card for it; the common pattern has no such gate. Now the
plan goes into the chat's existing Plan row (the plan-document status row Codex and ACP plan updates
already use) and the request is answered at once with 'abandoned' plus feedback telling Grok to stop and
wait for the person's feedback or a request to implement it in a later turn, so nothing is approved on
the person's behalf. Dialects gain settleRequest for requests answered without asking anyone.

* fix(orchestration): worker-start opens a Grok worker in a terminal, as before

With the structured chat setting on, worker-start decided 'structured' for Grok and then the structured
worker factory (Claude and Codex only) refused it, so the start failed; main opened a terminal Grok
worker. Worker-start now decides with no registered agents beyond Claude and Codex, so Grok gets a
terminal worker as before. agent.launch and the app's own launches still open Grok as a structured
chat. Temporary until structured workers take registered agents.

* fix(acp): a prompt answer Orca can't read ends the turn instead of hanging it

A session/prompt rejection that was not the agent's own error answer (an answer that fails Orca's
schema, or one too large to read) left the turn running: the next message became a steer with nothing
to cancel and was never sent or settled, and Stop waited its full grace. As in the common pattern, any
prompt failure now ends the turn as failed (a failed-turn row without words, since none are the
agent's) and settles the send, so the next message goes. Only a closed connection keeps the send
running, for the connection-loss path to settle.

* fix(acp): send Grok's prompt-identity extension only to agents that echo it

session/prompt carried _meta {promptId, requestId} for every ACP agent, though only Grok's dialect
echoes it (injectedPromptIdentity). Now only an agent whose dialect declares it gets the extension;
other ACP agents get a plain prompt.

* refactor(native-chat): the registered-agents capability lives in protocol-version again, as on main

This reverts 0ef6d21815. That commit moved the capability to its own module only because main's
protocol-version.ts was then one counted line over its limit; main now defines it there itself within
the limit, and main's new restart test imports it from there. Main's test also reads the desktop's
capability list as an older client; on this branch the desktop advertises registered agents, so its
older client is that list without this one capability.

* test(acp): read the sign-in method with a schema, not a type assertion

* refactor(native-chat): composer transport and Stop control in their own modules

Main's rewind change (#19338) brought NativeChatStructuredSession.tsx and use-structured-agent-session.ts
to their line limits, leaving no room for this branch's image-acceptance and unpublished-Stop lines.
The composer's transport (sends, commands, options, image acceptance) moves to
use-native-chat-structured-composer-transport.ts, and whether Stop shows and what it does moves to
structured-agent-session-stop-control.ts. Behavior is unchanged; the runtime cast on the composer's
'local' | 'remote' is now a typed return.

* test(native-chat): read registered agents by id, as main's structuredAgentsReadBy now takes

Main's A3 squash changed structuredAgentsReadBy to take agent ids; this branch's test still passed
{ agent } objects (CI typecheck TS2322).

* fix(acp): a first reopen warns when Grok forgets a chat that exchanged turns

A Grok session the chat created was treated as one Grok never saved, so
when Grok reported it missing on the chat's first reopen, Orca swapped in
a fresh session silently, even after completed exchanges: the person saw
the old messages while Grok had forgotten them.

The launch now counts a created session as never saved only when the
chat's journal, read at the failed reopen, holds no turn of that Grok
session; anything else, an unreadable journal included, takes the normal
path: the fresh session is recorded as replacing the old one and the one
warning row is written.

* fix(acp): a start writes the warning row an earlier attach failure dropped

The row saying Grok forgot the chat was written only into the attach's
deferred sink, while the fresh session's link was saved earlier. An attach
failure, quit or crash in between dropped the row forever.

Every start now derives the owed rows: each conversation the chain says was
lost to a failed restore gets its row unless the chat's journal already holds
it. The row now names the lost conversation rather than the fresh session, so
a row an unused replacement wrote still counts after Grok supersedes it.

* fix(acp): a question during a turn the agent began itself gets its cancelled reply

A question or other card-opening request the agent sends while no prompt of
Orca's runs (a turn it began itself, as when a background task wakes it) now
gets the agent's own cancelled reply and opens no card, the same rule
permissions already follow there. Nobody is waiting on that turn. A plan the
agent shares in it is still shown.

* test(native-chat): import the unfinished-work capture from the module that owns it

Main's reasoning sweep test (#19221) imported it from the dead-generation settlement, which
this branch split it out of.

* refactor(native-chat): keep the session host within its line budget after main's Stop work

Main's #25949 left the host at exactly its 300-line budget, and this branch's acquire-abort wiring
adds one line. The reveal module now comes in as a namespace import, as the host already does for
its other helper modules, and the earlier reorder of two type imports is undone.

* test(codex): move the stdout-before-exit test into its own file

Main's connection test file is at its 800-line test budget, and this branch's exit-order test
pushed it over (CI lint, max-lines).

* fix(acp): cancel a running turn before a close, dispose or quit ends the agent

Closing a tab, disposing a session or quitting while an ACP agent's turn ran
killed the process without asking the agent to cancel first. A requested
close now does what a Stop does: withdraw the agent's open requests and held
steers, send session/cancel once, and wait for the turn to end, bounded by
the Stop's grace (4 s), before closing the process. A close that follows a
Stop sends no second cancel and waits only what is left of that Stop's grace.
An idle close, a lost connection and a sink-failure force close are
unchanged. The cancel-and-wait moves to acp-structured-stop.ts, shared by
Stop and close.

* test(codex): pass resolveLaunchArgs in the stopped-send-order test so typecheck passes

Main's typecheck fails here too: #25721 made resolveLaunchArgs required and
this test (#25051) predates it. Same line, same place as the open main fix
(#25977), so merging main after it lands is a no-op.

* test(native-chat): run the close-aborts-start test on the Node runtime, as its SQLite journal fixture requires

* refactor(native-chat): the Stop control reads the outbox itself, so the chat hook stays within its line limit

* Keep the session host under the line cap after main's two new delegates

Pass the lifetime's conversation opener to reveal directly; it is already passed unbound to the mutation context.
2026-10-06 22:09:26 -07:00
8464b151c7 Claude account profiles: dormant routing and consumers (Step 2 of 4) (#24351)
* feat(claude): add dormant profile setup and history sharing

* fix(claude): make profile setup one gated, typed, fail-safe entry

Review round 1 of the dormant profile setup found that the pieces could
be called without their safety checks, that one failed write or an
unreadable bookkeeping file could silently stop sharing for good, and
that Windows prompt history could bring back history the user cleared.

- One entry, provisionClaudeAccountProfile: the profile gate (namespace,
  no linked components, outside ~/.claude and ~/.config/claude, and an
  ownership marker beside the home naming the account and target) runs
  first and refuses before creating anything; then history sharing,
  config provisioning, and the hook install after the settings merge.
  Results come back per surface with closed warning codes instead of
  message text.
- The sharing ledger is keyed by surface name, records a value only
  after its write succeeded, and an unreadable ledger starts empty and
  is rewritten instead of blocking every surface.
- The profile state file goes through the same locked writer as folder
  trust (Claude's <file>.lock plus the in-process queue), generalized as
  updateClaudeGlobalConfig. Onboarding and trust are still applied when
  the personal state file is unreadable.
- WSL descriptors build guest POSIX paths; the state-file path style
  follows the injected platform.
- Orca's managed statusLine has one owner in a profile: the settings
  merge never shares it, a user's own statusLine is shared over it, and
  the profile installer follows the default home's slot so a default
  opt-out reaches every profile. remove() takes the same destination;
  the remote installer cannot accept one.
- Prompt history compares file identity (bigint dev+ino) on every
  platform, never drains the shared file into itself, drains retained
  copies in generation order, never reuses a stale cursor, and on Windows
  keeps a replaced default's old copy aside instead of replaying it.
  Directory merges keep going past a failed entry.

* fix(claude): share the user's own hooks and keep merged history whole

A user's own Claude hooks in ~/.claude (notifications, formatters) did
not run under a managed account, because the whole hooks key stayed
private. They are now shared like any other settings key: Orca's own
hook entries and its managed statusLine are stripped from both the
personal value and the profile's current value before the per-key
ledger comparison, so they never travel through the merge and never make
the key look user-owned. Orca entries already in the profile are kept on
write, and the profile hook installer adds them on top as before.

Prompt history: merged bytes that lack a final newline are terminated,
so Claude's next record no longer fuses onto the last merged line. When
a CLI rewrote the profile's history file (old records plus new), only
the lines past the part it shares with the default history are added,
instead of the whole file again.

* fix(claude): close review round 2 gaps in profile setup

Hooks and statusLine sharing:
- When ~/.claude holds only Orca's hook entries, the user's shared hooks
  now read as an empty value instead of a missing key. Removing the
  user's last own hook in ~/.claude therefore reaches profiles that
  never edited it, and deleting the only shared hook inside a profile
  stays deleted.
- A custom statusLine Orca shared, and the profile never edited, goes
  away when the default home drops it. When a shared custom line
  replaced Orca's line in a profile, the profile's statusline marker is
  dropped so Orca's line comes back once the default returns to it; a
  profile that opted out stays opted out. No other key gains deletion.
- install/remove/getStatus with a profile directory refuse when it is
  the default home, or its settings.json resolves to the default one,
  instead of editing System Default's hooks and opt-out state.
- The profile statusline rule reads the default settings under the
  userHome passed to the setup entry, not os.homedir().

Profile state and ownership:
- A malformed `projects` value skips only folder trust (new warning
  code trust-refused); onboarding and shared keys still apply.
- The ownership marker stores only host-local facts (account, runtime,
  distro). The execution host id is the caller's view of the host, so
  it stays in the in-memory descriptor and is not compared.

Prompt history interruption paths:
- With no cursor yet, a retained copy starts past the bytes it shares
  with the default history, so an interrupted share no longer replays
  the whole history.
- A retained name for the shared file itself is removed with its cursor
  instead of lingering until a later scrub makes it look new.
- The Windows link record is read three-state: unreadable stops the
  share instead of reading as "no link". If the record cannot be
  written after linking, the fresh link is undone.
- An unreadable retained copy is reported and no longer blocks linking.

* build(cli): list the new Claude hook modules in the CLI project

hook-service.ts and hook-settings.ts are compiled into the packaged CLI
project, which lists every file explicitly. The statusline policy and
profile destination modules they now import were missing, so the CLI
typecheck failed with TS6307. The CLI still loads hook-service through
the existing managed-agent-hook-controls build entry, which bundles
both modules; neither imports electron.

* fix(claude): close review round 3 regressions in profile setup

- A profile whose hooks hold only Orca's entries and that sharing never
  recorded is no longer treated as a user edit, so the user's first own
  hook in ~/.claude reaches it (for example when the profile was set up
  before ~/.claude had any hooks).
- A retained prompt-history file is removed as a second name for the
  shared file only when the default history does not itself link to it;
  otherwise it holds the only copy and is kept.
- Default-home checks compare file identity: the profile hook
  destination check uses device and inode, and the profile/default
  separation check resolves on-disk case, so a case-only alias of
  ~/.claude is refused on case-insensitive filesystems.
- A test pins that an unreadable leftover session tree no longer blocks
  linking.

* fix(claude): let shared keys leave a profile when ~/.claude drops them

QA found that removing a setting from ~/.claude never reached a managed
account: deleting the whole `hooks` block left the user's hook running
there. Only statusLine followed the default away.

Every shared key now follows the same rule through the existing per-key
ledger: when a key disappears from ~/.claude/settings.json (or
mcpServers/theme from the personal state file), it is removed from the
profile if the profile still holds exactly what Orca last shared. A
value changed inside the account is kept. Keys Orca never shared,
including denylisted ones, are never touched. Deleting the whole hooks
block removes the user's shared hooks and keeps Orca's own entries. A
missing source counts as empty; an unreadable source removes nothing.

* feat(claude): add dormant profile routing and account consumers

* fix(claude): drop the dormant profile selection RPC; clients negotiate by capability

Restores the inline mobile allowlist so its source-scan guard sees every
accounts.* method again, and the generated params catalog to generator order.

* fix(claude): guard the claude shell function and honour a hand-exported config dir

The function is defined only in a routed pane where claude is a real
executable (the codex function's guard), re-reads the pointer only while
CLAUDE_CONFIG_DIR is unset or still Orca's injected twin, accepts Git Bash
drive paths, and starts on its own line after the fish/PowerShell codex text.

* fix(claude): spawn-time profile env, total account listing, setup at lifecycle triggers

Round-1 review fixes for the dormant profile routing:
- Panes get the selected profile's CLAUDE_CONFIG_DIR plus an Orca twin at
  spawn, so nested shells and scripts inherit the account; System Default
  injects nothing and its home is the inherited CLAUDE_CONFIG_DIR.
- An absent routing owner is System Default, never a throw; AI Vault and
  session-search scans receive profile roots from their parent, and the
  capability is advertised only where an owner is installed.
- Account listing never throws: per-account readiness, a stale pointer is
  republished in the background and reported on the snapshot.
- Profiles are set up at select and startup; a launch only sets up one that
  never was, and a worker fault on a prepared profile is a warning. The
  Claude version probe is cached per binary identity.
- Pre-trust goes through the existing deadline- and realpath-guarded writer
  against the launch env's profile config.
- Skill discovery keeps a caller's Claude root and a broken Claude selection
  no longer fails other providers.
- The durable record carries a provider-neutral launchAccountHome, read
  through one helper by the launch fallback and the model catalog.

* test(claude): pin the version-probe cache, launch-account record and temp-home readers

* test(claude): pin dormant bash rc text alongside fish and PowerShell

* test(claude): read the fish launch init without a nullable index

* fix(claude): withdraw the profile pointer when a selection cannot be published

A pointer left naming the previous account would launch it silently; a
missing pointer makes the claude function refuse visibly. A newer selection
that raced the failed one keeps its pointer.

* fix(claude): read the fish profile pointer with read -z for fish older than 3.4

Shell tests skip system config and abort unless claude resolves to the fake.

* fix(claude): only the newest publish withdraws the pointer; total config dir lookup

- An overtaken publish that fails leaves the newer selection's pointer.
- The runtime config dir falls back to the legacy home for an unresolvable
  account or a WSL target, so skill roots never fail for other providers.
- WSL guest reader roots merge verbatim, never realpathed on this thread.
- History readers include ~/.claude, where step-1 setup pools profile history.
- System Default ignores a config dir an outer Orca injected (twin-marked).

* fix(claude): System Default launches and probes use the structured create resolver

A Claude agent-env CLAUDE_CONFIG_DIR the create path stored is now the home
the launch pins and the model probe accepts.

* test(claude): type the System Default launch record as an agent-session record

* fix(claude): install profile hook scripts under the setup job's home

A worker thread's os.homedir() ignores its own env, so the hook and
statusline scripts now go under the home the job names. The worker test pins
the process HOME to a sentinel, refuses to run unless the worker sees it, and
asserts nothing lands there.

* test(claude): skip shell cases whose shell the runner lacks

* fix(claude): remove env vars in the PowerShell claude function instead of setting null

On .NET 9+ (pwsh 7.5+) SetEnvironmentVariable with $null creates an empty
variable, so stripped auth vars reached claude as empty strings and the
restore left CLAUDE_CONFIG_DIR empty in the user's session.

* fix(claude): give plain fish tabs the claude function through the codex hand-off

Main now gives a plain fish tab Orca's codex function through a vendor_conf.d
snippet instead of a -C init. The claude function only rode the -C path, so a
plain fish tab would not re-read the account selection per invocation once
profiles are on. Define it at the first prompt beside codex; it stays empty
while the profile gate is off.

* fix(claude): share personal rules, themes, workflows and keybindings into account profiles

A managed account launches Claude with its own config folder, so user-level
rules/, custom themes/ (which a shared `custom:<slug>` theme points at),
personal workflows/ and keybindings.json silently stopped applying. Link the
three directories like skills and commands, and copy keybindings.json with the
same edit-preserving ledger as CLAUDE.md. routines/ stays unshared: routines
belong to the claude.ai account and the folder holds per-run state.

* test(claude): wait for the running child to read its account before switching

The test switched the selection after a fixed 20 ms, so under load the backgrounded claude
had not yet read the pointer and picked up the new account. The stand-in now marks when it has
started, and the test waits for that mark (bounded) before switching.

* fix(claude): import the personal CLAUDE.md into account profiles instead of copying it

Claude also loads ~/.claude/CLAUDE.md as a parent folder's memory for any project under home,
so a copied account CLAUDE.md made every such session read the user's instructions twice
(checked live with Claude 2.1.288). An @~/.claude/CLAUDE.md import resolves to the same real
file, which Claude loads once from home, from projects under home and from folders outside it.

* refactor(claude): simplify account profile setup toward the prior art

- Windows keeps each account's history private; drop the hardlink, link
  record and conflict-copy machinery that only Windows reached.
- Share hooks and statusLine as ordinary settings keys: Orca writes the
  same entries into every folder, so the installer finds them present.
  Drops the Orca-entry carve-out, the per-profile statusline follow
  logic and its marker.
- Unreadable ledger is just an empty ledger.
- Share from the user's own CLAUDE_CONFIG_DIR when they set one (marked
  so Orca's injected value is never mistaken for it), and refuse a
  profile at or around it.
- Pin the one canonical profile path spelling in a test.

* refactor(claude): route launches through one account router, superset-shaped

Replace the routing service, owner interface, setup worker thread, reader-root
merging, persisted launch account and capability string with one
ClaudeProfileRouter: the pointer is written first and setup runs best-effort
after it (superset's order); a missing pointer means System default.

The claude shell function re-reads the pointer on every launch, refuses only
a selected account whose folder is missing, and prints a note when the user's
own CLAUDE_CONFIG_DIR overrides the selected account in that terminal.

Still dormant: claudeProfileRoutingEnabled() is false.

* test(claude): type router test settings instead of casting

* fix(claude): run account setup on a worker thread, never Electron main

publish() writes the pointer and starts setup in the background, so neither
startup nor an account switch blocks on a history merge. Each setup runs in a
one-shot worker (the profile-state backup worker's pattern); one setup per
account at a time, reused by later requests. A launch waits only for a folder
that was never set up, and refuses with a clear message if that setup fails.

* fix(claude): do not await the synchronous pointer publish

* test(claude): give the routing launch test the merged resolver deps and handle shape

* fix(claude-accounts): dedupe merged prompt history, drop drained copies, link setup folders by path

- Prompt-history drain appends only lines the shared file lacks, so a purge never re-adds lines.
- A set-aside history copy whose saved offset reaches its end is deleted on the next run.
- Setup folders link to the default home's own entry, not its resolved target.
- The profile gate and folder creation run once, in provisionClaudeAccountProfile.
- installHooks receives only configDir; drop a duplicate test key that fails CI.

* fix(claude-accounts): refuse a routed resume whose transcript is in another account; zsh claude function; setup timeout

- With account routing, a chat resume checks its transcript is in the launch folder; a missing one
  with a stored leaf refuses with historyInOtherAccount instead of starting fresh.
- The launch folder of a selected account comes from prepareLaunch(); the resolver stays for System default.
- zsh panes get the claude function like bash, fish and PowerShell (empty while routing is off).
- The setup worker is terminated after 60 s so a later launch can retry.
- Document that the setup marker means setup started, not finished.

* fix(claude-accounts): refuse a routed resume only when the transcript is found in another folder

A transcript found in no known folder keeps the old stored-leaf resume.

* fix(claude-accounts): record installed hooks as Orca-shared; skip symlink tests on Windows

After Orca installs its hooks into an account, record the account's hooks in
the settings ledger so a later run can still bring the user's own hooks in.
Tests that create real symlinks now skip on Windows.

* fix(claude-accounts): trim the which-account file in the PowerShell claude function

Co-Authored-By: Claude <noreply@anthropic.com>

* test(claude-accounts): spell the user's own config folder as an absolute path on every platform

Co-Authored-By: Claude <noreply@anthropic.com>

* test(claude): skip the POSIX-only WSL profile test on Windows

A WSL profile's data root is a POSIX path, so building one from a Windows
temp dir fails the absolute-path check there.

* fix(claude): check the transcript before the account-switch recheck when no router is installed

Keeps the switched-off path identical to main: the recheck stays the last await.

---------

Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
2026-10-07 01:00:12 -04:00
Jinwoo Hong 7f83e0a53f fix(serve): keep every persisted terminal on the phone after serve restarts (#26022)
On a windowless host the first hydrate pass after a cold start only keeps
serve-/SSH-owned terminals, stores that partial list, and the full pass then
skips the worktree because it already has tabs. Terminals made by
terminal.create, agent launches, or splits never reappear on the phone.

With no renderer window and no snapshot yet, the runtime-owned pass now builds
the full snapshot (including a persisted split layout), since nothing else
publishes the remaining tabs.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-10-07 00:49:35 -04:00
Neil 76d480f808 Fix unsafe test fixtures and the Bun version pin (#26051)
* Keep test interruption signals within owned processes

* Pin Bun and add optional unit runner shutdown diagnostics

* Unblock CI lint without changing session host runtime

* Avoid duplicating runtime import-check dependency bundles

* Leave unit runner diagnostics disabled by default

* test: keep runner incident follow-up focused on durable guards

* test: apply transcript replacements as authoritative snapshots
2026-10-06 21:42:02 -07:00
Jinwoo Hong d1fe11c6ba refactor(terminal): carry pane placement on terminal start, report-only (#25677)
* refactor(terminal): carry pane placement on pty spawn, report-only

Adds a placement type (new tab, split, or root) that says which tab and leaf a
new PTY joins. pty:spawn accepts it as an optional field, main's runtime creates
and PTY-backed splits supply it, and the binding write records on its trace span
whether placement names the tab today's write picks. The write never branches on
it, so bindings and saved state are unchanged. applyPtyBinding becomes
copy-on-write and returns the new session.

* test(terminal): typed handler registration in the placement threading test

* refactor(terminal): keep the binding write in place; placement stays report-only

Reverts applyPtyBinding to main's in-place write: placement is read before the
write and does not need a new session, so this PR keeps identical behavior.

Follow-up: build rows and splits from placement, delete persistHeadlessTerminalSplit,
the provisional-graft strip, and mobile's follow-up props write.

* fix(terminal): a throwing placement check never fails the binding write

The agreement check is report-only, so a throw records check_threw on the span
and the binding proceeds exactly as without placement.

* refactor(terminal): carry only the placement fields this change reads

New-tab rows, sizes, split ratios and proposed trees come back with the change that builds tabs and splits from placement. Leaf ids are capped at 128 like the other session schemas.

* fix(lint): bring structured-agent-session-host back under max-lines

(cherry picked from commit 1a5f687244)
2026-10-07 00:18:34 -04:00
Jinwoo Hong 9d8e3638de chore(relay): move Asia cell c34 to the general same-cap lists after promotion (#25758)
* feat(relay): give Asia cell c34 a promotion wave so it can become a general cell

c34 launched on 2026-10-05 as a migration-only spare with no promotion path. This adds
it to the Asia admission promotion waves, the workflow's promote and canary cases, and
the canary evidence map, so the reviewed Asia workflow can promote it with the same
five-minute canary c30 and c31 ran.

The same-cap migration-only list is deliberately unchanged: a same-cap job reads a
cell's class from that list, and c34 must be rolled to the director's image as a
migration-only cell before promotion can run. The list moves after promotion, in its
own change.

Claude-Session: 1145a80d-dec4-4a9b-9373-bbbb876b9041

* chore(relay): move Asia cell c34 to the general same-cap lists after promotion

A same-cap job reads a cell's class from SAME_CAP_MIGRATION_ONLY_CELLS, not from the
selector. Once c34's Asia canary promotes it, leaving it there would make a same-cap
rollback demote it to migration-only. This moves c34 to the general same-cap list and
the shadow gate's fleet pool list (its pool is the Asia 16).

Merge only after c34's promotion succeeds. This file set is trusted evidence code, so
merging invalidates any sealed monitor or same-cap canary: merge outside a
monitor-to-enable window and before the next same-cap canary seals.

Claude-Session: 1145a80d-dec4-4a9b-9373-bbbb876b9041

* docs(relay): scope the c34 same-cap pause to the window after promotion

Claude-Session: 1145a80d-dec4-4a9b-9373-bbbb876b9041

* docs(relay): rewrap the c34 paragraph

Claude-Session: 1145a80d-dec4-4a9b-9373-bbbb876b9041
2026-10-07 00:10:52 -04:00
Jinwoo Hong bfcbcfd59c refactor(terminal): move a dragged-out pane in main before the window opens its tab (#25380)
* refactor(terminal): move the topology revision and leaf-lookup helpers into terminal-topology

advanceTerminalTopologyRevision, findTerminalTabIdForLeaf and
hasHostAuthoritativeTerminalMembership move verbatim from the renderer-save
membership rebase into persistence/terminal-topology, the home of the commit
boundary. Importers are repointed; no behavior change.

* test(terminal): test-only guard for topology writes outside the commit boundary

The three session sinks now publish through one commitWorkspaceSessionPartition
helper, which hands the prior and published partition to the topology write
guard. Production never arms the guard, so each sink pays one global lookup.

The unit suite arms it in report mode: a sink write that changes class-(a)
topology (membership, root, bindings, titles, incarnations, sleeping records,
remote session ids, tombstones, default-applied, revision) outside a commit
scope is attributed to its writer's file by stack, and fails the test unless
the writer is on the unrouted-writers allowlist that later routing PRs shrink.
Renderer saves and test seeding are exempt. Deep-freeze is available but stays
off suite-wide until the in-place writers return new sessions.

* refactor(terminal): add the topology commit module with bindLeaf, closeLeaf and closeTab

terminal-topology-commit.ts is the boundary for class-(a) terminal topology.
bindLeaf forwards to persistPtyBinding, whose write now runs in a commit
scope; the spawn commits and the relay reattach bind through it. closeLeaf and
closeTab wrap the existing close mutation in a commit scope and one
persistence.terminal-topology span (kind, outcome, refusal reason; no ids),
and the runtime close goes through them. Session output is unchanged: tests
compare it byte for byte with the old writers for local, ssh: and folder
workspaces.

* chore(terminal): drop an unused lint suppression from the topology write guard

* fix(terminal): attribute topology guard writers relative to the repo root

The guard read a frame's file through its last /src/ segment, so a test under
tests/ (folder-upgrade-identity-persistence.unit.test.ts) had no source frame
and failed as an unknown writer. Frames are now taken relative to the repo
root the setup passes in, tests/ counts as test seeding, Windows backslashes
are normalized before the node_modules skip, and nested src/ paths keep their
full path. The two runtime funnel files skip only their funnel function, so
another writer in them still shows. The R9 allowlist key names the file whose
frame actually writes. The class-(a) diff and the attribution move into their
own files; the stack limit is restored in a finally.

* refactor(terminal): drop bindLeaf until binding reaches a sink; add the boundary ratchet

bindLeaf and the commit scope inside persistPtyBinding changed nothing: the
binding write never reaches a session sink, and a scope inside the Store method
would have admitted every direct caller once it did. Both return in B1-4; the
spawn commits and the relay reattach call persistPtyBinding directly again.

The runtime close now calls one closeLeafOrTab entry, so its callbacks keep
their contextual types. A census test is the primary enforcement: only
persistence/terminal-topology and the callers it lists may call
persistPtyBinding, the session setters or the three sinks, and every
unrouted-writer allowlist entry must name an existing file.

* fix(test): resolve the topology guard's repo root without the global URL

Under happy-dom the global URL is not Node's, so fileURLToPath(new URL(...))
threw in the setup file and failed every happy-dom test file.

* feat(terminal): moveLeaf commits a pane detach in main before its new tab mounts

Ported from fix/terminal-topology-stage-a (1bafc87, 3941bd4, 2ea2070, 3e79f31,
a9d428e minus the bridge comment) and exposed as the commit module's moveLeaf,
which runs inside the topology commit scope and span and publishes every
rewritten partition through the session sink, so the test-only write guard sees
it. Store.moveTerminalLeafToNewTab only hands it the store's state.

Detach-to-new-tab now asks main first: the leaf and its binding move into the
new tab in every owner partition (local and ssh:), with the incarnation,
sleeping session, remote session id, UI marks and SSH lease, and the topology
revision bumps. pty:moveLeafToNewTab then aliases the agent-status pane key and
re-keys orchestration worker resources; the renderer's repeat transfer is
skipped. A refused or failed move toasts and the pane stays; a move the
renderer cannot apply is undone (STA-9259).

Stage A review-2 fixes:
- SF1: an undo whose source tab has closed retires the moved tab instead of
  refusing, so no ghost tab comes back on restart.
- SF2: an undo restores the split direction, ratio and position, the source
  tab's PTY id and its SSH remote session id, from origins main kept for the
  move; it never adds a second copy of a leaf the source holds again.
- N1: layouts of tabs that no longer exist do not refuse a move.
- N2: an undo into a tab that gained a pane is refused as target_changed,
  not a silent not_held.
- N3: a repeat drop of a pane whose move is still committing is a no-op.

Note: the move bumps the topology revision, so a revision-0 repo enters
rebased membership on its first detach.

* fix(terminal): tighten moveLeaf undo and move across disagreeing SSH copies

- An undo ignores a target layout whose tab row is gone, as the move does.
- A retired undo leaves UI marks and SSH leases on the moved pane key instead
  of re-keying them to the closed source tab.
- When an SSH respawn bound one partition before the other caught up, the move
  follows the renderer's live PTY id and drops the stale copy's incarnation
  instead of refusing with a toast.
- A test pins the late-spawn graft the renderer's in-flight guard exists for.

* fix(terminal): a pane move the renderer cannot finish never leaves main holding it

- A pane whose PTY spawn is still in flight is not dragged out (no toast):
  the late spawn result would bind it to the source tab and graft the leaf
  back beside its moved copy. The IPC transport reports a pending connect.
- A thrown or timed-out commit (10 s bound) sends an undo, since the write may
  have landed; undo answers not_held when nothing moved. The in-flight entry
  clears with it, so a stalled main cannot wedge the pane's drag.
- A throw while opening the new tab after the pane left its source undoes the
  move too.
- No "stays where it was" toast when the undo found the source tab closed or
  the source pane is gone; a failed undo says the pane may open in a new tab
  after a restart instead.

* fix(terminal): drop a half-opened move tab so the window agrees with main's undo

If opening the moved pane's tab throws after createTab ran, the renderer now
closes that tab without killing its PTY or telling main (main never had it),
and restores the source layout, while main undoes the move.

* refactor(terminal): drop the runtime topology write guard; the boundary ratchet enforces

The AST boundary ratchet is the enforcement for B1. The stack-attributed
runtime guard, its class-(a) diff, the unrouted-writer allowlist, the vitest
setup and the freeze option are removed; it saw two writers in the whole unit
suite and its real value starts only once binding reaches a sink (B1-4).

Also: one closeLeafOrTab wraps the close in the span (no per-kind copies or
narrowed types), the span has one finish like persistence.pty-binding, the
sink helper is publishWorkspaceSessionPartition (it publishes; the commit
boundary is the module), the ratchet drops the private publishSession row and
checks that every listed caller file exists, and the close comparison keeps
its two meaningful cases with span cases chosen by name.

* refactor(terminal): roll a pane move forward instead of undoing it

Once main commits a pane move, the renderer only rolls forward:
- A repeat of a committed move answers moved and writes nothing, so a thrown
  commit (outcome unknown) is retried, up to three attempts; the IPC skips
  the agent-status transfer it already made.
- The renderer applies the move by leaf id. If the user closed the pane or
  its tab meanwhile, it closes main's new tab through the ordinary
  session:close-terminal-surface path; if a sibling closed and the pane is
  the source's last, the source tab closes without killing the PTY.
- One toast, for a move main refused or never answered.

Deleted: the undo planner, origin ledger, undo flag, retired and
target_changed results, the commit timeout, the half-opened-tab cleanup and
the second toast.

Also: one request validation (distinct tabs, stable leaf id, tab ids a pane
key can carry, built through makePaneKey); a two-line PTY check where the
live PTY id wins; the move and close share one span wrapper; the moved tab's
layout always takes the live PTY id (a stale saved id was kept before);
fallbackPtyId is now livePtyId; tests are organized by behavior.

* refactor(terminal): trim the B1-1 commit module and ratchet to what they enforce

- Point the acknowledged-tab-retirement audit fixture at the moved
  advanceTerminalTopologyRevision; its old import no longer resolved.
- Drop publishWorkspaceSessionPartition: it was the removed guard's
  interception point, so the three session sinks return to origin/main.
- One traced(kind, mutate) wrapper in the commit file replaces the span
  factory; closeLeafOrTab is one call.
- The ratchet walks src/main with the shared scanSourceTree, drops the
  loading-store-internal rows and the redundant file-exists test; exact-set
  equality already fails on a missing file.
- The commit test is a pure unit test of the span outcomes: no Store
  harness, electron mock or self-comparing close.

* fix(terminal): keep a moved last pane alive and trim moveLeaf to one attempt

A sibling closed while main committed the move persists the source layout
down to the moved leaf, so the renderer took that for "pane gone" and
closed main's new tab while the pane was live: after a restart it was held
nowhere. The renderer now reads the store layout: when the moved leaf is the
only one left it opens the new tab with that layout, syncs PTY ownership,
and closes the source tab without killing the PTY. Main's new tab is closed
only when main moved the pane and it is gone here.

Also:
- One commit attempt; the repeat answer, the retry loop and the IPC alias
  guard never ran in production (a write with an unknown outcome faults the
  writer, so every retry rethrows before it reaches the move).
- traced takes a typed refusalOf; one write-and-restore helper and a shared
  partition assign replace the duplicated local/remote and rollback checks.
- The planner's holders carry their source tab and layout, so the partition
  move cannot fail; pane-keyed UI marks are re-keyed in one loop; the request
  validator calls makePaneKey directly.
- applyMove looks the pane up once and takes one PTY id (main's when it moved
  the pane); the move-commit file is inlined; duplicated renderer tests are
  removed and the commit-path tests live together.

* refactor(terminal): census the runtime session controller's write and drop stage ids from comments

The controller's setter was named set, which the boundary census could not
list without matching every Map.set, so a new OrcaRuntime mixin could write
sessions through it unseen. Rename it setForWorktree and census it.
Comments now describe state instead of citing plan stage ids.

* refactor(terminal): census writer references and trace refusals by callback

- traced() takes refusalOf instead of assuming an Error refusal, and only
  mutate() sits in the try, so a span outcome of threw means the write threw.
- The boundary ratchet counts references, not just direct calls: non-null
  calls, bracket keys, aliases, destructures, .call/.bind and parenthesized
  callees all count; declared names and type positions do not.
- Census terminalSurfaceCloseMutation (boundary-only) and the partition sinks
  setLocalWorkspaceSession / setHostWorkspaceSession.

* test(terminal): count writer uses in extends clauses and instantiations, skip type-only imports and local declarations

The census skipped ExpressionWithTypeArguments as a type, which also holds
`extends f(x)` and `x<T>` value expressions. Type-only import/export
specifiers and declared names (variables, parameters, accessors, enum
members) no longer count as uses. The audit fixture is listed in the table
instead of a separate exemption.

* test(terminal): count quoted and assignment-pattern destructures of layout writers

* test(terminal): count every mention of a layout writer except its definition

Telling definitions from uses per syntax kind kept missing nested and
for-of destructures. Exempt only the writer's own function or class-member
definition; any other mention (including object-literal keys) counts, so the
census errs toward a loud false alarm rather than a silent miss. Quoted names
count only in member-name position.

* test(terminal): count every string literal naming a layout writer

Member-name positions missed wrapped keys like store[('name')] and
store['name' as const]. Counting every string literal outside types is
shorter and errs toward a loud false alarm.

* test(terminal): exempt only class members and functions as writer definitions

Object-literal methods and accessors were exempt while equivalent arrow
properties counted; all object-literal keys now count alike.

* test(terminal): parse files with unicode escapes in the writer census

A name spelled with a \u escape never appears verbatim, so the text
prefilter skipped it.

* test(terminal): parse any file with an escape in the writer census

\x, identity and line-continuation escapes also decode to a writer name
without it appearing verbatim.

* test(agent-hooks): stub isPaneAuthorityTransferredTo on the hook server fake

* refactor(terminal): one move path for every client and an idempotent status transfer

The pane manager decides whether the dragged pane is the last one, and a
refused detach closes main's new tab like a vanished pane. Paired web
clients answer the move as not held, so every client takes the same
moved / not_held / refused branch. Repeating an agent-status transfer that
is already in place is now a no-op where it happens, replacing the IPC
special case.

* fix(lint): bring structured-agent-session-host back under max-lines

(cherry picked from commit 1a5f687244)
2026-10-06 23:51:42 -04:00
Brennan BensonandJinwoo-H dff65d55a3 feat(claude): prepare account profiles and shared history (Step 1 of 4) (#24300)
* feat(claude): add dormant profile setup and history sharing

* fix(claude): make profile setup one gated, typed, fail-safe entry

Review round 1 of the dormant profile setup found that the pieces could
be called without their safety checks, that one failed write or an
unreadable bookkeeping file could silently stop sharing for good, and
that Windows prompt history could bring back history the user cleared.

- One entry, provisionClaudeAccountProfile: the profile gate (namespace,
  no linked components, outside ~/.claude and ~/.config/claude, and an
  ownership marker beside the home naming the account and target) runs
  first and refuses before creating anything; then history sharing,
  config provisioning, and the hook install after the settings merge.
  Results come back per surface with closed warning codes instead of
  message text.
- The sharing ledger is keyed by surface name, records a value only
  after its write succeeded, and an unreadable ledger starts empty and
  is rewritten instead of blocking every surface.
- The profile state file goes through the same locked writer as folder
  trust (Claude's <file>.lock plus the in-process queue), generalized as
  updateClaudeGlobalConfig. Onboarding and trust are still applied when
  the personal state file is unreadable.
- WSL descriptors build guest POSIX paths; the state-file path style
  follows the injected platform.
- Orca's managed statusLine has one owner in a profile: the settings
  merge never shares it, a user's own statusLine is shared over it, and
  the profile installer follows the default home's slot so a default
  opt-out reaches every profile. remove() takes the same destination;
  the remote installer cannot accept one.
- Prompt history compares file identity (bigint dev+ino) on every
  platform, never drains the shared file into itself, drains retained
  copies in generation order, never reuses a stale cursor, and on Windows
  keeps a replaced default's old copy aside instead of replaying it.
  Directory merges keep going past a failed entry.

* fix(claude): share the user's own hooks and keep merged history whole

A user's own Claude hooks in ~/.claude (notifications, formatters) did
not run under a managed account, because the whole hooks key stayed
private. They are now shared like any other settings key: Orca's own
hook entries and its managed statusLine are stripped from both the
personal value and the profile's current value before the per-key
ledger comparison, so they never travel through the merge and never make
the key look user-owned. Orca entries already in the profile are kept on
write, and the profile hook installer adds them on top as before.

Prompt history: merged bytes that lack a final newline are terminated,
so Claude's next record no longer fuses onto the last merged line. When
a CLI rewrote the profile's history file (old records plus new), only
the lines past the part it shares with the default history are added,
instead of the whole file again.

* fix(claude): close review round 2 gaps in profile setup

Hooks and statusLine sharing:
- When ~/.claude holds only Orca's hook entries, the user's shared hooks
  now read as an empty value instead of a missing key. Removing the
  user's last own hook in ~/.claude therefore reaches profiles that
  never edited it, and deleting the only shared hook inside a profile
  stays deleted.
- A custom statusLine Orca shared, and the profile never edited, goes
  away when the default home drops it. When a shared custom line
  replaced Orca's line in a profile, the profile's statusline marker is
  dropped so Orca's line comes back once the default returns to it; a
  profile that opted out stays opted out. No other key gains deletion.
- install/remove/getStatus with a profile directory refuse when it is
  the default home, or its settings.json resolves to the default one,
  instead of editing System Default's hooks and opt-out state.
- The profile statusline rule reads the default settings under the
  userHome passed to the setup entry, not os.homedir().

Profile state and ownership:
- A malformed `projects` value skips only folder trust (new warning
  code trust-refused); onboarding and shared keys still apply.
- The ownership marker stores only host-local facts (account, runtime,
  distro). The execution host id is the caller's view of the host, so
  it stays in the in-memory descriptor and is not compared.

Prompt history interruption paths:
- With no cursor yet, a retained copy starts past the bytes it shares
  with the default history, so an interrupted share no longer replays
  the whole history.
- A retained name for the shared file itself is removed with its cursor
  instead of lingering until a later scrub makes it look new.
- The Windows link record is read three-state: unreadable stops the
  share instead of reading as "no link". If the record cannot be
  written after linking, the fresh link is undone.
- An unreadable retained copy is reported and no longer blocks linking.

* build(cli): list the new Claude hook modules in the CLI project

hook-service.ts and hook-settings.ts are compiled into the packaged CLI
project, which lists every file explicitly. The statusline policy and
profile destination modules they now import were missing, so the CLI
typecheck failed with TS6307. The CLI still loads hook-service through
the existing managed-agent-hook-controls build entry, which bundles
both modules; neither imports electron.

* fix(claude): close review round 3 regressions in profile setup

- A profile whose hooks hold only Orca's entries and that sharing never
  recorded is no longer treated as a user edit, so the user's first own
  hook in ~/.claude reaches it (for example when the profile was set up
  before ~/.claude had any hooks).
- A retained prompt-history file is removed as a second name for the
  shared file only when the default history does not itself link to it;
  otherwise it holds the only copy and is kept.
- Default-home checks compare file identity: the profile hook
  destination check uses device and inode, and the profile/default
  separation check resolves on-disk case, so a case-only alias of
  ~/.claude is refused on case-insensitive filesystems.
- A test pins that an unreadable leftover session tree no longer blocks
  linking.

* fix(claude): let shared keys leave a profile when ~/.claude drops them

QA found that removing a setting from ~/.claude never reached a managed
account: deleting the whole `hooks` block left the user's hook running
there. Only statusLine followed the default away.

Every shared key now follows the same rule through the existing per-key
ledger: when a key disappears from ~/.claude/settings.json (or
mcpServers/theme from the personal state file), it is removed from the
profile if the profile still holds exactly what Orca last shared. A
value changed inside the account is kept. Keys Orca never shared,
including denylisted ones, are never touched. Deleting the whole hooks
block removes the user's shared hooks and keeps Orca's own entries. A
missing source counts as empty; an unreadable source removes nothing.

* fix(claude): share personal rules, themes, workflows and keybindings into account profiles

A managed account launches Claude with its own config folder, so user-level
rules/, custom themes/ (which a shared `custom:<slug>` theme points at),
personal workflows/ and keybindings.json silently stopped applying. Link the
three directories like skills and commands, and copy keybindings.json with the
same edit-preserving ledger as CLAUDE.md. routines/ stays unshared: routines
belong to the claude.ai account and the folder holds per-run state.

* fix(claude): import the personal CLAUDE.md into account profiles instead of copying it

Claude also loads ~/.claude/CLAUDE.md as a parent folder's memory for any project under home,
so a copied account CLAUDE.md made every such session read the user's instructions twice
(checked live with Claude 2.1.288). An @~/.claude/CLAUDE.md import resolves to the same real
file, which Claude loads once from home, from projects under home and from folders outside it.

* refactor(claude): simplify account profile setup toward the prior art

- Windows keeps each account's history private; drop the hardlink, link
  record and conflict-copy machinery that only Windows reached.
- Share hooks and statusLine as ordinary settings keys: Orca writes the
  same entries into every folder, so the installer finds them present.
  Drops the Orca-entry carve-out, the per-profile statusline follow
  logic and its marker.
- Unreadable ledger is just an empty ledger.
- Share from the user's own CLAUDE_CONFIG_DIR when they set one (marked
  so Orca's injected value is never mistaken for it), and refuse a
  profile at or around it.
- Pin the one canonical profile path spelling in a test.

* fix(claude-accounts): dedupe merged prompt history, drop drained copies, link setup folders by path

- Prompt-history drain appends only lines the shared file lacks, so a purge never re-adds lines.
- A set-aside history copy whose saved offset reaches its end is deleted on the next run.
- Setup folders link to the default home's own entry, not its resolved target.
- The profile gate and folder creation run once, in provisionClaudeAccountProfile.
- installHooks receives only configDir; drop a duplicate test key that fails CI.

* fix(claude-accounts): record installed hooks as Orca-shared; skip symlink tests on Windows

After Orca installs its hooks into an account, record the account's hooks in
the settings ledger so a later run can still bring the user's own hooks in.
Tests that create real symlinks now skip on Windows.

---------

Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
2026-10-06 23:46:35 -04:00
Jinwoo Hong 55c86ceb91 refactor(terminal): report breaches of the two layout invariants; add an unrun load repair (#25673)
* refactor(terminal): report breaches of the two layout invariants; add an unrun load repair

The binding write now checks, after the fast lane, that a terminal is bound to at most one
leaf and a leaf id is in at most one tab, and records a breach on the persistence.pty-binding
span as binding.owner_conflict. It never refuses or changes the write (D16). The load repair
for saved duplicates is added and unit-tested but not called on load; the B2 core turns on
refusal and repair together.

* refactor(terminal): keep the binding write unchanged if the owner check throws; plain-language comments

* fix(lint): merge duplicate removal import in delete-worktree failure toast

Main (#25668) introduced a duplicate import that fails the focused code-quality lint.

* refactor(terminal): keep only the report-only owner check; one definition of same terminal

Defer the unrun load repair to the change that runs it. Repeating relay ids without both
incarnations are no longer one terminal, a leaf held on another host is reported apart, and the
refusal check takes host partitions directly.

* fix(lint): bring structured-agent-session-host back under max-lines

(cherry picked from commit 1a5f687244)
2026-10-06 23:31:02 -04:00
Neil d3e1494674 test: bound memory used by the runtime Electron audit (#26049) 2026-10-06 20:21:34 -07:00
Jinwoo Hong 7c9c3d431a fix(lint): bring structured-agent-session-host back under max-lines (#26050) 2026-10-06 23:17:02 -04:00
Jinwoo Hong 70c6876635 fix(rovo): detect and launch Rovo Dev through acli (#25981)
* fix(rovo): detect and launch Rovo Dev through acli

Rovo Dev ships as the `acli rovodev` subcommand, with no `rovo` binary on
PATH, so Orca never listed it as installed. Detect `acli`, launch
`acli rovodev run` (prompt still typed after start, since a positional
instruction is one-shot), and build resume as the launch command plus
`--restore <id>` so a full-command settings override composes too.

Fixes STA-9645

* test(skills): detect Rovo Dev through acli in the skills CLI fixture
2026-10-06 22:49:18 -04:00
Brennan Benson 923ef7cff1 fix(native-chat): refuse a Command that can't run instead of silently running the stock CLI (#25667)
* fix(native-chat): honor custom Claude and Codex launch commands

* fix(native-chat): sync launch failure localization catalogs

* feat(native-chat): run Settings → Agents → Command as the native chat program

Native chat ignored the Command setting, so a user who pointed it at a
custom Claude or Codex build still got the stock CLI. The host now reads
the Command per acquisition and per model-catalog probe, resolves it on the
execution host as a program (an absolute path, ~ path, or a name on PATH),
and spawns it with the normal structured arguments. A set Command that is
not a runnable program refuses the start with a new agentCommandNotRunnable
failure that names the setting; it never falls back to the stock CLI.

* fix(native-chat): find a configured Claude Command on the launch PATH; clearer copy

A Claude Command given as a bare name was looked up on Orca's own PATH,
not the PATH the Claude child launches with (login shell plus Settings →
Agents environment), so a name only that environment provides was refused.
Codex already resolved against its launch environment. Claude now reads
its launch env first and resolves a configured name against its PATH and
home; the stock lookup with no Command is unchanged.

The failure copy now states the rule (a program path or name, no
arguments or variables) and names the Settings control (Reset). The
catalog fingerprint doc notes that the program is not keyed, so a chat on
an older program keeps refreshing the shared entry until it ends.

* fix(native-chat): one PATH key for the Claude lookup and the child on Windows

On Windows, an inherited `Path` and a Settings → Agents env `PATH` both
reached the Claude child. The configured-program lookup read the first
inserted twin (inherited), while Node's child_process keeps the
lexicographically first (`PATH`, the user's), so a bare-name Command found
only on the Settings PATH was refused.

The Claude launch env now drops inherited case-twins of the overlay's
variables (the rule the structured shell env already applies), so the
lookup and the child read one PATH. pathEnvOf follows Node's win32 key
rule, and Codex's lookup uses it too, so both agents read PATH one way.
es/fr/ja copy now says the Command must be, as the English does.

* fix(native-chat): refuse a Command that can't run instead of running the stock CLI

A saved Settings → Agents → Command that did not resolve to a program (a
command line such as `codex --profile work`, `FOO=1 claude` or `npx …`, a
missing or relative path, a non-executable file) silently ran the stock
Claude/Codex in native chat and its model-catalog probe. The user's choice
was ignored with no sign of it.

A set Command that does not resolve now refuses the start with a new
agentCommandNotRunnable failure naming the setting, beside the existing
Retry; the catalog probe refuses the same way and never lists the stock
CLI. A blank Command keeps the stock lookup. On Windows a resolved file
without a spawnable extension (.exe/.com/.cmd/.bat) refuses too, instead
of failing to spawn with no word about the setting.

* chore(native-chat): correct comments that say native chat ignores the Command

After #25721 native chat runs a runnable Command, and this PR refuses one it can't run; four comments still said the Command applies to terminal launches only. The pre-spawn test fixture also drops the saved value from its error message, matching the resolver.
2026-10-06 19:47:31 -07:00
Brennan Benson 5b1bb78d14 [STA-6940] Prepare file drops for element owners (PR 2 of 6) (#25748)
* feat(file-drop): add element owner preparation plumbing

* fix(file-drop): preserve feedback acceptance and destination ordering

* fix(file-drop): keep path resolution in filesystem namespace
2026-10-06 19:32:15 -07:00
Brennan Benson 0bdcaf36ed fix(claude): start a Claude chat with its saved options and send the first message at once (#25152)
* fix(claude): end a Claude start that never answers initialize after 120 s

* Read the Claude startup deadline inside startup; fix a stale test comment

* fix(claude): start a Claude chat on its initialize answer, not on a frame only a SessionStart hook sends

Startup waited for system/init or a SessionStart hook frame as well as the initialize answer.
Before the first turn only a SessionStart hook sends one, and Orca adds that hook only through
its optional status hooks, so with them off the first message was held forever. Startup now
lands on the initialize answer; a start frame already seen is still checked, and one naming
another session ends a started session. The deadline drops to 90 s so it fires inside the
host's 120 s start wait.

* fix(claude): time the Claude start by silence, and fail it at once on another session's frame

Claude answers initialize only after its SessionStart hooks finish, so a total-time deadline
would fail every start behind a slow hook. Each start frame now restarts the clock. A frame
naming another session fails a start still waiting on initialize at once, as before.

* test(claude): a real Claude chat starts and answers with every hook disabled

* Say what the start-frame re-arm covers, and check only start frames in the hook-less real test

* fix(native-chat): a Claude chat starts with its saved options and takes its first message at once

Saved model, effort, Fast and permission mode are passed as launch options, checked against the
account's cached model catalog, instead of restored by control requests after initialize. With
nothing left to restore, the host no longer holds a message until the CLI answers initialize, and
the 90 s startup deadline is gone. A Stop on a start that never answers ends the child and settles
what it was handed as stopped. A failed result that repeats the turn's own API error reply writes
no second row.

* fix(native-chat): a host stop of a Claude start fails the message it was handed, with one row

With no start-hold the delivery loop no longer sees a host stop of a start it waited on. The
child's end now rejects what it handed over with the host-stopped words and writes the one row,
as an exit of its own would; an idle start the host stops still goes quietly.

* test(native-chat): a Claude chat's first message is written before initialize answers

Rewrites the tests that encoded the start-hold, the startup deadline and the option restore to
the new contract, and adds: saved options at launch (catalog checks, bypass, fresh-session Fast),
a message written before initialize answers (adapter and runtime), Stop on a start that never
answers (stopped, child closed, nothing working), and an API error said once.

* revert(native-chat): keep a failed Claude turn's error row

The shared turn fold already shows a failed turn's error once after it settles, as an error;
dropping the row left the CLI's synthetic reply looking like something Claude said.

* fix(native-chat): pass saved Claude options unchecked, heal a retired model on the CLI's word, and never leave an unrun message in doubt

- Saved model, effort and Fast are launched as picked; only values no Claude can parse are left
  out. The pre-spawn cache check is gone.
- A saved Fast on for a new conversation is applied once the settings readback shows no
  per-session opt-in (dropped when there is one, or when the model is listed without Fast), with
  nothing waiting on it; the record keeps the pick.
- Under an Agent Permissions bypass, a saved narrower mode launches with the allow flag so bypass
  stays reachable.
- A turn whose reply is the CLI's model_not_found for the launched model drops that model from
  the record; the launch's own row for it is kept out of the account model cache.
- A child that ends before it answered initialize, for any reason, settles every message it was
  handed as not sent (cancelled for a Stop).
- A launched effort the CLI reports only as `applied.effort` is confirmed from there.
- The untimed-initialize comment is back to main's text.
- A real-CLI test for a message written before initialize answers, under saved options.

* fix(native-chat): type the close's ended event and the start-exit test fixtures

The close's ended event is typed as the adapter event so its optional startupUnanswered spread
fits exactOptionalPropertyTypes; two tests guard the fixture's optional generation, and the
hung-start fixture records initialize on the fake connection it holds.

* fix(native-chat): a Claude model heal keeps a later pick, a refused Fast is dropped, no allow flag

- `options-skipped` carries the retired value; the record drops it only while it still holds it.
- `started` carries the values a heal retired, and the record does not take them back from the
  CLI's report of the same value.
- A saved Fast on a new conversation is applied before `started`: a refusal drops it and records
  it skipped, as main's refused restore did; silence keeps it wanted and unconfirmed.
- A saved narrower mode under an Agent Permissions bypass launches without any bypass flag again:
  the allow flag is one older CLIs reject at start. Kept as a known limit.
- The real-CLI test asserts the message was written before initialize answered.
- The fake reports a launch effort only under `applied`, and a misplaced doc comment moves back.

* fix(native-chat): a new Claude chat reports started before its saved Fast is applied

The Fast apply on a new conversation now runs after `started`, so a Stop interrupts a running
first turn and an option write is not refused while the round trip is out. A refusal drops the
pick through `options-skipped`, in order after `started`; silence keeps it unconfirmed. The
launch's unreachable skipped-model branch is gone.

* fix(native-chat): a healed Claude chat goes back to the default model live; comments match the no-hold design

When the CLI says the launched model does not exist, the live child is also put back on the CLI's
own default (set_model with no model, fire-and-forget), so later messages in the same chat run; a
user's pick sent after it wins, and a refused or unanswered reset only logs. Comments that still
described the start-hold or the option restore now describe the launch options and the
handed-over, never-echoed rule.

* fix(native-chat): a message handed to a Claude start that never answered is kept as main keeps an unsent one

A child that ended before it answered initialize ran nothing it was handed, the same fact as a
send accepted and never handed over. Its end now settles those sends exactly as the chat settles
a queued send for that end: a quit keeps a person's message as a held card (restart words), a
close keeps it as a held card (closed words), a person's Stop withdraws it as cancelled, and a host
stop fails the start with one row.

* fix(native-chat): a quit during a Claude start that never answered offers no resume for the message it keeps as a card

The restart snapshot now reads the same never-answered fact the exit does, so a message handed
to such a start counts as queued work, not as work to resume. The retired-model reset comment
names the default it really applies.

* refactor(native-chat): the saved permission-mode launch helpers live with the spawn options that use them

Keeps claude-structured-launch-resolution.ts within max-lines once merged with main, and names
the hung-start test envelope's field type.

* fix(native-chat): a Claude chat's saved options take precedence over the agent Arguments' own flags

Main now passes the saved agent Arguments to the Claude child, and the SDK writes them after its
own options. An Arguments --model or --effort therefore reached the CLI as a second flag after the
chat's saved pick (a commander CLI keeps the last), and a saved Fast's launch settings replaced an
Arguments --settings file outright. The saved model and effort now stand in for the Arguments'
flags, and a saved Fast beside an Arguments --settings is applied by the start instead of at
launch.

* test(native-chat): the hand-built Claude session in the options test carries fastModeAtStart

* test(native-chat): the queued rig's start spy carries a named Mock type

An unannotated vi.fn() inferred @vitest/spy's internal Procedure, which CI's typecheck cannot
name in the factories' inferred return types (TS2883).

* fix(ci): run the PR's SQLite-backed tests in the Node runtime project

Lists the hung-start Stop test, renames the send-during-startup entry from its old name, and
carries main's own two entries from #26010 so the boundary test passes before the next merge.
2026-10-06 19:19:47 -07:00
Brennan Benson 7a59aa39d0 fix(worktrees): decide setup before git worktree add on the host's create (#26006)
* fix(worktrees): decide setup before git worktree add on the host's create, so an undecided ask repo leaves nothing behind

The host's create (agent.launch, worktree.create, the phone and the CLI) only checked the setup policy
after adding the worktree. An ask repo with no setup decision then threw with the worktree already on
disk and no worktrees-changed event. It now refuses before the add, as the desktop create does, and a
setup hook the new branch adds (one nobody decided on) is skipped with a warning instead of failing a
create that already exists.

* fix(worktrees): read the host create's setup decision in one place, after recording its host

The pre-add refusal now runs after the create records its execution host, so a refused create's
failure telemetry still says where it ran; it still runs before any git work. Both setup checks
read the decision from one helper so they cannot disagree, and the post-add skip logs once.
The refusal test now also proves no base resolution, fetch or branch naming ran and that the
hooks came from the main checkout.

* test(worktrees): cover a setup hook only the new branch adds through the managed create

An ask repo with no decision, whose setup hook exists only in the new worktree's orca.yaml, now
creates, refreshes the worktree list, reports setup as skipped and returns the warning.

* fix(cli): name the --setup flag when a worktree create needs a setup decision

The host refuses an undecided create in an ask repo with "Setup decision required for this
repository", which desktop and phone answer in their own UI. The CLI now adds the next step:
pass --setup run or --setup skip.
2026-10-06 19:11:48 -07:00
Brennan Benson b3b6c5dc13 Give native chat names one source for tabs, sidebar and AI Vault (list and search) (#25986)
* Give native chat names one renderer source and drop Vault's name repair copies

The host's saved conversation name now rides the structured session status
feed, which already exists per host, is keyed by the durable session id, and
keeps a closed chat's summary. Tab strip, sidebar rows and AI Vault (list and
search) read it through one hook and one display order (tab alias, saved name,
host label). Vault no longer copies names into its cached results, so the
projection, recovery and pending-title modules and their tab-snapshot lanes are
removed. Indexed search hits now carry the native owner and saved name from the
host that indexed them.

* Type the sidebar name test fixture without an assertion

* Keep Vault search working when the chat host will not install

Naming and owning search hits is bookkeeping: if the native chat host fails to
install, return the plain hits instead of failing the search. The runtime RPC
only installs the host for clients that will receive the owners.

* Publish chat names to the feed independently of the tab retitle

A failed feed publication no longer skips retitling the open tab. The publish
now lives in the naming deps, where a test covers it.

* Note why the status feed must keep closed chats' summaries

* Bound names and owner ids that come from a paired host

Drop a published chat name the record store would refuse, and cap a search
hit's owner workspace id at the same length the list row and record use.

* Let native chat search hits from a paired host open their chat

A paired host's search hits carry no resume command, so the row disabled
every open action even for a native chat it can open through its owner, as
its list row does.

* Ignore workspace ids that name object members in tab lookups

A paired host's row or search hit could carry a workspace id such as
"constructor", which read an Object.prototype member as a tab list and broke
the render. The shared tab index now reads only own workspace entries.
2026-10-06 19:04:04 -07:00
Brennan Benson 9fdd90ffa7 feat(orchestration): a native chat can be a dispatch worker, like a terminal agent (#22972)
* feat(native-chat): a queued card can record a Dispatch's task as its source

A chat worker's task is held in the chat's queue like any agent message, so
the card needs to say which Dispatch it is from: the sender, run, task and
Dispatch ids. An older build reads an unknown kind as the person's card and
sends it as written.

* feat(orchestration): a chat can be a dispatch worker, like a terminal agent

`dispatch --to orca_session_id:<id>` and `worker-start --terminal
orca_session_id:<id>` now accept a chat on this host instead of refusing it.

- The chat is refused only where mail to it would be: unknown, a provider id,
  another host, or closed. A chat can't be its own coordinator's worker.
- Its task goes through sendAgentTurn as a queued send, as mail notices do:
  an idle chat starts a turn, a busy one holds a card naming the Dispatch.
  The operation id is derived from the Dispatch, so a resend replays.
- worker-start reads that outcome as the terminal path reads its write; a
  card held behind a running turn is handed over with its start unobserved.
- The Dispatch names the chat by its /clear root, with no pane or process,
  and records its Orca session id, so its own sub-dispatches nest under it.
- Mail to dispatch:<id> reaches the chat; worker-show/list/read read the
  chat's session records; stop and abandon never close the chat; closing the
  chat fails its Dispatch as closing a terminal does.
- A dispatch preamble for a structured session names the CLI the structured
  mail lane names.

* test(orchestration): a chat worker's reach, and its Dispatch across a /clear

At rest is live, a closed chat has exited, and another host or nothing to
read is unverifiable. A /clear keeps the chat's Dispatch, and the session
that continues the chat reports as it.

* fix(orchestration): chat worker review round 1

- dispatch --inject is a keepalive-backed wait, so a chat that takes a while
  to accept its task no longer reads as a dead runtime at the 30 s idle cut.
- An injected task whose delivery is unknown keeps its Dispatch open, as an
  unknown worker-start does, instead of failing it while the chat may run it.
- A busy chat's worker-start receipt says the task waits as a card in the
  chat's queue, and that worker-abandon does not remove that card.
- A close that puts the chat's tab back no longer fails its Dispatch: the
  closed-chat settle runs after the close's outcome, not at its hide.
- worker-start adopting a chat says it gave the task to the chat, instead of
  claiming it started a terminal agent; the mode value is unchanged.
- Test: a chat whose agent has not taken its task reads outcome_unknown.

* test(native-chat): a close's hide defers its hidden notice until the close settles

The rollback suite pinned the hide's arguments; it now expects the deferred
notice, and that the notice is sent once, after a rolled-back tab is back.

* fix(orchestration): a chat worker whose successor is unknown is unverifiable, not exited

A /clear successor this host has no record of is missing evidence, not proof
the chat is gone. Only a closed chat reads exited, so only a closed chat
settles its Dispatch; a lost successor reads unverifiable with that reason.

* fix(orchestration): chat worker review round 2

- A chat's close notice settles only that chat's Dispatch. A scan of every
  chat Dispatch read another chat whose close was still in flight (its tab
  hidden, maybe to be put back) as closed and failed its Dispatch.
- Only a dispatch --inject into a chat is a long poll; a terminal inject
  writes and returns, and keeps its short-RPC slot.
- Receipt wording: worker-start says it gave the task to the chat, and a
  queued task is sent when the chat's queue reaches it.

* fix(orchestration): chat worker wording, review round 3

- worker-start's mode sentence for a chat states the placement, which is
  true whether the task is delivered, queued or refused.
- Comments and a test title no longer claim a close notice re-derives every
  chat worker, or that a queued task waits for the current turn.

* fix(orchestration): build a structured worker's task sender from the start's own ids

The minted worker's preamble names who its task is from with the run, task
and Dispatch ids worker-start already holds, so it reads no Dispatch row.

* fix(orchestration): point a chat worker at its held Dispatch mail again

A chat worker's coordinator mail lands in its Dispatch's mailbox, but a
chat's idle edge re-derived the Dispatch mailbox only for a party with a
terminal handle. A pointer lost with the provider then waited for new mail
after a restart or a /clear. The re-derivation now looks the Dispatch up by
the party's address, which names a worker by its handle and a chat by its
root.

* fix(native-chat): open a task card's sender by address; a task carries no mail

Opening an agent message's sender looked up the mail it carried, which only a
mail notice has; with the task kind in the union that read no longer typed.
A task's coordinator is found by its address alone.
2026-10-06 18:36:21 -07:00
Jinwoo Hong 825d7bd5a9 test(vitest): run agent-launch-instant-tab in the SQLite runtime project (#26028)
#25430 added a test that opens a real agent-session record store, but not to the
SQLite runtime list, so vitest-sqlite-runtime-boundary fails on main.
2026-10-06 21:33:40 -04:00
Brennan Benson cdd0b7a491 fix(native-chat): stop flashing a reconnecting line on a stream drop (#24898)
* fix(native-chat): drop the per-chat reconnecting line on a stream drop

A single transcript stream drop flashed 'Reconnecting to this chat…' above the
composer for about a second. An unnamed read failure now adds nothing: the
transcript and composer stay, and the host's own status shows reachability.
Named or final failures keep their line.

* fix(native-chat): keep a loaded chat as it is when its stream drops naming nothing

A loaded chat with no messages yet still flipped to the full-pane "Could not
load conversation" on every stream blip: the read owner stored any read
failure as status 'error', and the pane shows that error whenever there is no
transcript. Once the chat has loaded, a failure that names no reason is now the
transport's own retry and is not stored; a named or final failure, and any
failure before the first load, are reported as before.

The "names a reason" check moves next to the final-refusal check so the owner
and the failure notice share it. The older-page read moves to its own module to
keep the read owner within its line budget.

* fix(native-chat): word a read failure beside a chat that never loaded

A chat that never loaded but already shows the user's own message (a launch
prompt or a queued send) said nothing when its first read kept failing with a
failure that names no reason. The loaded case is handled at the read owner, so
any read failure the pane sees beside messages is named, final, or from before
the first load; the status area now says its words in every such case.

* fix(native-chat): keep a loaded chat as it is only when contact is lost

The read owner skipped any loaded-chat failure without a named reason, but an
older host sends every refusal without details, so a refusal it really sent
(such as a damaged history) went unshown while the read retried in silence.
Skip only failures the host sent no refusal for, which is lost contact; any
refusal is the host's answer and is shown. The shared 'named' check is no
longer needed and is removed.

* docs(native-chat): say what a read failure without a refusal is treated as

Comment-only: the read owner treats a failure with no host refusal as lost
contact, which covers hosts too old to attach refusals; the test's reasonless
case is a reason this build doesn't know.

* Say a chat's host outage once, above the composer

A loaded remote chat looked live through a long host outage while sends sat
in the outbox. Derive a host-scoped notice from the host's connection state:
'<host> is reconnecting…' after a 2 s grace, '<host> is offline' at once with
the status bar's Connect as Reconnect, and a composer placeholder saying
sends go out when the host reconnects. A lost read beside the notice adds no
line of its own, and older history waits for the host.

* Drop the outage placeholder; no Reconnect for a refused host

A send while the host's transport is down fails and waits for its own Retry,
so the composer must not promise it will go out on reconnect. A host that
refused us (auth, protocol) still reads offline but offers no Reconnect,
which would be turned away the same way. Name the host with the shared
display-label selector, and keep the notice's live region mounted so it is
announced.

* test(native-chat): mock the font-size hook main renamed in the host-outage test
2026-10-06 18:31:31 -07:00
Brennan Benson 0c96550ee9 fix(native-chat): queued messages carry on in order after any turn, and nothing sends by itself after a restart (#24586)
* fix(native-chat): drop the queue-paused header and Resume button

A Stop, a restart or /clear holds the queued cards. The hold stays; only the
header row naming why, and its Resume button, go. A held card shows no
caption, and its own Steer, or any new message, releases the queue.

* test(native-chat): type the unknown hold reason a newer host may publish

* fix(native-chat): a held card offers Send, not Steer, when no turn runs

Steer vs Send now follows whether a turn is running, not the card's hold,
so a card held after a Stop, a restart or /clear reads Send.

* fix(native-chat): the queue sends past held cards instead of stalling behind them

A card queued after a Stop (or written after a restart or /clear) sent only
once the cards held before it were released; with no header to explain or
release the hold, it sat silently. The next sendable card now skips held
cards; a returned card still blocks what is behind it.

* fix(native-chat): the queue's send of a card is the person's turn, so held cards follow it

After a Stop, a card queued later sent past the held cards, but the queue
recorded that send as Orca's own turn. It never ended the Stop's pause, so
the held cards then waited forever with nothing on the card saying why.

A queued card is always something the person wrote: only the client send
RPC may now create one. The queue's send of it is therefore recorded as the
person's turn, which ends the Stop's pause once the agent takes it, and the
held cards then drain in order.

* fix(native-chat): a queued card carries its author, so the queue's send of it is that author's turn

Main now lets Orca's own sends ask to queue (sendAgentTurn's 'queue' delivery),
so "every card is a person's" no longer holds by refusing host sends. Each card
records who wrote it (the submission's client/host vocabulary) in a new nullable
column; the drain records that origin, so a person's card ends a Stop's pause
and Orca's does not. /clear carries the author. Rows from before the column
read as a person's. The userSend-only admission gate is removed.

* docs(native-chat): state why an unrecorded card author reads as a person's

* fix(native-chat): a restart holds only cards written before it, and an idle held queue offers Resume

A restart's pause held every waiting card, including one a person typed after the restart while
Orca's own continuation ran, and nothing released it except a per-card Send. It now holds only
cards another host process wrote, the same way a Stop holds only cards queued before it.

The composer's primary button becomes Resume (Play) while nothing is typed, no turn runs and the
host holds a card Resume would send, whatever held it (Stop, restart or /clear). It calls the
existing agentSession.queuedMessagesResume, guarded against a second press in flight.

A card nothing holds keeps the run going between a turn's end and the queue's send of it, so its
Steer no longer flips to Send for the frame in between.

* fix(native-chat): the host publishes which pause holds each queued card

The host published one pause for the whole queue, so a client held every waiting card while it was
set. Between a turn's end and the queue's send of a card queued after a Stop or restart, the
composer could flash Resume and the cards Send, and a card queued after a Stop lost its
"Waiting for your answer" caption.

Each published card now carries an optional `heldBy`: the pause holding it, or null, derived from
the same rule the drain reads. A client holds only those cards; against a host without the field
it falls back to the queue-level pause.

* test(native-chat): Resume needs the queue capability and is disabled whenever Send is

* docs(native-chat): describe per-card holds in the queue contract and table comments

* fix(native-chat): the composer goes from Resume straight to Stop, and Resume returns focus

After Resume, the host lifts the hold in one update and sends the first card in a later one. In
between nothing was running, so the composer's button flashed a disabled Send. A card nothing
holds now keeps the queue's run going for the button too: an empty composer shows Stop, disabled
until the turn starts. Not when the host refuses every send (a rewind whose outcome is unknown,
read from its status), where nothing is coming. The same fix removes the Stop, Send, Stop flip
between queued turns.

Resume disables the button, which dropped keyboard focus; focus now returns to the composer.

* fix(native-chat): the host names the card its queue sends next, so the chat stays working across the gap

A turn's end, or a Resume, and the queue's send of the next card commit as two host updates. In
between nothing was running, so the working status, timer, pickers and composer button flipped
for one update. The client guessed the drain from its own copy of the host's gates, which missed a
/clear-replaced source and covered only the button.

The queue publication now carries `nextQueuedMessageId`: the drain's own next card through the
drain's own gate (`nextStructuredQueuedMessage`, which the drain step now calls), null whenever the
host would refuse the send. The client derives one fact, the queue is about to send, and every
working reader follows it; Stop stays disabled until a turn can be stopped. The client-side copy of
the gates and the status-feed rewind read are removed.

* test(native-chat): the queue's next card survives the coalescer, the reducer and a history page

* test(native-chat): build the snapshot that names the next card through its helper

* feat(native-chat): a held queue keeps its header row, and a new message asks before passing it

The queue's header row ("Queue paused because you interrupted", or Orca
restarted, or you cleared the conversation) comes back above the cards it
holds, with Resume; it names the oldest held card's pause, as the host
publishes it per card, and hides over cards held only on their own or
returned. The header's Resume and the composer's share one in-flight guard.

A held card reads Steer again whether or not a turn runs; a card held on its
own or returned keeps Send.

Sending a message while the header shows (Enter or the button) first asks
"Send message?": Clear queue deletes every card and then sends (a failed
delete sends nothing), Send message sends and keeps the cards, which follow
the new turn, and dismissing sends nothing and keeps the draft. Host
commands send as they are.

* fix(native-chat): the paused row goes while your own message is on its way to lift it

After "Send message" over a held queue, the row kept saying "Queue paused…"
until the agent accepted the new turn. The chat now reads that gap from the
outbox: while this composer's direct send is recorded by the host and not yet
accepted, the controller shows no paused row (and so no Resume or
confirmation). A refusal settles the entry and the row comes back, since the
hold did not lift. Orca's own sends never enter this outbox, and the queue's
send of a card goes under a fresh id, so neither hides it. Nothing is stored.

* fix(native-chat): a "Send message?" choice is taken once, and a failed Clear queue is one toast

The closing dialog stays mounted and clickable through its exit animation,
and a double-click or a held Enter lands twice before any re-render, so
Send message (or Clear queue) could send the captured message twice. The
pending send now lives in a ref that the first choice takes; a second one
finds nothing.

Clear queue deletes one card at a time and stops at the first failure, so a
failed press shows one toast instead of one per card.

The dialog keeps its compact width at desktop sizes and the primitive's
narrow-window gutter (`max-w-sm sm:max-w-sm`, as the other compact
confirmations).

* fix(native-chat): Clear queue's message goes out once, and keeps text typed while it waits

After Clear queue, the message waited in the composer while the cards were
deleted one by one. A second Enter in that window sent it again, and text
typed meanwhile was wiped when the chained send was accepted.

From the Clear queue choice until its message has gone out, the composer's
structured send does nothing. The chained send (and Send message's) now
carries the composition it was taken from, and the composer is cleared on
acceptance only if it still holds exactly that, as host commands already do.

Also: the v1 contract comment names `nextQueuedMessageId` and its absent-
means-null fallback, and the own-send check returns at once on an empty
outbox.

* fix(native-chat): the queue carries on after any turn, in order, and a restart sends nothing by itself

- Any accepted turn ends a Stop's or a /clear's pause, whoever sent it (a person,
  Orca's own messages, or the queue), and so does Resume. The card and submission
  author fields that only fed the old person-only rule are gone.
- The queue sends strictly in order: a card never overtakes a held one.
- After a restart nothing sends by itself and no paused row shows: the chat's next
  turn (the carry-on, or the person's own message) runs first, then the cards.
- Resume and "Send message?" are offered only while nothing runs and no prompt waits.

* fix(native-chat): after a restart no queue pause shows, and a card written before the next turn waits for it too

* fix(native-chat): a quit hands no queued card off, and the paused row goes while any turn that will lift it is on its way

- The queue stops handing cards off when the host tears down. A card sent during
  a quit was refused at close, and that refused send withdrew the chat's restart
  offer, so resuming after the relaunch sent nothing.
- The host publishes no pause while a turn sent after it (your message, Steer, or
  Orca's own) waits for the agent; a refusal shows it again. This replaces the
  client's own-send check.
- A card written after a restart is an ordinary card again: it waits while any
  card from before the restart still waits.

* refactor(native-chat): the host's paused-row-while-a-turn-is-on-its-way check in one expression

* refactor(native-chat): the composer's queue Resume rides the structured transport beside the held queue

* fix(native-chat): the "Send message?" choice ends with the pause it asked about; tests follow main's draft props

- The open dialog closes when the queue's pause lifts under it (Orca's mail, another client's
  Resume, any accepted turn): nothing is sent, the draft stays, and the next Enter sends as
  usual. The pending choice records the hold it was asked under; nothing new is stored.
- The composer-field Resume test passes main's dropScopeKey/draftScopeKey.
- The dialog test expects main's rule: only the sent text leaves the composer.

* test(mobile): a host-kept card's test stands in a Stop's pause, as this host publishes no restart pause

Main's #24660 test published queuePause 'restarted', which this branch's wire
type no longer lists, so the mobile tests typecheck ratchet failed.
2026-10-06 18:11:54 -07:00
Neil 3fb72d135d Run Node event-loop measurement after ordinary test suites (#26015) 2026-10-06 18:08:50 -07:00
Jinwoo Hong c0273b1ff7 test(e2e): move the Source Control reveal golden into its own spec so older release tags skip it (#26005) 2026-10-06 21:07:20 -04:00
Neil 2d08b3a1da Speed up expensive test fixtures (23–82% less time) (#26000)
* Advance Codex fixture deadlines with simulated clocks

* Build large file-listing fixtures without promise batches

* Compare large binary test results with native byte equality

* Speed up Claude stop-note deadline fixtures
2026-10-06 18:07:00 -07:00
Neil 37ff3873a0 Run combined localization catalog verification on Bun (#25999) 2026-10-06 18:03:47 -07:00
Brennan Benson 00e3027762 feat(agent-launch): show an agent's tab at once, where the caller asked (#25430)
* feat(agent-launch): host-assigned caller identity and a launch record written when the surface exists

Step 1 of the agent-launch unification, on main.

- The dispatcher stamps every request's caller from what its connection proved (runtime socket:
  the local CLI; the desktop's IPC: the desktop; a paired socket: its device). Params never set it.
- The launch record is written twice: once when the tab exists (what creation settled: on the
  launch command, a draft, or a submit still `unconfirmed`), and again once the prompt's fate is
  known. A restart in between finds the running agent instead of answering "unknown".
- A replay re-derives its terminal handle from the pane key in the running host, and shows
  `unconfirmed` only to callers that advertise agent.launch.prompt-unconfirmed.v1.
- The record store opens in its own slot, without building the chat host; the chat host is built
  on that same store.

Rebuilt from this PR's own commits (b9adf88da0, 0d7d9b5b30, c83f44dd72, 39d15d52c0) onto
main, without #24080/#24081. Conflicts: the delivery doc table (main's "line fits" row plus the
`unconfirmed` row), and main's journal-database open in install(), which now goes through the
record-store slot.

* feat(agent-launch): show an agent's tab at once, where the caller asked for it

An agent.launch now shows its terminal tab before admission and spawn, in
the requested placement (group and/or anchor tab), and the pane attaches to
the agent as soon as it runs. A pane whose agent can't start, or whose start
can't be confirmed, says so instead of becoming a plain shell, and keeps
saying so across restarts. A user's close of the tab or its pane during the
launch stops it and answers agent_launch_tab_closed. Whose view moves is
unchanged from main for every caller.

Rebuilt on main (with #24934) from the previous branch head 8f62858e60.

* style(runtime): one-line the tab-order map so the headless browser-tabs runtime stays under its line limit

The merge of main (#25724's emitMobileSessionTabsSnapshot metadata) plus this branch's close mark put
the file one line over max-lines. Formatting only.
2026-10-06 18:02:30 -07:00
Jinwoo HongandClaude 66c775fbf7 refactor(terminal): project main's terminal layout after each save (no reader yet) (#25682)
* refactor(terminal): publish main's terminal topology after each session write

Adds a by-value projection of each worktree's persisted terminal topology and a
publisher hooked on the two persistence funnels (scheduleSave and durable
mutations). Changed slices are pushed to the local window on
session:terminal-topology-changed with a monotonic publishSeq; a startup pull
(session:get-terminal-topology-slices) and publishSeq on pty:spawn and
session:close-terminal-surface replies are in place. No renderer consumer yet,
so behavior and saved state are unchanged. Observer failures are counted and
logged once and never reach the save.

* refactor(terminal): keep the topology publisher dormant until a reader pulls

Writes now cost nothing extra until the first session:get-terminal-topology-slices
pull (or subscribe()) takes the baseline; markDirty before that is a no-op.

* refactor(terminal): narrow topology publishing to an unwired publisher and save hook

Drop the window sink, pull handler and publishSeq replies; no listener attaches
in production. Filter sleeping records by worktree id only, and guard observer
failures once inside the publisher.

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-06 21:02:02 -04:00
Neil f0520851ab Keep mobile restore SQLite fixture in the Node test runtime 2026-10-06 17:46:39 -07:00
Neil 1040f673b1 Keep orchestration SQLite fixtures in the Node test runtime 2026-10-06 17:46:39 -07:00
Neil 14bae2e172 Route mirrored editor closes through their captured runtime owner 2026-10-06 17:46:39 -07:00
Neil 0c32e80cc8 test(codex): reuse unproven close fixture sequence 2026-10-06 17:46:39 -07:00
Neil 983098dd00 test(codex): synchronize fake clock with forced process close 2026-10-06 17:46:39 -07:00
Neil 31d85749b4 fix(editor): retain placeholders for drafts during bulk close 2026-10-06 17:46:39 -07:00
Neil 8bd567a015 fix(editor): protect document backing files and complete close cleanup 2026-10-06 17:46:39 -07:00
Neil 91bb636530 test(editor): align model and cursor custody fixtures with owner-close policy 2026-10-06 17:46:39 -07:00
NeilandChen c9f4cd7e27 fix(editor): close stale duplicate documents safely
Adapt the document-sibling cleanup from Pr1p's #23347 with exact owner and
captured provenance checks. Preserve divergent drafts and the backing bytes
needed by retained views, and keep one-pane close behavior intact.

Co-authored-by: Chen <zwq19980411@gmail.com>
2026-10-06 17:46:39 -07:00
Jinwoo Hong 9c6702bb07 test(codex): guard that turning hooks off wins over an in-flight turn-on (#26001)
Claude-Session: codex-hook-e
2026-10-06 20:36:29 -04:00
Neil 3308ff8b26 Bound YAML merge conversion and close SQLite routing review gaps (#25998)
* Close test runtime and YAML merge review gaps

* Bound YAML conversion inside explicitly tagged pairs

* Make completion notification fixture cadence deterministic
2026-10-06 17:31:56 -07:00
Jinwoo Hong 323f312819 refactor(terminal): mark layout updates from user gestures (#25680)
* refactor(terminal): mark layout updates from user gestures

Divider drag end, divider double-click reset, pane reorder drop, equalize,
and pane rename/title clear now report themselves as gestures. Their
remote pane-layout push carries an optional intent: 'gesture'; every
other persist is unmarked and byte-identical to before. Saved layouts
are unchanged, and no host reads the field yet: the params schema
accepts it and degrades any other value to unmarked, and the handler
does not forward it.

* refactor(terminal): name the layout gesture marker once, as the wire intent

Gesture sites pass onLayoutChanged('gesture') and persistLayoutSnapshot('gesture')
directly, dropping the per-gesture names and the options bag. The params schema
reads intent as z.literal('gesture') with .catch(undefined) so unknown values
still degrade to unmarked.
2026-10-06 20:30:41 -04:00
Brennan Benson 61bcca9fee fix: stop Codex and OpenCode helper servers when Orca quits or crashes (STA-9254, 2 of 2) (#25753)
* fix(supervisor): add a one-shot lifetime that runs past stdin end, and relay a provider's last output

A one-shot CLI (claude -p, codex exec -) reads its request until stdin ends, so the
supervisor's session rule (stdin end means the owner is closing) would stop it about
1 s into its answer. The one-shot lifetime passes stdin end through and leaves only
the owner-death watch and explicit signals to stop it.

The supervisor also exited as soon as its provider did, dropping output still in the
pipes when the owner reads slowly. It now waits, bounded, for the provider's output
to be relayed before exiting.

* fix(text-generation): run agent one-shots under the provider supervisor on POSIX

The Claude model-list probe, model discovery for every agent, and commit message,
pull request and branch name generation all spawned the agent CLI as a plain child
of Orca. If Orca quit, crashed or was killed while one was running, nothing stopped
it, and a CLI that hung kept running after Orca was gone.

They now run under the provider supervisor that native chat already uses, in its
one-shot lifetime, so Orca's exit stops the agent's whole process group however
Orca exits. A timeout or cancel asks the supervisor to stop (SIGTERM) and only
tears the tree down once it has had its full stop time; killing the supervisor
first would orphan the agent's group. A missing binary now reaches Orca as the
supervisor's exit 127, which is mapped back to the existing not-found message.
Windows and WSL keep spawning the agent directly.

* fix(supervisor): carry the provider argv on the supervisor's argv, map every spawn error back, and share one stop ladder

- The supervisor read the provider's command and arguments from one base64 JSON env string.
  Linux caps a single env string at 128 KiB, so an argv prompt of about 90-120 KiB, which
  passes the 120 KiB per-argument guard, failed execve with E2BIG. The provider argv now
  follows the supervisor script's '--' as real arguments; the env keeps only small fields.
- A supervisor that cannot start its provider reports Node's spawn error line and exits 127.
  That line now becomes the same error a direct spawn emits, so ENOENT still reads as
  'not found on PATH' and EACCES or any other spawn error reads as 'failed to start'.
- stopSupervisedProvider is the one ask, wait and force ladder: it gives a supervisor its full
  stop time before forcing. Agent one-shots, the Codex app-server close and the Claude child
  exit proof now share it, with the same requests, bounds and forced steps as before.

* fix(supervisor): report every provider spawn failure on one marked stderr line

Node throws most spawn failures (ENOEXEC, ENOTDIR, ELOOP, EPERM, ...) instead of emitting
them, and the supervisor had no catch, so it died with exit 1 and a stack trace that the
user saw as the agent's failure. A thrown or emitted spawn failure now exits 127 with one
marked line carrying whether it was thrown, its code and its message. Orca reads only the
last stderr line, so a runtime warning printed earlier cannot hide it, and maps it to the
message a direct spawn gave: thrown is 'could not be started', ENOENT is 'not found on
PATH', and any other emitted error is 'failed to start'.

* fix(supervisor): keep the user's Node options away from the supervisor and hand them to the provider

The supervisor runs Electron in Node mode, which honours NODE_OPTIONS and
NODE_REPL_EXTERNAL_MODULE. A user value such as a --require of a missing file stopped
the supervisor from starting, breaking even native CLIs like Codex that never load it.
The launch now takes both out of the supervisor's environment, carries them in its spec,
and restores them for the provider only, so a Node-based CLI still gets them.

* fix(text-generation): file a forced agent one-shot teardown under its own breadcrumb site

The forced tree teardown recorded every self-initiated kill as the Codex app-server's, so
a forced commit-message or model-discovery stop read as a Codex teardown in crash
breadcrumbs. The teardown now takes the caller's site; source-control stops pass the
site their Windows tree kill already uses.

* test(text-generation): cover the supervised stop under timeout and output limit; name the direct-child suites

The commit-message suites that drive fake children spawn them directly, the unsupervised
shape Windows and WSL use, so they now say so. The supervised POSIX stop gets its own
compositions: a timed-out Codex generation settles at once but holds the Codex home until
its supervisor has stopped (faithful fake, fake timers), and an agent that floods past the
output limit is stopped through its real supervisor with no process left behind.

* fix(codex): run short-lived app-server sessions under the provider supervisor on POSIX

The Codex model-list probe, the hook trust grant and the session index heal each
start a short-lived `codex app-server` as a plain child of Orca, and counted on it
exiting when its input closes. A wedged Codex (a cold model/list waiting on the
network, say) kept running after Orca quit, crashed or was killed.

These sessions now run under the provider supervisor that native chat's Codex
connection already uses, in its session lifetime, so Orca's exit stops the
server's whole process group. The session's end and its deadline share the one
stop ladder: end its input (and SIGTERM a session past its deadline), give the
supervisor its full stop time, and only then tear the tree down. A missing binary
reaches Orca as the supervisor's exit 127 and is mapped back to the spawn error,
so the trust-grant telemetry still reads it as a missing binary. Windows and WSL
keep spawning the server directly, with the same timings as before.

The supervisor now takes only the command line it starts, so the CLI build, which
also runs these sessions, no longer pulls in the native chat connection's types.

* refactor(supervisor): share the stop of a supervised child process

Agent one-shots stop their supervisor with SIGTERM through the shared stop ladder,
watching the child's own exit. That adapter moves into one helper so the Codex
backfill recovery can use it with its own stop request, instead of a copy.

* fix(codex): supervise the app-server that keeps a Codex index backfill alive

While Codex rebuilds its session index, Orca keeps a read-only `codex app-server`
running for up to an hour so Codex can finish. It was a plain child of Orca with
its input held open, so a quit, crash or kill left it running.

On POSIX it now runs under the provider supervisor in its session lifetime, so
Orca's exit stops its whole process group. Stopping it (done, aborted, or given
up) ends its input, as a Codex connection close does; the supervisor then
SIGTERMs the group and SIGKILLs it after the grace, and the tree is torn down
only if the supervisor outlives its full stop time. Windows and WSL keep the
direct spawn and the drain-first probe termination.

The spawn and stop of that process move into their own module.

* fix(opencode): stop the launch model preflight server with Orca on POSIX

Before an OpenCode launch, Orca starts `opencode serve` to read the configured
agent and models, then stops it. The server was detached into its own process
group and never exits when its input ends, so if Orca quit, crashed or was killed
during that preflight (up to 10 s), nothing ever stopped it.

On POSIX the server now runs under the provider supervisor in its one-shot
lifetime: the preflight's closed input does not stop it, and Orca's exit does.
The preflight's teardown asks the supervisor to stop (SIGTERM) and forces the
tree only after the supervisor's full stop time; signalling or SIGKILLing the
supervisor's own group would orphan the server's. A descendant that ignores
SIGTERM is now killed with the group rather than left running once the pipes
close. Windows keeps the direct spawn and its existing teardown.

* test(codex): pin when supervised session and backfill stops escalate

A session past its deadline is SIGTERMed through its supervisor rather than
waiting out the stdin-end grace, and its tree is torn down only after the
supervisor's full stop time; Windows keeps its deadline kill and 1.5 s close
wait. A supervised backfill app-server is stopped by ending its input, with the
same full stop time before any teardown.

* test(text-generation): run the direct-child suites on the Windows path and cover supervised discovery

The commit-message suites that drive fake children mocked the supervisor away on POSIX, so
they asserted a direct root SIGKILL that production no longer takes there. They now pin the
platform to Windows (with an empty PATH, so host installs cannot answer a bare agent name)
and assert the Windows kill, taskkill included. The three tests that check the host's own
discovery spawn shape run on the host and read the agent argv past the supervisor's '--'.
Model discovery gets its supervised composition: a timed-out Codex discovery settles at once
but holds the Codex home until its supervisor has stopped.

* fix(supervisor): show a supervised spawn failure in native chat as the spawn error it was

Native chat's exit errors carry the provider's stderr tail into Details. Under the supervisor
a missing CLI left the supervisor's internal spawn-failure report there instead of Node's own
'spawn <cmd> ENOENT'. The report, its parser and a display formatter now live in one module;
the Codex app-server and Claude stream-json exit errors pass the tail through the formatter,
which turns a report back into the spawn error and leaves any other stderr unchanged.

* fix(supervisor): report a spawn that failed without a pid instead of crashing on its missing pipes

When the provider spawn fails outright (EMFILE, ENFILE), Node emits 'error' later and leaves
the child with no pid and no stdio. Piping stdin into the missing pipe threw first, so the
supervisor died with exit 1 and a stack trace and never wrote its spawn-failure report. The
pipes are now wired only for a provider that started.

* refactor(supervisor): share the stop of a supervised child process

Agent one-shots stop their supervisor with SIGTERM through the shared stop ladder, watching
the child's own exit. That adapter moves into one helper beside the ladder, so other
supervised children can use it with their own stop request instead of a copy. The caller's
breadcrumb site still reaches the forced teardown. Same request, wait and force as before.

* fix(codex,opencode): file forced backfill and preflight teardowns under their own breadcrumb sites

A forced teardown of the Codex backfill app-server is filed under
'codex-state-db-backfill-recovery', and one of the OpenCode launch model
preflight under 'opencode-launch-model-preflight', instead of the generic
'codex-app-server-teardown'.

* test(codex): write the stand-in pid report atomically

A loaded host let the test read the pid file between its creation and its
write (Unexpected end of JSON input); the stand-in now renames it into place.

* fix(supervisor): give a session provider its stdin end and grace when its owner dies

The owner-death watch went straight to the group SIGTERM and cancelled any stdin-end grace,
so when Orca quit or crashed a session provider such as the Codex app-server never saw the
EOF that lets it finish writing its state (auth.json, the state database). A session whose
owner is gone now closes as an owner's stdin end does: the provider's stdin is ended, it
gets the stdin-end grace, and only then the SIGTERM and SIGKILL ladder. A one-shot already
had its EOF at the end of its request, so its owner's death still stops it at once.

* fix(opencode): stop the preflight server through the shared supervised stop, and trust only a proven stop

The preflight's own stop wrapper waited on the supervisor's pipes and counted a
forced teardown as proof, though the teardown reports success even when it found
no descendants to check, and a supervisor that failed to reap its group exits 1
with its pipes closed. It now uses stopSupervisedChildProcess with its breadcrumb
site, and counts the server stopped only when the supervisor ended on its stop
signal or relayed the server's own exit. A forced stop, or an exit of 1, returns
no context, as an unverified stop did before.

* chore(codex): state the supervised stop time in the probe and trust grant deadline comments

* test(codex): assert the signals a supervised stop sends itself, not a SIGKILL the mocked teardown never could

* fix(supervisor): kill the rest of the provider group once the provider exits on a stop

A requested stop waited out the whole SIGTERM grace for the provider's group even after the
provider itself had exited, so a SIGTERM-ignoring helper it left behind held every stop for
up to 3 s. Under a stop, the rest of the group is now SIGKILLed as soon as the provider has
exited, the same rule its own exit already follows.

* test(text-generation): cover a supervised Codex discovery past its output limit

Model discovery's supervised stop was covered only under timeout. A Codex discovery that
floods past the output limit now runs through a real supervisor: it settles with the
too-much-data error, its agent is stopped through the supervisor, and the next discovery on
the same Codex home starts only after that agent is gone. The direct-child suites' headers
now list exactly the supervised cases that are covered.

* fix(supervisor): close a provider whose owner is gone the way its owner closes it

Owner death gave every session provider the stdin-end grace, so after an Orca crash a
Claude session, whose close is a stdin end plus SIGTERM, could keep working on its turn
for a second with nobody watching. The spawn spec now names the provider's close request:
'stdin-end' (the Codex app-server drains and exits on EOF, then gets its grace) or
'stdin-end-and-sigterm' (Claude; the default). A gone owner gets that same request. One
constant per provider feeds both its spawn spec and its owner-side close, through one
requestProviderClose, so the two cannot drift. One-shots still stop at once.

* fix(codex): close short-lived sessions and the backfill app-server the way a Codex connection closes

Codex finishes its writes and exits on its stdin end, so the Codex connection's
supervisor closes it by ending stdin, and an owner that is gone now gets that
same close. The short-lived sessions and the backfill app-server now name the
same close request, through one shared constant, in their spawn spec and in
their own close: Orca quitting or crashing gives them the stdin end and its
grace before SIGTERM, instead of an immediate SIGTERM. A session past its
deadline still adds a SIGTERM, since it is wedged.

* test(codex,opencode): an owner's death drains Codex before SIGTERM, and the preflight stop no longer waits out the grace

The stand-in now records when its stdin ended and when SIGTERM arrived. A
SIGKILLed owner leaves a Codex session or backfill app-server its stdin-end grace
before SIGTERM, and the OpenCode preflight ends within 1 s of its server's SIGTERM
even with a descendant that ignores SIGTERM.

* test(codex): give the trust-grant deadline tests room for a supervised start on a loaded host

At 500 ms a loaded full run hit the deadline before the supervisor had started
the stub, which never wrote the pid the test reads.

* fix(supervisor): keep the SIGTERM grace for a session's group after its provider exits

Killing the rest of the group the moment the provider exited under a stop also reached
native chat's closes, so an MCP server, a tool's child or a dev server still in Claude's or
Codex's group was SIGKILLed mid-cleanup instead of getting the rest of the SIGTERM grace.
The early group kill now applies only to one-shots, where the saved wait was the point;
a session's stop is back to waiting out the grace for its group.

* test(claude): pin that Claude's spawn passes its close request explicitly

The spawn-spec assertion matched the default close request, so dropping Claude's explicit
request still passed. The test now checks that the spec is built with the exit-proof ladder's
own constant.

* refactor(codex): give the Codex app-server close request its own module

Other Codex app-server spawns will name the same close request as the connection does.
Holding it in its own small module lets them import it without the connection itself.

* build(cli): list the Codex close request and the provider supervisor in the CLI project

The command-line build runs the short-lived Codex app-server session, which will name the
same close request as the Codex connection. Listing the close request, the provider
supervisor it takes its type from, and the spawn-failure report the supervisor uses lets the
CLI project typecheck that import without pulling the connection in.

* test(codex): give a stand-in 20 s to report its pids on a loaded host

At a load average near 50 both owner-death tests timed out waiting for the
bundled owner's stand-in at 10 s; they pass alone.

* test(codex): start the deadline tests' clocks past the stand-in's start, and cover a server that ignores SIGTERM

The session and trust-grant deadline tests ran a 2-4 s deadline from spawn, so a
loaded host could stop the stand-in before it wrote its pid. The session test's
deadline now outlasts its pid-read budget, and the trust-grant deadlines are 8 s.
The trust-grant comment said a wedged server may ignore everything but SIGKILL,
but that stub dies on its stdin end; a new POSIX case pins that a server ignoring
both its stdin end and SIGTERM is SIGKILLed after the SIGTERM grace.

* test(text-generation): check the ENOEXEC start failure only where Node reports one

On Linux, glibc's execvp hands an executable that is not a program to /bin/sh, so both a
direct and a supervised spawn run it and it exits 127; only macOS throws ENOEXEC. The
not-a-program case now runs on macOS only; the path-through-a-file case (ENOTDIR) still
covers a thrown start failure everywhere.

* test(wsl): follow the backfill's wsl.exe spawn into its new process module

The WSL invocation boundary lists files that spawn wsl.exe directly. The
backfill recovery's spawn moved into codex-state-db-backfill-recovery-process.ts,
so the entry moves with it; the count is unchanged.

* test(codex): keep factory child_process mocks loadable now that Codex stops reach the process-table reader

The backfill recovery and Codex sessions now stop through the shared supervised
teardown, whose process-table reader binds execFile when it loads. Five rate-limit
fetcher tests mocked node:child_process with only spawn and failed at load: they
now mock the backfill recovery, as their sibling fetcher tests already do. The
account add-login tests' child_process mocks gain an execFile stub.

* fix(opencode): count no forced preflight stop as proof the server is gone

A forced teardown walks and group-kills the supervisor's tree, but the server
leads its own detached group, so a teardown verdict of 'exited' does not cover
members left in the server's group. Only a supervisor that reaped the group
itself, by its own stop or relaying the server's exit, now counts as a proven
stop. The Windows session-stop test now says why it sees no direct kill.
2026-10-06 17:24:19 -07:00
86d03ff908 fix(native-chat): sending brings the latest into view and follows the reply (#24514)
* fix(native-chat): sending brings the latest into view, with visible jumps to latest and top

Sending while scrolled up left the reader parked in old history, and the way back was a faint button. A send now resumes following, opening a row to read it stops following until it is closed, and Jump to latest and Jump to top sit above the composer. Refs #23797.

* fix(native-chat): a late or unsent answer no longer moves a reader who scrolled away

An answer's reveal is now held from the click and dropped if the reader acts before the host accepts. In the terminal lane, an answer with no terminal to write to reveals nothing.

* fix(native-chat): an empty answer no longer moves the reader

The terminal lane's reveal now uses the same check the send does, and the send-site guard test ignores comments.

* fix(native-chat): preserve upstream retry and question cancellation

* refactor(native-chat): leave Jump to top out of this change

Jump to top moves to its own PR so this one lands the send reveal and
follow behaviour on their own.

* refactor(native-chat): keep Jump to latest's existing look in this change

The solid, fading Jump to latest button moves to the UI PR (#25702) so this one carries only the scroll behaviour.

* Preserve native chat reader intent across submissions and disclosure layout

Replace open-row lifetime tracking with bounded position preservation, and tie delayed reveals to the originating reader and session. Keep queued drafts and picker actions from navigating the transcript.

Co-authored-by: Kelvin Amoaba <97001695+AmoabaKelvin@users.noreply.github.com>

* fix(native-chat): retain reader takeover until its frame ends

Pending reader takeover could keep an earlier end target once a duplicate 250 ms gate expired. Let the existing pending-frame check keep the reader's actual offset until that frame ends.

Move unchanged interactive reveal and approval projection into their existing modules so both view files meet the line limit.

Co-authored-by: Kelvin Amoaba <97001695+AmoabaKelvin@users.noreply.github.com>

* fix(native-chat): only a scroll gesture stops following; sends reveal at the press

Opening or closing a row used to stop the transcript from following the
newest output. A reader at the live end who expanded a tool run then lost
the stream. Sends that wait on the host (answers, commands, goals, option
changes) only brought the latest into view after the host replied, so the
reader watched nothing happen at the press.

Now following stops only on a reader gesture: an upward wheel or scroll key
when there is content above, a scrollbar press, a touch drag once it has
carried the view off the end, or a press on content while already away from
the end. Opening, closing, layout changes and the app's own scrolls never
stop it, and returning to the end resumes it. Every local send reveals the
latest at the press. A message that waits as a queued card, a message from
another device, and Stop leave a reader who scrolled up where they are.
Queue Resume still reveals only after the host lifts the pause, and only in
the pane that pressed it while that pane is shown.

Removes the disclosure position hold, the reader-opens wiring, the held
reveal and its reader generation tracking. The queued-card decision moves
into the outbox send so the session hook stays under the line limit. Also
restores the transcript label import and the approval card's verified send
that the merge with main dropped.

* chore: leave an unrelated fixture as main has it

* refactor(native-chat): share the terminal send paths' common options in the composer

* test(native-chat): drop a field the main merge declared twice

* test(native-chat): the pane's steer still steers the queue, and also reveals

---------

Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
2026-10-06 16:50:55 -07:00
Jinwoo Hong faed899cd3 fix(ci): run three new SQLite-backed tests in the Node runtime project (#26010)
#25888 and #25766 added tests that import the orchestration database or the
structured session runtime without registering them in the Node runtime list,
so vitest-sqlite-runtime-boundary fails on main and every PR.
2026-10-06 19:48:41 -04:00
e717ea6b4c feat(native-chat): copy chat images from the right-click menu (#24161)
* feat(native-chat): copy chat images from the right-click menu

Fixes #23904

* test(native-chat): cover Copy image and the open preview in Electron

* fix(native-chat): keep the image preview open by click target, confirm copies, pass PNGs through

The image preview stayed open through a chat-menu click by reading which
element had focus, which depends on the menu's exit-animation timing.
It now ignores an outside interaction whose target is the chat menu, the
same target check other Orca surfaces use for portaled menus. Both
previews drop their color and padding overrides on the dialog surface,
which the design system reserves to the dialog primitive.

A successful Copy image now shows "Image copied"; before, nothing told
the user the copy had finished.

An image that is already a PNG is copied as-is after the size checks,
instead of being decoded and re-encoded, which cost time and could grow
a screenshot past the clipboard size limit.

* fix(native-chat): pass an image through as PNG only when its bytes are PNG

Copy image skipped re-encoding any image whose type said PNG, but a chat
image's type comes from its file extension. A WebP or GIF saved under a
.png name was then sent unconverted, the main process could not decode
it, and the user saw "Image copied" with nothing on the clipboard. The
pass-through now checks the PNG file signature instead.

* test(native-chat): format combined image-copy and session-ID mocks

* Release chat image menu when the retained pane hides

* Validate inherited style directives in their owning scan

* Copy displayed chat images and refuse hidden menu capture

Use full-size sent-image sources, read HTTP images when copying is chosen,
and rasterize browser-decodable formats through the existing converter.
Preserve the current preview surface and prevent a hidden preview from
retaining another copy action.

Co-authored-by: Kelvin Amoaba <97001695+AmoabaKelvin@users.noreply.github.com>

* Revert "Validate inherited style directives in their owning scan"

This reverts commit fa2d91a3cd.

* Use standard dialog surfaces for chat image previews

Keep image previews open while their chat menu is used, and inherit the standard dialog padding, background and border.

Co-authored-by: Kelvin Amoaba <97001695+AmoabaKelvin@users.noreply.github.com>

* Use dialog-owned spacing in image previews

Inherit the standard dialog gap together with its padding and surface.

Co-authored-by: Kelvin Amoaba <97001695+AmoabaKelvin@users.noreply.github.com>

---------

Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
2026-10-06 16:48:24 -07:00
Brennan Benson f6448a1e26 fix(native-chat): say which retry a Claude retry is on and what it last failed with (#25818)
* fix(native-chat): say which attempt a Claude retry is on and what it last failed with

* fix(native-chat): count Claude's retries as retries, not attempts

Claude Code sends its first api_retry frame only after the first request
failed, so frame "attempt N of max_retries" is the Nth retry. Say
"Retry N of M." and store the bound as maxRetries.
2026-10-06 16:32:32 -07:00