Commit Graph
1079 Commits
Author SHA1 Message Date
Brennan Benson e4cd14d914 Remove legacy OS file-drop routing (STA-6940 PR6/6) (#26385)
* feat(file-drop): add element owner preparation plumbing

* Fix terminal and chat file drop destination ownership

* fix runtime terminal drop ownership and queued chat retries

* fix(file-drop): preserve feedback acceptance and destination ordering

* fix(file-drop): attach chat and composer files at the drop surface

* fix(file-drop): preserve project destinations and live chat availability

* fix(file-drop): keep path resolution in filesystem namespace

* test(file-drop): use filesystem path bridge in PR3 fixtures

* fix(file-drop): bind terminal drops to pane elements

* fix(file-drop): compare terminal pane identity across public views

* test(file-drop): align composer lifecycle with element-owned drops

* Fix interrupted chat composition and refuse unused quick-create drops

* Type the quick-create drop regression transfer

* fix(file-drop): let the explorer, project sidebar, tab strip and editor own OS drops

STA-6940 PR5. The file explorer tree, the project sidebar, each editor
group's tab strip and its editor area now take OS file drops on their own
element through the shared owner hook, with identity from their own render
(the explorer's shown workspace and target row, the tab strip's and editor
group's worktree and group) instead of the active worktree. The explorer and
sidebar broadcast subscribers and their legacy drop markers are deleted;
editor-open moves out of useGlobalFileDrop into editor-dropped-file-open,
which the legacy unmarked-chrome route still uses until PR6.

* fix(file-drop): keep editor delivery tied to its destination group

* fix(file-drop): guard delayed opens and share floating ownership

* fix(file-drop): remove legacy OS drop routing (STA-6940 PR6/6)

* test(file-drop): finish owner test and comment cleanup

* Clarify dropped-file default editor destination

* Retain UI test namespace for terminal resume coverage
2026-10-07 23:11:37 -07:00
Brennan Benson b1e0092d5d feat(opencode): open OpenCode 2.x in the structured chat too (#26395)
* feat(opencode): run OpenCode 2.x in the structured chat too

`opencode acp` on stable 2.x starts its own private `opencode serve --stdio`
child with the chat's environment and ends it when stdin closes, so the
chat's account pin reaches it just as it does on 1.x. Admit stable 2.x from
2.0.14 beside stable 1.x from 1.18.31; pre-releases, older releases, and
other major lines keep the terminal chat. The `opencode2` agent is unchanged.

Restart recovery already reads a 2.x session's own tables first; tests now
cover a database an upgrade left with both generations, and one whose 2.x
message table has an unknown shape (nothing found, no error).

* fix(opencode): show OpenCode's own permission option names

* test(opencode): run the 2.x recovery tests with the real-SQLite Node project
2026-10-07 23:09:40 -07:00
Neil b41e2730db Persist multiple workspace references with atomic edits 2026-10-07 22:44:25 -07:00
Brennan Benson 8fdad2a3af fix(claude): activate account profiles and remove credential replay (Step 4 of 4) (#24434)
Each saved Claude account now runs in its own CLAUDE_CONFIG_DIR folder, so Claude renews each login itself and switching no longer replays a copied credential. A claude shell function routes every launch to the selected account; the terminal daemon protocol moves to 42 so new terminals get it. After updating, each saved account signs in once: typed claude shows Claude's own sign-in, old terminals show an in-app banner, chats and the status-bar menu offer Sign in, and a one-time toast explains it. Orca prints nothing into terminals.
2026-10-07 22:51:34 -04:00
Neil 316822f2eb ci: skip cache warming for cache-test-only changes (#26399) 2026-10-07 19:46:10 -07:00
Neil acd298a263 feat(runtime): isolate the server behind a compatibility launcher (#26374)
* feat(runtime): isolate the server behind a compatibility launcher

* fix(runtime): preserve server crash exit semantics
2026-10-07 19:32:58 -07:00
Brennan Benson c71601f51c Add Pi structured native chat through its RPC mode (#25851)
* End a running call as its turn's journal row ends

A call still running when its turn ends takes the state of that turn's
row: a row another writer settled first (a person's Stop) stands, so its
calls read interrupted whatever the provider's later end reports. The
no-ending path that settled calls from the Stop row is gone, since a Stop
now leaves running calls to the provider. Adds the two Spanish strings.

* Say why a Grok turn failed, and keep task rows in Grok's own words

A failed Grok turn ended with no reason on screen: the translator dropped
every copy of Grok's message. The failed turn now gets one status row in
Orca's existing "provider did not accept this message" words with Grok's
reason, read from whichever copy arrives first (the given-up retry, the
turn's end, the prompt's completion notice, or the prompt's error answer);
later copies only fill a reason the row still lacks.

A running background command no longer reads "Background task <id>
started": a task's summary is mapped only once it has settled. A monitor
stays a monitor when the agent reads its output: a frame that names no
kind keeps the known one, and a "[monitor" command is a monitor.

A prompt's turn is marked started, so a late frame for an ended prompt
neither reopens it nor becomes the active turn. A tool's turn is held in
one place at a time.

* Read a monitor from Grok's exact output prefix

* Word a failed Grok turn in Grok's own text, not as a refused message

A turn that started and then failed was told "The provider did not accept
this message", Orca's sentence for a message refused before its turn. The
row now reads as a Codex turn-ending error does: an error status row with the
provider's own words. With no words, the dialect names the failure ("Grok
ended this turn with an error." / "Grok usage limit reached."), else the
agent's display name does.

* Settle a stopped turn's running call as its turn row ended after a restart too

The restart sweep ended every running call by the death evidence alone, so after
a person's Stop with no proof the child died the call read failed under a turn
that read interrupted. The sweep and the live dead-generation settlement now ask
the same rule the assembler does: a call in a turn already settled ends as that
row ended; only a turn still running leaves its calls to the evidence.

* Keep the dead-generation settlement under the line cap

* Register the ACP schema verify step in the PR preflight phase test

* refactor(agent-session): one required agent registry; declarations admit what they claim

A StructuredAgentRegistry, built once from the {definition, adapter}
registrations, is now a required host dependency and the router routes with
it. The adapter interface loses its optional router-only capabilities?() and
definition?(); every reader (options at rest, thread goal, rewind, the
/compact handover) asks the registry. A live session still narrows rewind
through the adapter, and the host combines declared and narrowed in one
helper. The registry refuses a registration that declares compact, a thread
goal or rewind without the adapter method behind it.

/compact is admitted by the declared capability, so an agent that declares
compact:false gets the commandRefused fact instead of a thrown error.

Also: the cut-turn notice names the agent from the catalog, create-support
builds the account home through agentSessionAccountHome, and the turn
status text older clients read names the session's agent instead of
defaulting to Codex (Claude/Codex text unchanged).

* Read ACP permissions, session events and prompt errors through the protocol client's own types

The translator now reads a permission request with the client's lenient reader, a session update
with its session-event reader, and takes only the agent's own error answer as a failed prompt's
reason, so an Orca-side error never reads as the provider's words. Tests cover protocol values
newer than this build.

* refactor(agent-session): the router applies the declared rewind itself

The router holds the registry, so its rewindSupport answers the owner's
declared rewind narrowed by the adapter, in one place no reader can bypass.
The host-side combining helper is gone; readers ask the adapter they hold.

* test(agent-session): register the agents the merged-in tests now need

The adoption replay test builds its host over this build's registered agents
instead of an adapter definition method, and the stored-form test passes the
stored agents the record guard now requires.

* chore(agent-session): the registry is the one lookup; drop the test-only empty capability record

* fix(agent-session): keep "dismiss all" restart offers dismissed for the desktop

The desktop renderer calls the restart methods as a paired runtime client
without the registered-agents capability, so it was handed a Claude/Codex
audience and its "dismiss all" took the scoped path: no dismissal fence and
no clearing of unwritten teardown witnesses, so a late teardown write could
bring a dismissed offer back.

The audience is now derived against this host's registry: a client that can
show every agent the host registered gets none, and dismisses exactly as
before (one fence for every offer). A client that truly cannot show some
agent still gets a final dismissal for what it sees: each cleared session is
fenced on its own (bounded, superseded by a dismiss-all fence) and only the
witnesses of agents it sees are dropped.

One predicate now answers which agents a client renders, for both tab
projection and restart offers; the Claude capability rule is folded in.

* fix(agent-session): a changed agent definition never hides that agent's chats

A record was readable only if every handle used the transport its agent's
definition declares today, and its account variable was declared by some
registered agent. A later build that drives the same agent over another
protocol, or under a renamed variable, would have set every existing chat of
that agent aside: no tab, no history, no error.

Readability now asks only that the record agree with itself: its agent is
registered, every handle shares one namespace owned by that agent, and the
account variable is a well-formed name. Whether this build can drive it (the
chain's transport is the one the agent speaks, and the account variable is
that agent's own, since it becomes the child's environment) is decided in
performAttach, the one admission every start of every agent passes through,
and refused there as hostUnsupported. A record pinning another agent's
variable is refused the same way rather than passing because some agent
declares it.

* refactor(agent-session): each agent's registration says where it runs and which account it pins

createSupport, create and the model catalog's account read still chose a
location rule and an account resolver by name (Claude, Codex, else none),
so a new agent would have needed a third branch beside the registration
list that already decides storage, routing and publication.

Each runtime registration now carries supportsLocation(location) and
resolveAccountHomePath({ launchEnv, location, purpose, workspacePath }),
and Claude's and Codex's rules move into their entries unchanged: a read
still syncs no home and starts no bridge, a Codex launch still trusts the
workspace first, and WSL locations resolve as before. The list exists
before the host is built, so none of these reads installs the host, and an
agent the list does not hold is answered no without opening the journal.

* test(agent-session): drop a duplicate registry key; keep fences out of the capsule state

A spread input already carries the registry, so the explicit key was
overwritten (TS2783). The session-fence helper now returns only the fence,
not the capsule state it was given.

* fix(agent-session): a scoped dismiss-all persists no per-session fence

The per-session dismissal fence added a forever-persisted capsule field
with an arbitrary cap that served no caller: only the local desktop calls
the restart methods, and on a host whose agents it all shows it gets no
audience and takes the unscoped, fenced dismissal. A scoped dismiss-all
(a caller that cannot show some registered agent) keeps that agent's
offers, clears only its audience's unwritten witnesses, and writes no
fence; a late write from this process is serialized behind it. The field
never left this unreleased branch.

* fix(agent-session): refuse an attach whose agent is not the session's own

The start check judges the record's agent while the router starts the adapter of params.agent, and the attach wire let the two differ. Admission now refuses agent !== provider as requestMalformed, so the agent checked is the agent started. Every host-built attach already sets them equal.

* fix(agent-session): offer to start a chat only when the start would accept it

The restart offer, a failure's retryable flag and the pre-send check asked only whether the adapter runs the chat's location, while the start also refuses a record this build cannot drive; such a chat was offered Resume and its Retry failed forever. One host predicate, hostCanStartRecord (location support and agentDrivesSession), now answers all four; adapterSupportsRecord stays the reader gate, and an undrivable chat's offer is kept, not retired.

* fix(agent-session): the model-catalog probe uses a record's account only if this build drives it

A record-scoped catalog read started the agent's lister under the record's account path whatever variable the record pinned. It now uses that path only when agentDrivesSession holds, and otherwise resolves the account as for a read with no record.

* refactor(agent-session): the record store admits agent ids; comments say where transport is checked

The store only ever asked whether an agent is registered, so its admission list is now the registered ids; the definition is the one home of an agent's transport and account variable. Comments that said the record store checks transport, and one that cited a nonexistent function, now point at agentDrivesSession and the attach admission. The isPersistedAgentSessionRecord(value, agents) call shape and the test fixture the cross-version probe imports are kept.

* docs(agent-session): the record store admits the registered agents' ids

* refactor(native-chat): Grok's registration declares where it runs; ACP no longer borrows Codex's location rule

The rule a self-supervised agent child runs under (this machine, no WSL, Windows only with process
start-time proof) is its own module that Codex and the ACP adapter both use. Grok's registration
takes the full account-home resolver signature, and D3's tests build hosts with the agent registry.

* fix(native-chat): Grok follows the ACP runtime's request contract and the managed process's close

A request the agent or a Stop cancels is answered with the agent's own cancelled reply by the code that
owns it (the runtime no longer answers a silent handler), so a Stop needs no separate decline pass. A
permission answer still being saved when the agent stopped waiting is reported unconfirmed, since the
protocol already answered it cancelled. Cancelling the agent's own turn is the plain cancel. Request
rows are matched under their generation-scoped ids. A refusal's reason comes from the dialect's wording
path. The child drops its own stderr tail and close policy for the managed process's, and a close
whose process tree was not proven gone is reported as the adapter contract asks.

* fix(native-chat): a Grok chat Orca already holds resumes without writing what Grok replays

A chat with a saved Grok session reattaches with session/resume where the agent offers it, else
session/load. Either way the call runs inside the translator's load window, so what Grok sends while
it reattaches (its saved exchange, a task the dead process left running, ended by the restart) opens
no turn and writes no row; only context usage reads on. A reply an Orca or Grok crash cut short is no
longer completed from Grok's saved history: it reads like a Claude or Codex chat's, with the existing
notice. The attach window also closes after a failed attach, and a created session that session/resume
reports missing is replaced like one session/load reports missing.

The replay reconciliation is removed: the lane no longer reads the journal, and D3's replayed-input
grammar test and completed-turn check in the assembler go with it.

* refactor(native-chat): a failed Grok reattach needs no window close of its own; its lane is replaced

* test(native-chat): D3's merged tests use the shipped declarations and the launch options main requires

* fix(native-chat): typecheck fallout of the base merges; any agent's empty chat is reusable

Main's idle-empty-chat lookup and launch join now take any registered agent, as the rest of the
launch path does. The refusal check moved into the prompt turns and the prompt-block conversion beside
the turns that send it, keeping both files in their line limit.

* fix(native-chat): a Grok Stop ends the process once Grok settles its turn; the next send resumes

Grok's session/cancel ends only the running turn: work it already moved to the background keeps
running and can begin a turn of its own after the person pressed Stop. Stop is now a session
boundary, as it is for Claude: the cancel answers open requests and lets Grok end the turn, the host
waits a bounded grace for that, then ends the process; the next send relaunches and resumes.
The adapter's own bounded close of a turn Grok began is gone. Its named-turn check stays: the host
ends the session unless the provider declines a Stop naming a turn that has since ended.

* test(native-chat): a Grok Stop ends the process only after Grok answered the cancel

* fix(native-chat): Steer on a Grok card cancels the running prompt, then sends it

A send that reached Grok while a prompt ran was held in the adapter until that turn ended: Steer
on a queued card took the card out of the host's editable queue and meant 'send after this turn'.
It now cancels the running prompt (session/cancel; the session stays) and sends as the next prompt
once Grok answers the cancel, as the common pattern does; a steer behind another cancels it in
turn, so the last one runs. The adapter holds a send only while that cancel lands, so its general
held-send queue and its holdsDispatch report are gone (every send it holds has its turn open in
the journal). An older client's mid-turn send takes the same path. capabilities.steering is
unchanged and still unread.

* refactor(native-chat): a close or Stop cancels a start through the acquire's own abort signal

The host owns the acquire it runs, so it now owns its cancellation: each attach's acquire gets an
AbortSignal, aborted from outside the session's queue by a close and by a Stop admitted now (the
same admission rule as before). The optional abandonStart adapter hook, the router's fan-out to
every adapter and the ACP adapter's session-keyed start map are gone; the ACP adapter keeps an
unkeyed set of starts only so quit can prove their children gone, and keeps a failed start's
unproven child until its exit is proven.
The hook also let a later close ask that child again. The host now does that from state it holds:
a close of a chat with no live child whose record still names an owner process with no death
evidence asks the adapter to release it. The answer is not recorded as proof (the lease probe
does that), so an owner pid an earlier Orca left is never killed or marked gone. Claude and Codex
ignore the signal and hold no such child; their release is a no-op (tested).

* fix(native-chat): a Grok crash that closes stdout before its exit still ends with Grok's last words

On macOS and Linux the agent's stdout ends before its exit is observed, with or without the
supervisor's EOF forwarding, so the connection's loss closed the journal first and its error text
became the session's ended reason, dropping Grok's stderr. The reason is now read at the proven
exit: the agent's last words when it left any, else why the connection closed. The failure already
carried them. Comments that assumed the exit comes first, that early frames past the cap refuse the
start, and that dispatch re-checks image support are corrected.

* fix(native-chat): nothing Grok sends while a held chat reattaches is written, marked as replay or not

The reattach window relied on the dialect's replay verdict, and Grok's frames read as live unless
they carry isReplay, so an unmarked chat frame during session/resume opened a turn that never
ended. D3 now marks every frame inside the window as replay before the translator reads it, so the
translator keeps only context usage whatever the agent marked; options and commands are still
adopted. The translator's load semantics are unchanged.

* test(native-chat): a Stop after a resume finds no turn an unmarked old reply opened

* test(native-chat): a resumed Grok chat keeps its last context reading; the resume refreshes only the window

* test(native-chat): a Grok background task a Stop ended reads as stopped reporting

* refactor(native-chat): quit's stop of each start answers through one promise kind

* fix(native-chat): quit aborts every start the host has in flight before draining attaches

A Grok that never answered its handshake held quit until the start's own 60 s bound, past the
20 s quit deadline. The host's teardown now aborts each in-flight acquire (and any the drain
still begins), so the adapter's own quit controller and its map of starts are gone: a start
has one canceller, the host's signal.

* fix(native-chat): a Grok start's abort stops reaching its child once the start has returned

The listener stayed on the host's signal until the attach finished committing, so a Close in that
window killed the now-live child behind the host's back and it read as Grok crashing. The start
now detaches it when it ends; a later Close goes through the session's own stop.

* fix(native-chat): a close or Stop during any attach phase stops the start before it launches

The attach began its abort controller only after reconciling leases, resolving recovery and
probing the previous owner, so a close or admitted Stop in those phases reached nothing and Grok
launched anyway. The controller now begins first, and the acquisition checks it before asking the
adapter to start.

* test(native-chat): a close during the attach's owner probe asks no adapter to start

Also renames the close test after the hook it no longer exercises.

* test(native-chat): a close's re-ask closes a Claude or Codex child a failed cleanup left

The re-ask is not a no-op for them: when the adapter still holds the child its cleanup could not
prove gone, the close stops it again as a requested close, and Claude persists the handle of the
conversation it ran so the next send resumes it. Corrects the tests' and comment's wording; the
close awaits the re-ask, bounded by each adapter's kill ladder.

* fix(native-chat): Steer during a turn Grok began itself cancels it and sends once it ends

A send while Grok ran a turn of its own (a background task waking it) went straight to Grok, which
queued it behind that turn where Orca could no longer withdraw it, while Stop treated the same turn
as the running reply. The send now waits as a steer, the turn is cancelled once, and the message
goes when the turn ends; a Stop withdraws it and an exit rejects it as never sent.

* test(native-chat): a steer whose cancel Grok never answers ends Grok and is rejected as never sent

Pins the bounded steer cancel kept from the runtime: past the bound the connection closes, the
running reply reads unverifiable, Grok's end reads as its exit, and the waiting steer is rejected
as never sent.

* fix(native-chat): a Grok crash stays a crash when a stop lands before its exit is proven

After the connection broke and the close could not prove Grok's exit, any later stop Orca asked
for (the next start, a Stop, a Close) marked the child as closed by Orca, so the crash read as a
requested close and Grok's last words were dropped; a send meanwhile was recorded unconfirmed.
The connection loss now decides the cause, and a send on that session is rejected as never sent.

* test(native-chat): fixtures this PR's registered Grok and desktop capability made stale

CI's unit shards failed on tests outside the PR's own lists. Each encodes something this PR changes
on purpose: Grok is now a registered agent (the seam test's unregistered agent is now Cursor); the
desktop now advertises registered agents (the restart-offer tests' older client drops that
capability explicitly); the attach context carries the start's abort controllers (the forget-status
double gains them); and the ACP real-host test rig sends to the host directly (listed beside the
other real-host rig in the send ratchet).

* fix(native-chat): a start quit stops is not the queued message's start failure

With quit now aborting a start it would have waited for, the delivery step recorded the aborted
start as the message's failure ("couldn't restart"). After quit has stopped delivery, the step
leaves the message to quit, which settles it as a close does ("The chat closed before this message
was sent."). The test that pinned quit waiting for that start and stopping its child now pins that
nothing is launched behind quit.

* fix(native-chat): a message sent after a Stop or close aborted a start gets its own start

A start the host aborts (an admitted Stop, a close, or quit) returned its refusal to the delivery
loop, which then rejected whatever was queued at that moment with "couldn't restart", including a
message the user sent after the Stop. The attach now reports that the host aborted it, and the loop
re-derives from the journal instead: what the Stop or close withdrew is already settled, a message
accepted since gets a start of its own, and quit's next step stops the loop. This replaces the
quit-only carve-out with the same rule for every abort and every agent.

* test(native-chat): the message sent after an aborted start is answered, so no settlement outlives the test

* fix(native-chat): a Grok model pick Grok never answers no longer holds Stop or Close

The pick runs on the session's queue. It now registers in the host's out-of-queue
abort registry beside a start, so a close, an admitted Stop or quit abandons it, and
the ACP adapter bounds it at 30 s like Claude and Codex. A late answer is still adopted.

* fix(agent-launch): a phone's launch opens a terminal for an agent whose chat it cannot show

agent.launch now reads the caller's capabilities by the rule tabs and restart offers
use (clientRendersStructuredAgent). A phone without registered-agents.v1 gets Grok as
a terminal again, as on main; the host's own callers and desktop clients are unchanged.

* fix(acp): strip every agent hook variable from the ACP child, from the shared list

ACP_CHILD_ENV_TO_DELETE was a second copy of the hook runtime keys that missed
ORCA_AGENT_HOOK_TRANSPORT; it now spreads AGENT_HOOK_RUNTIME_ENV_KEYS beside the pane
identity keys.

* refactor(native-chat): the mutation context carries the provider-wait registry itself

Keeps the host file within its line limit; one field instead of two closures over it.

* fix(agent-launch): agent.launch.v2 still vouches for Claude and Codex chats

The caller rule from the previous commit also turned Claude and Codex into terminals
for a client advertising only agent.launch.v2, whose contract says it opens a chat
(mobile retry-authority tests). Only an agent beyond those two now needs the client to
read it (clientRendersStructuredAgent); the test fixtures go back to what they were.

* refactor(native-chat): drop saved-history adoption from the timeline assembler

The common pattern discards the history a provider replays while loading a
session, so the assembler has no use for an input.history event.

* refactor(acp): drop session/load history adoption from the translator

The common pattern discards the history an agent replays during session/load,
keeping only what it says about the context window. Remove the adoption path
(acp-history-adoption.ts, the adopt option, and the historical background-task
liveness rewrite it fed) so load replay is always dropped except usage.

* refactor(native-chat): a pending input is only Orca's send now

Review follow-up to the adoption removal: drop the comment naming the
provider's saved message, and make requestedAt required since every pending
input comes from input.accepted.

* test(acp): keep the task-result status table on live frames

Review follow-up to the adoption removal: the result-status mapping was only
tested through adopted history, so run the same table on live frames, and
cover an unmarked task notice during a load being dropped.

* test(acp): a frame helper for a shell command Grok is running

* fix(acp): a Grok crash settles through the host's provider-exit batch, scoped to the turn it ended

A Grok crash ended the journal unverifiable before the adapter reported the exit, so the host's
provider-exit settlement found no running turn and wrote nothing: the adapter's failure (with
Grok's last words) never reached the journal, and a later stale-session pass wrote a bare,
thread-scoped cut-short row, so the partial reply was not folded as Claude's and Codex's are.

At a proven exit the ACP lane now ends its running turn interrupted at the exit instant, as the
host's exit contract expects of a child's own translator (Codex's does the same). When Grok's
stdout closed first (every POSIX crash), the turn is unverifiable only until the exit is proven:
the host's provider-exit settlement now takes the exit as proof naming the child's fence and
revises what that child left unverifiable in the same batch, with the turn-scoped row and the
adapter's failure. Claude and Codex write no unverifiable turn of a live child except a command
whose hand-off is in doubt; that turn is now revised at the exit instead of at the next open.

* test(acp): a crash seen first leaves the host no Grok turn to revise

* refactor(native-chat): what a gone generation left unfinished gets its own module

The settlement file passed 300 lines with the exit-proof revision. The unfinished-work reads
(capture, interrupted-by-the-exit, in-progress) are their own concept and move out unchanged,
apart from the exit proof they now take.

* refactor(native-chat): a watched exit revises what its child left unverifiable without reading Stop marks

An exit's own instant is the turn's end, so the revision needs only each row's fence: the
settlement's journal type gains itemFence alone, and the host test fakes say so.

* test(native-chat): drop the duplicate itemFence on the fake that already had one

* test(claude, codex): an exit whose stdout ended first still reports as it always did

The provider supervisor now ends Orca's stdout when the agent's ends, so on every crash EOF
arrives before the exit is seen. Claude's and Codex's connections report nothing at EOF and
report the exit, with its usual reason, once it is seen.

* fix(acp): reopen a chat with session/load, as the common pattern does

An agent that offers both now reloads its session instead of resuming it; the
reattach window still discards what it replays except context usage.

* fix(acp): drop the 60 s handshake bound; an abort fails the start's waits at once

Neither common design bounds an ACP handshake: Close, Stop and quit end a start
that never answers. The abort now also closes the connection, as a kill there
does, so the start settles even before the child's exit is proven. The
host-stopped start refusal only this bound produced goes with it; the idle
sweep keeps its words.

* fix(acp): a Stop naming an ended turn follows Claude's rule

It still stops nothing while another turn is live, but in the gap before a
follow-up's turn opens, which no client can name, it now stops what is in
flight and the session ends, as a Claude Stop does.

* fix(native-chat): a close no longer re-asks a failed start's unproven child

Neither common design retries that stop at Close, and Orca's Claude contract
re-asks only at the next start and at quit. The ACP adapter keeps the child
until its exit is proven and asks it again there, as Claude does.

* fix(acp): a message sent during a turn the agent began itself goes at once

Both common designs send it straight to the agent with no cancel; only Orca's
own running prompt is steered (cancelled, then re-prompted).

* test(native-chat): dismiss-all through a remote client's audience keeps a newer Orca's offer

Uses an audience production sends (one that cannot show every agent), per review.

* fix(acp): launch Grok as `grok agent stdio`, without the update and leader flags

The common pattern passes neither --no-auto-update, --no-leader nor
GROK_DISABLE_AUTOUPDATER; full access still adds --always-approve.

* fix(acp): an agent that ends its stdout, or answers unreadably, is not a lost connection

As in the common pattern, only a broken stdin (or Orca's own close) ends the
agent; one that closed its output but can still be written to stays until a
Stop, a close or its exit. The provider supervisor goes back to its base
content, so Claude and Codex no longer get the forwarded stdout end either.

* fix(native-chat): a person's close joining a failed one still binds the turn its child end cuts

On main every close of the chat writes its own Stop and settle. Here a later close joins the
first and writes no row, and the first's settle closed when its kill failed, so a turn that opened
in between and was cut by the next close read as failed. A person's close joining a person's close
whose Stop opened a settle now reopens that settle until its attempt is done.

Tests: a close whose kill failed still closes its settle; a turn opened between a failed close and
the next reads as the person's cancellation (each fails without its half of the fix).

* test(native-chat): Grok opens as a chat only behind the structured-chat setting

agent.launch and orchestration worker-start read the same setting as the
renderer route; pin both states for Grok on each. The setting's description no
longer names only Codex and Claude, in every catalog.

* docs(acp): generic ACP comments say what holds for every agent, not Grok

Stop ends the session for every ACP agent, as in the common pattern; the
adoption hook comment goes (adoption is not planned); a failed start's child is
retried at the next start or quit.

* test(claude, codex): type the EOF-before-exit test's streams; the supervisor no longer forwards EOF

The Claude test wrote to the child's stdout and stderr through their Readable
type, which the node typecheck rejects; it now holds its own PassThrough
streams. The comments no longer credit the reverted supervisor change.

* feat(acp): a steer's cancel asks once and never ends the agent

The runtime had one cancel: send session/cancel, wait at most 10 s for Orca's prompt to settle,
then close the connection, which ends the agent. A steer used it too, so a slow agent lost its
process just because the person added a message. requestSteerCancel() now sends session/cancel
once per prompt, cancels the agent's open requests and answers later permissions cancelled, and
never bounds or closes: the prompt's own reply ends it and the steer's prompt follows. cancel()
stays the Stop: bounded, then close. A Stop after a steer still bounds and closes. Both cancel
paths move into acp-prompt-cancel.ts over one cancel channel.

* chore(native-chat): keep the record store and recovery capsule under max-lines after the main merge

* fix(acp): a repeated steer shares the cancel in flight; say what the caller owns

Per review: a second steer before the first write lands returns that write instead of resolving
early. The steer's JSDoc says the wait for the prompt's reply is unbounded and that a prompt that
fails instead must not take the steer until the caller rebuilds the session; the Stop's says a
prompt that settles in time leaves the agent for the Stop's owner to end. The steer test now gives
the runtime a handler that would allow: the open permission's signal aborts and the late one never
reaches it.

* fix(native-chat): drop the stopDelivery the A3 merge doubled

* fix(acp): a steer's cancel asks Grok once and never ends it

A steer now uses D1's notify-only cancel. Two messages sent during a reply Grok began itself
cut that reply, as the common pattern does, and then both run; before, the queued first
message could not answer the bounded cancel and Orca ended Grok although Grok answered.
A Stop keeps the bounded cancel and its 4 s grace.

* fix(acp): a permission Grok asks with no prompt of Orca's running is declined

During a turn Grok began itself nobody asked it to act, so the request is answered
cancelled at once instead of opening a card that waits, as the common pattern does.

* fix(acp): a Grok that dies while starting is reported with its own last words

A dying process's stdout ends before its exit is seen, so the start failed as a closed
connection and Grok's stderr was lost. A start whose connection closed now waits, bounded by
the Stop grace (or a Close/Stop), for the exit before it is told.

* test: a Stop after a steer sends its own cancel; drop the import the A3 merge doubled

* test(native-chat): main's Stop-note test builds its turn context with the agent registry

* test(claude): say why the close test's fake child cast is safe

* test(native-chat): build the Stop-opened-turn test's identity and turn context the current way

The test (#25056) landed before the opaque provider handle (#24991), so main still built
the old {kind, threadId} handle; the turn context also needs this branch's agent registry.

* test(native-chat): build the Stop-opened-turn test's identity with the opaque handle

The test (#25056) landed before the opaque provider handle (#24991), so main still built
the old {kind, threadId} handle.

* test(ratchet): require src/main/provider-process now that it has landed

* test(native-chat): keep main's opaque-handle import in the Stop-opened-turn test

Main's #25706 and this branch both added the import at different lines; the merge kept both.

* test(native-chat): keep main's opaque-handle import in the Stop-opened-turn test

Main's #25706 made the same fix as this branch at a different line; the merge kept both imports.

* feat(native-chat): record a fresh provider conversation that replaced one the agent could not restore

A chat whose saved conversation the agent cannot reopen can now continue in a
fresh one: the handle chain records the new conversation as a creation that
replaces the lost one (which, why, and when), keeping every earlier link.

Rows keep a shape older builds read: the stored chain starts at the latest
replacement and carries the earlier links inside it.

* test(native-chat): build this stack's journal identities with main's opaque provider handle

Main's #24991 replaced the {kind, ...} handle with {transport, agent, nativeId}; three test
files from this stack still wrote the old shape. Same lines the downstream ACP branch uses.

* docs(acp): every reattach drops the agent's replay, not only for a chat the journal holds

* refactor(native-chat): store a replaced conversation flat; refuse it where older builds read the row

Older builds only read Claude and Codex records, so the nested stored form
protected rows no replacement can reach while adding a cap mismatch after a
downgrade. Store the chain as held, refuse a replacement in a Claude or Codex
chain until one has a stored shape older builds read, and refuse a supersession
key on a replacement that names no creation in the chain.

* refactor(native-chat): read hosts' structured agents from the app-shell services

Main grew the startup hydration hook to its line limit; the host agents sync is an app-lifetime subscription like the structured session tabs sync beside it, so it moves there.

* Use current provider handles in transition tests

* Use current provider handles in timeline fixtures

* test(native-chat): prove replacement rows survive downgrade and re-upgrade

* Require the ACP directory in the runtime import check

* test(ratchet): require src/main/acp now that this PR lands it

* feat(acp): a saved session the agent cannot reopen continues in a new one, with one warning row

When session/load (or session/resume) of a saved ACP session fails, the chat starts a new session and records it as a creation that replaces the lost one (#25747's 'replaces' link), and writes one warning row that the agent no longer remembers the earlier messages. A created session the agent reports missing is still superseded silently; a signed-out agent or a start that is over (Close, Stop, a lost agent) still fails the start.

* chore(acp): rewrap the acquire header comment

* test(acp): a start closed while the agent reopens fails without opening or announcing a new session

* Let ACP connections own their supervised agent process

* Preserve ACP cleanup evidence and isolate exit observers

* Expose ACP cleanup observations and type the permission fixture

* refactor(native-chat): the registered-agents capability lives in its own module

Main's growth put protocol-version.ts one counted line over its 300-line limit once the capability
was added; like main's other per-feature capabilities, it now has its own module, and importers read
it from there.

* refactor(acp): one connection owns the Grok process and its protocol

D3 now opens each ACP agent through createAcpAgentConnection (ACP-ALIGN #25810): one object spawns the
process on the execution host, owns its stdio and protocol, and reports its proven exit. It is built and
tracked before the handshake, so a start's abort (Close, Stop, quit) still reaches it, and a failed start
keeps that same connection for the next close to retry rather than spawning another process.

Deleted: the spawnAcpStructuredChild wrapper and its test, the raw-stream runtime assembly, the caller's
exit -> runtime.close wiring, the stdout-EOF heuristic (the connection no longer treats stdout EOF as
exit), and the 10 s steer/Stop cancel bound with requestSteerCancel. Reader control maps to
pauseReading/resumeReading; a close is connection.close after the host's existing 4 s Stop grace.

The adapter owns what the protocol no longer does: one session/cancel per running prompt however many
steers arrive (cleared with that send's settlement, retried after a failed write), and a Stop or steer
answers every open agent request the person has not already answered with the agent's own cancelled
reply. An answer already being saved when the Stop lands is sent.

Tests: blocked cancel write never holds Stop's grace, two quick steers send one cancel, a failed cancel
write is retried, a real process exiting while a child holds its stdout ends the session, and the
existing start-abort, retention, crash, connection-loss and reload-failure suites on the new rig.

* fix(acp): Grok signs in on its own machine with its API key or cached sign-in

When Grok reports that it needs authentication, Orca now names a sign-in method on the machine Grok runs
on, read from the same environment Grok was launched with: xai.api_key when XAI_API_KEY is set there and
Grok offers that method, else cached_token when Grok offers it, else none and the chat keeps the existing
not-signed-in refusal. The rule lives in Grok's launch spec; the adapter applies any agent's rule for new
and reopened sessions through the protocol client's caller-named method (authenticate, then retry once).
No new sign-in UI; interactive methods are never chosen.

* fix(acp): the adapter decides which of Grok's requests reach the person

The turn owner now admits every agent request, permission or question, from its own turn state: a
request reaches the person only while Orca's prompt runs and no steer or Stop is cutting it short (a
question may also come from a turn Grok began itself, until a Stop). Anything else gets the agent's
own cancelled reply and opens no card, so a question arriving after Stop or during a steer never
appears. A steer, like a Stop, withdraws the requests already open; an answer already being saved is
still sent. The protocol client's abort-on-cancel path is no longer used: after the connection
change its request signal aborts only when the connection closes.

* feat(native-chat): add inactive Pi RPC transport foundation

* fix(acp): a plan Grok proposes shows as a plan, with no approval card

When Grok leaves plan mode it asks the client to approve its plan (x.ai/exit_plan_mode). Orca showed a
blocking 'Approve plan / Request changes' card for it; the common pattern has no such gate. Now the
plan goes into the chat's existing Plan row (the plan-document status row Codex and ACP plan updates
already use) and the request is answered at once with 'abandoned' plus feedback telling Grok to stop and
wait for the person's feedback or a request to implement it in a later turn, so nothing is approved on
the person's behalf. Dialects gain settleRequest for requests answered without asking anyone.

* fix(orchestration): worker-start opens a Grok worker in a terminal, as before

With the structured chat setting on, worker-start decided 'structured' for Grok and then the structured
worker factory (Claude and Codex only) refused it, so the start failed; main opened a terminal Grok
worker. Worker-start now decides with no registered agents beyond Claude and Codex, so Grok gets a
terminal worker as before. agent.launch and the app's own launches still open Grok as a structured
chat. Temporary until structured workers take registered agents.

* fix(acp): a prompt answer Orca can't read ends the turn instead of hanging it

A session/prompt rejection that was not the agent's own error answer (an answer that fails Orca's
schema, or one too large to read) left the turn running: the next message became a steer with nothing
to cancel and was never sent or settled, and Stop waited its full grace. As in the common pattern, any
prompt failure now ends the turn as failed (a failed-turn row without words, since none are the
agent's) and settles the send, so the next message goes. Only a closed connection keeps the send
running, for the connection-loss path to settle.

* feat(native-chat): add bounded RPC reading control

* fix(acp): send Grok's prompt-identity extension only to agents that echo it

session/prompt carried _meta {promptId, requestId} for every ACP agent, though only Grok's dialect
echoes it (injectedPromptIdentity). Now only an agent whose dialect declares it gets the extension;
other ACP agents get a plain prompt.

* refactor(native-chat): the registered-agents capability lives in protocol-version again, as on main

This reverts 0ef6d21815. That commit moved the capability to its own module only because main's
protocol-version.ts was then one counted line over its limit; main now defines it there itself within
the limit, and main's new restart test imports it from there. Main's test also reads the desktop's
capability list as an older client; on this branch the desktop advertises registered agents, so its
older client is that list without this one capability.

* test(acp): read the sign-in method with a schema, not a type assertion

* refactor(native-chat): composer transport and Stop control in their own modules

Main's rewind change (#19338) brought NativeChatStructuredSession.tsx and use-structured-agent-session.ts
to their line limits, leaving no room for this branch's image-acceptance and unpublished-Stop lines.
The composer's transport (sends, commands, options, image acceptance) moves to
use-native-chat-structured-composer-transport.ts, and whether Stop shows and what it does moves to
structured-agent-session-stop-control.ts. Behavior is unchanged; the runtime cast on the composer's
'local' | 'remote' is now a typed return.

* test(native-chat): read registered agents by id, as main's structuredAgentsReadBy now takes

Main's A3 squash changed structuredAgentsReadBy to take agent ids; this branch's test still passed
{ agent } objects (CI typecheck TS2322).

* Add Pi structured chat over its native RPC mode

* Use shared provider spawn identity and base environment

* Complete Pi startup ownership and prompt delivery safeguards

* Use generic failure copy for Pi RPC sessions

* Admit finalized Pi output before session eviction

* Apply Pi client capability checks to all chat routes

* Keep older Pi versions on terminal chat

* Use supervised process support checks for Pi launches

* Complete the Pi location-test adapter fixture

* Retain the settled-work extraction in the main merge

* Use the settled-work module in the new reasoning sweep test

* Test that Pi prompt alerts skip clients that cannot read Pi

Main added prompt alerts to the turn-completion stream; the merge extended Pi's audience filter to them.

* Pin why a timed-out write closes the agent connection only once it is being written

A line still waiting in the queue is dropped and only its request fails, as before. A line already handed to the stream cannot be taken back and would hold every later write, a cancel included, so that timeout closes the connection; tests now cover both for the shared queue and for ACP.

* Type the prompt-alert fixture as the wire's prompt attention

Its host id is a branded type; the literal only checks under the wire type, as the completion fixture beside it does.

* Run the Pi launch-resolution test in the Node runtime project

It opens a real SQLite record store, which main's new boundary check now requires to run on Node.

* Pin that a Pi chat reopened after a restart resumes its stored session file

A fresh record store reading the journal back from disk must hand the launch the session file the chat proved before the restart, and the spawn must carry it as --session.

* Show an explicitly empty text answer as "Empty answer", not as unreadable

Live QA: answering a Pi input dialog with nothing (which Pi accepts) left the resolved card reading "Selected answer unavailable", the wording for an answer Orca cannot read. The receipt joined the answer's labels and text and turned an empty result into "no answer". It now keeps an answer whose text is present but empty (or whitespace) as an empty answer, shown as "Empty answer"; an answer naming neither text nor a known option still reads as unavailable.

Also pins that the editor dialog submits the text the user edited, not its prefill.

* refactor(pi): check and resolve Pi's binary through the shared version probe and command resolver

Pi's own `--version` runner and PATH lookup are gone. Create support and
every launch now resolve the binary through resolveStructuredAgentCommand,
so the Pi Command setting picks the binary both check and spawn, and an
unrunnable Command is refused at launch with agentCommandNotRunnable as for
the other agents. The version check runs through probeAgentCliVersion with
the same rule as before: stable 1.x of the current package; the older 0.x
package keeps the terminal chat. Pi's launch reads the shared per-agent
environment overlay and Command settings instead of Pi-only runtime deps.

* feat(pi): run stable Pi from 0.84.0 in the structured chat, not only 1.x
2026-10-07 17:44:45 -07:00
Jinwoo Hong b1e09a4c8e fix(native-chat): show an SSH session's chat history on the phone and desktop (#26334)
Fixes #26057. The hook reports the SSH host's transcript path, but native chat read it on the desktop's own disk, so SSH sessions showed an empty chat. The desktop now reads the transcript on the SSH host through its existing SSH filesystem provider, routed by the path layer that already handles WSL. A local row attesting the requested path keeps the session local; a transcript not yet written keeps the chat waiting. Desktop-only, no wire change.
2026-10-07 19:33:35 -04:00
Brennan Benson c785c986c2 feat(native-chat): teach chat agents to show inline visuals in their own folder (#26099)
* feat(native-chat): visual directive grammar and host read for a chat's visuals folder

A shared grammar for the ::orca-visual{file="..." title="..."} reply line,
the per-chat visuals folder location on the owning host, and the
agentSession.readVisual runtime method that reads one visual with lexical and
canonical containment, a 512 KiB bounded read and UTF-8 refusal.

* feat(native-chat): render chat visuals inline and in the right sidebar

Native-chat assistant replies render a ::orca-visual{...} line as the chat's
HTML visual in an opaque, scripts-only sandboxed frame: CSP first, the host
frame navigation guard registered before content runs, live theme without a
reload, fitted height, links opened in the viewer's browser only from a real
gesture, lazy mount, and one muted line when the visual cannot be shown.
Open in sidebar shows the same frame in the right sidebar, widened while it
is open and restored after.

* feat(native-chat): teach chat agents to show inline visuals in their own folder

Native-chat Claude and Codex sessions now get a per-chat visuals folder on
the host that runs them, write access to exactly that folder, its path in
ORCA_CHAT_VISUALS_DIR, and an Orca skill that teaches the ::orca-visual line.

- Skill ships as an unpacked plugin folder in desktop and headless builds.
- Claude: --plugin-dir via SDK plugins, behind the CLI version probe (now one
  shared probe for every version-gated flag); folder added to
  additionalDirectories beside the user's own.
- Codex: skills/extraRoots/set and the folder appended to the user's own
  writable roots, between initialize and the thread open, under a 2 s budget;
  unsupported, failed or hung setup opens the chat without visuals.
- A host sweep removes folders no chat record maps to, and folders whose
  local workspace is provably removed; anything unproven is kept.

* test(native-chat): the visuals skill names only theme variables the frame sets

* fix(native-chat): visual CI fixes, shared height governor, live-turn streaming hold

Registers agentSession.readVisual from the methods index so the structured
method file stays under its line budget, replaces reflective reads with checked
narrowing, moves the pure height governor to src/shared for mobile, and holds a
half-written directive tail while the turn works (structured text rows carry no
running state).

* fix(native-chat): harden the visual read and link opening

Re-checks after the open that the chat's visuals folder is still the real
directory at Orca's path, reports unexpected filesystem faults by code without
host paths, and lets one click in a visual open at most one page.

* fix(native-chat): keep visual lines out of plain-text reply surfaces; review fixes

One shared helper drops visual lines (outside fenced code) from reply text where
it becomes plain text: the structured status summary that feeds the sidebar row,
dashboard, notifications, phone rows and handoffs, and AI Vault reply previews.
Review fixes: height also counts a pinned body's overflow, only the live
frontier row holds a half-written visual line, the runaway-height stop needs the
same step repeated, and any host refusal evicts the cached revision.

* fix(native-chat): review round 1 for chat visuals delivery and cleanup

- Visuals sweep: a workspace counts as removed only when no profile's
  catalog holds it (chats and visuals are shared by every profile, catalogs
  are not); a worktree in a known project is removed only when git no
  longer records it; any unreadable profile decides nothing; one catalog
  snapshot per run; a symlinked visuals root is never walked.
- The other-profile catalog reader moves out of window/ and also returns
  project ids.
- A folder Orca itself inherited is stripped at the spawn layer for both
  agents, not only from the launch overlay.
- Claude version probe: a probe that gave no version is never remembered;
  the plugin check waits up to the probe's kill time so a slow first probe
  no longer costs a chat its skill.
- Skill: kept out of Claude's / menu, filename and theme guidance matched to
  the renderer, refused writes are not retried.
- Opt-in real Codex test for skill discovery and the writable root.

* feat(native-chat): ask the Claude CLI its version when native chat starts

The first chat after Orca starts usually finds the version known, so the
plugin and thinking-display checks answer at once instead of probing a
cold binary while the chat waits.

* fix(native-chat): resolve the visuals folder without the removed journal-paths helper

Main removed the per-chat journal paths and the journal database's state
directory; the visuals folder keeps the same sha256 layout on its own and the
read method uses the profile state directory the chat host is opened in.

* fix(native-chat): review round 2 fixes; copy a reply without its visual lines

Reply previews in Agent Session History drop visual lines per text part before
lines are folded; the frame adds a body's overflow only when the body really
overflows; fence tracking follows CommonMark closers and openers; the copy
button copies a reply without visual lines; a coded read fault keeps its cause.

* fix(native-chat): review round 2 for chat visuals delivery and cleanup

- Read git worktree records written relative to their own folder (git 2.48+),
  so a worktree on an unmounted drive in such a repo is still kept.
- Skip the running profile by its own storage folder, not the profile index a
  switch rewrites first; a profile never written to counts as empty, so the
  workspace rule is not switched off by a profile that was never opened.
- Ask for readable thinking again once the plugin check has waited for the
  version, so a slow first probe no longer drops it for the chat's life.
- Launch flag decisions move to their own module; tests use a typed record
  fixture instead of casts.

* fix(native-chat): review round 3 for the visuals sweep and version check

- Resolve a relative git worktree record against its folder's real path, so a
  project added through a link still matches and its worktree is kept.
- A profile with only backups of its data file is a lost file, not a fresh
  profile: it still stops the workspace rule.
- The readable-thinking re-check reads what is known and never starts a
  second version probe.

* fix(native-chat): update the frame's theme ref after render; read the visuals folder pair at once

* test(native-chat): declare agentSession.readVisual on the cross-version agent-session surface

* fix(native-chat): copying a reply keeps its code blocks and indentation

Removing visual lines now closes only the gap each removal leaves, instead of
collapsing blank lines across the whole reply and trimming its indentation; the
visuals folder is checked parent first again so a broken path answers the same
way every time.
2026-10-07 15:53:38 -07:00
Brennan Benson 995ef11ce7 feat(native-chat): say in the chat why Orca stopped a reply, and offer Continue (#25675)
* feat(native-chat): say in the chat why Orca stopped a reply, and offer Continue

When the Orca that runs a structured chat (this computer or a paired server) quits, updates or
crashes mid-reply, the chat's stopped row now names the cause and the machine, and a Continue
button sends the existing restart continuation for that cut turn, with or without a restart offer.

- Host: a quit/update writes one turn-scoped row for the turn its stop cut, in today's words, with
  an optional `orcaStop` cause on the providerExited fact; restart adjudication stamps how the
  previous runtime ended on the deaths it proves (crash, or the quit/update it began), so the
  crash row names it too. Older clients keep their single row.
- Host: agentSession.continueInterrupted, capability-gated, rechecks under the session lock that
  the chat still sits on that cut, so a second click or a retry sends nothing.
- Client: the row's copy names the cause and machine; Continue sits above the composer.

* test(native-chat): Continue is not held by a recovery file that never answers

* fix(native-chat): bind an Orca stop's cause to the runtime that held the agent; neutral row, Continue explains itself

- The cause now rides on the host's row itself (`orcaStop` on the status row, beside today's
  words), the same row family and id scheme as the reopen's death row.
- Each recorded owner is stamped with the Orca runtime that holds it; a death proven later (at
  restart, or when recovery stops a survivor) names how that runtime ended: the quit or update it
  began, else a crash. Owners an older build recorded, a terminal's claim, an agent that died while
  its Orca ran, and unreadable quit records all keep the generic words.
- The quitting runtime's word is written first in teardown, before the recovery wait, through a
  bounded asynchronous writer apart from the chat database.
- Row copy: one neutral sentence naming the machine and cause; it drops "You can continue in this
  conversation." while Continue is offered, and Continue's tooltip says what it does.

* test(native-chat): type the Orca-stop test fixtures so the typecheck passes

The cut turn's outcome takes the journal's outcome type, and the stand-in close reads the
provider sink through a checked lookup instead of an index that may be absent.

* fix(native-chat): call a cut a crash only when Orca's runtime started and never ended

A chat said "Orca stopped unexpectedly" whenever its runtime left no quit record, and only the
desktop quit wrote one, so a headless server's restart or update, the Settings relaunch, and a
Windows logoff all read as crashes.

Each runtime now records its own start when its chat store opens, and every graceful exit records
its end through one synchronous entry point: the desktop quit's teardown, the headless server's
stop, the in-app relaunch, the GPU-fallback restarts, the update-install watchdog, and Windows
session end. A crash is a runtime that started and never ended; a runtime with no readable record
(never written, pruned, unreadable) names no cause, so the chat keeps its generic words. One file
per runtime, written durably and only by that runtime, so a damaged file never blocks a later
write and two processes never lose each other's record.

* fix(native-chat): Continue answers once Orca accepts it, not once the agent has started

On a paired server, Continue waited for the agent to start before answering, and the client
gives a paired call 15 s. A slow start (account switch, login shell, a long resume) showed
"Couldn't continue this chat" while the agent was in fact continuing.

Continue now answers when Orca has accepted the message, as a send does. The agent's start and
answer settle afterwards, and a start that fails is the chat's own note, as before. The restart
dialog's batch still waits for the handover, which is where it counts a start as done. The
verdict helpers move to their own module to keep the continuation file within its size limit.

* fix(native-chat): a reply the user steered, or a command run after the cut, still offers Continue

The cut detector stopped at the first user message after the cut turn, so a steer the turn had
taken, or a conversation command such as /context run after the cut, removed Continue while the
row still named the cause.

The rule for what is no request of its own (a conversation command, a row its turn produced, a
send handed into a running turn) moves out of the latest-request reader into one shared
predicate, which both that reader and the cut detector use. The host's "still wanted?" check
reads the same detector, so the client and host agree.

* fix(native-chat): a death proven after an earlier settle explains the turn it ends

When a chat was read before the restart proved its old agent dead, the read could only call the
turn unverifiable. The proof then revised the turn to interrupted, but the death row was scoped
by the turn still marked running, and none was, so it landed on the conversation instead of the
turn. That cut never offered Continue, and the row did not name it as the turn's explanation.

The row is now scoped to the newest root turn the settle actually ends, running or revised.

* fix(native-chat): the cause row's words stay put, and Continue waits out a resume already running

The row's "You can continue in this conversation." came and went with the button: it showed
while a paired host's answer was still on its way, vanished when the button appeared, and came
back the moment Continue was clicked. Continue also appeared on chats the restart prompt or the
launch's own resume was already carrying on.

The row now drops that sentence wherever the chat's host can continue a cut, counting a host
that has not answered yet as able (a host that writes cause rows has Continue), so its words
never change on screen. Continue is hidden while a resume is carrying that chat on.

* fix(native-chat): a refused Continue says so once, in the composer

A Continue the host refused before accepting anything wrote a red note into the chat and brought
the button back, so each retry added another identical note; a chat the host had no record of
was refused with no word at all.

Continue now reports every refusal the same way as a failed request: the existing composer line
"Couldn't continue this chat. Try again, or send a message.", which a retry replaces rather than
repeats. A refusal before acceptance writes no note. A failure after the message was accepted
(the agent could not start) is still the chat's own note, as for any send.

* fix(native-chat): the row naming Orca's stop carries its own presentation and never folds

A client that re-words host rows it cannot name (the draft that makes these cuts read as
interruptions) treated the cause row as an older red row and replaced its words, so the cause
never showed there. The row was also folded away under its collapsed turn once shown muted.

The host's cause row now names the presentation 'orca-stop' beside today's words, failure fact and
red tone, so a client that predates both changes still prints exactly today's row, red and on
screen, and a client that re-words unnamed rows passes it through. This build shows it muted,
counts it as no failure (the reply it cut stays the turn's answer), and never folds it; the fold
field becomes `explainsTurn`, as the other change names it.

* test(native-chat): pass the session-end event without a type assertion

* test(native-chat): the row naming Orca's stop renders neutral, whoever re-presented it

Pins the rendered tone on this build: the stored red row, and the same row after a reader
re-presents it in the neutral tone with its presentation and cause kept, both render muted and
never fold. The phone draws chat rows without tone styling, so it needs no change.

* fix(native-chat): the "Couldn't continue" line goes once the chat is continued

The composer line a failed or refused Continue set stayed on screen while the agent carried on:
after an answer lost in transit, or once another client or the restart prompt continued the
chat. Only the next Continue click or the user's own send cleared it, and a click also wiped an
unrelated composer error.

The line is now derived: shown only while the chat still sits on the cut that Continue failed
on, so it goes as soon as the journal shows the chat continued, from anywhere. A Continue click
clears only its own line, and a retry answered "already continued" leaves none.

* fix(native-chat): Continue waits while an opted-in launch may still resume the chat

With "resume automatically" on, Continue showed on a quit or update cut while the launch was still
waiting for its settings and reading the restart offer, then vanished when the launch's own
resume began; a click in between sent a competing continuation.

The launch's one decision (nothing offered, ask, or resume) is now published, and the chats it
resumes are named the moment it decides, with no gap. Until it decides, and while the setting
has not loaded or is on, Continue stays hidden on this machine's chats; a paired server's chats
are not the launch's to resume and keep it.

* fix(native-chat): a runtime's end survives a late reinstall, a failed write and any clean exit

Three ways the runtime record could still read a graceful stop as a crash:

- A chat host reinstalled during the quit (a request landing after teardown began) recorded the
  runtime's start again and erased the end it had just written. A second start of the same
  runtime now keeps that end.
- When the end could not be written (a full disk), the start alone stayed and read as a crash.
  The runtime now removes its record, so its chats name no cause.
- Each `app.exit(0)` had to remember to record the end. A process 'exit' with code 0 now records
  a quit when nothing else did: Electron emits it on every quit and exit once its loop runs
  (`app.exit` -> Browser::Shutdown -> the app's 'quit' -> process 'exit'), and Node on every
  `process.exit`. The relaunch and GPU-fallback calls it covers are dropped; the quit teardown,
  the headless server's stop, the update watchdog and Windows session end keep theirs, which run
  earlier or say more.

* test(native-chat): build the re-presented row as the plain status item it is

* fix(native-chat): a Continue click clears the composer's old error, so its own failure shows

Since the "Couldn't continue" line became derived, an older composer error (such as "Remove
attachments before using a chat-session command.") outranked it: a failed Continue showed the
old error instead, and a Continue that went through left the old error on screen.

A Continue click is the user's newer action, so it clears the composer's error again, as before;
the line then shows the Continue's own failure, if any. That failure still goes away by itself
once the chat is continued, and nothing but the user's own Continue click clears an unrelated
composer error.

* fix(native-chat): a chat start compares the owner process, not the runtime stamped on it

A chat start checks that the process it just started is the one the record names, by a deep
comparison of the stored owner with the adapter's process. The store stamps that owner with the
Orca runtime holding it, so the check passed only because the store happened to return the record
from before the stamp; returning the published record would have refused every chat start with
agent_session_ownership_unknown.

The start now compares the process identity without the runtime stamp, which says who holds the
process rather than which process it is.

* test(native-chat): count agentSession.continueInterrupted among the structured methods

* refactor(native-chat): derive the structured chat's transcript session in its own hook

Main's appearance work and this branch's Continue wiring together put NativeChatStructuredSession
past the 400-line limit for components. The session the transcript reads moves, unchanged, to
use-structured-chat-live-session.ts.

* refactor(native-chat): keep the Continue capability in its own module

Main grew protocol-version.ts to its line limit; the Continue capability moves to its own module,
as other capability groups have, and the runtime list still names it.

* refactor(native-chat): keep two shared files within their line limit after the main merge

Main left agent-session-record.ts and structured-agent-session-params.ts just under 300 lines,
and this branch's additions put them over. The account-home shape check moves next to the
account-home type it checks (written without a type assertion), and the Continue params move to
their own contract module; the params catalog is regenerated. No behavior change.

* test(native-chat): compare the store directory's files without depending on listing order

The corruption test checks that no file was created or removed by comparing two recursive
listings. Their order is the runtime's: with the per-runtime record directory nested under the
store, Bun returns the same entries in a different order than Node. Both listings are now sorted.

* fix: share the path bound main's launch-directory check needs

* test: give the stop-row fold rows the draws flag main's fold now reads

* refactor: mark the launch's resume decision where the resume begins

* test: count main's new structured method alongside agentSession.continueInterrupted

* fix: the journal database keeps its folder, where runtime end records live

Main's #26038 dropped stateDirectory from JournalHostDatabase; this PR's
runtime end records are read from and written beside it.
2026-10-07 14:48:28 -07:00
Brennan Benson 0f9f199322 fix(native-chat): no saved outbox on the desktop; one send at a time, and the host owns what it accepted (#25959)
* fix(native-chat): the host owns the send queue; the window keeps no saved outbox

The desktop kept each structured chat's unsent messages in localStorage and
sent them in order, so one message whose fate was unknown froze every later
send, Retry dropped it silently, failed sends could not be discarded, and an
offline chat could send hours later. The host already records every message
and owns the queue; the window now only sends.

- One in-memory sender for composer, launch prompts and messages sent from
  outside the chat. One send in flight per chat; a transport failure resends
  the same id for up to 30 s; a refusal that proves nothing was recorded puts
  the text back in the composer with the reason; a send that went out and was
  never answered shows an in-doubt line with Send again and holds nothing up.
- The host's "unknown" rows get the same in-doubt line and Send again, from
  the journal, in every window.
- A message an older build left in localStorage is never sent: the host's
  conversation outline decides what goes back to the composer, and the copy is
  deleted once that is saved.

* fix(native-chat): send through the structured chat RPC wrapper, with its timeouts

The sender called the runtime RPC directly, skipping the per-method timeouts
every other structured chat call gets. Only a remote host's request skips the
compatibility check the sender already ran.

* fix(native-chat): an unconfirmed send goes back to the composer, with no new row line

The common pattern draws nothing extra on a message whose delivery is in
doubt once its turn is over, and puts a failed send's text back in the
composer with the reason. So a send nobody answered in time, or one a host
answers in a way that proves nothing, goes back to the composer worded as
unconfirmed, and a host "unknown" row shows nothing extra. Removes the in-doubt
phase, Send again, and its strings.

A host's made-up row for an id its journal lost now reads as unconfirmed, not
as recorded, so that message comes back instead of vanishing.

* test(native-chat): pin that nothing resends a send given back as unconfirmed

* fix(native-chat): sends survive a tab close, never resend after a Stop, and keep remote images

- A normal tab close lets sends on their way settle; what the host never took
  goes back to the conversation's draft. Only a cancelled launch or a worktree
  purge drops them.
- After any Stop, a send already on its way is never sent again under its id:
  a doubtful answer, or a resend that was due, hands its text back worded as
  unconfirmed.
- A returned image keeps the SSH connection it lives on; the connection never
  goes to the host.
- An older build's saved message the host recorded and then rejected is left
  to the host's own row, never handed back.

* fix(native-chat): a Stop or a tab close never resends a send already out, and a kept card is the card's

A Stop that landed while a same-id resend was being readied (its timer
fired, its request not out yet) let that resend go out after the Stop.
The sender now tracks whether a request is awaiting its answer: a Stop
hands back every send between attempts as unconfirmed, lets one whose
request is out settle from its answer, and never issues a request after
it. A normal tab close withdraws the same way instead of doing nothing,
so nothing more goes out and nothing is dropped.

A send the host rejected but kept as a card (keptAsQueuedMessageId) is
the card's, from its reply or the journal, and never goes back to the
composer.

Adds the freeze tests: a send whose fate is unknown holds later sends
no longer than its deadline, and one the host can neither confirm nor
deny releases the next at once.

* fix(native-chat): read a resend's turned-away call or reused id as unproven

Ports the send-answer proof contract. A call the host turned away before
running it (method_not_found, invalid_argument, unauthorized) proves only
that this request wrote nothing, so it reads as never sent on a first
attempt only; on a resend an earlier attempt may have landed, and it goes
again under the same id. A resent id the host says was already used for
other content (messageIdReused) proves nothing either, like an expired or
conflicting id.

Pins the rest of the contract: any row the host returns is its own, a
thrown error is no answer whatever its code, and an older build's
refused, held or outlived-Stop copy is handed back once and never sent.

* refactor(native-chat): type the structured composer's send with its attachment type

Keeps NativeChatStructuredSession.tsx within the file length limit.

* refactor(native-chat): drop the kept-card guards the sender never needed

A kept send's row is a rejection that was never a Stop's, so the sender
already reads it as recorded, from its reply or the journal. The test
that pins it stays; the two extra checks only covered a kept row that is
also a withdrawal, which the host never writes.

* test(native-chat): route launch tests' sends through the client wrapper the sender calls

The sender sends through callStructuredAgentSession, but these launch
tests replaced that module with a factory that answered nothing (and still
named a probe that no longer exists), so every launch prompt resent until
its 30 s deadline and the tests timed out. Each factory now hands sends to
the runtime RPC mock the tests already answer, and expectations of a local
send no longer ask for the remote-only compatibility option.

* fix(native-chat): hand a message back without importing the composer's attachment hook

The worktree purge reaches the launch prompt, which hands text back, and
the attachment hook's imports reach the store. A test that builds the
real store behind a mocked one then waited on itself and hung. Handing
back now writes images to the draft store directly, as the hook's helper
did, and a test pins that each returned image keeps its SSH connection.

* fix(native-chat): one send per chat, with no line of sends behind it

A chat with a send out took further messages into an in-memory line and
sent them one by one. A message typed behind one in doubt then hit its
own 30 s deadline and came back as not sent without ever going out.

Now, as the common pattern does, a chat takes one send at a time: while
it is out, Send is disabled and Enter leaves the text in the box. The
sender refuses a second send instead of lining it up, so the 30 s
deadline always runs from the send itself. A Stop or a tab close stops
the one send: settled from its answer if its request is out, handed back
as unconfirmed between attempts, or silently if it never went out.

Notes sent from outside the chat while its send is out stay with their
sender (not ready); a launch prompt that meets the person's own first
message waits in the composer instead of being lost.

* fix(native-chat): give a Stop-withdrawn note back to its chat once its notes were cleared

Notes sent from outside a chat clear once their message is recorded. A
message the host recorded as pending and a Stop then withdrew came back
only to its sender, which had already let go of it, so the text was
lost. The chat's draft now takes it, as an earlier build's outbox did.

* test(native-chat): type the composer-actions probe without a cast

* chore(native-chat): drop outbox wording left in comments and an empty locale group

* fix(native-chat): pace failed checks, keep the host's reason, and hold the chat for its launch prompt

- A send whose checks failed before its request went out (an unreachable
  or incompatible host, an unreadable history) was tried again at once,
  about a thousand times a second for 30 s. Attempts are now paced by
  the attempts made, whether or not their request went out.
- A host that refuses every resend by throwing (native chat turned off,
  a journal that won't open) gave back only "couldn't confirm". The
  host's reason now comes first, still without claiming not sent.
- A launch's prompt now holds the chat's one send from the click, drawn
  as sending, so a message typed while the chat starts can't overtake
  it; it goes out once the chat exists, and a cancelled launch frees it.
- Notes whose send nobody could confirm say so instead of "did not
  accept", and notes launched into a new chat let go of their text once
  that chat's composer holds it.
- An open chat keeps drawing a recorded send until its row arrives, so a
  reply that beats the history no longer makes the message flicker out.

* fix(native-chat): hold a chat's sends until it has started, and send a failed chat's message with its restart

A message sent to a chat still starting went out at once to a host that
had no record of the chat yet, read for a fence it could not get, and
came back as not sent. A message sent to a chat whose start failed
restarted it but no longer went with the restart.

While a chat starts, Send stays off and Enter leaves the text in the box,
as with a send already out. A text message sent to a chat whose start
failed restarts it and goes as the restart's first message, through the
same staged-prompt path a launch prompt takes: it holds the chat's one
send slot, drawn as sending, until the chat exists, and comes back to
the composer if the restart fails again. Notes wait while a chat starts
and ride a failed chat's restart, keeping their text until it is sent.

* chore(native-chat): test the queue request, rejection words and card hand-offs; drop an unused clear

- Pin which sends ask the host to queue, the moved rejection wording,
  and that a queue send whose card was handed off and then refused or
  withdrawn, or whose replay names a withdrawn card, is never handed
  back.
- The send-at-most-once gate now says what the desktop promises after a
  reload: it never sends the id again, and hands the text back.
- The legacy read names when it goes, and drops the notice clear nothing
  called.

* fix(native-chat): keep a sent launch prompt's entry, and give back at once what can't go out

- Cleaning up a launch prompt released its send slot even after the
  prompt had gone out, which deleted the entry the sender keeps once the
  host records it. A launch prompt recorded as pending and then
  withdrawn by a Stop was lost from both the chat and the box, and an
  open chat dropped the new chat's first message before its row arrived.
  The slot now gives back only a reservation that was never sent.
- A send stopped before its request went out by something trying again
  won't clear (this client and the server can't talk, or the host
  refused the history read) kept Send off for 30 s and then said Orca
  couldn't reach the agent. It now comes back at once with its own
  cause: not sent, since nothing went out, or unconfirmed if an earlier
  attempt did. Transport errors keep the paced retries.

* fix(native-chat): notes keep their own text through a new agent's launch

Notes sent to a New agent whose start failed came back to the notes and
also sat in the new chat's composer, so sending both delivered the text
twice. The notes keep their text until it goes out (they hold it from
the click), so the launch now says so and no composer gets a copy, on a
failed start or a refused prompt alike. The tests that asserted the copy
in the composer pinned the old double ownership and now assert the
notes are its only owner.

Notes that rode a failed chat's restart and ended unconfirmed now say
Orca couldn't confirm them, as notes sent directly do, instead of that
the agent did not accept them.

* chore(native-chat): the send-once gate says what the desktop keeps across a reload

The desktop keeps a send's id in memory only: a send still unsettled at a
reload or crash is not resent and not handed back. The coverage notes no
longer credit the renderer tests with durable identity across a remount.

* fix(native-chat): say once why notes sent to a new agent did not go

Since notes keep their own text through a new agent's launch, a prompt
the host refused, nobody could confirm, or that found the chat's send
taken came back to the notes with nothing said anywhere: the chat shows
no notice for text its caller keeps. The notes menu now reports it once,
with the toast a send to an existing chat already uses: not accepted,
couldn't confirm, or not ready. A start that failed still says so in the
new chat instead.

* fix(native-chat): retry a history refusal the host says clears, and name it at the deadline

A send stopped before its request went out came back at once for any
refusal of the history read, including ones the host names as clearing
(its journal briefly unavailable, a chat detached while the host quits),
so the automatic retry was lost. Only a version mismatch and refusals
that won't clear come back at once now; the rest go again on the paced
schedule, and if the deadline still finds nothing sent, the words are
the host's refusal rather than Orca couldn't reach the agent.

* feat(native-chat): a send makes one request, and nothing ever sends it again

The desktop resent a message under its own id for up to 30 s when the
answer was lost. A send now makes exactly one request, as the common
pattern's clients do:
- An answer that is lost, dropped or proves nothing hands the text back
  at once with "couldn't confirm… check the chat".
- A failure before the request goes out (the environment check, the
  history read for the fence, a refused connection) means nothing went
  out: the text comes back at once with its own reason, or as not sent.
- The 30 s cap stays on the one request, so a host that never answers
  can't hold Send.

Nothing ever resends, so nothing can go out after a Stop: a Stop only
takes back a send still in its pre-send checks. The resend state goes
with it (tries, generation, awaiting, stopped, the resend timer, the
send-answers-proof probe, and the first-attempt/resend split in the
evidence), and the send-once gate says the desktop never resends.

* fix(native-chat): notes keep the chat's line, and a send held behind a rewind says it was not sent

- Only a send whose text belongs to the chat's composer clears the chat's line. Notes sent from
  outside the chat no longer wipe a "couldn't confirm... Check the chat" line that explains text
  already back in the box.
- A send the host turns away behind a rewind it could not confirm went back as "not sent" but said
  "couldn't confirm what happened. Check the chat". It now gives the rewind's reason and says the
  message was not sent.
- Reliability gate names the single-request test and records a fresh evidence run; the sender
  test drops its leftover resend mocks and a duplicate Stop test.

* fix(native-chat): a send ends even when putting its text back fails

If writing the returned text into the chat's draft threw, the send never settled: it stayed
"sending", the chat refused every later send until a reload. The send now always ends after a
hand-back, the failure is logged, and the chat's line adds "Couldn't save your message."

* fix(native-chat): a message typed during /clear stays in the box and follows the chat

A /clear moves the chat to a new conversation, and its host refuses any send while it runs. A
message sent then went out, came back refused into the old conversation's draft with its line, and
vanished when the chat moved on: the box and the line now belong to the new conversation.

- A /clear holds the chat's one send slot while it runs: Enter does nothing, Send shows busy and
  the text stays in the box, as for a send that is out. Notes sent from outside get "busy".
- Once it moves the chat, the old conversation keeps taking no send until the view leaves it, and
  what is left of its draft moves into the new conversation's draft, after anything there: when the
  composer's /clear settles, and again when the view moves.

* fix(native-chat): no "Send message?" or queue clear while the chat's send is out

While a send is out (or a /clear runs) the chat takes no message, yet Enter over a held queue still
opened "Send message?", and Clear queue deleted every card before its message was refused. Enter now
does nothing there and the text stays; Clear queue re-checks and deletes nothing if a send went out
after the question opened.

* refactor(native-chat): take the /clear draft carry out of this change

Moving the old conversation's draft into the one a /clear replaces it with fixes a bug main has too
(text left in the box during a /clear stays with the old conversation), so it goes in its own change.
Kept here: a /clear holds the chat's send slot while it runs, and the conversation it moved away from
takes no send until the view leaves it.

* fix(native-chat): a /clear's hold on sends always ends

- A view that unmounted while a /clear that moves the chat was out left the old conversation
  holding its sends until a reload: the late reply kept the hold for a view that was gone. The
  reply now releases it.
- The local call for a conversation command has no deadline of its own, so a /clear that never
  answers kept Send off for good. The hold now also ends at the command's deadline (195 s, the one
  the remote call already uses), whichever comes first.

* test(native-chat): opening a chat an older build left stuck; hand back its copy in send order

Pins, through the real chat hook, sends, legacy recovery and draft store,
what opening such a chat does: the queued messages behind a message the
host recorded in doubt come back to the composer once, in order, with the
"not sent" or "couldn't confirm" wording; the chat is free to send under
its read's fence; nothing happens while the chat has no fence here.

The legacy reader now hands entries back in the order they were sent
(queuedAt), as the older build's reader did, instead of array order.

* fix(native-chat): a send this window turned away for a re-paired server comes back as not sent

A managed server's update rotates its pairing, and this window's main
process then answers the next call itself, before forwarding anything,
with runtime_environment_changed. The send read that thrown answer as
proving nothing, so the person was told Orca couldn't confirm a message
that never left. It is now handed back as not sent, with the reason.
Every other thrown answer still reads as unconfirmed once the request
may have gone out.

* test(native-chat): a message refused for an expired attachment comes back with its file and why

A paired server checks every stored file a message names when it admits it,
and refuses the whole message before recording it when one has expired. The
sender hands such a message back to the chat's draft, file included, with the
refusal's words, whether or not a view shows the chat, so it can be removed
and attached again. Ported from the saved-outbox test that came with
attaching files to a structured chat on a paired server.
2026-10-07 14:07:30 -07:00
Jinwoo Hong e05c69fe8f fix(session): at startup, a local copy never overrides an SSH-owned workspace's own copy (#26098)
* fix(session): at startup a local copy never overrides an SSH-owned workspace's own copy

Rows the local partition holds for a workspace whose repo the catalog places on
an SSH target are residue (pre-#19572 builds, relay reattach). Startup used to
keep them whenever they held a tab and skip the SSH partition for that
workspace, dropping live tabs, restoring closed ones and erasing agent-resume
records on the first save. Boot hydration now drops those local rows (keeping
open files and visit recency) so adoption takes the SSH partition whole.

* refactor(session): let adoption take a catalog-owned SSH workspace instead of pre-filtering local

Replaces the sentinel-partition split with one rule in adoption: the base's
terminal tabs no longer keep out the partition the repo catalog places the
workspace on (contested ids keep today's behavior).

* fix(session): an SSH-owned workspace's records win under shared keys; unsaved local drafts survive

Addresses review: a stale local tab sharing an id with the SSH tab kept its
layout and resume records (fill-only), and a replaced open-files row dropped
local-only unsaved drafts.

* fix(session): an SSH copy with no tabs never replaces local tabs

Addresses review: an owned workspace the SSH partition holds no tabs for keeps
today's rule, so its empty tab row cannot wipe the local tabs.

* refactor(session): drop a superseded local copy before adoption; publish path passes owned ids

Simplifies the rule: for a workspace the catalog places on this SSH host,
uncontested and with host tabs, the base's rows are dropped and the existing
gap-fill adoption runs unchanged; unsaved local drafts survive. One resolver
says which workspace each scoped entry belongs to, shared by the census, the
drop and adoption. The upload path (persistedSessionForTarget) now passes
owned ids too, so main cannot publish the stale local copy to the host.

* fix(session): match host tab rows by workspace id; a local draft beats a clean host entry

Addresses review: a host tab row under a workspace key now counts toward
superseding the local copy, and a local unsaved draft for a path the host
holds clean is kept instead of dropped.

* fix(session): scope catalog-owned ids to the ssh partition the catalog names

A workspace homed on one SSH target with leftover rows in another target's
partition no longer has the owner's adopted rows removed by the leftover's
pass.
2026-10-07 17:00:09 -04:00
Neil 7b26725ff4 Keep native watcher contracts in Node and reset runtime test caches (#26315)
* test: keep native filesystem watcher contract under Node

* test: reset canonical repository keys with runtime mocks
2026-10-07 13:39:28 -07:00
Neil 00984ccb25 ci: run full pull-request unit tests across ten shards (#26295)
* ci: run full pull-request unit tests across ten shards

* Bound required Linux package tooling setup
2026-10-07 13:12:02 -07:00
Brennan Benson 990c62e6c6 feat(native-chat): attach files to a structured chat on a paired server (#25146)
* feat(native-chat): a paired server keeps a store for chat attachments

A structured chat on a paired Orca server had nowhere to put a file the client
attached. The server now keeps a per-chat attachment store under its userData,
filled through agentSessionAttachment.uploadStart/Append/Commit/Abort and read
back for previews through agentSessionAttachment.read, behind the structured
session gate and advertised as agent-session.attachments.v1.

Nothing records cleanup as owed: a sweep re-derives what may go from the host's
chat records and journal (unfinished uploads after an hour, uploads for a chat
that never existed or that its journal never mentions after a day). The clipboard
RPC's in-flight bookkeeping moves into a shared ChunkedUploadRegistry both use.

* feat(native-chat): upload attached files into a paired server's chat store

Main streams dropped or picked files (fs:uploadPathsToAgentSessionAttachments)
and pasted images (clipboard:saveImageAsTempFile with agentSessionAttachment)
into the chat's store on its paired server, reusing the file-import slice
streamer and pinning every call to the pairing revision and server process.
The browser client's paste takes the same route.

* feat(native-chat): attach files to a structured chat on a paired server

Pasted images, files dropped from Finder/Explorer and the file picker no longer
refuse with "Local attachments are not available for remote sessions." in a
structured chat on a paired server: they upload into that server's chat store and
the chat gets only the stored path. Dropped files show as pending chips at once
so Send waits for them; every attached item carries the server, pairing and chat
it was stored for, and a send to anywhere else drops it with the existing
"changed hosts" notice. Chips and transcript images read stored files back
through the server, never from this machine's disk. An older server gets an
update notice instead of an upload. The terminal-backed chat keeps refusing.

* test(native-chat): type the attachment test doubles without bare casts

* refactor(native-chat): keep the clipboard RPC on its own upload bookkeeping

The chunked upload registry stays for the chat attachment store only, beside it.

* feat(native-chat): claim chat attachments when the host admits a message

A message's references into the host's attachment store are claimed in the same
journal transaction that makes it durable (a direct send, a queued draft, a /clear
carry). The sweep deletes an upload only after marking it in that database while
no claim exists, so a send and a sweep can never both win: a client's send that
names an expired or foreign upload is refused whole with the new attachmentExpired
reason, before anything is recorded. This replaces scanning chat transcripts for
paths.

The store is now <root>/<upload id>/<name>, refuses uploads for chats the host
does not hold, caps stored names in UTF-8 bytes and avoids Windows device names.
The preview read goes through the protected bounded read and the request's reply
budget, so an image too large for one reply is refused instead of closing the
connection.

* feat(native-chat): let structured Claude read the host's chat attachment store

Attached non-image files live outside the workspace; the store root is passed as an
added directory so the agent reads them without asking, the common pattern.

* fix(native-chat): skip a dropped folder before staging walks into it

* fix(native-chat): pin the browser client's chat paste to its server

Every upload call goes to the environment the paste was meant for and checks its
pairing and server process, as the desktop upload does; a re-pair ends the upload
instead of storing the image on another server.

* refactor(native-chat): let the host's claim be the only attachment check

The composer no longer records which server and pairing each attachment came from,
and Send no longer strips attachments it judges foreign: the host refuses an
expired or unknown stored path when it admits the message, which also covers
retries, Stop-restores and queued edits the client check missed. Previews of
stored files route through the chat's own server by path.

Removing a file's chip while it uploads now keeps its @path out of the draft. A
dropped or picked file shows its name and kind on its chip while it uploads, with
no separate progress toast, and files that did not attach are named in one notice.
A rich-text paste into a chat on an older server no longer shows the update notice
beside the pasted text.

* chore(native-chat): keep the claim hook beside the draft consume and fit the line budgets

The submission's claim runs in the same append hook as a draft's consume, owned by
the queued-message collaborator, so the journal store stays within its size limit.
Formats the new files and regenerates the runtime-required English catalog.

* test(native-chat): pin the attachment store grant in the Claude launch, restore upload result types

* fix(native-chat): create the attachment store at install and claim only exact store paths

Claude drops an added directory that does not exist when it starts, so the store
root is created when the host installs it, best effort. A commit keeps its upload
in flight until the rename lands, so a sweep never takes a slow upload's part file.
Only a path under this host's exact store root is claimed and required; any other
mention of a store path is plain text and never refuses the message.

* refactor(native-chat): reuse the newer-Orca notice and drop leftover attach options

An older server's refusal to store attachments uses the notice every other write
it refuses already shows, so one string fewer in every catalog; the attach
callbacks lose an options parameter no caller passes since the provenance check
went.

* fix(native-chat): give a message refused for an expired attachment back to the composer

Sending the same message again can never bring the attachment back, so instead of
a Retry that always fails the message's text and images return to the composer,
as a Stop's withdrawn message does, and the notice says what to remove.

* fix(native-chat): keep uploads across a prompt, say why a file did not attach

An upload that finishes while a prompt card has replaced the composer lands in the
scope's attachment and draft caches, which the composer reads back when it
returns, instead of the unmounted composer. The single failure notice names the
cause the files share, such as the size limit. Rich text pasted into a chat on an
older server no longer flashes an image chip: the chip waits for the server to
take the image. A picked command waits, as Send does, while an attachment is still
uploading.

* fix(native-chat): keep a paste's upload across a prompt; pin the Claude grant hand-off

A paste uploading into a paired server's store keeps its chip live when a prompt
card unmounts the composer, and its result is filed for the composer's return,
as a drop's is. A failure cause ending in full-width punctuation gets no extra
full stop. A runtime test pins that the store root reaches Claude's launch
resolver through the adapter.

* test(native-chat): type the grant hand-off captures without casts

* test(native-chat): pin that the composer files a paste's upload across a prompt

* fix(native-chat): retry a held rename at commit and keep cut names Windows-safe

A commit renames the part file with the Windows retry the app uses elsewhere, so
an antivirus or indexer holding it briefly no longer loses the upload. A name cut
to the byte limit is stripped of trailing dots and spaces again. The upload
pump's result cast states why it holds.

* fix(native-chat): keep attachments still on their way in the pane's attachment cache

A prompt card unmounts the composer, and an attachment still saving or uploading
lived only in that composer: one that came back showed no pending chip, so Send
went out without the file and the file then landed in the next draft. Pending
chips now live in the pane's attachment cache beside settled ones and settle
there whichever composer is showing, so Send waits for them after a remount too.
A stored file's @path goes into the pane's draft cache, which keeps it through an
input-method composition and a remount. A rich-text paste's image is registered
as a hidden pending chip before the server is asked, so Send waits for it from
the start, and it shows only once the server takes it.

* test(native-chat): a dropped file still uploading stays pending across a remount

* fix(native-chat): reveal a rich-text image in a composer that came back; keep caret inserts

A rich-text paste's held image is revealed through the pane's attachment cache,
so a composer a prompt card remounted during the server's answer shows it rather
than holding Send on an invisible chip. A stored file's @path goes in at the
caret while the composer is showing and not composing, as every other attach
does; only mid-composition or after a remount does it go to the pane's draft.

* fix(native-chat): never evict a pane's attachments while a composer shows them or one is on its way

The pane attachment cache is now where chips live, so its 128-scope bound passes
over a scope a mounted composer subscribes to or that still holds a pending
attachment, and evicts the oldest scope nobody uses instead. Protected scopes are
bounded by mounted composers and attachments in flight.

* fix(native-chat): never evict the scope just written when every older one is in use

* fix(native-chat): keep the attachment store out of the RPC method table's imports

The method table is imported far and wide (the SSH relay's CLI included), and the
attachment RPCs pulled in the store, whose preview read loads modules that read
fs constants at load: any test that mocks fs/promises without them failed to
import. The installed store now lives in a small registry module; only the host
wiring loads the store itself.

* refactor(native-chat): the submission hook gets its own module

Merging main added the operation receipt to the queued-message insert, which
put journal-queued-messages.ts over the 300-line limit. The submission hook
only uses the queued-messages object's public methods, so it moves out.

* refactor(native-chat): fit the attachment sweep and paste upload into main's line budgets

After merging main, the record store and the clipboard handlers each ran
a few lines over the 300-line budget. The sweep now asks the record store
whether a chat is recorded through the two lookups it already has
(readable or unreadable) instead of a new id-listing method, and the
paired-server paste upload lives beside the other attachment uploads.
No behavior changes.

* fix(native-chat): pending chips sit beside main's saved draft, which keeps server-stored images

Main now saves a composer's draft (text and settled images) in one store
that survives a reload or quit. This branch kept every chip, settled or
not, in its own pane cache. After the merge the saved draft owns settled
images and the pane cache holds only chips still on their way: they come
back pending in a composer a prompt card remounted, hold Send, settle into
the saved draft whichever composer is showing, and are never saved
themselves, so a restored draft cannot bring back an upload as if it were
attached.

An image a paired server stored for the chat is saved with the draft as
the real image (a pasted one is no longer turned into a "Not kept"
placeholder), and the restore check leaves it to that server, whose claim
at Send refuses one it no longer holds, instead of asking this machine's
disk for a path that only exists on the server.

Also: previews and the pending subscription move into their own small
hooks to fit the line budget, the "local attachments" notice moves to the
composer-target module so the attachments hook no longer pulls in the
upload module's imports, and tests follow main's new mocks.

* test(native-chat): give the claims test main's provider handle and message source

* fix(native-chat): a paste still uploading dies with its tab or workspace

A pasted image still uploading to a paired server sits in the pending
chip cache, which by design outlives the composer and its tab. Removing
the workspace deleted its saved drafts but left those chips, so the
upload finishing afterwards recreated and saved the deleted draft. With
the workspace's tabs gone it had no owner either, so no later workspace
cleanup could ever remove it.

A pending chip now records the owner its draft would have, resolved the
way the draft store resolves one when the chip is added. Workspace
removal and a user's tab close drop pending chips by the same owner,
conversation and tab matches they delete drafts by, and a chip that
settles after its chat's tab closed writes its draft under the owner it
was begun in.

* fix(native-chat): the claim type uses the journal's own SQLite type

* refactor(native-chat): read the pending-image handlers off the composer's attachments

Main's multi-file picker (#23956) and this branch together took NativeChatComposer over the 400-line limit; reading the two pending-image handlers the way the neighbouring ones already are keeps it at the limit.

* test(native-chat): pass the launch-args resolver main now requires in the attachment grant test

* fix(native-chat): grant the attachment store beside folders the saved Claude Arguments add

Main (#25721) now builds additionalDirectories from the user's --add-dir arguments; this branch's attachment-store grant replaced that list instead of adding to it. Both are granted now. Moves the Claude session-id derivation to its own module to keep the resolver under the line limit.

* test(native-chat): run the attachment store's SQLite tests in the Node runtime project

* test(native-chat): follow main's attach ownership recheck (#25749) in the attachment tests

* test(native-chat): a named upload chip still shows when the agent takes no images

* fix(native-chat): a paired-server image paste failure keeps its error apart, as main's notice card does

* fix(native-chat): the attachment sweep reads the chat list directly now that records have no import still owed
2026-10-07 11:56:55 -07:00
Jinwoo Hong d68f3bb5b2 fix(ssh): bind reattached SSH panes into the target's own session partition (#26088)
* fix(ssh): bind reattached SSH panes into the target's own session partition

A relay reattach persisted the pane binding into the local partition, minting a
minimal copy of every reattached SSH tab there. When nothing saved over it (a
reconnect with no window open), the next startup kept that copy and skipped the
SSH partition's rows for the workspace: its open editor tabs were missing for a
launch and its agent-resume records were dropped (STA-9544).

Bind into ssh:<target>, where the spawn bound the same pane.

* test(e2e): run the reattach home-partition spec only in the Docker SSH lane

* test(e2e): keep Electron running after its last window closes on Linux in the reattach spec

* test(e2e): sync the SSH reattach spec on state, not sleeps

Waits for the SSH partition to persist the open file and for main's SSH
state to report the reattached connection, instead of fixed 3s/30s sleeps
that could pass vacuously on a slow reconnect.
2026-10-07 14:47:53 -04:00
Neil 40d35fb7bf test: run real SSH command contracts under Node (#26226) 2026-10-07 06:53:40 -07:00
2584ed906e feat: play video and music files in mobile previews (#26148)
Play workspace video and music files on mobile using bounded authenticated downloads and native media controls.

Adversarial review fixes cover rewritten files and Android policy tests. Includes iOS simulator screenshots and a playback recording in PR #26148.

Co-authored-by: ChangJun Park <40492343+ckdwns9121@users.noreply.github.com>
Co-authored-by: Lirone Levy <lirone88@outlook.fr>
Co-authored-by: dupi <david.li.du@gmail.com>
Co-authored-by: John Cusack <5961784+John-Cusack@users.noreply.github.com>
2026-10-07 03:41:42 -07:00
5cafefe726 Phase 3: every SSH host runs a managed Orca server (orcad), replacing the relay (#24863)
* Revert "revert: take the 26 Phase 3 (#16741 port) PRs back out of main (#24559)"

This reverts commit 5f308bfa9c.

* feat(orcad): Windows remote primitives for managed orcad hosts (W1) (#24525)

* feat(orcad): Windows remote primitives for managed orcad hosts (W1)

* refactor(orcad): run Windows host ops as node.exe with plain argv, no PowerShell hop

* fix(orcad): refuse secret-shaped names on the breakaway launcher's --env

* fix(orcad): name the secret env guard for its role

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(orcad): stage and commit a dormant migration catalog on the managed server (#16741 T6-9) (#24521)

* feat(orcad): stage and commit a dormant migration catalog on the managed server (#16741 T6-9)

The destination half of a catalog migration: an orcad stages a T6-7 manifest
(repositories, project groups, folder workspaces, dormant session, client,
automation and worktree metadata, retired names, scrollback snapshots) with
exclusive claims, then commits it with a receipt so a retried commit returns
the same receipt and never imports twice. Served as orcad.migration.* runtime
RPC behind the orcad.migration-catalog.v1 capability; the client refuses a
host without it or with method-not-found, and any other failure is left for
the caller to recheck. Dormant only: no live PTY projection. Inert on the
desktop until T6-10.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(rpc): catalog the orcad.migration params in the shared contract; name the catalog-import install target

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(ssh): journal and fence an SSH host for dormant migration, gated on proven terminal exit (#16741 T8-c1+c2) (#24522)

A migration from a relay-hosted SSH target into a managed orcad now starts with
a journal in its own sidecar directory, then the target's managed-owner fence,
then a profile flush, before any remote call. A fence with no journal is
unverifiable and never released; a journal whose fence is gone is stale and
grants nothing; an unreadable journal fails closed. The fence requires every
terminal the target ever leased to be proven exited, checked before the fence
(with the relay's process list) and again under it. Same-owner claims now need
the durable record that explains them. Inert until T8-c4/T6-10.

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(orcad): run orcad itself on Windows hosts (W2) (#24529)

* feat(orcad): run orcad itself on Windows hosts (W2)

* test(orcad): load the ConPTY smoke's addon from out/orcad so the temp slot can be removed

* test(orcad): skip the foreign-uid lock case when running as root

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): a connected but unused SSH host previews as movable (#24609)

The untransferred-dependency census counted activeConnectionIdsAtShutdown naming the
target as workspace-session state. The renderer rewrites that list on every connection
change, so merely connecting to an empty host blocked the move. It is a reconnect hint;
the remote work it can stand for is counted on its own. The empty-target claim check
likewise ignores global-field copies inside the host's session partition.

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(ssh): stage, commit and abort a dormant migration against its destination (#16741 T8-c3) (#24523)

* feat(ssh): stage, commit and abort a dormant migration against its destination (#16741 T8-c3)

The coordinator re-checks before every stage and commit that the fenced source
still exports the journaled manifest, carries no untransferable state and
started no terminal. A lost answer is re-read from the destination's catalog
state; only a committed read whose receipt matches the journal advances it,
and the journal is on disk before anything returns. Abort releases the fence
only on proof the destination holds nothing, or on an unsupported destination
before anything was staged, and never once the destination committed. Codes
against T6-9's catalog client through an injected interface. Inert until
T8-c4/T6-10.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor(ssh): import the T6-9 client's unsupported refusal instead of mirroring it

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(orcad): deploy, activate and roll back orcad on Windows SSH hosts (W3) (#24563)

* feat(orcad): deploy, activate and roll back orcad on Windows SSH hosts (W3)

* test(orcad): exhaustive op switch in the Windows lifecycle fake

* test(ssh): narrow the Windows host-cell descriptor by lane before building a relay cell

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(serve): run orca serve on the local orcad slot behind ORCA_SERVE_RUNTIME=orcad (T6-11) (#24608)

* feat(serve): run orca serve on the local orcad slot behind ORCA_SERVE_RUNTIME=orcad (T6-11)

* fix(serve): keep orcad selection app-side and wait out Windows temp cleanup

* refactor(orcad): move the data-root privacy check out of the instance lock

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* test(serve): prove D7 and the profile lock across a real Electron/orcad serve switch (#24619)

* test(serve): prove D7 and the profile lock across a real Electron/orcad serve switch

* ci(e2e): install ripgrep for the serve mode-switch job's window-manager wait

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(ssh): convert an SSH host with Orca state into a managed server through a journaled migration (#16741 T8-c4) (#24562)

* feat(ssh): convert an SSH host with Orca state into a managed server through a journaled migration (#16741 T8-c4)

The conversion entry resumes or takes the fence, deploys and pairs the
managed server into it, marks the server as migrated, then stages and commits
the dormant catalog. Every step is keyed by the journal, so a repeat after a
crash, deferral or lost reply resumes the same migration. Status reports an
unfinished migration, and rollback is refused while one runs or when the
rollback snapshot predates the migrated catalog. Inert until T6-10.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* style: oxfmt the c4 conversion and maintenance files

* fix(ssh): name the fake migration destination's type so declarations stay portable

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(orcad): decommission, managed stop and GC on Windows SSH hosts (W4) (#24570)

* feat(orcad): decommission, managed stop and GC on Windows SSH hosts (W4)

* fix(orcad): accept a managed stop request whose lock path is spelled with Windows client separators

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(ssh): retire a migrated SSH host's source state only after a proven commit (#16741 T8-c5) (#24565)

* feat(ssh): retire a migrated SSH host's source state only after a proven commit (#16741 T8-c5)

Once the journal records destination-committed, the source profile drops the
manifest's repositories, folder workspaces and unreferenced project groups,
its dormant session, automation, client and worktree state, and the leases
the fence proved exited. The profile flushes, the retirement is verified,
the journal moves to source-retired and compacts once the server matches.
A retry after any crash repeats idempotent work. The fenced target stays: it
carries the managed server's tunnel. Conversion now ends retired. Inert until
T6-10.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(orcad): retirement drops the migrated host from the reconnect hint

The census no longer treats activeConnectionIdsAtShutdown as untransferable (#24609), so
retirement must remove the target from it; otherwise a restart dials a host that is now a
managed server.

* style: oxfmt the c5 conversion file

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(orcad): convert Windows relay-hosted SSH targets to managed orcad (W5 part 1) (#24579)

* feat(orcad): convert Windows relay-hosted SSH targets to managed orcad (W5 part 1)

* test(orcad): start the Windows lane's exec spy after the relay gate prelude restores its own

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(settings): managed servers and "Move to managed server", behind an experimental setting (#16741 T6-6 + T6-10 UI) (#24590)

* feat(settings): managed servers and "Move to managed server", behind an experimental setting (#16741 T6-6 + T6-10 UI)

Adds a Managed servers section under Remote servers (deploy an empty server,
status with deferred-update and migration states, update, rollback, recover,
stop and cancel-stop, and SSH access for paired servers), and a Move to managed
server action on connected macOS and Linux SSH hosts with a preflight summary,
a terminals-closed confirmation and a resumable progress view. Main wires the
conversion to the relay's process list, the direct session and the T6-9 catalog
client. Everything is hidden until the new experimental setting is turned on,
and Windows SSH hosts are never offered. Merges the T6-9 branch (#24521) until
it lands on the integration branch.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(settings): align managed-server form controls and name the section the setting reveals

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(orcad): ask the relay with an absolute deadline via the W5 terminal-gate lister

The conversion wiring passed a relative 10 s as listProcesses' deadlineMs, which the
provider reads as an absolute time, so every relay inventory timed out after 1 ms and
the terminal gate could never prove exit. Adopt #24579's lister verbatim so the stacks
merge cleanly.

* feat(settings): name blocking saved state in plain, localized words

The move preview listed internal dependency ids such as workspace-session; each kind
now has its own catalog entry.

* feat(settings): offer managed servers and the move on Windows SSH hosts

W1-W5 are on the integration branch, so a Windows relay-hosted host can deploy, convert
and retire like a POSIX one. The move still waits for a connected relay that reported
its platform.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(ci): list the serve mode-switch e2e job in the release-cut permissions matrix (#24865)

* test(ci): list the serve mode-switch e2e job in the release-cut permissions matrix

* test(ci): expect the orcad Windows host cells in the SSH Windows hosts workflow

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* test(orcad): wait for killed terminal daemons to exit before removing their temp profiles (#24871)

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(updater): read rollout kill switches from the update-campaign payload, all inactive (#24867)

The nudge request Orca already polls may now carry an optional versioned rollout block
naming the Node runtime flips. A typed reader resolves each flip with kill-switch, version
range and install-id-bucketed percent semantics, falling back to the baked value (every flip
inactive) when the block is absent, invalid or never read. No consumer reads it yet.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(telemetry): report Windows security-software refusals and unverifiable or failed runtime checks (#24866)

ssh_remote_runtime_resolved dropped the self-test's security_software refusal to 'none' and
sent nothing when a self-test was unverifiable or failed, because those attempts throw before
a rung settles. Add the refusal value, self_test 'unverifiable', and an outcome field
(resolved | unverifiable | failed) deduplicated per host and outcome per session.

Co-authored-by: m4air <m4air@Mac.localdomain>

* refactor(settings): move WorktreeVisibilityDefaults into its own module (#24884)

global-settings-types.ts sits at the 300-line max-lines ceiling; merging main's two new
agent-state-rules settings with experimentalManagedServers put it at 301. The worktree
visibility defaults type moves next to the other visibility types and is re-exported so its
38 importers are unchanged.

Co-authored-by: m4air <m4air@Mac.localdomain>

* refactor(watcher): move the supervisor's child, terminating child and canary into a child slot (#24887)

* refactor(watcher): move the supervisor's child, terminating child and canary into a child slot

* fix(watcher,runtime): take the child slot's child type from the shared wrapper, and stub main's title-display clear in the projection test

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>

* ci(adhoc): build and ship the orcad template in adhoc macOS and Windows builds (#24969)

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>

* feat(orcad): bound the desktop slot cache and prune proven-stopped orcad versions after each managed deploy (#24973)

* feat(orcad): bound the desktop slot cache and prune proven-stopped orcad versions after each managed deploy

* fix(orcad): keep the in-use slot plus the two most recent others, and prove same-version reuse survives eviction

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>

* Phase 3: managed orcad on every SSH connect, with downgrade-safe fence (#24975, #24979)

* feat(orcad): fence managed hosts outside owner and keep converted projects for downgrades

Shipped builds hide any SSH target with an owner, so a downgrade made a converted host and its
projects vanish. The managed fence moves to a new orcadFence field (older builds keep but ignore
it), phase-3 owner fences migrate on load, and managed hosts stay visible but refuse a direct
relay. Conversion now stops at destination-committed with sourceRetainedAt; source retirement
waits on the baked-off orcad-source-retirement rollout flag. This build hides retained rows, and
a start that finds an older build changed them marks the host sourceChangedAt (relay, needs a
new move) instead of merging a second manifest.

* feat(ssh): every SSH host runs managed orcad, decided on each connect (#24979)

* feat(settings): managed servers are no longer experimental

Remove experimentalManagedServers and its gates; the Managed servers section always shows.
Loading drops a stored value, which an older build reads back as its own default (off).

* feat(ssh): every SSH host runs managed orcad, decided on each connect

Before a relay is started, the connect decides the host's server: a converted host connects
through its tunnel (retiring a retained source once orcad-source-retirement is on), an empty
host deploys orcad, and a host with Orca state converts through the journaled migration. A host
with live or unproven relay terminals keeps the relay this session and converts later. A
refusal for any other reason keeps the relay and names the blocker. A host orcad can't run on
(unsupported target, no template, native preflight, runtime self-test) releases any claim or
fence it took, records why with this app version, and keeps the pinned-relay ladder.

Progress and the decision ride an optional SshConnectionState.managedServer field (dropped by
older clients' admission). SSH Hosts shows each host's server status; the manual move dialog
is removed.

* fix(ssh): always await the connect's server decision, after the provider authority rotates

The decision now runs after the old session and transport are torn down and the authority has
rotated synchronously, so concurrent connects still join one attempt; a shutdown that began
during the decision wins over the rotation. The IPC tests use an async double, no sync branch.

* test(renderer): the IPC events store double carries its SSH connection states

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>

* test(ssh): prove connect-time conversion to managed orcad on real Linux and Windows hosts (#24981)

* test(ssh): prove connect-time conversion to managed orcad on real Linux and Windows hosts, and the downgrade view

* fix(e2e): read the SSH host's session partition, provision the convert cell's account, and keep rollout file overrides out of packaged builds

* test(e2e): report the full relay state when the convert poll times out

* test(e2e): require the relay to list no terminals before the converting connect

* test(e2e): require no running-terminal lease before the converting connect

* test(e2e): report the host's terminal leases when the converting connect keeps the relay

* test(e2e): end the relay era with no relay shell left to respawn

* test(e2e): settle before the converting connect and name the tabs a blocking shell belongs to

* test(e2e): use the exited relay tab as the session tab, since any mounted tab starts a shell

* test(e2e): carry an editor tab through the conversion instead of a terminal tab

* test(e2e): log the conversion census inputs before the converting connect

* fix(orcad): log which state blocked a refused conversion

* test(e2e): log both session partitions before the converting connect

* fix(orcad): a source partition's copy of focus on another host no longer blocks conversion

* fix(ssh): an ssh2 forward on port 0 reports the port it bound, so managed tunnels pair

* test(e2e): give the conversion its full budget again

* test(e2e): report the migration journal phase when the conversion stalls

* test(e2e): report the connect's own result and main's state when the conversion stalls

* fix(ssh): a converted host's managed state reaches the renderer instead of staying on 'converting'

* test(e2e): print a failed server call's response

* test(e2e): give server calls the budget a fresh server's first inventory needs

* test(e2e): read the converted worktree's tabs with a scoped session.tabs.list

* test(e2e): log the converted worktree's tabs instead of asserting them, pending the server-side fix

* test(e2e): prove retirement by the dropped source rows; the journal compacts away after it

* test(e2e): drop the conversion diagnostics now the cell passes

* test(e2e): fail a hung disconnect or connect with main's state instead of the whole budget

* test(e2e): convert an upgraded relay-era profile's host on its first connect, on Docker and Windows

* test(e2e): seed the relay-era target the way addTarget registers it

* fix(ssh): a stale ssh2 forward drops a late connection instead of crashing main on 'Not connected'

* fix(orcad): the active-slot readiness probe reads Windows hosts through the host script

* ci(ssh-windows): let only the convert cell's account open the SSH local forward its managed server needs

* test(e2e): match server paths in their JSON-escaped form, for Windows backslashes

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(orcad): move what an older build added to a converted host, or keep the server's version (#24980)

* feat(orcad): move what an older build added to a converted host, or keep the server's version

A host an older build changed stays on the relay with two actions. 'Move the new projects' runs a
fresh journaled conversion of only what the host's earlier migrations didn't move: the source is
viewed with those migrations retired from it, so nothing that overlaps the server is merged, and
a row the server already holds fails the whole move at stage, before any commit. Its journal
supersedes the chain head; retained-source checks compare against its baseline, and retirement
retires every manifest in the chain only after the newest committed. 'Keep the server's version'
records the current source as the baseline and returns the host to its managed server.

* test(ipc): the runtime environment handler contract lists the delta-move channels

* test(native-chat): snapshot the journal directory after the attach's restart-offer lock is released

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): converted hosts keep their editor tabs, forwarding-refusing hosts stay on the relay, listAll settles (#25099)

* fix(orcad): publish the headless graph so session.tabs.listAll settles instead of hanging

* fix(ssh): a system SSH forward on port 0 picks a free port first and reports it

* fix(runtime): a headless host lists and closes the editor tabs its session holds, so migrated editors reach clients

* fix(ssh): keep a host that refuses TCP forwarding on the relay, and release a conversion it stranded

* test(e2e): assert the migrated editor tab and a settled listAll, and keep a forwarding-refusing host on the relay

* fix: restore the journal import after rebase, type the probe's failure code, and update headless-graph test seams

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): client focus and 'local'-stamped tabs no longer refuse a host's move to its managed server (#25110)

* fix(orcad): client focus and 'local'-stamped tabs no longer refuse a host's move to its managed server

* fix(orcad): a v1.4.218 profile focused on the SSH worktree converts, and a refusal names what blocks it

The debounced session writer never patched activeWorkspaceKey or activeWorkspaceExecutionHostId,
so the first focus stayed on disk. The all-dependency census also counted global focus copies in
every non-source partition.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(ssh): offer to move a host whose open terminals keep it on the relay (#25097)

A host with any open SSH terminal never converted: the connect gate keeps the relay while relay
terminals are live, and an open tab respawns them on every connect. The first such connect per
host per app version now marks the relay status with offerMove (recorded as
managedServerMoveOffered), which toasts "Move <host> ... Its N open terminals will restart." The
SSH Hosts status line keeps a "Move to managed server" action while terminals are live.

Confirming runs ssh:moveToManagedServer: stop the host's relay terminals through the extracted
ssh:terminateSessions path, re-run the connect gate's terminal census, refuse on anything but
exited, then reconnect so the connect-time decision runs the journaled conversion.

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(orcad): reach managed orcad over an SSH stdio bridge where sshd refuses port forwarding (#25120)

* feat(orcad): reach managed orcad over an SSH stdio bridge where sshd refuses port forwarding

Hosts with AllowTcpForwarding no were kept on the relay. The managed tunnel now
probes forwarding each time it starts and, on refusal, serves the same local port
through a second provider: each accepted socket opens one SSH exec channel running
a small bridge on the host's pinned Node, which dials orcad's loopback port.
POSIX hosts run it with node -e; Windows hosts run it as the content-addressed
host script's stdio-bridge op with base64 line framing. Bridges are capped at 8
per connection under sshd's MaxSessions default, and a lost channel only drops its
socket. Only a host where even the bridge cannot run keeps the relay, recorded as
ssh_tunnel_unavailable; the older tcp_forwarding_refused record is retried.

* test(e2e): prove the stdio bridge by refused forwarding plus a working call

* test(e2e): connect the refused-forwarding host without the relay-only repo step

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(ssh): report each connect's server decision, and show orcad.log's last lines when setup fails (#25118)

* telemetry: one enum-only event per connect decision (outcome, reason, tunnel transport, host platform, duration), plus conversion start/commit/fail, deploy failures, and the per-host move offer and its result
* errors: deploy, launch and rollback failures carry the last 40 redacted lines of the host's orcad.log, read over SSH (Windows through the node.exe host script)
* SSH Hosts: a deferred or failed setup offers its reason, log tail included, under Details

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(startup): Windows never crashes resolving userData when roaming AppData is unavailable (#25113)

* fix(startup): pin Windows appData and userData before anything resolves them

A Windows session without a loaded profile (e.g. orca serve over SSH) can fail the
roaming AppData known-folder lookup. Electron 43 then falls through to Chromium's
userData provider and crashes natively. Resolve appData first (falling back to
APPDATA, then USERPROFILE\AppData\Roaming), and set userData explicitly so
Electron's provider never runs.

* test(startup): remove the AppData fixture through the retrying helper

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): keep a terminal the previous Orca version's relay still runs instead of replacing it (#25124)

After an app update the new relay answers "not found" for a PTY the previous build's relay still
runs, because the old relay refuses this build's handshake. The client read that as absence: it
expired the lease and the pane cold-restored into an empty shell while the user's shell kept
running, unreachable. Each deploy now takes a census of this target's older relay endpoints; while
one is live or unverifiable, a not-found reattach keeps the lease and the pane binding, and the
pane says the terminal is still running under the previous Orca version. A detached lease also
keeps blocking managed-server conversion until that terminal exits.

The cross-version harness now extracts src/relay, and a new test drives v1.4.218's relay socket and
grace lifecycle with this build's endpoint probe: the probe leaves no grace deadline and reads the
old relay as live work.

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(orcad): update a managed server on connect when it runs an older build and is idle (#25122)

* feat(orcad): update a managed server on connect when it runs an older build and is idle

A connect to a managed SSH host now runs the Managed servers update when the host's
orcad differs from this app's bundled build, the template carries the host's target,
and the update planner finds no live or uncounted terminals. A rejected candidate is
restored through the activation journal; the reason is recorded per app version so
later connects don't retry it. A host a newer Orca activated is never downgraded:
the activation record now names the app version behind each build, and an explicit
rollback holds the build it left.

* test(e2e): connect without a racing disconnect after relaunch, and report each attempt

On launch the app already reaches the managed server through its tunnel; a disconnect racing that
restore cancelled the connect that runs the update.

* feat(orcad): run the update check when the launch restores a managed server's tunnel

An auto-restored host may never see an SSH connect, so its server would never update. The tunnel
restore now runs the same check, once per server per session and off the caller's path, through
the shared update-check module; a server mid-migration is left alone.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): retry the pinned Node download through transient network failures (#25154)

A Chromium network change (net::ERR_NETWORK_CHANGED) during the pinned Node
download failed the deploy and sent the host back to the relay. The archive
download now retries up to three more times, after 1s, 3s and 9s, on dropped
connections, timeouts, stalls and retryable HTTP statuses, removing the partial
file first; checksum mismatches, other HTTP errors and cancels stay final. The
transient-error classifier moves from the speech download to
src/main/network/transient-download-error.ts so both share it.

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(serve): run orca serve on orcad by default, with ORCA_SERVE_RUNTIME=electron as the opt-out (#24972)

* feat(serve): run orca serve on orcad by default, with ORCA_SERVE_RUNTIME=electron as the opt-out

* test(serve): read Electron serve's pretty-printed readiness in the CLI mode-switch e2e

* feat(serve): gate D7 on Windows and serve on orcad there by default

* test(serve): take the profile lock in the CLI mode-switch e2e, and keep Windows profile logs on failure

* test(serve): tell Electron and orcad serve apart by readiness health, and trace Windows startup

* ci(e2e): dump Electron's native log and stack on the Windows serve mode-switch job

* fix(serve): keep Windows on Electron serve until it can adopt orcad's daemon

The Windows D7 job shows Electron serve exiting before its window when it relaunches
onto a terminal daemon orcad forked. Restore the win32 fallback and skip that case
there as a known gap; the follow-up PR fixes it and re-flips Windows.

* fix(serve): let ORCA_SERVE_RUNTIME=orcad opt in on Windows while Electron stays the default

* test(serve): skip the Windows D7 cases where orcad forks the daemon, and stop cleanup hiding a failed relaunch

Test 3 hits the same Windows gap as test 2: orcad forks its own daemon there, and Electron
crashes at startup beside it. A failed relaunch also made dispose close the old, already
closed app, whose throw replaced the launch error.

* test(serve): run every Windows D7 case, with the isolated home's AppData in place

Electron 43 crashes natively (0xFFFF7003) when it resolves userData and Windows cannot find
roaming AppData. The e2e home isolation points USERPROFILE at a fresh folder with no AppData,
so later launches hit that. The harness now creates it, every D7 case runs on Windows again,
and a new case proves Electron serve starts beside another profile's live orcad daemon.

* test(serve): retry removing a Windows e2e profile while a killed daemon releases it

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(ssh): resume a terminal the previous Orca version's relay still runs, through that relay's own bridge (#25170)

* feat(ssh): resume a terminal the previous Orca version's relay still runs, through that relay's own bridge

After an app update the previous relay keeps the user's shells alive but refuses this build's
handshake. Its own relay.js --connect, run from its own version directory, presents its own bundle
hash, so on POSIX hosts the client now reaches it that way: a pane whose reattach the current relay
held for an older relay opens a route through the old bridge, takes the PTY owner role without
output flow control, reattaches the PTY with its replay, and routes every later operation on that
id to the old relay. When the last pane a route serves exits, the route hangs up and the old relay's
own idle grace retires it. Windows hosts, relocated short sockets and unreachable bridges keep the
held-pane behaviour.

The cross-version harness now builds v1.4.218's relay from its tagged sources and runs it as the
real detached daemon: a shipped client leaves a shell in it, this build resumes the pane through the
old bridge, types into it, sees its output, and watches the old relay exit on its own after the
shell does.

* fix(ssh): read the legacy relay router through its instance

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): a host re-upgraded after a downgrade can move what the older build added (#25179)

The delta view left the moved projects' session state in place whenever it held anything
unmovable, so it then counted against the delta. Relay PTY bindings, shutdown markers and the
relay consumer's recovery record blocked every move although the terminal gate already proves
those terminals exited before any move commits; the move now drops them instead.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(worktree-ps): report a lost-contact host's terminals as unverifiable, not live:0 pty:no (#25167)

During a network drop to a relay-served SSH host, every PTY on it reads as an
unconfirmed exit, so worktree ps skipped them and printed live:0 pty:no for a
terminal that was still running. The listing now counts terminals whose liveness
verdict is unverifiable into a new optional unverifiableTerminalCount, and the CLI
prints live:unverifiable pty:unverifiable (JSON carries the same word) instead of
zero. A host-confirmed exit still reads as zero.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(terminal): a held pane shows no client OS or shell, and the boundary doc says POSIX hosts resume it (#25194)

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): tunnel to the port managed orcad bound and verify it is ours (#25182)

When another Orca already listens on 6768, orcad binds a different port. The
tunnel kept forwarding to 6768, reached the other runtime, was rejected with
4001, and the connect still reported a managed server.

The tunnel now reads orcad's bound port from its active slot's readiness
(falling back to the persisted port for slots without one) and, after the
forward is up, proves the server answering is the paired runtime. On a
mismatch it re-reads the port once and fails with orcad_identity_mismatch
instead of reporting managed.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): a converted host shows only its managed server's rows, and its old editor tabs load there (#25193)

Retained relay-era project groups and the per-host SSH catalog now hide like repos and folders.
On the managed transition the renderer reloads server names, groups, folders and worktrees, then
drops the host's relay-era rows a local refresh would keep. A restored tab with no host stamp in
a workspace now owned by a managed server takes that server as owner, in place.

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude <noreply@anthropic.com>

* fix(orcad): ship the port-scan worker so managed servers detect workspace ports (#25197)

orcad resolves port-scan-command-worker-entry.js beside orcad.js, but the orcad
build never emitted it and ORCAD_ARTIFACTS never listed it, so every managed
server logged 'probe worker unavailable' and had no port detection. The build now
emits it with the other children, and the artifact list carries it, so it is
uploaded, hashed and covered by the template contract.

A new test bundles orcad and its children and fails when the bundle names a
worker or child entry the slot does not ship. Two existing gaps it found,
session-scanner-service-entry.js and wsl-transcript-fs-process-entry.js, are
listed as known and may only shrink.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): keep retrying when a reconnect loses the socket before the SSH banner (#25195)

After a network outage, a port forwarder or NAT can accept the TCP
connection and then close it before the server sends its banner. ssh2
reports that as 'Connection lost before handshake' with no errno, so the
reconnect ladder classified it as permanent and published 'error' with
no retry scheduled. Treat it as recoverable on the bounded ladder only;
the initial connect keeps its narrow classifier.

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(serve): serve on orcad by default on Windows too (#25162)

D7 now runs every case on Windows (orcad-serve-mode-switch-windows), including Electron
serve adopting a daemon orcad forked, so Windows no longer needs the Electron default.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): the terminal gate asks the relays before treating a detached terminal as running (#25200)

* fix(orcad): the terminal gate asks the relays before treating a detached terminal as running, and retiring a host drops its relay recovery record

* fix(ssh): earlier-relay census gaps and an unanswered relay stay unverifiable; asking the terminal gate changes nothing

- the gate is read-only; the conversion and delta move retire proven detached leases themselves
- an expired lease also needs every earlier-build relay to answer before it reads as exited
- a failed, input-less or truncated census marks its older relays unverifiable, so reattach holds
- a legacy relay route closes only when no attach or listing still awaits it
- a disposed session forgets the census it started

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(migration): recover interrupted conversions and delta moves, and trim dead migration code (#25226)

* refactor(migration): drop the unused full dependency census

* refactor(migration): one table for the routed UI fields a migration carries

* refactor(migration): share the session focus field list and fix stale claim comments

* fix(session): keep a fenced SSH host's source session partition intact through renderer saves

The renderer cannot see a fenced host's repos or folders, so hydration drops their
worktree rows; a save that still routes any row to ssh:<target> (a folder workspace
key stays valid) rewrote that partition without them. That changes the migration's
source manifest and loses the session a downgraded build reads back. Main now
ignores renderer writes to a fenced, unchanged host's partition.

* fix(migration): resume or back out unfinished conversions and delta moves, serialize delta moves, clean stale journals

- On connect, a registered server whose conversion never committed resumes the commit; a failure aborts it through the destination, unregisters the server and releases the fence.
- An interrupted delta move keeps its journal and gets its changed mark back; the next move resumes it from the journaled manifest or aborts it when the source has moved on.
- Delta moves run under the target lifecycle queue and re-check the head, so two concurrent moves cannot write two journal heads.
- Delta checks re-read the live source instead of a frozen copy.
- Stale journals from a stopped or never-registered server are cleared.
- Retiring a chain goes newest first and compacts only at the end.
- The journal schema tolerates fields a newer build adds.

* fix(migration): hide source rows only after commit, and retire what an older build added to a moved project

- Source rows and the renderer session guard share one rule: hidden once the server committed, shown while a migration is still in flight.
- A delta move a crash interrupted gets its changed mark back on the next start.
- Retirement removes worktree metadata and automations by moved-project scope, so an older build's additions no longer fail the leftover check forever; that check ignores rows of other migrations in the chain.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): survive a slow first daemon start and close review gaps (#25213)

- Give the terminal daemon 30 s to start on Windows and retry the spawn once before
  falling back, so a cold first deploy is not refused as daemonless.
- Say in the activation refusal that orcad.log holds the daemon's error and that the
  next connect retries.
- Managed stop completion now waits out daemon retirement plus the shutdown deadline.
- Publish the instance lock atomically and reclaim an abandoned torn lock.
- Fix the stop listener closing before it was defined on an already-present request.
- `orca serve`'s cache prune keeps the slot a running local orcad uses.
- A timed-out daemon retirement reopens admission once it is refused; a retirement that
  may have reached the daemon stays fenced.
- A headless host keeps an editor tab's unsaved draft unless the close is forced.
- Unverifiable terminals are attributed by the host's PTY record, like live ones.
- An explicit --user-data-dir is no longer overridden on Windows.
- Read Windows daemon process identity from the process table, not PowerShell.
- Remove the unread data-event incarnationId and daemon health runtime fields, share
  one process-alive and error-code check, and fold the websocket limits file back.

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(cli): stop, update, roll back and recover managed Orca servers from the CLI (#25201)

* feat(cli): stop, update, roll back and recover managed Orca servers from the CLI

orca environment status|update|rollback|recover|stop|cancel-stop call the same managed-server
actions as Settings > Managed servers, over new managedServer.* runtime RPC methods. The desktop
main process registers those actions; the runtime advertises managedServer.v1 only then, and the
CLI refuses on a runtime without it or one that answers method_not_found. stop requires --yes.

* fix(ipc): keep a missing managed-server selector a rejection, not a synchronous throw

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): Move to managed server converts the host it is offered for (#25196)

The move stops the relay terminals and passes the terminal census; with #25179 the conversion no
longer refuses on the stopped terminal's saved tab, layout leaf and pane incarnation. A new test
drives one live relay terminal through Move to a conversion whose manifest carries the tab without
its relay PTY, so it spawns a fresh shell on orcad.

ssh:terminateSessions also no longer records a not-found shutdown as terminated while an older
Orca build's relay may still run the PTY (#25124): that terminal is reported unverifiable and its
lease kept, so the move refuses instead of converting over a running shell.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): deploy the glibc 2.17 runtime on old-glibc hosts, and fall back to the pinned relay (#25180)

* fix(orcad): deploy the glibc 2.17 runtime on old-glibc hosts, and fall back to the pinned relay

Managed orcad picked its runtime from the host's libc flavour alone, so a CentOS 7 host
(glibc 2.17) got the default linux-x64-glibc Node, whose self-test fails there. The deploy
now picks the runtime by glibc the same way the relay ladder picks rung A or B, and the
host-side slot checks accept the compat target it ships.

When orcad still can't run, the host is recorded as such and the relay connect that
follows now runs the pinned-runtime ladder instead of defaulting to the host-Node relay,
so a host with no Node lands on rung B rather than failing.

The hostile-host lane deploys managed orcad on a fresh CentOS 7 host and asserts it runs on
the compat runtime with no default runtime uploaded.

* test(ssh): CentOS 7 cell asserts the compat runtime pick, then the refusal and relay rung B fallback

The compat template still ships the base @parcel/watcher binary, which needs GLIBCXX_3.4.20; CentOS 7's
libstdc++ stops at 3.4.19, so the candidate's preflight refuses. The cell now pins that end-to-end
behaviour: compat runtime picked and uploaded alone, refusal classified native_preflight, and the relay
that follows settles on rung B with no runtime setting.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* test(serve): prove D7 on an installed Windows app whose daemon runs from the relocated daemon-host (#24976)

* test(serve): prove D7 on an installed Windows app whose daemon runs from the relocated daemon-host

* test(serve): read the pre-switch scrollback best-effort and wait for it after reattach

* test(serve): opt the packaged Windows serve switch into orcad explicitly

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(migration): a retirement that fails after deleting rows resumes instead of reading as changed (#25234)

The chain head records sourceRetiringAt before any row is retired; the startup change check and the chain's own comparison skip a head that carries it.

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(orcad): managed orcad idles out after 15 minutes and is started again whenever it is down (#25121)

* feat(orcad): a managed orcad stops after 15 idle minutes and starts again on the next connect

A client-launched orcad now exits, like the relay, once no client, terminal,
working agent, staged migration or activation fence has been seen for 15
minutes. The exit is the normal graceful shutdown, which leaves the terminal
daemon running; the daemon retires only if it proves itself empty. A record in
the data root tells the next start (and its readiness health) that the stop
was an idle one rather than a crash.

On connect and after host resume, a fresh tunnel whose server does not answer
starts the activated slot under the activation fence, but only on a proven
exit, so a stopped server reads as not running rather than a failure.

* fix(orcad): keep orcad-entry under max-lines; idle e2e connects without a relay repo

* fix(orcad): deploy and rollback launches carry the managed idle-exit fence

The candidate launch in activation and the rollback launch built their
own launch spec without the activation root, so a freshly deployed orcad
never enabled idle exit; only the wake path did. The field is now required
on every launch spec, so the type system covers each launch site.

* feat(ssh): start a stopped managed orcad on connect, on restore and after resume

Once orcad stopped (idle, kill or host reboot), a connect still resolved
managed over a forward to a dead port and every call failed. Every connect
now checks the server behind its tunnel, as does a call through a restored
environment; a server proven stopped is started from its activated slot
under the activation fence, adopting a surviving daemon and its terminals.
The status line shows the start, and a start that fails keeps the host
managed with the reason and orcad.log's tail, never as a terminal verdict.

* fix(ssh): reuse a serving verdict only on the same SSH transport, for 5s

A reconnect right after a reboot was answered from the previous
transport's cached verdict, so the stopped server was never started.

* fix(ssh): key the serving verdict on the tunnel's remote port too

* test(ssh): a stopped server starts before the update counts its terminals

* feat(ssh): check serving at the bound port, and follow a restarted orcad to a new one

The serving check uses the port the tunnel forwards to (the one orcad bound).
A restart that binds a different port drops the forward and rebuilds it at
the new port, within the same ensure or on the explicit connect check. The
tunnel manager class moves to its own file to stay under max-lines.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a connect with no leases asks the host's relay endpoints before converting (#25223)

* fix(ssh): a connect with no leases and no relay session asks the host's relay endpoints before converting

* refactor(ssh): one isLiveSshPtyLease for every lease-liveness check

* fix(ssh): the host relay census asks a relay its PTYs before reading it as live work

An accepting relay whose holders or children the probe could not read (no lsof, an unrecognised
service child) read as live work, so a relay-era host that had exited its last shell never
converted. The relay's own bridge answers pty.listProcesses without the owner role.

* fix(ssh): the host relay census runs each relay's probe and bridge on the runtime it runs on

Pinned-ladder hosts often have no Node on PATH, so a PATH Node read every relay as unverifiable.
Each daemon's own argv names its pinned runtime, or the host Node a legacy relay started with.

* fix(ssh): an older relay's bridge runs on that relay's own runtime

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(orcad): build @parcel/watcher into the glibc 2.17 compat slot, so CentOS 7 runs managed orcad (#25199)

The compat target swapped in only node-pty, so it shipped the base target's upstream
watcher.node, which needs GLIBCXX_3.4.20; CentOS 7 stops at 3.4.19. orcad's preflight
refused it, and relay rung B lost file watching without saying so.

The compat slot now compiles @parcel/watcher from the package's own sources against the
pinned headers with the C++ runtime static, like node-pty. The slot gates (glibc 2.17
symbol floor, no shared libstdc++, N-API 8) and the smoke load cover it, the template
stages it into the compat target, and both orcad and relay rung B pick it up from there.

COMPAT_SLOT_ADDONS names a compat slot's own addons: a compat slot missing one fails
--require-slots, and a compat template target that would ship any native file without a
compat build fails the template build. The CentOS 7 cell expects an activated managed
server again, and every launched relay cell loads its watcher directly.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(phase3): close e2e routing holes, localize managed-server outcomes, trim harness and CI (#25206)

* fix(phase3): close e2e routing holes, localize managed-server outcomes, trim harness and CI

- Route the auto-convert spec into needs_build and route the new orcad/serve
  e2e helpers to the specs that use them, with routing tests.
- Run the auto-convert Docker lane only when routed; fold the missing-AppData
  check into the Windows mode-switch job and drop crash-hunt diagnostic env.
- Delta-move dialog keys its preview on the target id and cannot close or
  resubmit mid-move; Resume has an in-flight guard.
- Managed-server toasts and status lines show localized messages instead of
  raw codes or main-process English; add singular and zero-count wording.
- Settings style fixes (quiet Cancel, section header, labelled fields,
  progress labels); delete dead i18n keys and the unused previewConversion.
- Harness: kill the serve child on readiness timeout, share the isolated
  profile and spawn-until-ready helpers, hooks and a condition wait in the
  auto-convert spec, shared cross-version exec helper.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(phase3): localize the refusal toast, share serve readiness and liveness helpers, correct D7 docs

- The refusal toast no longer shows main's English blocker detail; the SSH
  Hosts status line keeps it under Details. A missing terminal count is no
  longer defaulted to 0.
- startOrcadServe uses spawnUntilReady, so a readiness timeout kills orcad;
  one isPidAlive replaces the spec's copy and the lock-holder loop.
- The port-6768 auto-convert test uses the shared hooks.
- The docs say the Windows D7 job checks rather than gates.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ci): give the idle-exit spec its app build and route the convert harness to it

Co-authored-by: m4air <m4air@Mac.localdomain>

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* ci(adhoc): pin every job of an adhoc build to the commit resolved at dispatch (#25311)

Each job read the requested branch name and checked out whatever it pointed at when that
job started. A push mid-run mixed commits: in run 37234548638 the glibc217 slot lane
built 2ddea8736e, then #25199 merged, and desktop_template checked out 70948d597f, whose
merge step requires the compat watcher that lane never built.

A first resolve-ref job (no secrets) resolves the branch, tag or full SHA once, and every
job checks out that commit. The mac job still vets it for reachability before signing;
the requested name stays the concurrency key and the release label.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(terminal): report a confirmed relay PTY exit as exited and give a cold Windows PTY probe more time (#25304)

A worktree terminal close returned as soon as the relay confirmed the stop, but the
exit frame reaches the runtime record only after the SSH output intake drains, so the
verdict read straight after saw a still-connected PTY and answered unverifiable. The
stop now waits, bounded by its deadline or 10 s, for that record.

The bundled runtime's PTY probe gives the first Windows spawn 15 s and retries once
after a timeout only; spawn errors and non-zero exits still fail at once.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(relay,orcad): keep new wire messages forward compatible, drop unshipped relay.reset, fix shutdown and serve gaps (#25298)

* fix(orcad-migration): keep migration wire forward compatible with newer peers

* docs(pty): record why pty.resumeClient negotiates by method-not-found

* fix(relay): remove the unused relay.reset method and keep a deferred shutdown serving

* refactor: drop dead hold API, release gate and migration pass-throughs; fix serve and delegation gaps

* fix(lint): keep reopen hooks within file size limits

* revert type dedupe in orcad-incumbent-recovery to avoid a parallel conflict

* fix(lint): name stop-reply request fields for their role

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* refactor(ssh): one managed-tunnel ownership check and forward bookkeeping (#25316)

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a stuck managed server can recover, cancels finish their change, unknown stays unknown; delete unused reset/drain code (#25227)

* refactor(ssh): delete the unused connection reset and drain machinery

* fix(ssh): surface a stuck managed server, finish activations past their first change, and fall back from Windows launch refusals

- A rejected build that changed profile state now refuses as orcad_recovery_changed_state
  (unverifiable) with a Recover path that restores the snapshot once the operator accepts;
  an interrupted activation is no longer reported as a quiet update deferral.
- Activation and rollback drop the abort signal after their first journaled change.
- Windows launch refusals and unsafe command lines send the host back to the relay.
- Fence refusals keep unverifiable blockers unverifiable; the POSIX liveness probe reads
  kill errors in the C locale and treats permission errors as unknown.
- A host with no relay fallback surfaces its real connect error; disconnect and removal
  clear the setting-up status, and a cancelled decision's progress is dropped.
- Remove tcp_forwarding_refused leftovers, dedupe incumbent stop/restore, exec-or-empty,
  census client, errorMessage and the blocking-blocker predicate; split activation and
  snapshot files under 300 lines.

* fix(ssh): one import per module in the rollback transition

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(migration): committed reads trust the receipt, conversions survive UI churn, journals are cached, and round-2 low items (#25327)

* fix(ssh): keep a newer build's fence and note fields, and validate a fence beside a legacy owner

Normalization now validates known fields but passes unknown ones through on orcadFence and the managed-server notes. A legacy managed owner next to a malformed fence falls back to the owner's environment id instead of keeping the malformed fence.

* fix(migration): delete a retired migration's scrollback files once retirement is durable

Retirement dropped the moved terminals' scrollback refs from session state but kept the files forever. After the journal records source-retired, the files the manifest names are deleted, except refs any session partition or pending export still names. A crash before the delete repeats it on the next retirement pass.

* fix(orcad): a committed migration reads as committed from its receipt, survives receipt eviction, and abandoned stages expire

- Committed-state reads trusted only a byte-identical copy of every moved row and snapshot, so a live server that had been used could never confirm its own commit. The receipt alone now proves it; full equality stays inside the commit.
- A receipt that ages out of the 64-entry list keeps a compact record, so the commit never reads as absent.
- A stage no client returns to within a week stops holding the server awake and is dropped at the next stage.

* fix(migration): freeze a host's session from the fence on, ignore UI churn in the frozen-source check, and cache parsed journals

- Renderer writes to a fenced host's ssh: partition are skipped from fence time, not only after commit, so tab work mid-conversion cannot change the source.
- The frozen-source check leaves out workspace session and client routing state, which the UI rewrites as the user works (including through the local partition and UI state); the server gets them as journaled.
- Parsed journals are cached per directory, keyed by each file's inode, size and mtime and dropped on every write and remove, so list and session calls stop re-parsing every manifest.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): an older relay a census cannot rule out stays held, and a reconnect resumes held PTYs through it (#25331)

- A census that does not know the host platform is unverifiable, never "no older relay". On Windows
  it now probes each older version directory's pipe for the target, the way relay GC does, so a
  live older Windows relay keeps its PTYs held instead of respawning their panes; those endpoints
  are held, never bridged.
- On reconnect, a PTY the current relay disowned while an older relay holds it is reattached through
  that relay's own bridge, with its runtime restored and its replay forwarded. When no route serves
  it, it is left for recovery like an exhausted reattach.
- A superseded or disposed deploy no longer starts the census, so it cannot replace the current
  attempt's.
- A route stops holding an unserved PTY that exits.
- Drop the unused censusPreviousRelays export.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(phase3): round-2 ui-infra review fixes (#25325)

- An unverifiable move refusal with no count says so instead of "0 terminals".
- One move per host: the dialog stays mounted while open, joins a run already
  in flight on remount, and the toast shares the same guard.
- A managed server's start, wake or update no longer reloads every host's
  catalog; only a new environment for the host loads, scoped to that host
  and the local catalog.
- A failed host-partition session write is no longer acknowledged as written,
  so the writer re-queues those fields.
- The deploy picker lists only hosts with no managed server or pending move.
- Harness: orcad and the released relay daemon are stopped when startup fails.
- Windows host CI: the convert cell always runs last and fails on a failed
  native switch or build; ssh-host-server unit tests no longer trigger it.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(relay,cli): never report an unreachable terminal as exited; keep PTYs through a deferred shutdown (#25328)

* fix(relay,cli): never read an unreachable relay or terminal as exited; keep PTYs on a deferred shutdown

* fix: stop an unrecorded breakaway child, route orcad serve through the spawn chokepoint, and drop review-flagged leftovers

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): reclaim a crashed desktop lock, close the stop-cancel race, drop dead daemon and RPC code (#25338)

* fix(orcad): reclaim a crashed desktop lock, close the stop-cancel race, drop dead daemon and RPC code

- A desktop reclaims a stale desktop lock record (Electron's lock proves it), and a real
  orcad hold shows a dialog instead of exiting silently.
- A failing quit handler no longer skips closing observability.
- Completion withdraws its request when a cancel lands mid-write; the listener removes a
  cancelled leftover.
- An idle stop's clean record is retracted when the shutdown fails or overruns.
- The managed-stop request tolerates unknown fields from a newer client.
- Shared error-code checks, one win32 coverage rule, no redundant isAlive filters.
- Remove the recovery-only daemon provider, requirePinnedWsPort/strictPort and the
  superseded orcad.migration.importCatalog RPC.

* fix(startup): keep main-process-preflight under the line cap

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): fail fast on a held fence, restore untouched rollbacks, atomic Windows host script, honest tunnel ensure (#25332)

* fix(ssh): fail fast on a held activation fence, restore a rollback the target never touched, stage the Windows host script atomically

- withOrcadActivationLock no longer waits up to 15 minutes inside the target lifecycle: a held
  fence answers at once (orcad_activation_recovery_required); deploy probes it before uploading.
- A rejected rollback target that left the restored snapshot untouched puts the newer build back
  unasked; only a real change waits for the operator. Crash recovery applies the same rule.
- The Windows host script is written only when missing, through a partial file and a rename; a
  bridge exit without a sentinel is unverifiable unless the shell could not find the command.
- The managed tunnel's ensure() throws when superseded, and a caller arriving after close()
  builds a fresh run instead of joining the doomed one.
- An unparseable stop answer after a stop was sent keeps the fence (new 'unconfirmed' outcome).
- stopRemote starts over instead of reporting live when another run settled the journal.
- Releasing an unreachable setup stops the orcad it activated and clears active on proven exit.
- A failed startup is torn down quietly so its error stays published; port-forward listeners
  keep an error handler; a linked SSH access connect is cancelled when its window closes.
- Shared errorMessage, one SFTP transfer helper, getConnectGeneration only, test-cell names.

* fix(ssh): stage the Windows host script through the pinned node.exe, which cmd.exe and PowerShell both run

* fix(ssh): host-script staging runs as host-script ops; a joiner of an overtaken tunnel run builds its own

- The presence check and the install are ops in the uploaded host script (script-present, and
  script-install run from the partial upload itself), invoked as node.exe <script> <op> <args>
  like every other Windows host op: no inline code.
- ensure() callers that joined an in-flight run no longer inherit its 'superseded' end; they
  start a fresh run, which also covers a close() in between.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a connect re-checks unverifiable relay terminals once its relay session can answer (#25385)

The server decision runs before any relay session exists, so a terminal this desktop left
detached could never be asked about and read unverifiable for as long as it ran. On Windows no
endpoint census can fill that gap. Once the session is up, the relay that holds the PTY answers
and a terminal it still runs is reported live.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(settings): a managed server whose status never loaded reads unknown, not "Not running" (#25389)

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(relay,runtime-env,cli): keep a deferred relay's AI Vault and skill uploads, make managed re-pair crash-safe, show unverifiable in worktree ps text (#25393)

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): keep a surviving daemon's slot through the cache prune, and drop an idle record a signal stop took over (#25387)

- The slot prune also protects any slot a live terminal daemon's PID record points into, and
  evicts nothing while a daemon record is unreadable.
- The shutdown trigger reports whether it took ownership; an idle stop that another source
  took over discards its clean-idle record.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): edit a managed host's connection, show rows no migration owns, drop the unused quit-drain predicate (#25403)

- SSH settings may edit a managed host's connection fields; its fence and generation are never written from the renderer, and its tunnel is closed so the next use redials. Removing it is refused with a pointer to Stop under Managed servers, which stops and removes the server; the renderer no longer ends the host's terminals before that refusal.
- A fenced host with no journal (an empty host's deploy, or one whose journal compacted after retirement) no longer hides its rows: no committed move owns them, so projects an older build added stay visible. An unreadable journal still hides them.
- The quit drain's mayDetach predicate and disconnectAll's filter had no production caller and are removed.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): input to a terminal an older relay holds but no route serves is refused, not dropped (#25404)

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): refresh a changed host's status once its delta move or keep-server choice lands (#25416)

The status line kept saying the host was changed on an older Orca until a manual reconnect. Main now publishes the host as managed by its server after a successful move or keep, with a disconnected state when the move released the relay session.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a stopped managed server stops showing on its SSH host without a reconnect (#25409)

* fix(ssh): a stopped managed server stops showing on its SSH host without a reconnect

On unlink, forget the host's managed-server decision and republish its connection state. The
republish also drops the host's cached worktree scans, so worktree.ps stops naming the removed
server.

* fix(runtime-environments): a removed server's workspace session partition goes with it

Host-scoped listings enumerate session partitions as known hosts, so worktree ps kept naming a
stopped or removed server as an omitted, unselectable host. Unlinking or removing a server now
drops its runtime:<id> partition.

* test: give the removal-storage fake store casts their SAFETY rationale

* fix(runtime-environments): drop orphaned runtime workspace sessions at startup

A crash between unlinking a server and dropping its session, or a build that unlinked before the
drop existed, left a runtime:<id> session listings name as an unselectable host. Startup now drops
runtime sessions whose server is not in the environment store; it never touches local or ssh
sessions, and skips entirely when that store is missing or unreadable.

* test(migration): retirement re-aims focus without creating the destination's session

Locks in the ordering the startup reconcile relies on: a runtime:<id> session is never written
before that server is registered.

* docs(runtime-environments): name the ordering invariant the startup reconcile relies on

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a connect whose relays prove its leases ended retires them, so the next connect converts (#25405)

Every connect decides before a relay session exists, so a detached or expired lease reads
unverifiable there. The post-session re-check proved such leases ended but left them in place, so
the host stayed unverifiable on every connect and each respawned pane added another lease.

Co-authored-by: m4air <m4air@Mac.localdomain>

* ci(ssh-windows): check out preload and renderer for the orcad-convert e2e build

Main's #25359 narrowed this workflow's checkout to the server and test trees. Phase 3's
orcad-convert cell builds the full e2e app with electron-vite, which also needs src/preload
and src/renderer, so the x64 inbox cell failed with UNRESOLVED_ENTRY.

* test(orcad): model a really converted host in the v1.4.218 downgrade wire test

#25403 shows a fenced host's rows when no migration journal explains the fence. The downgrade
test fenced the host without a journal, so this build showed the retained project it is meant
to hide. Write the destination-committed cutover journal a real conversion leaves.

* fix(ssh): a briefly held fence waits and is retried, not recorded as a failed update; terminate keeps held leases (#25420)

- The activation fence is retried for a few seconds. A fence still held answers
  orcad_activation_fence_busy (a waiting deferral, never an update failure) unless a journal or
  a lock past its stale age shows an interrupted run, which stays recovery-required.
- acquireInstallLock reports Busy only when a holder answered; a lock command that only ever
  failed surfaces its own error.
- Terminate detaches instead of disposing when any PTY was unverifiable, so the final teardown
  never bulk-marks a lease an older relay may hold as terminated.
- A wake whose connection dropped while holding the fence releases that fence on this client's
  next wake (no journal, slot proven exited), so a relaunch-then-connect is not left fenced.
- The orcad e2e reconnect helper surfaces a connect's error text instead of a JSON parse error.

Co-authored-by: m4air <m4air@Mac.localdomain>

* ci(ssh-windows): check out all of src for the orcad-convert cell

Its e2e global setup also compiles the bundled CLI from src/cli, which the narrowed checkout
left out.

* fix(orcad): idle stop — drop the record when a signal stop wins, and read the activation lock, not its root (#25464)

* fix(orcad): drop the idle-stop record when a signal stop finishes first

Every stop's cleanup now discards the record unless the idle trigger owns the stop, so a
takeover that exits before the idle request runs no longer leaves a false idle stop.

* fix(orcad): idle check reads the activation lock, not the transaction root

An interrupted acquire can leave the root without a lock; the client already treats that as
unfenced, and managed orcad now does too instead of never idling out.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): main refuses a managed server host's removal before ending its terminals (#25465)

The remove flow skipped ending terminals for a managed host only when the renderer's cached target list already showed the fence, so a fence that landed after the list loaded still lost the host's terminals before main refused the removal. The removal's terminate call now carries forRemoval, and main refuses it for a managed host with the Stop… message before touching anything; the renderer no longer decides.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(terminal): closing an SSH workspace counts a host-confirmed terminal exit as stopped (#25479)

* fix(terminal): a workspace close counts an SSH terminal's confirmed exit as stopped

The close's verdict treated any SSH terminal whose record was still present as
unconfirmed, even when that record held a host-confirmed exit. SSH records outlive
their exit, so every successful close of a relay terminal answered unverifiable.
Also, a relay reattach that finishes registering after the stop no longer revives an
incarnation whose exit is already recorded.

* test(terminal): cover a reattached relay exit confirmed without an incarnation

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(cli,ssh): managed-server actions outwait their own deadlines; cap the unverifiable serving detail (#25473)

orca environment update/rollback/recover/stop/cancel-stop waited the 60 s RPC default while the
desktop runs the whole action inline, so a slow host printed a timeout failure for an action that
kept going. They now wait 20 minutes, and a timeout says the action may still be running and points
at `orca environment status` instead of reporting failure. status keeps the default.

The retained managed-server state now clamps serving.detail to the same byte limit as its sibling
detail fields.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): keep the previous orcad.log on Windows across a restart (#25481)

The Windows breakaway launcher's addon recreates orcad.log on every start, while POSIX
appends, so a crash's log was gone once orcad restarted. The orcad launch now asks the
launcher to move the last run's log to orcad.log.1 first, capped to its last 1 MiB.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(sidebar): a stopped managed server leaves no empty project group behind (#25488)

The removed-runtime purge dropped the server's repos, setups and worktree rows but kept the project
groups and folder workspaces the renderer fetched from it, so an empty heading stayed in the
sidebar. It now drops those runtime-stamped rows too.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): run automations on orcad and keep a managed host up while they can fire (#25475)

* fix(orcad): run automations on orcad and keep a managed host up while they can fire

orcad never built an AutomationService, so with orca serve defaulting to orcad scheduled
runs never dispatched and Run now threw runtime_unavailable. The headless service setup
moves out of Electron startup into automations/runtime-automation-service.ts; orcad
installs, binds, starts and stops it, and managed idle exit counts an enabled schedule or
an unsettled run as busy.

* docs(orcad): list automations among the idle-exit conditions

* fix(automations): precheck reads the SSH manager from its registry, keeping electron out of orcad

The precheck imported getSshConnectionManager through ipc/ssh, which pulled 32 electron
modules and node:sqlite into the orcad bundle and failed build:orcad.

* test(orcad): stub the automation wiring in the push-startup runtime harness

That harness stubs OrcaRuntimeService without an automation surface, so starting the real
service threw setAutomationService is not a function and timed out the next test.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): Move to managed server survives a relay that hangs up on its last terminal exit (#25487)

BUG-15: with two live terminals, one served through an older build's relay, Move failed with
"Failed to terminate SSH host sessions: …: Multiplexer disposed" although both shells died. The
old relay reports the exit before the shutdown reply; that exit closes the route (its last served
PTY), and the disposed mux rejected the shutdown still awaiting its reply.

- The legacy relay route settles a shutdown whose PTY exit it already observed.
- Move no longer aborts on a failed stop: the terminal census (what the conversion trusts) decides.
  Exited closes the relay session and converts; live or unverifiable refuses and republishes the
  relay status, so the stale terminal count is replaced.
- The move dialog offers Try again after a refusal or failure.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): closing a terminal a previous Orca relay runs stops it there and confirms the exit (#25471)

A stop on a PTY an older relay runs reached the current relay whenever no route served it at that
moment (after a reconnect with the pane unmounted, or a shell no pane ever resumed), and the current
relay answers a stop for an id it never minted as done, so the PTY read stopped while its shell
kept running. A stop on a served PTY failed instead: its exit arrived before the stop's reply,
closing the route under the pending request.

- A served PTY stops on its route; a PTY no route serves is stopped through a short-lived route to
  the older relay that lists it, which hangs up once that PTY exits.
- When an older relay may hold the PTY but cannot be asked (incomplete census, Windows pipe, a
  bridge that will not open), or its bridge drops mid-stop, the stop is unverifiable, never reported
  done; terminate keeps the lease.
- A route stays open until its in-flight requests settle.
- The provider's exit stream includes the exits older relays report, so a stop observes the PTY it
  stopped exit on the relay that ran it.

The cross-version harness runs what terminal close --all runs per PTY against a real v1.4.218 relay,
for a pane resumed this connection and for a shell no pane resumed, and sees the old shell exit there
and the old relay retire on its own grace.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a live run's journal reads busy, a wake releases only its own fence, a cancelled re-check never reports connected (#25470)

* fix(ssh): a live run's journal reads busy, a wake releases only its own fence, a cancelled re-check never reports connected

- A held fence asks for Recover only when its lock is stale or a journal has no fence over it;
  a journal under a fresh fence is a run still working and reads orcad_activation_fence_busy.
- A wake writes an owner token into the fence it takes and later releases only a fence carrying
  that token; observing the fence gone forgets it.
- The connect re-checks ownership after the relay-terminal re-check, and the re-check itself
  neither records a decision nor retires leases for a cancelled attempt.

* fix(ssh): a tunnel caller that joined a run a disconnect cancelled builds its own

The launch-time restore's tunnel run connects over SSH; a disconnect then cancels that connect.
An explicit connect that had joined the run inherited its SshConnectAttemptCancelledError and
failed (seen as the idle-exit e2e's 'connect threw: ... cancelled'). A joiner now builds once
anew after any end of the joined run except an auth failure, which it shares rather than prompt again.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(runtime-env): a late reply from a replaced pairing never overwrites the re-paired device identity (#25490)

* fix(runtime-env): a reply from a replaced pairing never overwrites the re-paired device identity

* fix(types): narrow identity fields in markEnvironmentUsed

* fix(ssh): managed tunnel proves its server by runtime id; SSH access linking keeps the strict device check

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(migration): retained scrollback survives a closed tab, local acknowledgements stop blocking, close-intent retirement retries, mobile selections merge per workspace (#25508)

- A transfer reads a retained snapshot straight from storage once its tab closes, and releasing the retention deletes a ref no session names; the frozen-source check leaves the snapshot list to the journaled manifest.
- Only acknowledgements on panes the source host owns count toward the ui-routing blocker.
- Retiring close intents treats an absent source with an identical destination entry as already done.
- Importing a device's mobile selections keeps its selections for other workspaces.

Co-authored-by: m4air <m4air@Mac.localdomain>

* test(ssh): an expired lease an older relay still lists is never retired (#25449)

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(migration): keep each host's state its own, count only saved commits, and never wedge connect on a partial session (#25516)

* fix(migration): a host-qualified owner key belongs only to its own host

Owner matching stripped a key's host qualifier before matching its repo id, so converting host A claimed, moved and on retirement removed host B's session state when the two hosts share a repo id, and destination-qualified focus written by retarget read as leftover source state, failing retirement with orcad_migration_source_ui_routing_reappeared. A qualified key now matches only its own host, and an unqualified key in another host's session partition belongs to that host.

* fix(migration): a partially written session partition never wedges connect, and a marker two partitions agree on stops blocking the move

Real-host BUG-14: the renderer's per-host snapshot leaves out maps a host has no rows in, and main stored host partitions exactly as sent, so a runtime partition lacked tabsByWorktree and the dormant-state collector threw on every connect. Main now fills the required maps on every host-partition write, the migration collectors tolerate a partition persisted without them, and an unreadable session blocks the move instead of failing connect.

The same profile's workspace-session blocker was a false positive: the local and host partitions both carried defaultTerminalTabsApplied for the moved worktree with the same value, and the fragment merge refused any shared worktree key. It now refuses only when the partitions disagree.

* fix(migration): only a flushed commit acknowledgement moves a migration to committed

A committed state read may come from a receipt the server holds in memory but failed to flush. A retry took that read as proof, journaled destination-committed and went on to retire the source, so a later server restart could lose the catalog on both sides. A committed read in stage and in abort is now confirmed through the idempotent commit(), which flushes before it answers; a failure leaves the journal and the fence where they were.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): relay shells no lease here knows count as live, Move stops them, and Windows asks every relay pipe (#25518)

* fix(ssh): a relay shell no lease here knows counts as live, and terminate stops it

A CLI-created terminal has no lease, so with an attached lease the gate answered from leases alone,
a live decision was never re-counted once the relay could answer, and terminate stopped only the
shells it held leases or panes for. The relay's own listing is the authority on what runs.

* fix(ssh): a Windows connect asks every relay version's pipe for its PTYs before converting

Windows pipes cannot be listed, so the connect-time census answered 'unenumerable' and a shell no
lease here knew let the host convert under it. Each version directory's pipe for this target is
derived from its path, so the census probes them all, current included, and asks a live one
through its own bridge; a live pipe it cannot ask is unverifiable.

* fix(ssh): the terminal gate counts what earlier relays still run, leased or not

After an app update a shell a respawn superseded on its tab keeps running on the previous relay
with no live lease here, so a decision counted only the leased shells and Move could not see it.

* fix(ssh): a Windows relay folder with a live pipe it cannot ask stays unverifiable

Each pipe is probed and asked on its own, so one that answered with no PTYs can no longer stand in
for a live sibling the census could not reach.

* fix(ssh): terminate also stops shells only an earlier relay lists

provider.shutdown routes a held id to the older relay that runs it, so the terminate set now takes
listPreviousRelayPtyIds too; one it cannot reach is reported unverifiable as before.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a refused Move to managed server reconnects the host on its relay (#25543)

Stopping the terminals closes the relay session, so a move the census refused (another desktop's
terminals, or an older relay it can't rule out) left the host and its workspaces disconnected
until a manual Connect. The refusal now reconnects the host; the connect-time decision reads the
same census and keeps the relay. A reconnect that converts after all reports the move; a
reconnect that fails still reports the refusal.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(relay): an agent exec ends on its child's exit, not on pipes a background process still holds (#25544)

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): stop the automation scheduler first when a managed stop is dispatched (#25548)

A dispatch could otherwise race the daemon retirement census or write a run record that a
rollback restore then silently discards.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): count a degraded daemon's in-process terminals in the terminal census (#25545)

In degraded mode fresh terminals run on the local fallback inside orcad, but the census read
only daemon adapters, so an update or stop saw 0 live sessions and killed running agents.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): a retained source whose drafts or settings an older build changed is never hidden as unchanged or retired (#25549)

* fix(orcad): a retained source whose drafts or settings an older build changed is never hidden as unchanged or retired

The retained-source fingerprint covered only repo, folder and group identity, so an unsaved draft
edited on an older build read unchanged: the host kept serving the server's older draft and
retirement deleted the newer one. Retention now also records a versioned fingerprint of the source's
drafts and user-authored names and settings; a mismatch marks the host changed, and a journal
without one is never retired automatically.

* fix(orcad): a retained source's automations are part of its state fingerprint

An older build can edit an automation the source keeps; retirement would delete that edit as if
the server held it.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a rollback restores the older snapshot only behind orcad's terminal barrier (#25551)

The rollback's census is taken while orcad still admits work, so a terminal or automation that
starts before the stop had its state wiped while its PTY survived. The incumbent is now stopped
through its managed stop with idle-daemon retirement: orcad closes terminal admission on every
daemon generation, counts live sessions under that fence, and retires the daemon only when none
exist. Only 'retired' lets the older snapshot replace state; live, unverifiable or a missing
answer refuses and relaunches the incumbent on its untouched state. A build without managed stop
is refused before anything changes.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): failed managed setup leaves 'connecting', edited managed host redials, idled-out server starts before its census (#25474)

* fix(ssh): a failed managed setup leaves 'connecting', an edited managed host redials, an idled-out server starts before its census

- doConnect publishes the error and clears the 'setting up' status when the managed-server
  decision throws for a still-current attempt; a cancelled one still reports cancellation.
- Editing a managed host's connection fields closes its tunnel and disconnects its transport,
  serialized with the target's lifecycle, so the next use dials the edited target.
- The terminal census starts a server that idled out behind a forward still up, so Stop, Update,
  Rollback and status no longer refuse with 'census unavailable' on every retry.

* fix(ssh): a fenced failed setup publishes its cause, never-launched slots are collectable, a reused PID is not orcad

- doConnect publishes the relay decision's setup failure (and clears 'setting up') when a failed
  managed setup kept the host fenced, instead of throwing a bare 'serves a managed server'.
- The liveness probe answers NEVER_LAUNCHED for a slot with no process record and no readiness
  file; GC removes such a slot, and every other reader still reads it as UNKNOWN.
- On POSIX a PID whose command line does not run the slot's orcad.js reads DEAD, so a stopped
  orcad behind a reused PID is woken instead of reported serving or unverifiable.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a legacy relay route a new pane started serving stays open when a pending stop settles (#25542)

A served PTY's exit that lands before its stop's reply defers the route's hang-up until the stop
settles. A second pane the same older relay holds could start serving through the route in that
window, and the deferred hang-up then closed it anyway, sending that pane's input and stops to the
current relay. Serving a pane now cancels the deferred hang-up, and a settling request hangs up only
a route that serves nothing.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(migration): the journal holds a migration's scrollback until commit or abort, across failed uploads and restarts (#25550)

Retention was scoped to one transfer call, and its release in finally deleted a closed tab's
snapshot after an interrupted upload, so every retry failed with source_snapshot_changed.
Inline buffers had no file for the retained read at all. Retention now follows the cutover
journal: held from the journaled export through staging, rebuilt at startup, released on
commit or a removed journal. Inline bytes are written to their ref while held.

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude <noreply@anthropic.com>

* fix(migration): a repo id two hosts share never lets legacy keys cross hosts, and a dangling identity alias stops blocking the move (#25558)

- Repo ids are not unique across hosts: the same id may be registered on two SSH hosts. Stores that are not session partitions (worktree metadata, automations, lineage, client state, sparse presets, retired names) can hold legacy keys with no host qualifier, so an id both hosts register said nothing about whose a key was. The scope now records such shared ids; an unqualified key or bare repo id for one only matches with its row's own host evidence (worktree metadata's hostId, an automation's ssh target), so another host's rows are never moved, counted or retired. An automation's target generation now matches only alongside its target id.
- An identity alias whose identities hold no metadata (worktreeMetaByIdentity lost them, as on the B4 profile) is nothing to move rather than a worktree-metadata blocker; retirement drops it.

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(cli): orca environment recover --accept-changed-state --yes restores over changed state (#25597)

Recover refused when a rejected build changed profile state, and its refusal told the user to run
Recover, which the CLI could not do. --accept-changed-state (confirmed with --yes) maps to the same
acceptChangedState the Managed servers settings pass, and the refusal now names the flags.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): refuse an update that would end a degraded host's in-process terminals (#25598)

The census now reports inProcessSessions separately. Those terminals run inside orcad and
end with any restart, so planOrcadUpdate defers with a non-forceable
orcad_update_ends_in_process_terminals instead of claiming they survive on the daemon.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(migration): a delta move protects what it imported from rolling back across it (#25599)

A delta move never advanced the server's migration mark, so rolling back an update taken before the delta was admitted and dropped the delta's projects. The delta now records the mark before any commit can land, resumed commits included; the mark keeps the latest migration and never moves back; and the rollback gate also counts every journal into the server, so deltas finished before this change stay protected.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a missing relay inventory never proves terminals exited, and the Windows census covers every desktop's relays (#25611)

- The migration terminal gate returned `exited` whenever this desktop held no unresolved lease, even
  when the current relay or the earlier relays could not be asked. A failed or incomplete inventory
  is now `unverifiable` regardless of local leases. With no relay session at all the gate asks for a
  host census (`needsHostCensus`) instead of reading the silence as exit; the connect, conversion and
  delta move pass that census in, and the connect hands its own census result to the conversion it
  starts. A census that cannot list endpoints (`unenumerable`) is unverifiable too.
- The Windows connect-time census derived pipe names from this desktop's target id only, so another
  desktop's relay on the same account was never probed. It now lists every `orca-relay-*` pipe on
  the machine and maps each to the relay instance that owns it through that instance's credential
  file or active-pipe marker, asking each with its own credential. A pipe no version directory
  accounts for is unverifiable unless the host proves it another account's (or gone), and an
  inventory that could not be read is unverifiable.

Co-authored-by: m4air <m4air@Mac.localdomain>

* refactor: drop the unshipped pty.resumeClient relay method and unused SSH provider unregister guards (#25595)

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a connect still deciding its server holds the raw 'connected', closes a transport its cancelled decision opened, and an edit keeps a relay host's session (#25641)

- handleSshConnectionStateChange holds a raw 'connected' while a connect is in flight even before
  any relay session exists (published as 'connecting'), so the census, deploy or conversion that
  dials the pool no longer reports the host up with no providers.
- priorConnection is captured before the server decision; a connect cancelled after the decision
  closes a transport the decision opened, unless a newer connect is using it.
- Editing a fenced host an older build changed (it runs on the relay directly) no longer
  disconnects its transport; only a host reached through its managed server redials.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): a retained source is unchanged only against a pre-commit baseline of everything a user wrote (#25602)

The state fingerprint read drafts, automations and workspace metadata from what a move could
carry, so a session a move refuses hid an older build's draft edit, and retention hashed the source
after the commit, so a crash before retention blessed whatever an older build changed in between.
The baseline is now written with the fence, before any commit is possible, from the source read
directly; a session that cannot be read, or a journal without that baseline, is unverified.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): a delta move refuses a source that changed while it checked terminals (#25691)

The plan and manifest are taken before the terminal check and session release are awaited, but the
journal took its baseline after them, so a draft typed in between became the baseline while the
server received the older one, and retirement deleted the newer draft. The baseline now comes from
the plan's own snapshot, and a source that changed since it refuses the move.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(migration): a save landing mid-retirement no longer defers retirement (#25692)

Retirement removed the source rows, then awaited the profile flush, then checked nothing came back. A session save that landed during that flush re-added a source-owned row, the check failed, and the journal stayed committed until a later connect. Retirement is idempotent, so it now runs one more pass before deferring; a row back after that is reported with the partition and owner key it reappeared under.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(automations): a headless run never reads completed when its agent never ran (#25700)

* fix(automations): a headless run never reads completed when its agent never ran

orcad (and Electron serve) finished a dispatched run on a satisfied
tui-idle wait, and a ready shell prompt satisfies it: a run whose agent is
not installed read 'completed' within seconds. Like the desktop runner,
completion now needs the agent's own status for the run's pane after
dispatch; without it the run fails after the agent-start window with the
reason, instead of claiming completion.

* fix(automations): keep idle-means-done for agents without status; fail only a refused command

Not every automation agent reports status on orcad (no hooks on the host, no
recognised title), so requiring it would fail their runs. A run completes on
the agent's own status, fails when the shell refused the agent's command
(bash, zsh, dash, fish, PowerShell, cmd), and otherwise keeps the old
idle-means-done rule after the agent-start window.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(migration): the retained-source fingerprint sees all worktree metadata retirement owns (#25694)

* fix(migration): the retained-source fingerprint sees all worktree metadata retirement owns

The state view read only legacy worktreeMeta keys and attributed them without meta.hostId, so
identity-backed metadata and unqualified rows a shared repository id leaves to hostId were missing
from the fingerprint: an older build's edit read as unchanged and retirement deleted it. The view
now uses the same attribution as export and retirement, and metadata that claims the source host
but cannot be attributed leaves the source unverified. The fingerprint version moves to v2.

* fix(migration): the retained-source fingerprint skips automations on a repo id another host shares

The state view matched automations by scope.repoIds, which ignores sharedRepoIds, so host A's
fingerprint included host B's automation on a shared repository id; editing it marked A changed and
routed it back to the relay. The view now uses the move's automationTouchesScope.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): a failed readiness read no longer kills a healthy candidate; log unsettled activations (BUG-17a) (#25701)

* fix(orcad): a failed readiness read no longer fails a candidate's launch, and log every unsettled activation

BUG-17: the candidate went ready and the client SIGTERMed it ~1 s later through its reject path,
yet that readiness passes the gate, so the launch itself failed: one readiness-wait exec that
errored failed the launch outright. Retry such reads until the readiness deadline; an
unconfirmed termination still fails at once. The update's outcome never reached the app log,
so every update or rollback that does not go through now logs its code and reason.

* fix(orcad): the host-side readiness wait ends on a wall-clock deadline

On a loaded host each poll's reads outlasted its sleep, so the step-counted loop ran past
the client's 30 s exec timeout. That timeout failed the launch, the reject path SIGTERMed a
candidate still starting, and orcad, which defers a stop until startup completes, published
readiness and then exited (BUG-17).

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(migration): stop a converted host's stale tabs landing in local, and retry retirement only on an exact replay (#25712)

* fix(migration): retirement re-removes only an exact replay of moved session state

#25692's second pass re-ran retirement on any row that reappeared during the flush, which could delete a tab or draft written after the move. Retirement now records the session rows it removes before it runs; a row that reappears is removed again only when it is byte-identical to one of those (in any partition). A new or changed row defers retirement and stays.

* fix(ssh): a converted host's leftover session rows stay in its own partition, never local

After conversion the renderer drops the SSH host's projects and worktrees but keeps their session rows. With no catalog owner left, the next save routed those rows to the local partition, where retirement read them as moved source state reappearing and deferred. Converting now pins each dropped worktree's session key to the host's partition; main's fence guard keeps that partition frozen, so the stale rows are never written.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(test): restore the codexProviderHandle import main's #25078 dropped again

#25722 restored it, then #25078 removed it, so pnpm tc fails on main's tip.

* chore(sync): keep main's own cloud and mobile files byte-identical to main

Earlier syncs added lint-only brace and template fixes to these main-owned files; reverting
them keeps #24863's diff against main free of files Phase 3 does not own.

* fix(terminal): a remote pane's reattach error no longer shows this client's OS and shell (#25693)

* fix(terminal): a remote pane's reattach error no longer shows this client's OS and shell

The error toast appended the client's environment for every non-SSH error. It now shows it only
when the pane's known execution host is this client; an SSH, managed or not-yet-known host omits it.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(terminal): type the pane-host fixture as runtime owner state

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude <noreply@anthropic.com>

* fix(orcad): a retained activation fence is ownerless, so recovery can always take it over (BUG-17) (#25698)

* fix(orcad): a retained activation fence is ownerless, so recovery can always take it over

BUG-17: a recovery that took a stale fence over, failed and retained it left a fresh lock, so
every later recovery read it as still fresh and the host could never be recovered. A run that
keeps the fence once it is done now backdates the lock; one whose remote command may still be
running keeps it fresh.

Also run the in-process terminal deferral before the forced protocol check, so a degraded host
whose daemon is empty names its in-process terminals instead of an unreported protocol.

* test(orcad): the CLI's accepting recover takes over a fence a refused recover just retained

The fake host now answers a stale-only takeover busy while a recovery's own takeover is
fresh, which reproduces BUG-17's 'still fresh' loop without the ownerless mark.

* test(orcad): keep the fake host's fence acquisition void where callers expect it

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a cancelled connect closes only the transport its own decision opened and nothing newer adopted (#25696)

* fix(ssh): a cancelled connect closes only the transport its own decision opened and nothing newer adopted

#25641's cleanup took any transport that differed from the pre-decision one as the cancelled
decision's, guarded only by connectInFlight. A replacement connect that completed (and left
connectInFlight) then had its live transport disconnected by the stale attempt.

The pool now attributes a transport it opens inside a connect's server decision to that attempt
(AsyncLocalStorage), and the latest user adopts it: a connect that connects or publishes a managed
route, or a managed tunnel that records a forward. A cancelled attempt closes the transport only
when it opened it, nothing newer adopted it, and no current replacement is in flight.

* fix(ssh): a still-current connect whose server decision fails closes the transport that decision dialed

The decision's own failure (a throw, or a fenced relay refusal) published 'error' but left the
transport it dialed open, so getPublicSshState read 'connected' and a later auto-reconnect
broadcast a plain 'connected' with no relay. Both branches now close exactly the decision-owned
transport through abandonDecisionTransport, which treats the attempt's own in-flight entry as
no newer owner while that attempt is still current.

* test(ssh): name the stand-in transport type

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): Windows readiness identifies itself by one PID when the process-table snapshot times out (#25723)

* fix(orcad): Windows readiness identifies itself by one PID when the process-table snapshot times out

A loaded Windows runner timed the whole-table snapshot out during the bundled runtime's
readiness preflight, so the candidate failed to start and activation rejected it.

* fix(orcad): fall back to the one-PID query only for a slow process table, never an unreadable one

An unreadable table (EDR-hooked snapshot, restricted token) must still fail qualification.
The table now rejects slowness with a typed WindowsProcessTableTimeoutError, and only that
falls back. Review by win-serve.

* build(cli): list the process-table timeout error in the CLI project's file list

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a connected conversion proves the host idle with the account-wide census, not this target's lists (#25697)

* fix(ssh): a connected conversion proves the host idle with the account-wide census, not this target's lists

The migration terminal gate asked the host-wide census only when no relay session existed. With a
session, it trusted this target's relay listing and its earlier-relay census, both of which name
only this target's instances, and returned `exited` when they were empty, so another desktop's live
shell on the same account, under a different target id, let the host convert under it.

`exited` now always needs a complete host-wide census: this target's lists can prove `live`, but
empty lists only pass the question to the census, and a gate given none answers `unverifiable`
with `needsHostCensus`. The connect-time refinement passes the census too, so a connect retires its
leases only when no relay on the account holds work. The Windows host lane now runs a second
desktop's relay with a live shell and expects the connected gate to read the host live.

* test(ssh): a connected relay's empty lists still ask the account-wide census

* refactor(ssh): drop the gate's unread needsHostCensus flag; its unverifiable reason says why

* test(ssh): the delta snapshot fixtures give their account-wide census

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a reconnected transport starts managed orcad itself instead of inheriting a dropped start (#25689)

* fix(ssh): a reconnected transport runs its own orcad start instead of inheriting a dropped one

The serving check deduplicated in-flight starts by environment only. When a
connect dropped mid-start (as when the launch-time auto-connect is replaced
by a reconnect), the caller on the new transport joined the start bound to
the dead one, got its failure, and reported the host managed with no server
running. In-flight checks now join only on the same connection, connect
generation and port, and a wake's own fence token is cleared only by that
wake.

* fix(ssh): a reconnected wake releases the fence its dropped wake held at any point

The flake's real verdict was 'fenced': the dropped launch-time wake held the
activation fence, and the reconnected wake could not prove it its own. A
wake now claims its token before its first remote step, an absent owner
record under a held token is still its own, and a reconnected wake waits for
this client's dropped wake to settle before reading the fence.

* fix(ssh): release only a fence carrying this process's own wake token

A fence with no owner file could be another client's fresh one. Releasing it
now requires the owner token this process wrote; the token is claimed before
the write so a drop after it still proves ownership.

* fix(ssh): type the wake's fenced fallback

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* test(ssh): stop the exec-stdin test double from failing on EPIPE (#25739)

The truncation test's fake exec channel forwarded the local shell's
EPIPE (or 'Cannot call end after a stream was destroyed') as a channel
error. Whether that error or the shell's exit code won depended on
scheduling, so the test failed under full-suite load. ssh2 silently
drops writes once the remote stops reading; the double now does the
same, and a 1 MB payload makes the early-stop path deterministic.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): Move lets the reconnect's decision run its census, and a census outside a connect holds 'connected' and closes its own transport (#25735)

- Move to managed server no longer runs a separate host census after tearing the relay down,
  which dialed the pool with no connect in flight and broadcast a raw 'connected' with no
  session or providers. It reconnects, and reads the decision the reconnect's census recorded.
  A relay a failed stop left up is detached (leases kept), not disposed, before the reconnect.
- The CLI and delta-move census (censusHostRelayTerminalsFor) runs outside a connect under its
  own owner: the raw 'connected' it causes is held, and a transport it opened that nothing
  adopted is closed afterwards. Reusing a pooled transport inside a scope now adopts it.
- Drop the now-unused publishRelayTerminalsStatus.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(test): drop the restored codexProviderHandle import now that main restored it

* refactor(migration): keep a converted host's source rows instead of retiring them automatically (#25768)

Automatic source retirement leaves Phase 3: nothing deletes a converted host's retained rows on connect, delta move, keep-server's-version or restart. They stay hidden and are removed only by stopping the server, removing the host or uninstalling. Change detection goes back to the catalog-identity fingerprint, so an older build's edits inside an already-moved project stay preserved in the retained rows without marking the host changed. The copy-only helpers the delta view uses are renamed to subtract, and the converted-host session pin now also overrides a boot primary of local.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(automations): a headless run is watched past tui-idle timeouts by the one run observer (#25733)

* fix(automations): a headless run is watched past tui-idle timeouts by the one run observer

The headless dispatcher awaited a single tui-idle wait, which rejects after
its 5-minute default, so a healthy agent working longer was published as
dispatch_failed and never observed again. The dispatcher now hands the run
to its completion watcher, whose runtime observer already re-arms wait
timeouts, honours cancellation and bounds total observation; the agent
status and missing-command checks fold into that observer, and the separate
completion loop is gone.

* fix(automations): an already-idle pane completes when the start window passes

Real-host: a stub that exited before the window left an idle shell with no
agent status, and the observer re-armed a tui-idle wait that never resolves
for an already-idle shell, so it timed out instead of completing. The
observer now keeps judging the pane while its output is unchanged, and only
waits again once the pane changes.

* fix(automations): resolve a watched headless run by its launch handle first

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(runtime-env): a re-paired managed server's subscribers recover without a reload (#25752)

* fix(runtime-env): a re-paired managed server's subscribers recover without a reload

The renderer kept the pairing revision it last read, so after an on-connect update re-paired a
managed server every subscribe and request was refused as 'pairing changed' until a reload. The
first refusal now re-reads the environment catalog, so revision-keyed subscriptions resubscribe
and requests carry the new pairing. A managed server re-pairing for the same host registration is
the same peer, so its workspaces and tabs are no longer purged as a replaced environment.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(runtime-env): a re-paired managed server is the same machine only when its host proves the same identity

Same SSH target registration is not proof: a reinstalled host or a target now pointing elsewhere
keeps it. The runtime id the pairing handshake verifies must be known and unchanged; otherwise the
re-pair retires the environment as before. A proven runtime id change also counts as replaced.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(runtime-env): decide a managed re-pair's same machine by the host's proven key, not its runtime id

The runtime id is minted per process start, so every orcad restart would read as a new host.
The host's E2EE public key persists in its own profile across updates and its pairing handshake
proves it; main now lists a digest of it, and the renderer keeps a re-paired managed server only
when that digest is known and unchanged under the same SSH target registration.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(runtime-env): defer a managed re-pair's same-machine decision until the host key is known

A re-read that lands before the new pairing's host key is listed no longer purges: the decision
waits for a catalog that carries the key and retires only if it differs. Adds the update-flow
store test: same registration and key keeps workspaces and tabs, including a re-read that runs
before the key is known.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(runtime-env): watch for a pairing refusal on a side branch so requests settle on the same tick

Chaining .catch onto every subscribe and request delayed each success by a microtask, which let a
StrictMode cleanup run before a client-event subscription resolved, so its unsubscribe landed late.

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude <noreply@anthropic.com>

* refactor: consolidate pane-ownership, migration-catalog, activation-launch and parser helpers (#25738)

* refactor(orcad): one launch-and-judge helper for activation and rollback

* refactor(orcad-migration): one copy each of the destination projections, selectNewRows, assertSameValue, compareKeys and slotLiveness

* refactor(orcad-migration): one string-list validator and one uniqueness check, error codes passed in

* refactor: one shared pane-ownership and terminal-layout module for migration, profile transfer and split layout

* fix(orcad-migration): row-identity helpers in a leaf module (no import cycle); key order in its own module

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): reopen a managed tunnel to a host another desktop restarted (#25800)

The tunnel's identity check pinned the saved runtime id, which orcad mints per process. A host
updated or woken by another desktop, or restarted while this one was away, failed every reconnect
with orcad_identity_mismatch, and nothing could refresh the id because that needs the tunnel. The
E2EE handshake with the pinned host key and our accepted token now prove the server; the first
authenticated status reply records the new id. A different host is still refused.

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude <noreply@anthropic.com>

* fix(orcad): create the readiness file owner-only so its pairing token is not world-readable (#25809)

orcadLaunchCommand truncated .orcad-readiness before setting umask 077, so under a
login umask of 022 the file that receives the pairing offer (with a runtime-scope
device token) came out 0644. umask 077 now runs first, the readiness file is
chmod 600 after the truncate (a redirect keeps an earlier build's 0644), the pid
and log files are tightened too, and the slot dir and ~/.orca-remote are chmod
700 so files earlier builds left readable are no longer reachable. The state
snapshot capture also sets its umask before creating the snapshot directory.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(terminal): explain a pane whose saved session another host connection owns (#25814)

terminal_pane_owner_host_mismatch reached the user raw, with an issue link. It now reads as a
plain explanation with the open-a-new-terminal action, like the reattach failure.

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude <noreply@anthropic.com>

* test(topology): allow phase3's headless editor-tab retirement in main's boundary ratchet (#25823)

Main's #25329 added the ratchet; phase3's mobile-session-editor-projection.ts writes the host's
own session through setWorkspaceSessionForWorktree, the same way the listed headless
mobile-session tab writers do.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(automations): a headless host closes finished run terminals, keeping the newest few (#25831)

* fix(automations): a headless host closes finished run terminals, keeping the newest few

The desktop closes a run's terminal when the run completes; orcad had no
renderer to do it, so hourly automations left a shell and PTY per run open
forever (28 after ~6h on a real host). The headless service now closes a
finished run's terminal after a 10-minute grace, keeps the newest three per
automation viewable, and never touches a run that has not finished.

* fix(automations): never close a run terminal a client typed into or is viewing

Mirrors the desktop's take-over rule on headless hosts: a finished run's
terminal stays open when any client drove input to it since spawn, is
attached to or viewing it, or when this process cannot tell (it adopted
the PTY rather than spawned it).

* test(runtime): register a viewer through the public subscribe API

* fix(automations): close only completed runs' own panes

A failed run can still hold a live agent (blocked on a prompt, past the
watch window, or after an observer error), so like the desktop only a
completed run's terminal is closed. And only the run's own pane closes, so
a pane a user split into the same tab survives.

* fix(automations): close a run pane only while it still holds the run's PTY

The use check read run.terminalPtyId, but the close hit whatever PTY now
occupies the run's pane. Restart-exited-pane and the Codex account-switch
restart put a new PTY there, so a terminal a user was using could be killed.
The close now resolves the pane's current PTY and closes only when it is the
run's own; otherwise it closes nothing and only clears the run's terminal.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): a snapshot restore past the exec timeout keeps its fence instead of racing a second restore (#25811)

* fix(orcad): a snapshot restore past the exec timeout keeps its fence instead of racing a second restore

A capture, restore or clear ran under the generic 30s exec timeout. On ssh2 the timeout closes the
channel and reads as a confirmed failure, but sshd leaves a pty-less command running, so rollback
ran its rescue restore and recover orphaned the fence for a second one, both in the same stage.

- State mutations run through execOrcadStateMutation: no abort, a client wait past the host's
  deadline, and any closed channel or busy/deadline answer is unconfirmed, so the fence stays fresh.
- POSIX hosts wrap each one in `timeout -s KILL` (where present) and a pid-checked lock dir under
  ~/.orca-remote; the Windows host script takes the same lock.

* fix(orcad): a running state mutation keeps the activation fence fresh

The fence goes stale by its lock dir's mtime after 20 minutes, and a capture, restore or clear can
now run up to 15 under it, so a rollback's rescue capture plus restore could outlast the window and
let a recovery steal the fence from a live run. While a mutation runs, the host now touches the
fence every 60s (POSIX: a background beat that stops with its shell; Windows: an interval in the
host script, whose mutations are now async so the timer runs). A dead process stops refreshing, so
stale takeover still recovers it.

* fix(orcad): the state-mutation fence heartbeat never refreshes a wake's fence

A wake writes .orca-wake-owner into the fence dir and lets its fence age toward takeover; the
heartbeat now skips a fence that holds that token, and only ever changes the dir's mtime.

* fix(orcad): a state mutation releases its host lock before answering, and names its holder by pid and start time

On Windows answer() exits in the stdout write callback, so an op that answered before its first
await (MISSING, EMPTY, FAILED) exited before the wrapper's finally and leaked the lock; a reused
pid then read as alive and every later capture, restore and clear answered busy. Ops now return
their token and the wrapper answers after releasing the lock. The holder is pid plus creation time
(the slot's process-tree addon); one that cannot be identified is stale once its lock misses five
heartbeats. POSIX gets the same heartbeat-age check for a reused pid.

* fix(orcad): a state mutation's host lock is owned by its whole process group

The lock named only the shell's pid, so a shell killed while its rm or tar ran let the next
mutation take the lock and race that child. Each mutation now runs in its own process group
(setsid, or perl setpgrp on macOS), with timeout inside it so a deadline KILL reaches the children
too. The lock records the group, and is taken over only once no member is alive; a host that can
start no group records none, and its lock is never taken over. Windows ops run in-process, with no
children to outlive the holder.

* fix(orcad): record a state mutation's process group without ps -p, and never hold a groupless lock forever

BusyBox ps has no -p, so Alpine hosts recorded no group and their lock read busy forever after a
timeout kill, reboot or OOM. The group now comes from /proc/<pid>/stat (read after the comm field's
last paren), with ps -o pgid= -p as the fallback. A lock that still names no group is taken over
once its pid is dead and its heartbeat has missed three beats.

* fix(orcad): only proof of exit frees a state-mutation lock

A Windows holder whose creation time could not be read was taken over after five quiet minutes
though its pid was alive, so a suspended clear could resume and delete freshly restored profiles.
Both platforms now free the lock only on proof of exit: a dead pid, a different creation time, or
(POSIX) a group with no live member. A live holder of unknown identity stays busy until it exits.
The owner record is written exclusively, so a run that resumes after a takeover backs off.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): never offer Move for terminals another Orca desktop or session runs (#25815)

On a host where another desktop held a live relay shell, this desktop read relay_terminals_live
with offerMove, and its copy ("Its N open terminals will restart") implied they were its own. The
census already attributes them: terminals counted only by the host-wide census, with no lease or
listing of this target naming one, run under another target or session. That verdict now carries
elsewhere / terminalsElsewhere; no move is offered (no toast, no status-line action) and the status
line says the terminals belong to another Orca desktop or session.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(automations): an update releases finished, unused automation shells before counting terminals (#25844)

Hosts with schedules kept completed run shells (the newest three, and any not
yet past their grace), which counted as running terminals and deferred every
on-connect update with orcad_update_terminals_running. The update and
rollback census now ask the server to close completed automation run
terminals no client used, with no grace or keep rule, and count after the
daemon drops them. Used, unknown, failed and running ones still count.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): every activation fence holder carries a generation token its steps and release must match (#25834)

* fix(orcad): clear a bare stale activation fence instead of asking for Recover, and report a restarting update from status

BUG-21: a wake cut short leaves a stale fence with no journal. Every update then answered
'Recover it first' while Recover answered 'none'. The fence-hold check now takes such a fence
over and drops it, and the update retries once; Recover is asked for only over a journal.
The CLI also treats a connection closed by the server's own restart during update or rollback
as expected and reports what status shows once the runtime answers.

* fix(cli): type the reconnect status response explicitly

* fix(ssh): a wake's fence carries its owner token from the moment the lock exists

The idle-exit e2e still read 'fenced' on reconnect: the launch-time wake's
lock landed on the host but its connection dropped before the client saw OK,
so the wake body never ran and never wrote its owner token, leaving a fence
nothing could prove. The token is now claimed before the lock and written by
the same command that creates it, and a wake registers itself before any
remote step so a reconnected wake waits for it instead of racing it.

* fix(orcad): every activation fence holder carries a generation token its steps and release must still match

Astra pass 8: a holder suspended past the stale window resumed, kept acting, and its
unconditional release deleted the successor's fence and recovery journal mid-update. Every
holder (activation, rollback, stop, recover, wake) now writes a token into the lock it creates or
takes over. Each remote step it issues checks that token on the host, in the same command on
POSIX and inside the host script for Windows host ops; release is conditional on the token and
moves the lock aside instead of removing the root. A superseded holder aborts with
OrcadFenceLostError and its release is a no-op.

* fix(orcad): state mutations check the fence token before their lock, and refresh only a fence they still own

On POSIX the fence guard runs outermost in serializedStateMutationCommand, before the mutation
lock and the work, and the heartbeat touches the fence only while the token is still this run's.
The Windows host script records the --fence token and refreshFence compares it. A fence-lost
answer from a state mutation is a refusal, never a FAILED fallback. One owner-file constant
replaces the wake-owner copies.

* refactor(orcad): a state mutation's heartbeat touches the fence directory it checked ownership of (review)

* fix(orcad): a release moves the journal and lock aside and keeps only its own generation's

Astra pass 9: a release that passed its token check and stalled before deleting could, once a
takeover and a successor came and went, delete the successor's journal and lock. The journal is
now stamped with the writing run's fence token (a recovery takeover re-stamps the journal it
adopts), and the release renames the journal and the lock aside, deletes each only if it carries
this run's token, and otherwise moves it straight back.

* test(orcad): a successor restore stays busy beside a paused clear on the Windows host script

* fix(orcad): classify a lost fence from the step's exit and stdout, never the error message

The real exec error quotes the command, and every fenced command carries the
guard's marker text, so any failed or timed-out fenced step read as a lost
fence and dropped its unconfirmed flag. execCommand now attaches exitCode and
stdout to its exit error; execOrcadRemote rethrows unconfirmed terminations
before any reclassification.

* fix(orcad): a wake keeps its fence token until the fence is released

A disconnect fails a wake's next step without the unconfirmed flag, and the
fence release then fails over the dead connection. Forgetting the token on
that error left a fence the reconnected wake could not prove its own, so it
reported the host as held by an update.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* test(recovery): keep Phase 3's recovery-lifetime test on main's legacy-worker ports

Main added a required hasRequestedReleases port and now skips persist when a pass resolves
nothing, so the test mocks the new port and holds the pass at workspace resolution instead.

* fix(ssh): say "1 terminal" when another Orca desktop runs one on the host (#25853)

The terminalsElsewhere status line had no plural forms, so B9 read "while 1 terminals another
Orca desktop … are running". It gains _one/_other entries like the other terminal-count strings
on that line and in the move offer, which were already pluralized.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(automations): close failed and exited runs' terminals once the shell is proven alone (#25859)

* fix(automations): close failed and exited runs' terminals once the shell is proven alone

Failed (command not found, timeout) and forever-dispatched runs kept one
shell per run, unbounded and counted by the update gate. Their terminals now
close like completed ones (unused, past the grace, outside the newest few)
but only on fresh execution-host proof that the spawned shell is alone at
its prompt; a live or unprovable agent keeps its terminal. A still-dispatched
run closed this way is marked failed. Dead terminals no longer take one of
the newest-three keep slots. The update drain follows the same rules.

* fix(automations): prove a run shell alone from the process table, not the daemon's ownership flag

On a real daemon session the daemon's confirmShellForeground stays false
after a plain 'command not found' and after an agent that exited, because its
ownership flag only turns 'shell' after a full-screen command; failed runs
would never have closed. The proof now also reads the host's process table:
on POSIX the PTY's root shell must own the terminal foreground group with
nothing stopped under it, on Windows the host's job-based child census must
be empty. Anything unobservable still keeps the terminal.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* test(ssh): extensive orca CLI matrix on Windows hosts (#25114)

* test(ssh): extensive orca CLI matrix on Windows hosts

Adds three dispatch-only app cells to the ssh-windows-hosts lane that drive the
e2e build and the bundled orca CLI against the provisioned Win32-OpenSSH host:
empty-host deploy/terminal/reconnect/orcad-restart/decommission, seeded
relay-era conversion, and an open relay terminal keeping the host on the relay.

* test(ssh): pin the relay-kept cell's runtime; keep cleanup from masking failures

* test(ssh): run decommission before the orcad restart in the managed cell

* test(ssh): decommission through orca environment stop; accept an unverifiable relay close

* test(ssh): require a confirmed relay close; app cells must run last

* test(ssh): log and accept either relay-kept census reason; keep app-cell test results

* test(ssh): relay-kept requires a live census and its status line again

* test(ssh): match the pluralized relay-kept status line

* test(ssh): orcad restart proves a new process, terminal adoption, and a kill-then-connect relaunch

* test(ssh): restart kills only the orcad server, not its terminal daemon; wait for a released profile

* test(ssh): collect orcad.log.1 so a restarted orcad's previous run is kept

* test(ssh): the managed cell proves a workspace listener is detected and attributed

* test(ssh): start the port listener without $, so a PowerShell terminal doesn't expand it

* test(ssh): the port check proves Windows command-line attribution; retry a dropped version read

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* refactor(runtime): move run-terminal client-use and shell-alone checks out of the branch-cleanup runtime (#25876)

* refactor(runtime): move run-terminal client-use and shell-alone checks out of the branch-cleanup runtime

orca-runtime-preserved-branch-cleanup.ts had grown past max-lines (303) with
the headless run-terminal helpers. Their logic now lives in
run-terminal-client-use.ts and the runtime keeps one-line delegators, with
behavior unchanged.

* fix(ci): the runtime Electron ratchet bundles its entry points once, not 2.5k times

check-runtime-electron-ratchet bundled ~2,532 entry points each in full (format cjs, no
splitting), so esbuild held thousands of copies of the runtime graph: about 2.2GB RSS and 11s per
run, twice per test file. It was in flight in every unit shard that died with "The runner has
received a shutdown signal" (#25815 5/5 twice, #25876 2/5 twice). With esm + splitting the shared
modules land in one chunk: same metafile, about 200MB and 2s.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* test(ssh): wait for the busy relay's child before probing it (#25916)

The fake relay's spawn is not visible to pgrep at READY on Linux under Bun, so
the probe could count zero children. The sibling cases already wait.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a quit that aborts an upload whose read already ended no longer crashes main (BUG-23) (#25922)

Quitting while an on-connect orcad update was uploading its bundle aborted the connection's
teardown signal. sftp-upload's abort handler destroyed the local read stream with the signal's
reason, but once that read had ended, 'finished' had already removed its listeners, so the
stream emitted an unhandled 'error': [main_uncaught_exception] AbortError: This operation was
aborted. Electron's error dialog then blocked the main thread and the app never exited.

The read stream now always has a no-op error listener; the transfer's outcome still comes from
'finished'. Both the bare upload and the connection-level teardown abort are covered.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): a launch reads readiness at least once; fake hosts match the capture's tar flag, not any -cf (#25918)

A random fence token contains `-cf` about 1 time in 125, and the fake hosts
read any command containing it as a snapshot capture, so a rollback's restore
answered CAPTURED and the rollback never launched. Separately, a client
descheduled between computing the readiness deadline and checking it skipped
every read and failed a ready launch.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): a fence this desktop's exited process left is cleared without the 20-minute wait (#25941)

* fix(orcad): a fence this desktop's exited process left is cleared without the 20-minute wait

The client records the fence tokens its processes hold beside the profile. On
a later launch, a POSIX fence carrying a token from a process that has exited,
quiet for three heartbeats and with no live state mutation, is backdated so
the existing stale rules clear it or hand it to Recover at once. Another
desktop's fence, a live holder's, or one with a mutation still running keeps
the normal stale window.

* fix(orcad): held fence tokens are best effort, pinned to this machine and boot, and pruned after a day

A token-file write that fails no longer breaks a fence operation; an entry
recorded on another machine sharing the profile, or before a reboot, never
proves its holder exited; entries older than 24 hours are dropped. Tests cover
a journal kept for Recover, a successor freshened back after a racing backdate,
a Windows host, and the record across release, supersession, busy and a lost
connection.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): an install lock this desktop's exited process left mid-upload is taken over without the 20-minute wait (#25991)

A quit during the bundle upload leaves the version dir's install lock, not the
activation fence. The lock now carries this desktop's token, recorded in the
held-token store, and is forgotten only once its removal is confirmed. On a
later attempt, before each stale check, a POSIX lock whose token belongs to an
exited process of this machine and boot, quiet for three minutes, is backdated
so the existing stale takeover claims it at once. The fence path now shares
the same helper.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): a desktop that met another desktop's update fence clears its note once the host answers (#25995)

The serving note "holds this host" and a fence-busy update deferral stayed until a reconnect,
minutes after the other desktop's update finished. The connect now rechecks serving and the
update every 45s while the fence holds, and publishes the first answer without it. A recorded
deferral is dropped once the host runs its candidate or a newer release.

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude <noreply@anthropic.com>

* test(ci): run Phase 3's SQLite-backed tests in the Node runtime project

Main's #25967/#25998 boundary requires every test that opens real SQLite to be listed. This adds
Phase 3's eleven orcad and SSH migration tests, plus main's own agent-launch-instant-tab test
(#25430), which main's tip also leaves unlisted.

* fix(ci): keep Electron probes out of the node-server suites again (#26046)

The runner excluded *.electron.test.ts with a CLI --exclude, but main's switch to Vitest inline
projects (#25967) gave each project its own exclude list, which overrides the CLI one. The
directory selectors then pulled profile-state-writer-stall.electron.test.ts into the glibc-floor
and musl orcad-template jobs, which have no xvfb. Resolve the exact files with vitest list and
drop Electron and cross-runtime ones before running.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): keep the SSH host card quiet while its managed server is healthy (#26072)

The card showed "Runs a managed Orca server" under every healthy host. A managed server is the default, so the status line now appears only for setup progress, updates, the relay, or failures.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): Move to managed server keeps the host's terminal tabs (#26077)

* fix(ssh): Move to managed server keeps the host's terminal tabs

Move stops the relay shells; their exits read as a user exit and closed
the tabs before the conversion copied them to the server. Suppress those
exits for the move, restart stopped shells on the relay when the host
stays, re-home the open workspace onto the server, and report stopped
shells to the runtime so terminal list stops calling them connected.

* fix(ssh): mark Move's relay stops in main's intentional-stop register

The renderer-only exit suppression left main retiring the stopped tab from
the saved SSH session before the conversion copied it, left other viewers
unprotected, and swallowed real exits for the whole request. Register
exactly the shells the move stops, from just before each shutdown, as a
'replaced' stop with their incarnation; main keeps the surface and labels
the exit for every viewer, while a confirmed death stays 'exited'. Move
now returns the shells it stopped, and a host that stays on the relay
restarts only those tabs, discarding any buffered exit first.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(hosts): show an SSH host and its managed Orca server as one host (#26076)

* fix(hosts): show an SSH host and its managed Orca server as one host

Phase 3 registers the Orca server it deploys over SSH as its own runtime environment, so every
host list built from the execution-host registry listed the machine twice under the same name.
The registry now folds the pair into one row named after the SSH host. The row routes to the
server, since a managed host has no relay, unless main reports the host back on its relay; the
other id stays as an alias so selections and renames saved under it still resolve. A retired id
that workspaces still point at keeps its own row, and servers no configured SSH host deployed
(manual pairings, orphans) are untouched.

* fix(hosts): keep both ids of a merged SSH host and dedupe only in pickers

Deleting the merged-away id from the registry broke every consumer that matches hosts by exact
id: Add Project fell back to local after a connect, the composer lost ready projects and drafts
(and could swap in an unrelated local project), and a host scope hid folder-only workspaces.

The registry now keeps both entries and marks the pair (aliasHostIds on the row pickers show,
mergedIntoHostId on the other). Pickers show one row per machine, and a choice of that row
expands to both ids: sidebar host scope, jump palette filter, notification toggles, run-target
and repository host offers. Add Project resolves a saved SSH id to its server row and blocks the
actions while that server comes up instead of choosing local. The composer's resolver now fails
closed when a named draft repo isn't actionable rather than picking another project.

* fix(hosts): widen saved host scopes, both-way palette aliases, guard Add Project host

- A sidebar or agents host scope saved by an older build (or before a route flip) can hold one id
  of a merged SSH host; a background gate widens it to both ids so exact-id filters match either
  owner.
- The palette host filter now resolves a saved id to both owners whichever id it names.
- Add Project's create and clone refuse to run while the chosen host is unresolved, and their
  submit buttons stay disabled, instead of falling through to this computer.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(sync): reconcile main's ratchet bundling and cold-serve hydrate with phase3

The Electron-import ratchet keeps main's single-stdin bundle (cjs); the
auto-merge had also kept phase3's esm splitting, which broke main's
import-graph test. Editor tabs now follow the windowless full-seed rule
from #26022, so a cold serve restart lists persisted editors too.

* fix(ssh): reclaim this desktop's own exited lock on Windows hosts too (#26087)

* fix(ssh): reclaim this desktop's own exited lock on Windows hosts too

The relaunch after a quit mid-update now frees the activation fence and the
version-dir install lock on a Windows SSH host the same way it does on POSIX,
instead of waiting out the 20-minute stale window. The host script ages the
lock only when its token belongs to a desktop process proven exited, it has
been quiet for three heartbeats, and (for the fence) no state mutation is
live, where a mutation holder counts as gone only by pid plus creation time.

* fix(ssh): take an exited holder's lock only through the steal arbitration

Review found the reclaim backdated the lock by path after checking it, so a
live successor that replaced the lock in between could be aged and then
stolen, and an interrupted or failed restore left it aged for good.

The exited-holder check is now read-only. The steal command itself accepts
the proven token and, inside its steal claim and identity recheck, also takes
a lock whose owner file still names that token and that has been quiet for
three heartbeats. Nothing is written to a lock before the steal owns it.
POSIX uses the same path.

* fix(ssh): never take an exited holder's fence while a state mutation can start

Review round 2 found the fence's live-mutation guard ran only in the read-only
proof, so a mutation admitted after the proof, or one whose first heartbeat
landed after the steal sampled the fence's age, kept running under a fence
the steal had replaced.

For the fence, the steal now takes the state-mutation lock inside its claim
(mkdir on POSIX, the exclusive owner.json on Windows) and holds it until the
takeover is done; it refuses when any mutation lock exists. Holding it, it
rereads the owner and only then re-samples the fence identity. A mutation now
rechecks its fence token right after it takes the mutation lock and stops with
the fence-lost marker if it changed. The Windows proof also falls back to the
stale window when its command line would not fit cmd.exe.

* fix(ssh): record the exited-owner steal as a real mutation-lock holder

Review round 3 found the POSIX steal held the state-mutation lock as an empty
directory, which a mutation reclaims after a minute without any liveness
check; a steal stalled that long lost its exclusion and could replace the
fence under a running mutation.

The steal now writes its pid (and group, under the same rule) with the
mutation's own noclobber owner writer, so only proof of its exit frees the
lock, and it removes the lock only while the lock still names it. On
Windows the owner record is moved into place whole, so it never exists
empty, and is removed only while it still names the steal's pid.

* refactor(ssh): keep the relay lock commands off the orcad host-script graph

The mutation-lock owner writers moved into a leaf module, so the relay's
install-lock commands no longer import orcad-state-snapshot and, through it,
the Windows host script, orcad-instance-lock and the daemon process query.
Those modules evaluate imports at load time that existing suites mock
partially. No behavior change.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-07 03:15:06 -07:00
Jinwoo Hong 4077fb4a8d fix(ci): actually exclude Electron probes from the headless node-server lanes (#26093)
vitest 5 does not apply a CLI --exclude to inline projects, so the glibc floor
container ran profile-state-writer-stall.electron.test.ts and failed with
'spawn xvfb-run ENOENT'. Resolve the selectors to files and filter them before
vitest sees them.
2026-10-07 03:10:17 -04:00
NeilandShuhei Konno 67dc092f18 fix(windows): keep generated skills and snapshots LF (#25972)
Pin generated skill JSON and Vitest snapshots to LF so Windows autocrlf checkouts pass byte-for-byte verification. Cover the checkout behavior with a real Git regression test.

Co-authored-by: Shuhei Konno <shuhei.konno@gmail.com>
2026-10-07 00:09:40 -07:00
Neil 0acf039b5d Keep native contracts on Node and wait for Git upgrade completion (#26034)
* Keep native watcher and supervision contracts on Node

* Route real permission and new ledger contracts through Node

* Synchronize handshake cleanup with forced termination
2026-10-06 23:09:33 -07:00
Brennan Benson acea59c6d8 refactor(native-chat): remove the records-file and per-chat journal imports (#26038)
* refactor(native-chat): remove the records-file and per-chat journal imports

Every native chat user is on a build that already moved chat records and
history into the app-wide database, so the one-time copies are dead code:
the agent-sessions.json import and its owed-copy flag, the per-chat
journal.db import and the write-queue hold behind it, and the pre-SQLite
log.jsonl notice.

* refactor(native-chat): remove what only the deleted imports used

- JournalHostDatabase.unsyncedTransaction and the synchronous-pragma reset
  that undid it; JOURNAL_SYNCHRONOUS is now module-local. The stranded
  rollback test drives transaction() instead.
- deleteUnpublishedJournalRows and its SQL, plus its test.
- boundJournalStatusText.
- Stale comments naming the removed first-use copy (queued messages,
  send refusal example, JOURNAL_SYNCHRONOUS doc).
- Write-queue and readInOrder comments: a write also lags when issued
  behind one still waiting in line.
- Resolved-append ordering test binds the sink from inside a running read,
  so a resolver that read at handover now fails it.

* test(native-chat): run resolved-append in the SQLite runtime project
2026-10-06 23:03:32 -07:00
Brennan Benson d0b0f13b74 fix(native-chat): hold a queued message until the turn ahead opens (follow-up to the Stop-event plan, fixes STA-9348) (#25217)
* test(native-chat): a Stop over a card sent now into the running turn keeps it paused

Red on main: Codex's turn end withdraws the steered hand-off and the queue
sends the card again as a new host turn, with no pause recorded.

* fix(native-chat): a Stop's queue pause holds a card whose hand-off is still unanswered

A card sent now into the running turn was still pending when Stop judged the
pause, so nothing was recorded; the interrupt then withdrew the hand-off, the
card went back to waiting unpaused, and the queue sent it again as a new turn.
The pause now counts a hand-off that may still return to waiting, judged with
the appended row applied, so a withdrawal lands under the pause and an
acceptance retires it in that same write. Codex and Claude both hit it.

* test(native-chat): the Claude re-send case fails on its diff, inside the test's budget

* test(native-chat): a pause held by an unanswered hand-off ends on every path that ends it

The provider's answer, the provider dying, the chat closing, a restart, a
withdrawal still owed at open, and a /clear (refused until the hand-off ends,
then carrying every waiting card paused 'cleared'); each ends with the queue
sending again.

* fix(native-chat): narrow the pause's settled hand-off, and assert the queued receipt's card

* refactor(native-chat): derive the queue's pause from Stop and Resume journal rows

Stop now appends one journal row where it takes effect, before the interrupt,
whatever the queue holds; Resume appends its own. The pause is a pure function
of the fold: the latest Stop with no later Resume and no later accepted turn a
person asked for. A /clear's carried cards name their source, which is the
replacement's 'cleared' pause. Host-origin turns never lift either.

One predicate decides which cards a pause holds; by default every waiting card
without a hold of its own, including one queued after the Stop. The drain's
consume re-judges it inside its own transaction.

The rows are tombstones of an id no item takes, carrying the mark: a released
host reads an unknown row kind as corruption and truncates the journal there.

Deletes the stored pause (recordPause, the retire hook on every appended row,
the settle-before-record step, mayReturnToWaiting and its row overlay) and the
tests that only proved it retires. The queued_message_pauses table stays in the
schema, unread and unwritten, for downgrade safety.

* fix(native-chat): a card queued after a Stop sends normally, never ahead of held ones

A Stop's pause now holds only the cards queued before its row, plus a steer it
withdrew, which returns to its own place. Each card records the journal
position it was queued at, and the one hold rule compares that with the Stop
row. A card queued after the Stop is a new instruction: it sends as usual, but
the drain still stops at the first held card, so it never overtakes them.
/clear's pause holds the cards it carried. Holding every card again is a
one-line switch in that rule.

* fix(native-chat): the queue's own send re-checks the no-overtake rule in its transaction

The drain's pick and its consume now read one function, nextSendableQueuedCard,
so a Stop row that lands between them holds a newer card behind an older held
one exactly as the pick would. Notes why Stop and Resume ride a tombstone row.

* fix(native-chat): stop creating the unused queue pause table

The queue's pause is derived from journal rows, so nothing reads or writes
queued_message_pauses. It was still created on every open "for downgrade
safety", but an older build creates it itself when it opens the database, so
the table only sat empty in every new database. The tests now pin that no
pause table exists.

* fix(native-chat): a Stop's pause never hides the restart pause

A Stop holds only the cards queued before it. The pause derivation still
returned the Stop alone whenever it was in force, so the restart pause was
never considered: a card queued after the Stop, written by a host process
that has since exited, sent by itself after Orca restarted, with no pause
header and no Resume. A /clear pause that held nothing could hide it the
same way.

Every pause in force is now derived. A card is held if any of them holds
it, and it names the first that does. The drain's pick, the consume
transaction's re-check and the published header all read that one rule;
the header names the pause holding the first card Resume would send.

* test(native-chat): pin the Stop's no-resend, lift and held-card rules

- The Claude and Codex Stop-withdraws-a-steer tests checked "not sent
  again" at one instant, before a queue ignoring the pause re-sends. They
  now wait for the stopped turn to end and re-check after a quiet window.
- The deleted-card test read a card queued after the Stop, which sends
  whether or not a person's turn lifts it; it now reads the Stop's pause
  before and after that turn.
- Unit cases pin that a Stop holds a card with no recorded position and one
  queued before a rewind.

* refactor(native-chat): a Stop writes one Stop event with its reason, turn and caller

The Stop row that paused the queue becomes the general Stop event
{ reason, turnId?, at, caller? }, whose reason is the host's existing stop
cause. It still rides a tombstone of a host-only id (a released host deletes
the journal from the first unknown row kind), and Resume keeps its own marker
on its own id. Only a person's Stop (reason user-stop) pauses the queue.

* test(native-chat): a rewind keeps a lifted /clear pause lifted and restates the same Stop event

* test(native-chat): pin that Stop and Resume rows never reach apps or count as history

* test(native-chat): only a person's Stop event pauses the queue

* test(native-chat): pin that a Stop's event precedes the interrupt and the at-start stop

Through the real host: the event names the turn and who asked and is in the
journal when the interrupt reaches the agent; at an agent still starting it is
there before the start is ended and holds a card queued before it; an idle Stop
writes one only when it withdrew a send; and the queue's claim re-judges a
pause that landed after its pick.

* test(native-chat): a card held at a starting agent is checked before the Stop's timing

Also says precisely what the claim's in-transaction pause check defends
against: the Stop and the drain share one serialized lane.

* test(native-chat): a released build keeps and folds a journal holding Stop events

Replays this build's rows from the released build's own journal database: every
row is kept, the history after the Stop still folds, and an older client is sent
only removed ids no item uses.

* style(native-chat): format the Stop event changes

* test(native-chat): type the released build's exports through one checked helper

* fix(native-chat): the Stop/Resume row guard narrows to those tombstones only

* test(native-chat): run the Stop-event downgrade test in CI, and cover a writable downgrade

The Stop-event downgrade test ran in no CI lane: unit shards exclude the
cross-version folder, and the cross-version lane runs a fixed file list that
did not name it. It is now on that list.

Its only case replayed the rows into a release's own fresh database, because
that release cannot open the current host database. A second case opens the
journal this build wrote with a main build that shares the database: it opens
writable, keeps every row, appends, and this build then reopens it with the
person's Stop still pausing the queue.

* fix(native-chat): a Stop that stops nothing new writes no Stop event

A Stop reaching a running agent wrote a Stop event on every press. Two
presses before the first interrupt landed wrote two events, so a card
queued between them counted as before the latest Stop and was held,
though a card queued after a Stop should send normally. A Stop naming a
turn that had already ended, as a phone sends late, also wrote an event
for a turn it never stopped.

It now writes one only when it withdrew a queued send, or stops something
no event records yet: not a turn the journal no longer runs, and not the
live turn a Stop still in force already names, unless a card was handed
over into it since, which this Stop's interrupt sends back and must hold.
The interrupt and the "already finished" note are unchanged. A Stop at a
starting agent still always writes.

* test(native-chat): pin that a later host, eviction or close Stop never lifts a person's Stop

* chore(native-chat): put each Stop-row doc on its own declaration, and say only user-stop is journaled

* fix(native-chat): any later Stop event ends a person's Stop pause

A person's Stop paused the queue until their next accepted turn or Resume,
and a later Stop of another reason (the host stopping the agent, an
eviction, a close) was ignored. Now the pause is the latest Stop event's:
a later Stop of any reason ends a person's pause, and only a person's Stop
pauses. The fold keeps the latest Stop event whatever its reason.

An eviction of a resting chat writes no Stop event (a Stop that stops
nothing writes nothing), so it cannot release held cards; a test pins that
no event means no lift.

* fix(native-chat): a second Stop press is a repeat even when the first came before the turn showed

A Stop pressed before the agent's turn shows in the journal (before
Claude's echo, or before Codex opens the turn) records no turn. A second
press once the turn showed compared that missing turn with the live one,
wrote a second Stop event, and held a card queued between the presses.

A repeat is now judged by what was sent since the Stop in force: with
nothing sent after it (a refused send aside), a Stop that named no turn,
or named the live one, is repeated and writes nothing. Anything sent since
and not refused, including a send whose fate is unknown, makes the new
press write, since its interrupt may send that card back to waiting.

Tests: the two-press case across the turn showing; a steer between the
presses settled unknown; and a Stop naming a turn that ended while the next
card is sent but shows no turn yet, which writes and holds that card. The
fold test that claimed an eviction path is renamed.

* fix(native-chat): the queue's pause ignores a Stop or Resume row holding a value no build writes

A Stop or Resume row's value is read from disk with no shape check, and
the pause fold stored whatever it found. A stored `stopEvent: null` would
then throw on every pause check for that chat: the queue's pick, its
send, and every queue update to clients. No build writes such a row, so
this is hardening.

The fold now reads a Stop only when it is an object with a string reason
and a finite time, and a Resume only when it is `true`. Anything else is
ignored: it pauses nothing and ends nothing. The row is still not treated
as malformed, which could cut the history short.

* fix(native-chat): a Stop still reads as yours after Orca restarts before the turn ends

Every stop that ends work now writes the Stop's event before it ends the child: a
person's close of the chat, an eviction (worktree teardown, orchestration stop, tab
cleanup) and the idle sweep's stop of a start that never landed. A stop that ends
nothing writes nothing, and quit writes none: its resume marker records why.

The turn-end write reads the latest Stop event where every turn row is built, so the
adapter's settle, the host's fallback and the relaunch's settle all agree: a turn a
person's Stop or close named, ending with no verdict of its own after that Stop, ends
as their cancellation. A relaunch's probe-bounded end is no earlier than a Stop that
found the turn running. When the provider refuses the interrupt and the turn runs on,
a refusal row answers the Stop, so a later crash still reads Failed; pressing Stop
again after a refusal is a new Stop.

* refactor(native-chat): a stop no longer carries its cause; the turn's end reads the Stop event

The cause of a stop was threaded in memory from each entry through the host's stop
step, the adapter router and each adapter's close onto the `ended` it settled with,
and Claude kept a per-turn copy of a Stop it sent. All of that is gone: adapters
settle a turn they cut as interrupted with no verdict, the host's fallback does the
same, and the one rule where a turn row is built (`turnEndAfterStop`) reads the
journal's latest Stop event to say whether it was a person's.

- `closeSession` / `disposeSession` take no cause; `ended` has no `stopCause`.
- Claude reads an error result after a person's Stop as their cancellation from the
  journal's Stop event (through the event sink), not from a per-turn slot, and a
  refused interrupt is the host's refusal row, not `withdrawTurnStop`.
- An owed wind-down keeps no cause: its retry's fallback reads the Stop event.
- The mutation context's Stop passes no cause: its step already wrote the event, and
  the delivery loop's child-end reason is read back from it.
- A Stop pressed before its turn showed applies to the turn that opens under it,
  unless a send a person made since was accepted.

* test(native-chat): a turn a later send opened is no Stop's that named no turn

* test(native-chat): the restart test's death proof carries its detail

* refactor(native-chat): a refused Stop leaves no record; a Stop only ever ends the turn it names

The stop-refused mark is gone: its tombstone kind, its fold, the clock-keyed match that tied it to
a Stop, and the exception that let a second press after a refusal write a new Stop. A Stop that
stops nothing writes nothing. A Codex refusal names a turn that is no longer its active one, and
the Stop names that turn, so the turn running instead never reads as the person's by its id alone.

* fix(native-chat): a Stop pressed before any turn showed stops only the turn opened next

A Stop that named no turn read as the person's cancellation for every later turn that opened
after it, until a send a person made was accepted. The queue's drain, orchestration mail and a
restart continuation send as the host, so a turn they opened long after, cut by a crash, read
"Interrupted" as if the person had stopped it. The Stop now applies only to the first turn
opened after it.

* fix(native-chat): an older Claude's error end after a Stop pressed before its echo reads Interrupted

Claude CLIs before 2.1.91 end an interrupted turn with an error result that names no reason. The
translator judged whether a person's Stop explained it by its own copy of the Stop rule, which
ignored a Stop that named no turn, so a Stop pressed before Claude echoed the send read "Failed".
The translator now writes such an end as interrupted with no verdict and no error row whenever a
person's Stop may name the turn, and the journal's one rule decides as it writes the end.

* fix(native-chat): a person's Stop and /clear each name why they end the agent

The host's mutation path ended the agent with one "recorded" ending for every caller, which read
back the reason of whatever Stop event the journal held last, however old. /clear writes no Stop
event, so its end took an unrelated earlier reason. Each caller now names its own: the chat's Stop
`user-stop`, whose event its own step wrote, and /clear `user-close`, the user replacing this chat.

* fix(native-chat): a host stop judges whether it ends work after the provider's rows land

A close, eviction or host stop decided whether it ended a running turn from the journal as it
stood, while the provider's own rows (the turn its echo opened) could still be in the session's
event sink. A close landing in that gap wrote no Stop event, so the turn it cut read as news. It
now reads after the sink drains, as a person's Stop does, through the same check; a drain that
fails or takes over a second reads working.

* fix(native-chat): a Claude Stop naming a turn that just ended still marks the follow-up it cuts

A phone names the turn it last saw. When that turn had ended and a follow-up was still unechoed,
Claude's Stop interrupted the follow-up and ended the child, but the Stop's event named the ended
turn, so the follow-up's turn the child's end cut read "Failed" under "Cancellation requested.".
A Stop that ends the provider's session ends whatever is in flight, so its event now names the
live turn or none, and a Stop that names none binds the turn opened next. Codex keeps naming only
the turn the Stop names.

The Claude Stop turn-end tests move to their own file, since the session-ending Stop suite is at
its line budget.

* fix(native-chat): the idle sweep reads working by the same rule as a stop's event

The sweep judged a chat resting while a send whose reply was lost was still unanswered, but the
stop's event writer counts that send as work. So the sweep evicted it and wrote an evict event,
which ends a person's Stop pause and let the cards behind it drain on their own. The sweep's owed
work now reads the main agent working the way every session list and the event writer do.

* test(native-chat): an aborted eviction's injected drain failure lands on the eviction's own drain

A host stop now drains the session's sink once to judge whether it ends work, so the tests that
fail the eviction's drain-published step skip that first drain.

* fix(native-chat): the idle sweep's rest writes no Stop event; it evicts a send that never echoes

The previous commit made the sweep count an unanswered send as owed work, which pins a chat whose
admitted send Codex never echoes forever, and the sweep exists to retire exactly that. That rule
returns. The sweep stops only an agent it judged resting, so its eviction now writes no Stop
event, whatever send it retires: a person's Stop pause holds through it.

* fix(native-chat): stopping a start that carries no send writes no Stop event

A host stop, eviction or close of a starting child wrote a Stop event whatever the start carried.
A start with a send already reads working, so the clause only mattered for a start with none,
which ends no turn and no send: its event only lifted a person's Stop pause and bumped the idle
clock, which is why the idle sweep had been changed to close the conversation in the same pass.
The clause goes and the sweep is #24072's again. The child's end still reads host-stop, as before.

* test(native-chat): a Stop's pause across a restart is tested with a restart that writes no event

The rig's restart closes the chat with an eviction, which now writes a Stop event when work runs
and so ends a person's Stop pause. "A Stop never hides a restart's pause" then passed with no Stop
pause left to hide anything. Those tests, and the pause-lift test whose dropped assertion returns,
restart as a process that dies with no close, which like a quit writes no Stop event, and assert
that both the Stop's and the restart's pauses are in force first.

* fix(native-chat): a host stop of a turn a person's Stop is still ending keeps that Stop's reason

An eviction or host stop that landed while a person's Stop or close was already ending the same
turn wrote a newer Stop event, and the turn's end reads only the latest, so the person's Stop of
that turn read as news. A host reason now writes nothing while a person's Stop still decides what
runs: the live turn it names or bound, or, with none, the turn a send opens next. The person's
own close still writes. The E2 tests now open and end the stopped send's own turn, as Codex does,
so the mail turn after it is not the turnless Stop's.

* fix(native-chat): an older Claude's error on a later turn keeps its error text after a Stop

The translator left an error result that names no reason to the journal's Stop rule whenever a
person's Stop named the turn or none, but the rule binds a Stop naming no turn only to the turn
opened next. So a real error on a later turn read "Failed" with its error text dropped. The
translator now asks the journal's rule itself (`personStopDecidesTurn`, the one core
`turnEndAfterStop` and a host stop's in-force check share), so the two cannot disagree.

* fix(native-chat): a Stop of a start that never landed binds no later turn, whatever sent it

A person's Stop pressed while the agent starts names no turn, and the send it stopped is
cancelled before it opens one. The Stop then bound the next turn anything opened (orchestration
mail, a restart continuation, the queue's drain, all of which send as the host), so a host
eviction of that turn wrote nothing and its crash or close read as the person's cancellation. A
Stop that named no turn now binds only a turn no send journaled after it opened: any send since,
of any origin and not refused, opens its own. The E2 test's mail send is accepted as Codex
accepts it, instead of opening the stopped send's own turn first.

* test(native-chat): a rewind's restated turnless Stop binds no turn opened after the rewind

A Codex rewind restates a person's Stop still in force after the turns it keeps, at a new
sequence, so by sequence alone it would bind the next turn opened after the rewind. A send
journaled after the restated row voids that binding (the previous commit), which this pins.

* fix(native-chat): a relaunch settles a person's stopped turn with no "stopped while in progress" row

After a restart, a turn a person's Stop ended reads "Interrupted after N" with the muted mark, but
the relaunch still added the error row saying the provider stopped mid-response, which a live Stop
never writes. The settle now skips that row when every turn it interrupts is the person's Stop's
by the journal's one rule; a crash nobody stopped keeps it.

* test(native-chat): the unexpected-exit settle's journal fake answers whether a person's Stop decides a turn

* fix(native-chat): a host stop whose sink drain fails reads the journal as it stands

A host stop drains the session's sink before judging whether it ends work, and a failed or slow
drain read as working. So an eviction of an agent at rest wrote a Stop event that ended nothing,
which lifts a person's Stop pause, and a close wrote a person's event naming no turn. The drain is
now best effort: the stop goes ahead either way and only its record is at stake, so a failed or
slow drain leaves the journal's read as it stands. A person's Stop keeps its own rule.

* fix(native-chat): a Stop that named no turn applies only to a turn a send it stopped opened

A person's Stop pressed before any turn showed names no turn. It bound the first turn opened
after it, then (5ead1f6bcc) any turn opened by no later send, so a turn the host started for
a card the Stop held, or for orchestration mail, read as the person's cancellation, and a host
eviction of it wrote no Stop event when its send had been abandoned by the close first.

The rule is now the concept itself: a Stop naming no turn applies to a turn opened by a send it
stopped, one already handed to the agent at the Stop's position. Nothing new is stored. The turn's
row names the send that opened it (Codex: the submission's key; Claude: the echo, which the journal
aliases to the submission), and a handed-over send's item sits at its handover, so the target set
is derived from the journal. A card the Stop held is handed over after it, so it is no target; a
Stop of a start whose send never opens a turn binds nothing; a rewind keeps no submissions, so a
restated Stop binds no turn opened after it. With no turn running, a host stop defers to the
person's Stop only while every unanswered send is one it stopped. Claude's translator, which asks
before its echo row lands, passes the send its echo acknowledged.

* fix(native-chat): a host stop whose sink drain runs long reads the agent working; a failed one reads the journal

A drain past its bound may still hold the turn's row, while the echo's acceptance has already
landed, so the journal as it stands read nothing running: a person's close of that turn wrote no
Stop event and the turn read as news. The two drain outcomes now differ: one that failed has
nothing more to deliver, so the journal's read holds (as before); one still running reads working.

* test(native-chat): a host stop with no turn running defers only while every unanswered send is the Stop's

The branch had no test. An eviction with only the stopped send unanswered writes nothing; one
with a send made after the Stop still unanswered writes its event.

* fix(native-chat): a slow sink drain reads working only while an accepted send's turn row is due

The previous commit read every drain past its bound as working, so a host eviction or stop of an
agent at rest during a sink backlog wrote a Stop event that ended nothing and lifted a person's
Stop pause. A slow drain now reads working only when the latest send the agent accepted has opened
no turn the journal holds, the race it was for; otherwise the journal's read holds.

* test(native-chat): the host-stop control keeps the stopped send unanswered beside the later one

With both unanswered, the host writes only because not every unanswered send is the Stop's; a rule
that deferred when any one was would pass the old control.

* fix(native-chat): a steer is no send owed a turn when a slow drain judges a host stop

A slow drain reads working when the latest accepted send has opened no turn yet. A Codex steer or
a Claude fold is accepted into the running turn and never opens one, so a chat at rest whose last
send was a steer still read working, and an eviction lifted a person's Stop pause. Sends delivered
into a running turn, whose item carries that turn's scope, are skipped.

* feat(native-chat): a chat reads Stopping from the person's Stop until the work it stopped ends

The host derives it on each journal publish from the Stop's event, the live turn and the Stop's
own answer, and publishes it as an optional field on the session status and the main agent's row.
Clients present it: the chat's tail line and Stop control, the sidebar row, worktree ps and the
phone's row. The chat and the phone also read their own Stop press until its request answers.

* test(native-chat): pin Stopping on the phone and across mixed versions

* test: give touched fake journals and mocks their SAFETY notes

* test(native-chat): a turn waiting on the person reads attention, never Stopping

* test(native-chat): the sidebar row follows Stopping when it is the only field that moved

* test(mobile): the phone reads Stopping from its own Stop until the request answers

* test(mobile): type the held Stop request instead of casting it

* chore: keep the base lockfile (a local pnpm run rewrote it)

* fix(native-chat): narrow the Stop note's optional failure; type the phone test's reply

* test(native-chat): type the Stop test envelope's fields narrowly

* fix(native-chat): a Stop the agent declined, or whose child end failed, says so while the turn runs on

A Stop naming the turn that still runs, refused by the agent, now writes the
Stop's refused fact instead of 'already finished'. A session-ending Stop whose
child end fails while the work runs on revises its note to unconfirmed.

* fix(native-chat): Stopping holds while any press of the Stop took

A repeat press refused after an earlier press took no longer clears Stopping.
Exit early when the Stop named a turn that is not the live one.

* fix(native-chat): keep Stop enabled while the host says Stopping

Only this client's own Stop request in flight disables Stop and Esc. A repeat
Stop is how a stop the provider took but never answered escalates.

* fix(sidebar): every agent row says Stopping in place of its tool line

The dashboard row, which the sidebar's non-compact mode also draws, read the
last tool line while a person's Stop ended the turn. It now shares the compact
row's rule.

* fix(mobile): the worktree list sees Stopping change on its own

A snapshot whose only change was the host dropping Stopping compared equal and
was thrown away, leaving the row on Stopping.

* refactor(native-chat): fold the status feed in src/shared for both clients

The snapshot merge and the contact-loss strip move out of the renderer feed so
the phone folds the same stream the same way.

* feat(mobile): let phones read the structured session status stream

agentSession.subscribeStatus joins the mobile allowlist. The agent-session
methods move to their own file, which the at-cap allowlist spreads in, and the
allowlist test reads the Set instead of parsing the source.

* feat(mobile): the phone chat reads Stopping from the host, like the desktop

One status stream per client, opened on a host that advertises the status feed;
a refusal to phones reads as no feed. The chat reads Stopping from the host or
its own press, and holds Stop only while its own request is in flight.

* test: give the new fakes checked types or a SAFETY reason

* revert(native-chat): drop the refused-named-turn rewrite of a Stop's note

Codex can send its refusal before the turn's end frames, so reading the turn as
still live after a flush races; the Codex Stop that ends nothing is handled by
ending the process instead. The base's 'already finished' note and its test
expectation return. The Claude wind-down failure revise stays.

* fix(mobile): release the status stream when the host ends it; refusals last one connection

The feed now drops the handle of a stream the host ended or refused, so the
logical client never replays it on a later session. A refusal to phones holds
for one connection, so a host updated while the phone stays paired is asked
again.

* perf(native-chat): read the live turn's opener from its record when deriving Stopping

After a Stop that named no turn, every later commit walked and copied the whole
journal to find the live turn's record. The derivation now reads that record
off the rendered snapshot's tail and decides with the same rule.

* refactor(mobile): move the method-unavailable check into transport

The status feed imported it from the Files tab's fallback. No behaviour change.

* test(native-chat): a send after a Stop reads Working before its turn opens

Pins the derivation's running-only read of the newest turn: the stopped turn,
already ended, must not keep the next send on Stopping.

* fix(sidebar): the compact row leads with Stopping so a narrow sidebar keeps it whole

At the default width the row read 'Codex Chat - Stoppin…': the model and time
keep their room and the line truncates from the end. Stopping now leads the line
the way monitoring already does, so the chat name is what gets cut. Also pins
that the turn bar keeps its running clock while the tail line says Stopping.

* test(native-chat): the retry of a close whose exit was unproven writes no second Stop event

The idle sweep finishes a stop left owed with that stop's own cause. It is the same stop, so its
event stands alone and the child's end keeps the cause, for a person's close and an eviction.

* fix(native-chat): read and write a Stop's answer by the turn its event records

Stop notes are now one row per turn, keyed by the turn the Stop's event
records. Performing a Stop and deriving Stopping share that key. A refused or
unconfirmed answer never overwrites one that took at the same key; only a
session-ending Stop's failed wind-down downgrades it, and that step now revises
the note the Stop actually wrote, carried on the wind-down. Stopping reads the
turn's note whenever it was first written, plus newer notes no other turn owns.
Test fixtures gain the host logger and the phone's quietRepeatedStop.

* refactor(native-chat): read a Stop's target once for its event and its note

The chat's Stop now reads what it is aimed at (the named turn and whether the
Stop ends the provider session) once, and both its event and its note's key
derive their turn from that one value through the same rule. Adds the host test
for a Stop naming an ended turn on a provider whose Stop ends the session.

* fix(native-chat): the host never steers a message into a turn a Stop is ending

A queued card's Send-now, or a send made while a person's Stop ends the turn,
went to the agent as a steer into that turn. The delivery loop now holds any
waiting message while the host's own Stopping reading holds, and sends it as
its own turn once the turn ends. The Stopping reader also stops at the Stop's
position and looks the turn's note up by key, instead of walking the whole
journal.

* feat(native-chat): while Stopping, the composer says a message runs after the stop

Desktop and phone: the composer placeholder reads "Queue a message to run
after the stop" while the chat reads Stopping, and a queued card's Steer (and
the desktop's steer shortcut) is held. New key translated in all 6 catalogs.

* refactor(native-chat): the chat pane's Stop controls live in their own module

The pane went over its line limit once merged with main. Its Stopping reading,
the press that holds Stop, and the steer and placeholder it hands the composer
move to native-chat-structured-stop-controls.ts.

* fix(native-chat): hold a send at its handover, reading the feed's own Stopping

The hold was checked when the delivery step started, but the handover runs in a
later step after waiting on the agent's start, so a Stop landing in between let
a new send steer into the stopping turn. The check now runs at the handover.
It reads the status feed's projection for the commit instead of rendering the
journal again, so holding a send adds no journal read of its own.

* refactor(native-chat): one display status decides Stopping on every surface

agentStopDisplayStatus combines whether the agent works, the host's flag and
this client's own press. The chat pane, sidebar and dashboard rows, and the
phone's chat all read it, instead of each combining the flags.

* fix(native-chat): Stopping holds until the stopped turn ends, whatever the Stop's answer

A Stop the agent declined, or whose end went unconfirmed, used to drop the chat
back to Working. It now stays on Stopping until the turn ends, and Stop stays
enabled so a repeat press escalates. The Stop's answer is no longer read for
Stopping, so its note goes back to the key the base gives it (the restore of
queued-stop.ts and the removed key test landed in the previous commit). The
note still keeps a press that took over a later refusal, and a failed process
end still says the Stop went unconfirmed.

* fix(native-chat): a Stop binds only the turn it actually stopped

A Stop pressed before any turn showed used to claim, at end-write time,
whatever turn the stopped send later opened, even when the Stop stopped
nothing. A turn that then died on its own read as "Interrupted" (your
cancellation) instead of "Failed".

Now a person's Stop that named no turn binds, in memory only, every turn
that ends while the Stop settles, and afterwards only the turn its
interrupt took. The settle ends a still-running stopped turn once. A
relaunch finds nothing in memory, so an unsettled turnless Stop binds no
turn. A Codex Stop whose answered turn does not open within its wait, or
whose send's answer was lost, now answers refused, so the host ends the
child and the turn can never run.

* fix(native-chat): keep the person's queue pause and close binding after a Stop settles

A host stop or eviction with no turn running now defers to a person's
Stop while its queue pause still holds with nothing sent since, read from
rows, so a held card is not handed off on reopen after a Stop that did
nothing or whose kill failed.

A person's close that named no turn opens a settle around its child's
end, so a turn that end cuts reads as theirs.

A press opens its settle only when the latest Stop event is its own or
the one in force it repeats: a late Stop, a card's interrupt or a lost
event row reopens no earlier Stop.

A Codex Stop that cannot reach a turn still able to open says the Stop is
unconfirmed rather than that no turn ran, and a second Stop still reaches
a turn an earlier wait left unopened.

Also drops the unused openedBy plumbing and the unreachable "a written
cancellation stays one" rule, and pins a relaunch after a named Stop.

* fix(native-chat): keep a failed Stop's turn display-only, and settle edges off the commit path

A Stop that failed marks the turn it could not stop for "Stopping…" only
(JournalStopSettle.failedOn): no turn-end rule reads it, so that turn's
own end with no verdict reads as a failure, not the person's.

A settle edge writes no row, so it no longer goes through the journal's
commit listener, which also delivers history, counts as activity for the
idle sweep and schedules the queue drain. A narrow settle-edge hook
republishes the status row and wakes the steer hold's handover, and
nothing else.

* fix(native-chat): a Codex Stop agrees on both presses when a turn is still owed, and pin the close's settle

A Codex Stop that waited for a turn Codex answered a send into now answers
"may still open" whenever that turn neither opened nor ended and its send
is still owed, however the wait ended (it ran out, or the thread went
idle). Before, a first press after an idle thread said no turn was
running and kept Codex, while an identical second press ended it.

Adds a test that a person's close the conversation outlives (as /clear
does) closes its settle, so a later turn that ends on its own reads as a
failure.

* fix(native-chat): a Stop that failed before its turn showed still reads Stopping through that turn

A Stop that failed with no turn open marked nothing, so the chat dropped
to Working and the turn that then opened never read "Stopping…". The
display-only mark now also covers that case: the first turn that opens
after the Stop failed, provided no message was handed to the agent in
between. No turn-end rule reads the mark, so that turn's own end with no
verdict still reads as a failure.

* test(native-chat): a Stop whose event row failed binds no turn to an earlier Stop

With one ordered journal writer the Stop's event is in the fold when its
write returns, so the press reads whether it owns the latest Stop from the
fold instead of awaiting the write. Pins the case the read must refuse.

* test(native-chat): name the settle, not a stream drain, in the Stop's own-end test

* test(native-chat): a Codex Stop answered before Codex ends the turn reads interrupted throughout

Codex answers an interrupt it took before it sends turn/completed (interrupted):
on TurnAborted the app-server answers pending interrupts, then ends the turn, on
one channel. The test fake did the reverse. It now answers first and ends the
turn on a later read, and the tests that read the turn's end right after a Stop
wait for it.

New end-to-end test through the shipped host, journal and Codex adapter: with
the real order, every end row of the stopped turn reads interrupted by the Stop
(named, unnamed, and a Stop pressed while turn/start was in flight). Breaking the
settle window turns the in-flight case red: the Stop's own end row then has no
verdict, which reads as failed until Codex's end lands.

* test(native-chat): Stopping ends with a Codex turn whose interrupt is answered before its end

With Codex's real order (the interrupt's answer, then turn/completed interrupted),
the status shows Stopping while the Stop settles, drops it once the turn ends, and
never carries a verdict other than the person's cancellation.

* fix(native-chat): a Codex Stop interrupts a turn Codex answered but has not opened at once

A Stop that named no turn, made after Codex answered a send but before the turn
opened, used to wait up to 5 s for the turn to open before interrupting, and
ended the Codex process when it didn't. The stated reason, that Codex refuses an
interrupt until it opens the turn, holds only part of the time: with no turn
active, Codex takes an interrupt once its thread runs (turn_interrupt_inner),
and refuses it with -32600 "no active turn to interrupt" before that or once
the turn has ended.

The Stop now sends the interrupt at once. Only on that refusal, while the turn
has neither opened nor ended, does it wait for the turn to open (bounded at
5 s) and send it once more. A turn that ended meanwhile was nothing to stop. One
that never opens, or that an earlier wait already gave up on, fails the Stop,
and the host ends the child as before. Sends still wait for the turn to open
before steering into it.

The test fake models Codex taking an interrupt once the thread runs (run()).

* fix(mobile): name how the phone's status stream is released in the subscription inventory

Main made each inventory entry state its release; the status feed's stream is
released from its subscribe params, as the session event stream is.

* fix(native-chat): every Codex Stop waits for an answered turn to start, as the first did

A Stop whose interrupt Codex refused as finding no active turn skipped the wait
when an earlier wait, a Stop's or a send's, had already given up on that turn.
Every press now waits its own bound and retries once if the turn starts, so a
turn that opens during a later press is still stopped. Both presses still reach
the same verdict when it never starts.

* fix(codex): never steer a turn whose interrupt Codex answered

Codex answers an interrupt as the turn aborts, before it sends that turn's
turn/completed. In that gap the adapter still counted the turn as running, so a
message handed over right after a Stop settled (the Stop's own end row already
reads the turn ended) went out as turn/steer, which Codex refused with -32600
"no active turn to steer", and only then as turn/start. The adapter now marks a
turn whose interrupt Codex answered as aborted until its turn/completed, never
steers into it, and starts the message's own turn directly. Steering a turn
that is genuinely running is unchanged.

* fix(native-chat): a message sent while Stopping is queued as a card, whatever the setting

While the chat reads Stopping (the host's flag or this client's own Stop in
flight) there is no turn left to steer into, so the desktop asks the host to
queue the send even with the queueing setting off, and it is never drawn as a
bubble inside the turn being stopped. The phone already queued every send on a
capable host; a test now pins that it does so while Stopping.

* fix(native-chat): a message queued while Stopping is a card at once, not after the stop

A person's Stop holds the session's lane until Codex answers its interrupt, and
a send was admitted only behind it. By then the turn read ended, so a send that
asked to be queued went out plain: no card for the whole of Stopping, then a
bubble and a new turn.

While the host reads that a person's Stop is ending the work, a text send that
asks to be queued is admitted without waiting for the lane: the same ledger and
lease admission, and a plan that only writes the card through the journal's
ordered writer. The card runs when the stop lands; the Stop's pause holds only
cards queued before it. Anything else, including a Stop that settled by the
time the send runs, takes the lane as before.

The Codex test fake now drops the active turn when it takes an interrupt, as
Codex does before it answers, so a turn/start after the answer opens a new turn.

* fix(native-chat): a Codex Stop that ends the child before any turn opened withdraws its send

A Stop on a Codex turn that was answered but never opened ends the Codex
process. That end settled the send as in doubt (unknown, recovered), and the
client's outbox holds every later send behind a send in doubt until the person
presses Retry, which re-sends the very message they stopped. The chat looked
stuck.

Codex records a prompt only once its turn has started, so a send whose turn
never opened never ran. When the Stop's refusal says so (turnMayOpen), the child
end now settles the unanswered sends as withdrawn, the verdict Codex's own
interrupted-turn end already gives an unechoed send. The flag rides on the owed
wind-down, so a retry after a failed child end withdraws them too. Claude's
child end still leaves its unanswered send in doubt.

* fix(native-chat): derive the withdrawal of a Codex send whose turn never opened

Replaces the flag the Stop carried to the child's end, and its copy on the owed
wind-down, with a reading of the journal at the settlement that lands. A Codex
child's unanswered send is withdrawn when a person's Stop is in force since it
was sent and no turn row ran, or was written, after it; any other end (a turn
that opened, a host's close, a crash, Claude) still leaves it in doubt. A
retried wind-down reads the same rows, so it withdraws the same sends.

* fix(native-chat): a card queued while Stopping runs past the cards the Stop holds

A card queued before a person's Stop waits under its pause until Resume. One
queued after it, as a message sent while Stopping now is, was stuck behind them
too, since the queue never reorders. Such a card was asked for after the Stop,
so it runs when the stop lands, past the cards held only by that Stop's pause;
a returned card and the restart and /clear pauses still hold everything behind
them.

Also: the host's own Stopping reading is gated on working, as the published
flag is, so a failed Stop's mark never reads Stopping on an idle session; a send
that falls back to the lane re-reads the conversation's journal there; and the
end-to-end test asserts the queue's pause rather than a per-card field.

* fix(native-chat): a paused queue labels only the cards it holds

Since a card queued after a person's Stop runs past the cards the Stop holds,
labelling every card "paused" while the queue's pause is published misreads that
card. The host now marks each card its pause holds (heldByPause, a new optional
field), and the desktop and phone label only those. An older host marks none,
so a client keeps today's labels; an older client ignores the field.

Adds a test of the desktop's own send through the real outbox: while the chat
reads Stopping, the request asks the host to queue it and no bubble is drawn,
on a host that advertises the queue.

* revert(native-chat): defer the per-card queue pause label to the queue's rollout

The heldByPause field and its labels are visible only where the host
advertises the queued-messages capability, which shipped hosts do not yet do.
Deferred to that rollout; the real-outbox send test stays.

* fix(native-chat): while Stopping, say and show what a send does where the queue is dark

Shipped hosts do not advertise the queued-messages capability, so a message
sent while Stopping goes out plain: the host holds it until the stopped turn
ends and then runs it as its own turn. The composer still said "Queue a message
to run after the stop", and the message was drawn inside the turn being
stopped.

Now the placeholder reads "Send a message to run after the stop" where the host
does not queue sends, and "Queue a message…" only where it does (desktop and
phone, all six catalogs). A send this client made that the host has not
recorded yet is drawn after the Stopping line while the chat reads Stopping, as
a message held behind a running command already is; once the host hands it
over it opens its own turn. Client presentation only.

* fix(native-chat): keep a send in doubt when a turn was open for it

The derived withdrawal read a turn as open for a send only if it still ran or
was written after the send. A send steered into a running Codex turn whose
interrupt failed met neither once the adapter's end settled that turn ahead of
the host's settle, so it read withdrawn, though Codex drains a steer into the
running turn and may hold it. A turn that ended after the send was handed over
was open for it too: such a send stays in doubt, as before.

Pins that case, and that a send made after the Stop, to a child that then dies
before its turn opens, stays in doubt.

* fix(native-chat): restore the per-card queue pause label

Kept after all: a paused queue labels only the cards it holds (heldByPause),
which is visible only where the host advertises the queued-messages capability.

* fix(native-chat): draw only a send made while Stopping after the Stopping line

Every send the host had not recorded yet was drawn after the Stopping line,
including one made just before the Stop, which the host steers into the turn;
it then jumped up into that turn once recorded. The outbox now marks a send
made while the chat reads Stopping, and only those wait after the line.

* fix(native-chat): withdraw a Codex send by whether it started its own turn, not by timing

Whether a turn was open for a send was read from end times: a turn that ended
after the send's handover counted. A send made while a Stop ended the turn is
handed over once that turn reads ended, yet Codex's own end for it can arrive
later, so such a send whose own turn never opened read in doubt again, and the
chat's queue held behind it.

The handover already records where the send went: its message joins the turn
running then (a steer) or belongs to no turn (it starts its own). Only a send
that started its own turn, with none opened since, is withdrawn; one that
joined a running turn, or has no recorded place, stays in doubt.

The Codex test fake takes an answered interrupt as Codex does, dropping the
turn before its end arrives.

* refactor(native-chat): move queue-while-stopping to its own follow-up

The queued-messages capability is off on every shipped host (#21062), so the
parts of this PR that act only when it is on move to a follow-up stacked on
this one: admitting a queued card while a Stop holds the session's lane, a card
queued after a Stop running past the cards it holds, the per-card pause mark,
and asking the host to queue a send made while Stopping. This PR keeps the
Stopping state, the host's steer hold, the rule that never steers a turn whose
interrupt Codex answered, and what a send while Stopping looks like where the
queue is off.

* fix(native-chat): leave no Stop row when the Stop took back a send that never ran

A Stop on a Codex send whose turn never opened ends the child, and the child's end
takes the send back into the composer. The Stop still wrote "Cancellation
requested." at the conversation level, so with the send gone it sat under the
previous finished turn and read as if that turn had been stopped. A Stop that found
no turn running and whose child end took back every send it found now writes no
row; a Stop of a running turn, or one that leaves a send in doubt, still does.

* fix(native-chat): count a send whose answer was lost when a Stop takes it back

The no-row rule counted only pending sends, but the child's end also takes back a send this process left in doubt when Codex's turn/start answer was lost. That case still wrote "Cancellation requested." under the previous turn. Both now read one predicate, so they cannot drift apart.

* test(native-chat): pin which sends a Stop's child end can take back

A send an earlier process left in doubt is never withdrawn and never holds the row back, and a queued card's send is never counted.

* fix(native-chat): hold the next handover while the send ahead opens its turn

The delivery loop handed queued messages to the provider back to back. Codex steers a second send into the first one's turn once it opens, and Claude folds it into the running cycle, yet its handover row was written before that turn existed, so it read as belonging to no turn and was drawn ahead of the reply. The loop now hands the next message over only once the send ahead has opened its turn, settled, or been stopped, all read from the journal; each is a commit, which wakes the loop again. The message then goes in scoped to the opened turn.

* fix(native-chat): doubt a Codex send whose turn never opened once Codex goes idle

A Codex before 0.148 fails a turn before opening it with only an error. The send stayed pending, and the hold behind it waited on it. When Codex reports its thread not running with no turn open, a send answered into a turn it never opened or ended now settles as doubt with the existing idle reason, which releases the hold. No clock: a slow but healthy turn start is never doubted.

* fix(native-chat): end an unopened Codex turn on its final error, not at idle, and drop the Stop release

Codex publishes the thread idle ahead of turn/completed, and between an aborted turn and the next picked one, so releasing at idle doubted sends in turns Codex did open or was about to. A final error naming a turn Codex never opened (its only end before 0.148) now ends that turn as failed instead, settling its send as a turn/completed failure would. The test fake publishes idle before turn/completed, as Codex does. A Stop no longer releases the hold: the seven rig tests that needed it now settle the stopped send the way the provider does.

* docs(native-chat): put each turn-end settlement comment on its own function

* test(native-chat): echo the first send before the next in the restarted-child test

The test's child admitted the first message and never answered it, which no live provider does; the next send was then held behind a turn still opening. The child now echoes the first message, so the test still checks that the next send restarts nothing.

* fix(native-chat): draw a message held behind an opening turn after that turn's live status

While the send ahead was still opening its turn, the live turn was taken to be the newest user message, the held one, so its live status drew under the held message and the held message drew above it. While a send is opening, its turn is now the live one, and messages sent after it, queued or not yet recorded, wait behind the live turn as a message held behind /compact or a Stop does. The host's hold and the client's drawing read one shared predicate.

* fix(native-chat): keep held messages waiting while Stopping, and leave Retry-only sends in place

While Stopping, a message typed then returned early from the waiting rule, so a message held behind the opening turn moved back above the Stopping line. The rule now waits the union of both. A send only the user's Retry sends again is not held by the host, so it no longer waits behind an opening turn; the projection marks it.

* fix(native-chat): keep a turn on the send that opened it when Codex echoes a steer first

A message held behind an opening turn is steered in moments after that turn opens, and
Codex can echo the steer before the send that opened the turn. Both carry the turn's one
provider key, so the steer was read as the turn's opener: for that moment the live status
moved under it and restarted its clock. While the opener is still in flight ahead of the
turn record, a steer into that turn no longer takes it.

* fix(native-chat): never anchor an opening turn on a message still queued above its send

Two messages queued behind /compact, or sent while Stopping, are accepted above the first
send's handover, and the hold makes the first send's turn record land before the second is
handed over. The turn fell back to the first unechoed send ahead of its record, which was
the queued one, so its live status moved under it. A send still queued is not in flight.

The phone keeps no outbox, so its frame tests now draw only recorded rows, through the
phone's own fold rather than the desktop's transcript order.

* fix(native-chat): read the host's Stopping beside main's startup phase

Main now reads only the startup phase from the status feed and no longer publishes which
child is starting. The chat reads the host's Stopping from its own hook beside it, and the
Stopping bridge test mocks the execution-host lookup main's owner resolution now calls.

* test(orchestration): open the working send's turn before a `now` send joins it

The rig's working send was handed over with no turn record, so the running turn the test
names never existed; the hold rightly kept the `now` send until that turn opened. The setup
now opens it as a provider does. The assertions are unchanged.

* test(native-chat): pin the phone's own echo of an accepted send behind an opening turn

The phone keeps no outbox, but it does show its echo of a send the host accepted until that
send's row arrives, and the shared waiting rule moves that echo behind a turn still opening.
The phone test now builds its list as the phone's view does, echoes included, and covers it.

* test(native-chat): read the outbox reconcile from where main moved it

* test(native-chat): pass the projection's options after main's rejected-in-place rows

* fix(native-chat): a retried message no longer waits behind a later Stop

Retry dropped the Stop it had outlived but kept the mark that it was sent while
a Stop was ending a turn, so a retried message waited behind whatever later,
unrelated turn a Stop was ending. Retry is a new send: drop that mark too.

* fix(native-chat): word a send after a Stop by whether this send will queue

The 'queue a message to run after the stop' placeholder read the host's
queue capability alone. A send queues only when the host queues and this
send asks it to: the queue setting is on and no pending prompt blocks the
queue. Desktop and phone now word the placeholder from that same decision
their send uses.

* test(native-chat): one test per case for the words of a send after a Stop

* refactor(native-chat): the dictation hook owns the composer's dictation state

Keeps NativeChatComposer within its line limit after the Stop props and
main's /context answer both landed in it.

* refactor(native-chat): name the dictation hook for what it owns now

* fix(native-chat): a Stop's note says it took once a joined close proves the exit

A session-ending Stop whose child's end failed revises its note to
'unconfirmed'. Since main's #24862, the next Stop joins that close rather
than stopping again, and wrote no note, so a close that then proved the
exit left 'unconfirmed' under a turn that ended. The close now carries
the note it settles, and its proven exit revises it to 'Cancellation
requested.'. A join that fails again leaves it unconfirmed.

* fix(native-chat): a proven turn end says a Stop's unconfirmed note took

Replaces the note carried on the child's close. Every settlement that
ends turns interrupted on a proven exit, live or after a crash, also
revises an unconfirmed Stop note on those turns to 'Cancellation
requested.', found by the note's turn scope, in the same batch. A note a
Stop wrote before its turn showed is re-keyed onto the running turn when
it becomes unconfirmed, in one batch, so that end finds it. Known limit:
with no turn open yet, the note keeps its key and no turn's end revises
it.

* test(native-chat): a Stop's unconfirmed note says it took when the agent exits on its own

* test(native-chat): a message that joins a Stop's unproven Claude close goes to the resumed child, the next waits for its echo

Covers the case main's #24862 retired with its unproven-stop test: the
first message after an unproven close joins it and reaches the resumed
child, and a second waits for that message's turn to open.

* test(native-chat): name the Codex handle as main's opaque handle does

* test: restore the provider handle import the main merge dropped

* fix: derive Stop note wording from interrupted turns

* test: name the raw replay case for what it covers

* refactor(native-chat): move the waiting-slot split into its own hook

* test: give the opening-send hold host the agent registry main now requires

* test: follow main's chat font-size rename in the stopping tests

* test: follow main's chat font-size change in the opening-send test

* test: follow main's single live-line value in the Stopping tests

* test: give the android live-line fixtures the stopping field

* chore: keep the session host under its line limit after the main merge

* chore: keep the composer test and the phone chat view under their line limits after the main merge

The Stop control now disables itself while Stopping, so the composer passes
the flag through and its test file stays as main has it. The phone chat
header's Stop moves to its own component.

* perf(native-chat): read a Stop note's fields before parsing its key on every snapshot

Every snapshot projects each item through the Stop-note read, and each new
snapshot rebuilds the index of Stop notes by turn. Both parsed every item's
key first; they now check the row's kind and turn scope (and, for the
projection, its unconfirmed-stop failure) before the parse. Every Stop note
is a status row, so what each finds is unchanged.

* fix(native-chat): a Stop ends the hold on the next message, even when the stopped send's turn never opened

* test: a Stop's note ends the hold on the next message

* fix(native-chat): only a Stop that took ends the hold; a refused or unconfirmed one leaves the turn opening

* fix: import the moved Stop-note helpers where the host module uses them; type the test's failure note

* test(native-chat): pin that a taken Stop's note lands after the turn the provider opened before answering it

* test(native-chat): give the opening-send hold test's runtime the launch arguments main now requires

* fix(native-chat): a message a Stop takes back while the turn ahead opens stays after that turn, with its stop row; move the rows that wait behind the live turn into their own module

* test(native-chat): run the two opening-send hold tests, which open a real journal database, in the Node runtime project

* test(native-chat): with a send still opening its turn, a starting child gets only the first message; a failed start still rejects both in its words
2026-10-06 22:57:51 -07:00
64bb9373da Claude account profiles: dormant WSL guest setup (Step 3 of 4) (#24384)
* feat(claude): add dormant profile setup and history sharing

* fix(claude): make profile setup one gated, typed, fail-safe entry

Review round 1 of the dormant profile setup found that the pieces could
be called without their safety checks, that one failed write or an
unreadable bookkeeping file could silently stop sharing for good, and
that Windows prompt history could bring back history the user cleared.

- One entry, provisionClaudeAccountProfile: the profile gate (namespace,
  no linked components, outside ~/.claude and ~/.config/claude, and an
  ownership marker beside the home naming the account and target) runs
  first and refuses before creating anything; then history sharing,
  config provisioning, and the hook install after the settings merge.
  Results come back per surface with closed warning codes instead of
  message text.
- The sharing ledger is keyed by surface name, records a value only
  after its write succeeded, and an unreadable ledger starts empty and
  is rewritten instead of blocking every surface.
- The profile state file goes through the same locked writer as folder
  trust (Claude's <file>.lock plus the in-process queue), generalized as
  updateClaudeGlobalConfig. Onboarding and trust are still applied when
  the personal state file is unreadable.
- WSL descriptors build guest POSIX paths; the state-file path style
  follows the injected platform.
- Orca's managed statusLine has one owner in a profile: the settings
  merge never shares it, a user's own statusLine is shared over it, and
  the profile installer follows the default home's slot so a default
  opt-out reaches every profile. remove() takes the same destination;
  the remote installer cannot accept one.
- Prompt history compares file identity (bigint dev+ino) on every
  platform, never drains the shared file into itself, drains retained
  copies in generation order, never reuses a stale cursor, and on Windows
  keeps a replaced default's old copy aside instead of replaying it.
  Directory merges keep going past a failed entry.

* fix(claude): share the user's own hooks and keep merged history whole

A user's own Claude hooks in ~/.claude (notifications, formatters) did
not run under a managed account, because the whole hooks key stayed
private. They are now shared like any other settings key: Orca's own
hook entries and its managed statusLine are stripped from both the
personal value and the profile's current value before the per-key
ledger comparison, so they never travel through the merge and never make
the key look user-owned. Orca entries already in the profile are kept on
write, and the profile hook installer adds them on top as before.

Prompt history: merged bytes that lack a final newline are terminated,
so Claude's next record no longer fuses onto the last merged line. When
a CLI rewrote the profile's history file (old records plus new), only
the lines past the part it shares with the default history are added,
instead of the whole file again.

* fix(claude): close review round 2 gaps in profile setup

Hooks and statusLine sharing:
- When ~/.claude holds only Orca's hook entries, the user's shared hooks
  now read as an empty value instead of a missing key. Removing the
  user's last own hook in ~/.claude therefore reaches profiles that
  never edited it, and deleting the only shared hook inside a profile
  stays deleted.
- A custom statusLine Orca shared, and the profile never edited, goes
  away when the default home drops it. When a shared custom line
  replaced Orca's line in a profile, the profile's statusline marker is
  dropped so Orca's line comes back once the default returns to it; a
  profile that opted out stays opted out. No other key gains deletion.
- install/remove/getStatus with a profile directory refuse when it is
  the default home, or its settings.json resolves to the default one,
  instead of editing System Default's hooks and opt-out state.
- The profile statusline rule reads the default settings under the
  userHome passed to the setup entry, not os.homedir().

Profile state and ownership:
- A malformed `projects` value skips only folder trust (new warning
  code trust-refused); onboarding and shared keys still apply.
- The ownership marker stores only host-local facts (account, runtime,
  distro). The execution host id is the caller's view of the host, so
  it stays in the in-memory descriptor and is not compared.

Prompt history interruption paths:
- With no cursor yet, a retained copy starts past the bytes it shares
  with the default history, so an interrupted share no longer replays
  the whole history.
- A retained name for the shared file itself is removed with its cursor
  instead of lingering until a later scrub makes it look new.
- The Windows link record is read three-state: unreadable stops the
  share instead of reading as "no link". If the record cannot be
  written after linking, the fresh link is undone.
- An unreadable retained copy is reported and no longer blocks linking.

* build(cli): list the new Claude hook modules in the CLI project

hook-service.ts and hook-settings.ts are compiled into the packaged CLI
project, which lists every file explicitly. The statusline policy and
profile destination modules they now import were missing, so the CLI
typecheck failed with TS6307. The CLI still loads hook-service through
the existing managed-agent-hook-controls build entry, which bundles
both modules; neither imports electron.

* fix(claude): close review round 3 regressions in profile setup

- A profile whose hooks hold only Orca's entries and that sharing never
  recorded is no longer treated as a user edit, so the user's first own
  hook in ~/.claude reaches it (for example when the profile was set up
  before ~/.claude had any hooks).
- A retained prompt-history file is removed as a second name for the
  shared file only when the default history does not itself link to it;
  otherwise it holds the only copy and is kept.
- Default-home checks compare file identity: the profile hook
  destination check uses device and inode, and the profile/default
  separation check resolves on-disk case, so a case-only alias of
  ~/.claude is refused on case-insensitive filesystems.
- A test pins that an unreadable leftover session tree no longer blocks
  linking.

* fix(claude): let shared keys leave a profile when ~/.claude drops them

QA found that removing a setting from ~/.claude never reached a managed
account: deleting the whole `hooks` block left the user's hook running
there. Only statusLine followed the default away.

Every shared key now follows the same rule through the existing per-key
ledger: when a key disappears from ~/.claude/settings.json (or
mcpServers/theme from the personal state file), it is removed from the
profile if the profile still holds exactly what Orca last shared. A
value changed inside the account is kept. Keys Orca never shared,
including denylisted ones, are never touched. Deleting the whole hooks
block removes the user's shared hooks and keeps Orca's own entries. A
missing source counts as empty; an unreadable source removes nothing.

* feat(claude): add dormant profile routing and account consumers

* fix(claude): drop the dormant profile selection RPC; clients negotiate by capability

Restores the inline mobile allowlist so its source-scan guard sees every
accounts.* method again, and the generated params catalog to generator order.

* fix(claude): guard the claude shell function and honour a hand-exported config dir

The function is defined only in a routed pane where claude is a real
executable (the codex function's guard), re-reads the pointer only while
CLAUDE_CONFIG_DIR is unset or still Orca's injected twin, accepts Git Bash
drive paths, and starts on its own line after the fish/PowerShell codex text.

* fix(claude): spawn-time profile env, total account listing, setup at lifecycle triggers

Round-1 review fixes for the dormant profile routing:
- Panes get the selected profile's CLAUDE_CONFIG_DIR plus an Orca twin at
  spawn, so nested shells and scripts inherit the account; System Default
  injects nothing and its home is the inherited CLAUDE_CONFIG_DIR.
- An absent routing owner is System Default, never a throw; AI Vault and
  session-search scans receive profile roots from their parent, and the
  capability is advertised only where an owner is installed.
- Account listing never throws: per-account readiness, a stale pointer is
  republished in the background and reported on the snapshot.
- Profiles are set up at select and startup; a launch only sets up one that
  never was, and a worker fault on a prepared profile is a warning. The
  Claude version probe is cached per binary identity.
- Pre-trust goes through the existing deadline- and realpath-guarded writer
  against the launch env's profile config.
- Skill discovery keeps a caller's Claude root and a broken Claude selection
  no longer fails other providers.
- The durable record carries a provider-neutral launchAccountHome, read
  through one helper by the launch fallback and the model catalog.

* test(claude): pin the version-probe cache, launch-account record and temp-home readers

* test(claude): pin dormant bash rc text alongside fish and PowerShell

* test(claude): read the fish launch init without a nullable index

* fix(claude): withdraw the profile pointer when a selection cannot be published

A pointer left naming the previous account would launch it silently; a
missing pointer makes the claude function refuse visibly. A newer selection
that raced the failed one keeps its pointer.

* fix(claude): read the fish profile pointer with read -z for fish older than 3.4

Shell tests skip system config and abort unless claude resolves to the fake.

* fix(claude): only the newest publish withdraws the pointer; total config dir lookup

- An overtaken publish that fails leaves the newer selection's pointer.
- The runtime config dir falls back to the legacy home for an unresolvable
  account or a WSL target, so skill roots never fail for other providers.
- WSL guest reader roots merge verbatim, never realpathed on this thread.
- History readers include ~/.claude, where step-1 setup pools profile history.
- System Default ignores a config dir an outer Orca injected (twin-marked).

* fix(claude): System Default launches and probes use the structured create resolver

A Claude agent-env CLAUDE_CONFIG_DIR the create path stored is now the home
the launch pins and the model probe accepts.

* test(claude): type the System Default launch record as an agent-session record

* fix(claude): install profile hook scripts under the setup job's home

A worker thread's os.homedir() ignores its own env, so the hook and
statusline scripts now go under the home the job names. The worker test pins
the process HOME to a sentinel, refuses to run unless the worker sees it, and
asserts nothing lands there.

* test(claude): skip shell cases whose shell the runner lacks

* fix(claude): remove env vars in the PowerShell claude function instead of setting null

On .NET 9+ (pwsh 7.5+) SetEnvironmentVariable with $null creates an empty
variable, so stripped auth vars reached claude as empty strings and the
restore left CLAUDE_CONFIG_DIR empty in the user's session.

* feat(claude): add dormant WSL guest profile setup

* fix(claude): open WSL panes without guest calls and coalesce same-profile publishes

A WSL pane now gets the same non-throwing, guest-free spawn env as a host
pane; only select, startup and Claude launches publish into the guest.
Overlapping publishes of one target share the newest publish while the
selection still names the same profile, instead of failing as superseded.
Publish issues name their WSL distro and drop out when the target is no
longer routed. A late inspect from an older selection no longer replaces
the newer one's verification, a failed guest request evicts the cached
guest, and readiness is derived per account from the guest's owned homes.

* fix(claude): roll back only the target whose selection failed

With profiles, a failed select or remove republishes just its own target
instead of running startup over every WSL distro, and a rollback failure is
logged instead of replacing the error that caused the rollback.

* fix(claude): scan WSL profile history only in running distros

Vault and usage scans pass Claude profile roots through the same
running-distro filter as every other WSL root, so a stopped distro's UNC
paths are never walked.

* fix(wsl): ship the Claude profile helper only in the WSL bundle dir

The helper only ever runs inside WSL from the desktop, so it moves out of
the SSH relay artifacts (no upload, no relay version change) into
out/relay/wsl beside the other WSL-only guest bundles. The three WSL bundle
resolvers share one candidate list.

* fix(wsl): refuse old glibc before downloading, and keep the shared download per caller

The pinned Node runtime needs glibc 2.28, so a distro below the floor is
refused before any download with a message naming both versions, as SSH
hosts are. The shared download again owns its own deadline and each caller
waits on its own signal, and the OpenCode reader keeps its architecture
error text.

* fix(claude): bound each WSL guest operation and run the helper through the WSL runner

A cached guest no longer carries its 180 s preparation deadline into later
requests. The helper runs through runWslProcess (stdin payload, WSL_UTF8),
the distro is confirmed running once per preparation and once per request,
a failed `claude --version` probe continues with an unknown version like
native setup, the helper resolves from the WSL bundle dir, and the guest
entry decodes stdin once so split UTF-8 survives.

* test(claude): cover WSL profile pre-trust routing and its deadline

* refactor(claude): drop WSL refresh cleanup that the failed publish's withdraw already does

* fix(claude): catch rollback failures only when profiles route the selection

With the gate off, select and remove surface the rollback error exactly as
before; only profile routing logs it and keeps the original error.

* fix(claude): give every WSL pane a guest-relative Claude profile pointer

WSL panes now always carry `~/.local/share/orca/claude-profiles/selected-wsl`,
which the bash/zsh and fish claude functions expand against the guest $HOME
at each invocation, so a pane opened before Orca has met the distro still
follows the selected account instead of falling back to ~/.claude. Absolute
pointers are untouched, PowerShell is unchanged, and a missing pointer file or
profile still refuses visibly. CLAUDE_CONFIG_DIR is set at spawn only when the
selection resolves without a guest call.

* test(claude): assert a missing guest-relative pointer refuses with a visible message

* test(claude): type the WSL runner mock in the transport test

* fix(claude): route only WSL distros that hold an Orca account, and re-derive their publish

A WSL distro is routed only while host settings hold an Orca Claude account
for it, decided from settings with no guest call. An unrouted distro behaves
as before profiles: its panes get no pointer or profile env, and a Claude
launch is System Default with no guest prepare. A distro that loses its last
account has its pointer withdrawn best-effort so older panes stop launching
the removed account.

A routed distro without a current publish (for example stopped at startup)
gets one non-blocking background publish from its next pane spawn, coalesced
per target; its failure stays that distro's issue and a later success clears
it. A late setup result from an older publish no longer replaces the newer
selection's verification. The owner contract moves to its own module so the
routing service stays under the size limit.

* fix(claude): read WSL profile history in native chat and adoption only in running distros

Native chat resolves Claude transcripts from host roots first and reads WSL
profile roots only after a miss, filtered to running distros like Codex's WSL
homes. Structured adoption candidates go through the same filter.

* fix(claude): target registration rollbacks and keep their errors in profile mode

A failed add or re-authentication rolls back only the account's own target.
With profiles, a failed re-authentication rollback is logged instead of
replacing the original error; with the gate off both behave as before.

* fix(claude): spell the guest pointer location once and keep set -u safe

The guest helper, the withdraw script and the pane pointer all derive from
one home-relative constant, and the posix claude function reads ${HOME:-}
so `set -u` with HOME unset refuses cleanly instead of aborting.

* test(claude): cover the IPC preflight and daemon WSLENV paths for WSL profile env

The renderer preflight is tested for wsl.exe and Windows shells with a \\wsl$
cwd (which always launch wsl.exe) and with the gate off, the daemon launch
plan imports the pointer and profile home without a WSLENV flag, and the
Windows launch test uses the guest-relative pointer production sends.

* fix(wsl): report why the guest runtime failed, with download context and trimmed stderr

The install's promote output is classified with the SSH classifier, so a
self-test failure shows the exit code and the loader's words (for example a
missing libstdc++ on Alpine) and a security-software change is named. A failed
runtime download says it was Orca's Node runtime for WSL, while a checksum
mismatch keeps its own text. Guest stderr is trimmed before it reaches a
refusal message.

* test(claude): pin that pointer retirement never runs for host targets or with the gate off

* test(claude): give the routed WSL preflight fixture its required authMethod

* fix(claude): let the pane-triggered WSL publish repair a distro stopped at startup

"Distro not running" is now a typed refusal: it never withdraws the pointer
(the distro's last pointer cannot be stale, and a withdraw racing the boot
could delete a valid one) and never records a distro issue. The background
publish a pane fires now waits a few seconds for the pane's own spawn to boot
the distro, probing three times, and is dropped silently and re-armed if the
distro stays down. It joins any publish already in flight for that target
instead of preparing the guest a second time. Per-target generations and
pointer-write ordering move to ClaudeProfilePointerQueue so the routing
service stays under the size limit.

* fix(claude): remove the last selected WSL account without a guest publish

With profiles, removal writes the account list and the selection in one
update, so a distro losing its last account is already unrouted when it syncs
and its pointer is retired best-effort. Removal no longer needs the distro to
be running or able to run Orca's runtime. The gate-off order is unchanged.

* fix(claude): keep native chat's legacy Claude roots first and unfiltered

Only roots added by WSL profiles are read after a miss and filtered to
running distros; a host CLAUDE_CONFIG_DIR on a \\wsl$ share is searched first
and unfiltered, as before profiles.

* test(claude): cover stopped-at-startup repair, launch join and last-account removal end to end

* test(claude): assert no running probe before the pane has had a turn to boot the distro

* fix(claude): let user-initiated profile work boot an idle-stopped WSL distro

WSL distros idle-stop on their own, and the legacy path boots them with its
spawn or \\wsl$ write. With profiles on, a Claude launch, a select, a remove,
a failed-change rollback and the retire after removing a distro's last
account now skip the running pre-check and let their first bounded guest
command (`wsl -d <distro> --exec ...` through runWslProcess) boot the
distro. They refuse only if that command fails, with wsl.exe's own reason,
for example a distro that does not exist. Startup, the pane-triggered repair
and the history readers keep the running pre-check and its typed refusal, so
background work never boots a distro. With the gate off nothing changes.

* fix(claude): let startup join a launch or select already publishing a WSL distro

Startup no longer overtakes a user's in-flight publish for the same target,
so a launch that is booting an idle-stopped distro is not handed startup's
"not running" refusal.

* fix(claude): remove accounts of a WSL distro that no longer exists, and name the helper once

wsl.exe's own failures (exit 0xFFFFFFFF, empty stderr, the diagnostic and its
WSL_E_* code on stdout) are now read by one shared reader used by the git
runner and the WSL profile transport, so profile refusals show wsl.exe's
message. WSL_E_DISTRO_NOT_FOUND becomes ClaudeProfileHostMissingError: with
profiles, removing an account from a distro that no longer exists keeps the
removal and logs a warning, while select and launch still refuse visibly.
The helper's file name is defined once in shared/relay-artifacts.ts and used
by the relay build and the transport.

* fix(claude): give plain fish tabs the claude function through the codex hand-off

Main now gives a plain fish tab Orca's codex function through a vendor_conf.d
snippet instead of a -C init. The claude function only rode the -C path, so a
plain fish tab would not re-read the account selection per invocation once
profiles are on. Define it at the first prompt beside codex; it stays empty
while the profile gate is off.

* fix(claude): share personal rules, themes, workflows and keybindings into account profiles

A managed account launches Claude with its own config folder, so user-level
rules/, custom themes/ (which a shared `custom:<slug>` theme points at),
personal workflows/ and keybindings.json silently stopped applying. Link the
three directories like skills and commands, and copy keybindings.json with the
same edit-preserving ledger as CLAUDE.md. routines/ stays unshared: routines
belong to the claude.ai account and the folder holds per-run state.

* test(claude): wait for the running child to read its account before switching

The test switched the selection after a fixed 20 ms, so under load the backgrounded claude
had not yet read the pointer and picked up the new account. The stand-in now marks when it has
started, and the test waits for that mark (bounded) before switching.

* fix(claude): accept WSL setup warnings for every shared Claude file

The guest reply schema listed CLAUDE.md by name, so a warning about the newly shared
keybindings.json would have rejected the whole reply. It now takes the shared-file list
from provisioning, like the shared folders.

* fix(claude): import the personal CLAUDE.md into account profiles instead of copying it

Claude also loads ~/.claude/CLAUDE.md as a parent folder's memory for any project under home,
so a copied account CLAUDE.md made every such session read the user's instructions twice
(checked live with Claude 2.1.288). An @~/.claude/CLAUDE.md import resolves to the same real
file, which Claude loads once from home, from projects under home and from folders outside it.

* refactor(claude): simplify account profile setup toward the prior art

- Windows keeps each account's history private; drop the hardlink, link
  record and conflict-copy machinery that only Windows reached.
- Share hooks and statusLine as ordinary settings keys: Orca writes the
  same entries into every folder, so the installer finds them present.
  Drops the Orca-entry carve-out, the per-profile statusline follow
  logic and its marker.
- Unreadable ledger is just an empty ledger.
- Share from the user's own CLAUDE_CONFIG_DIR when they set one (marked
  so Orca's injected value is never mistaken for it), and refuse a
  profile at or around it.
- Pin the one canonical profile path spelling in a test.

* refactor(claude): route launches through one account router, superset-shaped

Replace the routing service, owner interface, setup worker thread, reader-root
merging, persisted launch account and capability string with one
ClaudeProfileRouter: the pointer is written first and setup runs best-effort
after it (superset's order); a missing pointer means System default.

The claude shell function re-reads the pointer on every launch, refuses only
a selected account whose folder is missing, and prints a note when the user's
own CLAUDE_CONFIG_DIR overrides the selected account in that terminal.

Still dormant: claudeProfileRoutingEnabled() is false.

* test(claude): type router test settings instead of casting

* fix(claude): run account setup on a worker thread, never Electron main

publish() writes the pointer and starts setup in the background, so neither
startup nor an account switch blocks on a history merge. Each setup runs in a
one-shot worker (the profile-state backup worker's pattern); one setup per
account at a time, reused by later requests. A launch waits only for a folder
that was never set up, and refuses with a clear message if that setup fails.

* fix(claude): do not await the synchronous pointer publish

* refactor(claude): route WSL distros through a small guest router on the Step 2 shape

Replaces the WSL owner/transport/guest-inspect stack with ClaudeWslProfileRouter:
publish writes the guest pointer with one sh command and kicks Step 1's setup
best-effort; prepareLaunch checks the folder over the distro share and waits only
for a never-set-up folder; preparation returns main's WSL shape, so trust, rate
limits and readers need no new code. Setup runs as Linux in the guest on Orca's
pinned Node via a bundled helper (argv in, exit code out), without hooks.

Restores OpenCode's WSL runtime prep, git's wsl-host-failure, wsl-runner,
workspace trust, readers and account selection/registration to Step 2.
Names the guest pointer per Orca build so dev and packaged never share it.

* test(claude): give the routing launch test the merged resolver deps and handle shape

* test(claude): type the WSL routing mock's original() without an inline import()

* fix(claude-accounts): dedupe merged prompt history, drop drained copies, link setup folders by path

- Prompt-history drain appends only lines the shared file lacks, so a purge never re-adds lines.
- A set-aside history copy whose saved offset reaches its end is deleted on the next run.
- Setup folders link to the default home's own entry, not its resolved target.
- The profile gate and folder creation run once, in provisionClaudeAccountProfile.
- installHooks receives only configDir; drop a duplicate test key that fails CI.

* fix(claude-accounts): refuse a routed resume whose transcript is in another account; zsh claude function; setup timeout

- With account routing, a chat resume checks its transcript is in the launch folder; a missing one
  with a stored leaf refuses with historyInOtherAccount instead of starting fresh.
- The launch folder of a selected account comes from prepareLaunch(); the resolver stays for System default.
- zsh panes get the claude function like bash, fish and PowerShell (empty while routing is off).
- The setup worker is terminated after 60 s so a later launch can retry.
- Document that the setup marker means setup started, not finished.

* fix(claude-accounts): write the WSL account pointer before a launch returns; one relay bundle candidate list

- prepareLaunch awaits writePointer, so a missing or stale guest pointer cannot run another account.
- Startup's WSL republish runs inside serializeMutation, like rollback.
- relayBundleCandidates takes 'wsl'; the hook relay, browser relay and Claude helper use it, and
  wsl-relay-bundle-dirs.ts is gone.
- One setup-marker path helper for host and WSL; the guest pointer path is home-relative and only
  the pane value carries '~/'; drop a no-op esbuild external.

* fix(claude-accounts): refuse a routed resume only when the transcript is found in another folder

A transcript found in no known folder keeps the old stored-leaf resume.

* fix(claude-accounts): a WSL launch writes the pointer for the selection current at write time; bound the pointer read

A selection made while a launch waited on setup was overwritten by the launch's stale account.
A hung \\wsl.localhost read no longer stalls startup's serialized publish.

* fix(claude-accounts): record installed hooks as Orca-shared; skip symlink tests on Windows

After Orca installs its hooks into an account, record the account's hooks in
the settings ledger so a later run can still bring the user's own hooks in.
Tests that create real symlinks now skip on Windows.

* fix(claude-accounts): trim the which-account file in the PowerShell claude function

Co-Authored-By: Claude <noreply@anthropic.com>

* test(claude-accounts): spell the user's own config folder as an absolute path on every platform

Co-Authored-By: Claude <noreply@anthropic.com>

* test(claude): skip the POSIX-only WSL profile test on Windows

A WSL profile's data root is a POSIX path, so building one from a Windows
temp dir fails the absolute-path check there.

---------

Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
2026-10-07 01:25:41 -04:00
Brennan Benson 2f49377425 feat(native-chat): Grok as a structured chat over the Agent Client Protocol (#25225)
* Leave a stopped turn's running tools to the agent's own end

When another writer settles the open turn (a person's Stop), the assembler
now only stops that turn's text and cancels its pending prompts. Running tool
calls stay the agent's: a progress update or completion it reports after the
Stop lands as reported, and whatever is still running settles at the agent's
turn end for that turn, the next turn's open, or the session's end.

An agent's end for an earlier turn while a newer one is open no longer clears
the open turn's activity line or ends its anonymous reply. An unnamed end right
after a Stop ends the stopped turn instead of being dropped. The test rig's
restart no longer writes the dead assembler's window text, matching dispose.

* Pin that a stopped turn's running tools hold budget until the agent's end

* Type the stopped turn's tool progress update as a tool body

* List every event the assembler hands to the decision step

The type-aware lint requires an exhaustive switch with no default case.
Also retitle a Stop test to say what it asserts.

* End a running call as its turn's journal row ends

A call still running when its turn ends takes the state of that turn's
row: a row another writer settled first (a person's Stop) stands, so its
calls read interrupted whatever the provider's later end reports. The
no-ending path that settled calls from the Stop row is gone, since a Stop
now leaves running calls to the provider. Adds the two Spanish strings.

* Say why a Grok turn failed, and keep task rows in Grok's own words

A failed Grok turn ended with no reason on screen: the translator dropped
every copy of Grok's message. The failed turn now gets one status row in
Orca's existing "provider did not accept this message" words with Grok's
reason, read from whichever copy arrives first (the given-up retry, the
turn's end, the prompt's completion notice, or the prompt's error answer);
later copies only fill a reason the row still lacks.

A running background command no longer reads "Background task <id>
started": a task's summary is mapped only once it has settled. A monitor
stays a monitor when the agent reads its output: a frame that names no
kind keeps the known one, and a "[monitor" command is a monitor.

A prompt's turn is marked started, so a late frame for an ended prompt
neither reopens it nor becomes the active turn. A tool's turn is held in
one place at a time.

* Read a monitor from Grok's exact output prefix

* Word a failed Grok turn in Grok's own text, not as a refused message

A turn that started and then failed was told "The provider did not accept
this message", Orca's sentence for a message refused before its turn. The
row now reads as a Codex turn-ending error does: an error status row with the
provider's own words. With no words, the dialect names the failure ("Grok
ended this turn with an error." / "Grok usage limit reached."), else the
agent's display name does.

* Settle a stopped turn's running call as its turn row ended after a restart too

The restart sweep ended every running call by the death evidence alone, so after
a person's Stop with no proof the child died the call read failed under a turn
that read interrupted. The sweep and the live dead-generation settlement now ask
the same rule the assembler does: a call in a turn already settled ends as that
row ended; only a turn still running leaves its calls to the evidence.

* Keep the dead-generation settlement under the line cap

* Register the ACP schema verify step in the PR preflight phase test

* refactor(agent-session): one required agent registry; declarations admit what they claim

A StructuredAgentRegistry, built once from the {definition, adapter}
registrations, is now a required host dependency and the router routes with
it. The adapter interface loses its optional router-only capabilities?() and
definition?(); every reader (options at rest, thread goal, rewind, the
/compact handover) asks the registry. A live session still narrows rewind
through the adapter, and the host combines declared and narrowed in one
helper. The registry refuses a registration that declares compact, a thread
goal or rewind without the adapter method behind it.

/compact is admitted by the declared capability, so an agent that declares
compact:false gets the commandRefused fact instead of a thrown error.

Also: the cut-turn notice names the agent from the catalog, create-support
builds the account home through agentSessionAccountHome, and the turn
status text older clients read names the session's agent instead of
defaulting to Codex (Claude/Codex text unchanged).

* Read ACP permissions, session events and prompt errors through the protocol client's own types

The translator now reads a permission request with the client's lenient reader, a session update
with its session-event reader, and takes only the agent's own error answer as a failed prompt's
reason, so an Orca-side error never reads as the provider's words. Tests cover protocol values
newer than this build.

* refactor(agent-session): the router applies the declared rewind itself

The router holds the registry, so its rewindSupport answers the owner's
declared rewind narrowed by the adapter, in one place no reader can bypass.
The host-side combining helper is gone; readers ask the adapter they hold.

* test(agent-session): register the agents the merged-in tests now need

The adoption replay test builds its host over this build's registered agents
instead of an adapter definition method, and the stored-form test passes the
stored agents the record guard now requires.

* chore(agent-session): the registry is the one lookup; drop the test-only empty capability record

* fix(agent-session): keep "dismiss all" restart offers dismissed for the desktop

The desktop renderer calls the restart methods as a paired runtime client
without the registered-agents capability, so it was handed a Claude/Codex
audience and its "dismiss all" took the scoped path: no dismissal fence and
no clearing of unwritten teardown witnesses, so a late teardown write could
bring a dismissed offer back.

The audience is now derived against this host's registry: a client that can
show every agent the host registered gets none, and dismisses exactly as
before (one fence for every offer). A client that truly cannot show some
agent still gets a final dismissal for what it sees: each cleared session is
fenced on its own (bounded, superseded by a dismiss-all fence) and only the
witnesses of agents it sees are dropped.

One predicate now answers which agents a client renders, for both tab
projection and restart offers; the Claude capability rule is folded in.

* fix(agent-session): a changed agent definition never hides that agent's chats

A record was readable only if every handle used the transport its agent's
definition declares today, and its account variable was declared by some
registered agent. A later build that drives the same agent over another
protocol, or under a renamed variable, would have set every existing chat of
that agent aside: no tab, no history, no error.

Readability now asks only that the record agree with itself: its agent is
registered, every handle shares one namespace owned by that agent, and the
account variable is a well-formed name. Whether this build can drive it (the
chain's transport is the one the agent speaks, and the account variable is
that agent's own, since it becomes the child's environment) is decided in
performAttach, the one admission every start of every agent passes through,
and refused there as hostUnsupported. A record pinning another agent's
variable is refused the same way rather than passing because some agent
declares it.

* refactor(agent-session): each agent's registration says where it runs and which account it pins

createSupport, create and the model catalog's account read still chose a
location rule and an account resolver by name (Claude, Codex, else none),
so a new agent would have needed a third branch beside the registration
list that already decides storage, routing and publication.

Each runtime registration now carries supportsLocation(location) and
resolveAccountHomePath({ launchEnv, location, purpose, workspacePath }),
and Claude's and Codex's rules move into their entries unchanged: a read
still syncs no home and starts no bridge, a Codex launch still trusts the
workspace first, and WSL locations resolve as before. The list exists
before the host is built, so none of these reads installs the host, and an
agent the list does not hold is answered no without opening the journal.

* test(agent-session): drop a duplicate registry key; keep fences out of the capsule state

A spread input already carries the registry, so the explicit key was
overwritten (TS2783). The session-fence helper now returns only the fence,
not the capsule state it was given.

* fix(agent-session): a scoped dismiss-all persists no per-session fence

The per-session dismissal fence added a forever-persisted capsule field
with an arbitrary cap that served no caller: only the local desktop calls
the restart methods, and on a host whose agents it all shows it gets no
audience and takes the unscoped, fenced dismissal. A scoped dismiss-all
(a caller that cannot show some registered agent) keeps that agent's
offers, clears only its audience's unwritten witnesses, and writes no
fence; a late write from this process is serialized behind it. The field
never left this unreleased branch.

* fix(agent-session): refuse an attach whose agent is not the session's own

The start check judges the record's agent while the router starts the adapter of params.agent, and the attach wire let the two differ. Admission now refuses agent !== provider as requestMalformed, so the agent checked is the agent started. Every host-built attach already sets them equal.

* fix(agent-session): offer to start a chat only when the start would accept it

The restart offer, a failure's retryable flag and the pre-send check asked only whether the adapter runs the chat's location, while the start also refuses a record this build cannot drive; such a chat was offered Resume and its Retry failed forever. One host predicate, hostCanStartRecord (location support and agentDrivesSession), now answers all four; adapterSupportsRecord stays the reader gate, and an undrivable chat's offer is kept, not retired.

* fix(agent-session): the model-catalog probe uses a record's account only if this build drives it

A record-scoped catalog read started the agent's lister under the record's account path whatever variable the record pinned. It now uses that path only when agentDrivesSession holds, and otherwise resolves the account as for a read with no record.

* refactor(agent-session): the record store admits agent ids; comments say where transport is checked

The store only ever asked whether an agent is registered, so its admission list is now the registered ids; the definition is the one home of an agent's transport and account variable. Comments that said the record store checks transport, and one that cited a nonexistent function, now point at agentDrivesSession and the attach admission. The isPersistedAgentSessionRecord(value, agents) call shape and the test fixture the cross-version probe imports are kept.

* docs(agent-session): the record store admits the registered agents' ids

* refactor(native-chat): Grok's registration declares where it runs; ACP no longer borrows Codex's location rule

The rule a self-supervised agent child runs under (this machine, no WSL, Windows only with process
start-time proof) is its own module that Codex and the ACP adapter both use. Grok's registration
takes the full account-home resolver signature, and D3's tests build hosts with the agent registry.

* fix(native-chat): Grok follows the ACP runtime's request contract and the managed process's close

A request the agent or a Stop cancels is answered with the agent's own cancelled reply by the code that
owns it (the runtime no longer answers a silent handler), so a Stop needs no separate decline pass. A
permission answer still being saved when the agent stopped waiting is reported unconfirmed, since the
protocol already answered it cancelled. Cancelling the agent's own turn is the plain cancel. Request
rows are matched under their generation-scoped ids. A refusal's reason comes from the dialect's wording
path. The child drops its own stderr tail and close policy for the managed process's, and a close
whose process tree was not proven gone is reported as the adapter contract asks.

* fix(native-chat): a Grok chat Orca already holds resumes without writing what Grok replays

A chat with a saved Grok session reattaches with session/resume where the agent offers it, else
session/load. Either way the call runs inside the translator's load window, so what Grok sends while
it reattaches (its saved exchange, a task the dead process left running, ended by the restart) opens
no turn and writes no row; only context usage reads on. A reply an Orca or Grok crash cut short is no
longer completed from Grok's saved history: it reads like a Claude or Codex chat's, with the existing
notice. The attach window also closes after a failed attach, and a created session that session/resume
reports missing is replaced like one session/load reports missing.

The replay reconciliation is removed: the lane no longer reads the journal, and D3's replayed-input
grammar test and completed-turn check in the assembler go with it.

* refactor(native-chat): a failed Grok reattach needs no window close of its own; its lane is replaced

* test(native-chat): D3's merged tests use the shipped declarations and the launch options main requires

* fix(native-chat): typecheck fallout of the base merges; any agent's empty chat is reusable

Main's idle-empty-chat lookup and launch join now take any registered agent, as the rest of the
launch path does. The refusal check moved into the prompt turns and the prompt-block conversion beside
the turns that send it, keeping both files in their line limit.

* fix(native-chat): a Grok Stop ends the process once Grok settles its turn; the next send resumes

Grok's session/cancel ends only the running turn: work it already moved to the background keeps
running and can begin a turn of its own after the person pressed Stop. Stop is now a session
boundary, as it is for Claude: the cancel answers open requests and lets Grok end the turn, the host
waits a bounded grace for that, then ends the process; the next send relaunches and resumes.
The adapter's own bounded close of a turn Grok began is gone. Its named-turn check stays: the host
ends the session unless the provider declines a Stop naming a turn that has since ended.

* test(native-chat): a Grok Stop ends the process only after Grok answered the cancel

* fix(native-chat): Steer on a Grok card cancels the running prompt, then sends it

A send that reached Grok while a prompt ran was held in the adapter until that turn ended: Steer
on a queued card took the card out of the host's editable queue and meant 'send after this turn'.
It now cancels the running prompt (session/cancel; the session stays) and sends as the next prompt
once Grok answers the cancel, as the common pattern does; a steer behind another cancels it in
turn, so the last one runs. The adapter holds a send only while that cancel lands, so its general
held-send queue and its holdsDispatch report are gone (every send it holds has its turn open in
the journal). An older client's mid-turn send takes the same path. capabilities.steering is
unchanged and still unread.

* refactor(native-chat): a close or Stop cancels a start through the acquire's own abort signal

The host owns the acquire it runs, so it now owns its cancellation: each attach's acquire gets an
AbortSignal, aborted from outside the session's queue by a close and by a Stop admitted now (the
same admission rule as before). The optional abandonStart adapter hook, the router's fan-out to
every adapter and the ACP adapter's session-keyed start map are gone; the ACP adapter keeps an
unkeyed set of starts only so quit can prove their children gone, and keeps a failed start's
unproven child until its exit is proven.
The hook also let a later close ask that child again. The host now does that from state it holds:
a close of a chat with no live child whose record still names an owner process with no death
evidence asks the adapter to release it. The answer is not recorded as proof (the lease probe
does that), so an owner pid an earlier Orca left is never killed or marked gone. Claude and Codex
ignore the signal and hold no such child; their release is a no-op (tested).

* fix(native-chat): a Grok crash that closes stdout before its exit still ends with Grok's last words

On macOS and Linux the agent's stdout ends before its exit is observed, with or without the
supervisor's EOF forwarding, so the connection's loss closed the journal first and its error text
became the session's ended reason, dropping Grok's stderr. The reason is now read at the proven
exit: the agent's last words when it left any, else why the connection closed. The failure already
carried them. Comments that assumed the exit comes first, that early frames past the cap refuse the
start, and that dispatch re-checks image support are corrected.

* fix(native-chat): nothing Grok sends while a held chat reattaches is written, marked as replay or not

The reattach window relied on the dialect's replay verdict, and Grok's frames read as live unless
they carry isReplay, so an unmarked chat frame during session/resume opened a turn that never
ended. D3 now marks every frame inside the window as replay before the translator reads it, so the
translator keeps only context usage whatever the agent marked; options and commands are still
adopted. The translator's load semantics are unchanged.

* test(native-chat): a Stop after a resume finds no turn an unmarked old reply opened

* test(native-chat): a resumed Grok chat keeps its last context reading; the resume refreshes only the window

* test(native-chat): a Grok background task a Stop ended reads as stopped reporting

* refactor(native-chat): quit's stop of each start answers through one promise kind

* fix(native-chat): quit aborts every start the host has in flight before draining attaches

A Grok that never answered its handshake held quit until the start's own 60 s bound, past the
20 s quit deadline. The host's teardown now aborts each in-flight acquire (and any the drain
still begins), so the adapter's own quit controller and its map of starts are gone: a start
has one canceller, the host's signal.

* fix(native-chat): a Grok start's abort stops reaching its child once the start has returned

The listener stayed on the host's signal until the attach finished committing, so a Close in that
window killed the now-live child behind the host's back and it read as Grok crashing. The start
now detaches it when it ends; a later Close goes through the session's own stop.

* fix(native-chat): a close or Stop during any attach phase stops the start before it launches

The attach began its abort controller only after reconciling leases, resolving recovery and
probing the previous owner, so a close or admitted Stop in those phases reached nothing and Grok
launched anyway. The controller now begins first, and the acquisition checks it before asking the
adapter to start.

* test(native-chat): a close during the attach's owner probe asks no adapter to start

Also renames the close test after the hook it no longer exercises.

* test(native-chat): a close's re-ask closes a Claude or Codex child a failed cleanup left

The re-ask is not a no-op for them: when the adapter still holds the child its cleanup could not
prove gone, the close stops it again as a requested close, and Claude persists the handle of the
conversation it ran so the next send resumes it. Corrects the tests' and comment's wording; the
close awaits the re-ask, bounded by each adapter's kill ladder.

* fix(native-chat): Steer during a turn Grok began itself cancels it and sends once it ends

A send while Grok ran a turn of its own (a background task waking it) went straight to Grok, which
queued it behind that turn where Orca could no longer withdraw it, while Stop treated the same turn
as the running reply. The send now waits as a steer, the turn is cancelled once, and the message
goes when the turn ends; a Stop withdraws it and an exit rejects it as never sent.

* test(native-chat): a steer whose cancel Grok never answers ends Grok and is rejected as never sent

Pins the bounded steer cancel kept from the runtime: past the bound the connection closes, the
running reply reads unverifiable, Grok's end reads as its exit, and the waiting steer is rejected
as never sent.

* fix(native-chat): a Grok crash stays a crash when a stop lands before its exit is proven

After the connection broke and the close could not prove Grok's exit, any later stop Orca asked
for (the next start, a Stop, a Close) marked the child as closed by Orca, so the crash read as a
requested close and Grok's last words were dropped; a send meanwhile was recorded unconfirmed.
The connection loss now decides the cause, and a send on that session is rejected as never sent.

* test(native-chat): fixtures this PR's registered Grok and desktop capability made stale

CI's unit shards failed on tests outside the PR's own lists. Each encodes something this PR changes
on purpose: Grok is now a registered agent (the seam test's unregistered agent is now Cursor); the
desktop now advertises registered agents (the restart-offer tests' older client drops that
capability explicitly); the attach context carries the start's abort controllers (the forget-status
double gains them); and the ACP real-host test rig sends to the host directly (listed beside the
other real-host rig in the send ratchet).

* fix(native-chat): a start quit stops is not the queued message's start failure

With quit now aborting a start it would have waited for, the delivery step recorded the aborted
start as the message's failure ("couldn't restart"). After quit has stopped delivery, the step
leaves the message to quit, which settles it as a close does ("The chat closed before this message
was sent."). The test that pinned quit waiting for that start and stopping its child now pins that
nothing is launched behind quit.

* fix(native-chat): a message sent after a Stop or close aborted a start gets its own start

A start the host aborts (an admitted Stop, a close, or quit) returned its refusal to the delivery
loop, which then rejected whatever was queued at that moment with "couldn't restart", including a
message the user sent after the Stop. The attach now reports that the host aborted it, and the loop
re-derives from the journal instead: what the Stop or close withdrew is already settled, a message
accepted since gets a start of its own, and quit's next step stops the loop. This replaces the
quit-only carve-out with the same rule for every abort and every agent.

* test(native-chat): the message sent after an aborted start is answered, so no settlement outlives the test

* fix(native-chat): a Grok model pick Grok never answers no longer holds Stop or Close

The pick runs on the session's queue. It now registers in the host's out-of-queue
abort registry beside a start, so a close, an admitted Stop or quit abandons it, and
the ACP adapter bounds it at 30 s like Claude and Codex. A late answer is still adopted.

* fix(agent-launch): a phone's launch opens a terminal for an agent whose chat it cannot show

agent.launch now reads the caller's capabilities by the rule tabs and restart offers
use (clientRendersStructuredAgent). A phone without registered-agents.v1 gets Grok as
a terminal again, as on main; the host's own callers and desktop clients are unchanged.

* fix(acp): strip every agent hook variable from the ACP child, from the shared list

ACP_CHILD_ENV_TO_DELETE was a second copy of the hook runtime keys that missed
ORCA_AGENT_HOOK_TRANSPORT; it now spreads AGENT_HOOK_RUNTIME_ENV_KEYS beside the pane
identity keys.

* refactor(native-chat): the mutation context carries the provider-wait registry itself

Keeps the host file within its line limit; one field instead of two closures over it.

* fix(agent-launch): agent.launch.v2 still vouches for Claude and Codex chats

The caller rule from the previous commit also turned Claude and Codex into terminals
for a client advertising only agent.launch.v2, whose contract says it opens a chat
(mobile retry-authority tests). Only an agent beyond those two now needs the client to
read it (clientRendersStructuredAgent); the test fixtures go back to what they were.

* refactor(native-chat): drop saved-history adoption from the timeline assembler

The common pattern discards the history a provider replays while loading a
session, so the assembler has no use for an input.history event.

* refactor(acp): drop session/load history adoption from the translator

The common pattern discards the history an agent replays during session/load,
keeping only what it says about the context window. Remove the adoption path
(acp-history-adoption.ts, the adopt option, and the historical background-task
liveness rewrite it fed) so load replay is always dropped except usage.

* refactor(native-chat): a pending input is only Orca's send now

Review follow-up to the adoption removal: drop the comment naming the
provider's saved message, and make requestedAt required since every pending
input comes from input.accepted.

* test(acp): keep the task-result status table on live frames

Review follow-up to the adoption removal: the result-status mapping was only
tested through adopted history, so run the same table on live frames, and
cover an unmarked task notice during a load being dropped.

* test(acp): a frame helper for a shell command Grok is running

* fix(acp): a Grok crash settles through the host's provider-exit batch, scoped to the turn it ended

A Grok crash ended the journal unverifiable before the adapter reported the exit, so the host's
provider-exit settlement found no running turn and wrote nothing: the adapter's failure (with
Grok's last words) never reached the journal, and a later stale-session pass wrote a bare,
thread-scoped cut-short row, so the partial reply was not folded as Claude's and Codex's are.

At a proven exit the ACP lane now ends its running turn interrupted at the exit instant, as the
host's exit contract expects of a child's own translator (Codex's does the same). When Grok's
stdout closed first (every POSIX crash), the turn is unverifiable only until the exit is proven:
the host's provider-exit settlement now takes the exit as proof naming the child's fence and
revises what that child left unverifiable in the same batch, with the turn-scoped row and the
adapter's failure. Claude and Codex write no unverifiable turn of a live child except a command
whose hand-off is in doubt; that turn is now revised at the exit instead of at the next open.

* test(acp): a crash seen first leaves the host no Grok turn to revise

* refactor(native-chat): what a gone generation left unfinished gets its own module

The settlement file passed 300 lines with the exit-proof revision. The unfinished-work reads
(capture, interrupted-by-the-exit, in-progress) are their own concept and move out unchanged,
apart from the exit proof they now take.

* refactor(native-chat): a watched exit revises what its child left unverifiable without reading Stop marks

An exit's own instant is the turn's end, so the revision needs only each row's fence: the
settlement's journal type gains itemFence alone, and the host test fakes say so.

* test(native-chat): drop the duplicate itemFence on the fake that already had one

* test(claude, codex): an exit whose stdout ended first still reports as it always did

The provider supervisor now ends Orca's stdout when the agent's ends, so on every crash EOF
arrives before the exit is seen. Claude's and Codex's connections report nothing at EOF and
report the exit, with its usual reason, once it is seen.

* fix(acp): reopen a chat with session/load, as the common pattern does

An agent that offers both now reloads its session instead of resuming it; the
reattach window still discards what it replays except context usage.

* fix(acp): drop the 60 s handshake bound; an abort fails the start's waits at once

Neither common design bounds an ACP handshake: Close, Stop and quit end a start
that never answers. The abort now also closes the connection, as a kill there
does, so the start settles even before the child's exit is proven. The
host-stopped start refusal only this bound produced goes with it; the idle
sweep keeps its words.

* fix(acp): a Stop naming an ended turn follows Claude's rule

It still stops nothing while another turn is live, but in the gap before a
follow-up's turn opens, which no client can name, it now stops what is in
flight and the session ends, as a Claude Stop does.

* fix(native-chat): a close no longer re-asks a failed start's unproven child

Neither common design retries that stop at Close, and Orca's Claude contract
re-asks only at the next start and at quit. The ACP adapter keeps the child
until its exit is proven and asks it again there, as Claude does.

* fix(acp): a message sent during a turn the agent began itself goes at once

Both common designs send it straight to the agent with no cancel; only Orca's
own running prompt is steered (cancelled, then re-prompted).

* test(native-chat): dismiss-all through a remote client's audience keeps a newer Orca's offer

Uses an audience production sends (one that cannot show every agent), per review.

* fix(acp): launch Grok as `grok agent stdio`, without the update and leader flags

The common pattern passes neither --no-auto-update, --no-leader nor
GROK_DISABLE_AUTOUPDATER; full access still adds --always-approve.

* fix(acp): an agent that ends its stdout, or answers unreadably, is not a lost connection

As in the common pattern, only a broken stdin (or Orca's own close) ends the
agent; one that closed its output but can still be written to stays until a
Stop, a close or its exit. The provider supervisor goes back to its base
content, so Claude and Codex no longer get the forwarded stdout end either.

* fix(native-chat): a person's close joining a failed one still binds the turn its child end cuts

On main every close of the chat writes its own Stop and settle. Here a later close joins the
first and writes no row, and the first's settle closed when its kill failed, so a turn that opened
in between and was cut by the next close read as failed. A person's close joining a person's close
whose Stop opened a settle now reopens that settle until its attempt is done.

Tests: a close whose kill failed still closes its settle; a turn opened between a failed close and
the next reads as the person's cancellation (each fails without its half of the fix).

* test(native-chat): Grok opens as a chat only behind the structured-chat setting

agent.launch and orchestration worker-start read the same setting as the
renderer route; pin both states for Grok on each. The setting's description no
longer names only Codex and Claude, in every catalog.

* docs(acp): generic ACP comments say what holds for every agent, not Grok

Stop ends the session for every ACP agent, as in the common pattern; the
adoption hook comment goes (adoption is not planned); a failed start's child is
retried at the next start or quit.

* test(claude, codex): type the EOF-before-exit test's streams; the supervisor no longer forwards EOF

The Claude test wrote to the child's stdout and stderr through their Readable
type, which the node typecheck rejects; it now holds its own PassThrough
streams. The comments no longer credit the reverted supervisor change.

* feat(acp): a steer's cancel asks once and never ends the agent

The runtime had one cancel: send session/cancel, wait at most 10 s for Orca's prompt to settle,
then close the connection, which ends the agent. A steer used it too, so a slow agent lost its
process just because the person added a message. requestSteerCancel() now sends session/cancel
once per prompt, cancels the agent's open requests and answers later permissions cancelled, and
never bounds or closes: the prompt's own reply ends it and the steer's prompt follows. cancel()
stays the Stop: bounded, then close. A Stop after a steer still bounds and closes. Both cancel
paths move into acp-prompt-cancel.ts over one cancel channel.

* chore(native-chat): keep the record store and recovery capsule under max-lines after the main merge

* fix(acp): a repeated steer shares the cancel in flight; say what the caller owns

Per review: a second steer before the first write lands returns that write instead of resolving
early. The steer's JSDoc says the wait for the prompt's reply is unbounded and that a prompt that
fails instead must not take the steer until the caller rebuilds the session; the Stop's says a
prompt that settles in time leaves the agent for the Stop's owner to end. The steer test now gives
the runtime a handler that would allow: the open permission's signal aborts and the late one never
reaches it.

* fix(native-chat): drop the stopDelivery the A3 merge doubled

* fix(acp): a steer's cancel asks Grok once and never ends it

A steer now uses D1's notify-only cancel. Two messages sent during a reply Grok began itself
cut that reply, as the common pattern does, and then both run; before, the queued first
message could not answer the bounded cancel and Orca ended Grok although Grok answered.
A Stop keeps the bounded cancel and its 4 s grace.

* fix(acp): a permission Grok asks with no prompt of Orca's running is declined

During a turn Grok began itself nobody asked it to act, so the request is answered
cancelled at once instead of opening a card that waits, as the common pattern does.

* fix(acp): a Grok that dies while starting is reported with its own last words

A dying process's stdout ends before its exit is seen, so the start failed as a closed
connection and Grok's stderr was lost. A start whose connection closed now waits, bounded by
the Stop grace (or a Close/Stop), for the exit before it is told.

* test: a Stop after a steer sends its own cancel; drop the import the A3 merge doubled

* test(native-chat): main's Stop-note test builds its turn context with the agent registry

* test(claude): say why the close test's fake child cast is safe

* test(native-chat): build the Stop-opened-turn test's identity and turn context the current way

The test (#25056) landed before the opaque provider handle (#24991), so main still built
the old {kind, threadId} handle; the turn context also needs this branch's agent registry.

* test(native-chat): build the Stop-opened-turn test's identity with the opaque handle

The test (#25056) landed before the opaque provider handle (#24991), so main still built
the old {kind, threadId} handle.

* test(ratchet): require src/main/provider-process now that it has landed

* test(native-chat): keep main's opaque-handle import in the Stop-opened-turn test

Main's #25706 and this branch both added the import at different lines; the merge kept both.

* test(native-chat): keep main's opaque-handle import in the Stop-opened-turn test

Main's #25706 made the same fix as this branch at a different line; the merge kept both imports.

* feat(native-chat): record a fresh provider conversation that replaced one the agent could not restore

A chat whose saved conversation the agent cannot reopen can now continue in a
fresh one: the handle chain records the new conversation as a creation that
replaces the lost one (which, why, and when), keeping every earlier link.

Rows keep a shape older builds read: the stored chain starts at the latest
replacement and carries the earlier links inside it.

* test(native-chat): build this stack's journal identities with main's opaque provider handle

Main's #24991 replaced the {kind, ...} handle with {transport, agent, nativeId}; three test
files from this stack still wrote the old shape. Same lines the downstream ACP branch uses.

* docs(acp): every reattach drops the agent's replay, not only for a chat the journal holds

* refactor(native-chat): store a replaced conversation flat; refuse it where older builds read the row

Older builds only read Claude and Codex records, so the nested stored form
protected rows no replacement can reach while adding a cap mismatch after a
downgrade. Store the chain as held, refuse a replacement in a Claude or Codex
chain until one has a stored shape older builds read, and refuse a supersession
key on a replacement that names no creation in the chain.

* refactor(native-chat): read hosts' structured agents from the app-shell services

Main grew the startup hydration hook to its line limit; the host agents sync is an app-lifetime subscription like the structured session tabs sync beside it, so it moves there.

* Use current provider handles in transition tests

* Use current provider handles in timeline fixtures

* test(native-chat): prove replacement rows survive downgrade and re-upgrade

* Require the ACP directory in the runtime import check

* test(ratchet): require src/main/acp now that this PR lands it

* feat(acp): a saved session the agent cannot reopen continues in a new one, with one warning row

When session/load (or session/resume) of a saved ACP session fails, the chat starts a new session and records it as a creation that replaces the lost one (#25747's 'replaces' link), and writes one warning row that the agent no longer remembers the earlier messages. A created session the agent reports missing is still superseded silently; a signed-out agent or a start that is over (Close, Stop, a lost agent) still fails the start.

* chore(acp): rewrap the acquire header comment

* test(acp): a start closed while the agent reopens fails without opening or announcing a new session

* Let ACP connections own their supervised agent process

* Preserve ACP cleanup evidence and isolate exit observers

* Expose ACP cleanup observations and type the permission fixture

* refactor(native-chat): the registered-agents capability lives in its own module

Main's growth put protocol-version.ts one counted line over its 300-line limit once the capability
was added; like main's other per-feature capabilities, it now has its own module, and importers read
it from there.

* refactor(acp): one connection owns the Grok process and its protocol

D3 now opens each ACP agent through createAcpAgentConnection (ACP-ALIGN #25810): one object spawns the
process on the execution host, owns its stdio and protocol, and reports its proven exit. It is built and
tracked before the handshake, so a start's abort (Close, Stop, quit) still reaches it, and a failed start
keeps that same connection for the next close to retry rather than spawning another process.

Deleted: the spawnAcpStructuredChild wrapper and its test, the raw-stream runtime assembly, the caller's
exit -> runtime.close wiring, the stdout-EOF heuristic (the connection no longer treats stdout EOF as
exit), and the 10 s steer/Stop cancel bound with requestSteerCancel. Reader control maps to
pauseReading/resumeReading; a close is connection.close after the host's existing 4 s Stop grace.

The adapter owns what the protocol no longer does: one session/cancel per running prompt however many
steers arrive (cleared with that send's settlement, retried after a failed write), and a Stop or steer
answers every open agent request the person has not already answered with the agent's own cancelled
reply. An answer already being saved when the Stop lands is sent.

Tests: blocked cancel write never holds Stop's grace, two quick steers send one cancel, a failed cancel
write is retried, a real process exiting while a child holds its stdout ends the session, and the
existing start-abort, retention, crash, connection-loss and reload-failure suites on the new rig.

* fix(acp): Grok signs in on its own machine with its API key or cached sign-in

When Grok reports that it needs authentication, Orca now names a sign-in method on the machine Grok runs
on, read from the same environment Grok was launched with: xai.api_key when XAI_API_KEY is set there and
Grok offers that method, else cached_token when Grok offers it, else none and the chat keeps the existing
not-signed-in refusal. The rule lives in Grok's launch spec; the adapter applies any agent's rule for new
and reopened sessions through the protocol client's caller-named method (authenticate, then retry once).
No new sign-in UI; interactive methods are never chosen.

* fix(acp): the adapter decides which of Grok's requests reach the person

The turn owner now admits every agent request, permission or question, from its own turn state: a
request reaches the person only while Orca's prompt runs and no steer or Stop is cutting it short (a
question may also come from a turn Grok began itself, until a Stop). Anything else gets the agent's
own cancelled reply and opens no card, so a question arriving after Stop or during a steer never
appears. A steer, like a Stop, withdraws the requests already open; an answer already being saved is
still sent. The protocol client's abort-on-cancel path is no longer used: after the connection
change its request signal aborts only when the connection closes.

* fix(acp): a plan Grok proposes shows as a plan, with no approval card

When Grok leaves plan mode it asks the client to approve its plan (x.ai/exit_plan_mode). Orca showed a
blocking 'Approve plan / Request changes' card for it; the common pattern has no such gate. Now the
plan goes into the chat's existing Plan row (the plan-document status row Codex and ACP plan updates
already use) and the request is answered at once with 'abandoned' plus feedback telling Grok to stop and
wait for the person's feedback or a request to implement it in a later turn, so nothing is approved on
the person's behalf. Dialects gain settleRequest for requests answered without asking anyone.

* fix(orchestration): worker-start opens a Grok worker in a terminal, as before

With the structured chat setting on, worker-start decided 'structured' for Grok and then the structured
worker factory (Claude and Codex only) refused it, so the start failed; main opened a terminal Grok
worker. Worker-start now decides with no registered agents beyond Claude and Codex, so Grok gets a
terminal worker as before. agent.launch and the app's own launches still open Grok as a structured
chat. Temporary until structured workers take registered agents.

* fix(acp): a prompt answer Orca can't read ends the turn instead of hanging it

A session/prompt rejection that was not the agent's own error answer (an answer that fails Orca's
schema, or one too large to read) left the turn running: the next message became a steer with nothing
to cancel and was never sent or settled, and Stop waited its full grace. As in the common pattern, any
prompt failure now ends the turn as failed (a failed-turn row without words, since none are the
agent's) and settles the send, so the next message goes. Only a closed connection keeps the send
running, for the connection-loss path to settle.

* fix(acp): send Grok's prompt-identity extension only to agents that echo it

session/prompt carried _meta {promptId, requestId} for every ACP agent, though only Grok's dialect
echoes it (injectedPromptIdentity). Now only an agent whose dialect declares it gets the extension;
other ACP agents get a plain prompt.

* refactor(native-chat): the registered-agents capability lives in protocol-version again, as on main

This reverts 0ef6d21815. That commit moved the capability to its own module only because main's
protocol-version.ts was then one counted line over its limit; main now defines it there itself within
the limit, and main's new restart test imports it from there. Main's test also reads the desktop's
capability list as an older client; on this branch the desktop advertises registered agents, so its
older client is that list without this one capability.

* test(acp): read the sign-in method with a schema, not a type assertion

* refactor(native-chat): composer transport and Stop control in their own modules

Main's rewind change (#19338) brought NativeChatStructuredSession.tsx and use-structured-agent-session.ts
to their line limits, leaving no room for this branch's image-acceptance and unpublished-Stop lines.
The composer's transport (sends, commands, options, image acceptance) moves to
use-native-chat-structured-composer-transport.ts, and whether Stop shows and what it does moves to
structured-agent-session-stop-control.ts. Behavior is unchanged; the runtime cast on the composer's
'local' | 'remote' is now a typed return.

* test(native-chat): read registered agents by id, as main's structuredAgentsReadBy now takes

Main's A3 squash changed structuredAgentsReadBy to take agent ids; this branch's test still passed
{ agent } objects (CI typecheck TS2322).

* fix(acp): a first reopen warns when Grok forgets a chat that exchanged turns

A Grok session the chat created was treated as one Grok never saved, so
when Grok reported it missing on the chat's first reopen, Orca swapped in
a fresh session silently, even after completed exchanges: the person saw
the old messages while Grok had forgotten them.

The launch now counts a created session as never saved only when the
chat's journal, read at the failed reopen, holds no turn of that Grok
session; anything else, an unreadable journal included, takes the normal
path: the fresh session is recorded as replacing the old one and the one
warning row is written.

* fix(acp): a start writes the warning row an earlier attach failure dropped

The row saying Grok forgot the chat was written only into the attach's
deferred sink, while the fresh session's link was saved earlier. An attach
failure, quit or crash in between dropped the row forever.

Every start now derives the owed rows: each conversation the chain says was
lost to a failed restore gets its row unless the chat's journal already holds
it. The row now names the lost conversation rather than the fresh session, so
a row an unused replacement wrote still counts after Grok supersedes it.

* fix(acp): a question during a turn the agent began itself gets its cancelled reply

A question or other card-opening request the agent sends while no prompt of
Orca's runs (a turn it began itself, as when a background task wakes it) now
gets the agent's own cancelled reply and opens no card, the same rule
permissions already follow there. Nobody is waiting on that turn. A plan the
agent shares in it is still shown.

* test(native-chat): import the unfinished-work capture from the module that owns it

Main's reasoning sweep test (#19221) imported it from the dead-generation settlement, which
this branch split it out of.

* refactor(native-chat): keep the session host within its line budget after main's Stop work

Main's #25949 left the host at exactly its 300-line budget, and this branch's acquire-abort wiring
adds one line. The reveal module now comes in as a namespace import, as the host already does for
its other helper modules, and the earlier reorder of two type imports is undone.

* test(codex): move the stdout-before-exit test into its own file

Main's connection test file is at its 800-line test budget, and this branch's exit-order test
pushed it over (CI lint, max-lines).

* fix(acp): cancel a running turn before a close, dispose or quit ends the agent

Closing a tab, disposing a session or quitting while an ACP agent's turn ran
killed the process without asking the agent to cancel first. A requested
close now does what a Stop does: withdraw the agent's open requests and held
steers, send session/cancel once, and wait for the turn to end, bounded by
the Stop's grace (4 s), before closing the process. A close that follows a
Stop sends no second cancel and waits only what is left of that Stop's grace.
An idle close, a lost connection and a sink-failure force close are
unchanged. The cancel-and-wait moves to acp-structured-stop.ts, shared by
Stop and close.

* test(codex): pass resolveLaunchArgs in the stopped-send-order test so typecheck passes

Main's typecheck fails here too: #25721 made resolveLaunchArgs required and
this test (#25051) predates it. Same line, same place as the open main fix
(#25977), so merging main after it lands is a no-op.

* test(native-chat): run the close-aborts-start test on the Node runtime, as its SQLite journal fixture requires

* refactor(native-chat): the Stop control reads the outbox itself, so the chat hook stays within its line limit

* Keep the session host under the line cap after main's two new delegates

Pass the lifetime's conversation opener to reveal directly; it is already passed unbound to the mutation context.
2026-10-06 22:09:26 -07:00
Neil 76d480f808 Fix unsafe test fixtures and the Bun version pin (#26051)
* Keep test interruption signals within owned processes

* Pin Bun and add optional unit runner shutdown diagnostics

* Unblock CI lint without changing session host runtime

* Avoid duplicating runtime import-check dependency bundles

* Leave unit runner diagnostics disabled by default

* test: keep runner incident follow-up focused on durable guards

* test: apply transcript replacements as authoritative snapshots
2026-10-06 21:42:02 -07:00
Neil d3e1494674 test: bound memory used by the runtime Electron audit (#26049) 2026-10-06 20:21:34 -07:00
Brennan Benson 0bdcaf36ed fix(claude): start a Claude chat with its saved options and send the first message at once (#25152)
* fix(claude): end a Claude start that never answers initialize after 120 s

* Read the Claude startup deadline inside startup; fix a stale test comment

* fix(claude): start a Claude chat on its initialize answer, not on a frame only a SessionStart hook sends

Startup waited for system/init or a SessionStart hook frame as well as the initialize answer.
Before the first turn only a SessionStart hook sends one, and Orca adds that hook only through
its optional status hooks, so with them off the first message was held forever. Startup now
lands on the initialize answer; a start frame already seen is still checked, and one naming
another session ends a started session. The deadline drops to 90 s so it fires inside the
host's 120 s start wait.

* fix(claude): time the Claude start by silence, and fail it at once on another session's frame

Claude answers initialize only after its SessionStart hooks finish, so a total-time deadline
would fail every start behind a slow hook. Each start frame now restarts the clock. A frame
naming another session fails a start still waiting on initialize at once, as before.

* test(claude): a real Claude chat starts and answers with every hook disabled

* Say what the start-frame re-arm covers, and check only start frames in the hook-less real test

* fix(native-chat): a Claude chat starts with its saved options and takes its first message at once

Saved model, effort, Fast and permission mode are passed as launch options, checked against the
account's cached model catalog, instead of restored by control requests after initialize. With
nothing left to restore, the host no longer holds a message until the CLI answers initialize, and
the 90 s startup deadline is gone. A Stop on a start that never answers ends the child and settles
what it was handed as stopped. A failed result that repeats the turn's own API error reply writes
no second row.

* fix(native-chat): a host stop of a Claude start fails the message it was handed, with one row

With no start-hold the delivery loop no longer sees a host stop of a start it waited on. The
child's end now rejects what it handed over with the host-stopped words and writes the one row,
as an exit of its own would; an idle start the host stops still goes quietly.

* test(native-chat): a Claude chat's first message is written before initialize answers

Rewrites the tests that encoded the start-hold, the startup deadline and the option restore to
the new contract, and adds: saved options at launch (catalog checks, bypass, fresh-session Fast),
a message written before initialize answers (adapter and runtime), Stop on a start that never
answers (stopped, child closed, nothing working), and an API error said once.

* revert(native-chat): keep a failed Claude turn's error row

The shared turn fold already shows a failed turn's error once after it settles, as an error;
dropping the row left the CLI's synthetic reply looking like something Claude said.

* fix(native-chat): pass saved Claude options unchecked, heal a retired model on the CLI's word, and never leave an unrun message in doubt

- Saved model, effort and Fast are launched as picked; only values no Claude can parse are left
  out. The pre-spawn cache check is gone.
- A saved Fast on for a new conversation is applied once the settings readback shows no
  per-session opt-in (dropped when there is one, or when the model is listed without Fast), with
  nothing waiting on it; the record keeps the pick.
- Under an Agent Permissions bypass, a saved narrower mode launches with the allow flag so bypass
  stays reachable.
- A turn whose reply is the CLI's model_not_found for the launched model drops that model from
  the record; the launch's own row for it is kept out of the account model cache.
- A child that ends before it answered initialize, for any reason, settles every message it was
  handed as not sent (cancelled for a Stop).
- A launched effort the CLI reports only as `applied.effort` is confirmed from there.
- The untimed-initialize comment is back to main's text.
- A real-CLI test for a message written before initialize answers, under saved options.

* fix(native-chat): type the close's ended event and the start-exit test fixtures

The close's ended event is typed as the adapter event so its optional startupUnanswered spread
fits exactOptionalPropertyTypes; two tests guard the fixture's optional generation, and the
hung-start fixture records initialize on the fake connection it holds.

* fix(native-chat): a Claude model heal keeps a later pick, a refused Fast is dropped, no allow flag

- `options-skipped` carries the retired value; the record drops it only while it still holds it.
- `started` carries the values a heal retired, and the record does not take them back from the
  CLI's report of the same value.
- A saved Fast on a new conversation is applied before `started`: a refusal drops it and records
  it skipped, as main's refused restore did; silence keeps it wanted and unconfirmed.
- A saved narrower mode under an Agent Permissions bypass launches without any bypass flag again:
  the allow flag is one older CLIs reject at start. Kept as a known limit.
- The real-CLI test asserts the message was written before initialize answered.
- The fake reports a launch effort only under `applied`, and a misplaced doc comment moves back.

* fix(native-chat): a new Claude chat reports started before its saved Fast is applied

The Fast apply on a new conversation now runs after `started`, so a Stop interrupts a running
first turn and an option write is not refused while the round trip is out. A refusal drops the
pick through `options-skipped`, in order after `started`; silence keeps it unconfirmed. The
launch's unreachable skipped-model branch is gone.

* fix(native-chat): a healed Claude chat goes back to the default model live; comments match the no-hold design

When the CLI says the launched model does not exist, the live child is also put back on the CLI's
own default (set_model with no model, fire-and-forget), so later messages in the same chat run; a
user's pick sent after it wins, and a refused or unanswered reset only logs. Comments that still
described the start-hold or the option restore now describe the launch options and the
handed-over, never-echoed rule.

* fix(native-chat): a message handed to a Claude start that never answered is kept as main keeps an unsent one

A child that ended before it answered initialize ran nothing it was handed, the same fact as a
send accepted and never handed over. Its end now settles those sends exactly as the chat settles
a queued send for that end: a quit keeps a person's message as a held card (restart words), a
close keeps it as a held card (closed words), a person's Stop withdraws it as cancelled, and a host
stop fails the start with one row.

* fix(native-chat): a quit during a Claude start that never answered offers no resume for the message it keeps as a card

The restart snapshot now reads the same never-answered fact the exit does, so a message handed
to such a start counts as queued work, not as work to resume. The retired-model reset comment
names the default it really applies.

* refactor(native-chat): the saved permission-mode launch helpers live with the spawn options that use them

Keeps claude-structured-launch-resolution.ts within max-lines once merged with main, and names
the hung-start test envelope's field type.

* fix(native-chat): a Claude chat's saved options take precedence over the agent Arguments' own flags

Main now passes the saved agent Arguments to the Claude child, and the SDK writes them after its
own options. An Arguments --model or --effort therefore reached the CLI as a second flag after the
chat's saved pick (a commander CLI keeps the last), and a saved Fast's launch settings replaced an
Arguments --settings file outright. The saved model and effort now stand in for the Arguments'
flags, and a saved Fast beside an Arguments --settings is applied by the start instead of at
launch.

* test(native-chat): the hand-built Claude session in the options test carries fastModeAtStart

* test(native-chat): the queued rig's start spy carries a named Mock type

An unannotated vi.fn() inferred @vitest/spy's internal Procedure, which CI's typecheck cannot
name in the factories' inferred return types (TS2883).

* fix(ci): run the PR's SQLite-backed tests in the Node runtime project

Lists the hung-start Stop test, renames the send-during-startup entry from its old name, and
carries main's own two entries from #26010 so the boundary test passes before the next merge.
2026-10-06 19:19:47 -07:00
Brennan Benson b3b6c5dc13 Give native chat names one source for tabs, sidebar and AI Vault (list and search) (#25986)
* Give native chat names one renderer source and drop Vault's name repair copies

The host's saved conversation name now rides the structured session status
feed, which already exists per host, is keyed by the durable session id, and
keeps a closed chat's summary. Tab strip, sidebar rows and AI Vault (list and
search) read it through one hook and one display order (tab alias, saved name,
host label). Vault no longer copies names into its cached results, so the
projection, recovery and pending-title modules and their tab-snapshot lanes are
removed. Indexed search hits now carry the native owner and saved name from the
host that indexed them.

* Type the sidebar name test fixture without an assertion

* Keep Vault search working when the chat host will not install

Naming and owning search hits is bookkeeping: if the native chat host fails to
install, return the plain hits instead of failing the search. The runtime RPC
only installs the host for clients that will receive the owners.

* Publish chat names to the feed independently of the tab retitle

A failed feed publication no longer skips retitling the open tab. The publish
now lives in the naming deps, where a test covers it.

* Note why the status feed must keep closed chats' summaries

* Bound names and owner ids that come from a paired host

Drop a published chat name the record store would refuse, and cap a search
hit's owner workspace id at the same length the list row and record use.

* Let native chat search hits from a paired host open their chat

A paired host's search hits carry no resume command, so the row disabled
every open action even for a native chat it can open through its owner, as
its list row does.

* Ignore workspace ids that name object members in tab lookups

A paired host's row or search hit could carry a workspace id such as
"constructor", which read an Object.prototype member as a tab list and broke
the render. The shared tab index now reads only own workspace entries.
2026-10-06 19:04:04 -07:00
Jinwoo Hong 825d7bd5a9 test(vitest): run agent-launch-instant-tab in the SQLite runtime project (#26028)
#25430 added a test that opens a real agent-session record store, but not to the
SQLite runtime list, so vitest-sqlite-runtime-boundary fails on main.
2026-10-06 21:33:40 -04:00
Neil 3fb72d135d Run Node event-loop measurement after ordinary test suites (#26015) 2026-10-06 18:08:50 -07:00
Neil 37ff3873a0 Run combined localization catalog verification on Bun (#25999) 2026-10-06 18:03:47 -07:00
Neil f0520851ab Keep mobile restore SQLite fixture in the Node test runtime 2026-10-06 17:46:39 -07:00
Neil 1040f673b1 Keep orchestration SQLite fixtures in the Node test runtime 2026-10-06 17:46:39 -07:00
Neil 3308ff8b26 Bound YAML merge conversion and close SQLite routing review gaps (#25998)
* Close test runtime and YAML merge review gaps

* Bound YAML conversion inside explicitly tagged pairs

* Make completion notification fixture cadence deterministic
2026-10-06 17:31:56 -07:00
Jinwoo Hong faed899cd3 fix(ci): run three new SQLite-backed tests in the Node runtime project (#26010)
#25888 and #25766 added tests that import the orchestration database or the
structured session runtime without registering them in the Node runtime list,
so vitest-sqlite-runtime-boundary fails on main and every PR.
2026-10-06 19:48:41 -04:00
Neil d8c871a1f0 Speed up unit tests with cross-runtime duration scheduling (#25967)
* Schedule unit tests across runtimes by measured duration

* Keep new SQLite fixture suites on Node after updating main

* Inject scheduling timings instead of mocking the module

* Keep agent-session database lifecycle contracts on Node
2026-10-06 15:33:40 -07:00
Jinwoo HongandClaude 72b84118e0 test(e2e): terminal layout parity check against main (#25681)
* test(e2e): add terminal layout parity check for topology refactor PRs

Runs fixed terminal-layout journeys in the real app on two builds, captures
the renderer topology and the saved workspace session, normalizes volatile
values and fails on any difference not declared for a named bug.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(e2e): record quit exit status and report paths main does not reproduce

Co-Authored-By: Claude <noreply@anthropic.com>

* test(e2e): invoke pnpm correctly under corepack and silence the typeless-module warning

Co-Authored-By: Claude <noreply@anthropic.com>

* test(e2e): accept pnpm's forwarded -- in the layout parity runner

Co-Authored-By: Claude <noreply@anthropic.com>

* test(e2e): close parity panes through the user's chord and treat absent maps as empty

Driving PaneManager.closePane directly left main to learn of the close from the
PTY exit, which raced quit; the keyboard path commits the close in main first.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(e2e): build parity checkouts before running, forward -g, and add a drag-out scenario

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-06 18:26:49 -04:00
Brennan Benson 6aa12c30ef Name native Claude and Codex chats after their first message (#25724)
* feat: generate chat names through configured text agents

* Project structured chat names across stored tabs and session rows

* Name native chats from their first live message

* Restore the journal test provider handle import

* Fix first-message naming and live Vault title updates

* Preserve Unicode characters in bounded chat naming prompts

* Read chat naming settings only when a turn needs them

* fix(chat): store only generated conversation names

* fix(chat): keep ordinary labels across unnamed chat surfaces

* fix(chat): preserve naming after fast first turns

* Preserve native chat names across command-first sends and Vault lifecycle

* Route native chat SQLite contracts through the existing Node test pool

* feat(settings): add chat naming controls to Chat page

* Preserve chat naming drafts and configure Custom commands through host settings

* Scope synthetic command output to its journal thread in naming test

* Keep naming test fixtures within their typed project boundaries

* Resolve the real Vault hook directly in its integration test

* Publish chat naming save refs after render commits

* fix: keep Japanese chat naming copy stable during repair

* Update test contracts for the naming integration

* Keep journal fixture reads on the host clock

* Give Chat names a separate settings section
2026-10-06 14:42:59 -07:00
Jinjing f7c542c7a3 Keep profile saving alive after a stalled main loop (#25318)
* Keep profile saving alive after a stalled main loop

After a long main-loop stall (overnight sleep, dark wakes), the profile
writer's overdue 30s timeout could run before an acknowledgement that was
already queued, permanently retiring the writer until restart. Terminal
creation then failed because pane bindings could not be saved.

- Writer deadlines measure lateness on the monotonic clock and grant a
  bounded fresh window when the callback is overdue or the system reports
  suspended; resume re-arms without spending grace. Applies to
  initialization, every command, and the close/exit wait.
- The "Saving stopped" alert is parented to a visible main window (never a
  parentless synchronous macOS alert), deferred until shown, deduplicated,
  and says whether the latest change is unconfirmed.
- Timeouts, grace, writer faults, and alert presentation leave sanitized
  durable breadcrumbs.

* Fix profile writer timeout and shutdown races

* Run profile writer stall regression on Linux and Windows

* Keep Electron probes out of headless runtime qualification
2026-10-06 14:13:39 -07:00
Neil 83cf7cf5e2 perf(ci): compile release JavaScript once for all packaging hosts (#25828)
* perf(ci): share release JavaScript across packaging hosts

* fix(ci): verify the projected web entry in release archives

* fix(ci): use the Windows system archive tool for release bundles

* test(ci): retain stylesheet evidence in release build comparisons

* test(ci): verify release parity across native color rounding

* test(ci): normalize manifest asset references without changing import order

* test(ci): compare portable outputs across Windows text and color formatting

* test(ci): preserve module identity across dependent asset hashes

* fix(ci): keep SVG build inputs identical across release hosts

* fix(ci): stabilize compiler inputs and projected web bindings

* fix(ci): retain vendor minification in projected web output

* test(ci): normalize platform-specific pnpm manifest source paths

* fix(packaging): exclude shared build staging from application files

* test: align thinking-state fixtures with the current source shape

* test(mobile): reuse message fixtures within the line limit
2026-10-06 13:22:04 -07:00
Jinwoo Hong ac8ea9f958 fix(codex): write Orca's hook into ~/.codex only when something changed (#25743)
* fix(codex): write Orca's hook in ~/.codex only when something changed

One reconcile replaces the per-launch writer and its background approval
session. It runs at app start after PATH hydration, when the setting turns on,
once per native pane spawn, and (with a bounded 3 s wait) on Codex launches and
resumes. Orca's entry lives alone in a matcherless group, appended last unless
already in place; other copies are removed and the shifted user approvals move
verbatim under both key spellings. The approval, with Codex's own hash, goes in
first. The routing gate closes only for a hooks.json with a bad shape.

* fix(codex): report hook status for the home the next pane uses

Status reads ~/.codex (both key spellings) when launches use it, or the CLI
has no selection, and the selected managed home otherwise. It explains: update
Codex, Codex not found, not asked yet, approved by Orca but not yet confirmed
by Codex, a hooks.json shape that moves panes to Orca's own home, and inline
config.toml approvals Orca cannot add to.

* test(codex): pin the ~/.codex approval against a real Codex, and run it on those files

The contract now also writes Orca's entry into a throwaway ~/.codex and checks
that Codex lists it trusted and enabled, turns it back on over a /hooks
switch-off, lists it for review after a user inserts a hook ahead until the next
check, keeps an inline config.toml loadable, and shows no review in a real TUI
start. The real-binary CI job now runs when the ~/.codex writer changes.

* test(ci): keep the scope test under its line cap; assert the ~/.codex paths beside the contract job

* test(codex): type the hoisted test holders instead of asserting

* test(codex): cover the matcher rule, an older build's event, and the spawn trigger

Adds tests that failed against mutants which survived the first pass: an
entry alone in a matcher group, a current copy's approval in an event left to
an older build, one run per spawn, a spawn riding a running reconcile, and the
native-only spawn trigger. The opt-out's Codex-hash cleanup now runs once,
after the sweep, instead of twice.

* docs(codex): name the ~/.codex reconcile in the legacy sweep's lane comment

* fix(codex): approve ~/.codex with Orca's own hash while Codex's answer is pending

The reconcile waits at most 0.5 s for Codex's hash. If the lookup is still
running, an entry already in place keeps its approval and nothing is written;
a missing entry or approval gets Orca's computed hash, approval first, inside a
launch's 3 s wait, as main's did. When the lookup lands, the reconcile runs
again and Codex's hash replaces it.

* refactor(codex): one start for the hook lookup and the ~/.codex reconcile

startCodexHooks replaces the two start functions and takes the PATH wait that
the reconcile request's after option carried. A spawn reconciles unless one is
running, without a scheduled flag, launches call one reconcileCodexHooksForLaunch
on the shared withTimeout, and the reconcile alone reads the hooks setting. The
startup test now runs the ready phase instead of matching its source text.

* refactor(codex): plan ~/.codex with the shared Orca-hook pruner and approval reader

The planner prunes with removeManagedCommands, as the opt-out does (so a hook
that runs Orca's script through its args goes too), and returns a prune or
settle union. While Codex has not answered, the kept approval comes from the
shared per-event reader. The reconcile result is its outcome alone, a
concurrent save spends a pass of the one bound, the opt-out removes Orca's
approvals once, the approval-first writer takes the hooks path, and
getRealHomeConfigTomlPath gives way to getSystemCodexConfigTomlPath.

* fix(codex): ~/.codex keeps only an approval holding a hash Orca's entry may carry

* refactor(codex): name the ~/.codex hooks-file check for what it reports

* test(codex): the ~/.codex entry tests know Codex's earlier hashes, as the app's lookup does

* refactor(codex): the ~/.codex reconcile uses the lookup's answer names

* test(codex): ~/.codex status tests keep a Codex on PATH unless one is missing

* refactor(codex): one stopgap for both homes; the ~/.codex pass finds its own home, hash and user data

Also passes every approval to the approval-first writer, which already skips the ones in place.

* refactor(codex): the reconcile start owns the warm-up; no rerun-on-answer flag

The lookup start only records the hydrated PATH; one catch, one request type, and the app config is kept as given.

* refactor(codex): the ~/.codex pass returns nothing; tests read the files

Also makes the ~/.codex opt-out take Codex's hashes, as every production caller passes them.

* refactor(codex): the ~/.codex approval cleanup takes Codex's hashes

* test(startup): check the Codex hook start waits for PATH by behavior, not identity

* fix(codex): a status-hook problem never moves the system default off ~/.codex

* chore(ci): run the real-Codex contract when the ~/.codex reconcile changes

* docs(codex): the ~/.codex reconcile's refused branch covers every definitive refusal

* test(codex): start the failure-memo lookup with the reconcile's PATH-only start

* fix(codex): ~/.codex gets nothing while Codex is missing, a pending answer uses the saved one, a conversion waits for a run that writes, and status and the reconcile read one home and one Orca-hash rule

* test(codex): ~/.codex and managed-home status tests name their home; ~/.codex rewrites the backslash key Codex on Windows reads

* test(codex): the Windows upgrade test reads status for the managed home it installs

* test(codex): the real-TUI contract trusts its workdir by its real path, and its userData exists

* fix(codex): keep Orca's approval at every slot its entry holds in ~/.codex
2026-10-06 16:12:23 -04:00
Brennan Benson ab6389045e fix(native-chat): start Windows chats without reading process creation times (#25718)
* fix(native-chat): start Windows chats without reading process creation times

Native chat on Windows refused to start ("Orca can't run this agent in a chat
here") whenever the process-table addon could not report process creation
times. Chat never needed them; only the bookkeeping around stopping the agent
did.

Windows now follows the common pattern: Stop ends the agent's tree with
`taskkill /T /F` on the child Orca still holds, and reports the tree gone only
when taskkill exits 0. A saved pid is never signalled after a restart.

- Remove the Claude and Codex location gate and the Codex launch refusal.
- A start time that cannot be read is recorded as unknown instead of refusing
  the session; recovery already releases such an owner without signalling it.
- Delete Claude's Windows creation-time descendant snapshot and its verifier.
- Codex's Windows teardown reports the real taskkill outcome.
- The renderer no longer waits on the capability flag; the host keeps
  publishing it for older clients (temporary).

* fix(native-chat): treat a Windows Claude exit after stdin end as a proven close

On Windows an idle Claude leaves on its own once its stdin ends, so every Stop
and close read as unproven: live background work settled as unknown and the
trace logged a close that "did not finish cleanly". As with the Codex close,
that exit is now the close and Orca makes no claim about processes Claude
started; a forced close still rests on taskkill's own report.

Also drop the Settings clause about Windows needing process start times, and
fix comments and the tracked process-enumeration doc that still described the
removed start-time gate and descendant snapshot.

* fix(native-chat): renew a held child's lease without a PID probe

An owner recorded without a process start time could never renew: the renewer
re-proves every live owner by PID identity, and with no start time and no
spawn-token echo that probe is indeterminate. The lease's last renewal then
stayed at the spawn, so a turn cut by an Orca crash was dated to its own start,
and every tick logged a failed renewal and split the batch into one store
transaction per chat.

The runtime now renews a lease for a child it still holds at the record's
fence and whose exit it has not received, recorded as a `held-child` match.
Receipt of the exit ends that proof before the exit is settled, and a restart
holds no child, so a dead owner's lease still expires and restart adjudication
is unchanged. Records this runtime does not hold keep the PID probe.

Also correct the identity probe's comment about shipped addons and note that
recovery's stop ladder is POSIX-only.

* fix(native-chat): keep a failed Windows taskkill unproven across a retried close

A Windows close counts Claude leaving on its own after its stdin ends as the
close. That shortcut also caught a retried close whose first attempt forced a
taskkill that failed: once the root exited, the retry returned true and the
failure read as a proven close. The tree reaper now records whether a reap
ever reached the live root, here, on an earlier close or from a transport
failure, and the shortcut applies only when none did; otherwise taskkill's
verdict stands and no new taskkill runs against the exited root.

Pin the platform on the existing tests that assume the POSIX close, and say
what `exit-proven` means on Windows in the acquisition-failure docs.

* fix(native-chat): derive a held child's liveness from the adapter's own handle

Lease renewal trusted a held child until the host settled its exit, and the
host hears of an exit late: Claude runs its close ladder and a store write
first, unexpected exits wait on one delivery chain shared by every chat, and a
Claude close that cannot prove its tree publishes nothing at all. A dead root
could keep renewing through that window, so a crash in it dated the cut turn
late. Renewal now asks the adapter, which owns the process handle and sees the
exit first: a child is held only while it is on record at the lease's fence
and its adapter still runs that exact acquisition with no root exit seen. An
adapter that cannot answer falls back to the PID probe. The stored
exit-received mark is gone.

The held-child read is now required by the runtime state and the renewer, and
a host-level test drives renewal through the real host wiring.
2026-10-06 12:09:58 -07:00
Jinwoo Hong 8d049b594d fix(codex): approve Orca's hook in managed Codex homes with Codex's own hash (#25742)
* feat(codex): ask Codex for its hash of Orca's hook in a throwaway home, cross-checked by position and path

* feat(codex): cache Codex's hook hashes per binary and version, asked one at a time and only by the app

* feat(codex): write a hook approval before its entry, and take back only its own on failure

* feat(codex): approve Orca's hook in managed Codex homes with Codex's own hash, written first

Managed homes (the shared mirror and per-account homes) no longer run a
background approval session. Status reads the home's files against Codex's
answer, and turning hooks off recognizes every saved version's hashes.

* feat(codex): managed homes approve Orca's hook with Codex's hash; drop their background approval

The previous commit carried only the managed resume's wait; this one holds
the managed install it relies on. Managed homes (the shared mirror and
per-account homes) write Codex's hash before the entry, fall back to their
own approvals when the answer is late, and strip Orca's entry only when
Codex itself answered with nothing to approve. Status reads the home's files
against Codex's answer, and turning hooks off recognizes every saved
version's hashes.

* feat(codex): only an Orca-launched Codex waits up to 3 s for the hook hash; warm it after PATH hydration

* feat(cli): name the file each agent hook status reports on

* test(codex): real-Codex contract for the derived hook hash in managed homes, on both pins and latest

* test(codex): type the hook-hash test fixtures and drop a duplicate import

* fix(codex): give a Codex launch its own install run instead of joining a plain terminal's

* test(codex): a user hook's approval stays put in an event Codex does not list

* test(codex): cover late answers, first-install mirroring, stale approvals and opt-out re-asking

* test(codex): a long managed home gets the daemon guard on its first install

* fix(codex): until Codex answers, approve a managed home's hook with Orca's own hash, as main did

A late, temporary or missing answer with no earlier approval in the home now
writes main's self-computed approval instead of leaving the hook out. Codex's
answer replaces it at the next install, a definitive answer (no hooks/list,
a refused cross-check, 0.128) never uses it, and status says the approval
is Orca's until Codex confirms it.

* test(codex): status flags an unapproved entry while Codex has not answered

* fix(codex): managed stopgap fills each missing event

Until Codex answers, a managed home kept only the events it had already
approved and dropped Orca's entry from the rest. Each event now keeps the
home's approval, else gets Orca's own hash, as main wrote every event. One
reader of the approval at Orca's entry serves the stopgap and status.

* refactor(codex): one Codex answer type, one in-process answer map, a disk-only memo

- One answer type with a kind (hashes, refused, pending) replaces two types
  and the three-field decoding at each caller.
- The lookup keeps one in-process answer per binary path, replacing the
  process memo, the global latest answer and the transient-failure map;
  status now reads the answer for the codex on PATH, not the last one asked.
- The memo file keeps Codex's refusals per version, like its hashes.
- Derivation takes the version it is given; one 30 s version-probe timeout.
- The launch wait reuses withTimeout, and launch prep passes launchesCodex
  down instead of a wait in milliseconds.
- Turning hooks off no longer forgets Codex's answer.
- Tests mock the derivation instead of a test-only resolver in production.

* chore(codex): list the approval reader for the CLI build; fold two identical scope checks

* fix(codex): count an approval at Orca's key only when it holds a hash Orca's entry may carry

* refactor(codex): one append for hook trust tables

* refactor(codex): the lookup keeps no entry for a missing Codex, and status checks the binary's fingerprint

Also names the lookup functions for the answer they return.

* refactor(codex): one stopgap reader for the managed home; the refused branch reads its own status

* refactor(codex): drop defaults and exports only tests relied on

* fix(codex): a failed ask of Codex stays pending instead of refusing its version

* chore(ci): run the real-Codex contract when the approval reader changes

* fix(codex): only a scratch home Codex loaded can refuse; the memo takes any hash and writes only on change

* fix(codex): hooks turned off during a launch's wait win, Off re-keys mirrored user approvals, and one rule says which hashes are Orca's

* fix(codex): an approval counts only under every key spelling Orca writes, as Codex on Windows reads only the backslash one
2026-10-06 14:15:35 -04:00
Brennan Benson 468e4e1167 fix(native-chat): a prompt card owns the chat input until its answer lands (terminal-backed chat, desktop and phone) (#25761)
* fix(native-chat): an answerable prompt card owns the chat input until its answer lands

* fix(mobile): a terminal chat's composer waits while its prompt card is up

* test(native-chat): type the prompt card fixtures without casts

* fix(native-chat): scope replies to acknowledged prompt occurrences

* fix(native-chat): preserve answer ordering and verified delivery

* test(native-chat): keep mock RPC client inside test boundary

* test(native-chat): place mock fixtures in the test-only scope

* Keep runtime comments within the module size limit

* test: preserve prompt delivery coverage in desktop CI

* Treat an older host's accepted write as delivered

A newer desktop or phone talking to a host that predates the write
settlement field read every accepted reply as "unconfirmed". Prompt cards
never dismissed, the phone showed "Response unconfirmed" on every tap and
ordinary chat messages were held as "Delivery unconfirmed".

The reader now uses writeSettlement when present and otherwise keeps the
host's whole-write accepted/refused verdict, exactly as before this branch.
Only prompt answers ask for provider settlement; ordinary callers
(follow-up delivery, paste drafts, option commands, composer sends) are
back on the original contract, so the legacy-handoff error class, its
flag, the sequence-only send helper and the mobile handoff hook are gone.

* Keep terminal-pane Escape on the plain accepted write

Every pane's Escape/Ctrl+C goes through pty:writeAccepted. This branch had
switched that IPC to wait for provider settlement, which dropped the
"remount this pane" signal for a daemon session awaiting recovery and could
stall later keystrokes behind a slow daemon acknowledgment.

pty:writeAccepted is back to its original local-only, synchronous write.
Prompt answers opt into settlement with requireWriteSettlement on the same
channel, and a settled refusal while the daemon recovers now sends the same
remount signal. Ordinary verified sends regain their original fallback write.

* Report a partly accepted local paste as unconfirmed

A settled local write split into chunks returned plain false when a later
chunk was refused after earlier ones were accepted. Callers read false as
"nothing was written", so chat showed "Message not sent" with a prefix
already in the agent's input. It now reports the write as unconfirmed,
the same verdict the paired host gives for a partial write.

* Hide the chat composer under a prompt card instead of unmounting it

When an approval or question card took the input region, the composer
unmounted. A message still waiting for its Enter was cancelled and its
bubble deleted after the draft had already been cleared, so the message
vanished without a notice; composer history was also wiped each time.

The composer now stays mounted but hidden while a card owns input, so its
state survives. A send that has not submitted yet is still stopped (its
Enter would answer the card), but its bubble stays with "Message not sent"
so the text is not lost. The composer ref is detached while hidden, so
root typing, paste and reveal focus never reach it.

* Keep an answered prompt hidden after the chat view remounts

The "answered" dismissal lived in component state. Toggling chat to
terminal and back, a PTY reconnect, or leaving the phone session and coming
back while the approved tool was still running brought the answered
approval back, and it then took over the input again.

Desktop now keeps the answered occurrence per pane outside the view;
phone keeps it per chat tab outside the controller. Both still retire it
when the pane observes the prompt clear or change, desktop also when the
tab retires, and both maps are size-bounded.

* Update the prompt-reply reliability gate for the review fixes

Older hosts' accepted answers now dismiss like acknowledged ones, the
composer stays mounted under a card, and answered prompts survive a view
remount. The gate's invariant, oracle, assertion list, new test files and
the two latest evidence runs now describe that contract.

* Let users hide a prompt card, keep Escape from denying, and gate only Send on the phone

The chat input could stay locked behind a card the host never closes (for
example after a Deny typed in the terminal), and Escape on a focused
approval card denied the tool even when the user meant to close a picker.

- A Hide control (chevron) on terminal approval and question cards, desktop
  and phone, hides that prompt occurrence and gives the input back. It writes
  nothing to the agent and uses the same per-occurrence dismissal as an
  acknowledged answer, so a new occurrence shows the card again.
- On desktop, Escape on a card now does the same Hide instead of Deny, and a
  card that appears while the user is typing no longer takes focus.
- On the phone, a card blocks only Send: typing, dictation and image attach
  keep working on the draft. The placeholder is back to the normal one.

* Fix two comments that still called older-host replies unconfirmed

Since an older host's accepted write now counts as delivered, the
requireWriteSettlement comment and the reliability gate's oracle said the
opposite of the code. Both now describe the current rule.

* Collapse prompt cards to a strip instead of hiding them, and close the round-2 gaps

Hide removed a card completely, so nothing on screen said a prompt was still
waiting, and several edges let the chat type into a live prompt.

- Collapse (the header chevron, or Escape on desktop) folds the card to a
  one-line strip above the composer; the strip's chevron expands it back.
  Collapsing writes nothing, frees the composer, and is disabled while an
  answer is still being written. Each pane or tab keeps the occurrence as
  answered or collapsed, so a remount restores the same view.
- Questions now carry the host wait's start like approvals, so an identical
  question in a new wait shows again (desktop and phone). A transcript-only
  prompt, which has no wait start, is dropped when the view stops observing
  it, and a transcript still loading no longer clears a dismissal.
- Desktop: while a card owns the input, the hidden composer cannot send or
  interrupt even if it still has keyboard focus, and the card takes focus in
  the same commit. A send the card retires no longer types Ctrl+U under it.
- Phone: an Ask hides the heuristic card read from the same waiting status,
  and the dismissal store is scoped by host, worktree and tab.

* Keep a collapsed card's partial answer, and scope its focus to its own pane

Collapsing a question card unmounted it, so expanding it again lost the
chosen step, selections and typed "Other" text; Escape typed in that text
field collapsed the card. A card arriving while the user typed in another
surface (sidebar, notes, a browser URL bar) also took the keyboard.

- The collapsed card now stays mounted but hidden (and inert on desktop)
  under its strip, on desktop and phone, so a partial answer survives
  collapse and expand. Escape inside the card's text field no longer
  collapses it. The question card shows the same focus ring as the approval
  card.
- A card takes focus only from inside its own pane (its hidden composer) or
  from the page body, never from a text field elsewhere.
- Desktop and phone share one dismissal store in src/shared, bounded by the
  existing scope-cache helper, which moves to src/shared with it.
- The card send imports the verified helper from its own module, and the
  phone files are split so each name matches its contents (header action,
  strip, lane selector).

* Return focus to the composer after a prompt card collapses

Since a collapsed card stays mounted, Escape or the chevron left keyboard
focus inside the now hidden, inert card. The composer's reveal-focus took
that as focus already in the pane and stood down, then the browser dropped
focus to the page body, so typed keys went nowhere.

Reveal-focus now treats focus inside a hidden or inert subtree as not in the
pane and focuses the composer. On the phone, collapsing a card also
dismisses the keyboard so a hidden reply field does not keep it.

* Keep the question card's collapse chevron beside its Cancel button

The question card header spread its three items with justify-between, which
put the new chevron in the middle of the header. The title now takes the
free space, as in the approval card, so the chevron sits next to Cancel at
the right edge.

* Run the prompt tests on the merged main

Main now runs Vitest under Bun, which resolves a long data: URL import as a
package name, so the SSH delivery test loads its bundled mobile module from a
temp file instead. The phone prompt harnesses mock the live line that main's
view now renders, and add Platform, which main's text-selection helper reads,
the same way main's own view tests do.
2026-10-06 10:57:58 -07:00
Neil 3ec38b8c6d Run Vitest on Bun with Node runtime contracts (#25840)
* Run Vitest on Bun while preserving Node runtime contracts

* Preserve runtime timing provenance and keep the Bun pin in config

* Scope builtin compatibility mocks to test-only lint exceptions

* Give capture retention fixtures distinct filesystem timestamps

* Await the copy button success state in the React fixture

* Bound Node test worker shutdown and tighten migration fixtures
2026-10-06 03:18:06 -07:00
Neil 13ea35973c Stop expensive checks when an unmerged PR closes (#25829)
* Cancel active checks when an unmerged PR closes

* Register owned-branch cancellation qualification

* Keep temporary cancellation qualification outside the review diff
2026-10-06 01:03:24 -07:00
Neil f6f96db6be Build SSH hostile-host Linux slots independently (#25821) 2026-10-06 01:02:53 -07:00
Brennan Benson eaaae0196f feat(native-chat): record fresh sessions after failed restoration (#25747)
* feat(native-chat): record a fresh provider conversation that replaced one the agent could not restore

A chat whose saved conversation the agent cannot reopen can now continue in a
fresh one: the handle chain records the new conversation as a creation that
replaces the lost one (which, why, and when), keeping every earlier link.

Rows keep a shape older builds read: the stored chain starts at the latest
replacement and carries the earlier links inside it.

* refactor(native-chat): store a replaced conversation flat; refuse it where older builds read the row

Older builds only read Claude and Codex records, so the nested stored form
protected rows no replacement can reach while adding a cap mismatch after a
downgrade. Store the chain as held, refuse a replacement in a Claude or Codex
chain until one has a stored shape older builds read, and refuse a supersession
key on a replacement that names no creation in the chain.

* test(native-chat): prove replacement rows survive downgrade and re-upgrade
2026-10-05 23:22:24 -07:00