Files
orca/tests/e2e/cross-version-wire
Brennan Benson a551aa5fe6 OpenCode structured chat over the Agent Client Protocol (#25845)
* refactor(native-chat): Grok's registration declares where it runs; ACP no longer borrows Codex's location rule

The rule a self-supervised agent child runs under (this machine, no WSL, Windows only with process
start-time proof) is its own module that Codex and the ACP adapter both use. Grok's registration
takes the full account-home resolver signature, and D3's tests build hosts with the agent registry.

* fix(native-chat): Grok follows the ACP runtime's request contract and the managed process's close

A request the agent or a Stop cancels is answered with the agent's own cancelled reply by the code that
owns it (the runtime no longer answers a silent handler), so a Stop needs no separate decline pass. A
permission answer still being saved when the agent stopped waiting is reported unconfirmed, since the
protocol already answered it cancelled. Cancelling the agent's own turn is the plain cancel. Request
rows are matched under their generation-scoped ids. A refusal's reason comes from the dialect's wording
path. The child drops its own stderr tail and close policy for the managed process's, and a close
whose process tree was not proven gone is reported as the adapter contract asks.

* fix(native-chat): a Grok chat Orca already holds resumes without writing what Grok replays

A chat with a saved Grok session reattaches with session/resume where the agent offers it, else
session/load. Either way the call runs inside the translator's load window, so what Grok sends while
it reattaches (its saved exchange, a task the dead process left running, ended by the restart) opens
no turn and writes no row; only context usage reads on. A reply an Orca or Grok crash cut short is no
longer completed from Grok's saved history: it reads like a Claude or Codex chat's, with the existing
notice. The attach window also closes after a failed attach, and a created session that session/resume
reports missing is replaced like one session/load reports missing.

The replay reconciliation is removed: the lane no longer reads the journal, and D3's replayed-input
grammar test and completed-turn check in the assembler go with it.

* refactor(native-chat): a failed Grok reattach needs no window close of its own; its lane is replaced

* test(native-chat): D3's merged tests use the shipped declarations and the launch options main requires

* fix(native-chat): typecheck fallout of the base merges; any agent's empty chat is reusable

Main's idle-empty-chat lookup and launch join now take any registered agent, as the rest of the
launch path does. The refusal check moved into the prompt turns and the prompt-block conversion beside
the turns that send it, keeping both files in their line limit.

* fix(native-chat): a Grok Stop ends the process once Grok settles its turn; the next send resumes

Grok's session/cancel ends only the running turn: work it already moved to the background keeps
running and can begin a turn of its own after the person pressed Stop. Stop is now a session
boundary, as it is for Claude: the cancel answers open requests and lets Grok end the turn, the host
waits a bounded grace for that, then ends the process; the next send relaunches and resumes.
The adapter's own bounded close of a turn Grok began is gone. Its named-turn check stays: the host
ends the session unless the provider declines a Stop naming a turn that has since ended.

* test(native-chat): a Grok Stop ends the process only after Grok answered the cancel

* fix(native-chat): Steer on a Grok card cancels the running prompt, then sends it

A send that reached Grok while a prompt ran was held in the adapter until that turn ended: Steer
on a queued card took the card out of the host's editable queue and meant 'send after this turn'.
It now cancels the running prompt (session/cancel; the session stays) and sends as the next prompt
once Grok answers the cancel, as the common pattern does; a steer behind another cancels it in
turn, so the last one runs. The adapter holds a send only while that cancel lands, so its general
held-send queue and its holdsDispatch report are gone (every send it holds has its turn open in
the journal). An older client's mid-turn send takes the same path. capabilities.steering is
unchanged and still unread.

* refactor(native-chat): a close or Stop cancels a start through the acquire's own abort signal

The host owns the acquire it runs, so it now owns its cancellation: each attach's acquire gets an
AbortSignal, aborted from outside the session's queue by a close and by a Stop admitted now (the
same admission rule as before). The optional abandonStart adapter hook, the router's fan-out to
every adapter and the ACP adapter's session-keyed start map are gone; the ACP adapter keeps an
unkeyed set of starts only so quit can prove their children gone, and keeps a failed start's
unproven child until its exit is proven.
The hook also let a later close ask that child again. The host now does that from state it holds:
a close of a chat with no live child whose record still names an owner process with no death
evidence asks the adapter to release it. The answer is not recorded as proof (the lease probe
does that), so an owner pid an earlier Orca left is never killed or marked gone. Claude and Codex
ignore the signal and hold no such child; their release is a no-op (tested).

* fix(native-chat): a Grok crash that closes stdout before its exit still ends with Grok's last words

On macOS and Linux the agent's stdout ends before its exit is observed, with or without the
supervisor's EOF forwarding, so the connection's loss closed the journal first and its error text
became the session's ended reason, dropping Grok's stderr. The reason is now read at the proven
exit: the agent's last words when it left any, else why the connection closed. The failure already
carried them. Comments that assumed the exit comes first, that early frames past the cap refuse the
start, and that dispatch re-checks image support are corrected.

* fix(native-chat): nothing Grok sends while a held chat reattaches is written, marked as replay or not

The reattach window relied on the dialect's replay verdict, and Grok's frames read as live unless
they carry isReplay, so an unmarked chat frame during session/resume opened a turn that never
ended. D3 now marks every frame inside the window as replay before the translator reads it, so the
translator keeps only context usage whatever the agent marked; options and commands are still
adopted. The translator's load semantics are unchanged.

* test(native-chat): a Stop after a resume finds no turn an unmarked old reply opened

* test(native-chat): a resumed Grok chat keeps its last context reading; the resume refreshes only the window

* test(native-chat): a Grok background task a Stop ended reads as stopped reporting

* refactor(native-chat): quit's stop of each start answers through one promise kind

* fix(native-chat): quit aborts every start the host has in flight before draining attaches

A Grok that never answered its handshake held quit until the start's own 60 s bound, past the
20 s quit deadline. The host's teardown now aborts each in-flight acquire (and any the drain
still begins), so the adapter's own quit controller and its map of starts are gone: a start
has one canceller, the host's signal.

* fix(native-chat): a Grok start's abort stops reaching its child once the start has returned

The listener stayed on the host's signal until the attach finished committing, so a Close in that
window killed the now-live child behind the host's back and it read as Grok crashing. The start
now detaches it when it ends; a later Close goes through the session's own stop.

* fix(native-chat): a close or Stop during any attach phase stops the start before it launches

The attach began its abort controller only after reconciling leases, resolving recovery and
probing the previous owner, so a close or admitted Stop in those phases reached nothing and Grok
launched anyway. The controller now begins first, and the acquisition checks it before asking the
adapter to start.

* test(native-chat): a close during the attach's owner probe asks no adapter to start

Also renames the close test after the hook it no longer exercises.

* test(native-chat): a close's re-ask closes a Claude or Codex child a failed cleanup left

The re-ask is not a no-op for them: when the adapter still holds the child its cleanup could not
prove gone, the close stops it again as a requested close, and Claude persists the handle of the
conversation it ran so the next send resumes it. Corrects the tests' and comment's wording; the
close awaits the re-ask, bounded by each adapter's kill ladder.

* fix(native-chat): Steer during a turn Grok began itself cancels it and sends once it ends

A send while Grok ran a turn of its own (a background task waking it) went straight to Grok, which
queued it behind that turn where Orca could no longer withdraw it, while Stop treated the same turn
as the running reply. The send now waits as a steer, the turn is cancelled once, and the message
goes when the turn ends; a Stop withdraws it and an exit rejects it as never sent.

* test(native-chat): a steer whose cancel Grok never answers ends Grok and is rejected as never sent

Pins the bounded steer cancel kept from the runtime: past the bound the connection closes, the
running reply reads unverifiable, Grok's end reads as its exit, and the waiting steer is rejected
as never sent.

* fix(native-chat): a Grok crash stays a crash when a stop lands before its exit is proven

After the connection broke and the close could not prove Grok's exit, any later stop Orca asked
for (the next start, a Stop, a Close) marked the child as closed by Orca, so the crash read as a
requested close and Grok's last words were dropped; a send meanwhile was recorded unconfirmed.
The connection loss now decides the cause, and a send on that session is rejected as never sent.

* test(native-chat): fixtures this PR's registered Grok and desktop capability made stale

CI's unit shards failed on tests outside the PR's own lists. Each encodes something this PR changes
on purpose: Grok is now a registered agent (the seam test's unregistered agent is now Cursor); the
desktop now advertises registered agents (the restart-offer tests' older client drops that
capability explicitly); the attach context carries the start's abort controllers (the forget-status
double gains them); and the ACP real-host test rig sends to the host directly (listed beside the
other real-host rig in the send ratchet).

* fix(native-chat): a start quit stops is not the queued message's start failure

With quit now aborting a start it would have waited for, the delivery step recorded the aborted
start as the message's failure ("couldn't restart"). After quit has stopped delivery, the step
leaves the message to quit, which settles it as a close does ("The chat closed before this message
was sent."). The test that pinned quit waiting for that start and stopping its child now pins that
nothing is launched behind quit.

* fix(native-chat): a message sent after a Stop or close aborted a start gets its own start

A start the host aborts (an admitted Stop, a close, or quit) returned its refusal to the delivery
loop, which then rejected whatever was queued at that moment with "couldn't restart", including a
message the user sent after the Stop. The attach now reports that the host aborted it, and the loop
re-derives from the journal instead: what the Stop or close withdrew is already settled, a message
accepted since gets a start of its own, and quit's next step stops the loop. This replaces the
quit-only carve-out with the same rule for every abort and every agent.

* test(native-chat): the message sent after an aborted start is answered, so no settlement outlives the test

* fix(native-chat): a Grok model pick Grok never answers no longer holds Stop or Close

The pick runs on the session's queue. It now registers in the host's out-of-queue
abort registry beside a start, so a close, an admitted Stop or quit abandons it, and
the ACP adapter bounds it at 30 s like Claude and Codex. A late answer is still adopted.

* fix(agent-launch): a phone's launch opens a terminal for an agent whose chat it cannot show

agent.launch now reads the caller's capabilities by the rule tabs and restart offers
use (clientRendersStructuredAgent). A phone without registered-agents.v1 gets Grok as
a terminal again, as on main; the host's own callers and desktop clients are unchanged.

* fix(acp): strip every agent hook variable from the ACP child, from the shared list

ACP_CHILD_ENV_TO_DELETE was a second copy of the hook runtime keys that missed
ORCA_AGENT_HOOK_TRANSPORT; it now spreads AGENT_HOOK_RUNTIME_ENV_KEYS beside the pane
identity keys.

* refactor(native-chat): the mutation context carries the provider-wait registry itself

Keeps the host file within its line limit; one field instead of two closures over it.

* fix(agent-launch): agent.launch.v2 still vouches for Claude and Codex chats

The caller rule from the previous commit also turned Claude and Codex into terminals
for a client advertising only agent.launch.v2, whose contract says it opens a chat
(mobile retry-authority tests). Only an agent beyond those two now needs the client to
read it (clientRendersStructuredAgent); the test fixtures go back to what they were.

* refactor(native-chat): drop saved-history adoption from the timeline assembler

The common pattern discards the history a provider replays while loading a
session, so the assembler has no use for an input.history event.

* refactor(acp): drop session/load history adoption from the translator

The common pattern discards the history an agent replays during session/load,
keeping only what it says about the context window. Remove the adoption path
(acp-history-adoption.ts, the adopt option, and the historical background-task
liveness rewrite it fed) so load replay is always dropped except usage.

* refactor(native-chat): a pending input is only Orca's send now

Review follow-up to the adoption removal: drop the comment naming the
provider's saved message, and make requestedAt required since every pending
input comes from input.accepted.

* test(acp): keep the task-result status table on live frames

Review follow-up to the adoption removal: the result-status mapping was only
tested through adopted history, so run the same table on live frames, and
cover an unmarked task notice during a load being dropped.

* test(acp): a frame helper for a shell command Grok is running

* fix(acp): a Grok crash settles through the host's provider-exit batch, scoped to the turn it ended

A Grok crash ended the journal unverifiable before the adapter reported the exit, so the host's
provider-exit settlement found no running turn and wrote nothing: the adapter's failure (with
Grok's last words) never reached the journal, and a later stale-session pass wrote a bare,
thread-scoped cut-short row, so the partial reply was not folded as Claude's and Codex's are.

At a proven exit the ACP lane now ends its running turn interrupted at the exit instant, as the
host's exit contract expects of a child's own translator (Codex's does the same). When Grok's
stdout closed first (every POSIX crash), the turn is unverifiable only until the exit is proven:
the host's provider-exit settlement now takes the exit as proof naming the child's fence and
revises what that child left unverifiable in the same batch, with the turn-scoped row and the
adapter's failure. Claude and Codex write no unverifiable turn of a live child except a command
whose hand-off is in doubt; that turn is now revised at the exit instead of at the next open.

* test(acp): a crash seen first leaves the host no Grok turn to revise

* refactor(native-chat): what a gone generation left unfinished gets its own module

The settlement file passed 300 lines with the exit-proof revision. The unfinished-work reads
(capture, interrupted-by-the-exit, in-progress) are their own concept and move out unchanged,
apart from the exit proof they now take.

* refactor(native-chat): a watched exit revises what its child left unverifiable without reading Stop marks

An exit's own instant is the turn's end, so the revision needs only each row's fence: the
settlement's journal type gains itemFence alone, and the host test fakes say so.

* test(native-chat): drop the duplicate itemFence on the fake that already had one

* test(claude, codex): an exit whose stdout ended first still reports as it always did

The provider supervisor now ends Orca's stdout when the agent's ends, so on every crash EOF
arrives before the exit is seen. Claude's and Codex's connections report nothing at EOF and
report the exit, with its usual reason, once it is seen.

* fix(acp): reopen a chat with session/load, as the common pattern does

An agent that offers both now reloads its session instead of resuming it; the
reattach window still discards what it replays except context usage.

* fix(acp): drop the 60 s handshake bound; an abort fails the start's waits at once

Neither common design bounds an ACP handshake: Close, Stop and quit end a start
that never answers. The abort now also closes the connection, as a kill there
does, so the start settles even before the child's exit is proven. The
host-stopped start refusal only this bound produced goes with it; the idle
sweep keeps its words.

* fix(acp): a Stop naming an ended turn follows Claude's rule

It still stops nothing while another turn is live, but in the gap before a
follow-up's turn opens, which no client can name, it now stops what is in
flight and the session ends, as a Claude Stop does.

* fix(native-chat): a close no longer re-asks a failed start's unproven child

Neither common design retries that stop at Close, and Orca's Claude contract
re-asks only at the next start and at quit. The ACP adapter keeps the child
until its exit is proven and asks it again there, as Claude does.

* fix(acp): a message sent during a turn the agent began itself goes at once

Both common designs send it straight to the agent with no cancel; only Orca's
own running prompt is steered (cancelled, then re-prompted).

* test(native-chat): dismiss-all through a remote client's audience keeps a newer Orca's offer

Uses an audience production sends (one that cannot show every agent), per review.

* fix(acp): launch Grok as `grok agent stdio`, without the update and leader flags

The common pattern passes neither --no-auto-update, --no-leader nor
GROK_DISABLE_AUTOUPDATER; full access still adds --always-approve.

* fix(acp): an agent that ends its stdout, or answers unreadably, is not a lost connection

As in the common pattern, only a broken stdin (or Orca's own close) ends the
agent; one that closed its output but can still be written to stays until a
Stop, a close or its exit. The provider supervisor goes back to its base
content, so Claude and Codex no longer get the forwarded stdout end either.

* fix(native-chat): a person's close joining a failed one still binds the turn its child end cuts

On main every close of the chat writes its own Stop and settle. Here a later close joins the
first and writes no row, and the first's settle closed when its kill failed, so a turn that opened
in between and was cut by the next close read as failed. A person's close joining a person's close
whose Stop opened a settle now reopens that settle until its attempt is done.

Tests: a close whose kill failed still closes its settle; a turn opened between a failed close and
the next reads as the person's cancellation (each fails without its half of the fix).

* test(native-chat): Grok opens as a chat only behind the structured-chat setting

agent.launch and orchestration worker-start read the same setting as the
renderer route; pin both states for Grok on each. The setting's description no
longer names only Codex and Claude, in every catalog.

* docs(acp): generic ACP comments say what holds for every agent, not Grok

Stop ends the session for every ACP agent, as in the common pattern; the
adoption hook comment goes (adoption is not planned); a failed start's child is
retried at the next start or quit.

* test(claude, codex): type the EOF-before-exit test's streams; the supervisor no longer forwards EOF

The Claude test wrote to the child's stdout and stderr through their Readable
type, which the node typecheck rejects; it now holds its own PassThrough
streams. The comments no longer credit the reverted supervisor change.

* feat(acp): a steer's cancel asks once and never ends the agent

The runtime had one cancel: send session/cancel, wait at most 10 s for Orca's prompt to settle,
then close the connection, which ends the agent. A steer used it too, so a slow agent lost its
process just because the person added a message. requestSteerCancel() now sends session/cancel
once per prompt, cancels the agent's open requests and answers later permissions cancelled, and
never bounds or closes: the prompt's own reply ends it and the steer's prompt follows. cancel()
stays the Stop: bounded, then close. A Stop after a steer still bounds and closes. Both cancel
paths move into acp-prompt-cancel.ts over one cancel channel.

* chore(native-chat): keep the record store and recovery capsule under max-lines after the main merge

* fix(acp): a repeated steer shares the cancel in flight; say what the caller owns

Per review: a second steer before the first write lands returns that write instead of resolving
early. The steer's JSDoc says the wait for the prompt's reply is unbounded and that a prompt that
fails instead must not take the steer until the caller rebuilds the session; the Stop's says a
prompt that settles in time leaves the agent for the Stop's owner to end. The steer test now gives
the runtime a handler that would allow: the open permission's signal aborts and the late one never
reaches it.

* fix(native-chat): drop the stopDelivery the A3 merge doubled

* fix(acp): a steer's cancel asks Grok once and never ends it

A steer now uses D1's notify-only cancel. Two messages sent during a reply Grok began itself
cut that reply, as the common pattern does, and then both run; before, the queued first
message could not answer the bounded cancel and Orca ended Grok although Grok answered.
A Stop keeps the bounded cancel and its 4 s grace.

* fix(acp): a permission Grok asks with no prompt of Orca's running is declined

During a turn Grok began itself nobody asked it to act, so the request is answered
cancelled at once instead of opening a card that waits, as the common pattern does.

* fix(acp): a Grok that dies while starting is reported with its own last words

A dying process's stdout ends before its exit is seen, so the start failed as a closed
connection and Grok's stderr was lost. A start whose connection closed now waits, bounded by
the Stop grace (or a Close/Stop), for the exit before it is told.

* test: a Stop after a steer sends its own cancel; drop the import the A3 merge doubled

* test(native-chat): main's Stop-note test builds its turn context with the agent registry

* test(claude): say why the close test's fake child cast is safe

* test(native-chat): build the Stop-opened-turn test's identity and turn context the current way

The test (#25056) landed before the opaque provider handle (#24991), so main still built
the old {kind, threadId} handle; the turn context also needs this branch's agent registry.

* test(native-chat): build the Stop-opened-turn test's identity with the opaque handle

The test (#25056) landed before the opaque provider handle (#24991), so main still built
the old {kind, threadId} handle.

* test(ratchet): require src/main/provider-process now that it has landed

* test(native-chat): keep main's opaque-handle import in the Stop-opened-turn test

Main's #25706 and this branch both added the import at different lines; the merge kept both.

* test(native-chat): keep main's opaque-handle import in the Stop-opened-turn test

Main's #25706 made the same fix as this branch at a different line; the merge kept both imports.

* feat(native-chat): record a fresh provider conversation that replaced one the agent could not restore

A chat whose saved conversation the agent cannot reopen can now continue in a
fresh one: the handle chain records the new conversation as a creation that
replaces the lost one (which, why, and when), keeping every earlier link.

Rows keep a shape older builds read: the stored chain starts at the latest
replacement and carries the earlier links inside it.

* test(native-chat): build this stack's journal identities with main's opaque provider handle

Main's #24991 replaced the {kind, ...} handle with {transport, agent, nativeId}; three test
files from this stack still wrote the old shape. Same lines the downstream ACP branch uses.

* docs(acp): every reattach drops the agent's replay, not only for a chat the journal holds

* refactor(native-chat): store a replaced conversation flat; refuse it where older builds read the row

Older builds only read Claude and Codex records, so the nested stored form
protected rows no replacement can reach while adding a cap mismatch after a
downgrade. Store the chain as held, refuse a replacement in a Claude or Codex
chain until one has a stored shape older builds read, and refuse a supersession
key on a replacement that names no creation in the chain.

* refactor(native-chat): read hosts' structured agents from the app-shell services

Main grew the startup hydration hook to its line limit; the host agents sync is an app-lifetime subscription like the structured session tabs sync beside it, so it moves there.

* Use current provider handles in transition tests

* Use current provider handles in timeline fixtures

* test(native-chat): prove replacement rows survive downgrade and re-upgrade

* Require the ACP directory in the runtime import check

* test(ratchet): require src/main/acp now that this PR lands it

* feat(acp): a saved session the agent cannot reopen continues in a new one, with one warning row

When session/load (or session/resume) of a saved ACP session fails, the chat starts a new session and records it as a creation that replaces the lost one (#25747's 'replaces' link), and writes one warning row that the agent no longer remembers the earlier messages. A created session the agent reports missing is still superseded silently; a signed-out agent or a start that is over (Close, Stop, a lost agent) still fails the start.

* chore(acp): rewrap the acquire header comment

* test(acp): a start closed while the agent reopens fails without opening or announcing a new session

* Let ACP connections own their supervised agent process

* Preserve ACP cleanup evidence and isolate exit observers

* Expose ACP cleanup observations and type the permission fixture

* refactor(native-chat): the registered-agents capability lives in its own module

Main's growth put protocol-version.ts one counted line over its 300-line limit once the capability
was added; like main's other per-feature capabilities, it now has its own module, and importers read
it from there.

* refactor(acp): one connection owns the Grok process and its protocol

D3 now opens each ACP agent through createAcpAgentConnection (ACP-ALIGN #25810): one object spawns the
process on the execution host, owns its stdio and protocol, and reports its proven exit. It is built and
tracked before the handshake, so a start's abort (Close, Stop, quit) still reaches it, and a failed start
keeps that same connection for the next close to retry rather than spawning another process.

Deleted: the spawnAcpStructuredChild wrapper and its test, the raw-stream runtime assembly, the caller's
exit -> runtime.close wiring, the stdout-EOF heuristic (the connection no longer treats stdout EOF as
exit), and the 10 s steer/Stop cancel bound with requestSteerCancel. Reader control maps to
pauseReading/resumeReading; a close is connection.close after the host's existing 4 s Stop grace.

The adapter owns what the protocol no longer does: one session/cancel per running prompt however many
steers arrive (cleared with that send's settlement, retried after a failed write), and a Stop or steer
answers every open agent request the person has not already answered with the agent's own cancelled
reply. An answer already being saved when the Stop lands is sent.

Tests: blocked cancel write never holds Stop's grace, two quick steers send one cancel, a failed cancel
write is retried, a real process exiting while a child holds its stdout ends the session, and the
existing start-abort, retention, crash, connection-loss and reload-failure suites on the new rig.

* Add chat-owned OpenCode server transport foundation

* Combine server fixture type imports for CI lint

* fix(acp): Grok signs in on its own machine with its API key or cached sign-in

When Grok reports that it needs authentication, Orca now names a sign-in method on the machine Grok runs
on, read from the same environment Grok was launched with: xai.api_key when XAI_API_KEY is set there and
Grok offers that method, else cached_token when Grok offers it, else none and the chat keeps the existing
not-signed-in refusal. The rule lives in Grok's launch spec; the adapter applies any agent's rule for new
and reopened sessions through the protocol client's caller-named method (authenticate, then retry once).
No new sign-in UI; interactive methods are never chosen.

* Bound request serialization and correct transport fixture types

* fix(acp): the adapter decides which of Grok's requests reach the person

The turn owner now admits every agent request, permission or question, from its own turn state: a
request reaches the person only while Orca's prompt runs and no steer or Stop is cutting it short (a
question may also come from a turn Grok began itself, until a Stop). Anything else gets the agent's
own cancelled reply and opens no card, so a question arriving after Stop or during a steer never
appears. A steer, like a Stop, withdraws the requests already open; an answer already being saved is
still sent. The protocol client's abort-on-cancel path is no longer used: after the connection
change its request signal aborts only when the connection closes.

* fix(acp): a plan Grok proposes shows as a plan, with no approval card

When Grok leaves plan mode it asks the client to approve its plan (x.ai/exit_plan_mode). Orca showed a
blocking 'Approve plan / Request changes' card for it; the common pattern has no such gate. Now the
plan goes into the chat's existing Plan row (the plan-document status row Codex and ACP plan updates
already use) and the request is answered at once with 'abandoned' plus feedback telling Grok to stop and
wait for the person's feedback or a request to implement it in a later turn, so nothing is approved on
the person's behalf. Dialects gain settleRequest for requests answered without asking anyone.

* fix(orchestration): worker-start opens a Grok worker in a terminal, as before

With the structured chat setting on, worker-start decided 'structured' for Grok and then the structured
worker factory (Claude and Codex only) refused it, so the start failed; main opened a terminal Grok
worker. Worker-start now decides with no registered agents beyond Claude and Codex, so Grok gets a
terminal worker as before. agent.launch and the app's own launches still open Grok as a structured
chat. Temporary until structured workers take registered agents.

* Bound OpenCode event consumers and isolate chat caller identity

* fix(acp): a prompt answer Orca can't read ends the turn instead of hanging it

A session/prompt rejection that was not the agent's own error answer (an answer that fails Orca's
schema, or one too large to read) left the turn running: the next message became a steer with nothing
to cancel and was never sent or settled, and Stop waited its full grace. As in the common pattern, any
prompt failure now ends the turn as failed (a failed-turn row without words, since none are the
agent's) and settles the send, so the next message goes. Only a closed connection keeps the send
running, for the connection-loss path to settle.

* fix(acp): send Grok's prompt-identity extension only to agents that echo it

session/prompt carried _meta {promptId, requestId} for every ACP agent, though only Grok's dialect
echoes it (injectedPromptIdentity). Now only an agent whose dialect declares it gets the extension;
other ACP agents get a plain prompt.

* refactor(native-chat): the registered-agents capability lives in protocol-version again, as on main

This reverts 0ef6d21815. That commit moved the capability to its own module only because main's
protocol-version.ts was then one counted line over its limit; main now defines it there itself within
the limit, and main's new restart test imports it from there. Main's test also reads the desktop's
capability list as an older client; on this branch the desktop advertises registered agents, so its
older client is that list without this one capability.

* test(acp): read the sign-in method with a schema, not a type assertion

* refactor(native-chat): composer transport and Stop control in their own modules

Main's rewind change (#19338) brought NativeChatStructuredSession.tsx and use-structured-agent-session.ts
to their line limits, leaving no room for this branch's image-acceptance and unpublished-Stop lines.
The composer's transport (sends, commands, options, image acceptance) moves to
use-native-chat-structured-composer-transport.ts, and whether Stop shows and what it does moves to
structured-agent-session-stop-control.ts. Behavior is unchanged; the runtime cast on the composer's
'local' | 'remote' is now a typed return.

* test(native-chat): read registered agents by id, as main's structuredAgentsReadBy now takes

Main's A3 squash changed structuredAgentsReadBy to take agent ids; this branch's test still passed
{ agent } objects (CI typecheck TS2322).

* fix(acp): a first reopen warns when Grok forgets a chat that exchanged turns

A Grok session the chat created was treated as one Grok never saved, so
when Grok reported it missing on the chat's first reopen, Orca swapped in
a fresh session silently, even after completed exchanges: the person saw
the old messages while Grok had forgotten them.

The launch now counts a created session as never saved only when the
chat's journal, read at the failed reopen, holds no turn of that Grok
session; anything else, an unreadable journal included, takes the normal
path: the fresh session is recorded as replacing the old one and the one
warning row is written.

* Add setting-gated OpenCode server native chat integration

* Use checked native event ordinals in timeline identities

* Validate nested account tags and preserve bounded restore failures

* Keep native history restoration within its session module

* Check native protocol records and support registered-agent drafts

* fix(acp): a start writes the warning row an earlier attach failure dropped

The row saying Grok forgot the chat was written only into the attach's
deferred sink, while the fresh session's link was saved earlier. An attach
failure, quit or crash in between dropped the row forever.

Every start now derives the owed rows: each conversation the chain says was
lost to a failed restore gets its row unless the chat's journal already holds
it. The row now names the lost conversation rather than the fresh session, so
a row an unused replacement wrote still counts after Grok supersedes it.

* fix(acp): a question during a turn the agent began itself gets its cancelled reply

A question or other card-opening request the agent sends while no prompt of
Orca's runs (a turn it began itself, as when a background task wakes it) now
gets the agent's own cancelled reply and opens no card, the same rule
permissions already follow there. Nobody is waiting on that turn. A plan the
agent shares in it is still shown.

* Update launch and status tests for registered structured agents

* test(native-chat): import the unfinished-work capture from the module that owns it

Main's reasoning sweep test (#19221) imported it from the dead-generation settlement, which
this branch split it out of.

* Run OpenCode structured chat over ACP

OpenCode joins the ACP lane as two launch rows (opencode and opencode2 run
'acp'), with a dialect that reads command output and exit codes from
OpenCode's metadata and labels its project-wide 'always' grant honestly.
The launch forces OPENCODE_CLIENT=acp and turns OpenCode's question tool
off, strips Orca's terminal status overlay from the child, and pins the
chat's OpenCode account as a tagged locator (managed profile, or data and
state directories). Image prompts go as ACP image blocks when the agent
advertises them, and restart recovery reads OpenCode's own store for
presence only.

* OpenCode 2 over ACP only for its default account

OpenCode 2's ACP command runs inside the user's own background service,
which the chat's environment never reaches, so a managed profile cannot be
honoured: while one is selected a new OpenCode 2 chat opens in the terminal,
and a record pinned to one never launches.

* Test OpenCode over ACP against its recorded traffic

Seven scrubbed recordings of opencode acp (1.18.31 and 2.0.14) replay through
the adapter: an approved shell command with its exit code and the
project-wide always label, deny on both majors, the 1.x subagent approval
that never arrives (Stop still ends the turn), the 2.x subagent approval
answered on the chat's session, the stray 1.x file write, and a provider
error shown in OpenCode's own words. The replay now plays recorded error
answers too.

* Name the normalized tool update for what it is

* OpenCode over ACP: 1.x only, honest grants, sound recovery

Review round 1:
- OpenCode 2 stays on its terminal-backed chat: its acp command runs in the
  user's own background service, which a chat's environment and account pin
  don't reach, and a declined local launch has no terminal fallback yet.
- The model-pick launch change is out of this PR (it could leave a worktree
  behind when the host later chose the terminal).
- 'Always allow' reads 'Allow for this chat': OpenCode 1.x keeps it in the
  chat's own process.
- A file write OpenCode repeats to a client with no file system leaves no
  row.
- Images are bounded by the encoded prompt line, not just raw bytes.
- Restart recovery: each accepted send claims its own stored copy first, and
  queued cards never trigger the read.
- A send whose child ended during the image read settles as not sent.
- An account Orca cannot pin (inline auth, relative directories) is refused
  as unsupported; default XDG directories the user never set stay unset; the
  retired shared plugin directory and data-account variables never reach the
  child; newer locator fields are tolerated.

* Round 2: recovery reads the chat's own database, honest grant label

- Restart recovery resolves OpenCode's database under the chat's own home,
  not Orca's, now that default data directories stay unset.
- 'Always allow' reads 'Allow until OpenCode restarts': Stop, quitting and an
  idle chat's stopped agent all end it.
- Only a message with images is refused for passing one protocol line, so a
  long text message never reads as an image problem.

* refactor(native-chat): keep the session host within its line budget after main's Stop work

Main's #25949 left the host at exactly its 300-line budget, and this branch's acquire-abort wiring
adds one line. The reveal module now comes in as a namespace import, as the host already does for
its other helper modules, and the earlier reorder of two type imports is undone.

* test(codex): move the stdout-before-exit test into its own file

Main's connection test file is at its 800-line test budget, and this branch's exit-order test
pushed it over (CI lint, max-lines).

* Keep the record's path bound, which the launch directory now uses

* Narrow the refusal in the prompt content test instead of casting it

* Type the replayed error answer as the protocol's error

* Read the replayed message through the replay's own schema

* Match recovered OpenCode messages through main's send fingerprint

Main gave every stored send one fingerprint helper (it drops the sender field); restart recovery now hashes stored messages through it instead of rebuilding the hash by hand.

* Keep an `opencode` that is 2.x on the terminal chat

Homebrew's `opencode` is 2.x now, so the binary, not only the agent name, decides whether OpenCode's acp runs in-process. An ACP launch spec can name the --version releases it runs on; OpenCode takes stable 1.x from 1.18.31. createSupport asks the binary a launch would spawn, with the environment a launch starts from, so a desktop launch of 2.x opens the terminal chat; every launch asks again before spawning, so a binary swapped after create never starts. The --version probe is shared and bounded, and a version it cannot read counts as unsupported.

* Type the mocked login-shell environment instead of asserting it

* Run the OpenCode the Command setting names, and say why a version check refused

The structured OpenCode and Grok chats resolved a bare `opencode`/`grok` from PATH while detection and the terminal chat honoured the per-agent Command setting, so a user who pointed Orca at a specific binary got another one (and, with the version gate, another version) in the structured chat. resolveStructuredAgentCommand now takes any agent id and the stock binary name; the ACP launch and createSupport's version check both resolve through it, so they ask about and spawn the same file. An unrunnable Command is admitted at create and refused at launch with agentCommandNotRunnable, as Claude and Codex do, instead of silently opening a terminal.

The --version probe now logs why it refused (exit code, timeout, oversized or versionless output, unsupported release), with the binary's own first words, so a refusal on a real host can be diagnosed. A real-probe test drives createSupport and the launch against stand-in binaries: per-agent PATH, Command setting, 2.x first on PATH, and an unrunnable Command.

* Say why a new OpenCode tab opened a terminal, and ask the host again for its agents

Live QA run 5 opened the terminal-backed OpenCode chat with the private 1.18.31 set as the Command and on the per-agent PATH, and Orca's main log had no version-check line. The host never decided: a local launch routes to the structured chat only once the renderer has learned the host's registered agents, and that list was read once at startup, so a read that failed (or was skipped while the local runtime's capabilities were unknown) kept OpenCode on the terminal for the whole session. A launch that finds the list unlearned now asks the host again, so the next launch opens the chat.

Every point that sends a launch to a terminal now says so, naming the check and never a value: the renderer's route decision ([agent-launch-route], in the renderer console), createSupport's refusing check ([structured-create-support]), a version check that passed as well as one that refused ([agent-cli-version]), an OpenCode account that cannot be pinned ([opencode-account]), and a host launch the default asked to be a chat ([agent-launch]). createSupport's composition moves out of the unchecked runtime class into structured-agent-launch-support.ts.

Tests reproduce run 5: the host's create check with HOME and every XDG directory in the rig, the Command and per-agent PATH naming a stand-in 1.18.31, and the account pinned to the rig's XDG directories; and the renderer route with those settings, before and after the host's list is learned (ablation: without the re-ask the route stays on the terminal chat).

* Type the mocked login-shell environment in the OpenCode probe tests instead of asserting it

* Pass the OpenCode Command and environment through the route test's store helper, which types its settings

* Open OpenCode's chat on the first new tab, and learn the host's agents when they become available

The renderer routes an OpenCode tab to the structured chat only once it has learned the host's registered agents, and that list was read once at startup with nothing to read it again, so a read that came too early or failed sent every OpenCode tab to the terminal chat.

- agentSession.agents answers from the registrations the host is built from, as createSupport already does, instead of installing the host first: the answer is fixed for this build, so it no longer waits on, or fails with, opening the chat journal.
- The local list is read again when the local runtime's capabilities land, so a startup read made while they were unknown is redone on the event that makes it possible.
- A new tab whose chat route waits only on that list asks the host for it and opens nothing until it answers or 3 seconds pass (the local admission bound), then decides: the first launch opens the chat. A host that does not publish its agents is never waited on.

* Type the new-tab route request as a launch request, which carries the delivery callback

* Declare that agentSession.agents answers from the registrations in the cross-version manifest

Older builds read the list from the installed host, newer ones from the registrations it is built from; the manifest now asks only that Codex is listed, which holds for both.

* Wait for this computer's runtime too before routing a new tab during a slow startup

Field QA: soon after a slow startup, a new tab opened the terminal chat because the local runtime had not reported its capabilities yet, which routes straight to the terminal; the bounded wait covered only an unlearned agent list. A new tab whose chat route waits only on either answer now asks for it and opens nothing until it comes or 3 seconds pass, then decides; the [agent-launch-route] line still names the blocker if it falls back. Tests that never meant to exercise startup now start with the runtime's answer known.

* Wait for the probe that lands the runtime capabilities, not one that failed at startup

When the startup probe fails (a rig whose terminal daemon failed to start), asking again answers null at once, so the new-tab wait ended early and the tab opened the terminal chat. The wait now subscribes to the event that delivers the capabilities, probes once, and stops at the 3 s cap; it applies to both routes, so OpenCode (whose blocker is the unlearned agent list) and Claude (unknown capabilities) both wait for it. Tests run the real capability module with a failed startup probe: Claude and OpenCode open their chat when a later probe lands within the wait, and the terminal chat after the cap otherwise.

* Answer an OpenCode account Orca cannot pin at create support, not as a failed chat

Live QA with inline OpenCode credentials (OPENCODE_AUTH_CONTENT) passed the version check at create support and was refused only at create, so the window showed a failed "OpenCode Chat" with Retry where a declined launch opens the terminal chat. Create support now resolves the account a new chat would pin through the same read-only resolver the model catalog and create use, and answers no, logging the account-pin check, when the host refuses to pin it; any other failure is left for create to state. Create's own refusal stays as the backstop.

* fix(opencode): start the chat with the user's OpenCode environment as-is

An OpenCode chat with no managed profile no longer records data/state
folders or a database choice, no longer refuses inline credentials or
relative paths, and launches with the user's environment unchanged.
Drops the create-support account-pin probe that only those refusals used.

* test(cross-version): build each method's params in their own file, under the line limit
2026-10-07 16:26:27 -07:00
..