Files
orca/docs/reference/agent-status-store.md
T
Brennan Benson 2bd656c384 feat(native-chat): a Claude subagent waiting on a permission prompt reads as waiting (#22634)
* feat(native-chat): a Claude subagent waiting on a permission prompt reads as waiting

A subagent's permission request reaches the parent session's callback naming the
subagent that asked (agent_id) and the tool call it gates (tool_use_id). The
pending request is recorded with the asking agent. On every drain the child-work
producer re-derives which children a pending request blocks and hands that set
to the Claude child decoder, the one owner of each child's live edges: a blocked
child reads waiting on every live edge it reports, and a child that starts or
stops waiting is a live edge of its own. Answering, denying or cancelling the
request returns the child to its prior live state; nothing is stored beyond the
pending requests.

A live task_updated carrying an error now reaches the record as the child's last
message, without an ending or a new state.

The replay test drives a scrubbed capture of the real CLI (foreground allow,
deny, interrupt, background allow, and the main agent's own request) through the
real adapter into the host's child records.

* test(native-chat): a subagent's request names it before its tool call is read

* docs(agent-status): a subagent asking for approval waits in every lane; the parent row keeps the session's own attention

* test(native-chat): hand canUseTool the asking agent without widening the helper's cast

* test(native-chat): an interrupted Claude subagent settles cancelled, not failed

A captured interrupt shows the spawn call's error result ("The user doesn't want
to proceed…") arriving before the subagent's own `task_updated {status: killed}`.
The spawn result ends nothing (the child ends only on its own terminal frame), so
the child stays live until its `killed` status settles it cancelled. A genuine
failure, captured with the subagent on a model that does not exist, sends its
`failed` status before the error result and still ends failed. Both captures now
replay through the real adapter into the host's records.

* fix(native-chat): a Claude subagent's prompt makes the parent row wait, not block

A subagent's pending prompt made the whole session `attention`, which reads as
the main agent's own `blocked` and outranks the fold's waiting arm, so the
parent row read blocked where a CLI Claude parent reads waiting. The main
agent's state now reads only its own pending prompts.

- A Claude prompt row carries the linkage of the agent that raised it: the one
  the permission request names, or the owner of the tool call it gates. The
  same join decides which child reads waiting, so the two cannot disagree.
- The status summary projects the session's own status from root prompts only;
  every other reader (delivery gates, teardown, restart) still asks whether
  anyone is waiting on a human.
- An answer keeps the prompt row's linkage by the journal's own rule: a
  revision that names no producer keeps the row's existing one.
- The child-tool queries gain the prompt's producer, so a prompt row and a
  child record answer "which agent" from the same join.

* test(native-chat): say which ids the permission capture scrubs and which are its own

* fix(native-chat): the status clock dates attention by the session's own asks only

The session's status is now `attention` only for its own pending prompt, so the
clock's fallback to a subagent's ask could no longer be reached, and it read the
journal by a different rule than the status it dates. Both now read root prompts.

The journal also stamps a Codex subagent's prompt with its thread (#22532), so a
Codex child's approval is that child's wait in the Codex lane too. Two tests
written for the earlier rule are updated: a subagent's ask leaves a running
session `working` on its turn's clock, and a Codex child's answered approval
leaves the settled parent's Activity row done with nothing unread.

* fix(native-chat): a completion still says the user is asked when a subagent asks

The turn-completion feed marked a completion `awaitingUser` from the status
summary's `attention`. The status now means the session's own agent is waiting,
so a subagent's pending approval stopped reaching the completion. The projection
now also says whether anyone is waiting on the user, as the delivery gates,
teardown and restart ask it, and the completion reads that.

The waiting-subagent replay answers its prompt with the adapter's current
response shape.

* revert(native-chat): a live Claude task's error stays out of the child's last message

No capture shows a live task_updated carrying an error, and it is unrelated to
a subagent waiting on a permission request; it leaves this PR.

* test(native-chat): settle the Claude session's startup before replaying a subagent's request

A startup frame drained child work during the first await, so answering a
request freed the child even with the answer's own republish removed.

* fix(native-chat): a Claude subagent's prompt row names it as its other rows do

The prompt row stamped only the asking agent's id, so a nested subagent's
request lost the agent that spawned it, its spawn call and its run. It now takes
the linkage the asker's own rows take: the gated tool call's, when that names the
same agent, else the one resolved through the agent's spawn call. The provider's
agent id stays the asker's id.

* fix(native-chat): a subagent's request makes the parent row wait without a child record

The parent row learned that a subagent needed the user only from that
subagent's child record, so a request no record carried (a Codex child the host
never registered, a Claude task past the live cap) left the row working or done
while the approval card sat in the chat.

"Someone in this session must answer" is now one derived session fact. The
projection names two facts instead of a mode flag: the main agent's own status
(attention only for its own request) and structuredAgentSessionAwaitsUser (any
pending prompt). The status summary publishes the second as an optional
awaitsUser, and the shared fold reads it: the main agent's own ask is blocked,
otherwise awaitsUser or a waiting child record makes the row wait. Every caller
picks the fact it means: the completion edge's awaitingUser and the delivery
gate read awaitsUser; the quit snapshot folds the same two inputs the sidebar
does.

* fix(native-chat): a client that predates awaitsUser still reads a subagent's request as attention

A status summary's status is now the main agent's own, so a client built before
the split would read a subagent's request as working (or idle) and fold it with
code that has no awaitsUser input. Clients advertise
agent-session.status-awaits-user.v1; at agentSession.subscribeStatus the host
sends any client that does not the pre-split summary: attention whenever
awaitsUser is set, without the main agent's own tool line, verdict and clock.
The feed and every in-process reader keep the canonical summary. Transitional,
like the turn-item downgrade.

* test(native-chat): a Codex subagent's approval makes its settled parent's Activity row wait

The test pinned the parent row done while a Codex child asked, through a harness
that fed no child records, so it proved nothing about the ask. It now drives the
ask twice through the real host status store: with no child record (the
session's awaitsUser alone) and with the child's own record waiting from
thread/status/changed. Both read waiting with needsAttention while the ask is
open, then done with nothing unread.

* docs(agent-status): a subagent's request reaches the parent row through awaitsUser in every structured lane

The store reference said a Codex child's request still read as the main agent's
blocked and that only the Codex hook lane fed a waiting child. Both structured
lanes stamp the asking child and feed child records, and awaitsUser carries the
request when no record does. The liveness comment goes back to main's: a child's
blocked is a failed task on an older host's legacy rows.

* fix(native-chat): the restart dialog still headlines a subagent's pending approval

The quit snapshot now records the main agent's own state, so a subagent asking
while the main agent worked recorded `working` and the dialog said "Was
mid-reply" where it used to say "Waiting for your approval". The headline now
comes from the snapshot's pending prompt, whoever raised it, with the existing
copy; `state` stays the main agent's own.

* test(orchestration): a subagent's pending approval holds structured mail delivery

Scoping the delivery gate to the main agent's own request left every gate test
green; a subagent's request now has its own case.

* fix(native-chat): a subagent's request is dated by when it was raised, on every client

Since the summary's clock became the main agent's own, nothing dated a wait
that only a subagent's request held: a pre-split client was sent attention with
no clock, where the old host dated it by the subagent's prompt, and a new
client's waiting row fell back to the time it first saw it, so after a reload a
request the user had already read could read unread again.

The session fact is now when someone started being asked: awaitsUserSince, the
oldest pending prompt whoever raised it, and its presence is what awaitsUser
meant. A row waiting on someone else's request takes that as its clock; the
downgrade for a client without the capability dates its attention by it, which
is what the old host published. A cross-version test pinned to the last
pre-split release runs the same journals through that release's projection and
through this one plus the downgrade, and compares the whole summary. The Codex
end-to-end test also reads the host's own status row, and keeps a read ask read
through a later row and a reload.

* test(native-chat): the pre-split parity check compares only the fields the split owns

An additive summary field is safe for old clients, so comparing whole summaries
against the pinned release would redden on one. The wire comment now says how
the downgrade dates attention: the main agent's own oldest ask, else
awaitsUserSince.

* test(runtime): an aged host-held working summary states that nobody is asked

The test built its working summary by overriding the status of a published
approval summary, which still carried awaitsUserSince, so the row correctly
read waiting. It now drops the request as its scenario says.

* fix(native-chat): the chat's subagent block says waiting when the strip does

While a Claude subagent's request was open, the sidebar and the composer strip
read waiting but the subagent block in the chat history a few pixels above
still read "Kicked off 1 subagent working": it shows the journal's roster
state, and the journal records no wait.

The structured chat now hands its transcript the subagents the strip shows
waiting, read from the host's child records through the strip's own row model
and matched by the provider id the roster names each one by. A running entry
the host says is waiting reads waiting in the group row, its entry and its
section head, with the strip's word and the question colour; it reads the
journal's state again as soon as the host stops reporting the wait.

* fix(native-chat): a collapsed subagent group shows a wait beside a failed sibling

A failed sibling took the group row's one alert slot, so a group with a waiting,
a working and a failed child read "1 working +1 failed" and hid the wait; it
now reads "1 working +1 waiting +1 failed". The waiting set keeps its identity
while a child frame changes no wait, so the transcript's subagent rows do not
re-render on every frame, and the test of a wait ending now updates one mounted
row instead of remounting it.

* refactor(claude): one needs-input state on the parent; the asking subagent alone reads waiting

Drop the split of the main agent's own status from a session-wide "someone must
answer" fact: awaitsUserSince, the agent-session.status-awaits-user.v1
capability and its old-client downgrade, and every reader change that only
consumed them (fold, equality, ingest, delivery gate, turn-completion feed,
quit snapshot, resume headline, status clock, status bridge, attention
dispatch) go back to main. The parent row again reads one needs-input state
for a pending request whoever asked, dated as before.

Kept: a request's owner recorded once on its prompt row with full producer
linkage; the asking subagent's own record reads waiting, re-derived on every
update; the chat history's subagent block reads that same state; an answered
subagent request stays in its subagent's group.

A subagent now waits only on a request the user can still answer (its card
open, no answer underway), and the adapter frees it before the host records an
answer or dismissal. So a waiting child record always sits beside the pending
card, and main's fold never reads the parent as waiting on it: no window after
an answer, and no ~3 s wait after a card dismissed by Stop.

* fix(claude): a subagent waits only beside its committed card

A subagent's wait was pushed to the host as soon as its request arrived,
while the request's card row reached the journal at least a microtask later.
So every subagent request published the parent row as waiting before
blocked (the main agent's own fold reads a waiting child that way), and
Activity got an extra unread "waiting" event that main never shows.

The card is now the one record of an open request. The translator records
the asker on the card once (its row's linkage) and counts the card open only
after the sink confirms its rows landed, then publishes the wait; anything
that closes the card (an answer underway, a dismissal handed to the host,
Claude's own withdrawal, the session's end) frees the subagent first. So
every publish that shows a subagent waiting also shows its pending card, and
the parent reads one needs-input state, exactly as on main.

This retires the registry's view of pending requests (unclaimed(), the
asking-child join) and the translator's holdsOpen. The prompt row's linkage
takes one rule: the agent the provider names, else the gated call's owner.

The parent-row proof now runs through the real deferred sink, durable
journal and status feed, publishing as production does, and checks at every
publish that waiting subagents have pending cards and that the parent row
matches a host fed no waits.

* fix(native-chat): a closed sink's dropped writes never read as landed

The sink's written() resolved ok when the sink was closed with writes still
queued, so a subagent's prompt card could count as open with no row in the
journal. written() now reports a close that dropped writes admitted so far
as not landed; drained() and lifecycleBarrier() keep reading a closed sink as
settled.

Tests: a card never opens when its sink closes first; a card Claude withdraws
while the sink holds the cancelled row back closes at once; two subagents
asking at once, and the main agent asking beside a subagent, keep the parent
row as before with each waiting subagent beside its own card; a process that
dies mid-request leaves no subagent waiting.

* fix(claude): a withdrawn subagent request frees its child before its card closes

Main's sink now hands each write to the journal as it is submitted, and an idle journal commits it
and runs the publication at once. Claude's own withdrawal of a subagent's request therefore closed
the card and published the parent row before the child's wait was freed, so one publish showed the
subagent waiting beside no pending card (fg-interrupt replay). The child's wait now also requires the
request to still be open in the registry, and a withdrawal republishes child work before the
journal takes the close.

* refactor(claude): trim subagent request waiting to the common pattern and its essential tests

The chat history no longer marks a subagent block as waiting: the approval card itself carries the
request, and the asking subagent's row in the sidebar and composer strip reads waiting, as before.
NativeChatWaitingSubagentsProvider, native-chat-waiting-subagents.ts and their renderer changes go.

A subagent waits while its request is still open and unanswered in the prompt registry and its card
has landed in the journal. The registry check also covers a withdrawal under backpressure, so the
card list no longer filters pending cancellations itself.

Tests: one integration file replays the captured CLI frames through the real adapter, sink, journal
and status feed (renamed claude-subagent-permission-request.test.ts), with the asking subagent's
state timeline, attribution, nested linkage and a card write that waits for the journal. The
producer-harness waiting test, the redundant prompt-card cases, the harness reducer swap and three
unused captures (deny, interrupt, main agent, failed subagent) are removed.

* test(claude): pin the parent row's dating when a subagent asked first, and narrow the oracle's claim
2026-10-03 12:45:51 -07:00

38 KiB

Agent status store

Status

The current boundary is PR 2A: structured sessions use the hook server's fully scoped canonical store; unbound PTY/relay evidence remains in an isolated legacy adapter. Do not remove the renderer bridge or its publication filters in this slice: they still carry native-chat child rows.

The sections below record the original 2026-09-09 rollout. Its PR 1a and PR 1b have landed; its proposed PR 2/3 sequence is superseded by that boundary:

  1. main-only: every producer writes into one store and worktree ps reads it, split into 1a (structured sessions join the store) and 1b (the runtime's duplicate retained store is deleted);
  2. renderer: the sidebar becomes a subscriber and stops re-deriving rows;
  3. shared: one worktree-status rollup and one freshness rule for every reader.

The problem this solves

Orca shows "what is this agent doing" in four places: the desktop sidebar, the orca worktree ps command, the mobile app, and the agent dashboard. Before #19217 those readers did not even share their inputs. After #19217 they share the structured-session mapping and nothing else.

An audit on 2026-09-09 found six producers and three consumers, and three separate copies of the same row inside the main process alone:

Main-process copy Keyed by Owned by Persisted Evicted
hook server lastStatusByPaneKey paneKey src/main/agent-hooks/server.ts last-status.json tab close, pty exit, hydrate
runtime RuntimeAgentRowStore paneKey runtime-agent-row-store.ts (deleted in PR 1b) no pty exit only
structured feed published sessionId src/main/native-chat/agent-session-wire/structured-agent-session-status-feed.ts no never (a broadcast cache)

The second copy is a duplicate write: the OSC status parsed in main is forwarded to the hook server and retained in the runtime store from the same call (orca-runtime-create-terminal-side-effect-command-code-detector.ts). The third copy is keyed differently and never reaches the hook server at all, which is why worktree ps grew its own adapter for it in #19217.

Each reader then applies its own precedence and freshness rules, so the same pane can legitimately read differently on the desktop, on the phone, and in the CLI.

The rule

The execution host owns agent status, in one store, and every reader subscribes to it. This follows the boundary in ssh-execution-boundary.md: the host that runs the process is the only party that can observe it, and the client is never authoritative for execution state.

Three consequences:

  • One store per execution host. A remote host keeps its own store and the client mirrors it down, as the web-session mirror already does. Mirroring is not merging: a client never writes its observations back to a host.
  • Precedence is decided once, at write time, with provenance recorded on the row. Readers never re-adjudicate hook versus terminal versus structured.
  • Readers keep only presentation policy and user facts: the 30-minute display decay, acknowledgements, dismissals, unread. Those stay reader-side but become one shared implementation (PR 3).

The store already exists

The hook server's state is that store today for every PTY-based agent. The audit established:

  • hook HTTP posts, the WSL and SSH relay receivers, and main's own OSC parse all converge on the same applyNormalizedStatus path, stamped with the authority id main-agent-hooks;
  • it alone holds pane authority: launch tokens and their hashed commitments, retired-pane fences, pane-key aliases, per-connection ordering watermarks, and the evidence-age map that must outlive a transport clear;
  • it alone persists, with a seven-day hydrate window and the restoredUnconfirmed stamp that keeps a hydrated row from ever reading as live truth;
  • it already fans out to both renderer windows over agentStatus:set and agentStatus:clear, and serves agentStatus:getSnapshot.

Nothing else in main carries those guarantees, and building a second store with them would be the wrong direction. So the design is not "add a store". It is: route the two producers that bypass the hook server through it, then delete the copies.

PR 1a: structured sessions publish into the store

No renderer behavior changes. The sidebar keeps receiving the same IPC events it receives today, plus structured-session rows it currently derives itself.

Structured sessions publish into the hook server

The structured feed keeps its job of projecting a session's journal into a summary and streaming it to subscribers. On every publish it additionally ingests the summary into the hook server as a status row:

Row field From
paneKey structuredAgentSessionPaneKey(tabId, sessionId), the key the renderer already uses; its leaf is UUID-shaped so pane-key validation accepts it
tabId structuredAgentSessionTabId(sessionId)
worktreeId summary.workspaceId (a folder workspace id is a valid value)
state structuredAgentSessionAgentStatus(summary).state: the lead's own status folded with the live child records the store holds for the session (not the summary's task list), so a settled lead whose subagent still runs reads working
workingMode 'monitoring' from the same fold when watch loops are the only live child work; omitted otherwise, which clears it on the row
mainAgent the main agent's own state before the fold, its last-turn verdict (summary.turnOutcome, present only while idle) and its own clock; see "The main agent fact" below
structuredHost 'owned' while summary.hostExecutionOwned is set, otherwise 'held'; worktree ps derives its row's structuredHostOwned from it
prompt, tool, last message, model, provider session the summary's fields

Sessions with no request (status === null) produce no row. A request is a turn record, an assistant message, a user message the provider journaled itself (history, an older host), an accepted or unanswered send, or a send the agent or its start refused; a send that was withdrawn, or left undelivered by a restart or a close, fails nobody and makes nothing listable. summary.turnOutcome is the latest request's verdict: its turn's outcome, or failure for a send the agent or its start refused (a send that joined a running turn is answered by that turn). A turn the provider gave no outcome reads as its host-observed end through agentTurnVerdict: interruption for an interrupted lifecycle, unconfirmed for an unverifiable one. It is derived on each read, never journaled. The row also publishes interrupted from mainAgent.outcome, exactly as the hook lanes do. When the host revokes live ownership the row is re-set without the flag; when the host closes or evicts the session the row is dropped. Both already exist as feed events (revokeLive and the roster filter in liveSessionSummaries); PR 1 turns them into store writes.

Dropping the session from the host's map and dropping its row are one operation, forgetStructuredAgentSession. The store keeps a row until told, and a host-owned row bypasses the staleness check, so a deletion path that forgot the row would strand a permanently working-looking agent.

Two rules the ingest must keep:

  • Never persist a structured row. The journal is the durable truth for a structured session and the host republishes on restore. A structured row in last-status.json would hydrate as restoredUnconfirmed and then fight the live republish. The serializer skips rows carrying structuredHost, and hydrate drops any such row found on disk. Applying one therefore also skips the persist schedule: the walk and stringify could only reproduce the file that is already on disk, once per debounce window for every streaming chat.
  • Never let it fight a hook row. A structured session has no PTY, so no hook or OSC event carries its pane key. The ingest still goes through the disposition gate so a retired pane key is refused like any other.

Applying one does still run both status fan-outs, and that is intended rather than incidental. notifyStatusChangeListeners is what feeds agentAwakeService's power-save blocker, and subscribeEnrichedStatus is what feeds AgentSessionTransitionRecorder's stats, so joining the store enrolls native chats in both. A working native chat is real work and should hold the machine awake exactly like a PTY agent does.

The drop side routes through dropStatusEntry, not clearPaneState: a pane-status-clear reaches the renderer, and until PR 2 the renderer's own feed bridge is that pane key's writer. It also passes preserveResumeIdentity: false — the providerSessionOnly remnant a dismissed pane keeps exists so the agent can be resumed in that pane, and a structured session has no pane and keeps its resume identity in the record store. Like every other dropStatusEntry caller, it emits no pane clear, so a session dropped mid-working leaves AgentSessionTransitionRecorder holding an open stats session until its LRU evicts it; that gap is shared with the user-dismissal path and is not specific to structured rows.

The ingest lives in the feed, not in structured-agent-session-host.ts, which sits at the file-length cap.

worktree ps becomes a reader

The structured adapter added in #19217 is deleted, and structured rows reach worktree ps through the same snapshot as every other row. The retained-versus-hook reconciliation in collectRuntimeWorktreePtyAgentSources stayed until PR 1b removed the store that fed it. What this step settles is the admission gate that decides which rows a worktree listing may show:

  • a hook or OSC row needs its tab mirrored or a connected pty, as today, and SSH rows stay exempt because their tabs may exist only remotely;
  • a row carrying structuredHost is admitted while the host holds the session, and the host's drop on close is what removes it. No tab-mirror requirement: a structured session's tab lives in the renderer's own tab state, and a headless host has no renderer to mirror it from. That argument only holds if the headless host is itself wired to the store, which is a separate obligation per entry point: the Electron hosts (desktop and orca serve) share main-process-runtime-service.ts, and orcad constructs its own runtime in src/main/orcad/orcad-entry.ts. A host missing that wiring lists no agents at all, not just no structured ones, because worktree ps reads the same snapshot for every row.

The freshness bypass for host-owned structured rows already exists in isFreshNonDoneAgentStatus; with the flag now on the row it becomes the only path, and the hand-rolled check in runtime-worktree-agent-rows.ts goes.

Wire compatibility

AgentStatusIpcPayload gains one optional field, structuredHost, and the worktree ps row gains structuredHostOwned. Under rule 1 of remote-wire-compatibility.md both are safe: an old client ignores them. worktree ps rows keep their shape and vocabulary, so the mobile app sees no change.

Until PR 2 the main process does not forward structured rows to the renderer over agentStatus:set or agentStatus:getSnapshot. The renderer's feed bridge still writes those rows itself, and forwarding them too would give one pane key two writers. Removing that filter is the first step of PR 2.

The main agent fact

Claude, Codex and Grok hook rows and structured-session rows publish the combined state and, beside it, the main agent's own state as payload.mainAgent. Other agents' rows and terminal-title-only rows carry none, and readers fall back to state:

mainAgent?: { state: AgentStatusState; outcome?: AgentTurnOutcome; stateStartedAt: number }

state still answers "what should the user see" and folds live child work in, so a settled main agent whose subagent still runs reads working. mainAgent answers "what is the main agent itself doing", which the fold used to destroy at publish time; every guard that reconstructed a fragment of it (fromChildWork, the persisted claudeLeadBoundaryChildOnly flag) now reads mainAgent instead of a stored copy. A Claude row whose mainAgent is done while a child agent still works (including a child's permission wait) refuses OSC, which carries no child identity; the children's own lifecycle hooks settle it. outcome is the recorded verdict on the main agent's most recent finished turn, present only while mainAgent.state is done. It is reported by the provider, or is a cancellation Orca inferred from the user's own interrupt keystroke, or a superseded the host recorded when a newer request replaced a structured Claude turn before it ended (it names no sender, and sets no legacy flag), or, on a structured row whose turn the provider gave no verdict, is what the host observed of its end: interruption (a proven death nobody asked for) or unconfirmed (an end it cannot prove, never success). The journal's turn outcome, by contrast, stores only recorded verdicts: the provider's, a cancellation, or the host's superseded; interruption and unconfirmed are derived from the turn's lifecycle state and never stored. A plain end of turn carries none, because absent means unknown and a provider that omits its interrupt flag must not turn a cancel into a success. In the Claude hook lane the cancellation comes primarily from Orca's own inferred interrupt (markClaudeLeadTurnInterrupted), because current Claude sends no hook at all on a cancel and no is_interrupt on Stop; that flag on a turn boundary remains a secondary source for builds that send it, and StopFailure maps to failure.

Readers decode the verdict through one accessor, agentMainAgentVerdict, which reads the main agent's own state, not the combined row's: mainAgent.outcome while mainAgent.state is done, then the legacy interrupted flag as a cancellation, which alone needs the combined done. So a main agent that failed while its subagents still run has a verdict on a working row. Every copy of a row (state-history entries, sleep records, worktree ps rows) takes the verdict through agentVerdictFields, which carries interrupted and the whole mainAgent (state, outcome and its own clock) together, so a copy agrees with the row and can date a failure by mainAgent.stateStartedAt.

Display reads the verdict through agentVerdictDisplayMark. A fault marks the agent failed whatever the combined state, because it is news the user must see even while subagents run: a failure, and an interruption, a turn cut short by anything other than the user or a newer request. A user's stop (cancellation) marks it interrupted, drawn in the muted tone with the row text "Interrupted by user"; a turn a newer request replaced (superseded) marks it interrupted in the same muted tone with the row text "Interrupted"; and unconfirmed marks it unconfirmed, all only on a done row, so a stopped or finished main agent with live child work still reads working. The folded turn header follows the same mark: "Failed after N", "Interrupted after N", or "Worked for N". Each subagent keeps its own row and state. Container rollups (worktree card, terminal tab, Cmd+J) rank a pending question first, then a failure, then live work, then an unconfirmed end, then a user's stop, then done. On the worktree card, a failure retained after its agent's pane went away has no expiry, so it ranks below live work and above an unconfirmed end. Lifecycle waiters keep reading the combined state.

Policy splits the verdict two ways. Clean-finish policy (hibernation, pane ownership, the star-nag value moment) treats a failure, an interruption and an unconfirmed end like a cancellation (agentTurnEndedUncleanly). Attention (completion time, Smart Sort, sticky retention, Cmd+J Recent) demotes only a turn ended on purpose, the user's stop or a newer request that replaced it (agentTurnEndedOnPurpose); a failure, an interruption or an unconfirmed end ranks like a completion.

Admission is one function, normalizeAgentStatusPayload, on the relay wire, IPC and disk. A malformed mainAgent drops the field and keeps the row. Old hosts send none and readers fall back to state. Hook rows persist it inside the payload; hydration maps an older row's claudeLeadBoundaryChildOnly: true onto mainAgent: { state: 'done' } when the row has no mainAgent, and never writes the flag again. Hydration seeds the Claude main agent record straight from a saved mainAgent that is done, so the children's drain can still settle the row after a restart. claudeRunningNonAgentTask is persisted alongside because it is the one child-work fact mainAgent cannot express: a shell running beside the main agent, whose liveness hydration does not restore. Hydration seeds only a row that says false; a row silent about it stays unseeded. The row builder pairs the two facts in one place: a listener event restates the shell fact, and any other write (an OSC repaint, an inferred answer) keeps it only while mainAgent is unchanged. A child's sticky permission prompt still records the main agent's own progress and background evidence in the held row, and pushes the held row to subscribers when mainAgent changes.

Every lane, Codex included, combines through the fold. A child waiting on a human is a fold input (childWorkLiveness: 'waiting', derived from the child's own waiting state; a child's blocked means it failed and stays live work) and makes the row wait whatever the main agent is doing, unless the main agent is itself asking. The Codex hook lane feeds it from its child transcripts, and the structured lanes from child records, which read waiting for a Codex child thread's approval or input flag and for a Claude subagent's open permission request. Known divergences, pinned by name in the parity table (src/shared/main-agent-status-parity.test.ts) where they are reachable, so a reader does not mistake them for drift:

  • The Claude hook lane holds a child's permission wait in one slot on the displaced main agent record (waitingAgentId, stateBeforeWait), not on the child. It publishes the displaced state as mainAgent, but the next main agent event overwrites the slot, so the row stops reading waiting while the child is still asking, and a second asking child replaces the first.
  • In the structured lane a child's pending prompt also makes the session attention, which reads as the main agent's own blocked: one needs-input state whoever asked. A Claude subagent reads waiting only while the journal holds its card pending: from after the card's row is written until just before anyone closes it, so every publish that shows the child waiting also shows the session's attention, and the row never reads waiting for a Claude subagent's request.
  • The Codex hook lane drops its roster on a root Stop when it tracks no child transcripts, so a still-running or still-asking child stops holding the row.

How the main agent's turn ended is not a fold input. A cancel is a verdict on the main agent, carried as mainAgent.outcome: 'cancellation' (and, for readers that predate mainAgent, as the row's interrupted flag on a done row); it never retires a shell, scheduled check or subagent the turn left running. That work leaves the row only when its own inventory omits it or the session ends, so a cancelled turn with a still-running shell reads monitoring in every lane, and the parity table in src/shared/main-agent-status-parity.test.ts drives that story through all of them. The same rule governs the cancel Orca infers from Ctrl+C: for any row that publishes mainAgent, the inference is admitted only when mainAgent.state is working, so Orca does not treat a Ctrl+C at the idle prompt of a row held open by child work as a turn cancel (Codex also keeps the child-evidence guard, and a row without mainAgent keeps only that guard). The keypress itself is not inert, though: measured live, Claude 2.1.280 stops its background subagents on a single idle-prompt Ctrl+C (shells survive) and Codex 0.156.1 quits outright, so refusing the inference can leave the row showing a subagent its CLI already stopped. The synthesized row is the fold of the cancelled main agent with the child work the pane's owner can see: the local listener's roster for a local pane, the row's own subagents and shell fact for a relayed one, whose provider records live on the relay.

The store holds that verdict against restatements that predate it (server-cancel-verdict-latch.ts), because a relay never learns of a cancel the desktop infers and some TUIs emit late same-turn hooks. The hold is read off the row (mainAgent.outcome: 'cancellation'), never stored beside it, and dies on a new turn (a main agent prompt submission, a changed or explicit prompt, a session start) or the provider's own settled mainAgent. Child and replayed events under the hold keep the cancelled main agent and are re-folded with their own child evidence.

PR 1b: the runtime's retained row store is deleted

Landed. RuntimeAgentRowStore is gone, and with it the retained-versus-hook reconciliation in collectRuntimeWorktreePtyAgentSources. The hook server's store is now the only main-process copy of a PTY agent's row.

The five call sites

Call site Before After
orca-runtime-create-terminal-side-effect-command-code-detector.ts retain() second write of the OSC payload already sent to the hook server deleted; the event now carries the pane's terminalHandle and the hook ingest keeps the only copy
...command-code-detector.ts clearPty() drops rows on pty exit deleted; pane teardown already clears the hook row
orca-runtime-get-worktree-ps.ts values() fed retainedSnapshots deleted; the reader keeps only hookSnapshots
orca-runtime-serialize-agent-prompt-submission.ts getFreshExplicit() retained row first, hook rows second selectFreshExplicitAgentStatus, hook rows only
orca-runtime-prune-mobile-session-tab-group-layout.ts getFreshForMobile() pane key, then pty id selectFreshAgentRowForMobileTab: pane key, then terminalHandle

Both readers moved into runtime-hook-agent-row-selection.ts, which also owns RuntimeAgentRowSnapshot now that nothing retains one.

terminalHandle is the row's join back to its terminal

The retained store's only real extra was the pty id, and two readers used it. The plan said to stamp the event's ptyId into terminalHandle; that was wrong. A terminal handle (term_<uuid>) and a pty id are different identifiers, and getFreshExplicit was already comparing hook rows against a real handle. What landed instead:

  • AgentHookEventPayload and the runtime's terminal-status event gained an optional terminalHandle. The detector resolves it once per chunk through getAgentStatusTerminalHandleForPaneKey — the same lookup the renderer-facing IPC boundary already runs for every row, so the two surfaces cannot disagree about which terminal a pane is.
  • applyNormalizedStatus carries the handle forward when an incoming event resolves none. Only main's OSC parse can resolve one, so an HTTP hook post for the same pane would otherwise erase it.
  • It is never persisted. A handle belongs to the runtime that issued it, and a hydrated one could only rejoin a row to somebody else's terminal.
  • toAgentStatusIpcPayload publishes it, which also makes getFreshExplicit's long-dead handle comparison live: the runtime reads raw snapshot rows, and before this nothing ever stamped the field on them.

worktree ps uses it too. ConnectedPtyEvidence traded its flat ptyIds set for ptyIdByTerminalHandle, so a row still resolves the connected PTY behind it — which is both the working-terminal rollup's match key and the last rescue for a row whose pane binding was nulled by a controller incarnation change.

The change detector had to move with the store

retain() was not only a store: its boolean return was the signal that republished session.tabs for a status-only transition, which no title change covers (#7970). hook-status-session-tabs-invalidation.ts already mirrors that projection change set, including restore provenance and terminal-handle joins, so the replacement was to route the signal off the store rather than build a second comparator. installHookStatusSessionTabsRepublish now owns all three arms — enriched status, pane clear, and the status-drop tap a dismissal emits — and both hosts install it.

Both hosts, not just the desktop one

orcad constructed its runtime with no onTerminalAgentStatus, so main's OSC parse never reached the store there and the retained copy was the only carrier. Deleting it without wiring orcad would have made a headless host list no PTY agents at all. orcad-entry.ts now binds the producer and installs the republish signal, alongside the snapshot and structured sink it already had.

The intended behavior change

A row the user dismisses on the desktop leaves worktree ps and the phone at once, instead of lingering until the pty exits. One store means one dismissal.

Legacy numeric pane keys remain a bounded compatibility case. Persisted layouts register aliases to their stable leaf owners; an in-process OSC observation may also retain a numeric key only when the runtime supplies the matching tab, PTY, and terminal handle. HTTP and relay ingress still require a stable key or a registered alias, and numeric rows are never persisted.

PR 2: the renderer subscribes

With structured rows arriving over agentStatus:set, the renderer's StructuredAgentSessionStatusBridge no longer needs to write status; its unmount cleanup becomes a tab-close signal to the host. The IPC applicator is the single writer for observed status. The 2026-09-09 audit sorted the other writers:

Writer Disposition
Command Code output seeds, parked-pane seeds, pty-exit removal delete; main already emits the same facts
structured bridge status writes delete; main now publishes the row
launch placeholder seeds (a user launched an agent with a prompt) keep for now; main holds the launch config and can seed later
dismissal, acknowledgement, unmount keep; user facts and component lifecycle
remote-runtime OSC parse (bytes never transit local main) keep, fenced behind the host's published row once the host is new enough; rule 3 of the wire doc applies
web-session mirror receipt clock keep; the decay rule needs both clocks from one machine

The Command Code done-settle window is renderer policy with no main equivalent. PR 2 either moves it into main's detector or leaves it, and says which.

PR 3: one rollup, one clock

The worktree card status is derived three times: lib/worktree-status.ts in the renderer, runtime-worktree-status-projection.ts in main, and agent-row-display.ts in mobile, which hand-copies the 30-minute constant. PR 3 moves the rollup and the decay into src/shared and makes all three call it.

Readiness reads the store

terminal wait --for tui-idle is a reader too. Before STA-9100 hook state reached it only through the <Agent> ready titles the window writes, so a headless orca serve never saw it (#16095). Now an agent whose rule file says profile.hooks: "authoritative" (OpenCode, OpenCode 2, Pi, OMP) or "turn-end" (Codex) has its fresh row read straight from the store, through the same selectFreshExplicitAgentStatusRow join prompt-receipt verification uses (src/main/runtime/tui-idle-hook-lane.ts):

  • the main agent's turn, not the combined row, decides: mainAgent.state when published, so a subagent's Stop does not end the lead turn. done settles the wait, working holds it, and a permission wait never settles. The tail's blocked text goes through the existing permission arbiter with the turn as its explicit status, so a denied prompt's dialog left in the tail no longer blocks a turn the hook says ended;
  • the row joins on any pane key or terminal handle the PTY owns; a pane neither reaches, a stale or restored row, a session-start done, a row from before the PTY respawned, and a done received before the pane's latest input all leave the decision to the screen and text rules, which is also how startup readiness works before an agent's first hook. The input is the PTY run's lastInputAt (terminal-run-facts.ts), which both write funnels record, so a key the user typed counts like a prompt Orca sent: the next turn's first hook may still be in flight, and an agent restarted in the same shell has not posted one. A shell command marker is no process boundary: Pi paints OSC 133 zones itself;
  • every other agent stays identity-only: Claude sends no event when an approval is denied or Esc stops a tool, so its row can sit at waiting or working forever, and the rules keep deciding;
  • Codex is turn-end: only a done decides (it settles the wait), and a working or permission row leaves the decision to the rules. Before its Interrupt hook an Esc mid-turn can leave the row working, and an older TUI can hand its hooks to a newer shared app server, so no version check tells which Codex posts it. A done is a real turn end on every version, so trusting only that one keeps the headless gain (no quiet window after the turn) without letting a missing cancel hang the wait.

The titles stay for display; remote clients read them.

What does not change

  • The hook scripts, the OSC 9999 wire format, and the relay protocol.
  • The status vocabulary. working / blocked / done for rows, working / attention / idle for structured summaries, mapped once.
  • The live / unverifiable / exited verdicts for remote work. Loss of contact clears nothing; the SSH exemptions in the admission gate stay.
  • Hydration honesty: a restored non-done row is restoredUnconfirmed and is never fresh.

PR 1b reliability contract

  • Invariant (agent-session.status-host-ownership): each execution host has one agent-status store; OSC, hooks, and structured sessions write it, while desktop, worktree ps, and mobile only project it. Dismissal, certified PTY exit, and provider-generation replacement remove the same row everywhere; transport loss alone removes nothing.
  • Failure source: the deleted runtime row store duplicated OSC observations, keyed them by a different terminal identity, and outlived a dismissal from the hook store. Relay replay could also make old evidence look fresh when readers used its new delivery timestamp.
  • Oracle: one OSC observation appears through the hook snapshot in worktree ps and mobile, and one store dismissal removes it from both without stopping the PTY. Focused tests also require leaf/incarnation-handle rejoin, legacy numeric-pane compatibility, certified-exit and provider-generation cleanup, evidence-age freshness, and exactly-once startup/stop teardown.
  • Gate: terminal-performance.osc-status-scan-budget covers the unchanged bounded OSC parser and the runtime projection. There is not yet a dedicated blocking multi-surface status-store gate; the focused suites below are the accepted gap until they accumulate reliability-gate soak evidence.
  • Provider/platform coverage: local and daemon-backed PTYs are covered by runtime tests, and SSH relay loss/replay semantics by relay integration tests. The projection is shared by git worktrees and folder workspaces. WSL uses the same store and admission code but has no live run here; Linux and Windows runtime execution, native mobile clients, and mixed-version paired clients remain validation gaps.
  • Performance budget: publication stays event-driven with no new polling or subprocesses. One mobile projection clones the status snapshot once, builds pane/handle indexes once, and has a deterministic call-count test; lifecycle cleanup is bounded by the existing status and handle inventories, and orcad tests prove listeners clean up once on failed startup and repeated stop.
  • Diagnostics: existing hook-listener errors name the pane and PTY, while status-store tests pin delivery versus evidence clocks. No new telemetry or raw terminal data is emitted.
  • Residual gaps: rendered Electron/mobile behavior, live SSH reconnect, and Linux/Windows/WSL execution require the platform QA pass. The current cross-version gate does not cover session.tabs content.

Verification

  • Unit: ingest a structured summary and read it back through getStatusSnapshot, worktree ps, and the mobile projection; assert the serializer never writes a row carrying structuredHost; assert a hydrated file that somehow contains one is dropped.
  • Unit: the worktree ps suites written against the retained store are rewired to a real AgentHookServer (agent-status-store-wiring.test-fixture.ts) rather than deleted, so each still asserts the listing behavior it named. The dismissal change is pinned end to end in orca-runtime-tests/worktree-ps-agent-row-dismissal.spec.ts, which fails with the retained store restored.
  • Live: the parity check from #19217 (working, done, close, reload) repeated against the merged store, with both surfaces read from the one row.

Retired OMP pane recovery

A desktop renderer retirement carries an optional UUID through the existing agentStatus:retirePaneAuthority IPC message. The hook server retains it with its bounded retirement fence. A validated live OMP new turn consumes that UUID and echoes authorityRestartId only in the live notification. Cached rows, persistence and startup replay never carry the acknowledgement. Older peers omit or ignore it and retain explicit attach restoration.

The renderer keeps the UUID in its existing non-persisted retirement tombstone; every re-retirement mints a new one. A matching acknowledgement may clear that tombstone only with a successful status write for the existing pane and matching workspace/connection. Closed tombstones remain true, including after the tab LRU evicts its entry. Closing a retired physical alias revokes its whole group. This is control-plane retirement correlation, not a second agent-status store.

Fallback restores the hook server's recorded status aliases through the existing attach-restoration path. The accepted renderer write restores the matching status alias routes too, preserving group membership for the next retirement. It does not restore orchestration or launch credentials. It is scoped to the requesting desktop renderer. A different window's retirement UUID cannot be cleared by the acknowledgement, and web mirrors keep their existing host-snapshot/attach behavior.