* feat(native-chat): publish the host's child records to the status summary and the chat strip
The status summary and the background-task channel now read a session's child
records from the host's canonical store, through the status sink its row
landed in, and derive the legacy task and subagent shapes from the same views.
The parent row folds its child-work liveness from those records at ingest,
not from the summary's task list. The adapters no longer push their task DTO
to clients: the onBackgroundTasksChanged path is gone, and a child-work ingest
is what republishes both the summary and the strip. Finished children stay
listed until the session's own next turn starts. A reader that predates child
views never receives a roster whose rows are all settled.
* feat(native-chat): the chat strip reads the host's child records with its parent's verdict
The strip's roster now renders from the child views its channel carries, and
passes the verdict the session's own status row gives its children, built the
way the sidebar builds it (the row's freshness and the status feed's
observation). So one child reads the same in the strip and the sidebar, live,
after the transport drops, and once the row goes stale. A roster of finished
children stays shown until the next turn but no longer animates the monitoring
indicator or blocks conversation commands.
Tests: an end-to-end run on a host with no renderer (a real hook server as the
status sink) shows the summary and the strip channel carrying the same records
at every step, the parent row folded from them, retention, and an older
client's task list holding live work only; a wired renderer test shows both
surfaces agree when live, lost and stale.
* test(native-chat): a newer host's view degrades, an older reader keeps its live roster, a finished roster holds nothing open
- The view decoder ignores unknown keys, degrades unknown kinds, states,
outcomes and memberships, and drops only rows it cannot identify.
- At the RPC boundary a reader that predates child views gets no strip for a
roster of finished children and never the views themselves; a stop-only
reader keeps rows whose host offers no targeted stop.
- The strip shows finished children without reading them as live work.
- The row keeps its child list's identity when a summary repeats it.
* refactor(native-chat): the summary's task list is the legacy projection's live rows, unfiltered
* test(native-chat): type the switch tests' mocks instead of asserting them
* test: remote clients advertise reading child views
* docs(agent-status): the structured row folds the store's child records
* fix(agent-status): keep the view reader's header from reading as a value import to the renderer boundary
The renderer node-builtin boundary test scans raw text, so a header comment
that said "imports" ahead of the import block turned the type-only import of
agent-status-child-work into a value edge that reaches node:crypto.
* refactor(native-chat): the status summary's broadcast equality gets its own module
The status feed crossed the file-size limit once the summary gained the main agent's turn
outcome beside the child views. Which summary changes reach every session list now lives in
structured-agent-session-status-summary-equality.ts.
* fix(native-chat): command admission reads the strip's child records
A conversation command was refused on the provider tracker's own roster
while the strip read the host's child records, so a drift between the two
rule sets could refuse /clear with a stop instruction the strip had no
button for. Admission now reads the same records through the same read as
the strip, uses the strip's liveness fold, and asks for a stop only when
the strip renders one. The adapter contract no longer exposes the tracker
roster, so no host decision can read it.
Also records when the legacy child shapes die, every earlier death of a
settled child, and the display-precision invariant behind the summary's
clock tolerance.
* refactor(native-chat): command admission takes only what it reads of a turn
* fix(native-chat): the session list drops a session's children when the store does
A session's end no longer removes its child records: a child still running
settles with an outcome nobody reported, and a finished one stays listed.
Records now leave only at the session's own next turn, at the cap on settled
records, or when the host lets go of the session and its row leaves the store.
The summary kept after the host lets go used to strip its children on close,
a rule of its own. It now re-reads them from the store when the row leaves,
through the same read every live summary uses, so the session list and the
chat strip list the same children at each step, including a forget with no
close. Closing only revokes ownership, as before the child records existed.
* test(native-chat): write the Codex frame script's parent row out step by step
Once every surface reads the child records, the provider tracker's roster is
no oracle: it and the records read the same child executions, so agreeing
with it cannot catch a defect in either. Each frame now states the child
liveness and the parent row it must fold to.
* fix(native-chat): the idle sweep and the restart snapshot read the host's child records
The idle sweep (keep an agent running while its subagents or commands run) and
the restart-resume snapshot (what a chat was doing when Orca stopped it) both
read the provider tracker's roster through the adapter interface, which no
longer carries it. Both now take the host's one child-record read, the same one
the status summary, the chat strip and command admission use.
The snapshot's working test also folded that roster through the shared fold's
old `backgroundTasks` input, which the fold no longer reads, so a settled lead
whose subagent was still running would have been offered nothing. It now hands
the fold the records.
* test(native-chat): the child-record tests follow the merged command lifecycle
A command is a live child record from its start and is removed, not settled,
when it stops, whatever Codex tagged it. The end-to-end switch now shows the
child's `npm test` as a live row beside its dev server, and both are gone once
they exit; only the finished subagent stays listed until the next turn. Letting
go of the session is its tab closing, since a closed conversation whose tab
remains keeps its row.
Command admission's finished row is a subagent, the one kind that settles, and
the failed-verdict row test admits its live subagent as a host record, the only
thing the row folds.
* refactor(native-chat): the status feed's journal projection cache gets its own module
The status feed crossed the file-size limit once the child records joined the
agent-start signal and the completion feed's status read. The per-journal
projection, cached per commit, now lives in
structured-agent-session-status-journal-projection.ts.
* test(native-chat): the admission test's compaction resolves with a real outcome
Main's compaction result is a tagged outcome; the host-level admission test
resolved its mock compaction with an empty object.
* feat(native-chat): the sidebar lists running subagents; the strip, running then the newest finished
The host keeps every child record; what each surface lists is picked from them on every read, so
nothing is stored twice. The status summary, which every session list reads, now carries only
running children (and a finished one whose shell still runs, which reads monitoring): a finished
or failed subagent leaves the sidebar and stays in the chat's strip. The strip lists every running
child, then the newest finished ones, 100 rows in all; more than 100 running all show.
This matches common practice: sidebars show live subagents, and finished ones stay in the chat's
panel, newest first. No wire field is added. An older client reads fewer rows: its legacy task
lists were already live-only in the summary, and the strip's settled tasks come from the same
bounded roster.
* fix(sidebar): one rule for what the worktree sidebar lists: running children, from every source
`worktreeSidebarListsChild` is the one definition: a child that runs, counting a finished one
whose own shell still runs (it reads monitoring). The sidebar's row builder applies it to every
child source it reads, a terminal agent's hook roster and a chat session's records alike, and the
host's status summary applies the same predicate, so the sidebar's payload stays small. The chat's
strip keeps finished children, newest first.
A terminal agent's hook roster already drops a child on its own stop, so nothing changes there:
a teammate between turns and a child gone quiet still run, and still show. The selection module
moves to `agent-child-work-listing.ts`, since it now covers every source, not only chat sessions.
* test(native-chat): the switch test passes the startup child key main's status bar takes
* fix(native-chat): a finished child stays until the user's next send, not a turn Claude opens on its own
Claude wakes the agent on its own when a background task ends, and that wake is a
new root turn. Keying retention on the newest root turn retired every finished
child about two seconds after a background agent or shell finished, so its
outcome never showed in the strip.
Retention now keys on the user's newest send the provider accepted (a message, a
steer or a command), with the journal epoch so a rewind still retires. A turn the
provider opens itself and a subagent's turn carry no send. Replayed captured wake
orders through the real adapter, hook server and status feed.
* fix(native-chat): a background Stop reaches the tasks the child records show
The strip draws a row's Stop, and /clear, /compact and rewind wait for background
work, from the host's child records, but the Claude adapter still resolved which
tasks a Stop reached from its own tracker's roster, and refused to stop at all
once that roster was empty. A task the records kept live after the roster dropped
it showed a Stop that sent nothing and blocked those commands until the chat tab
closed.
The host now resolves the provider ids a Stop sends from the records (the same
per-row rule the strip and admission use; every such row for stop-all), and the
adapter stops exactly those, with no tracker guard. An acknowledged stop ends the
record: a running task sends its own stopped frame first, and the CLI answers
success with no frame for a task it no longer knows. A refused stop leaves the
record live. No production code reads the tracker's roster any more.
* fix(native-chat): one rule for a finished child that still owns live work, at any depth
The listing kept a finished child whose work ran through any depth of ownership,
but retention at the user's next turn protected only the direct owner, so a
finished agent whose finished subagent still ran a shell was removed and that
subagent jumped to the top level. Both now read settledOwnersOfLiveWork.
* fix(native-chat): an older client sees a Codex child's shell as it did before views
Clients that predate child views read a flat task roster derived from the views.
It listed a Codex child agent's shell as an extra row beside the running agent,
then as a bare command once the agent finished. The derivation now hides a
running agent's commands and names a finished agent's as "<agent> — <command>",
as the Codex tracker did; the label rule moves to a shared module both use.
* fix(native-chat): the chat decodes a roster's child rows once, as the frame arrives
The client reducer compared raw wire rows, so a row shaped by a newer host could
throw there, and the strip decoded a new array on every render, which defeated its
grouped-rows memo while a turn streamed. Rows are now decoded where the frame
enters the reducer, an unchanged roster keeps its identity, and the strip's parent
context is rebuilt only when one of its values changes.
* fix(native-chat): the strip channel forgets a closed conversation's roster
It kept the last roster fingerprint of every conversation for the host's lifetime.
The conversation map now tells observers when one leaves it. Also corrects the
summary's children comment: it carries running children only.
* docs(native-chat): rewrap the retention comment
* fix(native-chat): a task's own ending replaces a Stop's, and a child finished after the user wrote stays
Two lifecycle gaps from the round-1 fixes.
A Stop acknowledged ahead of the task's own ending relabelled it. The SDK hands
Orca a control answer as soon as it reads it and queues other frames, so when a
task finished just as the user pressed Stop, the acknowledgement arrived before
the task's completion the CLI wrote first, and the task read "Stopped" with its
result lost. An acknowledged Stop now ends a record provisionally (outcome basis
`stop-acknowledged`); the task's own terminal frame replaces it, and nothing
replaces an ending the task reported itself.
A finished child still vanished with no new action from the user when the send
the provider took was written before the child finished: a steer Claude takes at
its next boundary, or a queued draft handed over at the end of the turn. The
user's next turn now carries when they acted (the send's written time, or the
draft's queued time, both on the host clock), and only children that finished at
or before it retire; a later one stays until the user's next send. A rewind
still retires every finished child.
Also: a Stop-all keeps stopping the remaining tasks after one request fails, then
reports the failure.
* fix(native-chat): the strip keeps one empty list for a roster that omits one
A roster with only running or only finished rows made a new empty array on every
render, so the strip regrouped its rows each time. One shared empty list keeps
its memo.
* fix(native-chat): a strip row whose owner the 100-row budget cut renders under the main agent
The budget can keep a finished child and cut its finished owner. The child still
named that owner, so it rendered nowhere. The selection now clears an owner it
did not keep, as the view contract says for an owner outside the projection.
* fix(native-chat): a stop-all that times out stops asking, and "no children" is sent once
A stop-all kept asking after a request timed out, so a Claude CLI that stopped
answering control requests cost one full deadline per task, while the chat's
sends, Stop and /compact waited behind it. A timeout now ends the loop; other
request failures still let the remaining tasks be stopped.
The chat strip channel never remembered that it had sent "no children", so
every change in a chat with none re-sent that frame to each subscriber. It now
remembers it, and forgets only when the conversation closes.
Also renames agent-child-work-stop.ts to agent-child-work-stop-targets.ts, which
says what it answers: the provider ids a background Stop reaches.
* fix(native-chat): the status feed reads a journal snapshot with no submissions, and e2e tests use the child-work reader
CI on dbd2439cd4 was red in three places:
- first-work-branch-rename and the agentSession.subscribeStatus RPC test feed the status feed a
journal whose snapshot lists items only. The projection reads the user's newest accepted send
from `snapshot.submissions`; it now tolerates their absence, as the status projection beside
it already did.
- the cross-version downgrade test still passed `backgroundTasks` to the teardown's working
marker, which now takes `childWork`.
- an e2e unit test still gave the status feed the removed `readBackgroundTasks` dependency
(harmless at run time, a type error in the tests/ project).
* fix(native-chat): the chat strip lists running children only, by the sidebar's rule, and hides when none runs
A finished subagent's result is already in the transcript ("Ran N subagents ·
completed"), so the strip is for work that runs. It now lists exactly what the
sidebar lists, by one predicate (a running child, or a finished one whose own
shell still runs, which reads monitoring), and the host sends no roster once none
runs, so the strip hides.
Gone with it: the 100-row budget and the running-then-newest-finished
selection, the re-homing of a child whose owner the budget cut, and the RPC
gate's rule for a roster of finished rows only (no such roster exists now).
Older clients still get their derived task list, running work only.
Finished records still stay in the host's store until the user's next accepted
message: they refuse a late frame of their run, let a task's own ending replace
an acknowledged Stop's, and keep a running shell's owner. Dating that retention
by when the user wrote the message only kept finished rows visible longer, so it
is removed.
* fix(native-chat): the strip shows running work only from any host, and hides after a released session's last child
- A new app paired with an older host no longer shows that host's finished task
rows: the strip lists running work only, whatever host sent it, and hides when
an older host's roster has only finished rows left.
- A test for the path that hides the strip when a session's last running child
settles after the provider let go of the session (Claude's release path): the
channel sends `null` though no provider answers for the session any more.
- A test comment still described the strip keeping finished children.
35 KiB
Agent status store
Status
The current boundary is PR 2A: structured sessions use the hook server's fully scoped canonical store; unbound PTY/relay evidence remains in an isolated legacy adapter. Do not remove the renderer bridge or its publication filters in this slice: they still carry native-chat child rows.
The sections below record the original 2026-09-09 rollout. Its PR 1a and PR 1b have landed; its proposed PR 2/3 sequence is superseded by that boundary:
- main-only: every producer writes into one store and
worktree psreads it, split into 1a (structured sessions join the store) and 1b (the runtime's duplicate retained store is deleted); - renderer: the sidebar becomes a subscriber and stops re-deriving rows;
- shared: one worktree-status rollup and one freshness rule for every reader.
The problem this solves
Orca shows "what is this agent doing" in four places: the desktop sidebar, the
orca worktree ps command, the mobile app, and the agent dashboard. Before
#19217 those readers did not even share their inputs. After #19217 they share
the structured-session mapping and nothing else.
An audit on 2026-09-09 found six producers and three consumers, and three separate copies of the same row inside the main process alone:
| Main-process copy | Keyed by | Owned by | Persisted | Evicted |
|---|---|---|---|---|
hook server lastStatusByPaneKey |
paneKey | src/main/agent-hooks/server.ts |
last-status.json |
tab close, pty exit, hydrate |
runtime RuntimeAgentRowStore |
paneKey | runtime-agent-row-store.ts (deleted in PR 1b) |
no | pty exit only |
structured feed published |
sessionId | src/main/native-chat/agent-session-wire/structured-agent-session-status-feed.ts |
no | never (a broadcast cache) |
The second copy is a duplicate write: the OSC status parsed in main is
forwarded to the hook server and retained in the runtime store from the same
call (orca-runtime-create-terminal-side-effect-command-code-detector.ts).
The third copy is keyed differently and never reaches the hook server at all,
which is why worktree ps grew its own adapter for it in #19217.
Each reader then applies its own precedence and freshness rules, so the same pane can legitimately read differently on the desktop, on the phone, and in the CLI.
The rule
The execution host owns agent status, in one store, and every reader
subscribes to it. This follows the boundary in
ssh-execution-boundary.md: the host that runs
the process is the only party that can observe it, and the client is never
authoritative for execution state.
Three consequences:
- One store per execution host. A remote host keeps its own store and the client mirrors it down, as the web-session mirror already does. Mirroring is not merging: a client never writes its observations back to a host.
- Precedence is decided once, at write time, with provenance recorded on the row. Readers never re-adjudicate hook versus terminal versus structured.
- Readers keep only presentation policy and user facts: the 30-minute display decay, acknowledgements, dismissals, unread. Those stay reader-side but become one shared implementation (PR 3).
The store already exists
The hook server's state is that store today for every PTY-based agent. The audit established:
- hook HTTP posts, the WSL and SSH relay receivers, and main's own OSC parse
all converge on the same
applyNormalizedStatuspath, stamped with the authority idmain-agent-hooks; - it alone holds pane authority: launch tokens and their hashed commitments, retired-pane fences, pane-key aliases, per-connection ordering watermarks, and the evidence-age map that must outlive a transport clear;
- it alone persists, with a seven-day hydrate window and the
restoredUnconfirmedstamp that keeps a hydrated row from ever reading as live truth; - it already fans out to both renderer windows over
agentStatus:setandagentStatus:clear, and servesagentStatus:getSnapshot.
Nothing else in main carries those guarantees, and building a second store with them would be the wrong direction. So the design is not "add a store". It is: route the two producers that bypass the hook server through it, then delete the copies.
PR 1a: structured sessions publish into the store
No renderer behavior changes. The sidebar keeps receiving the same IPC events it receives today, plus structured-session rows it currently derives itself.
Structured sessions publish into the hook server
The structured feed keeps its job of projecting a session's journal into a summary and streaming it to subscribers. On every publish it additionally ingests the summary into the hook server as a status row:
| Row field | From |
|---|---|
paneKey |
structuredAgentSessionPaneKey(tabId, sessionId), the key the renderer already uses; its leaf is UUID-shaped so pane-key validation accepts it |
tabId |
structuredAgentSessionTabId(sessionId) |
worktreeId |
summary.workspaceId (a folder workspace id is a valid value) |
state |
structuredAgentSessionAgentStatus(summary).state: the lead's own status folded with the live child records the store holds for the session (not the summary's task list), so a settled lead whose subagent still runs reads working |
workingMode |
'monitoring' from the same fold when watch loops are the only live child work; omitted otherwise, which clears it on the row |
mainAgent |
the main agent's own state before the fold, its last-turn verdict (summary.turnOutcome, present only while idle) and its own clock; see "The main agent fact" below |
structuredHost |
'owned' while summary.hostExecutionOwned is set, otherwise 'held'; worktree ps derives its row's structuredHostOwned from it |
| prompt, tool, last message, model, provider session | the summary's fields |
Sessions with no request (status === null) produce no row. A request is a
turn record, an assistant message, a user message the provider journaled itself
(history, an older host), an accepted or unanswered send, or a send the agent or
its start refused; a send that was withdrawn, or left undelivered by a
restart or a close, fails nobody and makes nothing listable.
summary.turnOutcome is the latest request's verdict: its turn's outcome, or
failure for a send the agent or its start refused (a send that joined a running
turn is answered by that turn). A turn the provider gave no outcome reads as its
host-observed end through agentTurnVerdict: interruption for an interrupted
lifecycle, unconfirmed for an unverifiable one. It is derived on each read,
never journaled. The row also publishes interrupted from
mainAgent.outcome, exactly as the hook lanes do. When the host revokes live ownership the row is re-set
without the flag; when the host closes or evicts the session the row is
dropped. Both already exist as feed events (revokeLive and the roster
filter in liveSessionSummaries); PR 1 turns them into store writes.
Dropping the session from the host's map and dropping its row are one
operation, forgetStructuredAgentSession. The store keeps a row until told,
and a host-owned row bypasses the staleness check, so a deletion path that
forgot the row would strand a permanently working-looking agent.
Two rules the ingest must keep:
- Never persist a structured row. The journal is the durable truth for a
structured session and the host republishes on restore. A structured row in
last-status.jsonwould hydrate asrestoredUnconfirmedand then fight the live republish. The serializer skips rows carryingstructuredHost, and hydrate drops any such row found on disk. Applying one therefore also skips the persist schedule: the walk and stringify could only reproduce the file that is already on disk, once per debounce window for every streaming chat. - Never let it fight a hook row. A structured session has no PTY, so no hook or OSC event carries its pane key. The ingest still goes through the disposition gate so a retired pane key is refused like any other.
Applying one does still run both status fan-outs, and that is intended rather
than incidental. notifyStatusChangeListeners is what feeds
agentAwakeService's power-save blocker, and subscribeEnrichedStatus is what
feeds AgentSessionTransitionRecorder's stats, so joining the store enrolls
native chats in both. A working native chat is real work and should hold the
machine awake exactly like a PTY agent does.
The drop side routes through dropStatusEntry, not clearPaneState: a
pane-status-clear reaches the renderer, and until PR 2 the renderer's own feed
bridge is that pane key's writer. It also passes preserveResumeIdentity: false — the providerSessionOnly remnant a dismissed pane keeps exists so the
agent can be resumed in that pane, and a structured session has no pane and
keeps its resume identity in the record store. Like every other
dropStatusEntry caller, it emits no pane clear, so a session dropped
mid-working leaves AgentSessionTransitionRecorder holding an open stats
session until its LRU evicts it; that gap is shared with the user-dismissal
path and is not specific to structured rows.
The ingest lives in the feed, not in structured-agent-session-host.ts, which
sits at the file-length cap.
worktree ps becomes a reader
The structured adapter added in #19217 is deleted, and structured rows reach
worktree ps through the same snapshot as every other row. The
retained-versus-hook reconciliation in collectRuntimeWorktreePtyAgentSources
stayed until PR 1b removed the store that fed it. What this step settles is
the admission gate that decides which rows a worktree listing may show:
- a hook or OSC row needs its tab mirrored or a connected pty, as today, and SSH rows stay exempt because their tabs may exist only remotely;
- a row carrying
structuredHostis admitted while the host holds the session, and the host's drop on close is what removes it. No tab-mirror requirement: a structured session's tab lives in the renderer's own tab state, and a headless host has no renderer to mirror it from. That argument only holds if the headless host is itself wired to the store, which is a separate obligation per entry point: the Electron hosts (desktop andorca serve) sharemain-process-runtime-service.ts, andorcadconstructs its own runtime insrc/main/orcad/orcad-entry.ts. A host missing that wiring lists no agents at all, not just no structured ones, becauseworktree psreads the same snapshot for every row.
The freshness bypass for host-owned structured rows already exists in
isFreshNonDoneAgentStatus; with the flag now on the row it becomes the only
path, and the hand-rolled check in runtime-worktree-agent-rows.ts goes.
Wire compatibility
AgentStatusIpcPayload gains one optional field, structuredHost, and the
worktree ps row gains structuredHostOwned. Under rule 1 of
remote-wire-compatibility.md both are safe:
an old client ignores them. worktree ps rows keep their shape and vocabulary,
so the mobile app sees no change.
Until PR 2 the main process does not forward structured rows to the renderer
over agentStatus:set or agentStatus:getSnapshot. The renderer's feed
bridge still writes those rows itself, and forwarding them too would give one
pane key two writers. Removing that filter is the first step of PR 2.
The main agent fact
Claude, Codex and Grok hook rows and structured-session rows publish the combined
state and, beside it, the main agent's own state as payload.mainAgent. Other agents'
rows and terminal-title-only rows carry none, and readers fall back to state:
mainAgent?: { state: AgentStatusState; outcome?: AgentTurnOutcome; stateStartedAt: number }
state still answers "what should the user see" and folds live child work in,
so a settled main agent whose subagent still runs reads working. mainAgent answers
"what is the main agent itself doing", which the fold used to destroy at publish
time; every guard that reconstructed a fragment of it (fromChildWork, the
persisted claudeLeadBoundaryChildOnly flag) now reads mainAgent instead of a
stored copy. A Claude row whose mainAgent is done while a child agent still
works (including a child's permission wait) refuses OSC, which carries no child
identity; the children's own lifecycle hooks settle it. outcome is the recorded verdict on
the main agent's most recent finished turn, present only while mainAgent.state is
done. It is reported by the provider, or is a cancellation Orca inferred
from the user's own interrupt keystroke, or a superseded the host recorded when
a newer request replaced a structured Claude turn before it ended (it names no
sender, and sets no legacy flag), or, on a structured row whose turn the
provider gave no verdict, is what the host observed of its end: interruption
(a proven death nobody asked for) or unconfirmed (an end it cannot prove,
never success). The journal's turn outcome, by contrast, stores only recorded verdicts: the provider's, a cancellation, or the host's superseded; interruption and unconfirmed are derived from the turn's lifecycle state and never stored. A plain end of turn carries none, because absent
means unknown and a provider that omits its interrupt flag must not turn a
cancel into a success.
In the Claude hook lane the cancellation comes primarily from Orca's own
inferred interrupt (markClaudeLeadTurnInterrupted), because current Claude
sends no hook at all on a cancel and no is_interrupt on Stop; that flag on a
turn boundary remains a secondary source for builds that send it, and
StopFailure maps to failure.
Readers decode the verdict through one accessor, agentMainAgentVerdict, which
reads the main agent's own state, not the combined row's: mainAgent.outcome
while mainAgent.state is done, then the legacy interrupted flag as a
cancellation, which alone needs the combined done. So a main agent that
failed while its subagents still run has a verdict on a working row. Every
copy of a row (state-history entries, sleep records, worktree ps rows) takes
the verdict through agentVerdictFields, which carries interrupted and the
whole mainAgent (state, outcome and its own clock) together, so a copy agrees
with the row and can date a failure by mainAgent.stateStartedAt.
Display reads the verdict through agentVerdictDisplayMark. A fault marks the
agent failed whatever the combined state, because it is news the user must see
even while subagents run: a failure, and an interruption, a turn cut short
by anything other than the user or a newer request. A user's stop (cancellation)
marks it interrupted, drawn in the muted tone with the row text "Interrupted by user";
a turn a newer request replaced (superseded) marks it interrupted in the same muted
tone with the row text "Interrupted"; and unconfirmed marks it unconfirmed, all only
on a done row, so a stopped
or finished main agent with live child work still reads working. The folded
turn header follows the same mark: "Failed after N", "Interrupted after N", or
"Worked for N".
Each subagent keeps its own row and state. Container rollups (worktree card,
terminal tab, Cmd+J) rank a pending question first, then a failure, then live
work, then an unconfirmed end, then a user's stop, then done. On the worktree
card, a failure retained after its agent's pane went away has no expiry, so it
ranks below live work and above an unconfirmed end. Lifecycle waiters keep
reading the combined state.
Policy splits the verdict two ways. Clean-finish policy (hibernation, pane
ownership, the star-nag value moment) treats a failure, an interruption and an
unconfirmed end like a cancellation (agentTurnEndedUncleanly). Attention
(completion time, Smart Sort, sticky retention, Cmd+J Recent) demotes only a
turn ended on purpose, the user's stop or a newer request that replaced it
(agentTurnEndedOnPurpose); a failure, an interruption or
an unconfirmed end ranks like a completion.
Admission is one function, normalizeAgentStatusPayload, on the relay wire,
IPC and disk. A malformed mainAgent drops the field and keeps the row. Old hosts
send none and readers fall back to state. Hook rows persist it inside the
payload; hydration maps an older row's claudeLeadBoundaryChildOnly: true
onto mainAgent: { state: 'done' } when the row has no mainAgent, and never writes the
flag again. Hydration seeds the Claude main agent record straight from a saved
mainAgent that is done, so the children's drain can still settle the row after
a restart. claudeRunningNonAgentTask is persisted alongside because it is the one
child-work fact mainAgent cannot express: a shell running beside the main agent,
whose liveness hydration does not restore. Hydration seeds only a row that says
false; a row silent about it stays unseeded. The row builder pairs the two facts in
one place: a listener event restates the shell fact, and any other write (an OSC
repaint, an inferred answer) keeps it only while mainAgent is unchanged. A child's
sticky permission prompt still records the main agent's own progress and background
evidence in the held row, and pushes the held row to subscribers when mainAgent changes.
Every lane, Codex included, combines through the fold. A child waiting on a
human is a fold input (childWorkLiveness: 'waiting', derived from the child's
own waiting state; a child's blocked means it failed and stays live work)
and makes the row wait whatever the main agent is doing, unless the main agent
is itself asking. Only the Codex hook lane feeds that input today. Known
divergences, pinned by name in the parity table
(src/shared/main-agent-status-parity.test.ts) where they are reachable, so a
reader does not mistake them for drift:
- The Claude hook lane holds a child's permission wait in one slot on the
displaced main agent record (
waitingAgentId,stateBeforeWait), not on the child. It publishes the displaced state asmainAgent, but the next main agent event overwrites the slot, so the row stops readingwaitingwhile the child is still asking, and a second asking child replaces the first. - The structured lane has no per-child wait: a child's pending prompt makes
the session
attention, which reads as the main agent's ownblocked. - The Codex hook lane drops its roster on a root
Stopwhen it tracks no child transcripts, so a still-running or still-asking child stops holding the row.
How the main agent's turn ended is not a fold input. A cancel is a verdict on
the main agent, carried as mainAgent.outcome: 'cancellation' (and, for
readers that predate mainAgent, as the row's interrupted flag on a done
row); it never retires a shell, scheduled check or subagent the turn left
running. That work leaves the row only when its own inventory omits it or the
session ends, so a cancelled turn with a still-running shell reads
monitoring in every lane, and the parity table in
src/shared/main-agent-status-parity.test.ts drives that story through all of
them. The same rule governs the cancel Orca infers from Ctrl+C: for any row
that publishes mainAgent, the inference is admitted only when
mainAgent.state is working, so Orca does not treat a Ctrl+C at the idle
prompt of a row held open by child work as a turn cancel (Codex also keeps the
child-evidence guard, and a row without mainAgent keeps only that guard).
The keypress itself is not inert, though: measured live, Claude 2.1.280 stops
its background subagents on a single idle-prompt Ctrl+C (shells survive) and
Codex 0.156.1 quits outright, so refusing the inference can leave the row
showing a subagent its CLI already stopped. The synthesized row is the fold
of the cancelled main agent with the child work the pane's owner can see: the
local listener's roster for a local pane, the row's own subagents and shell fact
for a relayed one, whose provider records live on the relay.
The store holds that verdict against restatements that predate it
(server-cancel-verdict-latch.ts), because a relay never learns of a cancel
the desktop infers and some TUIs emit late same-turn hooks. The hold is read
off the row (mainAgent.outcome: 'cancellation'), never stored beside it, and
dies on a new turn (a main agent prompt submission, a changed or explicit
prompt, a session start) or the provider's own settled mainAgent. Child and
replayed events under the hold keep the cancelled main agent and are re-folded
with their own child evidence.
PR 1b: the runtime's retained row store is deleted
Landed. RuntimeAgentRowStore is gone, and with it the retained-versus-hook
reconciliation in collectRuntimeWorktreePtyAgentSources. The hook server's
store is now the only main-process copy of a PTY agent's row.
The five call sites
| Call site | Before | After |
|---|---|---|
orca-runtime-create-terminal-side-effect-command-code-detector.ts retain() |
second write of the OSC payload already sent to the hook server | deleted; the event now carries the pane's terminalHandle and the hook ingest keeps the only copy |
...command-code-detector.ts clearPty() |
drops rows on pty exit | deleted; pane teardown already clears the hook row |
orca-runtime-get-worktree-ps.ts values() |
fed retainedSnapshots |
deleted; the reader keeps only hookSnapshots |
orca-runtime-serialize-agent-prompt-submission.ts getFreshExplicit() |
retained row first, hook rows second | selectFreshExplicitAgentStatus, hook rows only |
orca-runtime-prune-mobile-session-tab-group-layout.ts getFreshForMobile() |
pane key, then pty id | selectFreshAgentRowForMobileTab: pane key, then terminalHandle |
Both readers moved into runtime-hook-agent-row-selection.ts, which also owns
RuntimeAgentRowSnapshot now that nothing retains one.
terminalHandle is the row's join back to its terminal
The retained store's only real extra was the pty id, and two readers used it.
The plan said to stamp the event's ptyId into terminalHandle; that was
wrong. A terminal handle (term_<uuid>) and a pty id are different
identifiers, and getFreshExplicit was already comparing hook rows against a
real handle. What landed instead:
AgentHookEventPayloadand the runtime's terminal-status event gained an optionalterminalHandle. The detector resolves it once per chunk throughgetAgentStatusTerminalHandleForPaneKey— the same lookup the renderer-facing IPC boundary already runs for every row, so the two surfaces cannot disagree about which terminal a pane is.applyNormalizedStatuscarries the handle forward when an incoming event resolves none. Only main's OSC parse can resolve one, so an HTTP hook post for the same pane would otherwise erase it.- It is never persisted. A handle belongs to the runtime that issued it, and a hydrated one could only rejoin a row to somebody else's terminal.
toAgentStatusIpcPayloadpublishes it, which also makesgetFreshExplicit's long-dead handle comparison live: the runtime reads raw snapshot rows, and before this nothing ever stamped the field on them.
worktree ps uses it too. ConnectedPtyEvidence traded its flat ptyIds set
for ptyIdByTerminalHandle, so a row still resolves the connected PTY behind
it — which is both the working-terminal rollup's match key and the last rescue
for a row whose pane binding was nulled by a controller incarnation change.
The change detector had to move with the store
retain() was not only a store: its boolean return was the signal that
republished session.tabs for a status-only transition, which no title change
covers (#7970). hook-status-session-tabs-invalidation.ts already mirrors that
projection change set, including restore provenance and terminal-handle joins,
so the replacement was to route the signal off the store rather than build a
second comparator.
installHookStatusSessionTabsRepublish now owns all three arms — enriched
status, pane clear, and the status-drop tap a dismissal emits — and both hosts
install it.
Both hosts, not just the desktop one
orcad constructed its runtime with no onTerminalAgentStatus, so main's OSC
parse never reached the store there and the retained copy was the only carrier.
Deleting it without wiring orcad would have made a headless host list no PTY
agents at all. orcad-entry.ts now binds the producer and installs the
republish signal, alongside the snapshot and structured sink it already had.
The intended behavior change
A row the user dismisses on the desktop leaves worktree ps and the phone at
once, instead of lingering until the pty exits. One store means one dismissal.
Legacy numeric pane keys remain a bounded compatibility case. Persisted layouts register aliases to their stable leaf owners; an in-process OSC observation may also retain a numeric key only when the runtime supplies the matching tab, PTY, and terminal handle. HTTP and relay ingress still require a stable key or a registered alias, and numeric rows are never persisted.
PR 2: the renderer subscribes
With structured rows arriving over agentStatus:set, the renderer's
StructuredAgentSessionStatusBridge no longer needs to write status; its
unmount cleanup becomes a tab-close signal to the host. The IPC applicator is
the single writer for observed status. The 2026-09-09 audit sorted the other
writers:
| Writer | Disposition |
|---|---|
| Command Code output seeds, parked-pane seeds, pty-exit removal | delete; main already emits the same facts |
| structured bridge status writes | delete; main now publishes the row |
| launch placeholder seeds (a user launched an agent with a prompt) | keep for now; main holds the launch config and can seed later |
| dismissal, acknowledgement, unmount | keep; user facts and component lifecycle |
| remote-runtime OSC parse (bytes never transit local main) | keep, fenced behind the host's published row once the host is new enough; rule 3 of the wire doc applies |
| web-session mirror receipt clock | keep; the decay rule needs both clocks from one machine |
The Command Code done-settle window is renderer policy with no main equivalent. PR 2 either moves it into main's detector or leaves it, and says which.
PR 3: one rollup, one clock
The worktree card status is derived three times: lib/worktree-status.ts in
the renderer, runtime-worktree-status-projection.ts in main, and
agent-row-display.ts in mobile, which hand-copies the 30-minute constant.
PR 3 moves the rollup and the decay into src/shared and makes all three
call it.
What does not change
- The hook scripts, the OSC 9999 wire format, and the relay protocol.
- The status vocabulary.
working / blocked / donefor rows,working / attention / idlefor structured summaries, mapped once. - The
live / unverifiable / exitedverdicts for remote work. Loss of contact clears nothing; the SSH exemptions in the admission gate stay. - Hydration honesty: a restored non-done row is
restoredUnconfirmedand is never fresh.
PR 1b reliability contract
- Invariant (
agent-session.status-host-ownership): each execution host has one agent-status store; OSC, hooks, and structured sessions write it, while desktop,worktree ps, and mobile only project it. Dismissal, certified PTY exit, and provider-generation replacement remove the same row everywhere; transport loss alone removes nothing. - Failure source: the deleted runtime row store duplicated OSC observations, keyed them by a different terminal identity, and outlived a dismissal from the hook store. Relay replay could also make old evidence look fresh when readers used its new delivery timestamp.
- Oracle: one OSC observation appears through the hook snapshot in
worktree psand mobile, and one store dismissal removes it from both without stopping the PTY. Focused tests also require leaf/incarnation-handle rejoin, legacy numeric-pane compatibility, certified-exit and provider-generation cleanup, evidence-age freshness, and exactly-once startup/stop teardown. - Gate:
terminal-performance.osc-status-scan-budgetcovers the unchanged bounded OSC parser and the runtime projection. There is not yet a dedicated blocking multi-surface status-store gate; the focused suites below are the accepted gap until they accumulate reliability-gate soak evidence. - Provider/platform coverage: local and daemon-backed PTYs are covered by runtime tests, and SSH relay loss/replay semantics by relay integration tests. The projection is shared by git worktrees and folder workspaces. WSL uses the same store and admission code but has no live run here; Linux and Windows runtime execution, native mobile clients, and mixed-version paired clients remain validation gaps.
- Performance budget: publication stays event-driven with no new polling or subprocesses. One mobile projection clones the status snapshot once, builds pane/handle indexes once, and has a deterministic call-count test; lifecycle cleanup is bounded by the existing status and handle inventories, and orcad tests prove listeners clean up once on failed startup and repeated stop.
- Diagnostics: existing hook-listener errors name the pane and PTY, while status-store tests pin delivery versus evidence clocks. No new telemetry or raw terminal data is emitted.
- Residual gaps: rendered Electron/mobile behavior, live SSH reconnect, and
Linux/Windows/WSL execution require the platform QA pass. The current
cross-version gate does not cover
session.tabscontent.
Verification
- Unit: ingest a structured summary and read it back through
getStatusSnapshot,worktree ps, and the mobile projection; assert the serializer never writes a row carryingstructuredHost; assert a hydrated file that somehow contains one is dropped. - Unit: the
worktree pssuites written against the retained store are rewired to a realAgentHookServer(agent-status-store-wiring.test-fixture.ts) rather than deleted, so each still asserts the listing behavior it named. The dismissal change is pinned end to end inorca-runtime-tests/worktree-ps-agent-row-dismissal.spec.ts, which fails with the retained store restored. - Live: the parity check from #19217 (working, done, close, reload) repeated against the merged store, with both surfaces read from the one row.
Retired OMP pane recovery
A desktop renderer retirement carries an optional UUID through the existing
agentStatus:retirePaneAuthority IPC message. The hook server retains it with
its bounded retirement fence. A validated live OMP new turn consumes that UUID
and echoes authorityRestartId only in the live notification. Cached rows,
persistence and startup replay never carry the acknowledgement. Older peers
omit or ignore it and retain explicit attach restoration.
The renderer keeps the UUID in its existing non-persisted retirement tombstone;
every re-retirement mints a new one. A matching acknowledgement may clear that
tombstone only with a successful status write for the existing pane and matching
workspace/connection. Closed tombstones remain true, including after the tab
LRU evicts its entry. Closing a retired physical alias revokes its whole group.
This is control-plane retirement correlation, not a second agent-status store.
Fallback restores the hook server's recorded status aliases through the existing attach-restoration path. The accepted renderer write restores the matching status alias routes too, preserving group membership for the next retirement. It does not restore orchestration or launch credentials. It is scoped to the requesting desktop renderer. A different window's retirement UUID cannot be cleared by the acknowledgement, and web mirrors keep their existing host-snapshot/attach behavior.