Files
orca/docs/reference/agent-status-store.md
T
Jinwoo HongandClaude cdfdadf9ea fix(runtime): settle tui-idle on hook state for agents whose hooks cover the whole turn (#24388)
* fix(codex): install Codex's Interrupt hook so an Esc-cancelled turn settles

Codex 0.150+ fires an Interrupt hook when the user presses Esc on an
approval prompt or mid-tool, and nothing else. Orca did not install it, so
the pane stayed blocked/working until the next prompt.

- Add Interrupt to the managed Codex events and label maps, written with
  Codex's 3s cap (a larger value triggers a startup clamp warning).
- Hash the timeout Codex hashes (Interrupt is clamped to [1,3], default 1)
  so self-computed trust matches Codex; pinned against a real 0.159.3 hash.
- Map a root Interrupt to the existing cancelled-turn record
  (markCodexLeadTurnInterrupted), keeping child work in the fold; a
  child-scoped Interrupt is ignored. Relayed rows take the same path.

* test(runtime): add a readiness census pinning every tui-idle verdict

Replays every recorded agent PTY transcript frame by frame through a real
runtime pane (agent-known and agent-unknown, clocked and clockless) and a
synthetic evidence matrix for all 43 TuiAgents, and compares each verdict
and tui-idle wait outcome to committed run-length-encoded baselines.

Refs STA-9098

* test(runtime): pin the census quiet probes to literal windows

A census that read TUI_IDLE_QUIESCENCE_MS would move with it; fixed 2999/3000 ms
reads and a fixed 2000 ms poll step make a changed window show as changed verdicts.

Refs STA-9098

* test(runtime): say which census probe writes runtime state

Refs STA-9098

* refactor(codex): let the hook builder own Codex's per-event timeout

The managed hook's timeout is now Codex's own normalization of the shared
budget, and every installer derives its trust entry from the hook it wrote,
so no installer repeats the Interrupt special case.

Claude-Session: codex-interrupt-hook review

* refactor(codex): route Interrupt through the Stop lead update with an outcome

Interrupt now writes the lead record through the same setCodexMainAgentTurnState
call as Stop, so markCodexLeadTurnInterrupted keeps its original signature.
Drops the child-scoped Interrupt guard: Codex never runs Interrupt hooks for
subagents and its input schema has no agent_id.

Claude-Session: codex-interrupt-hook review

* test(runtime): observe the census through settled panes and caller-visible waits

- Read each verdict through the runtime's own settle seam (evaluateTuiIdleForLeaf) instead
  of re-wiring evaluateTuiIdle/leafTuiIdleEvidence/buildTerminalWaitText, so the census is
  coupled to one runtime method, not to the module STA-9098 rewrites.
- Let the runtime finish each chunk (one macrotask turn) before reading. The old read raced
  work chained on the paint, so 14 frames pinned a microtask-ordering artefact.
- Record when a wait settles (@start vs @poll), not just its outcome.
- Exit each pane's PTY after reading it so its emulator is freed.
- Replace the hand-grouped families, literal fixture list and per-pane split flag with a
  directory-scanned catalog, one baseline per replayed pane, and size-balanced shards.
- Run the synthetic matrix in one file; it takes about 2 s.

* test(runtime): cross dialog-versus-ready-screen order with every title in the census matrix

Blocked detection is position-ordered (design doc 11.5): the later of a blocker and a ready
anchor wins. The matrix now paints a workspace-trust dialog after, and before, each agent's
ready screen under every title, so a rule engine that loses that ordering fails per agent.

* test(runtime): read the census baseline field without Reflect.get

The anti-slop lint rejects Reflect.get on parsed input.

* refactor(runtime): read Antigravity, Cline, Prime Agent and Cursor readiness from rule files

Adds agent-state-rules/: a zod-validated JSON file per agent, one priority list of
screen rules per agent (idle with strength and requiresQuiet, or hold), and text
anchors that feed the shared, position-ordered blocked layer every pane reads first.

The three screen-ruled agents and Cursor's approval menu and prompt move to data;
the Antigravity text scan stays code as a named anchor. Their old code paths are
deleted. Every other agent still runs through the existing lanes, unchanged.
The readiness census baselines are untouched and pass.

Refs STA-9098

* test(runtime): cover the agent state rule engine's schema, priority, rows, anchors and lanes

Refs STA-9098

* fix(runtime): refuse rule patterns that repeat an optional or alternating group

The load-time regex check only flagged a repeated group whose body held * + or {,
so (a?)* and (a|aa)+ passed though both backtrack exponentially. A repeated
group's body must now be fixed: no quantifier of any kind and no alternation.
The comment states the remaining polynomial gap instead of claiming linearity.

* refactor(runtime): give agent state rules and text anchors one when/answer shape

Every rule and text anchor is now when (a region and what it must show) plus
answer, each a discriminated union, so part (b) adds title, text and status
regions and working or blocked answers as new variants instead of new fields.

- Cursor's prompt is two anchors answering working and idle; the one-off
  workingIfAfter and followedBy fields become a general after test.
- Anchor literals and the probe banner must be lowercase, since they are
  matched against the lowercased tail.
- screenProbeBanner moves under profile, the place for non-detection facts.
- why is required on every rule and anchor.
- A blocked anchor must name a lastOf literal, which the prefilter keys on.

* docs: point the readiness evidence docs at the agent state rule files

* refactor(runtime): read Codex, Claude, OpenCode, Pi, OMP and Gemini readiness from rule files

The rule engine gains the regions and answers these agents need, as closed-list entries:
- rule regions `title` (the classified title status) and `text` (one of the file's idle text
  anchors, settled), and a `predicate` form of the screen region for named engine scans;
- `withoutClock: skip` for strong quiet rules a clockless pane must not believe;
- anchors (renamed from textAnchors) gain a `title` region, and `live` and `hold` answers;
- `profile.screenSource` (trusted grid or live screen), and an `unknown-pane` file for panes
  with no known agent.

Codex's header, composer and provisional-startup checks become named predicates referenced
from codex.json; its ready header, header and startup hold become shared text anchors. Native
idle title markers become shared title anchors; name-only title handling becomes each agent's
idle-title rule. The agent-specific branches in terminal-wait-detection.ts and
tui-idle-evidence.ts are deleted, and the "later live prompt cancels a blocker" rule now reads
only rule-file anchors (plus Muse, which moves in part b2).

No behaviour change: the readiness census baselines are untouched and pass.

Refs STA-9098

* test(runtime): cover the rule engine's title, text and predicate regions and the bundled anchors

Refs STA-9098

* fix(runtime): reject a rule file that repeats an anchor or rule id

A text rule names its anchor by id, so a repeated id let a file pass validation and then throw
while compiling. Also states that engineVersion bumps once a version ships; version 1 is still
being defined.

* refactor(runtime): fold the working anchor answer into live

The engine treated an anchor's working and live answers identically: both mark a live prompt
that cancels an earlier blocker and settles nothing. Cursor's busy prompt now answers live, so
anchors have one non-settling prompt answer.

Refs STA-9098

* refactor(runtime): read the shared π title anchor from pi.json alone

Pi and OMP paint the same `π - <session>` rest title, and title anchors apply to every pane,
so one copy covers both.

Refs STA-9098

* refactor(runtime): key every rule file and read the trusted screen from screenSource alone

readsTrustedScreen no longer also asks for a screen rule (every trusted file has one, and the
schema requires screenSource where it matters), so rule-less files need no filter. A rule's
match is a plain boolean, and compileTitleAnchors is module-private.

Refs STA-9098

* test(runtime): pin that a clocked Codex pane takes no other agent's ready text

No test failed when holdsReadyTextToQuiet was removed; this one does.

Refs STA-9098

* fix(agent-hooks): keep an OMP approval wait until omp resolves it

omp posts tool_execution_start a few milliseconds after
tool_approval_requested, while its Approve/Deny select still holds the
human. Both mapped onto the pane row, so the working event overwrote the
blocked one and the pane read as busy for the whole prompt.

A working event now leaves an OMP approval wait in place; only
tool_approval_resolved or a new turn ends it. An ask row is unchanged:
its own tool_execution_end ends it. The test replays the order a live
omp 17 run posted for a denied bash call.

Refs STA-9100

* refactor(runtime): select the fresh hook row on any of a terminal's handles or pane keys

selectFreshExplicitAgentStatus matched one handle and one pane key and
returned only the mapped status. The row selection now takes sets of
handles and pane keys, an optional received-at floor, and returns the
row itself, so a reader can see the main agent's own state. The old
function keeps its signature and result on top of it.

Refs STA-9100

* feat(runtime): let tui-idle read hook state for agents whose hooks cover the whole turn

tui-idle read no hook state. Hook state reached readiness only through
the `<Agent> ready` titles the window writes, so a headless `orca serve`
never saw it (#16095), and Codex settled only once its screen had been
quiet for three seconds.

Rule files gain `profile.hooks: "authoritative" | "identity-only"`,
defaulting to identity-only. Codex (with its Interrupt hook), OpenCode,
OpenCode 2, Pi and OMP are authoritative. For them a fresh hook-store
row decides ahead of every other lane:

- the main agent's turn decides (`mainAgent.state` when published), so a
  subagent's Stop does not end the lead turn: done settles strong,
  working holds, a permission wait never settles;
- the tail's blocked text goes through the existing permission arbiter
  with the turn as its explicit status, so a denied prompt's dialog left
  in the tail no longer blocks a turn the hook says ended;
- the row joins on every pane key and terminal handle the PTY owns.

No row, a stale, restored or other agent's row, a session-start done,
and a row from before a PTY respawn all fall back to today's lanes. That
keeps startup on the screen and text rules: Codex posts SessionStart
only with the first prompt. Claude, Cursor, Gemini and the rest stay
identity-only.

The readiness census has no hook server, so its frames are unchanged.

Refs STA-9100

* docs(agent-status): record readiness as a reader of the hook store

Refs STA-9100

* fix(runtime): ignore a hook done older than the latest input Orca wrote

A finished turn leaves a fresh `done` row. A caller that sends the next
prompt and waits at once could settle on it before the new turn's first
hook arrives, so the wait returned while the agent was starting work.

Orca's own input writes (terminal send, agent prompts, mailbox pointers)
now stamp a per-PTY input clock, and the hook lane reads no `done`
received before it; the pane falls back to the screen and text rules
until the agent reports again. A `working` row is unaffected.

Refs STA-9100

* docs(agent-status): note the input floor on the hook lane's done

Refs STA-9100

* fix(runtime): take the hook lane's input floor from the PTY run's input record

The hook lane ignored a done older than Orca's latest write to the pane, kept in a
new per-PTY map stamped by a wrapper threaded through four write sites. The PTY
run register already sits on both write funnels, so it now records the last
input (launch writes included, terminal replies not) and the lane reads it.
Keys the user types now count too, which closes the restart-in-the-same-shell
gap: typing `codex` to relaunch no longer lets the previous process's done read
ready while the new one boots.

The respawn floor moves from the shared row join into the lane, beside the
input floor; the freshest row predates a floor exactly when every row does.

* test(runtime): drop runtime hook-lane cases the unit suite already proves

Working over a ready title, a permission wait, and an identity-only agent are
decided inside evaluateTuiIdle and covered there; the runtime suite keeps the
wiring: the join, both floors, Pi's own OSC 133 markers and the arbiter.

* fix(runtime): record a PTY's last input even when main adopted it without a spawn commit

A materialized pane re-adopted by the renderer returns before the spawn-commit
site, so it had no run record and its input never moved the hook lane's floor.
The last input now lives beside the run records: any PTY's input counts, and a
new process's commit still clears it.

* fix(runtime): keep a running process's input time when main reattaches or adopts it

A reattach or adoption commit without an incarnation id cleared the PTY's
last-input time, so a prompt sent just before an SSH adoption was forgotten
and the hook lane could accept the previous turn's done as ready. Only a new
process (or a reattach naming a different incarnation) now starts clean; the
first-input fact follows the same rule.

* docs(runtime): say why a lead turn that ended reads ready while a subagent runs

* fix(runtime): refuse uppercase contains terms in text anchors, which read the lowercased tail

A text anchor's after and lines tests run on the lowercased tail, so an
uppercase contains term loaded and then never matched. Build the text test
schema from the literal it accepts and give anchors the lowercase one. Also
drop a probe-banner early return that no bundled catalog reaches.

* refactor(runtime): state Codex's provisional startup and title anchors as plain rules

The provisional-startup hold becomes a lastOf anchor with an all/none test, so
its TypeScript scan goes. Title anchors drop their status field (every caller
already gates on an idle title), and withoutClock keeps only the value a rule
can set.

* fix(runtime): leave Codex readiness to its title and screen rules

Codex before its Interrupt hook posts nothing for an Esc mid-turn, so its hook
row stays working and a hook-authoritative tui-idle wait hangs until the row
goes stale. Current Codex already settles fast through its ready title.

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-02 01:36:30 -04:00

566 lines
37 KiB
Markdown

# Agent status store
## Status
The current boundary is PR 2A: structured sessions use the hook server's fully
scoped canonical store; unbound PTY/relay evidence remains in an isolated legacy
adapter. Do not remove the renderer bridge or its publication filters in this
slice: they still carry native-chat child rows.
The sections below record the original 2026-09-09 rollout. Its PR 1a and PR 1b
have landed; its proposed PR 2/3 sequence is superseded by that boundary:
1. main-only: every producer writes into one store and `worktree ps` reads it,
split into 1a (structured sessions join the store) and 1b (the runtime's
duplicate retained store is deleted);
2. renderer: the sidebar becomes a subscriber and stops re-deriving rows;
3. shared: one worktree-status rollup and one freshness rule for every reader.
## The problem this solves
Orca shows "what is this agent doing" in four places: the desktop sidebar, the
`orca worktree ps` command, the mobile app, and the agent dashboard. Before
#19217 those readers did not even share their inputs. After #19217 they share
the structured-session mapping and nothing else.
An audit on 2026-09-09 found six producers and three consumers, and three
separate copies of the same row inside the main process alone:
| Main-process copy | Keyed by | Owned by | Persisted | Evicted |
| --------------------------------- | --------- | --------------------------------------------------------------------------------- | ------------------ | ---------------------------- |
| hook server `lastStatusByPaneKey` | paneKey | `src/main/agent-hooks/server.ts` | `last-status.json` | tab close, pty exit, hydrate |
| runtime `RuntimeAgentRowStore` | paneKey | `runtime-agent-row-store.ts` (deleted in PR 1b) | no | pty exit only |
| structured feed `published` | sessionId | `src/main/native-chat/agent-session-wire/structured-agent-session-status-feed.ts` | no | never (a broadcast cache) |
The second copy is a duplicate write: the OSC status parsed in main is
forwarded to the hook server _and_ retained in the runtime store from the same
call (`orca-runtime-create-terminal-side-effect-command-code-detector.ts`).
The third copy is keyed differently and never reaches the hook server at all,
which is why `worktree ps` grew its own adapter for it in #19217.
Each reader then applies its own precedence and freshness rules, so the same
pane can legitimately read differently on the desktop, on the phone, and in
the CLI.
## The rule
**The execution host owns agent status, in one store, and every reader
subscribes to it.** This follows the boundary in
[`ssh-execution-boundary.md`](./ssh-execution-boundary.md): the host that runs
the process is the only party that can observe it, and the client is never
authoritative for execution state.
Three consequences:
- One store per execution host. A remote host keeps its own store and the
client mirrors it down, as the web-session mirror already does. Mirroring is
not merging: a client never writes its observations back to a host.
- Precedence is decided once, at write time, with provenance recorded on the
row. Readers never re-adjudicate hook versus terminal versus structured.
- Readers keep only presentation policy and user facts: the 30-minute display
decay, acknowledgements, dismissals, unread. Those stay reader-side but
become one shared implementation (PR 3).
## The store already exists
The hook server's state is that store today for every PTY-based agent. The
audit established:
- hook HTTP posts, the WSL and SSH relay receivers, and main's own OSC parse
all converge on the same `applyNormalizedStatus` path, stamped with the
authority id `main-agent-hooks`;
- it alone holds pane authority: launch tokens and their hashed commitments,
retired-pane fences, pane-key aliases, per-connection ordering watermarks,
and the evidence-age map that must outlive a transport clear;
- it alone persists, with a seven-day hydrate window and the
`restoredUnconfirmed` stamp that keeps a hydrated row from ever reading as
live truth;
- it already fans out to both renderer windows over `agentStatus:set` and
`agentStatus:clear`, and serves `agentStatus:getSnapshot`.
Nothing else in main carries those guarantees, and building a second store
with them would be the wrong direction. So the design is not "add a store". It
is: **route the two producers that bypass the hook server through it, then
delete the copies.**
## PR 1a: structured sessions publish into the store
No renderer behavior changes. The sidebar keeps receiving the same IPC events
it receives today, plus structured-session rows it currently derives itself.
### Structured sessions publish into the hook server
The structured feed keeps its job of projecting a session's journal into a
summary and streaming it to subscribers. On every publish it additionally
ingests the summary into the hook server as a status row:
| Row field | From |
| --------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `paneKey` | `structuredAgentSessionPaneKey(tabId, sessionId)`, the key the renderer already uses; its leaf is UUID-shaped so pane-key validation accepts it |
| `tabId` | `structuredAgentSessionTabId(sessionId)` |
| `worktreeId` | `summary.workspaceId` (a folder workspace id is a valid value) |
| `state` | `structuredAgentSessionAgentStatus(summary).state`: the lead's own status folded with the live child records the store holds for the session (not the summary's task list), so a settled lead whose subagent still runs reads `working` |
| `workingMode` | `'monitoring'` from the same fold when watch loops are the only live child work; omitted otherwise, which clears it on the row |
| `mainAgent` | the main agent's own state before the fold, its last-turn verdict (`summary.turnOutcome`, present only while idle) and its own clock; see "The main agent fact" below |
| `structuredHost` | `'owned'` while `summary.hostExecutionOwned` is set, otherwise `'held'`; `worktree ps` derives its row's `structuredHostOwned` from it |
| prompt, tool, last message, model, provider session | the summary's fields |
Sessions with no request (`status === null`) produce no row. A request is a
turn record, an assistant message, a user message the provider journaled itself
(history, an older host), an accepted or unanswered send, or a send the agent or
its start refused; a send that was withdrawn, or left undelivered by a
restart or a close, fails nobody and makes nothing listable.
`summary.turnOutcome` is the latest request's verdict: its turn's outcome, or
`failure` for a send the agent or its start refused (a send that joined a running
turn is answered by that turn). A turn the provider gave no outcome reads as its
host-observed end through `agentTurnVerdict`: `interruption` for an `interrupted`
lifecycle, `unconfirmed` for an `unverifiable` one. It is derived on each read,
never journaled. The row also publishes `interrupted` from
`mainAgent.outcome`, exactly as the hook lanes do. When the host revokes live ownership the row is re-set
without the flag; when the host closes or evicts the session the row is
dropped. Both already exist as feed events (`revokeLive` and the roster
filter in `liveSessionSummaries`); PR 1 turns them into store writes.
Dropping the session from the host's map and dropping its row are one
operation, `forgetStructuredAgentSession`. The store keeps a row until told,
and a host-owned row bypasses the staleness check, so a deletion path that
forgot the row would strand a permanently working-looking agent.
Two rules the ingest must keep:
- **Never persist a structured row.** The journal is the durable truth for a
structured session and the host republishes on restore. A structured row in
`last-status.json` would hydrate as `restoredUnconfirmed` and then fight the
live republish. The serializer skips rows carrying `structuredHost`, and
hydrate drops any such row found on disk. Applying one therefore also skips
the persist schedule: the walk and stringify could only reproduce the file
that is already on disk, once per debounce window for every streaming chat.
- **Never let it fight a hook row.** A structured session has no PTY, so no
hook or OSC event carries its pane key. The ingest still goes through the
disposition gate so a retired pane key is refused like any other.
Applying one does still run both status fan-outs, and that is intended rather
than incidental. `notifyStatusChangeListeners` is what feeds
`agentAwakeService`'s power-save blocker, and `subscribeEnrichedStatus` is what
feeds `AgentSessionTransitionRecorder`'s stats, so joining the store enrolls
native chats in both. A working native chat is real work and should hold the
machine awake exactly like a PTY agent does.
The drop side routes through `dropStatusEntry`, not `clearPaneState`: a
pane-status-clear reaches the renderer, and until PR 2 the renderer's own feed
bridge is that pane key's writer. It also passes `preserveResumeIdentity:
false` — the `providerSessionOnly` remnant a dismissed pane keeps exists so the
agent can be resumed in that pane, and a structured session has no pane and
keeps its resume identity in the record store. Like every other
`dropStatusEntry` caller, it emits no pane clear, so a session dropped
mid-`working` leaves `AgentSessionTransitionRecorder` holding an open stats
session until its LRU evicts it; that gap is shared with the user-dismissal
path and is not specific to structured rows.
The ingest lives in the feed, not in `structured-agent-session-host.ts`, which
sits at the file-length cap.
### `worktree ps` becomes a reader
The structured adapter added in #19217 is deleted, and structured rows reach
`worktree ps` through the same snapshot as every other row. The
retained-versus-hook reconciliation in `collectRuntimeWorktreePtyAgentSources`
stayed until PR 1b removed the store that fed it. What this step settles is
the admission gate that decides which rows a worktree listing may show:
- a hook or OSC row needs its tab mirrored or a connected pty, as today, and
SSH rows stay exempt because their tabs may exist only remotely;
- a row carrying `structuredHost` is admitted while the host holds the session, and
the host's drop on close is what removes it. No tab-mirror requirement: a
structured session's tab lives in the renderer's own tab state, and a
headless host has no renderer to mirror it from. That argument only holds if
the headless host is itself wired to the store, which is a separate
obligation per entry point: the Electron hosts (desktop and `orca serve`)
share `main-process-runtime-service.ts`, and `orcad` constructs its own
runtime in `src/main/orcad/orcad-entry.ts`. A host missing that wiring lists
no agents at all, not just no structured ones, because `worktree ps` reads
the same snapshot for every row.
The freshness bypass for host-owned structured rows already exists in
`isFreshNonDoneAgentStatus`; with the flag now on the row it becomes the only
path, and the hand-rolled check in `runtime-worktree-agent-rows.ts` goes.
### Wire compatibility
`AgentStatusIpcPayload` gains one optional field, `structuredHost`, and the
`worktree ps` row gains `structuredHostOwned`. Under rule 1 of
[`remote-wire-compatibility.md`](./remote-wire-compatibility.md) both are safe:
an old client ignores them. `worktree ps` rows keep their shape and vocabulary,
so the mobile app sees no change.
Until PR 2 the main process does not forward structured rows to the renderer
over `agentStatus:set` or `agentStatus:getSnapshot`. The renderer's feed
bridge still writes those rows itself, and forwarding them too would give one
pane key two writers. Removing that filter is the first step of PR 2.
### The main agent fact
Claude, Codex and Grok hook rows and structured-session rows publish the combined
`state` and, beside it, the main agent's own state as `payload.mainAgent`. Other agents'
rows and terminal-title-only rows carry none, and readers fall back to `state`:
```ts
mainAgent?: { state: AgentStatusState; outcome?: AgentTurnOutcome; stateStartedAt: number }
```
`state` still answers "what should the user see" and folds live child work in,
so a settled main agent whose subagent still runs reads `working`. `mainAgent` answers
"what is the main agent itself doing", which the fold used to destroy at publish
time; every guard that reconstructed a fragment of it (`fromChildWork`, the
persisted `claudeLeadBoundaryChildOnly` flag) now reads `mainAgent` instead of a
stored copy. A Claude row whose `mainAgent` is `done` while a child agent still
works (including a child's permission wait) refuses OSC, which carries no child
identity; the children's own lifecycle hooks settle it. `outcome` is the recorded verdict on
the main agent's most recent finished turn, present only while `mainAgent.state` is
`done`. It is reported by the provider, or is a `cancellation` Orca inferred
from the user's own interrupt keystroke, or a `superseded` the host recorded when
a newer request replaced a structured Claude turn before it ended (it names no
sender, and sets no legacy flag), or, on a structured row whose turn the
provider gave no verdict, is what the host observed of its end: `interruption`
(a proven death nobody asked for) or `unconfirmed` (an end it cannot prove,
never success). The journal's turn outcome, by contrast, stores only recorded verdicts: the provider's, a `cancellation`, or the host's `superseded`; `interruption` and `unconfirmed` are derived from the turn's lifecycle state and never stored. A plain end of turn carries none, because absent
means unknown and a provider that omits its interrupt flag must not turn a
cancel into a success.
In the Claude hook lane the cancellation comes primarily from Orca's own
inferred interrupt (`markClaudeLeadTurnInterrupted`), because current Claude
sends no hook at all on a cancel and no `is_interrupt` on Stop; that flag on a
turn boundary remains a secondary source for builds that send it, and
`StopFailure` maps to `failure`.
Readers decode the verdict through one accessor, `agentMainAgentVerdict`, which
reads the main agent's own state, not the combined row's: `mainAgent.outcome`
while `mainAgent.state` is `done`, then the legacy `interrupted` flag as a
cancellation, which alone needs the combined `done`. So a main agent that
failed while its subagents still run has a verdict on a `working` row. Every
copy of a row (state-history entries, sleep records, `worktree ps` rows) takes
the verdict through `agentVerdictFields`, which carries `interrupted` and the
whole `mainAgent` (state, outcome and its own clock) together, so a copy agrees
with the row and can date a failure by `mainAgent.stateStartedAt`.
Display reads the verdict through `agentVerdictDisplayMark`. A fault marks the
agent failed whatever the combined state, because it is news the user must see
even while subagents run: a `failure`, and an `interruption`, a turn cut short
by anything other than the user or a newer request. A user's stop (`cancellation`)
marks it interrupted, drawn in the muted tone with the row text "Interrupted by user";
a turn a newer request replaced (`superseded`) marks it interrupted in the same muted
tone with the row text "Interrupted"; and `unconfirmed` marks it unconfirmed, all only
on a `done` row, so a stopped
or finished main agent with live child work still reads working. The folded
turn header follows the same mark: "Failed after N", "Interrupted after N", or
"Worked for N".
Each subagent keeps its own row and state. Container rollups (worktree card,
terminal tab, Cmd+J) rank a pending question first, then a failure, then live
work, then an unconfirmed end, then a user's stop, then done. On the worktree
card, a failure retained after its agent's pane went away has no expiry, so it
ranks below live work and above an unconfirmed end. Lifecycle waiters keep
reading the combined `state`.
Policy splits the verdict two ways. Clean-finish policy (hibernation, pane
ownership, the star-nag value moment) treats a failure, an interruption and an
unconfirmed end like a cancellation (`agentTurnEndedUncleanly`). Attention
(completion time, Smart Sort, sticky retention, Cmd+J Recent) demotes only a
turn ended on purpose, the user's stop or a newer request that replaced it
(`agentTurnEndedOnPurpose`); a failure, an interruption or
an unconfirmed end ranks like a completion.
Admission is one function, `normalizeAgentStatusPayload`, on the relay wire,
IPC and disk. A malformed `mainAgent` drops the field and keeps the row. Old hosts
send none and readers fall back to `state`. Hook rows persist it inside the
payload; hydration maps an older row's `claudeLeadBoundaryChildOnly: true`
onto `mainAgent: { state: 'done' }` when the row has no `mainAgent`, and never writes the
flag again. Hydration seeds the Claude main agent record straight from a saved
`mainAgent` that is `done`, so the children's drain can still settle the row after
a restart. `claudeRunningNonAgentTask` is persisted alongside because it is the one
child-work fact `mainAgent` cannot express: a shell running beside the main agent,
whose liveness hydration does not restore. Hydration seeds only a row that says
`false`; a row silent about it stays unseeded. The row builder pairs the two facts in
one place: a listener event restates the shell fact, and any other write (an OSC
repaint, an inferred answer) keeps it only while `mainAgent` is unchanged. A child's
sticky permission prompt still records the main agent's own progress and background
evidence in the held row, and pushes the held row to subscribers when `mainAgent` changes.
Every lane, Codex included, combines through the fold. A child waiting on a
human is a fold input (`childWorkLiveness: 'waiting'`, derived from the child's
own `waiting` state; a child's `blocked` means it failed and stays live work)
and makes the row wait whatever the main agent is doing, unless the main agent
is itself asking. Only the Codex hook lane feeds that input today. Known
divergences, pinned by name in the parity table
(`src/shared/main-agent-status-parity.test.ts`) where they are reachable, so a
reader does not mistake them for drift:
- The Claude hook lane holds a child's permission wait in one slot on the
displaced main agent record (`waitingAgentId`, `stateBeforeWait`), not on
the child. It publishes the displaced state as `mainAgent`, but the next
main agent event overwrites the slot, so the row stops reading `waiting`
while the child is still asking, and a second asking child replaces the
first.
- The structured lane has no per-child wait: a child's pending prompt makes
the session `attention`, which reads as the main agent's own `blocked`.
- The Codex hook lane drops its roster on a root `Stop` when it tracks no
child transcripts, so a still-running or still-asking child stops holding
the row.
How the main agent's turn ended is not a fold input. A cancel is a verdict on
the main agent, carried as `mainAgent.outcome: 'cancellation'` (and, for
readers that predate `mainAgent`, as the row's `interrupted` flag on a `done`
row); it never retires a shell, scheduled check or subagent the turn left
running. That work leaves the row only when its own inventory omits it or the
session ends, so a cancelled turn with a still-running shell reads
`monitoring` in every lane, and the parity table in
`src/shared/main-agent-status-parity.test.ts` drives that story through all of
them. The same rule governs the cancel Orca infers from Ctrl+C: for any row
that publishes `mainAgent`, the inference is admitted only when
`mainAgent.state` is `working`, so Orca does not treat a Ctrl+C at the idle
prompt of a row held open by child work as a turn cancel (Codex also keeps the
child-evidence guard, and a row without `mainAgent` keeps only that guard).
The keypress itself is not inert, though: measured live, Claude 2.1.280 stops
its background subagents on a single idle-prompt Ctrl+C (shells survive) and
Codex 0.156.1 quits outright, so refusing the inference can leave the row
showing a subagent its CLI already stopped. The synthesized row is the fold
of the cancelled main agent with the child work the pane's owner can see: the
local listener's roster for a local pane, the row's own subagents and shell fact
for a relayed one, whose provider records live on the relay.
The store holds that verdict against restatements that predate it
(`server-cancel-verdict-latch.ts`), because a relay never learns of a cancel
the desktop infers and some TUIs emit late same-turn hooks. The hold is read
off the row (`mainAgent.outcome: 'cancellation'`), never stored beside it, and
dies on a new turn (a main agent prompt submission, a changed or explicit
prompt, a session start) or the provider's own settled `mainAgent`. Child and
replayed events under the hold keep the cancelled main agent and are re-folded
with their own child evidence.
## PR 1b: the runtime's retained row store is deleted
Landed. `RuntimeAgentRowStore` is gone, and with it the retained-versus-hook
reconciliation in `collectRuntimeWorktreePtyAgentSources`. The hook server's
store is now the only main-process copy of a PTY agent's row.
### The five call sites
| Call site | Before | After |
| ------------------------------------------------------------------------------ | --------------------------------------------------------------- | -------------------------------------------------------------------------------------------------- |
| `orca-runtime-create-terminal-side-effect-command-code-detector.ts` `retain()` | second write of the OSC payload already sent to the hook server | deleted; the event now carries the pane's `terminalHandle` and the hook ingest keeps the only copy |
| `...command-code-detector.ts` `clearPty()` | drops rows on pty exit | deleted; pane teardown already clears the hook row |
| `orca-runtime-get-worktree-ps.ts` `values()` | fed `retainedSnapshots` | deleted; the reader keeps only `hookSnapshots` |
| `orca-runtime-serialize-agent-prompt-submission.ts` `getFreshExplicit()` | retained row first, hook rows second | `selectFreshExplicitAgentStatus`, hook rows only |
| `orca-runtime-prune-mobile-session-tab-group-layout.ts` `getFreshForMobile()` | pane key, then pty id | `selectFreshAgentRowForMobileTab`: pane key, then `terminalHandle` |
Both readers moved into `runtime-hook-agent-row-selection.ts`, which also owns
`RuntimeAgentRowSnapshot` now that nothing retains one.
### `terminalHandle` is the row's join back to its terminal
The retained store's only real extra was the pty id, and two readers used it.
The plan said to stamp the event's `ptyId` into `terminalHandle`; that was
wrong. A terminal handle (`term_<uuid>`) and a pty id are different
identifiers, and `getFreshExplicit` was already comparing hook rows against a
real handle. What landed instead:
- `AgentHookEventPayload` and the runtime's terminal-status event gained an
optional `terminalHandle`. The detector resolves it once per chunk through
`getAgentStatusTerminalHandleForPaneKey` — the same lookup the renderer-facing
IPC boundary already runs for every row, so the two surfaces cannot disagree
about which terminal a pane is.
- `applyNormalizedStatus` carries the handle forward when an incoming event
resolves none. Only main's OSC parse can resolve one, so an HTTP hook post for
the same pane would otherwise erase it.
- It is never persisted. A handle belongs to the runtime that issued it, and a
hydrated one could only rejoin a row to somebody else's terminal.
- `toAgentStatusIpcPayload` publishes it, which also makes `getFreshExplicit`'s
long-dead handle comparison live: the runtime reads raw snapshot rows, and
before this nothing ever stamped the field on them.
`worktree ps` uses it too. `ConnectedPtyEvidence` traded its flat `ptyIds` set
for `ptyIdByTerminalHandle`, so a row still resolves the connected PTY behind
it — which is both the working-terminal rollup's match key and the last rescue
for a row whose pane binding was nulled by a controller incarnation change.
### The change detector had to move with the store
`retain()` was not only a store: its boolean return was the signal that
republished `session.tabs` for a status-only transition, which no title change
covers (#7970). `hook-status-session-tabs-invalidation.ts` already mirrors that
projection change set, including restore provenance and terminal-handle joins,
so the replacement was to route the signal off the store rather than build a
second comparator.
`installHookStatusSessionTabsRepublish` now owns all three arms — enriched
status, pane clear, and the status-drop tap a dismissal emits — and both hosts
install it.
### Both hosts, not just the desktop one
`orcad` constructed its runtime with no `onTerminalAgentStatus`, so main's OSC
parse never reached the store there and the retained copy was the only carrier.
Deleting it without wiring orcad would have made a headless host list no PTY
agents at all. `orcad-entry.ts` now binds the producer and installs the
republish signal, alongside the snapshot and structured sink it already had.
### The intended behavior change
A row the user dismisses on the desktop leaves `worktree ps` and the phone at
once, instead of lingering until the pty exits. One store means one dismissal.
Legacy numeric pane keys remain a bounded compatibility case. Persisted layouts
register aliases to their stable leaf owners; an in-process OSC observation may
also retain a numeric key only when the runtime supplies the matching tab, PTY,
and terminal handle. HTTP and relay ingress still require a stable key or a
registered alias, and numeric rows are never persisted.
## PR 2: the renderer subscribes
With structured rows arriving over `agentStatus:set`, the renderer's
`StructuredAgentSessionStatusBridge` no longer needs to write status; its
unmount cleanup becomes a tab-close signal to the host. The IPC applicator is
the single writer for observed status. The 2026-09-09 audit sorted the other
writers:
| Writer | Disposition |
| ----------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------- |
| Command Code output seeds, parked-pane seeds, pty-exit removal | delete; main already emits the same facts |
| structured bridge status writes | delete; main now publishes the row |
| launch placeholder seeds (a user launched an agent with a prompt) | keep for now; main holds the launch config and can seed later |
| dismissal, acknowledgement, unmount | keep; user facts and component lifecycle |
| remote-runtime OSC parse (bytes never transit local main) | keep, fenced behind the host's published row once the host is new enough; rule 3 of the wire doc applies |
| web-session mirror receipt clock | keep; the decay rule needs both clocks from one machine |
The Command Code done-settle window is renderer policy with no main
equivalent. PR 2 either moves it into main's detector or leaves it, and says
which.
## PR 3: one rollup, one clock
The worktree card status is derived three times: `lib/worktree-status.ts` in
the renderer, `runtime-worktree-status-projection.ts` in main, and
`agent-row-display.ts` in mobile, which hand-copies the 30-minute constant.
PR 3 moves the rollup and the decay into `src/shared` and makes all three
call it.
## Readiness reads the store
`terminal wait --for tui-idle` is a reader too. Before STA-9100 hook state
reached it only through the `<Agent> ready` titles the window writes, so a
headless `orca serve` never saw it (#16095). Now an agent whose rule file says
`profile.hooks: "authoritative"` (OpenCode, OpenCode 2, Pi, OMP) has its
fresh row read straight from the store, through the same
`selectFreshExplicitAgentStatusRow` join prompt-receipt verification uses
(`src/main/runtime/tui-idle-hook-lane.ts`):
- the main agent's turn, not the combined row, decides: `mainAgent.state` when
published, so a subagent's Stop does not end the lead turn. `done` settles
the wait, `working` holds it, and a permission wait never settles. The tail's
blocked text goes through the existing permission arbiter with the turn as
its explicit status, so a denied prompt's dialog left in the tail no longer
blocks a turn the hook says ended;
- the row joins on any pane key or terminal handle the PTY owns; a pane neither
reaches, a stale or restored row, a session-start `done`, a row from before
the PTY respawned, and a `done` received before the pane's latest input all
leave the decision to the screen and text rules, which is also how startup
readiness works before an agent's first hook. The input is the PTY run's
`lastInputAt` (`terminal-run-facts.ts`), which both write funnels record, so
a key the user typed counts like a prompt Orca sent: the next turn's first
hook may still be in flight, and an agent restarted in the same shell has
not posted one. A shell command marker is no process boundary: Pi paints
OSC 133 zones itself;
- every other agent stays `identity-only`: Claude sends no event when an
approval is denied or Esc stops a tool, so its row can sit at `waiting` or
`working` forever, and the rules keep deciding. Codex is identity-only too:
before its `Interrupt` hook an Esc mid-turn leaves the row `working`, and an
older TUI can hand its hooks to a newer shared app server, so no version
check tells which Codex posts it. Current Codex settles fast anyway, since
`Interrupt` drives its `Codex ready` title.
The titles stay for display; remote clients read them.
## What does not change
- The hook scripts, the OSC 9999 wire format, and the relay protocol.
- The status vocabulary. `working / blocked / done` for rows,
`working / attention / idle` for structured summaries, mapped once.
- The `live / unverifiable / exited` verdicts for remote work. Loss of contact
clears nothing; the SSH exemptions in the admission gate stay.
- Hydration honesty: a restored non-done row is `restoredUnconfirmed` and is
never fresh.
## PR 1b reliability contract
- **Invariant (`agent-session.status-host-ownership`):** each execution host has
one agent-status store; OSC, hooks, and structured sessions write it, while
desktop, `worktree ps`, and mobile only project it. Dismissal, certified PTY
exit, and provider-generation replacement remove the same row everywhere;
transport loss alone removes nothing.
- **Failure source:** the deleted runtime row store duplicated OSC observations,
keyed them by a different terminal identity, and outlived a dismissal from the
hook store. Relay replay could also make old evidence look fresh when readers
used its new delivery timestamp.
- **Oracle:** one OSC observation appears through the hook snapshot in
`worktree ps` and mobile, and one store dismissal removes it from both without
stopping the PTY. Focused tests also require leaf/incarnation-handle rejoin,
legacy numeric-pane compatibility, certified-exit and provider-generation
cleanup, evidence-age freshness, and exactly-once startup/stop teardown.
- **Gate:** `terminal-performance.osc-status-scan-budget` covers the unchanged
bounded OSC parser and the runtime projection. There is not yet a dedicated
blocking multi-surface status-store gate; the focused suites below are the
accepted gap until they accumulate reliability-gate soak evidence.
- **Provider/platform coverage:** local and daemon-backed PTYs are covered by
runtime tests, and SSH relay loss/replay semantics by relay integration tests.
The projection is shared by git worktrees and folder workspaces. WSL uses the
same store and admission code but has no live run here; Linux and Windows
runtime execution, native mobile clients, and mixed-version paired clients
remain validation gaps.
- **Performance budget:** publication stays event-driven with no new polling or
subprocesses. One mobile projection clones the status snapshot once, builds
pane/handle indexes once, and has a deterministic call-count test; lifecycle
cleanup is bounded by the existing status and handle inventories, and orcad
tests prove listeners clean up once on failed startup and repeated stop.
- **Diagnostics:** existing hook-listener errors name the pane and PTY, while
status-store tests pin delivery versus evidence clocks. No new telemetry or
raw terminal data is emitted.
- **Residual gaps:** rendered Electron/mobile behavior, live SSH reconnect, and
Linux/Windows/WSL execution require the platform QA pass. The current
cross-version gate does not cover `session.tabs` content.
## Verification
- Unit: ingest a structured summary and read it back through
`getStatusSnapshot`, `worktree ps`, and the mobile projection; assert the
serializer never writes a row carrying `structuredHost`; assert a hydrated
file that somehow contains one is dropped.
- Unit: the `worktree ps` suites written against the retained store are rewired
to a real `AgentHookServer` (`agent-status-store-wiring.test-fixture.ts`)
rather than deleted, so each still asserts the listing behavior it named. The
dismissal change is pinned end to end in
`orca-runtime-tests/worktree-ps-agent-row-dismissal.spec.ts`, which fails with
the retained store restored.
- Live: the parity check from #19217 (working, done, close, reload) repeated
against the merged store, with both surfaces read from the one row.
## Retired OMP pane recovery
A desktop renderer retirement carries an optional UUID through the existing
`agentStatus:retirePaneAuthority` IPC message. The hook server retains it with
its bounded retirement fence. A validated live OMP new turn consumes that UUID
and echoes `authorityRestartId` only in the live notification. Cached rows,
persistence and startup replay never carry the acknowledgement. Older peers
omit or ignore it and retain explicit attach restoration.
The renderer keeps the UUID in its existing non-persisted retirement tombstone;
every re-retirement mints a new one. A matching acknowledgement may clear that
tombstone only with a successful status write for the existing pane and matching
workspace/connection. Closed tombstones remain `true`, including after the tab
LRU evicts its entry. Closing a retired physical alias revokes its whole group.
This is control-plane retirement correlation, not a second agent-status store.
Fallback restores the hook server's recorded status aliases through the existing
attach-restoration path. The accepted renderer write restores the matching status
alias routes too, preserving group membership for the next retirement. It does
not restore orchestration or launch credentials.
It is scoped to the requesting desktop renderer. A different window's retirement
UUID cannot be cleared by the acknowledgement, and web mirrors keep their existing
host-snapshot/attach behavior.