Files
orca/docs/reference
Brennan Benson 2bd656c384 feat(native-chat): a Claude subagent waiting on a permission prompt reads as waiting (#22634)
* feat(native-chat): a Claude subagent waiting on a permission prompt reads as waiting

A subagent's permission request reaches the parent session's callback naming the
subagent that asked (agent_id) and the tool call it gates (tool_use_id). The
pending request is recorded with the asking agent. On every drain the child-work
producer re-derives which children a pending request blocks and hands that set
to the Claude child decoder, the one owner of each child's live edges: a blocked
child reads waiting on every live edge it reports, and a child that starts or
stops waiting is a live edge of its own. Answering, denying or cancelling the
request returns the child to its prior live state; nothing is stored beyond the
pending requests.

A live task_updated carrying an error now reaches the record as the child's last
message, without an ending or a new state.

The replay test drives a scrubbed capture of the real CLI (foreground allow,
deny, interrupt, background allow, and the main agent's own request) through the
real adapter into the host's child records.

* test(native-chat): a subagent's request names it before its tool call is read

* docs(agent-status): a subagent asking for approval waits in every lane; the parent row keeps the session's own attention

* test(native-chat): hand canUseTool the asking agent without widening the helper's cast

* test(native-chat): an interrupted Claude subagent settles cancelled, not failed

A captured interrupt shows the spawn call's error result ("The user doesn't want
to proceed…") arriving before the subagent's own `task_updated {status: killed}`.
The spawn result ends nothing (the child ends only on its own terminal frame), so
the child stays live until its `killed` status settles it cancelled. A genuine
failure, captured with the subagent on a model that does not exist, sends its
`failed` status before the error result and still ends failed. Both captures now
replay through the real adapter into the host's records.

* fix(native-chat): a Claude subagent's prompt makes the parent row wait, not block

A subagent's pending prompt made the whole session `attention`, which reads as
the main agent's own `blocked` and outranks the fold's waiting arm, so the
parent row read blocked where a CLI Claude parent reads waiting. The main
agent's state now reads only its own pending prompts.

- A Claude prompt row carries the linkage of the agent that raised it: the one
  the permission request names, or the owner of the tool call it gates. The
  same join decides which child reads waiting, so the two cannot disagree.
- The status summary projects the session's own status from root prompts only;
  every other reader (delivery gates, teardown, restart) still asks whether
  anyone is waiting on a human.
- An answer keeps the prompt row's linkage by the journal's own rule: a
  revision that names no producer keeps the row's existing one.
- The child-tool queries gain the prompt's producer, so a prompt row and a
  child record answer "which agent" from the same join.

* test(native-chat): say which ids the permission capture scrubs and which are its own

* fix(native-chat): the status clock dates attention by the session's own asks only

The session's status is now `attention` only for its own pending prompt, so the
clock's fallback to a subagent's ask could no longer be reached, and it read the
journal by a different rule than the status it dates. Both now read root prompts.

The journal also stamps a Codex subagent's prompt with its thread (#22532), so a
Codex child's approval is that child's wait in the Codex lane too. Two tests
written for the earlier rule are updated: a subagent's ask leaves a running
session `working` on its turn's clock, and a Codex child's answered approval
leaves the settled parent's Activity row done with nothing unread.

* fix(native-chat): a completion still says the user is asked when a subagent asks

The turn-completion feed marked a completion `awaitingUser` from the status
summary's `attention`. The status now means the session's own agent is waiting,
so a subagent's pending approval stopped reaching the completion. The projection
now also says whether anyone is waiting on the user, as the delivery gates,
teardown and restart ask it, and the completion reads that.

The waiting-subagent replay answers its prompt with the adapter's current
response shape.

* revert(native-chat): a live Claude task's error stays out of the child's last message

No capture shows a live task_updated carrying an error, and it is unrelated to
a subagent waiting on a permission request; it leaves this PR.

* test(native-chat): settle the Claude session's startup before replaying a subagent's request

A startup frame drained child work during the first await, so answering a
request freed the child even with the answer's own republish removed.

* fix(native-chat): a Claude subagent's prompt row names it as its other rows do

The prompt row stamped only the asking agent's id, so a nested subagent's
request lost the agent that spawned it, its spawn call and its run. It now takes
the linkage the asker's own rows take: the gated tool call's, when that names the
same agent, else the one resolved through the agent's spawn call. The provider's
agent id stays the asker's id.

* fix(native-chat): a subagent's request makes the parent row wait without a child record

The parent row learned that a subagent needed the user only from that
subagent's child record, so a request no record carried (a Codex child the host
never registered, a Claude task past the live cap) left the row working or done
while the approval card sat in the chat.

"Someone in this session must answer" is now one derived session fact. The
projection names two facts instead of a mode flag: the main agent's own status
(attention only for its own request) and structuredAgentSessionAwaitsUser (any
pending prompt). The status summary publishes the second as an optional
awaitsUser, and the shared fold reads it: the main agent's own ask is blocked,
otherwise awaitsUser or a waiting child record makes the row wait. Every caller
picks the fact it means: the completion edge's awaitingUser and the delivery
gate read awaitsUser; the quit snapshot folds the same two inputs the sidebar
does.

* fix(native-chat): a client that predates awaitsUser still reads a subagent's request as attention

A status summary's status is now the main agent's own, so a client built before
the split would read a subagent's request as working (or idle) and fold it with
code that has no awaitsUser input. Clients advertise
agent-session.status-awaits-user.v1; at agentSession.subscribeStatus the host
sends any client that does not the pre-split summary: attention whenever
awaitsUser is set, without the main agent's own tool line, verdict and clock.
The feed and every in-process reader keep the canonical summary. Transitional,
like the turn-item downgrade.

* test(native-chat): a Codex subagent's approval makes its settled parent's Activity row wait

The test pinned the parent row done while a Codex child asked, through a harness
that fed no child records, so it proved nothing about the ask. It now drives the
ask twice through the real host status store: with no child record (the
session's awaitsUser alone) and with the child's own record waiting from
thread/status/changed. Both read waiting with needsAttention while the ask is
open, then done with nothing unread.

* docs(agent-status): a subagent's request reaches the parent row through awaitsUser in every structured lane

The store reference said a Codex child's request still read as the main agent's
blocked and that only the Codex hook lane fed a waiting child. Both structured
lanes stamp the asking child and feed child records, and awaitsUser carries the
request when no record does. The liveness comment goes back to main's: a child's
blocked is a failed task on an older host's legacy rows.

* fix(native-chat): the restart dialog still headlines a subagent's pending approval

The quit snapshot now records the main agent's own state, so a subagent asking
while the main agent worked recorded `working` and the dialog said "Was
mid-reply" where it used to say "Waiting for your approval". The headline now
comes from the snapshot's pending prompt, whoever raised it, with the existing
copy; `state` stays the main agent's own.

* test(orchestration): a subagent's pending approval holds structured mail delivery

Scoping the delivery gate to the main agent's own request left every gate test
green; a subagent's request now has its own case.

* fix(native-chat): a subagent's request is dated by when it was raised, on every client

Since the summary's clock became the main agent's own, nothing dated a wait
that only a subagent's request held: a pre-split client was sent attention with
no clock, where the old host dated it by the subagent's prompt, and a new
client's waiting row fell back to the time it first saw it, so after a reload a
request the user had already read could read unread again.

The session fact is now when someone started being asked: awaitsUserSince, the
oldest pending prompt whoever raised it, and its presence is what awaitsUser
meant. A row waiting on someone else's request takes that as its clock; the
downgrade for a client without the capability dates its attention by it, which
is what the old host published. A cross-version test pinned to the last
pre-split release runs the same journals through that release's projection and
through this one plus the downgrade, and compares the whole summary. The Codex
end-to-end test also reads the host's own status row, and keeps a read ask read
through a later row and a reload.

* test(native-chat): the pre-split parity check compares only the fields the split owns

An additive summary field is safe for old clients, so comparing whole summaries
against the pinned release would redden on one. The wire comment now says how
the downgrade dates attention: the main agent's own oldest ask, else
awaitsUserSince.

* test(runtime): an aged host-held working summary states that nobody is asked

The test built its working summary by overriding the status of a published
approval summary, which still carried awaitsUserSince, so the row correctly
read waiting. It now drops the request as its scenario says.

* fix(native-chat): the chat's subagent block says waiting when the strip does

While a Claude subagent's request was open, the sidebar and the composer strip
read waiting but the subagent block in the chat history a few pixels above
still read "Kicked off 1 subagent working": it shows the journal's roster
state, and the journal records no wait.

The structured chat now hands its transcript the subagents the strip shows
waiting, read from the host's child records through the strip's own row model
and matched by the provider id the roster names each one by. A running entry
the host says is waiting reads waiting in the group row, its entry and its
section head, with the strip's word and the question colour; it reads the
journal's state again as soon as the host stops reporting the wait.

* fix(native-chat): a collapsed subagent group shows a wait beside a failed sibling

A failed sibling took the group row's one alert slot, so a group with a waiting,
a working and a failed child read "1 working +1 failed" and hid the wait; it
now reads "1 working +1 waiting +1 failed". The waiting set keeps its identity
while a child frame changes no wait, so the transcript's subagent rows do not
re-render on every frame, and the test of a wait ending now updates one mounted
row instead of remounting it.

* refactor(claude): one needs-input state on the parent; the asking subagent alone reads waiting

Drop the split of the main agent's own status from a session-wide "someone must
answer" fact: awaitsUserSince, the agent-session.status-awaits-user.v1
capability and its old-client downgrade, and every reader change that only
consumed them (fold, equality, ingest, delivery gate, turn-completion feed,
quit snapshot, resume headline, status clock, status bridge, attention
dispatch) go back to main. The parent row again reads one needs-input state
for a pending request whoever asked, dated as before.

Kept: a request's owner recorded once on its prompt row with full producer
linkage; the asking subagent's own record reads waiting, re-derived on every
update; the chat history's subagent block reads that same state; an answered
subagent request stays in its subagent's group.

A subagent now waits only on a request the user can still answer (its card
open, no answer underway), and the adapter frees it before the host records an
answer or dismissal. So a waiting child record always sits beside the pending
card, and main's fold never reads the parent as waiting on it: no window after
an answer, and no ~3 s wait after a card dismissed by Stop.

* fix(claude): a subagent waits only beside its committed card

A subagent's wait was pushed to the host as soon as its request arrived,
while the request's card row reached the journal at least a microtask later.
So every subagent request published the parent row as waiting before
blocked (the main agent's own fold reads a waiting child that way), and
Activity got an extra unread "waiting" event that main never shows.

The card is now the one record of an open request. The translator records
the asker on the card once (its row's linkage) and counts the card open only
after the sink confirms its rows landed, then publishes the wait; anything
that closes the card (an answer underway, a dismissal handed to the host,
Claude's own withdrawal, the session's end) frees the subagent first. So
every publish that shows a subagent waiting also shows its pending card, and
the parent reads one needs-input state, exactly as on main.

This retires the registry's view of pending requests (unclaimed(), the
asking-child join) and the translator's holdsOpen. The prompt row's linkage
takes one rule: the agent the provider names, else the gated call's owner.

The parent-row proof now runs through the real deferred sink, durable
journal and status feed, publishing as production does, and checks at every
publish that waiting subagents have pending cards and that the parent row
matches a host fed no waits.

* fix(native-chat): a closed sink's dropped writes never read as landed

The sink's written() resolved ok when the sink was closed with writes still
queued, so a subagent's prompt card could count as open with no row in the
journal. written() now reports a close that dropped writes admitted so far
as not landed; drained() and lifecycleBarrier() keep reading a closed sink as
settled.

Tests: a card never opens when its sink closes first; a card Claude withdraws
while the sink holds the cancelled row back closes at once; two subagents
asking at once, and the main agent asking beside a subagent, keep the parent
row as before with each waiting subagent beside its own card; a process that
dies mid-request leaves no subagent waiting.

* fix(claude): a withdrawn subagent request frees its child before its card closes

Main's sink now hands each write to the journal as it is submitted, and an idle journal commits it
and runs the publication at once. Claude's own withdrawal of a subagent's request therefore closed
the card and published the parent row before the child's wait was freed, so one publish showed the
subagent waiting beside no pending card (fg-interrupt replay). The child's wait now also requires the
request to still be open in the registry, and a withdrawal republishes child work before the
journal takes the close.

* refactor(claude): trim subagent request waiting to the common pattern and its essential tests

The chat history no longer marks a subagent block as waiting: the approval card itself carries the
request, and the asking subagent's row in the sidebar and composer strip reads waiting, as before.
NativeChatWaitingSubagentsProvider, native-chat-waiting-subagents.ts and their renderer changes go.

A subagent waits while its request is still open and unanswered in the prompt registry and its card
has landed in the journal. The registry check also covers a withdrawal under backpressure, so the
card list no longer filters pending cancellations itself.

Tests: one integration file replays the captured CLI frames through the real adapter, sink, journal
and status feed (renamed claude-subagent-permission-request.test.ts), with the asking subagent's
state timeline, attribution, nested linkage and a card write that waits for the journal. The
producer-harness waiting test, the redundant prompt-card cases, the harness reducer swap and three
unused captures (deny, interrupt, main agent, failed subagent) are removed.

* test(claude): pin the parent row's dating when a subagent asked first, and narrow the oracle's claim
2026-10-03 12:45:51 -07:00
..