mirror of
https://github.com/stablyai/orca.git
synced 2026-10-03 16:02:11 +00:00
cd8d03bc06398b856f6db734aef11df4e2afe15a
12
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
cdfdadf9ea |
fix(runtime): settle tui-idle on hook state for agents whose hooks cover the whole turn (#24388)
* fix(codex): install Codex's Interrupt hook so an Esc-cancelled turn settles
Codex 0.150+ fires an Interrupt hook when the user presses Esc on an
approval prompt or mid-tool, and nothing else. Orca did not install it, so
the pane stayed blocked/working until the next prompt.
- Add Interrupt to the managed Codex events and label maps, written with
Codex's 3s cap (a larger value triggers a startup clamp warning).
- Hash the timeout Codex hashes (Interrupt is clamped to [1,3], default 1)
so self-computed trust matches Codex; pinned against a real 0.159.3 hash.
- Map a root Interrupt to the existing cancelled-turn record
(markCodexLeadTurnInterrupted), keeping child work in the fold; a
child-scoped Interrupt is ignored. Relayed rows take the same path.
* test(runtime): add a readiness census pinning every tui-idle verdict
Replays every recorded agent PTY transcript frame by frame through a real
runtime pane (agent-known and agent-unknown, clocked and clockless) and a
synthetic evidence matrix for all 43 TuiAgents, and compares each verdict
and tui-idle wait outcome to committed run-length-encoded baselines.
Refs STA-9098
* test(runtime): pin the census quiet probes to literal windows
A census that read TUI_IDLE_QUIESCENCE_MS would move with it; fixed 2999/3000 ms
reads and a fixed 2000 ms poll step make a changed window show as changed verdicts.
Refs STA-9098
* test(runtime): say which census probe writes runtime state
Refs STA-9098
* refactor(codex): let the hook builder own Codex's per-event timeout
The managed hook's timeout is now Codex's own normalization of the shared
budget, and every installer derives its trust entry from the hook it wrote,
so no installer repeats the Interrupt special case.
Claude-Session: codex-interrupt-hook review
* refactor(codex): route Interrupt through the Stop lead update with an outcome
Interrupt now writes the lead record through the same setCodexMainAgentTurnState
call as Stop, so markCodexLeadTurnInterrupted keeps its original signature.
Drops the child-scoped Interrupt guard: Codex never runs Interrupt hooks for
subagents and its input schema has no agent_id.
Claude-Session: codex-interrupt-hook review
* test(runtime): observe the census through settled panes and caller-visible waits
- Read each verdict through the runtime's own settle seam (evaluateTuiIdleForLeaf) instead
of re-wiring evaluateTuiIdle/leafTuiIdleEvidence/buildTerminalWaitText, so the census is
coupled to one runtime method, not to the module STA-9098 rewrites.
- Let the runtime finish each chunk (one macrotask turn) before reading. The old read raced
work chained on the paint, so 14 frames pinned a microtask-ordering artefact.
- Record when a wait settles (@start vs @poll), not just its outcome.
- Exit each pane's PTY after reading it so its emulator is freed.
- Replace the hand-grouped families, literal fixture list and per-pane split flag with a
directory-scanned catalog, one baseline per replayed pane, and size-balanced shards.
- Run the synthetic matrix in one file; it takes about 2 s.
* test(runtime): cross dialog-versus-ready-screen order with every title in the census matrix
Blocked detection is position-ordered (design doc 11.5): the later of a blocker and a ready
anchor wins. The matrix now paints a workspace-trust dialog after, and before, each agent's
ready screen under every title, so a rule engine that loses that ordering fails per agent.
* test(runtime): read the census baseline field without Reflect.get
The anti-slop lint rejects Reflect.get on parsed input.
* refactor(runtime): read Antigravity, Cline, Prime Agent and Cursor readiness from rule files
Adds agent-state-rules/: a zod-validated JSON file per agent, one priority list of
screen rules per agent (idle with strength and requiresQuiet, or hold), and text
anchors that feed the shared, position-ordered blocked layer every pane reads first.
The three screen-ruled agents and Cursor's approval menu and prompt move to data;
the Antigravity text scan stays code as a named anchor. Their old code paths are
deleted. Every other agent still runs through the existing lanes, unchanged.
The readiness census baselines are untouched and pass.
Refs STA-9098
* test(runtime): cover the agent state rule engine's schema, priority, rows, anchors and lanes
Refs STA-9098
* fix(runtime): refuse rule patterns that repeat an optional or alternating group
The load-time regex check only flagged a repeated group whose body held * + or {,
so (a?)* and (a|aa)+ passed though both backtrack exponentially. A repeated
group's body must now be fixed: no quantifier of any kind and no alternation.
The comment states the remaining polynomial gap instead of claiming linearity.
* refactor(runtime): give agent state rules and text anchors one when/answer shape
Every rule and text anchor is now when (a region and what it must show) plus
answer, each a discriminated union, so part (b) adds title, text and status
regions and working or blocked answers as new variants instead of new fields.
- Cursor's prompt is two anchors answering working and idle; the one-off
workingIfAfter and followedBy fields become a general after test.
- Anchor literals and the probe banner must be lowercase, since they are
matched against the lowercased tail.
- screenProbeBanner moves under profile, the place for non-detection facts.
- why is required on every rule and anchor.
- A blocked anchor must name a lastOf literal, which the prefilter keys on.
* docs: point the readiness evidence docs at the agent state rule files
* refactor(runtime): read Codex, Claude, OpenCode, Pi, OMP and Gemini readiness from rule files
The rule engine gains the regions and answers these agents need, as closed-list entries:
- rule regions `title` (the classified title status) and `text` (one of the file's idle text
anchors, settled), and a `predicate` form of the screen region for named engine scans;
- `withoutClock: skip` for strong quiet rules a clockless pane must not believe;
- anchors (renamed from textAnchors) gain a `title` region, and `live` and `hold` answers;
- `profile.screenSource` (trusted grid or live screen), and an `unknown-pane` file for panes
with no known agent.
Codex's header, composer and provisional-startup checks become named predicates referenced
from codex.json; its ready header, header and startup hold become shared text anchors. Native
idle title markers become shared title anchors; name-only title handling becomes each agent's
idle-title rule. The agent-specific branches in terminal-wait-detection.ts and
tui-idle-evidence.ts are deleted, and the "later live prompt cancels a blocker" rule now reads
only rule-file anchors (plus Muse, which moves in part b2).
No behaviour change: the readiness census baselines are untouched and pass.
Refs STA-9098
* test(runtime): cover the rule engine's title, text and predicate regions and the bundled anchors
Refs STA-9098
* fix(runtime): reject a rule file that repeats an anchor or rule id
A text rule names its anchor by id, so a repeated id let a file pass validation and then throw
while compiling. Also states that engineVersion bumps once a version ships; version 1 is still
being defined.
* refactor(runtime): fold the working anchor answer into live
The engine treated an anchor's working and live answers identically: both mark a live prompt
that cancels an earlier blocker and settles nothing. Cursor's busy prompt now answers live, so
anchors have one non-settling prompt answer.
Refs STA-9098
* refactor(runtime): read the shared π title anchor from pi.json alone
Pi and OMP paint the same `π - <session>` rest title, and title anchors apply to every pane,
so one copy covers both.
Refs STA-9098
* refactor(runtime): key every rule file and read the trusted screen from screenSource alone
readsTrustedScreen no longer also asks for a screen rule (every trusted file has one, and the
schema requires screenSource where it matters), so rule-less files need no filter. A rule's
match is a plain boolean, and compileTitleAnchors is module-private.
Refs STA-9098
* test(runtime): pin that a clocked Codex pane takes no other agent's ready text
No test failed when holdsReadyTextToQuiet was removed; this one does.
Refs STA-9098
* fix(agent-hooks): keep an OMP approval wait until omp resolves it
omp posts tool_execution_start a few milliseconds after
tool_approval_requested, while its Approve/Deny select still holds the
human. Both mapped onto the pane row, so the working event overwrote the
blocked one and the pane read as busy for the whole prompt.
A working event now leaves an OMP approval wait in place; only
tool_approval_resolved or a new turn ends it. An ask row is unchanged:
its own tool_execution_end ends it. The test replays the order a live
omp 17 run posted for a denied bash call.
Refs STA-9100
* refactor(runtime): select the fresh hook row on any of a terminal's handles or pane keys
selectFreshExplicitAgentStatus matched one handle and one pane key and
returned only the mapped status. The row selection now takes sets of
handles and pane keys, an optional received-at floor, and returns the
row itself, so a reader can see the main agent's own state. The old
function keeps its signature and result on top of it.
Refs STA-9100
* feat(runtime): let tui-idle read hook state for agents whose hooks cover the whole turn
tui-idle read no hook state. Hook state reached readiness only through
the `<Agent> ready` titles the window writes, so a headless `orca serve`
never saw it (#16095), and Codex settled only once its screen had been
quiet for three seconds.
Rule files gain `profile.hooks: "authoritative" | "identity-only"`,
defaulting to identity-only. Codex (with its Interrupt hook), OpenCode,
OpenCode 2, Pi and OMP are authoritative. For them a fresh hook-store
row decides ahead of every other lane:
- the main agent's turn decides (`mainAgent.state` when published), so a
subagent's Stop does not end the lead turn: done settles strong,
working holds, a permission wait never settles;
- the tail's blocked text goes through the existing permission arbiter
with the turn as its explicit status, so a denied prompt's dialog left
in the tail no longer blocks a turn the hook says ended;
- the row joins on every pane key and terminal handle the PTY owns.
No row, a stale, restored or other agent's row, a session-start done,
and a row from before a PTY respawn all fall back to today's lanes. That
keeps startup on the screen and text rules: Codex posts SessionStart
only with the first prompt. Claude, Cursor, Gemini and the rest stay
identity-only.
The readiness census has no hook server, so its frames are unchanged.
Refs STA-9100
* docs(agent-status): record readiness as a reader of the hook store
Refs STA-9100
* fix(runtime): ignore a hook done older than the latest input Orca wrote
A finished turn leaves a fresh `done` row. A caller that sends the next
prompt and waits at once could settle on it before the new turn's first
hook arrives, so the wait returned while the agent was starting work.
Orca's own input writes (terminal send, agent prompts, mailbox pointers)
now stamp a per-PTY input clock, and the hook lane reads no `done`
received before it; the pane falls back to the screen and text rules
until the agent reports again. A `working` row is unaffected.
Refs STA-9100
* docs(agent-status): note the input floor on the hook lane's done
Refs STA-9100
* fix(runtime): take the hook lane's input floor from the PTY run's input record
The hook lane ignored a done older than Orca's latest write to the pane, kept in a
new per-PTY map stamped by a wrapper threaded through four write sites. The PTY
run register already sits on both write funnels, so it now records the last
input (launch writes included, terminal replies not) and the lane reads it.
Keys the user types now count too, which closes the restart-in-the-same-shell
gap: typing `codex` to relaunch no longer lets the previous process's done read
ready while the new one boots.
The respawn floor moves from the shared row join into the lane, beside the
input floor; the freshest row predates a floor exactly when every row does.
* test(runtime): drop runtime hook-lane cases the unit suite already proves
Working over a ready title, a permission wait, and an identity-only agent are
decided inside evaluateTuiIdle and covered there; the runtime suite keeps the
wiring: the join, both floors, Pi's own OSC 133 markers and the arbiter.
* fix(runtime): record a PTY's last input even when main adopted it without a spawn commit
A materialized pane re-adopted by the renderer returns before the spawn-commit
site, so it had no run record and its input never moved the hook lane's floor.
The last input now lives beside the run records: any PTY's input counts, and a
new process's commit still clears it.
* fix(runtime): keep a running process's input time when main reattaches or adopts it
A reattach or adoption commit without an incarnation id cleared the PTY's
last-input time, so a prompt sent just before an SSH adoption was forgotten
and the hook lane could accept the previous turn's done as ready. Only a new
process (or a reattach naming a different incarnation) now starts clean; the
first-input fact follows the same rule.
* docs(runtime): say why a lead turn that ended reads ready while a subagent runs
* fix(runtime): refuse uppercase contains terms in text anchors, which read the lowercased tail
A text anchor's after and lines tests run on the lowercased tail, so an
uppercase contains term loaded and then never matched. Build the text test
schema from the literal it accepts and give anchors the lowercase one. Also
drop a probe-banner early return that no bundled catalog reaches.
* refactor(runtime): state Codex's provisional startup and title anchors as plain rules
The provisional-startup hold becomes a lastOf anchor with an all/none test, so
its TypeScript scan goes. Title anchors drop their status field (every caller
already gates on an idle title), and withoutClock keeps only the value a rule
can set.
* fix(runtime): leave Codex readiness to its title and screen rules
Codex before its Interrupt hook posts nothing for an Esc mid-turn, so its hook
row stays working and a hook-authoritative tui-idle wait hangs until the row
goes stale. Current Codex already settles fast through its ready title.
Co-Authored-By: Claude <noreply@anthropic.com>
---------
Co-authored-by: Claude <noreply@anthropic.com>
|
||
|
|
0b79720c2e |
feat(native-chat): the chat strip and the sidebar read the host's child records (#22614)
* feat(native-chat): publish the host's child records to the status summary and the chat strip
The status summary and the background-task channel now read a session's child
records from the host's canonical store, through the status sink its row
landed in, and derive the legacy task and subagent shapes from the same views.
The parent row folds its child-work liveness from those records at ingest,
not from the summary's task list. The adapters no longer push their task DTO
to clients: the onBackgroundTasksChanged path is gone, and a child-work ingest
is what republishes both the summary and the strip. Finished children stay
listed until the session's own next turn starts. A reader that predates child
views never receives a roster whose rows are all settled.
* feat(native-chat): the chat strip reads the host's child records with its parent's verdict
The strip's roster now renders from the child views its channel carries, and
passes the verdict the session's own status row gives its children, built the
way the sidebar builds it (the row's freshness and the status feed's
observation). So one child reads the same in the strip and the sidebar, live,
after the transport drops, and once the row goes stale. A roster of finished
children stays shown until the next turn but no longer animates the monitoring
indicator or blocks conversation commands.
Tests: an end-to-end run on a host with no renderer (a real hook server as the
status sink) shows the summary and the strip channel carrying the same records
at every step, the parent row folded from them, retention, and an older
client's task list holding live work only; a wired renderer test shows both
surfaces agree when live, lost and stale.
* test(native-chat): a newer host's view degrades, an older reader keeps its live roster, a finished roster holds nothing open
- The view decoder ignores unknown keys, degrades unknown kinds, states,
outcomes and memberships, and drops only rows it cannot identify.
- At the RPC boundary a reader that predates child views gets no strip for a
roster of finished children and never the views themselves; a stop-only
reader keeps rows whose host offers no targeted stop.
- The strip shows finished children without reading them as live work.
- The row keeps its child list's identity when a summary repeats it.
* refactor(native-chat): the summary's task list is the legacy projection's live rows, unfiltered
* test(native-chat): type the switch tests' mocks instead of asserting them
* test: remote clients advertise reading child views
* docs(agent-status): the structured row folds the store's child records
* fix(agent-status): keep the view reader's header from reading as a value import to the renderer boundary
The renderer node-builtin boundary test scans raw text, so a header comment
that said "imports" ahead of the import block turned the type-only import of
agent-status-child-work into a value edge that reaches node:crypto.
* refactor(native-chat): the status summary's broadcast equality gets its own module
The status feed crossed the file-size limit once the summary gained the main agent's turn
outcome beside the child views. Which summary changes reach every session list now lives in
structured-agent-session-status-summary-equality.ts.
* fix(native-chat): command admission reads the strip's child records
A conversation command was refused on the provider tracker's own roster
while the strip read the host's child records, so a drift between the two
rule sets could refuse /clear with a stop instruction the strip had no
button for. Admission now reads the same records through the same read as
the strip, uses the strip's liveness fold, and asks for a stop only when
the strip renders one. The adapter contract no longer exposes the tracker
roster, so no host decision can read it.
Also records when the legacy child shapes die, every earlier death of a
settled child, and the display-precision invariant behind the summary's
clock tolerance.
* refactor(native-chat): command admission takes only what it reads of a turn
* fix(native-chat): the session list drops a session's children when the store does
A session's end no longer removes its child records: a child still running
settles with an outcome nobody reported, and a finished one stays listed.
Records now leave only at the session's own next turn, at the cap on settled
records, or when the host lets go of the session and its row leaves the store.
The summary kept after the host lets go used to strip its children on close,
a rule of its own. It now re-reads them from the store when the row leaves,
through the same read every live summary uses, so the session list and the
chat strip list the same children at each step, including a forget with no
close. Closing only revokes ownership, as before the child records existed.
* test(native-chat): write the Codex frame script's parent row out step by step
Once every surface reads the child records, the provider tracker's roster is
no oracle: it and the records read the same child executions, so agreeing
with it cannot catch a defect in either. Each frame now states the child
liveness and the parent row it must fold to.
* fix(native-chat): the idle sweep and the restart snapshot read the host's child records
The idle sweep (keep an agent running while its subagents or commands run) and
the restart-resume snapshot (what a chat was doing when Orca stopped it) both
read the provider tracker's roster through the adapter interface, which no
longer carries it. Both now take the host's one child-record read, the same one
the status summary, the chat strip and command admission use.
The snapshot's working test also folded that roster through the shared fold's
old `backgroundTasks` input, which the fold no longer reads, so a settled lead
whose subagent was still running would have been offered nothing. It now hands
the fold the records.
* test(native-chat): the child-record tests follow the merged command lifecycle
A command is a live child record from its start and is removed, not settled,
when it stops, whatever Codex tagged it. The end-to-end switch now shows the
child's `npm test` as a live row beside its dev server, and both are gone once
they exit; only the finished subagent stays listed until the next turn. Letting
go of the session is its tab closing, since a closed conversation whose tab
remains keeps its row.
Command admission's finished row is a subagent, the one kind that settles, and
the failed-verdict row test admits its live subagent as a host record, the only
thing the row folds.
* refactor(native-chat): the status feed's journal projection cache gets its own module
The status feed crossed the file-size limit once the child records joined the
agent-start signal and the completion feed's status read. The per-journal
projection, cached per commit, now lives in
structured-agent-session-status-journal-projection.ts.
* test(native-chat): the admission test's compaction resolves with a real outcome
Main's compaction result is a tagged outcome; the host-level admission test
resolved its mock compaction with an empty object.
* feat(native-chat): the sidebar lists running subagents; the strip, running then the newest finished
The host keeps every child record; what each surface lists is picked from them on every read, so
nothing is stored twice. The status summary, which every session list reads, now carries only
running children (and a finished one whose shell still runs, which reads monitoring): a finished
or failed subagent leaves the sidebar and stays in the chat's strip. The strip lists every running
child, then the newest finished ones, 100 rows in all; more than 100 running all show.
This matches common practice: sidebars show live subagents, and finished ones stay in the chat's
panel, newest first. No wire field is added. An older client reads fewer rows: its legacy task
lists were already live-only in the summary, and the strip's settled tasks come from the same
bounded roster.
* fix(sidebar): one rule for what the worktree sidebar lists: running children, from every source
`worktreeSidebarListsChild` is the one definition: a child that runs, counting a finished one
whose own shell still runs (it reads monitoring). The sidebar's row builder applies it to every
child source it reads, a terminal agent's hook roster and a chat session's records alike, and the
host's status summary applies the same predicate, so the sidebar's payload stays small. The chat's
strip keeps finished children, newest first.
A terminal agent's hook roster already drops a child on its own stop, so nothing changes there:
a teammate between turns and a child gone quiet still run, and still show. The selection module
moves to `agent-child-work-listing.ts`, since it now covers every source, not only chat sessions.
* test(native-chat): the switch test passes the startup child key main's status bar takes
* fix(native-chat): a finished child stays until the user's next send, not a turn Claude opens on its own
Claude wakes the agent on its own when a background task ends, and that wake is a
new root turn. Keying retention on the newest root turn retired every finished
child about two seconds after a background agent or shell finished, so its
outcome never showed in the strip.
Retention now keys on the user's newest send the provider accepted (a message, a
steer or a command), with the journal epoch so a rewind still retires. A turn the
provider opens itself and a subagent's turn carry no send. Replayed captured wake
orders through the real adapter, hook server and status feed.
* fix(native-chat): a background Stop reaches the tasks the child records show
The strip draws a row's Stop, and /clear, /compact and rewind wait for background
work, from the host's child records, but the Claude adapter still resolved which
tasks a Stop reached from its own tracker's roster, and refused to stop at all
once that roster was empty. A task the records kept live after the roster dropped
it showed a Stop that sent nothing and blocked those commands until the chat tab
closed.
The host now resolves the provider ids a Stop sends from the records (the same
per-row rule the strip and admission use; every such row for stop-all), and the
adapter stops exactly those, with no tracker guard. An acknowledged stop ends the
record: a running task sends its own stopped frame first, and the CLI answers
success with no frame for a task it no longer knows. A refused stop leaves the
record live. No production code reads the tracker's roster any more.
* fix(native-chat): one rule for a finished child that still owns live work, at any depth
The listing kept a finished child whose work ran through any depth of ownership,
but retention at the user's next turn protected only the direct owner, so a
finished agent whose finished subagent still ran a shell was removed and that
subagent jumped to the top level. Both now read settledOwnersOfLiveWork.
* fix(native-chat): an older client sees a Codex child's shell as it did before views
Clients that predate child views read a flat task roster derived from the views.
It listed a Codex child agent's shell as an extra row beside the running agent,
then as a bare command once the agent finished. The derivation now hides a
running agent's commands and names a finished agent's as "<agent> — <command>",
as the Codex tracker did; the label rule moves to a shared module both use.
* fix(native-chat): the chat decodes a roster's child rows once, as the frame arrives
The client reducer compared raw wire rows, so a row shaped by a newer host could
throw there, and the strip decoded a new array on every render, which defeated its
grouped-rows memo while a turn streamed. Rows are now decoded where the frame
enters the reducer, an unchanged roster keeps its identity, and the strip's parent
context is rebuilt only when one of its values changes.
* fix(native-chat): the strip channel forgets a closed conversation's roster
It kept the last roster fingerprint of every conversation for the host's lifetime.
The conversation map now tells observers when one leaves it. Also corrects the
summary's children comment: it carries running children only.
* docs(native-chat): rewrap the retention comment
* fix(native-chat): a task's own ending replaces a Stop's, and a child finished after the user wrote stays
Two lifecycle gaps from the round-1 fixes.
A Stop acknowledged ahead of the task's own ending relabelled it. The SDK hands
Orca a control answer as soon as it reads it and queues other frames, so when a
task finished just as the user pressed Stop, the acknowledgement arrived before
the task's completion the CLI wrote first, and the task read "Stopped" with its
result lost. An acknowledged Stop now ends a record provisionally (outcome basis
`stop-acknowledged`); the task's own terminal frame replaces it, and nothing
replaces an ending the task reported itself.
A finished child still vanished with no new action from the user when the send
the provider took was written before the child finished: a steer Claude takes at
its next boundary, or a queued draft handed over at the end of the turn. The
user's next turn now carries when they acted (the send's written time, or the
draft's queued time, both on the host clock), and only children that finished at
or before it retire; a later one stays until the user's next send. A rewind
still retires every finished child.
Also: a Stop-all keeps stopping the remaining tasks after one request fails, then
reports the failure.
* fix(native-chat): the strip keeps one empty list for a roster that omits one
A roster with only running or only finished rows made a new empty array on every
render, so the strip regrouped its rows each time. One shared empty list keeps
its memo.
* fix(native-chat): a strip row whose owner the 100-row budget cut renders under the main agent
The budget can keep a finished child and cut its finished owner. The child still
named that owner, so it rendered nowhere. The selection now clears an owner it
did not keep, as the view contract says for an owner outside the projection.
* fix(native-chat): a stop-all that times out stops asking, and "no children" is sent once
A stop-all kept asking after a request timed out, so a Claude CLI that stopped
answering control requests cost one full deadline per task, while the chat's
sends, Stop and /compact waited behind it. A timeout now ends the loop; other
request failures still let the remaining tasks be stopped.
The chat strip channel never remembered that it had sent "no children", so
every change in a chat with none re-sent that frame to each subscriber. It now
remembers it, and forgets only when the conversation closes.
Also renames agent-child-work-stop.ts to agent-child-work-stop-targets.ts, which
says what it answers: the provider ids a background Stop reaches.
* fix(native-chat): the status feed reads a journal snapshot with no submissions, and e2e tests use the child-work reader
CI on
|
||
|
|
24edf0f64b |
fix(agent-status): a turn a crash cut off reads Interrupted, an unproven end Couldn't confirm (#23467)
* refactor(native-chat): remove the unused terminal handoff
No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.
* fix(native-chat): never let the pre-stop snapshot hold a chat's stop
Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.
Co-Authored-By: Claude <noreply@anthropic.com>
* refactor(native-chat): drop helpers only the terminal handoff called
`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.
Co-Authored-By: Claude <noreply@anthropic.com>
* docs(native-chat): stop citing the removed handoff in lifecycle comments
Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.
Co-Authored-By: Claude <noreply@anthropic.com>
* test(native-chat): type the stalled snapshot drain without a cast
Co-Authored-By: Claude <noreply@anthropic.com>
* test(native-chat): pin that a start dead before proving owes no settlement
The removed restart handoff test pinned this branch; nothing else did.
Co-Authored-By: Claude <noreply@anthropic.com>
* fix(native-chat): keep the owner-status read behind an in-flight attach
The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.
Co-Authored-By: Claude <noreply@anthropic.com>
* refactor(terminal): remove the agent-session PTY write gate
The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.
Co-Authored-By: Claude <noreply@anthropic.com>
* refactor(native-chat): drop the transcript helpers only the handoff called
appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.
Co-Authored-By: Claude <noreply@anthropic.com>
* fix(native-chat): stop calling a starting chat "mid-handoff"
A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.
Co-Authored-By: Claude <noreply@anthropic.com>
* test(native-chat): type the stand-in roster decoder without a cast
Co-Authored-By: Claude <noreply@anthropic.com>
* refactor(codex): name the pinned rollout lookup for what it does
With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.
* refactor(native-chat): type the owner-status reply as the host sends it
The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.
* refactor(native-chat): normalize terminal-handoff lease values once at decode
Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.
The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:
- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
`conflicted`, the claim every build probes but never stops. A plain native
owner would be stopped by restart recovery, here and in older builds.
Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.
The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.
* refactor(native-chat): stop threading the owner kind through a reservation
A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.
* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else
Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.
* fix(native-chat): name a chat write by its target, not the owner generation
A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.
Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.
Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.
* fix(native-chat): every journal append reaches the chats that are open
A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.
A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.
* test(native-chat): an epoch replacement reaches the open chat
* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map
* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite
The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.
* test(worktree-activation): restore the OMP surfaced-agent resume test
The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.
* perf(native-chat): a publish behind a delivered commit reads nothing
Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.
* test(native-chat): state why the teardown test's fake journal is safe to cast
* docs(native-chat): say mutation admission checks only the writer lease
* docs(native-chat): drop the send rebase from comments that still described it
* fix(native-chat): a message is accepted, then delivered
A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".
A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.
Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.
A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.
Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.
* fix(native-chat): settle queued messages only for the child that ended
A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.
A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.
The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.
* fix(native-chat): an adoption that fails to import keeps the conversation open
The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.
* perf(native-chat): the recovering open reads the journal once
Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.
* fix(native-chat): an attach that fails after indexing its child leaves no child behind
A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.
* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer
The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.
A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.
* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down
The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.
* fix(native-chat): a message rejected while its chat was closed reads as not sent
A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.
* test(orchestration): name why the readiness settlement fakes are cast
* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent
* docs(native-chat): drop the fence from the admission the send effects run behind
* docs(native-chat): give the fence move on release the reason that still holds
* docs(native-chat): stop citing a write fence check in launch and mailbox comments
Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.
* refactor(native-chat): the provider child is its own record
A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.
- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.
* fix(native-chat): the delivery loop alone settles a message its start or child failed
A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.
- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
reads how it ended: a Stop continues; anything else writes one failure row and rejects every
queued message with the same words, then stops. A child still starting whose start the adapter
says did not land fails the same way. The exit, eviction and the settlement retry only settle
the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
closed, with or without a child, and a start the loop already has in flight is waited for so the
child it produces is stopped rather than left behind.
* refactor(native-chat): a stopped child ends on the one reading of its stop
The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.
* feat(native-chat): the host says it accepts a send before any agent has it
The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.
* refactor(native-chat): an attach never opens a journal of its own
The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.
* fix(native-chat): a moved fence resends nothing on a host that accepts first
The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.
The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.
* refactor(native-chat): a child's end says whether the user or the host stopped it
The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.
* fix(native-chat): a chat whose only work is a queued message is not offered for resume
A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.
* test(native-chat): type the queued-message fixtures in the resume-offer tests
* fix(native-chat): a start that dies while a message waits on it is that message's failed start
Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.
* fix(native-chat): a request that failed reads as failed
A structured chat whose only message the agent's start refused read as a
green finish, and a cancelled structured turn did too: the host published a
verdict only for turn records, and structured rows carried no `interrupted`.
The host projection now reads the session's latest request: its turn's
outcome, or `failure` for a send the agent or its start refused. A send
that was withdrawn, or left undelivered by a restart or a close, fails
nobody and makes nothing listable. The ingest publishes `interrupted` as the
hook lanes do, and every reader decodes the verdict through one accessor, so
a failure reads Failed on the dot, the rollups, history and `worktree ps`,
behaves like a cancellation in every clean-finish policy, and notifies as
"failed".
* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now
* test(native-chat): a verdict change republishes the mobile status projection
* refactor(native-chat): the store's retention trigger keeps its flag compare
A verdict change always moves the completion clock the same check already
reads, so a second verdict compare there caught nothing new.
* test(native-chat): a user message the provider journaled keeps its session listed
* test(native-chat): pin what a failed start settles, and what a resume offer names
A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.
* test(native-chat): the failed-start pins fail on what the message became, not on a timeout
* fix(native-chat): a late provider-session update keeps a failed recovery record failed
A provider-session heartbeat that rewrites a completed recovery record kept
its interrupted flag but dropped the outcome it was copied with, so a live
failed checkpoint read as a clean finish until the next status write.
* test(orchestration): the preamble's host stub is typed, not cast
The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.
* test(native-chat): the terminal-bell check asserts the renamed verdict field
The bell notification test still checked for agentInterrupted, which no
longer exists, so it could not catch a verdict leaking into a bell dispatch.
* fix(native-chat): a failed turn ranks like a completion for attention
Attention readers (completion time, Smart Sort, sticky retention, Cmd+J
Recent) now demote only a turn the user stopped. A failure is news the
user has not seen, so it keeps its completion time, ranks in the Done
class, stays retained after its pane goes away, and a retained failure
reads failed in the worktree rollup instead of done. Clean-finish
policy (hibernation, pane ownership, the value moment) still treats a
failure like a stop.
The retention trigger compares verdicts again: success -> failure no
longer moves the completion clock.
* fix(native-chat): a failed main agent reads failed while its subagents still work
The verdict is now read from the main agent's own state, not the folded
row: a main agent that is done and failed has a verdict even while its
subagents keep the row working. Without mainAgent (history, worktree ps,
older hosts) the old combined-done rule stands.
Display marks the verdict through agentVerdictDisplayMark: a failure
outranks every combined state on the agent's dot, label, tab badge,
dashboard and activity rows; a stop marks only a done row, so a
successful or stopped main agent with live subagents still reads
working. Subagent rows keep their own state. The worktree card, terminal
tab and Cmd+J rollups share one pane fold and rank a pending question,
then failed, then working, monitoring, interrupted and done.
worktree ps publishes the main agent's outcome on a working row, and the
mobile mirror reads it. The store's change check, the paired-client
mirror's equality and its epoch now see a verdict change on a working
row, which otherwise moves no state or clock and left the worktree card
reading working. Clean-finish policy is unchanged: a working row is never
hibernated and has no completion time.
* docs(native-chat): the worktree ps outcome comment no longer claims old hosts send it
The field is new: an old host sends no outcome at all, so a reader falls
back to interrupted. The removed clause said old hosts send it on done
rows, which never shipped.
* docs(native-chat): the status-store listing rule names provider-journaled user messages
* fix(native-chat): a refused send notifies failed through the completion feed
The host's completion feed followed only the newest turn, so a send the
agent or its start refused, which creates no turn, read Failed on its row
but sent no notification. The feed now follows the session's latest
request, read from the projection the status feed already makes for the
commit: a turn keeps its id, a refused send is named by its journal item
key. It announces only while the session is idle, as the row reports a
verdict, so queued sends refused one commit at a time notify once, and a
withdrawn send falls back to a request already announced.
* fix(native-chat): every copy of a row carries the main agent's own status
History entries, sleep records and `worktree ps` rows carried a flattened
top-level `outcome`, copied under different gates and without the main agent's
clock. They now carry `mainAgent` (state, outcome, stateStartedAt), the type
the live row already persists and sends, and every copy site takes it with
`interrupted` through one function, `agentVerdictFields`.
- The accessor reads `mainAgent` then the legacy flag; the mobile mirror
matches it line for line.
- Sleep records admit `mainAgent` with `normalizeMainAgentStatusField`, so a
malformed value drops the field, never the record.
- Mobile dates a main agent that failed under live subagents by its own clock,
as desktop does, and its row equality compares `mainAgent`.
- The activity feed reads a history entry's own `mainAgent` instead of
rebuilding one; the sync key and history equality compare it.
* test(native-chat): pin the worktree ps verdict across host and phone versions
Pairs the real v1.4.212 host and phone row reader with this build: an old phone
reads a new host's rows by `interrupted`, a new phone reads an old host's rows
(no `mainAgent`) the same way, and a new phone reads a failure under live
subagents as Failed, dated by `mainAgent.stateStartedAt`. The release checkout
now carries the phone's self-contained row reader, and the lane runs when the
`worktree ps` row producers change.
* test(mobile): name the parity table's row for its role
* fix(native-chat): a request that settles while the user is asked something notifies once
The completion edge waited for an idle session, and a pending prompt (including a
subagent's approval) is not idle. Structured chat has no other attention producer,
so a main turn that finished while a subagent waited on the user sent nothing
until the prompt was answered.
The edge now waits only on owed work (a running turn or an unanswered send), which
the projection reports even beneath a pending prompt. A request that settles with
a prompt pending announces once; the renderer words it "needs input" from the
host status mirror's `attention`, and answering the prompt keeps the same request
identity, so it does not announce again. The wire shape is unchanged.
* fix(native-chat): the completion says when the user is being asked
A request that settles while a prompt waits on the user was worded "needs input"
from the renderer's status-feed mirror. Remote clients receive the status and
completion streams over separate sockets, so they can arrive in either order and
the wording could be wrong both ways.
The host already knows at emit time, so the completion now carries an optional
`awaitingUser: true` in that case and omits it otherwise. The renderer words the
notification from that field alone and no longer reads the status mirror. Old
clients ignore the field and word by outcome; old hosts never send it.
* fix(worktree-status): a departed agent's failure yields to live work on the worktree card
A retained failed agent has no expiry, so ranking it with a live failure pinned the card to Failed over other panes' live work. It now ranks below working, monitoring and permission, and above every finished outcome.
* docs(agent-status): a departed agent's failure ranks below live work on the worktree card
* fix(native-chat): a view never restarts a chat whose last start failed
A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.
* test(native-chat): start the child the loop waits on with an attach, not a second view
A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.
* fix(native-chat): settle a gone generation's turn wherever a conversation opens
A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.
* test(native-chat): prove the next child's start settles the turn an earlier child left
The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.
* test(native-chat): count a failed start's rows by row, not by text
Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.
* test(cross-version): load the phone row readers without mobile's toolchain
Vite transforms a file against its nearest tsconfig, and mobile/tsconfig.json
extends expo/tsconfig.base.json, which the root-only cross-version lane never
installs. The worktree ps verdict suite imported the current phone row reader
from mobile/ directly, so CI failed with TSConfckParseError before any test ran.
The harness now imports a copy of the working-tree reader placed under the
checkout cache, where the root tsconfig applies, as it already does for the
release checkout's copy. Both readers are still the real files.
* test(cross-version): keep the checkout path-guard message and justify the copy import's cast
* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget
* test(native-chat): pin the open's and the send's start and row counts, however the view binds
Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.
* fix(agent-status): a turn a crash cut off reads Interrupted, an unproven end Couldn't confirm
When the provider gave no verdict, the structured status projection now derives
one from the newest turn's lifecycle: interrupted -> interruption, unverifiable ->
unconfirmed. Nothing new is journaled, the completion feed stays provider-only, and
the legacy interrupted flag stays a user stop only. Every verdict reader handles
both arms explicitly.
* fix(native-chat): settle a gone generation's turn at every open but an acquisition's
The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.
* fix(native-chat): a folded turn a crash cut off reads Interrupted after N
The settled-turn timing now carries the turn's verdict, derived by the same
agentTurnVerdict the status row uses. A turn that ended interrupted with no
provider verdict heads its fold 'Interrupted after N'; a user's stop keeps
'Worked for N'.
* test(native-chat): hold the create's start open until the views bind
The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.
* fix(native-chat): the user's close of a chat records the turn it cuts short as their cancellation
The expected-close settle writes outcome cancellation when the user aimed the
stop at this chat: a Stop while the agent starts, the chat's tab closed (the
agentSession.close RPC, or session.tabs.close with a user reason), or /clear.
A quit, an idle eviction, a worktree teardown or an orchestration stop leaves the
turn with no verdict, so it still reads Interrupted.
* test(native-chat): a Claude turn a newer send superseded reads Interrupted
The supersede fires for any send Orca dispatched, the user's or another agent's,
and nothing at that site records the sender, so the turn keeps no verdict and
folds as Interrupted after N.
* fix(native-chat): the user's close records cancellation on the turn the provider settled on its way out
The Codex and Claude adapters settle their open turn as interrupted, with no
verdict, while the host stops them. The expected-close settle then found no
running turn, so a user's close of a mid-turn chat read Interrupted. The stop
now reads the running turns before it reaches the provider and records the
user's cancellation on each one it cut short, unless the provider gave a
verdict of its own. The close-verdict test's adapter now settles its turn on
close the way the real adapters do.
* test(native-chat): update the close and settled-turn expectations for the host-observed verdict
agentSession.close now passes the user's word to the host, and a settled
interrupted turn with no provider verdict carries `interruption`. Also merge a
duplicate import the code-quality gate rejects.
* refactor(native-chat): drop the composer's second error formatter
After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.
* test(native-chat): pin the reason on a message rejected while its chat was closed
The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.
* fix(native-chat): the user's close cancels a turn whose start landed as the provider stopped
The close read which turns were running before the stop. A turn whose start was
still in flight (a send echo not yet journaled) was absent from that read, so the
provider's verdict-less settle on the way out left it Interrupted. The close now
reads which turns were already over instead, and records the user's cancellation
on every other turn the stop left running or interrupted with no verdict. A turn
cut off earlier, or finished during the stop, keeps its end.
* fix(native-chat): a send the provider never received after a restart has no verdict
Restart reconciliation rejects a crash-stranded send that is absent from a
trustworthy provider history with reason 'not_delivered'. Nobody failed that
send, but the verdict allowlist did not name it, so after a crash the chat
read Failed, was listed, and could notify "failed". Give the reason a shared
constant (persisted value unchanged), add it to the no-verdict set, and treat
it as an internal marker so the Retry row no longer shows the raw string.
* refactor(native-chat): the adapter settles the turn a stop cuts with the stop's typed cause
The host hands its stop's cause to the adapter's close. Each adapter settles its own
open turn on 'ended' through one mapping, turnVerdictForChildEnd, and the host's
dead-generation fallback uses the same mapping for any turn no adapter settled. A
user's close or stop of this chat is their cancellation; a quit, eviction, teardown
or an exit the adapter saw first is news.
Deletes the snapshot-and-diff reconstruction (endedTurnItemIds,
userStoppedTurnRevisions, settleUserStoppedTurns) and the requestedByUser flag.
host.close now takes a required cause.
* test(native-chat): a user's close drops the chat's status row like an eviction
* test(native-chat): expect the eviction cause on the host closes of idle release, worker stop and worker discard
The stop's typed cause now travels into host.close and the adapter's close, so these
three non-user closes assert the 'evict' they pass.
* refactor(native-chat): every stop names its cause, so none defaults to the user's cancellation
stopStructuredAgentSessionAgentUnderSerialize defaulted its ending to 'user-stop', which now
settles the cut turn as the user's cancellation. Every caller already passes a cause; the
parameter is now required, and a type-level test fails to compile if the default returns.
* test(native-chat): pin who a chat's session.tabs.close speaks for, older clients' reasonless close included
The mapping lived inline in a type-unchecked file, and only the explicit user reason had a test: an older client's reasonless close, or a lifecycle echo read as the user's, stayed green. It is now one exhaustive, type-checked function with a case per reason.
* style(mobile): draw the unconfirmed dot in the theme's status amber, not an inline hex
* docs(agent-status): the main agent's outcome also carries the host-observed end, interruption or unconfirmed
* fix(native-chat): a chat the user closed while its agent started is not a failed start
A still-starting child the user's close cut counted as a failed start, since only 'user-stop' was excluded: a start-failure row, and queued messages rejected as a provider failure. Whether an ending fails its start is now one exhaustive switch, shared by the failed-start read and the delivery loop's handover: a user's stop or close never does; an exit, a failed attach and the host's own stops still do.
* fix(native-chat): a chat the user closed closes its queued messages, and starts no agent for them
After
|
||
|
|
85067494a1 |
fix(native-chat): a request that failed reads as failed (#22944)
* refactor(native-chat): remove the unused terminal handoff No client ever called agentSession.requestHandoff or mounted the handoff chrome. Delete the handoff coordinator, the terminal-owner runtime, the proof write path and the unmounted UI. Keep agentSession.handoffStatus, which released desktop clients read for worktree activation, and let records an older build left mid handoff reconcile through the ordinary restart and recovery paths. * fix(native-chat): never let the pre-stop snapshot hold a chat's stop Eviction now drains delivered events before quit's resume-offer snapshot. An unbounded wait there sits ahead of the provider stop, so a sink whose journal write stalls kept the child running until the step deadline aborted the eviction. The offer is advisory: bound the drain and stop the child regardless. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(native-chat): drop helpers only the terminal handoff called `claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and `queryWindowsProcessRowsFresh` lost their last caller with the handoff. The fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`, the teardown path that still depends on that contract. Co-Authored-By: Claude <noreply@anthropic.com> * docs(native-chat): stop citing the removed handoff in lifecycle comments Six comments still named the handoff coordinator, a handoff suspend, or a terminal-owned session as live participants in the flows they describe. Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): type the stalled snapshot drain without a cast Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): pin that a start dead before proving owes no settlement The removed restart handoff test pinned this branch; nothing else did. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): keep the owner-status read behind an in-flight attach The handoff removal dropped the per-session queue from `handoffStatus`, so a read landing mid-start reported the reservation (no owner) instead of the settled chat owner, and shipped desktop clients blocked worktree activation on it. The read is queued again, as it was before the removal. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(terminal): remove the agent-session PTY write gate The gate only refused a write when a PTY had been bound to a chat session, and the only code that ever bound one was the terminal handoff this branch removes. With it gone, every admit/readmit returned "admitted" unconditionally, so the checks on the renderer write path, the runtime controller backstop, terminal.send, agent prompts, preview input and orchestration pointers, the refusal fields on terminal.send and worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane orchestration routing could no longer run. Ordinary writes take the same path in the same order as before. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(native-chat): drop the transcript helpers only the handoff called appendLegacyTranscriptMessages fed the terminal transcript catch-up and proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost their last caller with the handoff. Their tests now go through the live entry points instead: the roster bounds through the legacy import, the pinned-read and growth tests through the ancestry replay the history window uses, and the marker rules through the string proof in their own file rather than the session-file resolver's. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): stop calling a starting chat "mid-handoff" A send refused because the chat's owner is not settled showed "The session is mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that reach it are a chat that is still starting, or one whose previous agent process has not yet been confirmed stopped. The message now says which of the two it is. The refusal code is unchanged. Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): type the stand-in roster decoder without a cast Co-Authored-By: Claude <noreply@anthropic.com> * refactor(codex): name the pinned rollout lookup for what it does With the terminal handoff gone, the module named codex-tui-rollout-proof holds only the pinned rollout lookup that structured Codex launches use to resume a thread, so the name described code that no longer exists. Rename the module and its options type. Also drop a mobile allowlist assertion that pinned the removed agentSession.requestHandoff method, which no longer exists to allow. * refactor(native-chat): type the owner-status reply as the host sends it The handoffStatus reply type still listed the terminal handoff's fields and states (terminal placement, host label, proof retry, queued and waiting phases, the to-terminal direction). No host writes them any more and the only client reader parses the reply as unknown, so they described nothing. The reply on the wire is unchanged. * refactor(native-chat): normalize terminal-handoff lease values once at decode Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types still admitted them, so readers across the host kept branches for values no path produces and the compiler could not point at them. The store now validates the on-disk shape, which still accepts those values so an older record is not quarantined, and maps them once while parsing: - `preparing` and `old-owner-stopped` become `recovering` - a `tui` lease becomes `native`; when it records a process it also becomes `conflicted`, the claim every build probes but never stops. A plain native owner would be stopped by restart recovery, here and in older builds. Revisions are taken over the normalized state on both sides of every compare, and the mapped record reaches disk with the store's first transaction, the same way the tab-id backfill does. The in-memory types narrow to what this build writes, and the branches that existed only for the removed values go. Structured-worker identity keeps its verdict for a former terminal owner by refusing a conflicted claim rather than a non-native kind. * refactor(native-chat): stop threading the owner kind through a reservation A reservation only ever names a native owner now, so the request no longer carries a kind and the reserved lease records `native` directly. The attach params keep `runtimeKind`: agentSession.ensure and create accept it, and the operation fingerprint stored in the ledger covers it. * test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else Hiding a tab also committed the visibility index, so the no-op transaction wrote the file even when its open-time revision was wrong. Committing the index first leaves the pending rewrite as the only reason to write. * fix(native-chat): name a chat write by its target, not the owner generation A write carried the fence of the last frame the pane read, and the host refused it unless that fence was still current. An idle release and the restart after it each move the fence, and the release publishes nothing, so a send after a release was refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a cold start was refused as stale. Every write already names what it acts on: a send its conversation, a cancel its turn, a prompt answer its item revision, a rewind its epoch; an option is last-writer-wins. So admission stops comparing the client's fence, and the rebase that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it. The writer-lease check stays, and so does the attach's compare-and-swap. Frames now stamp the fence read when each frame is sent instead of a copy each subscriber kept, which went stale on the same release. * fix(native-chat): every journal append reaches the chats that are open A journal write and its delivery to open readers were two calls, and some writers made only the first. A failed start whose lease could not be handed back, a provider revision with no frame behind it, and eviction's settlement were all journaled without reaching an open chat. A journal handle now reports every durable change, and the host's session map binds that report to the session's readers when the handle is set. Writers no longer publish what they append; the per-writer publish calls are deleted. * test(native-chat): an epoch replacement reaches the open chat * test(native-chat): each row reaches an open chat once, and a live handle enters only through the map * test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite The seeded record had no surface tab id, so the next open backfilled one and that rewrite alone made the no-op transaction write. The test passed with the legacy-lease rewrite signal removed. * test(worktree-activation): restore the OMP surfaced-agent resume test The handoff removal deleted it alongside the terminal-owner tests, but it covers the surfaced-PTY block that still guards resume, including an agent whose ownership is unknown. * perf(native-chat): a publish behind a delivered commit reads nothing Each commit now delivers itself, so the publish a provider frame still sends afterwards found every reader caught up but still read rows and rebuilt the timeline for each one. A caught-up reader now skips the read. * test(native-chat): state why the teardown test's fake journal is safe to cast * docs(native-chat): say mutation admission checks only the writer lease * docs(native-chat): drop the send rebase from comments that still described it * fix(native-chat): a message is accepted, then delivered A send to a chat with no running agent restarted the agent inside the send call, before the message was recorded, so the client waited for the whole start and a failed restart refused the message. Claude held prompts sent during startup, and those could settle as "unconfirmed". A send is now accepted inside the session's serialized queue: one ledger row and one submission row marked handoverRecorded, published, answered pending. A per-session delivery loop exists while a message is queued. It starts the agent through the same serialized attach a hold uses, waits outside the queue for a Claude child to prove its start, and hands the oldest queued message over as its own serialized step, writing dispatch{pending} before the adapter call. A start it needed and did not get writes one error-tone row and rejects every queued message with the same words; a start Stop cancelled writes none. Settlement follows from the rows. A queued message is provably unwritten, so a close, an eviction or an exit rejects it. A handed-over message stays in doubt. A queued row at or below the sequence a handle found when it opened was left by an earlier process and is rejected at open, with no latch. Stop withdraws queued messages with no writer lease and no fence. An attach failure keeps the conversation open, and the attach adopts its journal. Owed work counts the loop and queued rows. A compaction or rewind found prepared when a conversation opens was started under a child this process no longer has, so the open settles it rather than leaving it to refuse every send until a view attaches. The open cursor is scoped to its epoch, because sequences restart when an epoch is replaced. Deleted: restart-before-admission, recordFailedRestart, the fence rebase, Claude's startup gate, the attach's forget on failure and its own crash boundary. Clients without agent-session.accepted-send.v1 get their reply held until the handover; the desktop and paired desktop lists advertise it. * fix(native-chat): settle queued messages only for the child that ended A child that proved its start and then exited before its message was handed over left the message queued: the exit settlement returned early when nothing else was in flight. Delivery then started another child for it, and a child that died the same way started another, without end and without a row. A retried settlement for an earlier generation, run by the attach that delivery started, did the opposite: with that generation's turn unfinished it rejected the message queued for the child being attached. The settlement now takes the rejection for queued messages from its caller. The unexpected exit and the eviction pass one, and it applies even with no other work in flight; the retry for an earlier generation passes none. * fix(native-chat): an adoption that fails to import keeps the conversation open The attach now writes into the conversation's own open journal, but a failed transcript import still closed it as if it were the attach's provisional one. The conversation stayed indexed with a closed journal, so every later send answered "could not be recorded" and every attach failed again until the app restarted. The import now closes only a journal the attach opened for itself. * perf(native-chat): the recovering open reads the journal once Every conversation open now goes through the recovering open, including the read restore of every chat at startup, which used to replay its journal once. The recovering open replayed it twice: once to probe it and again inside the open. The probe is now handed to the open as its load. * fix(native-chat): an attach that fails after indexing its child leaves no child behind A failed attach now keeps the conversation open, but a failure after `onAttached` indexed the child (the rewind or compaction recovery, or the attach's own success record) left that entry claiming a child the failure path had already released. The next send found the phantom, skipped the start, and wrote at a fence the journal had moved past, so the message stayed queued for good. The entry now drops the released child and its event sink, and follows the record's fence, as a failure before indexing already did. * fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer The error strip for a message the host accepted and then did not deliver matched the entry before the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not send your message" with nothing to retry. It now reads the reconciled entry. A rejection the journal records before the send's own pending answer lands is final as well: that answer no longer puts the entry back to dispatching with no Retry. * fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down The preamble waits for its submission to be delivered while the worker's agent starts. When that wait ran out it threw operation_unknown, and the failed-start teardown then closed the session, which rejected the very preamble the host was about to deliver. It now reports a turn start nobody observed yet: the worker is start-unknown with its session kept, the host delivers the preamble when the agent starts, and the worker's report settles the dispatch as for any unobserved start. The receipt no longer suggests reading a screen a structured worker lacks. * fix(native-chat): a message rejected while its chat was closed reads as not sent A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked every later message behind a Retry and no reason, and the delivery probe, seeing the journal already answered, never ran. The reconcile now settles it as rejected like a dispatching one. * test(orchestration): name why the readiness settlement fakes are cast * fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent * docs(native-chat): drop the fence from the admission the send effects run behind * docs(native-chat): give the fence move on release the reason that still holds * docs(native-chat): stop citing a write fence check in launch and mailbox comments Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease. * refactor(native-chat): the provider child is its own record A conversation now outlives any number of provider children, so the child is one record on the conversation's entry instead of five loose fields beside its journal. It is written in one place: indexed only once an attach has fully succeeded, and ended through one function that an exit, a failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence. - A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence patch after it are gone. - Conversation writes read the record's fence, the way mutation admission already does; a child's own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of the conversation's fence, are gone. - The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer dropped when an attach replaced the whole entry. - Stop on a child still proving its start stops only the child: its lease goes back and the chat is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus the conversation's close. - The settlement retry uses the conversation's own journal, opened through the host's one open. * fix(native-chat): the delivery loop alone settles a message its start or child failed A queued message was settled by whichever path happened to end the child first: the loop, the unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that rejected every pending row. That gave two failure rows with different tones for one start, a loop that could hand over to a different child than the one it waited on, and a Claude start that died while starting reading unlike every other failed start. - The loop remembers the child it waited on. At handover, if that child is gone or replaced, it reads how it ended: a Stop continues; anything else writes one failure row and rejects every queued message with the same words, then stops. A child still starting whose start the adapter says did not land fails the same way. The exit, eviction and the settlement retry only settle the handed-over and legacy rows of the child that ended. - One failure row, always an error, keyed by the start. A start a view began that dies with nothing queued writes the same row through the same builder, so a second report revises it. - The open no longer rejects leftovers; the loop's first step does, and the open wakes it. - `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the failure before the exit is processed. - Quit closes every conversation the way closing a chat does: what is still queued is rejected as closed, with or without a child, and a start the loop already has in flight is waited for so the child it produces is stopped rather than left behind. * refactor(native-chat): a stopped child ends on the one reading of its stop The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that verdict to the child's ending, so the host never forms a second view of whether the root is gone. Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition, and a failed re-attach passes what its release saw. The end-of-child record can therefore also carry a stop whose root was not seen to go, which nothing ends on yet. * feat(native-chat): the host says it accepts a send before any agent has it The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same string capable clients already send. A client can then tell a host that answers a send at acceptance, and admits a Stop with no writer before a turn starts, from an older one that still restarts the agent inside the send. Additive: an older client ignores a capability it does not know. * refactor(native-chat): an attach never opens a journal of its own The attach adopts the conversation's open journal, which outlives it, so it no longer opens one for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag that told the two cases apart is gone. Tests that attach without a host open the conversation the way a host does. * fix(native-chat): a moved fence resends nothing on a host that accepts first The outbox treated any fence change as a new owner: it dropped the answer of a send in flight, queued that send to go out again under the same id, and unblocked a refused head. On an older host that is how a send the restart refused, unrecorded, gets another try. On a host that records every send before it starts an agent, a fence moves because that start ran, so the same rule resent into every failed start. With a fence stamped on every frame, that became a loop. The outbox now reacts to a fence change only when the host has not advertised that it accepts a send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed start reaches the client as a rejected message it keeps with its Retry. Against an older host, or before one has answered, the outbox behaves as it did. Desktop and paired web share this hook. * refactor(native-chat): a child's end says whether the user or the host stopped it The end-of-child record's cause now tells a user's Stop from the host stopping the child for a cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a user's Stop, as before, and fails the start it was waiting on after a host stop, with the one error row and every queued message rejected, in the stop's reason when it gave one. The reason stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet. * fix(native-chat): a chat whose only work is a queued message is not offered for resume A message accepted while the agent was starting counts as working in the chat, and quit rejects it as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a chat whose agent never had the message. The snapshot now reads only what was handed over. * test(native-chat): type the queued-message fixtures in the resume-offer tests * fix(native-chat): a start that dies while a message waits on it is that message's failed start Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When that start died, its exit wrote the start's error row and left the message queued, so the delivery loop started a second agent into the same failure and wrote a second row. A child's end now records where the conversation's journal stood, and the loop settles a message accepted before a failed start ended with that start: one row, under its key, and no second start. A message sent after the failure still gets a fresh start. * fix(native-chat): a request that failed reads as failed A structured chat whose only message the agent's start refused read as a green finish, and a cancelled structured turn did too: the host published a verdict only for turn records, and structured rows carried no `interrupted`. The host projection now reads the session's latest request: its turn's outcome, or `failure` for a send the agent or its start refused. A send that was withdrawn, or left undelivered by a restart or a close, fails nobody and makes nothing listable. The ingest publishes `interrupted` as the hook lanes do, and every reader decodes the verdict through one accessor, so a failure reads Failed on the dot, the rollups, history and `worktree ps`, behaves like a cancellation in every clean-finish policy, and notifies as "failed". * docs(native-chat): say what an attach's open conversation and unconfirmed ids are now * test(native-chat): a verdict change republishes the mobile status projection * refactor(native-chat): the store's retention trigger keeps its flag compare A verdict change always moves the completion clock the same check already reads, so a second verdict compare there caught nothing new. * test(native-chat): a user message the provider journaled keeps its session listed * test(native-chat): pin what a failed start settles, and what a resume offer names A view's child that dies while a sent message waits settles that message only when it died starting and no child has taken its place: a proven child's crash, or a second start since, gets the message delivered. The resume offer names the handed-over message, never a newer one still queued. * test(native-chat): the failed-start pins fail on what the message became, not on a timeout * fix(native-chat): a late provider-session update keeps a failed recovery record failed A provider-session heartbeat that rewrites a completed recovery record kept its interrupted flag but dropped the outcome it was copied with, so a live failed checkpoint read as a clean finish until the next status write. * test(orchestration): the preamble's host stub is typed, not cast The preamble send now takes only what it reads of the host, the send, the settlement wait and the record's fence, so its test builds that host with real types instead of `as never`. * test(native-chat): the terminal-bell check asserts the renamed verdict field The bell notification test still checked for agentInterrupted, which no longer exists, so it could not catch a verdict leaking into a bell dispatch. * fix(native-chat): a failed turn ranks like a completion for attention Attention readers (completion time, Smart Sort, sticky retention, Cmd+J Recent) now demote only a turn the user stopped. A failure is news the user has not seen, so it keeps its completion time, ranks in the Done class, stays retained after its pane goes away, and a retained failure reads failed in the worktree rollup instead of done. Clean-finish policy (hibernation, pane ownership, the value moment) still treats a failure like a stop. The retention trigger compares verdicts again: success -> failure no longer moves the completion clock. * fix(native-chat): a failed main agent reads failed while its subagents still work The verdict is now read from the main agent's own state, not the folded row: a main agent that is done and failed has a verdict even while its subagents keep the row working. Without mainAgent (history, worktree ps, older hosts) the old combined-done rule stands. Display marks the verdict through agentVerdictDisplayMark: a failure outranks every combined state on the agent's dot, label, tab badge, dashboard and activity rows; a stop marks only a done row, so a successful or stopped main agent with live subagents still reads working. Subagent rows keep their own state. The worktree card, terminal tab and Cmd+J rollups share one pane fold and rank a pending question, then failed, then working, monitoring, interrupted and done. worktree ps publishes the main agent's outcome on a working row, and the mobile mirror reads it. The store's change check, the paired-client mirror's equality and its epoch now see a verdict change on a working row, which otherwise moves no state or clock and left the worktree card reading working. Clean-finish policy is unchanged: a working row is never hibernated and has no completion time. * docs(native-chat): the worktree ps outcome comment no longer claims old hosts send it The field is new: an old host sends no outcome at all, so a reader falls back to interrupted. The removed clause said old hosts send it on done rows, which never shipped. * docs(native-chat): the status-store listing rule names provider-journaled user messages * fix(native-chat): a refused send notifies failed through the completion feed The host's completion feed followed only the newest turn, so a send the agent or its start refused, which creates no turn, read Failed on its row but sent no notification. The feed now follows the session's latest request, read from the projection the status feed already makes for the commit: a turn keeps its id, a refused send is named by its journal item key. It announces only while the session is idle, as the row reports a verdict, so queued sends refused one commit at a time notify once, and a withdrawn send falls back to a request already announced. * fix(native-chat): every copy of a row carries the main agent's own status History entries, sleep records and `worktree ps` rows carried a flattened top-level `outcome`, copied under different gates and without the main agent's clock. They now carry `mainAgent` (state, outcome, stateStartedAt), the type the live row already persists and sends, and every copy site takes it with `interrupted` through one function, `agentVerdictFields`. - The accessor reads `mainAgent` then the legacy flag; the mobile mirror matches it line for line. - Sleep records admit `mainAgent` with `normalizeMainAgentStatusField`, so a malformed value drops the field, never the record. - Mobile dates a main agent that failed under live subagents by its own clock, as desktop does, and its row equality compares `mainAgent`. - The activity feed reads a history entry's own `mainAgent` instead of rebuilding one; the sync key and history equality compare it. * test(native-chat): pin the worktree ps verdict across host and phone versions Pairs the real v1.4.212 host and phone row reader with this build: an old phone reads a new host's rows by `interrupted`, a new phone reads an old host's rows (no `mainAgent`) the same way, and a new phone reads a failure under live subagents as Failed, dated by `mainAgent.stateStartedAt`. The release checkout now carries the phone's self-contained row reader, and the lane runs when the `worktree ps` row producers change. * test(mobile): name the parity table's row for its role * fix(native-chat): a request that settles while the user is asked something notifies once The completion edge waited for an idle session, and a pending prompt (including a subagent's approval) is not idle. Structured chat has no other attention producer, so a main turn that finished while a subagent waited on the user sent nothing until the prompt was answered. The edge now waits only on owed work (a running turn or an unanswered send), which the projection reports even beneath a pending prompt. A request that settles with a prompt pending announces once; the renderer words it "needs input" from the host status mirror's `attention`, and answering the prompt keeps the same request identity, so it does not announce again. The wire shape is unchanged. * fix(native-chat): the completion says when the user is being asked A request that settles while a prompt waits on the user was worded "needs input" from the renderer's status-feed mirror. Remote clients receive the status and completion streams over separate sockets, so they can arrive in either order and the wording could be wrong both ways. The host already knows at emit time, so the completion now carries an optional `awaitingUser: true` in that case and omits it otherwise. The renderer words the notification from that field alone and no longer reads the status mirror. Old clients ignore the field and word by outcome; old hosts never send it. * fix(worktree-status): a departed agent's failure yields to live work on the worktree card A retained failed agent has no expiry, so ranking it with a live failure pinned the card to Failed over other panes' live work. It now ranks below working, monitoring and permission, and above every finished outcome. * docs(agent-status): a departed agent's failure ranks below live work on the worktree card * fix(native-chat): a view never restarts a chat whose last start failed A Claude chat whose CLI exits during startup left one red row per start, and every time a view bound to it (the chat opening right after its create died, or the user switching back to it) the hold started the CLI again, so the same launch-failure row repeated. Only a send retries a failed start now, the same rule provider-exit recovery already applied; the rule lives in one predicate the hold, exit recovery and the delivery loop share. * test(native-chat): start the child the loop waits on with an attach, not a second view A view no longer starts a child whose last start failed, so the R2 case that waits on a child started since the failure now gets that child from a client attach, the one non-send starter left. * fix(native-chat): settle a gone generation's turn wherever a conversation opens A send that opens a chat this process had not read yet (after a crash, from a phone or the CLI) went through the delivery open, which never settled what the dead generation left running; only the read restore and a successful acquire did. When the send's start then failed, the turn stayed running for every reader. The settlement now runs in the one journal open, at the crash boundary, for every opener except an acquisition, which settles from the evidence it read before its reserve; the read restore's separate step is gone. * test(native-chat): prove the next child's start settles the turn an earlier child left The R1 case lost its only settlement assertion when the latch it checked was deleted. It now seeds the running turn the earlier child left and asserts it ends at the exit's receipt, with the exit's row, before the message is handed to the new child. * test(native-chat): count a failed start's rows by row, not by text Comparing the set of texts passed when two different rows carried the same words, which is the duplicate the test exists to catch. * test(cross-version): load the phone row readers without mobile's toolchain Vite transforms a file against its nearest tsconfig, and mobile/tsconfig.json extends expo/tsconfig.base.json, which the root-only cross-version lane never installs. The worktree ps verdict suite imported the current phone row reader from mobile/ directly, so CI failed with TSConfckParseError before any test ran. The harness now imports a copy of the working-tree reader placed under the checkout cache, where the root tsconfig applies, as it already does for the release checkout's copy. Both readers are still the real files. * test(cross-version): keep the checkout path-guard message and justify the copy import's cast * test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget * test(native-chat): pin the open's and the send's start and row counts, however the view binds Opening a fresh chat whose starts fail makes one start and one row, with two views bound before or after the create's child died; one send makes one more of each. * fix(native-chat): settle a gone generation's turn at every open but an acquisition's The journal open skipped the settlement whenever the lease read reserved or live, to leave an acquisition's own open to the acquisition. But a lease a crashed process left in recovery also reads live, until the next acquire resolves it. A send that opened such a chat, from a phone or the CLI after a crash on a host that could not prove the old owner gone, skipped the settlement; when its start then failed, the dead turn stayed running for every reader. The acquisition now says it is the opener, and every other open settles, whatever the lease still claims. * test(native-chat): hold the create's start open until the views bind The "view binds while the create is still starting" case gave the create a 300 ms head start and asserted the views bound before it died. On a loaded runner the holds took longer, the create's exit landed first, and the case failed its own precondition. The create's initialize now waits on a gate the test releases once the views are bound. * refactor(native-chat): drop the composer's second error formatter After the merge with main, every chat write in the composer path reports its failure as a typed outcome worded by the refusal-notice table, so the send's catch sees only a local throw. The {code, message} formatter this branch added for it has no payload left to format, and its claim to be the one way a chat words a failure is no longer true. The composer send is main's again. * test(native-chat): pin the reason on a message rejected while its chat was closed The reopen test checked only that the message reads as not sent; it now also checks the Retry row carries the host's reason. * fix(native-chat): a send the provider never received after a restart has no verdict Restart reconciliation rejects a crash-stranded send that is absent from a trustworthy provider history with reason 'not_delivered'. Nobody failed that send, but the verdict allowlist did not name it, so after a crash the chat read Failed, was listed, and could notify "failed". Give the reason a shared constant (persisted value unchanged), add it to the no-verdict set, and treat it as an internal marker so the Retry row no longer shows the raw string. --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
a0e24905f6 |
fix(agent-status): a cancel never hides live work (#22476)
* fix(agent-status): a cancel never hides live work After the user cancels a turn, a background shell, scheduled check or subagent that is still running keeps reading as it truly is in both lanes. The fold no longer takes a verdict input; the cancellation survives only as lead.outcome, restated as the row's interrupted flag on a settled row for readers that predate lead. * fix(agent-status): keep a cancel's verdict and clock on every settle path A Grok turn cancelled while a task ran now reads monitoring, and the idle_prompt backstop that later settles it restated done without the row's `interrupted` flag, so notification readers announced the cancelled turn as a clean finish. Derive `interrupted` from the main agent's outcome, as the Claude builder already does. The inferred Claude cancel now folds through the host's local main agent record, which a relayed pane never refreshes, so a second cancel on an SSH pane inherited the first cancel's clock. The caller admits only a working main agent, so the cancel always starts a new done clock. * fix(agent-status): keep the shell fact on an inferred cancel so restart can seed it An inferred Ctrl+C cancel beside a working subagent publishes a row held open by child work, but the synthesized event dropped the row's paired claudeRunningNonAgentTask fact because mainAgent changed. Hydration seeds a settled main agent only when that fact says no shell ran, so after a restart the child's drain left the row working with no mainAgent. Carry the fact forward: a cancel does not change what the shell inventory said. * fix(agent-status): a Ctrl+C at an idle main agent's prompt cancels nothing Every row that publishes the main agent fact now admits an inferred cancel only while that main agent is working. Grok's Ctrl+C at the idle prompt leaves its background task running, so settling the monitoring row to done hid live work. Rows without the fact keep the evidence guard, and Codex keeps it too because its synthesized row is a plain done. * fix(agent-status): fold a relayed pane's cancel from its row, not the desktop's records The inferred Claude cancel read and wrote the desktop's own listener records for every pane. For an SSH pane those records are not the relay's: hydration seeds them from the saved row and nothing reaps them, so a subagent that finished on the remote after a desktop restart kept a cancelled row spinning with nothing running. A local pane still records the verdict on its listener and folds its own roster; a relayed pane folds only the child work its row carries. The relayed-pane parameter and forced clock the shared record path grew for this are gone. * fix(agent-status): hold a cancel verdict in the store until a new turn or the provider's own A relay never learns of the cancel the desktop infers from Ctrl+C, so its next child hook or reconnect replay restated the main agent as working and flipped the row back. The late-hook suppression that guarded this keyed on a done row flagged interrupted, which a cancel held open by a shell or subagent no longer is; it also dropped Grok's own stop_cancelled when the inference won the settle race, hiding the task that hook reported. The suppression is replaced by a latch derived from the row: its main agent reads cancelled (or, from an older host, a done row flagged interrupted). A settled incoming main agent, another prompt, an explicit prompt or a session start releases it. Child and replayed events keep the latched main agent and are re-folded with their own child evidence; late main agent work is held as before, and Codex keeps its record re-mark. * test(agent-status): pin Codex's evidence guard beside the main agent fact * fix(agent-status): a prompt submission ends the cancel verdict latch The task notification Claude starts when background work ends is a real turn, but it keeps the cached prompt and carries no explicit prompt, so within 15 s of a cancel the latch held its prompt submission and every tool event after it: the turn read as monitoring under a cancelled main agent until its Stop. The captured shell cancel has exactly this: the notification lands 0.17 s after the cancel key. * fix(agent-status): derive a Codex row's interrupted flag from its main agent The cancel verdict latch lets any settled mainAgent through, so a late root Stop after an inferred Codex cancel now applies where the old same-prompt window held it. It restates the cancellation on mainAgent but, unlike Claude and Grok rows, carried no interrupted flag, so mobile, the dashboard and notification text read the cancelled turn as finished. Codex rows (local and relayed) now derive the flag from the main agent record, like the other providers that publish one. * docs(agent-status): describe cancel admission for every provider and the store's cancel-verdict hold * docs(agent-status): correct the idle-prompt Ctrl+C claim to the measured CLI behavior * fix(agent-status): preserve waiting relay children on cancel * fix(agent-status): resolve the cancel hold before a child's permission card adopts a relayed main agent The permission-card hold took the incoming event's mainAgent before the cancel hold ran, so on an SSH pane a child's next tool under a sticky card restated the relay's stale working main agent and dropped the cancellation the desktop had inferred. * test(agent-status): pin that a cancelled turn's drained subagent settles as stopped, not completed * fix(agent-status): keep a cancel through a restarted relay's child hook and a teammate's idle A relay that restarts after a desktop-inferred cancel has lost its prompt cache, so the child's next hook arrived with an empty prompt, read as a new turn, and replaced the cancelled main agent with none; the row then stayed working after every child stopped. A child's empty prompt is now unknown, not another turn; a non-empty different one still releases, since it is the listener's newer prompt. TeammateIdle names its child by teammate_name and carries no agent id, so the latch treated it as the main agent's and let the late-hook window apply it after 15 s, reviving the cancelled turn. It is now re-folded as child work. |
||
|
|
b4d732685c |
feat(agent-status): combine Codex child work through the shared main-agent status fold (#22475)
* feat(agent-status): combine Codex child work through the shared main-agent status fold * docs(agent-status): correct two comments the waiting child-work arm made stale A child failure reported in place as `blocked` now pins the row `waiting`, not `working`; and no relay ever sent an unfolded `working` beside a waiting child. * fix(agent-status): only a waiting child asks for a human A child's `blocked` state means its task failed (the only producer maps a failed background task to it, and the background-task view labels it "failed"), not that a human must act. Folding it into the waiting arm would surface a failed child as needs-you. It stays live work, as before this series. * docs(agent-status): say a waiting child, not a blocked one, makes the row wait A child's blocked state means it failed; only its waiting state feeds the waiting arm. Two fold comments, a test describe and two parity story names still called the waiting child blocked. * docs(agent-status): name where a child's wait is still lost, and pin the structured lane's real input The doc said the Claude hook lane's rows match Codex and that every lane feeds a child's wait into the fold. Neither holds: Claude keeps the wait in one slot the next main agent event overwrites, the structured lane turns a child's prompt into the main agent's own attention, and Codex drops its roster on a root Stop when it tracks no child transcripts. The parity story now drives the structured lane with the input it actually receives. |
||
|
|
80f5aae0f9 |
feat(agent-status): publish the main agent's own state beside the combined row state (#22452)
* feat(agent-status): publish the lead agent's own state beside the combined row state
Every status producer folded the main agent's state together with live child
work into one `state`, so a lead that had finished while a subagent still ran
read `working` and its own state was lost. The row now also carries
`lead: { state, outcome?, stateStartedAt }`, admitted by the one payload
normalizer on the relay wire, IPC and disk, and published from the Claude hook
lane, the structured host ingest and renderer bridge, Grok (now on the shared
fold) and Codex (own combine kept). The persisted child-only boundary flag is
derived from `lead` plus child evidence and no longer written; old rows map
onto `lead` at hydrate. Combined `state` and `workingMode` are unchanged for
every reader; a cross-lane parity table pins that, with the cancelled-turn
watch-loop story recorded as a known divergence.
* fix(agent-status): make Orca's inferred interrupt the primary source of a Claude lead cancellation
Current Claude Code sends no hook at all on a cancel and no is_interrupt on
Stop, so the cancellation enters the lead record from the server's inferred
interrupt and rides into the next real Stop; is_interrupt on a turn boundary
stays as the secondary source for builds that send it. Comments, the store
reference and the parity table say so; no suppression changes.
* docs(agent-status): the child-only boundary comment now describes the persisted shell fact
The old sentence said a hydrated row no longer carries the shell fact, which is
the opposite of the mechanism: claudeRunningNonAgentTask is persisted precisely
so hydration can read it, and only a pre-lead row lacks it — reading as
shell-free, the same assertion its legacy flag made at write time.
* rename the lead fact to mainAgent: the main agent's own state
* docs(agent-status): the inferred cancel comes from Ctrl+C, not Esc
* fix(agent-status): an inferred interrupt keeps an already settled main agent, and the row verdict docs name its inferred source
* fix(agent-status): a child-induced wait publishes the main agent state it displaced
* fix(agent-status): decide child-held Claude rows from the saved main agent fact
Restart seeds the Claude main agent from the row's saved mainAgent whenever it
settled and no shell held the row, instead of re-deriving a child-only shape.
OSC cannot settle or repaint a row child agents hold open, including a row
waiting on a child's permission prompt. A sticky child permission prompt still
records the main agent's own progress, and OSC repaints and inferred answers
keep the shell fact beside the main agent they preserve.
* fix(agent-status): keep a finished turn's main agent verdict and clock with that turn
A Claude SessionStart restarts the main agent's clock instead of inheriting the
previous session's last Stop. A Grok idle prompt or session end, and a late
Codex root Stop after an inferred cancel, restate the same finished turn, so
they keep its recorded verdict; only a new turn clears it.
* test(agent-status): publish the Grok verdict restatement past the late-event window
* docs(agent-status): describe hydrate seeding and the OSC refusal from the saved main agent fact
* fix(agent-status): push a held child permission row when its main agent changes
* fix(agent-status): keep the shell fact on a held child permission row so restart does not settle it
* docs(agent-status): note the held child permission row carries the shell fact and is pushed
* fix(agent-status): pair the Claude shell fact with the main agent at the one row-build point
Every non-hook rewrite (terminal-title repaint, inferred answer, held child
permission) had to re-carry the shell fact beside `mainAgent`, and each one that
forgot let a restart settle a row while a shell still ran. The row builder now
pairs the fact once: a listener event restates it, any other write keeps it only
while `mainAgent` is unchanged. Restart seeds a settled main agent only when the
row says no shell ran, and legacy child-only rows map to that explicitly.
A held child permission now also accepts the main agent event's background
evidence, as it already accepts its `mainAgent`, so the child's drain no longer
settles a row a shell still holds. The renderer keeps a previous `mainAgent`
only for writers that never carry one, so a hook row without it matches the
host snapshot.
* test(agent-status): pin that restart never seeds a main agent from a row silent about its shell
* docs(agent-status): the row builder pairs the shell fact with the main agent, and restart seeds only on an explicit no-shell
* test(agent-status): name the legacy-row case parameter for what it holds
* docs(agent-status): name which rows carry the main agent fact
|
||
|
|
4c696a1e2a |
fix(agent-status): a structured session with live child work reads as working (#22295)
* fix(agent-status): a structured session with live child work reads as working An idle native-chat session whose subagent was still running showed a green check in the sidebar, the collapsed worktree pill, and worktree ps, while a terminal Claude session in the same situation showed working. The two lanes folded child work into the parent's status with different code: the hook listener did, the structured lane did not. Both lanes now share one child-work liveness vocabulary and one lead-status fold. Live agent work makes a settled lead working; shells and monitors alone make it monitoring. The structured lane derives liveness from the background task list already on the wire, in both its readers, so the sidebar, the CLI, the dashboard and mobile agree. The Claude task-kind table is one shared file covering both the hook inventory and SDK stream names, and the renderer bridge reuses the shared child-work projection instead of carrying its own copy. * fix(agent-status): a blocked or out-of-contact subagent still holds its session working Child-work liveness retired an agent-kind child on any state but working/monitoring, while the shell beside it stayed live on everything except done/idle. A subagent waiting on a permission prompt, or one whose host lost contact, therefore counted for less than a backgrounded sleep and let the session read done. Both kinds now share the settlement rule `resolveAgentChildWorkFreshness` already reads rows by: only an explicit done/idle retires child work. Also keep empty task labels out of the shared background-task projection candidate, so a host that publishes `name: ''` cannot beat the child-row fallbacks. * test(agent-status): pin the widened hook-inventory agent names, and correct two stale claims The hook inventory now classifies through the shared kind table, which also maps the SDK stream's `local_agent` / `local_subagent`. Nothing pinned that widening, so add cases for all four agent names — including `teammate`, whose pane state stays `done` under the #8825 idle-squat rule. Two comments the fold made false: - the teardown marker rule's comment claimed it could not disagree with what the UI calls working; it is deliberately lead-only, so now it says that and why; - the agent-status store reference still described the structured row's `state` as the deleted `structuredAgentSessionStatusState`, and omitted the `workingMode` the ingest now writes. * fix(agent-status): the state clock restarts when monitoring becomes a real turn `stateStartedAt` carried forward whenever the prior `state` matched, which was sound while `state` meant "a turn is running". Now that it folds in child work, an idle lead watching a `sleep 3600` publishes `working`/`monitoring`; the user's prompt 45 minutes later keeps `state: 'working'`, so the row inherited the watch loop's clock and read "Working for 45m" the instant the turn began. Monitoring is its own displayed label (`worktree-card-compact-agent-row.tsx:40`), so the continuity key is now the whole published work identity — state AND workingMode — in both writers. Also record two facts the code stated wrongly: the structured lane's `interrupted: false` is inert (a projected session status has no interrupted member) rather than a decision, and the child-work liveness rule's escape hatch is the roster's session lifetime, not a settled state. * fix(agent-status): a workflow is watch work, and child work dates itself Two defects the fold introduced. `isAgentChildWorkKind` counted `workflow` as agent work, so a structured session whose only live task was a backgrounded `local_workflow` published a full working spinner while the children projection — which admits `kind === 'agent'` only — rendered nothing to expand, and the same workflow in a terminal pane showed the monitoring badge instead. The repo already decides this: `isClaudeSubagentTask` excludes workflows by name, and MATERIALIZED_TASK_KINDS leaves "the backgrounded shell command and the workflow" to the non-agent owner. The predicate is now `kind === 'agent'`, and the three sites that restated the same test route through it, so a new kind is decided in one place instead of three that merely agree. `evidenceObservedAt` dated every row by `summary.updatedAt`, the journal's last activity. The journal cannot date child work: its clock stopped when the lead's turn did, so a genuinely live roster aged past the 30-minute staleness window and mobile's dot decayed a running session to idle. The fold now reports whether child work alone holds the row open, and only then does the host's observation clock stand in — keeping "a restart's republish is not new evidence" for lead turns. * fix(agent-status): the sidebar dates child work the same way the host does `fromChildWork` reached the host ingest but not the renderer bridge, so after ~30 minutes of live child work with no journal activity the sidebar's row aged into staleness while `worktree ps` and mobile stayed fresh — two writers for one session answering differently, which is the defect this PR exists to remove. For a remote host the client's own receipt time is also the more honest clock, since the journal stamp is the host's and is never comparable against this machine's now. |
||
|
|
e8a7be4ce2 |
fix(omp): recover retired pane status with validated restart authority
Merged after fresh run 35448889017 passed all required checks, including static analysis, typecheck, package jobs, all test shards, changed E2E, Docker SSH E2E, and verify. |
||
|
|
291b4ddd6f |
feat(agent-status): route structured sessions through canonical ownership (#20718)
* feat(agent-status): route structured status through canonical ownership and fence child lifetimes Restacked onto the canonical store and child-work contract. Completing that restack drops the `reopenStructuredParent` mutation flag this change had carried, along with its contract field, its codec branch, and its single call site in structured ingest, which passed a hardcoded `true`. The flag was a narrow escape hatch from the absolute `tombstones.has(...)` rule that governed parent upserts in this branch's original base. The canonical store replaces that rule with a revision envelope, because a bounded store compacts tombstones away and a presence-based guard silently stops fencing once one is evicted. With the envelope deciding the outcome, the escape hatch has nothing left to escape from, so removing it changes no production behaviour. `agent-status-store-reopen.test.ts` is rewritten against the envelope: the reopen case now pins that an unflagged republication succeeds while replay from before the reopen stays fenced even after the parent tombstone is compacted away, and the second case pins where the guard genuinely bites — a republication inside the removing mutation itself, for every subject kind. * fix(agent-status): re-admit unchanged structured owners after teardown * fix(agent-status): clear anti-slop object-param and Reflect.apply findings - agent-status-store-byte-budget.ts: type the byte-budget helper's record parameter as the union of what its call sites actually pass (the snapshot header plus each store entity record) instead of the broad `object`. - server-structured-canonical-status.test.ts: replace `Reflect.apply` with a typed, explicitly-bound call that models a caller at an untyped boundary omitting the trusted owner subject. * docs(agent-status): drop the 2A progress doc from docs/reference docs/reference/ holds implementation detail, not rollout progress. The canonical-boundary notes move to the effort's working directory; the agent-status-store status section keeps the boundary statement and loses the now-dangling link. * fix(agent-status): mint the canonical epoch on first use, not at construction The hook server's canonical store was built in an instance-member initializer, so constructing AgentHookServer — which happens at import time for the module singleton — demanded a live randomUUID. Any importer that stubs node:crypto threw 'Invalid agent status store epoch' before a single test ran. The store is now created on first canonical access and reset by dropping it, so construction owes nothing to a crypto implementation and the epoch still rotates per authority incarnation. * fix(agent-status): drop the orphaned snapshot budget and a duplicated pane guard Two leftovers from the canonical-store routing change: agent-status-store-snapshot-budget.ts lost its only caller when the store state switched to agentStatusStoreFitsByteBudget. Nothing in the repo imports it now, so the module goes with the caller it existed for. The replacement is not a straight copy: it only memoises a record's measured size once the record is frozen, so a still-mutable record can no longer return a stale byte count. persistedStructuredWorkerPaneKeyIsValid repeated its public-pane-key rejection verbatim three lines below the first one. The tests covering that rejection pass on the first occurrence alone, so the second decided nothing and only obscured which predicate was load-bearing. * fix(agent-status): stop a failed structured publish from latching as owned Three defects found reviewing the structured routing path. combinedStatusEntries defaulted a missing listing order to 0, but the counter it compares against starts at 1, so any unordered row sorted above every ordered one. Unknown order now sorts last. The owner map recorded a session as owned before the sink ran. A publish that threw therefore left matchesLocation reporting an owned location for a row that was never written, and the unchanged-projection path — the only thing that would re-offer it — stopped. The address still has to survive a throw so teardown can forget a row that did land, so the two facts are now separate: the address is recorded up front, and only a publish that returned marks the row as landed. The reopen test claimed the revision envelope rather than the tombstone fences a stale replay. It cannot tell: transport consecutiveness, the parent-revision validator and the tombstone guard each refuse that replay alone, and ablating any two leaves the test green. It now asserts the outcome and says so. |
||
|
|
da5d555259 |
refactor(agent-status): delete the runtime's retained row store (PR 1b) (#19785)
* docs(agent-status): plan PR 1b at file level Names the five RuntimeAgentRowStore call sites and what each becomes, why terminalHandle has to be stamped before the store can go, and the one intended behavior change. * feat(agent-status): stamp the pane terminal handle on hook-server rows The runtime's retained row store carried the pty binding two readers need. Put that fact on the row that already owns the pane instead, resolved through the same lookup the renderer-facing IPC boundary runs, so the two surfaces cannot disagree about which terminal a pane is. Carried forward when a later write resolves no handle (only main's OSC parse can), and never persisted: a handle belongs to the runtime that issued it. * refactor(agent-status): route the session-tabs republish off the store `retain()` was not only a duplicate store: its boolean return was the signal that republished `session.tabs` for a status-only transition, which no title change covers (#7970). `hook-status-session-tabs-invalidation.ts` already mirrors that change set plus hook restore provenance, so route the signal off the store rather than keep a second comparator. Adds the status-drop arm a user dismissal emits, which the pane-clear fan-out deliberately skips — now load-bearing, because a dismissed row leaves the listing at once. Installed on both hosts. orcad had neither the OSC producer nor this signal, so its runtime observed agent status and published it nowhere; deleting the retained copy without wiring it would list no PTY agents there at all. * refactor(agent-status): delete the runtime's duplicate retained row store `RuntimeAgentRowStore` held the same payload the hook server already holds, so the same pane could legitimately read differently in the sidebar, in `worktree ps`, and on the phone. Both of its readers move onto the store's snapshot in `runtime-hook-agent-row-selection.ts`, and `collectRuntimeWorktreePtyAgentSources` loses the retained-versus-hook reconciliation that only existed because two stores could disagree. `ConnectedPtyEvidence` trades its flat pty-id set for `ptyIdByTerminalHandle`, which is how a row still resolves the connected PTY behind it — the working-terminal rollup's match key, and the last rescue for a row whose pane binding a controller incarnation nulled under it. The one intended behavior change: a row the user dismisses on the desktop leaves `worktree ps` and mobile at once instead of lingering until the pty exits. One store means one dismissal. The suites written against the retained store are rewired to a real AgentHookServer rather than deleted, so each still asserts the listing behavior it named. * docs(agent-status): record what PR 1b landed Past tense, plus two corrections to the plan: `terminalHandle` is not the pty id (they are different identifiers, and the explicit-status reader was already comparing against a real handle), and the legacy numeric pane key is a consequence the plan did not name. * fix(agent-status): harden single-store lifecycle * fix(agent-status): preserve mobile terminal rejoin * fix(agent-status): preserve unverifiable remote rows * fix(agent-status): own PTY row lifecycle in hook server * fix(agent-status): preserve state and renew freshness * fix(agent-status): ignore freshness for dismissed identity rows * fix(agent-status): fence orcad observed identities * fix(orcad): always release daemon adapter on cleanup * fix(agent-status): cover remint and headless lifecycle edges * fix agent status identity recovery gaps * fix(agent-status): suppress duplicate child-only row mutation * test(runtime): preserve hook store wiring in transcript harness --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
ebb1acfa37 |
refactor(agent-status): publish structured sessions into the hook server store (#19683)
* refactor(agent-status): publish structured sessions into the hook server store Structured (native chat) sessions have no PTY and no hook script, so their status never reached the hook server's store; #19217 gave `worktree ps` its own adapter over the structured feed instead. The feed now writes every projection into that store through a status sink the runtime wires, drops the row when the host closes the session, and `worktree ps` reads the one snapshot like every other agent. Rows carry a `structuredHost` marker and the journal clock; they are never persisted to last-status.json, and the main process does not forward them to the renderer yet, whose feed bridge still owns them until it is retired. Design and the two follow-ups: docs/reference/agent-status-store.md. * chore: drop stray @pnpm/exe lockfile entry An unrelated local pnpm run added @pnpm/exe as a packageManagerDependency with no package.json change, so CI's --frozen-lockfile install failed before any job ran. * docs(agent-status): describe the step that actually landed The design record claimed PR 1 deletes RuntimeAgentRowStore, drops the retained-versus-hook reconciliation, stamps terminalHandle on OSC rows, and tags rows with a source field of 'structured-host'. None of that is true of the shipped code: the retained store and its reconciliation are still in place, and the row field is structuredHost: 'held' | 'owned'. AGENTS.md points every future contributor here before they touch agent status, so split the roadmap into the 1a that landed and the 1b that has not, and name the fields the code actually writes. * fix(agent-status): pair session removal with the status-row forget A session dropped from the host's map without an explicit forget left its row in the store forever: `structuredHostOwned` bypasses the staleness check, so a failed re-attach (the Claude rewind path reaches one) stranded a permanently working agent in `worktree ps` and on mobile with no UI able to clear it. Deletion and forget are now one operation both callers route through. * fix(agent-status): give orcad the store worktree ps reads from `orcad` constructed its runtime with neither `getAgentStatusSnapshot` nor `structuredAgentStatusSink`, so once `worktree ps` sourced rows only from that snapshot the headless host published nowhere and listed nothing. The hook server's store is a module singleton whose import tree never reaches Electron, and its file paths come from `start()`, which orcad never calls. * fix(agent-status): drop a structured row without a renderer clear `dropStructuredStatus` went through `clearPaneState`, which fans a pane clear out to the renderer for a pane key the renderer's own feed bridge still writes - so 'exactly one writer per pane key' held for writes and not for deletes. `dropStatusEntry` routes through the status-drop tap instead, and skips the resume-identity remnant: a structured session has no pane to resume into, and every null-status publish would otherwise re-mint one. * test(agent-status): pin both half-migration structured-row filters Neither the `agentStatus:getSnapshot` filter nor the main-window listener's had a single assertion, so deleting either — the first step of PR 2 — was green everywhere. Also covers the perf skip and the drop's lack of a renderer clear. * docs(agent-status): correct three statements this PR made false The sink JSDoc claimed only tests construct a host without one; `orcad` did. The doc argued a structured row needs no tab mirror 'because headless serve has no renderer', reasoning about exactly the topology the wiring had not reached. The deleted runtime adapter's warning that the pane key must be the DERIVED one - never a bearer handle or minted worker key - was lost with it. * test(agent-status): declare orcad in the hook-row producer census Wiring the hook store into the orcad runtime added a production site that hands hook rows to a consumer, which the census ratchet pins deliberately. --------- Co-authored-by: Merge Sim <sim@local> |