Commit Graph
172 Commits
Author SHA1 Message Date
f2af92b2fa feat(native-chat): show live background work and name each row by kind (#19705)
* fix(codex): reserve the label's share of a qualified command row

A child's label is raw provider text and was spliced into the command
row unbounded, then the pair clipped to the description cap. A label at
or past that cap clipped the command away entirely, leaving a row of
kind 'command' that named an agent and showed no command - the failure
qualification exists to remove, inverted. The same clip could also cut a
surrogate pair, which boundSubagentField already guards against on the
agent row two lines away.

Give the label a reserved share and clip it the way the agent row does.

* feat(native-chat): show live background work and name each row by kind

The strip suppressed itself in three places: the Claude tracker blanked
its roster for the whole of any turn, the Codex tracker returned nothing
while a primary turn was open, and the renderer view gated on
`turnId === null`. Between them, work in flight was never shown — and a
task backgrounded in an earlier turn vanished from the strip as soon as
the next prompt was sent. Claude additionally dropped every foreground
subagent, so a fan-out reported nothing at all.

Report work while it is live, in all three layers. Foreground Claude
work is turn-scoped, so `result` retires it — that is the provider's own
outcome for a task it marked foreground, not a roster sweep. Nothing
settles a Codex child on turn end: those keep reporting well past their
parent, so turn frames only prompt a republish.

Name each ROW by kind — Subagent, Shell command, Workflow, Monitor —
instead of a generic "Background <kind>", each drawing the glyph the
shared tool-icon table already uses for that category. A row that
carries a provider description still shows it unchanged. The collapsed
header summary is deliberately untouched; it is owned elsewhere.

The conversation-command gate is unchanged in effect: an open turn
already refuses first, and Claude foreground work never reaches the
backgrounded set the gate reads.

* fix(native-chat): withhold the row stop Claude foreground work cannot honour

The strip now publishes foreground rows, but `stoppableTaskIds` still filters
on `backgrounded`, so `stopClaudeBackgroundTasks` resolved an empty target list
and returned `{ cancelled: false }` that no renderer reads: the user clicked
"Stop Subagent" and nothing ever happened.

Carry stoppability per row instead of widening the stop to a target the SDK has
no way to reach. `AgentSessionBackgroundTask.stoppable` is absent-means-yes, so
hosts that predate it keep their working control, Claude emits `false` only on
foreground rows, and the strip hides that row's button the same way it already
hides the stop-all a provider cannot honour.

* fix(claude): scope aggregate-roster authority to the work it enumerates

`background_tasks_changed` lists BACKGROUNDED tasks, so a foreground subagent
can never appear in it. Treating it as the whole world meant any such frame
cleared every live foreground row mid-flight and then dropped every later
foreground `task_started` for the rest of the session, killing the in-turn
fan-out the strip exists to show in any session that ever backgrounds anything.

Decide `backgrounded` before the staleness guard and apply the guard only to a
backgrounded start, and retain live foreground entries across a roster replace.
Retained rows count against MAX_TRACKED_TASKS, so the map stays bounded, and a
stale backgrounded start the roster no longer lists is still dropped.

* test(native-chat): pin the strip's monitor amber to the constant that defines it

`MONITOR_GLYPH_COLOR`'s comment claimed a test held it and AgentStateDot's amber
together, but no test imported it — the assertions hardcoded 'text-yellow-500',
so the two could drift with every test still green. Read the colour from the
module, which is what the comment always said was happening. Drop the unused
`BackgroundTaskGlyph` export too: nothing outside the module names it.

* fix(native-chat): keep the task list open across a gap in live work

The strip is now mounted on live work, so a sequential fan-out unmounts it
between one subagent finishing and the next starting: local `useState` meant
the expanded list collapsed itself on every such gap, on top of the strip
flickering above the composer.

Hand the disclosure to the session, keyed by session id so it does not leak
across a session switch. The strip is now controlled and holds no state of its
own, which is what makes it survive its own mount churn.

* fix(codex): route every command-row cut through one surrogate-safe clip

`boundLabel` avoided splitting a pair, then `qualifiedDescription` re-cut the
COMPOSED string with a raw slice: label (<=96) plus separator plus description
(<=512) is up to 611 chars, so that second cut landed at an arbitrary index
inside the description and could publish a lone high surrogate — lossy through
any non-JSON UTF-8 hop. `parse` had the identical hazard on an unqualified
primary-thread command.

One `boundText` helper now owns all three cuts, so no path in the file can emit
a lone surrogate from well-formed input.

* fix(claude): keep terminal evidence for ids an aggregate roster never lists

Narrowing the admission guard to backgrounded starts left a finished FOREGROUND
id with no defence: `replaceAggregateRoster` wiped `terminalTaskIds` wholesale,
so after any `background_tasks_changed` a replayed `task_started` revived a task
whose completion had already been seen — and only a later `result` could settle
it again.

Scope the wipe the same way the guard was scoped: delete only the ids the
incoming roster actually enumerates. A roster still overrules terminal evidence
for the work it lists, which is what that behaviour was added for.

* fix(claude): keep retained rows in place and evict the stalest, not the newest

Re-adding retained foreground entries after the roster made a live row the user
is reading jump below the backgrounded rows on every `background_tasks_changed`,
and the cap `break` kept the STALEST retained rows while dropping the newest.

Merge in the tracked map's own order so a surviving row holds its position, and
count the overflow up front so eviction takes the oldest retained rows. Roster
entries are never starved and the map stays bounded either way.

* fix(claude): retire leftover foreground rows when the next turn starts

A foreground `task_started` arriving with no turn open has no `result` coming
to retire it, so it sat in the strip indefinitely — with no per-row stop, since
foreground rows are not stoppable — and refused conversation commands behind an
instruction nobody could follow.

Settle on turn start as well as on `result`. This is cleanup only: visibility
never consults `startsTurn`, so a missed one degrades to today's behaviour and
can never switch the feature off. It shortens the row's life to the next turn;
the case where no further turn is ever sent is filed separately.

* fix(agent-session): withhold unstoppable rows from readers that predate them

Rule 3 of remote-wire-compatibility: changing what the host publishes reaches
old clients with no wire change. The Claude host published no foreground rows
before this feature; it does now, and a client that cannot read `stoppable`
draws a per-row Stop on every one of them — Claude always sets
`supportsTaskStop` — which filters to the backgrounded ids, stops nothing, and
returns a result no renderer inspects. That is the dead button `stoppable` was
added to remove, reappearing across a version skew.

Negotiate it. A client can advertise the existing background-task-stop
capability and still predate `stoppable`, so this needs its own constant.
Readers that do not advertise it get unstoppable rows dropped, and a state whose
every row is dropped becomes no strip — exactly their pre-feature view.

RUNTIME_PROTOCOL_VERSION is not bumped: this adds an optional field and a new
negotiated capability, and changes no existing field's meaning, which is the
explicit do-not-bump case in protocol-version.ts.

* test(agent-session): name the projected rows so the fixture typechecks

An indexed lookup into the fixture's task list is possibly-undefined under
`pnpm tc`; the rows are more readable named anyway.

* test(web): advertise the row-stop capability in the e2ee auth expectation

The web e2ee handshake started sending
AGENT_SESSION_BACKGROUND_TASK_ROW_STOP_CAPABILITY, and this test asserts the
advertised list by deep equality, so it went red on CI while every targeted
test run stayed green. Add the capability in the position the router sends it.

* test(claude): pin why the roster empties mid-turn in a sequential fan-out

The strip unmounting between two sequential subagents is truthful, not a swept
row: A leaves on the provider's own terminal frame, B does not exist yet, and
backgrounded work spanning the same gap holds the roster open — so an empty
roster is never work the strip is hiding.

Also pins the previous-turn rule against the one the subagent roster already
applies on the same frame: a still-working FOREGROUND child becomes
`unverifiable` there and a backgrounded one is left alone, so the strip drops
the first and keeps the second rather than asserting `live` for either.

---------

Co-authored-by: Merge Sim <merge-sim@users.noreply.github.com>
Co-authored-by: Merge Sim <sim@local>
2026-09-10 12:16:30 -07:00
Brennan BensonandMerge Sim 5868fdc9e3 feat(native-chat): report Codex background tasks in the chat strip (#19346)
* feat(native-chat): report Codex background tasks in the chat strip

The background-tasks strip works for Claude only; a structured Codex
session shows nothing in it. Feed it from the Codex app-server stream.

The strip stands for work that OUTLIVED a turn, which is what the
monitoring header, Claude's foreground suppression, and the conversation
command gate all already assume. Codex has no `is_backgrounded` flag, so
that fact is derived from the turn boundary: a `subAgentActivity` child or
a primary-thread `commandExecution` becomes visible once the turn it
belongs to completes and it is still unsettled.

`turn/completed` only reveals a task here, never settles one — measured on
`codex app-server` 0.153.4, a spawn_agent child reported `completed` 95.8s
after its parent turn ended. Only a child's own activity kind settles it.

Codex exposes no honest stop: `turn/interrupt` on a child ends its turn
without emitting a terminal activity item and leaves its shell running. So
the state carries a new optional `supportsStopAll: false`, the strip hides
a control that could not act, and the blocked-command message asks the user
to wait rather than to press a button that does not exist.

* refactor(codex): move session teardown out of the structured adapter

Merging main crossed the 300-line cap on
`codex-structured-session-adapter.ts`: the rewind backend (#19235) and this
branch's close-time strip clear both landed in it. The four close paths move
verbatim into `codex-structured-session-teardown.ts`, where they funnel
through one `settled` helper instead of repeating the notification-retry and
background-task cleanup at each call site. No ratchet bump.

Also normalize a background task's description once at receipt rather than on
every projection; the roster is re-projected on each observed frame.

* fix(codex): drop the shell row the journal already settles

A `commandExecution` still `inProgress` when its turn ends was reported as a
`command` task. But `settleCodexJournalTurn` writes exactly those items to the
journal as `state: 'failed'` on `turn/completed` and forgets them, so the strip
row would have claimed a shell was still running at the same instant Orca
recorded that it was not — two surfaces contradicting each other about the same
process.

A subagent is the opposite case and stays: the roster pointedly does not sweep
at a turn boundary, because children measurably outlive it. That leaves the
producer making exactly one claim — these spawn_agent children are still live
after their turn — which the durable roster row corroborates.

* fix(native-chat): track Codex background execution lifetimes

* fix(native-chat): keep running tool groups from claiming completion

* Fix runtime catalog and capability expectation

* fix(codex): keep a child's name on the command row that outlives it

A child agent's commands stay hidden behind its agent row while the child
works. Once the child's turn settles with a command still running, that
command surfaces as its own row labelled from the raw command string, so
'long_probe' became "/bin/zsh -lc 'ping -c 300 127.0.0.1 > /dev/null'"
at the moment that row was the only remaining signal for the work.

Qualify a child's command row with the child's label. Resolved on read,
so a label registered after the command still lands, and bounded by the
existing description cap so admission accounting stays valid. Primary-
thread commands are left unqualified: they have no child to name.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-09 00:13:20 -07:00
Brennan BensonandMerge Sim d15a6df224 Render task checklists with update diffs and a composer progress panel (#19230)
* Render native chat task lists with incremental checklist updates

* Keep native chat task-list review plan out of repository root

* Render live Codex plan notifications through task checklists

* fix(native-chat): keep checklist test fixture within shared boundary

* Keep agent task progress in one composer panel

* Restore inline task checklists and historical update diffs

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-08 23:50:32 -07:00
eedaf2bdfc perf: check backfill date cardinality before expanding ranges (#19456)
* perf: check backfill date cardinality before expanding ranges

* test(codex): pin the backfill cardinality gate to the enumerated range

Differential coverage at maxDates === length and length - 1 across leap days,
century rules, year rollover and DST switch dates.

* test(codex): type the backfill cardinality table as date tuples

Untyped it.each rows widen to string[], which tsc rejects when cast to the
3-tuple CodexSessionBackfillDate.

---------

Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
Co-authored-by: Neil <neil@stably.ai>
2026-09-08 19:53:24 -07:00
6d2d63104a perf: reuse naturally ordered unique Codex trust ranges (#19492)
* perf: reuse naturally ordered unique Codex trust ranges

* test(codex): pin the trust-range ordering the dedup removal relies on

Removing the pairwise dedup+sort is only sound while the scanner emits
strictly ascending, non-overlapping spans. Guard that precondition so a
future scanner change cannot silently widen or drop a trust block.

---------

Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
2026-09-08 19:44:31 -07:00
OrcaWinandm4air 9039cd522e perf: index forgotten Codex turns for bounded ordinal retention (#19488)
Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
2026-09-07 22:21:39 -07:00
Brennan BensonandMerge Sim 9f044031fc fix(native-chat): render compaction notices, plan documents, and images (#19228)
* fix(native-chat): render compaction notices, plan documents, and images

* fix(native-chat): avoid repeating notice text in details

* fix(native-chat): journal canonical and legacy compaction events

* test: add digest to native chat notice payload fixture

* chore(native-chat): drop the planning doc from the PR

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-07 16:21:47 -07:00
Brennan BensonandMerge Sim d0506bf5de feat(native-chat): add execution details and tool row identity (#19226)
* feat(native-chat): annotate tool rows with execution and source details

* fix(native-chat): require explicit MCP identity for tool annotations

* test: add required state to MCP projection fixture

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-07 13:51:47 -07:00
Brennan BensonandMerge Sim ce4a3a4186 feat(chat): add structured session rewind backend (#19235)
* feat(chat): add structured session rewind backend

* fix(chat): make interrupted session rewinds recover safely

* fix(native-chat): negotiate rewind runtime capability

* fix(native-chat): consolidate remaining adapter imports

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-07 12:20:24 -07:00
Brennan BensonandMerge Sim 0252fe5c36 feat(native-chat): show Codex subagent activity instead of opcode rows (#18773)
* feat(native-chat): show Codex subagent activity instead of opcode rows

Codex spawns subagents and reports their lifecycle, but Orca rendered only
gray `codex · item:subAgentActivity` opcode rows. Build the real display: one
summary row per spawn group with a live working count and token usage.

State is accumulated from `subAgentActivity.kind` alone. A live probe against
app-server 0.152.1 showed `agentsStates` arrives empty even in a real subagent
run, and that every activity item is delivered twice (item/started and
item/completed), so every transition is idempotent and terminal states latch.

Children never receive `thread/started`, so there is no nickname, role, or
depth to read; the row labels from the trailing segment of `agentPath`.

Two sweeps keep a row from claiming work forever: the parent turn's terminal
event settles still-running children, and session start marks a pre-restart
roster unverifiable rather than exited, since Codex resume replays no
non-message items and no event can ever settle them.

The roster rides a new NativeChatBlock variant paired with a plain-text twin.
A journal item kind could not be used: that union is closed, and an unknown
kind parses as malformed, which is the corrupt-journal class that can hide the
chat tab. Block types are explicitly admissible when unknown, so an older
client drops the block and renders the sentence.

MessageRow moves out of NativeChatMessageList to keep both files under the
max-lines budget without a disable.

* feat(native-chat): give the subagent summary row its bot glyph

The row led with a glyph that swapped on state — a check once every child
completed, a group icon otherwise — so a group appeared to change identity
the moment it settled. Per the approved mock, the glyph names the category
and never moves: state is carried by the status dot and the tone of the
words beside it.

Use lucide `bot`, the same glyph the individual `subAgentActivity` rows take
in the eight-category vocabulary, so the summary reads as their parent. Slot
and glyph are the mock's 16px/14px, muted by default, and the svg is
`aria-hidden` — the headline is what a screen reader announces, so the icon
never stands alone.

* fix(native-chat): correct the Codex subagent roster's build, journal write, and failure reporting

* Restore the exhaustive block handling that adding `subagent-group` to
  `NativeChatBlock` broke. `formatWorkerTranscriptMessage` and `boundBlock`
  both fell through to `image-ref` field access, so `tsc -p` failed for the
  CLI and node projects and `build:cli` could not emit. Both now guard on
  `image-ref` explicitly and give the roster block its own branch.

* Stop the roster's publish from evicting its own append. The sink queue
  coalesces by `coalescingKey` alone with no op-kind check, so passing the
  append's key to `tryPublish` spliced the queued append out and the row
  never reached the journal — permanently, since `lastSerialized` was
  already set. `tryPublish()` now takes no argument, matching every other
  call site. The regression test's fake sink honours the key, which the
  previous fake did not.

* Keep `collabAgentToolCall` substantive. Only the MultiAgentV2 path emits
  `subAgentActivity`, so a V1 turn has no roster row; suppressing its collab
  tool calls too would have left a V1 fan-out showing nothing at all.

* Surface a settled failure while siblings still work. The summary now
  reports the worst adverse outcome independently of the group verdict, so
  the row shows `3 working +1 failed` with a failed-coloured dot instead of
  a neutral pulsing dot. The plain-text twin names it too.

* Treat `/morpheus` as a child. Only `/root` is the turn itself; the old
  segment-count test silently dropped a valid single-segment agent.

* Refresh token-usage recency on update so an active thread is not evicted
  as the oldest entry, and scope the `agentsStates` comment to the V2 path.

* fix(native-chat): stop the subagent roster announcing a new duration every second

The roster row is an `aria-live="polite"` region and it contains the elapsed
clock, which reticks once a second for as long as the fan-out runs. A screen
reader therefore reads out a fresh duration every second, burying the state
changes the live region exists to report — the headline, the verdict, and the
`+1 failed` alert.

No other live region in the transcript does this. `NativeChatToolRun`'s live
button holds only the active tool label, and in `NativeChatWorkingStatus` the
variant that shows a duration is precisely the one with no `aria-live`.

Hide the clock from the accessibility tree only while it is moving. Once the
group settles the duration is fixed, so it stays readable and costs no
announcements.

* fix(native-chat): retry a refused roster publish, and stop two wrong readings

Four defects from a third review pass over the Codex subagent roster.

`write()` set `lastSerialized` before the append and rolled it back only when
the APPEND was refused. A refused PUBLISH left it set, so an identical replay
short-circuited and the revision was never published again. The repo's own
pattern is the opposite: `codex-structured-item-streams.ts` advances
`checkpointLengths` only once the append AND the publish are both accepted.
Roll back on either half.

That alone did not cover the sweep, which is the LAST event a group ever gets:
its `changed` guard skips the write on a retry because every child has already
latched, stranding the settled roster's final revision. Write when the previous
attempt was refused part-way, too.

`formatWorkerTranscriptMessage` read `block.agents` as its exhaustive fallback.
The journal schema deliberately admits block types this build does not know and
`client.call` casts the RPC result instead of validating it, so a newer remote
host's block reached that line and threw `agents is not iterable`, taking down
the whole `worker read`. It printed a harmless `[image omitted]` before. Match
`subagent-group` explicitly and degrade the unknown case.

The elapsed clock measured to `now` whenever no child carried a terminal
timestamp. That is exactly the roster restored from the journal after the host
died: the reconciler latches `unverifiable` without a `settledAt`, so a child
that ran four seconds reported the time since the crash as its run length, on a
row that is not even counting. Show no duration when none is known.

Also restores package.json to origin/main: the merge had deleted one of main's
two duplicate `bench:terminal-partial-escape-tail` keys. Behaviour-preserving
(JSON is last-wins and the deleted line was the dead one), but unrelated to this
PR and better left to its own change. No gate rejects duplicate JSON keys.

The new refusal tests also cover the append-side rollback, which had none.

* fix(native-chat): stop the subagent roster vanishing from every settled turn

`NativeChatToolRun` bailed out for a completed turn whose activity disclosure
is collapsed before it reached the branch that draws a roster-only run. That
guard exists to push TOOL activity behind the turn-status disclosure, and it
fires on exactly the shape a spawn group has: a roster message carries no tool
blocks, so `selectActiveToolCall` returns null and `isSettled` is true, while
the list passes `expandOverride={expandedTurnIds.has(turnKey)}` — false until
the reader opens that turn — and `activeTurnIsWorking={false}`.

That is the default state of every finished turn in the transcript, so the one
compact row this feature exists to leave behind ("Ran 3 subagents") disappeared
the moment its turn ended. Worse, `MessageRow` counts a spawn group as
renderable specifically so the row survives, then rendered a wrapper around a
component that returned null — the empty ghost bubble its own guard is written
to prevent.

Order the roster branch before the disclosure guard. A roster has no tool
activity to hide, and the guard's reasoning ("a failed child command looked
like the whole response was still running") does not reach it. Runs that do
carry tool blocks still fall through to the guard unchanged, and in practice a
roster never shares a message with them: it is its own `role: 'system'` journal
row and `isToolOnlyMessage` is false for it, so `foldToolMessages` never merges
tool blocks into it.

Also drop childless groups when building the rows, so `subagentRows.length`
stays an honest test of "something will draw" — the roster-only branch returns
a margin-bearing wrapper on the strength of it, and a group with no children
renders null.

Both tests fail with their fix reverted; the existing NativeChatToolRun suite
still passes, so the completed-turn disclosure behaviour is unchanged.

* test(native-chat): cover the subagent roster at the message-list level

Every defect this feature has shipped so far lived in the assembly between
rows, and the row-level suites kept passing through all of them. Loop 4's
regression — a settled roster swallowed by the completed-turn disclosure —
was found by reading the code, not by a test, and an independent visual-proof
run observed the same symptom in the real UI and routed around it rather than
reporting it. `NativeChatToolRun` rendered alone is handed `expandOverride`
and `activeTurnIsWorking` by the test author, so it agrees with whatever the
caller was assumed to pass.

Drive the real component instead. The roster is its own `role: 'system'`
journal row carrying the producer's two blocks (structured + plain-text twin),
so what reaches the DOM depends on `foldToolMessages`, the turn-key mapping
and the disclosure state `NativeChatMessageList` owns — none of which a row
test exercises.

Three cases, on one assembled transcript that holds tool calls AND a roster:
  - a settled turn with activity collapsed, the resting state of the whole
    transcript, still shows the row (fails with loop 4's reorder reverted);
  - tool activity stays behind that disclosure and appears only on expand,
    and expanding draws no second roster (fails with the guard removed);
  - a working turn reads as a live spawn.

The first also pins that the plain-text twin is dropped rather than printed
beside the row it stands in for.

Timestamps are explicit and ascending: the list re-sorts by (timestamp, id),
so rows sharing a millisecond tie-break alphabetically and the user turn can
sort last, stranding the roster outside its own turn and reconciling live
children to `unverifiable`.

No production code changed.

* fix(native-chat): make "counts as renderable" and "actually draws" agree for a spawn group

`MessageRow` counts any `subagent-group` block as renderable, but
`NativeChatSubagentRun` renders null for a childless roster. A group with
`agents: []` therefore mounted a row that drew nothing — an empty div that still
costs the transcript one `gap-5` slot. The Codex producer never writes one (every
`write()` call site operates on a group that already holds an entry), but the
block schema admits `agents: []` with no `.min(1)`, and the wire is where such a
shape would arrive.

Narrow `subagentGroupBlocks` — whose only production caller IS that renderable
check — to the groups that will draw, behind a named `isRenderableSubagentGroup`
that `NativeChatToolRun` now shares in place of its own copy of the predicate, so
the two guards cannot drift apart again. A childless group carrying its
plain-text twin now prints the twin, which is what the twin is for; a bare one
skips the row entirely.

Also correct four comments that had stopped describing the code:
  - the roster header called `agentsStates` "always empty", contradicting the
    probe note in `codex-subagent-activity.ts` — it is empty on the MultiAgentV2
    path that emits these items, and the V1 path does populate it;
  - `tokensByThread` was documented "retained UNCONDITIONALLY" while
    `handleTokenUsage` LRU-caps it 65 lines below;
  - the sweep is not "the LAST event a group ever gets": neither `settleTurn` nor
    `settleSession` removes the group, so a later `thread/tokenUsage/updated`
    naming a swept child still writes it. The retry condition is right; only its
    stated reason was wrong;
  - the `subAgentActivity` classification is not reached "for every event — and
    every one of them arrives twice". `handleSubagentItem` intercepts those items
    before `items.handle`, so the live path never consults the catalog;
    `restoreThread` replays them straight through, and is the real consumer.

Comment-only apart from the childless-group guard.

* fix(cli): stop `worker read` printing the subagent roster sentence twice

The producer ALWAYS writes a roster block beside a plain-text twin carrying the
same sentence, for clients that cannot draw the block. The renderer honours that
contract from one side — it draws the block and drops the twin. The CLI honoured
neither side: it printed the twin as prose AND rendered the block as
`[subagents] <same sentence>`, so a real roster message read

    [system] Ran 2 subagents (1 failed)
    [subagents] Ran 2 subagents (1 failed)

Take the mirror of the renderer's rule, which is the cleaner half for a text
client: the twin IS the sentence, so print it and drop the block it stands in
for. A block that arrives WITHOUT its twin — a shape the wire admits and no
producer writes — still stands in for itself, because dropping it
unconditionally would lose the roster entirely. Either way the sentence prints
exactly once, off the same `subagentGroupFallbackText` helper both sides use.

Unreachable through `readWorkerTranscript` today, whose provider rollout decoder
never emits a `subagent-group` block — but the formatter is the CLI's contract
for any transcript source, and the shape is already producible.

The test pinned a TWIN-LESS group, a body `codexSubagentGroupBody` never writes:
it asserted the exact double-print this fixes was correct output, and would have
blessed either behaviour. Rebuild the fixture as the producer's real two-block
row, with the sentence taken from the shared helper rather than hardcoded so it
cannot drift, and assert the sentence appears exactly once. The twin-less shape
keeps a test of its own, labelled as the wire-only fallback it is.

Also record why `settleTurn` keys on the RAW `turnId` while `groupFor` remaps
off-primary activity onto the primary's active turn. The asymmetry is
load-bearing, not an oversight: were `settleTurn` to remap, a child thread
ending its own turn would sweep the parent group and settle every still-working
sibling to `unverifiable`. The lookup missing is the intended no-op.

* fix(native-chat): add the subagent roster's localization keys and narrow its twin filters

The roster row called 16 `components.native-chat.subagents.*` keys that were
never added to the catalog, failing the localization gate. Synced en.json; the
English strings are the component's own inline fallbacks, so nothing renders
differently.

Also tightens the twin/block handoff on both readers. The renderer dropped
every text block once a roster was present, which is safe only because Codex
writes a roster as its own message — the block is provider-agnostic, so a lane
folding prose in beside one would have lost it on desktop while mobile kept it.
And both readers decided "the twin is already printing" by recomputing the
sentence and comparing bytes, which a roster from a newer build never matches:
its unknown state normalizes to `unverifiable` here, so the CLI printed the
roster twice with two different verdicts. Both now recognize a twin by shape.

* test(native-chat): pin the roster twin recognizer against prose

Both readers use it to decide the twin is already printing, so a false positive
eats a message's real prose and a false negative prints the roster twice.

* docs(codex): restore the roster's evictionated trigger to its KNOWN LIMITATION

The previous rewrite dropped both triggers the old comment named and kept only
the restart one, but eviction is the reachable half: `groupFor` caps `groups` at
MAX_CODEX_SUBAGENT_GROUPS and drops the oldest-INSERTED entry (it returns an
existing group without re-inserting, so this is not LRU), which can evict a
still-live group in-process. The row identity is keyed on the group id alone, so
the next activity item rebuilds that row from one child — the same N-to-1
rewrite, with no restart, and with the sweep skipped so the children never latch
`unverifiable`. Also softens "every real turn id is freshly minted" to the
provider assumption it is: turn ids are read verbatim off provider frames and
nothing in this repo mints or asserts them.

* docs(codex): justify the subagent wire notes from the live probe alone

The roster and disposition comments explained themselves in terms of a
provider-internal path taxonomy rather than anything this repo can observe.
Restate them from the evidence Orca actually has: the live app-server probe
saw `agentsStates` arrive empty, so nothing reads it; and `collabAgentToolCall`
stays substantive because nothing guarantees a session reports subagent work as
`subAgentActivity` at all — one that only emits the collab tool call gets no
roster row, and suppressing that too would leave its fan-out blank.

Same behaviour, same tests; comments and one test name only.

* fix(native-chat): stop the roster's durable twin from claiming live subagents

The spawn-group row is written once and revised in place, but the row itself
is durable and replayed on every reconnect. Its plain-text twin — the only
thing a client that cannot draw the block ever sees — froze a live count into
that row: `Kicked off 4 subagents — 2 working`. The desktop renderer never
shows it, and reconciles the block's `working` to `unverifiable` outside the
live turn. A text-only reader does neither. When the writing process dies
mid-flight the turn-end sweep never runs, so the sentence keeps asserting two
running children forever, with nothing left that could re-check them. That is
the collapse `docs/reference/ssh-execution-boundary.md` forbids: loss of
contact reported as a live state.

Fix it at the source rather than per client: the durable sentence now states
only what survives its process — that the group was spawned, plus whatever
outcome had latched. `Kicked off` vs `Ran` stays, because it reports whether an
outcome was recorded at write time; saying `Ran` while children were in flight
would assert they exited, the same error inverted. The adverse count stays so a
failing fan-out still reads as failing. Reconciliation stays in the renderer,
where the block still needs it.

The twin recognizer keeps matching the legacy `— N working` shape: journals
already hold those sentences and their rows replay forever, so dropping the
branch would print every one of them twice, once as the block and once as prose
the reader meant to drop.

Also align the two functions that read `agentPath`. The root check compared the
raw string while the label normalized separators, so `/root/` was both the turn
itself and a child of it — a phantom row labelled `root` inflating the group by
one. Compare normalized segments instead, keeping `/morpheus` a child. And a
trailing segment with nothing visible in it survives the empty-segment filter
and would draw a nameless row, so it now reads as no label and falls back to the
placeholder.

* fix(codex): key the subagent label collision ordinal on what the row draws

`codexSubagentLabel` tested the trailing segment trimmed but returned it
untrimmed, and `claimLabel` keys its collision ordinal on that string. Two
children at `/root/read` and `/root/ read ` therefore both drew as `read`
with no ordinal — the one thing the ordinal exists to prevent. Return the
trimmed segment so labels that render identically collide.

Also correct the legacy-clause note on the twin recognizer. It claimed shipped
journals hold the old `— N working` sentence; the feature is unreleased, so the
only journals holding one are dev worktrees of this branch. The branch still
earns its place — those rows replay too, and it adds no false-positive surface
the bare shape does not already carry — but the stated reason was wrong.

* test(native-chat): retire the subagent-visibility guards now the roster renders

Two tests from the sibling item-coverage PR asserted that subagent items stay
on the generic gray row, explicitly gated on "until a real renderer exists".
This branch is that renderer, so both guards fire on merge — the handoff they
were written to mark rather than a regression.

They now pin the other side of it: subAgentActivity is suppressed because the
spawn-group roster renders it, and collabAgentToolCall deliberately stays
visible, since nothing guarantees a session reports subagent work as
subAgentActivity at all.

Git merged both files without conflict; only running the suite surfaced this.

* fix(native-chat): let a subagent swept at turn end still report what it did

The turn-end sweep marks still-running children `unverifiable`, and the
producer latched on any state that was not `working` — so `unverifiable`
latched too. A subagent that outlived its turn then reported `completed`, the
latch refused it, and a child that finished successfully read as one we never
saw finish, permanently.

One predicate was doing two jobs. `isTerminalSubagentState` is right for
counting — `unverifiable` is not working — and wrong for latching, because
`unverifiable` records that we stopped being able to see the child, not what
it did. Split them: a child's own verdict latches, the sweep's guess does not.

The reverse stays refused. Nothing returns to `working` once we have given up
on it, so a straggler progress tick cannot re-light a settled row.

Neither the latch nor the sweep was wrong alone, and both were tested; the
defect lived only in their interaction, and only when a subagent outlives its
turn — which the probe that drove this design never produced, because the
parent it captured waited on its child.

* fix: drop the @pnpm/exe lockfile drift a merge staged

`git add -A` swept up the pnpm-lock.yaml mutation that every pnpm invocation
leaves in this repo. Nineteen lines, thirteen of them @pnpm/exe, and it fails
sixteen unrelated CI checks — native smoke, typecheck, packaging, xterm patch
sync — none of which name the lockfile.

* fix(native-chat): restore the item fall-through an inline dropped

Inlining the subagent routing helper lost its null check: the roster returning
null means it did not claim the item, and the translator must keep looking.
Returning unconditionally once any thread item parsed swallowed every ordinary
item — twelve settlement tests, none of them about subagents.

* fix(orchestration): rebind the subagent block arm to the renamed bound state

Main renamed clipMetadata's second parameter from a warnings set to a
TranscriptBoundState. The subagent-group arm still passed `warnings`, and git
merged both sides without a conflict because the lines never overlapped — the
rename and the new arm are in different hunks. Typecheck was the only thing
that could catch it, and did.

* fix(codex): publish the turn tail for a subagent item the roster claims

Main's #19055 added a `subAgentActivity` arm to the provider activity table,
which is reached only through `publishActivity`. The roster's admission returned
above that call, so every `subAgentActivity` item bypassed it and a fan-out that
reports nothing else left the turn tail stuck on the previous frame's text.

`publishActivity` already no-ops on a refused admission and on a non-primary
thread, so routing the roster's admission through it is safe.

Also corrects a docstring the frames extraction copy-pasted onto
`settleOversizedNotification`.

* fix(native-chat): bound the subagent roster on every boundary that carries it

The spawn-group arm was the one collection in the worker-transcript payload with
no cap, and the one block type mobile's `sanitizeBlock` forwarded verbatim. The
producer's `MAX_CODEX_SUBAGENTS_PER_GROUP` does not reach either boundary: the
journal schema declares no maximum on `agents`, and a remote host may run a build
with a different cap. Both transports now cap the roster and bound `id`, `label`
and the open `state` string; `label` and `id` also take the standard inline bound
on the journal write path, where every other provider string already does.

A token count is now persisted onto its entry at write time. `write` rebuilt
`tokens` from the LRU-capped thread map on every write, so an eviction silently
retracted a count the durable row had already shown.

Adds the first coverage of the three roster caps, including the group eviction
that rewrites a row from N children down to one.

* fix(native-chat): keep the roster drawn beside tool calls and its clock honest

The roster-only escape is keyed on `blocks.length === 0`, so a spawn group
sharing its message with tool-call blocks fell through to the settled-turn guard,
which returned bare null and took the roster with it — the exact regression the
escape above was written to avoid, after the message row had already counted the
group as renderable. Unreachable for Codex today; the block type is deliberately
provider-agnostic, so it is live for the Claude lane.

The elapsed clock also froze at a sibling's timestamp on a partial sweep: in a
group where one child completed and another is unaccounted for, the ended turn
left `working === 0` with the completed child's `settledAt`, and the row showed
that child's duration as the group's run length. No clock is drawn while any
child is `unverifiable` with no terminal timestamp.

* perf(native-chat): bound the roster's provider strings without digesting them

`boundInlineText` computes a sha256 and a Buffer BEFORE it checks the length,
so the roster paid two digests per child on every write even when nothing was
truncated — and `write()` runs on every claimed activity item (each delivered
twice) and again from `handleTokenUsage`, which streams. A same-process A/B over
a 64-child group: 76.5 us/write before, 2.0 us/write after (plain, unbounded row
is 1.2 us).

The cap changes with the mechanism. 16 KB is the tool-output bound; both readers
of this row already clip the same fields to 512, so the producer was admitting
~2 MB per durable roster row for consumers to throw ~97% of away. One
`MAX_SUBAGENT_FIELD_CHARS` now serves the producer and both readers, and the
marker is an ellipsis rather than the tool-output truncation sentence — `id` is
the roster key and the renderer's React key.

Also raises the orchestration arm's per-group bound from 20 to the producer's
64, matching the mobile arm: a 21-64 child group is routinely producible here,
so that arm clipped children and warned while its sibling clipped none. The
slice and warning stay as the transport's own defence against a remote host with
a larger cap.

* fix(orchestration): suppress one roster block per twin, not all of them

`hasTwin` was a single boolean over the whole message, so a message carrying two
`subagent-group` blocks and one plain-text twin printed one sentence and dropped
the second roster with no marker. Count the twins and claim one per group
instead. Not reachable from this branch's producer, which writes one group per
journal item, but the surrounding reasoning is explicitly about wire shapes the
producer never writes and this is the adjacent one it missed.

* fix(native-chat): loop-3 fixes to the Codex subagent worklog

Five defects loop 2's own fixes introduced.

Twin claiming was order-blind: the count-based claim silenced whichever
roster block came first, so a lone twin belonging to a LATER group erased
an earlier group's roster and printed the later sentence twice. Exact-text
claims are now settled for every group before any leftover twin is claimed
by position; the positional fallback stays for a newer build's frozen twin,
which can never equal a recomputed sentence.

`boundSubagentField` sliced UTF-16 units and could leave a lone high
surrogate in a durable row, and the clip removed exactly the tail that told
two children apart — `id` is the renderer's React key and `claimLabel`
writes its repeat ordinal at the end. It now backs off a split pair and
reserves the child index inside the bound, so both readers' re-clip cannot
cut the disambiguator off again.

`MAX_SUBAGENT_FIELD_CHARS`'s doc claimed a `groupId` bound the producer
never applies; the doc now says so and why. The worker-transcript metadata
cap is a separate literal again: it governs message ids, turn ids, tool-call
names and image urls, so a roster-motivated change must not move it.

* fix(native-chat): never infer a lost subagent from a turn boundary

QA drove a real Codex session with three live `spawn_agent` children and sent
a mid-turn correction. The roster row immediately read "Ran 3 subagents /
3 unverifiable" with no clock, while all three were still running — they
reported `completed` 57-87s after that turn ended.

Both sites rested on the same false premise: that a turn ending means no
event will ever settle a child. Children outlive their turn and keep
reporting into the same group.

- Renderer: drop `reconcileSubagentRoster`. Nothing plumbed to the component
  distinguishes a row written by a dead host from a turn that merely ended —
  journal render items carry no epoch, and a new epoch deletes the rows of the
  one it supersedes — so the row now draws the state the journal recorded.
  Under-claiming beats over-claiming.
- Main: stop sweeping on `turn/completed`. That sweep wrote `unverifiable`
  into the DURABLE journal, which mobile reads with no reconciliation.
  `turn/completed` is Codex's only turn-end notification, so an abort cannot
  be told apart from a clean finish; the safe default is not to sweep.

`settleSession` — the provider actually being gone — is unchanged and is now
the only sweep. `unverifiable` stays non-latching so a late verdict still lands.

* test(native-chat): pin the roster at the seam the QA defect came from

The mid-turn correction opens a new turn, so the fan-out's row stops being
the current turn and the list hands the roster `activeTurnIsWorking={false}`.
Asserted through the list, not the component, because that prop is what
carried the wrong claim.

* fix(native-chat): settle a roster the dying host never got to sweep

`settleSession` only fires when the provider goes away while this process is
alive. If the host itself dies, nothing sweeps and nothing reconciles on
restore, so a `subagent-group` row persisted as `working` claimed live children
forever — the mirror of the defect the previous commit fixed, and the same
`ssh-execution-boundary.md` violation in the other direction.

Reconciled host-side, at journal open, not in the renderer: mobile shows only
the durable text twin and reconciles nothing, so a renderer-only fix would
leave it claiming live children indefinitely. Opening the journal is also the
one moment a host can honestly say the previous writer is gone.

- `staleSubagentRosterRevisions` rewrites every child still reading `working`
  to `unverifiable` and regenerates the twin from the same summary, so the
  block and the sentence cannot disagree.
- No terminal timestamp: the child stopped being observable at an unknown
  moment, and stamping the reopen would report the downtime as its run length.
- Revises in place under the parsed identity, so a reopen upserts the row
  rather than appending a duplicate, and a second reopen writes nothing.
- Skipped on a corrupt load: that journal is still owed a rebuild from provider
  history, and content past the repair's free sequence retires the demand.

Reconciles journal ROWS, not roster state — the producer's in-process group map
is untouched, so the roster's known seeding limitation is unchanged, as is
`canReplaceSubagentState`: `unverifiable` still does not latch.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-07 09:38:16 -07:00
a899f92402 feat(windows): enable structured Codex chat on native Windows (#18519)
* feat(native-chat): enable Windows structured sessions

* fix(codex): prove native Windows process identity

* style(codex): format Windows session seam

* fix Windows structured Codex admission

* fix(windows): reprobe missing process identity capability

* fix(windows): decide folder-workspace WSL routing before the click

Review found pathUsesWslUnc exported but unused, and the folder composer
hardcoding worktreeUsesWslPath:false. Together those meant a folder picked
under a \\wsl.localhost\ parent routed to structured chat, then got refused
by the host and fell back AFTER the click -- which defeats the lane's own
design goal that create cannot fail after the click.

The group's parentPath is in scope at submit and the workspace is created
under it, so the parent decides WSL-ness pre-click. Wires pathUsesWslUnc
there and adds tests for the helper, including the unhydrated-store case
that previously threw.

* fix(windows): collapse the gate derivation to one call, restoring max-lines

CI static analysis failed: launch-agent-in-new-tab.ts crossed the 300-line
oxlint ceiling. Adding a max-lines disable is forbidden, so the two gate
derivations collapse into one readWindowsStructuredGateInputs() call --
a store-backed site now adds one line and one import name instead of two.
Better shape anyway: one derivation entry point rather than two reads a
call site must remember to pair.

* fix(windows): engage the legacy fallback when the host THROWS a refusal

Review found a P1 this merge composes: neither parent could reach it. At the
lane head the only structured entry was launch-agent-in-new-tab (full
store-backed WSL check); on main all win32 was refused. The merge enables
win32 in creation flows that pass no projectRuntime, so a WSL folder
workspace, a WSL-configured repo, or a repair-required runtime now routes
structured -- and the host refuses correctly, but by THROWING rather than
returning {ok:false, refusal}.

Callers engage their legacy-terminal fallback on the refusal CLASS, so an
unmapped throw arrives as a generic RPC rejection: no fallback, empty
workspace, error toast, prompt stranded in the launch outbox. Pre-merge the
same action opened a legacy terminal agent.

Map the host's thrown definitive refusals onto the refusal class at the
launch boundary, so every creation flow -- present and future -- degrades to
the legacy terminal instead of stranding. Narrow predicate: unrelated
failures (ECONNRESET, empty message, non-Error) still propagate untouched.

Ablation-proven: removing the mapping reddens the fallback test.

* fix(windows): teach the mobile RPC double the status probe the lane added

CI's first-ever run on this lane caught a pre-existing lane defect. The lane
changed status.get to resolve through
runtime.getStatusAfterWindowsProcessStartTimeProbe(), but never taught the
mobile-surface runtime double about it, so status.get failed for mobile
clients with "not a function". The lane's own test list did not include this
file and the lane had zero CI, so nothing ever ran it.

The real runtime always implements the method; the double omitted it.

* chore: merge current main and regenerate the localization runtime catalog

CI static analysis failed on a stale en-runtime-required.json: main added
onboarding integration-capability keys, and the generated catalog is checked
against the PR MERGE result, not the branch alone -- so it read clean locally
while failing in CI. Merging current main (90780acb85) and regenerating.

Gates after the merge: pnpm tc 0, oxlint 0, changed-code quality 0/56,
7 gate/lane test files 69 tests green.

* fix: route structured launches by execution host platform

* fix: recover paired structured session mirror on host swap

* Revert "fix: recover paired structured session mirror on host swap"

This reverts commit 81bfca0007.

* Revert "fix: route structured launches by execution host platform"

This reverts commit 47abbd354a.

* fix(windows): refuse structured chat in a paired web client

Reverts the two review-loop commits (restoring a tree byte-identical to the
validated head) and closes the hole they were aiming at, without their cost.

A paired web client's `platform` describes the browser's machine, not the host
that will run the agent, so the Windows gate cannot be evaluated there. Before
this, a browser on macOS driving a Windows runtime read "not win32", skipped the
creation-time proof entirely and allowed structured chat — fail-OPEN, the
dangerous direction, bypassing the guarantee this lane is built on.

`isWebClient` is a required input like the other gate fields, so the compiler
enumerated all seven call sites. Refusal is synchronous and fail-closed: no
async round-trip, no null window, no cache to invalidate — unlike keying on an
asynchronously-fetched host platform, which would have made every desktop
launch wait on a round-trip to fix a paired-web-only hole.

Paired web therefore gets the legacy chat until the host publishes eligibility
itself; that is the proper fix and belongs in its own PR.

Ablation-proven: removing the guard reddens both refusal tests; the
desktop-unaffected test is a preservation check and passes either way.
Gates: tc 0, oxlint 0.

Known open: repos-onboarding-folder-startup.test.ts fails on this branch and
passes on plain main — under investigation, NOT caused by this commit.

* test(onboarding): mock the web-client check the store path now reaches

The web-client refusal added `isWebClientLocation()` to the launch-route
inputs, which this suite's store path reaches while adding the FIRST folder.
The suite stubs `window` as `{ api }` with no `location`, so the function
cleared its `typeof window === 'undefined'` guard and then threw on
`window.location.pathname`.

That threw inside addNonGitFolder's own catch, so folder-1 never activated;
folder-2 then returned early (a project already existed) before reaching the
call at all, leaving exactly one activation with no startup seed.

Test artifact, not a product defect: a real renderer always has
`window.location`, so the seeding path is intact for users. Mocking the module
is the convention 7 other suites already use, and keeps product code free of
defensive branches that only exist to satisfy a stub.

Ablation-proven: removing the mock reproduces the original failure exactly.

* fix(renderer): make the web-client check total over a partial window

isWebClientLocation() guarded `typeof window === 'undefined'` and then assumed
`window.location` existed. A window stubbed without a location cleared the
guard and threw on `.pathname`.

That matters because this branch put the call on the launch-routing path,
where the throw is swallowed by the caller's catch and silently becomes a
FAILED LAUNCH rather than a visible error. CI caught it as 9 failures in
launch-work-item-direct.test.ts.

I previously "fixed" this by mocking the module in the one suite I knew about.
That was whack-a-mole against an unbounded set, and it missed this one. The
defect is the partial-window assumption, so fix it there: the mock is removed
from the onboarding suite and both suites now pass on the hardening alone.

Ablation-proven: reverting to the unguarded form reddens 11 tests across the
new unit suite and launch-work-item-direct.

Gates: tc 0, oxlint 0, changed-code quality 0/58.

* Move Codex's Windows structured-chat eligibility onto the host createSupport probe

The renderer no longer decides Codex win32 eligibility: launchStructuredAgentSession
probes agentSession.createSupport for both providers, the host answers via
supportsCodexStructuredLocation (process start-time proof + WSL refusal), and the
create path re-checks live. Deletes the client-side windows gate module and its
routing inputs (windowsProcessStartTime, worktreeUsesWslPath, isWebClient, platform)
from six call sites. Splits killCodexAppServerProcessTree out of
codex-app-server-session to hold the max-lines ceiling without a disable.

* fix(ci): keep pnpm lockfile stable

* test(windows): align foreground snapshot flags

* Restore main's pane-snapshot flag contract

Main asks for CreationTime on both projections; this branch's hot-path
isolation went away with the async probe it served.

---------

Co-authored-by: Orca Worker <orca-worker@localhost>
Co-authored-by: Merge Sim <sim@local>
Co-authored-by: Merge Sim <merge@localhost>
2026-09-07 09:18:38 -07:00
Brennan BensonandMerge Sim f1d8545024 feat(chat): support structured /clear and /compact commands (#19164)
* feat(chat): support structured clear and compact commands

* fix(chat): authorize mobile commands and bound clear-chain projection

* fix(chat): localize conversation command send errors

* fix(chat): retain clear pane identity with reopened history

* test: account for combined structured session RPC additions

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-06 21:17:00 -07:00
Brennan BensonandMerge Sim 68dd3909c7 feat(orchestration): orchestrate native-born structured chat sessions (#18827)
* feat(orchestration): orchestrate native-born structured chat sessions

Orchestration resolves every worker through a terminal handle and a pane key
backed by a live PTY. A session created directly as structured has neither, so
it was not refused by orchestration — it was invisible. A coordinator could not
start one, address one, or receive `worker_done` from one.

Add a second authority source rather than a parameter channel. A registry maps a
session id to the same three facts the PTY path supplies — a bearer handle, a
pane key and a host scope — and the four runtime getters consult it before
giving up on `ptysById`. `orchestration.send` and `verifyDispatchCapability` are
untouched: authority stays host-derived and the CLI still cannot assert who it
is. PTY handles short-circuit on the handle prefix, so the terminal path is
unchanged.

Mail travels as a session turn instead of as bytes, on a sibling lane that keeps
the PTY lane's outstanding-run, waiter, reserved-type and batch rules.
Orchestration's database stays the source of truth; the send is best-effort,
exactly as the byte write is, and mail is consumed only on a proven-accepted
dispatch. Delivery waits for the session to be between turns, because one
provider refuses a mid-turn start outright and the other cannot acknowledge one
inside the ack window.

Security properties, each pinned by test: the pane key's leaf is random and
persisted rather than derived, since `check` is identity-gated and accepts a
caller-supplied pane key; the handle is a random bearer token; the child env
carries no pane key, which would otherwise flow into hook pipelines that assume
a PTY leaf; hook attestation stays closed for structured handles; and process
continuity comes from record lineage, never the runtime fence, which the host
bumps during its own crash recovery.

Also remove the "Orchestration paused" notice, which gated only on dispatch
status and rendered over bridge chat where orchestration always worked; refuse
the implicit-sender fallback when a worktree has more than one candidate leaf
instead of guessing; and collapse the archive kinds to one named type with a
compile-time assertion that the capture set cannot drift ahead of the storable
set.

* fix(orchestration): answer the structured idle gate from the reduced timeline

The structured pointer gate read a bounded 40-item tail page. A settled turn is
tombstoned rather than rewritten, so an idle worker with any real history carries
no turnLifecycle item at all and the "full page, no lifecycle item" guard read it
as busy forever: every nudge after the worker's first substantial turn parked on a
settle edge that had already passed, and the preamble tells workers not to poll.
The attention gate had the mirror bug — a prompt older than the tail window was
missed and the nudge was delivered into a session blocked on a human.

Both facts now come from `journal.snapshot()`, the fully reduced timeline, via a
new narrow `readGateFacts` host read; the policy module stays pure and still
projects through the shared helpers the chat view reads.

Also:
- Park `session-not-attached` on the journal edge, so mail that arrives during a
  transient detach is redriven by the re-attach reset instead of sitting unread.
- Resolve a structured worker's provider from the durable agent-session record
  when the registry entry was rehydrated, so a restarted Codex worker is no longer
  reported and archived as Claude.
- Clear `structured_pointer_operations` in every `orchestration reset` scope.
- Drop the per-chat-pane dispatch-status store subscription left behind by the
  removed paused notice, and re-pin the two terminal-pane ratchets it moves.
- Hoist the identical pointer batch selection out of both delivery lanes into
  `selectOrchestrationPointerBatch`.
- Refuse the pre-graph-ready focus-based guess for `requireUnambiguous` callers,
  matching the ready path.
- Move the host teardown phase list into the teardown module it belongs to, which
  is what keeps the host inside its max-lines budget.

* fix(orchestration): discard a structured worker session whose create settled unknown

`commitStructuredAgentSessionCreate` answers `agent_session_operation_unknown` when
`attach` SUCCEEDED and only the tab publish failed, so `created.ok === false` is not
proof that nothing exists. The worker start read it that way and skipped
`discardCreatedSession`, leaving a live provider child that took no hold, has no
`bindingsByDispatchId` entry and no published tab — the outer
`releaseStructuredWorkerSession` no-ops without a binding, and a session that never had
a holder never starts the eviction clock, so nothing in the runtime ever retires it. A
throw out of the commit half is past `attach` for the same reason; the pre-commit half
refuses rather than throwing. Cleanup now asks whether the create MAY have committed,
via the existing `isDefinitiveAgentSessionCreateRefusal` predicate.

Also:
- Strengthen the pre-ready `requireUnambiguous` test so it actually pins the guard: the
  snapshot now carries a focused terminal, so deleting the `? [] :` ternary turns the
  test red instead of leaving the refusal to the ambiguous `listTerminals` fallback.
- Correct the guard's justification comment, which cited `orchestration check` as
  covered. `check` resolves through the `--terminal` scope and still guesses; the
  guard covers the implicit `--from` sender, and a structured worker is covered by the
  `ORCA_TERMINAL_HANDLE` baked into its child.

* docs(orchestration): stop two structured-worker comments claiming guarantees the code does not give

The send-time owner re-check reads `target.refusal`, the snapshot the resolver
already admitted, so `decideStructuredPointerDelivery` can only agree with the
resolve-time answer and `owner-not-settled-native` is unreachable from that call
site. What actually fences an owner that moved is `expectedRuntimeFence`, which a
handoff bumps. Say that, so nobody later drops the fence trusting a re-check that
is structurally a tautology.

`discardCreatedSession` was credited with retiring "a published background tab
that no dispatch owns". It hides the DURABLE tab reference and closes the
session; the live tab snapshot keeps the row, so the background tab this start
published stays on screen until the app restarts. Same for stop and release. The
comment now describes what the two calls do — including that both are no-ops on a
session that was never attached, which is what makes the non-definitive-refusal
path safe to reach unconditionally.

* fix(orchestration): retire a structured worker's chat tab when the worker settles

Starting a structured worker always publishes a real `agent-session:<id>` tab, but
every settlement path only called `setSessionTabVisibility(sessionId, false)` plus
`host.close(sessionId)`. That clears the DURABLE restore index and leaves the LIVE
snapshot untouched, so stop, release and the half-started discard all left a dead
"Claude Chat" / "Codex Chat" tab in the worktree's tab bar for the rest of the app
session — five dispatches, five dead tabs — and opening one re-attached the released
session, respawning a provider child outside orchestration's hold accounting.

The snapshot-pruning half of `closeStructuredAgentSessionTab` is extracted into
`structured-agent-session-tab-retirement.ts` and exposed on the runtime as
`retireStructuredAgentSessionTabFromSnapshot`, so the user-initiated tab close and
the three settlements share one implementation instead of a second copy.

The settlement side is best-effort BY CONSTRUCTION: it runs only after the close is
already proven, calls the runtime method optionally, and swallows any throw. It
talks to no renderer, so the startup release reconciler can call it too. Nothing
here can turn a proven stop into `release_unknown`.

* fix(orchestration): stop a structured worker's nudges, archive and liveness from lying

Five defects in the structured-worker lanes, each with the same shape: a check
that answered from something other than what it claimed to measure.

- The pointer lane gated a WORKER's `dispatch:` mailbox on its RUN's outstanding
  delivery. Delivery rows exist only for a `run:` address, so that row belongs to
  the coordinator — and a coordinator holds one for exactly as long as it is
  acting on received mail, which is when it replies to its workers. The gate is
  gone; there is no coordinator mailbox in this lane to protect.
- `dispatch-rejected` now parks on the journal edge. A rejection consumes no mail
  and nothing else redrives the mailbox, so an unparked pointer left the worker
  idle on durable mail until unrelated mail happened to arrive.
- The released journal archive bounded forward — keeping the HEAD — before
  capping newest-first, so a long worker's archive ended at its early exploration
  and dropped the answer it was released for, under a warning that said the
  oldest messages had gone. One newest-first pass now, and the warning is true.
- The durable pointer operation id was reused on a matching BODY fingerprint, and
  the body names only the unread count. Two unrelated same-size batches collided,
  the host replayed its ledger answer as `accepted` with no turn sent, and the
  lane marked the new mail delivered. Reuse is keyed on the batch's message ids.
- `worker-read` on a structured worker hardcoded `terminal: 'running'` and
  emitted no `liveness`, so a runtime that could not see the session reported the
  worker as alive. It now carries the observed verdict, as the PTY branch does.

Also: the live journal cursor is an index into a re-derived tail window, so the
page's oldest item joins its source identity — a slid window now answers
`source_changed` instead of silently resuming past the items it skipped. And a
stop that reached no host reports `processAction: 'none'`, after installing the
host the way release already does.

* fix(orchestration): stop a released structured archive claiming a close that never landed

`worker-read` on a released structured worker hardcoded `liveness: 'exited'`. The
archive is frozen BEFORE the close, so it proves nothing about the provider child,
and the read is served for `release_state` in `releasing` / `unknown` too — the two
states that exist precisely to record a close that did NOT land. A coordinator that
read `exited` from a `release_unknown` worker would start a replacement over the same
worktree while the original child was still attached, which is the outcome
docs/reference/ssh-execution-boundary.md rule 2 exists to prevent, and it contradicts
the release receipt's own "the structured session close was not proven" text.

The verdict now comes from the resource row the read already holds: only a settled
`released` row is `exited`, everything else is `unverifiable` — which the existing
mapping renders as `terminal: 'unknown'`, the same way the live branch does.

* fix(orchestration): stop a structured worker-start reporting a preamble it never delivered

Two ways a structured `worker-start` handed the coordinator a receipt that did not
describe the worker it got.

`sendStructuredWorkerPreamble` threw only on a refusal and on `rejected`, so a
submission that settled `unknown` fell through as success: the start pushed
`dispatch_input: accepted` and marked the dispatch ready. `unknown` is not rare —
`dispatchSafely` converts ANY thrown adapter call (provider child gone, transport
dropped, ack window missed) into it, and `performSend` still returns ok. The worker
then has no task spec while its coordinator blocks in `check --wait --types
worker_done` until timeout. This PR's own mail lane already states the rule —
"`pending` is not yet an acknowledgement; only `accepted` may consume mail" — so the
preamble now applies it too, and raises `operation_unknown` for the states that
prove neither delivery nor failure, which is the code `failWorkerStartWithReceipt`
turns into the `outcome_unknown` receipt whose nextCommands send the coordinator to
look. `rejected` stays a proven failure.

`--structured` also accepted `--model` / `--effort` and dropped them: structured
session creation takes no launch preferences, while `launch.receipt.effective`
echoes whatever was requested either way, so `--model opus` ran on the workspace
default and the receipt still said `opus`. Refused now, for the same reason
`--terminal` refuses them, and the spec note records that refusal along with the
new-child/new-top-level one it never mentioned.

Tests: the refusal guard had no coverage at all, and `structured-mailbox-pointer-host`
— where the full-timeline gate read lives — had none either; reinstating the bounded
tail there left the whole repo green. Both are covered now, and the vacuous
"never selects an exact provider session" case is re-pointed at the absent
`ORCA_PANE_KEY` that actually keeps that selector shut.

* fix(orchestration): let a structured worker actually reach the Orca CLI, and stop four settlements lying

A structured worker's provider child runs `orca orchestration ...` exactly like a PTY worker's
agent does, but it was handed the ambient PATH. On packaged Linux the CLI installs as `orca-ide`
so it never claims GNOME Orca's /usr/bin/orca (#7904), so bare `orca` execs the screen reader and
the worker can never read mail, reply or send worker_done; on packaged macOS/Windows the bundled
launcher is only reachable from the app's own resources dir. The PTY lane already solves this
inside `buildPtyHostEnv`; that block is now its own module and both lanes call it.

Also:
- a worker start that fails AFTER its session exists now discards the session, so a failed start
  stops stranding a dead chat tab that the durable restore index republishes on every launch;
- a structured worker's resource reconciles to `released` after settlement forgot its identity,
  instead of answering `unverifiable` for the life of the DB;
- `closeAttempted` is set only once a close is issued, so a tab-visibility failure can no longer
  report `closed_agent_terminal` for a running child;
- `forgetSession` prunes only what the settled worker parked, not every sibling whose target
  momentarily fails to resolve;
- release settles with an explicitly empty, warned archive when the journal is unreadable AND the
  session is proven exited — closing the chat tab is routine, and `archive_failed` there wedged
  release on evidence that could never arrive;
- the new migration test uses mkdtemp and cleans up, so it stops failing Windows CI and leaking.

* fix(orchestration): merge the duplicated release-receipts import

The release-completion module imported ./orchestration-worker-release-receipts
twice, which trips import/no-duplicates in audit:code-quality:native. The
changed-file gate does not load that config, so only whole-tree CI saw it.

* docs(runtime): note that a background structured tab re-publish is a no-op

The activate:false branch for an already-published session returns without
writing the snapshot or emitting, so it cannot re-surface a client whose
mirror lost the tab. Orchestration is safe from this only incidentally.

* feat(orchestration): make the worker mode the user's own default, not a flag

`worker-start --structured` was an explicit opt-in that REFUSED --on, --terminal,
--model/--effort and worktree-creating placements. The flag, its spec entry and the
`structured` RPC param are gone: the mode now follows the user's setting for new agent
tabs, so a local claude/codex worker is a structured chat session whenever the user's
own default says agent tabs open as one.

A setting is a preference, not a demand, so none of those combinations refuses any more.
A dispatch that cannot be structured starts an ordinary PTY terminal worker and the
receipt names the mode that ran and why, so the fallback is never silent:

- a remote --on, an existing --terminal, a new-child/new-top-level worktree and
  --model/--effort are decided from the request;
- the agent, TUI launch customization, Codex-on-Windows and the runtime capability are
  decided by the shared launch route;
- WSL, remoteness and the Windows start-time gate are settled by the executing host's own
  agentSession.createSupport, asked once the worktree resolves and before anything is
  created, so a refusal is a terminal worker rather than a failed start.

The decision is the renderer's, lifted rather than copied: `resolveAgentLaunchRoute`'s
structured half and the settings predicate now live in
shared/structured-native-chat-launch-route, which both surfaces call, and the TUI launch
customization test moves to shared beside it. `getClientSettings` gains the two native-chat
default booleans it was missing.

No security invariant moves: the structured worker registry, bearer handle, persisted pane
key, the absence of ORCA_PANE_KEY from the child env, hook attestation and lineage-derived
process incarnation are untouched.

* fix(orchestration): stop the worker mode leaking into the agent contract

The mode a worker runs in is a runtime implementation detail. An agent should be
taught the same verbs, run the same commands and read the same receipts whether it
is a structured chat session or a PTY terminal — otherwise a settings-driven
fallback silently changes what the agent can do.

The real leak was `canDispatchSubWorkers`, which was forced false for a structured
worker. That was not a wording choice: `worker-start` resolved `--from` through
`showTerminal`, which needs a live PTY or renderer leaf, so a `structworker_`
coordinator genuinely could not dispatch. Rather than withhold the capability, the
one fact the command needs from `--from` — its worktree id — now comes from
`getOrchestrationDispatchAuthority`, the same authority the pane-key and
process-incarnation getters already answer structured handles from. Sub-dispatch is
gated on depth alone, identically for both modes.

`showTerminal` itself is deliberately NOT taught structured handles: it returns a
ptyId, a leaf id and a pane runtime id, and synthesising those for a session with no
PTY would hand every caller of a public terminal verb something that looks writable
and is not. `inspectWorkerTerminal` already returns `terminal: null` for exactly
that reason.

Also neutralised three agent-visible refusals that named the worker's kind: a
`worker-read --source terminal` on a worker with no terminal now names the sources
that do work, and both archive refusals say "transcript output" rather than
"structured chat output" (the PTY `transcript_pin` branch said "structured" too).

New tests pin both properties: the two preambles are byte-identical once the handle
and per-dispatch ids are normalised, and a structured coordinator starts a worker
with `showTerminal` rejecting.

* fix(orchestration): stop claiming a structured worker was checked for a prompt

worker-show reported observation.agentWait: null for every structured worker. The
field's own contract says null means Orca looked and found no wait, and absent means
it never looked — and nothing looks here: a structured worker parks on a journal
question item, which no terminal prompt scan can see.

So null was a false negative on the one field a coordinator is explicitly told to
read, and it was mode-dependent: the same worker as a PTY would have reported the
wait. Absent is both the honest value and a state a PTY worker already reaches (an
older host, an unreadable pane, a probe that did not answer), so it discloses
nothing about which mode ran.

* docs(cli): stop the worker-start spec pointing a caller at the worker kind

The note said "the receipt mode field names the mode used and why", which is an
instruction to read a field no verb behaves differently for — the one thing the
mode was not supposed to become. It now says what a caller actually needs: the
dispatch always starts, the options passed are the ones honoured, and every worker
is driven the same way. The receipt still carries the mode for operators and
telemetry; nothing tells an agent to look at it.

* perf(orchestration): coalesce the structured redrive edge

Every journal batch is a redrive candidate, because a settled turn is tombstoned
rather than rewritten — there is no completed row to watch for. That is free while
nothing is parked on the session, but once mail IS parked each batch re-resolved the
dispatch, queried unread mail and read the host's gate facts, only to re-park because
the turn was still running. A turn streaming tool calls paid that per batch.

The edge now coalesces on a 300ms quiet window with a 2s starvation cap, so a
streaming turn costs a handful of evaluations instead of one per batch and a settled
turn still nudges promptly. Delivery semantics are untouched: the gate, the
accepted/rejected/unknown handling and the retain rules all still run exactly as
before, just fewer times. Nor is this the path fresh mail takes to an idle worker —
that is `deliverForHandle` at enqueue time, which this does not touch — so the
common case gains no latency.

The mechanism is the session.tabs notify coalescer, generalised into
`keyed-trailing-edge-coalescer` and called by both rather than duplicated; the
session.tabs windows stay where they were, since 50ms is right for a spinner title
and far too tight for a journal stream. Disposal drops the pending timer rather than
flushing it, on the existing subscription disposer that every settlement already
reaches, so a redrive can never fire for a session no dispatch owns.

* fix(orchestration): deliver direct peer mail to a structured worker, and let a peer read it

Two agent-to-agent verbs had no answer for a worker that IS a structured agent
session, and both failed quietly.

Mail addressed to a worker's own bearer handle — how agents mail each other
outside a dispatch — fell between the lanes. The send stored durably and
reported success, `getLiveTerminalPaneKey` resolved the recipient, and then
neither lane claimed the mailbox: the structured resolver answered only
`dispatch:` addresses, and the PTY lane refuses a structured handle outright.
Nothing errored and nothing logged, so the worker never reacted and the peer
waiting on a reply hung. The resolver now also answers a bare worker handle,
preferring that worker's active dispatch so peer and coordinator nudges share
one operation-ledger budget. A worker BETWEEN dispatches is still nudged, under
a session-scoped key: a dispatch says nothing about whether delivery is safe —
the idle gate and the lease fence do — and its own `check` reads exactly the
direct mailbox the mail is sitting in. The dispatch caller key is left
byte-identical, because the ledger is keyed on (callerKey, operationId) and
reshaping it would re-mint nudges already in flight as second turns.

`terminal read` had no structured branch, so the only peer-accessible read verb
answered `terminal_handle_stale` for a live worker; `worker-read` is closed to a
peer, which holds neither coordinator standing nor a dispatch id. It now serves
the session's journal, projected to LINES and paged by the same reader the PTY
tail uses, so the result stays a plain RuntimeTerminalRead and nothing an agent
reads discloses which kind of worker answered. Bounding and dispatch-capability
redaction are the archive path's, reused rather than rebuilt. A session that is
not attached refuses with the existing not-attached code rather than returning
an empty tail, which would read as "this worker has said nothing".

`terminal.show` still refuses a structured handle. This is read-only on purpose:
synthesising a ptyId/leafId/paneRuntimeId would hand every public terminal verb
something that looks writable and is not.

* fix(orchestration): stop three PTY-only probes answering for structured sessions

Three defects, one shape: a probe that enumerates PTYs or resolves a pane was
standing in for a question that is not about panes at all.

`worktree rm` destroyed a live structured worker. `killAllProcessesForWorktree`
sweeps the renderer graph, the provider session list and the local pty-registry,
and a structured session is registered on none of them — so all three counted
zero, nothing errored, and removal deleted the checkout out from under a running
provider child, which kept running with its `cwd` gone while the dispatch still
reported the worker live and exact. A fourth sweep now asks what the other three
cannot: membership by `location.workspaceId`, which covers a plain chat session
as well as a dispatched worker, and liveness by the same
`live`/`unverifiable`/`exited` observation the rest of the structured surface
uses. It REFUSES a destructive removal rather than auto-closing, on the same
bargain and the same `--force` escape hatch as the unstopped-PTY gate — this is
the verb that deletes a user's work, and a running agent is exactly what they
would want to be told about. Force closes the sessions properly instead of
orphaning a child. Best-effort reconciliation callers are excluded: they repair
state, delete nothing, and must never be failed closed.

Twelve coordinator verbs failed for a structured worker running as itself.
`isLiveTerminalHandle` validated `ORCA_TERMINAL_HANDLE` with `terminal.show`, a
PTY verb whose leaf lookup misses for a session that never had a pane; the pane
remint that would have recovered it needs `ORCA_PANE_KEY`, which a structured
child deliberately does not carry, so every one of them died on
`no_active_sender_terminal` — including the ones the worker's own dispatch
preamble tells it to run. The identity question gets its own probe,
`terminal.resolveIdentity`: a handle and a boolean and nothing writable.
`terminal.show` still refuses a structured handle, because synthesising
ptyId/leafId/paneRuntimeId would hand every public terminal verb something that
looks writable and is not. The PTY half is byte-for-byte today's check,
`getLiveLeafForHandle` included, so its `rendererGraphEpoch` re-check still runs
— that check is the whole reason the sender is validated at all, and a cheaper
probe would have quietly started passing stale post-reload handles. A host that
predates the method answers `method_not_found` and the client falls back to
`terminal.show`, which is correct for that host: one without the identity probe
has no structured workers to miss.

`dispatch --inject` reported `no_agent_detected` for a structured worker, because
`isTerminalRunningAgent` reaches `getLiveLeaf`, throws, and the catch returns
false. A structured session IS the agent; there is no foreground process to
recognise, so it answers before the PTY probes rather than through them.

Also: a Run whose coordinator is structured now gets its `run:` mail. Both lanes
declined and neither logged — the PTY lane because the owner is structured, the
structured lane because the mailbox was not `dispatch:` — so each half believed
the other owned it. The PTY lane's reasoning (a coordinator blocks in
`check --wait`, where a waiter preempts pointer delivery) does not transfer: a
structured coordinator is a chat session whose turn ends. Its `run:` deliveries
take the `hasOutstandingRunDelivery` gate the PTY lane applies for exactly that
mailbox, and only for that mailbox.

The test that would have caught the twelve drives the CLI with
`ORCA_TERMINAL_HANDLE=structworker_…` and no `--from`. Every existing
orchestration CLI test passes `--from` explicitly, so the resolver a real worker
goes through was never exercised — which is why the suite stayed green while the
preamble failed on its first line.

Two files crossed their line ceiling and are split rather than waived:
`worktree-teardown.ts` sheds its two PTY-surface sweeps and the deadline
arithmetic they share, and `orchestration.test.ts` — which sat exactly on 800 —
sheds the two caller-identity suites this change rewrote.

* fix(orchestration): arm the takeover signal for structured chat input

`worker-release` closed a structured session a user had taken over, losing work
mid-conversation, while `orchestration-worker-specs.ts:106` promised "Never
closes … user-taken-over terminals".

Every guard was already correct and simply never armed.
`reportWorkerTerminalUserInput` has exactly one call site — the real-user-input
signal on a PTY connection — so structured chat input never reached
`orchestration.workerTerminalUserInput`, `markWorkerTerminalUserOwned` never ran,
ownership stayed `owned` instead of `user_owned`, `retainedReason` never returned
`user_takeover`, and `stopStructuredWorker` proceeded. The durable flag is reused
as-is rather than given a parallel mechanism: it exists precisely so a restart,
an SSH drop or a renderer remount cannot erase a takeover.

Addressed by SESSION, never by pane key. A structured worker's pane key is a
random identity credential — anyone holding it can read and consume that worker's
mailbox, and session ids are embedded in tab ids in plain text — so it stays in
main and the runtime resolves the session to it. Handing it to a renderer to echo
back would make it learnable by anyone who can see a chat pane. The RPC gains an
optional `sessionId` alongside `paneKey`; a host that predates it rejects the
call, and the report is already best-effort with a catch, so that host degrades
to exactly today's behaviour rather than failing a send.

The signal fires from the composer send hook and only past `accepted`: the outbox
dispatcher retries, and orchestration's own pointer nudges never pass through the
composer at all — so neither can be mistaken for a user takeover.

* fix(orchestration): reach structured workers through group addresses

`orca orchestration send --to @all` — and `@idle`, `@claude`, `@codex`,
`@worktree:<id>` — silently skipped every structured worker. Recipients came
from `listTerminals`, which enumerates leaves and PTYs, and a structured session
is on neither. The exclusion happened BEFORE per-recipient resolution, so the
`SendRecipientWarning` machinery never ran: the caller got exit 0 and a receipt
naming the workers that did resolve, and a broadcast "stop work" or "base moved"
reached the PTY workers and nobody else. With every worker structured it
degraded to `terminal_not_found`, which reads as "the group was empty".

Fixed at the group-resolution site rather than inside `listTerminals`. That
result is published to paired mobile and remote clients and to consumers that
assume a summary carries a `ptyId` or is writable, so widening it is its own
change under `docs/reference/remote-wire-compatibility.md`. Group addressing
reads exactly three fields off a recipient, and `RuntimeTerminalSummary` already
satisfies them structurally, so the resolver widens to that smaller shape and
nothing here invents a `worktreePath` or a `branch`. Candidates are liveness-
gated on the same observation the rest of the structured surface uses — mail
addressed to a settled worker would be stored for a lane that will never deliver
it — and once a worker IS a candidate, the existing per-recipient warnings cover
it, so an unresolvable one is reported rather than dropped.

`@idle` needed more than enumeration: `getAgentStatusForHandle` reaches a PTY
probe that throws for a handle with no pane, so a structured worker would have
been enumerated and then silently dropped from the one group address that
selects on status. It now answers from the session's journal — and off the FULL
reduced timeline, never a bounded tail. Settlement tombstones the running turn's
lifecycle item rather than rewriting it, so on any page-sized read a long
tool-calling turn looks identical to an idle session; `@idle` would then
broadcast into a running turn, which Codex answers with `turn already running`
and Claude queues behind. An unreadable session answers null, never idle.

`terminal list` and `worktree ps` still omit structured workers; that is the
wire-visible half and is deliberately not in this change.

* fix(orchestration): refuse rather than guess when a chat session has no identity

An ordinary structured chat session — not a dispatched worker — is spawned with
no `ORCA_TERMINAL_HANDLE`, because `structuredWorkerChildIdentityEnv` early-
returns for any session outside the worker registry. `orca orchestration check`
then fell through to `terminal.resolveActive`, which picks the focused tab's
active leaf or the first leaf in the worktree. It returned a valid handle, so
nothing errored — and `check` is destructive by default, so it consumed another
pane's oldest unacknowledged batch and marked it read. The rightful worker never
saw that mail.

`requireUnambiguous` does not fix this, only narrows it: it refuses when MULTIPLE
leaves could be meant, and with exactly one terminal pane in the worktree the
guess still resolves — to a sibling. "One terminal pane plus one chat tab" is a
normal layout, so the common case stayed broken. The pinned test is that case.

So the child now carries `ORCA_STRUCTURED_SESSION`, and every remaining route
that would GUESS an implicit terminal refuses on it with an error naming the flag
to pass. The marker names NOTHING — no handle, no pane key, no session id, no
token — which is the whole reason it is safe: it cannot be replayed, cannot
impersonate, and cannot flow into the hook-attestation, agent-row or
mobile-projection pipelines the way a pane key would. That makes it a different
decision from withholding `ORCA_PANE_KEY`, not a reversal of it. It also grants
no CLI reachability, so packaged builds keep exactly today's exposure.

The comment at `orca-runtime-adopt-terminal-orphans-from-inventory.ts` that
justified the guess — "a structured worker is covered instead by the
`ORCA_TERMINAL_HANDLE` its child is spawned with" — was true only for dispatched
workers and false for every other structured session, a population this branch
creates. It now says which case it covers and which case it does not.

* fix(orchestration): stop two surfaces lying about a worker with no terminal

`orca terminal <verb>` answered `terminal_handle_stale` for a structured
worker's handle. Nothing went stale: the session is live and simply has no
terminal, and it never had one — so callers acted on a false claim and went
hunting for a remint that cannot exist. The refusal now carries its own code and
names the structured equivalents (`orca terminal read`, `worker-read --source
transcript`, `orca orchestration send`), so an agent that lands there learns
what to run rather than what failed. A PTY handle that really did go stale keeps
the old error, and so does a session this runtime no longer owns — that handle
IS dead. `terminal.show` stays non-resolving: synthesising a
ptyId/leafId/paneRuntimeId would hand every public terminal verb something that
looks writable and is not.

`orchestration-worker-specs.ts` promised "the same verbs, the same handle, and
the same worker-read sources", and all three clauses were false for a worker with
no terminal. A spec agents read must not carry a false promise, so it now states
the limitation and the alternative that always works.

Note this had to be reconciled with an invariant this branch already holds: the
worker MODE must stay opaque, or a coordinator starts branching on something no
verb it runs behaves differently for. So the note says "not every worker has a
terminal" and points at `--source auto`/`--source transcript` WITHOUT naming a
kind — the same mode-neutral wording `readStructuredWorkerOutput` already uses
when it refuses `--source terminal`. Both properties are now pinned by tests, so
neither can be restored by breaking the other.

* fix(orchestration): close the review findings on the structured parity work

Four defects and two follow-ups from the delta review.

The `worktree rm` refusal was a dead end in the desktop UI. Its message matched
no matcher in `classifyWorktreeForceDeleteReason`, and an ordinary desktop delete
already passes `force=true` for the dirty-file skip, so classification returned
null unconditionally: the toast showed raw CLI wording with no Force Delete
button, and a user with a live chat session was stuck unless they knew to reach
for the CLI. That is the #11960 shape `shared/worktree/removal.ts` documents, so
the refusal now has its own prefix, matcher, `WorktreeForceDeleteReason` and
toast copy, classified BEFORE the `force` guard and nulled once the waiver is
spent — exactly how `unstopped-pty` is handled, with matcher and hint kept in
the same file as that contract requires. The copy says Force Delete will close a
running conversation rather than borrowing the "could not confirm" wording,
because Orca watched these sessions stay attached; there is no doubt to waive.

Structured `terminal read` cursors were unsound and are now refused. The PTY
cursor indexes an append-only completed-line buffer with a monotone count; a
session journal is a BOUNDED tail re-projected on every read, so a saved index
addressed different lines as the journal grew — and `truncated` could never fire
to say so, because it tests `cursor < oldestCursor` and `oldestCursor` was always
0. A poller got wrong or duplicated lines under `truncated:false`. Separately, a
streaming turn's lines counted as completed with `partialLine` hardcoded empty,
so a mid-turn cursor consumed a half-written line whose growth was never
redelivered — the `"hel"`/`"hello"` hazard the PTY reader guards against. The
journal does have stable item identity, but `terminal.read`'s cursor is a number
on the wire and cannot carry it, so a cursor read now refuses and names
`worker-read --source transcript`, which already has that contract including
`source_changed`. No cursor space is advertised either: `nextCursor` is null and
the cursor fields are absent, rather than claiming an index the next read cannot
honour. The header claim that all four fields kept their meanings was true of the
shape and false of the invariants; it now says which ones hold.

Two fixes had no test at their real seam, which is the same failure that produced
this whole set — the runtime tested directly, the seam tested by neither. The
group-addressing test hand-composed the recipient list itself, so deleting the
composition at the call site left it green; it now drives `sendGroupMessage` with
no PTY terminals at all. Nothing referenced `isLiveStructuredAgent`, so the
`dispatch --inject` fix had no red-then-green at all; it now has one driving
`RuntimeTerminalAgentPresence.isRunning`. Both were ablated and confirmed red.

Folder-workspace removals sweep and kill PTYs without `requirePhysicalStop`, so
the structured sweep no-opped there and left a live session bound to a workspace
about to be forgotten. They now close best-effort under an explicit
`closeStructuredSessions` flag, kept separate from `requirePhysicalStop` because
the two questions differ: that one asks whether a stop must be PROVEN before
files are touched, and it is what licenses a refusal. These paths do not refuse —
the root is shared so no checkout vanishes under the child, and one of them is a
never-throw forget a refusal would wedge. Reconciliation sweeps set neither and
still close nothing.

Also: the force close is raced against the same sweep deadline every PTY surface
is bounded by, so a wedged provider close reports the timeout instead of hanging
`worktree rm --force` forever; and the refusal now prints a count and the
providers instead of raw session ids, which our own marker rationale treats as
one tab-id hop from a credential.

* test: pin structured-session close on the folder-workspace removal path

The folder and orphan removal callers now pass closeStructuredSessions so a
live structured session is closed best-effort rather than left bound to a
workspace Orca has forgotten. These three exact-args characterizations describe
that call and had not been updated.

* fix(orchestration): stop the structured worker-read cursor misdelivering silently

`worker-read --source transcript` for a structured worker fingerprinted only the
oldest item's id, so `source_changed` fired when the window slid off the front
and could NOT fire when the page's contents changed under a stable oldest item —
which is the normal case, because the journal is a reduced, mutable timeline. A
`running` tool item gains its `[tool result]` at its original sequence once later
items exist, the 60ms delta coalescer revises a message in place, settlement can
rewrite an item smaller, and a pending approval projects to null until it
resolves and then appears in the MIDDLE of the array.

Two silent failures followed, both returning ok. Omission: a caller handed a
coalesced `hel`, resuming past it, never received the revision to `hello world`
— the same defect we refused to ship on the terminal read path, already shipped
here. Duplication: a resolved approval inserted ahead of a saved index, which was
still accepted, so the caller re-read content it already had. The blast radius is
the coordinator polling loop, the verb's primary consumer.

The anchor is now the oldest item PLUS every item whose projected message sits
below the caller's position, by id and revision. `createWorkerOutputSourceIdentity`
already takes an arbitrary string array and the cursor is already opaque
base64url carrying its own position, so neither the wire shape nor the
`source_changed` contract changes.

Prefix-scoped rather than whole-page deliberately: fingerprinting every item on
the page would flip the identity every 60ms with the coalescer window during an
active turn, making the cursor unusable exactly while the worker is working —
that trades a silent bug for a useless verb. Tail growth the caller has not read
cannot invalidate; any change to what it already holds does. Position-dependence
is safe because `p` rides in the same opaque payload as the identity, and the
returned cursor is stamped with the identity of its own end, which is precisely
what the next read recomputes. The frozen archive keeps a constant identity: no
item can be revised under a caller there, so it has no prefix to fingerprint.

Both silent shapes are pinned across a page boundary with the journal mutating
between reads — a static-journal test passes either way. Two ablations at the
real call site: reverting to the oldest-item-only anchor turns both red, and
widening the prefix to the whole page turns the tail-growth case red, which is
what proves the scoping is real in both directions.

* docs(orchestration): stop the structured terminal-read refusal recommending a dead end

The refusal told a peer to "page it with `orca orchestration worker-read
--source transcript`", which is wrong three ways and this file said so itself:
its own header explains that this verb exists BECAUSE `worker-read` demands a
dispatch id and coordinator standing "a peer does not have" — and then the
refusal sent that same peer there. The verb it named is also a window index over
the same bounded page, so it is not a paging answer even for a caller who can
reach it; under load it now answers `source_changed` on most polls, which is
better than the silent hole it had before but still not what the sentence
promised.

The refusal now says what actually works — the tail is bounded and newest-last,
so poll it and diff — and names no alternative, because there is none. That is
the honest framing: a durable cursor is not achievable here at all, rather than
blocked on the wire shape. The journal is a reduced, MUTABLE timeline: an item's
projected text changes at its original sequence after later items exist, the
delta coalescer revises repeatedly, settlement can rewrite an item smaller, a
pending approval renders as nothing and then as something, and `sequence` resets
on epoch rollover. No index, numeric or opaque, survives that.

So the docstring's "pagination with a real anchor lives on `worker-read --source
transcript`" is gone too — there is no real anchor there — and the file now
records why no windowed alternative should be built later: a broken cursor fails
UNSAFE, as a silent hole in a poller's output, while diffing a bounded tail fails
safe as a harmless re-read, and a second paging-shaped verb would invite the PTY
assumptions this one cannot honour.

The test asserted the old advice, so it now pins the contract instead: the
refusal explains the working approach and must never name `worker-read`.
`worker-read --source transcript` remains a good bounded snapshot for a
coordinator reading a worker it dispatched; only the "or page it with" clause was
false.

* fix(i18n): add the missing worktree-removal agent-session refusal string

The structured-session removal refusal introduced a translate() key with no
en.json entry. Nothing local catches that: typecheck passes, and the full
suite passes, because a missing key falls back to its inline default at
runtime. Only verify:localization-catalog fails on it, which is why CI's
static analysis reddened on a branch that was green everywhere else.

Fallback wording mirrors the sibling unstoppedPtyLive string, since the two
refusals differ only in what is still running and what Force Delete does to it.

* test(codex): expect the no-identity marker on an unregistered structured child

The refuse-rather-than-guess marker landed after these expectations were
written, and all three assert exact env equality on the unregistered path —
the one branch that now carries ORCA_STRUCTURED_SESSION. One of the two files
was added by this same branch, so this is a self-inflicted drift; the other
predates the branch and was broken by it.

The marker's presence is still pinned positively by
structured-worker-child-identity-env.test.ts and the CLI's
orchestration-structured-session-no-identity.test.ts, so relaxing these three
exact-equality checks loses no coverage of the security property.

* fix(orchestration): require exit evidence before settling structured close

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-06 20:24:21 -07:00
Brennan BensonandMerge Sim b7b6ea3942 fix(native-chat): auto-rename the workspace on a structured chat's first turn (#19138)
* fix(native-chat): auto-rename the workspace on a structured chat's first turn

Structured native chat (Claude and Codex) never reached the first-work
workspace rename. The orchestrator has a single production caller, the
agent-hook server listener, and structured sessions never set
ORCA_PANE_KEY, so no hook event could ever be attributed to one. The
renderer knew this and suppressed pendingFirstAgentMessageRename for
structured launches at three sites, which also closed the gate the
folder-workspace title rename depends on.

The host's status feed already computes the exact edge: status 'working'
with a latestPrompt normalized the same way the hook payload is, and a
workspaceId that IS the worktree id. Publish that projection to the host,
thread it out to the runtime, and hand it to the same orchestrator the
hook path uses.

Re-projections of state the host already knew (restore, an arriving
subscriber) are flagged as replays and map to the orchestrator's existing
isReplay gate, so a host restart cannot rename off a stale journal.

One host and one journal serve both providers, so this covers Claude and
Codex together.

Verified in a live Electron instance, worktrees created through the real
composer and prompts sent through the real chat composer:
  Codex  langouste -> retry-helper-exponential-backoff
  Claude prowfish  -> parse-csv-headers

* fix(native-chat): preserve first-work rename across runtime and queued turns

* fix(native-chat): skip branch rename for folder projects

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-06 17:38:06 -07:00
Brennan BensonandMerge Sim f7d5216016 Show provider activity in chat turn tails (#19055)
* feat(chat): show turn-scoped activity tail

* fix(chat): keep turn activity broad

* feat(chat): surface provider activity in turn tail

* fix(chat): keep reasoning headline as activity and widen redaction

A Codex reasoning summary streams as a bold headline followed by body text.
Folding the whole summary into the tail leaked literal ** markers and body
prose; only the first non-empty line is activity copy, and an unterminated
bold header mid-stream is unwrapped too.

Redaction used a hyphen for GitHub token prefixes (they use an underscore),
and missed fine-grained GitHub tokens, AWS access key ids, JWTs, URL
userinfo passwords, and bare token= values.

* fix(chat): wait for a complete reasoning headline

A bold headline still streaming has no closing marker yet; holding the
previous activity copy until it lands avoids flashing a half word.

* refactor(chat): drop bespoke secret redaction from activity copy

Reference agent hosts render provider-derived status text unredacted;
this table was the only one of its kind and its GitHub pattern matched
no real token. Bounding and the reasoning-headline extraction stay.

* Bound provider headline updates and clear activity on reconnect

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-06 16:42:18 -07:00
Brennan BensonandMerge Sim ade9718557 fix(native-chat): suppress provider user echoes in Claude and Codex (#19136)
* fix(native-chat): keep provider user echoes out of the conversation

* fix(native-chat): retain input beside Codex skill context

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-06 16:28:10 -07:00
Brennan BensonandMerge Sim 298571ad9f fix(codex): uncap app-server stdio records (#18590)
Co-authored-by: Merge Sim <sim@local>
2026-09-06 11:49:42 -07:00
Brennan BensonandMerge Sim 471a5f4aa7 feat(native-chat): model Codex MCP and web-search items instead of leaking opcodes (#18763)
* feat(native-chat): model Codex MCP and web-search items instead of leaking opcodes

Codex's app-server sends 19 thread-item types; the structured translator handled
six. The rest fell through to a generic gray `codex · item:<type>` row, even
though the disposition table's own comment says it exists so a new item type
cannot leak like that — the table had one entry.

Give `mcpToolCall` and `webSearch` real tool-call bodies, and chrome `sleep`,
which carries only a duration and renders as nothing in Codex's own TUI.

`subAgentActivity` and `collabAgentToolCall` deliberately keep their generic
rows. They arrive in real sessions today and are currently the only visible
sign a subagent is running; hiding them before the subagent UI lands would
render minutes of work as an idle turn. Tests pin that they stay visible.

MCP tool names pass through verbatim when they contain `:`, `.`, `/` or `__`,
so `mcp__server__tool` survives instead of being title-cased into nonsense.

* fix(native-chat): keep Codex MCP tool identity and web-search results on the row

Four fixes to the Codex MCP / web-search item bodies:

- Drop the title-casing display name. `get_forecast` became `Get Forecast`,
  which no longer matches the raw snake_case identifiers that the diff
  renderer, question parsers, and tool-input previews dispatch on, and does not
  match how the Claude lane or the sibling `shell`/`apply_patch`/`web_search`
  bodies name a tool. The row name is now `server/tool` verbatim, the bare
  `tool` when no server is given, and `mcp` when the item names no tool at all.
  Server-qualifying also stops an MCP tool that happens to be called
  `apply_patch` from hijacking the diff renderer.

- Pass the MCP call's own `arguments` as the tool input instead of wrapping it
  in `{server, tool, arguments}`. Row-label derivation only reads top-level
  keys, so the wrapper degraded every MCP row to a truncated raw JSON blob.
  A non-object `arguments` stays addressable under a key rather than being
  dropped; an absent one becomes null, which labels as empty rather than `{}`.

- Carry a web search's `results` as the call output, bounded like every other
  inline payload and omitted when there are none. They were being dropped
  entirely, which showed less than the generic fallback row it replaced.

- No streaming branches were added for these two item types: the Codex delta
  stream is a closed set of six methods that neither can reach, so such
  branches would be unreachable.

* fix(native-chat): label Codex web searches and argument-less MCP calls

A row label is derived from top-level `input` keys only, so a webSearch
whose detail lives inside `action` — an opened page, an in-page find, or
a bare `other` — fell through to the raw JSON of the whole input, as did
the empty `query` Codex leaves on a completed search. Hoist the action's
`url`, `pattern` and `type` beside the query, keep the full `action`
object so the expanded detail loses nothing, and emit no input at all for
the start frame.

An MCP tool that takes no arguments sends `arguments: {}`, which passed
straight through and labelled the row a literal `{}`; treat it as absent
so the row reads as a bare `server/tool`.

Split the durable-identity half of the item translator into
`codex-thread-item-identity.ts`, re-exported so every existing import is
unchanged, to keep both files under the max-lines cap.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-05 15:33:04 -07:00
Brennan BensonandMerge Sim ddc5b75ac7 feat(native-chat): label Codex tool rows by what the command actually did (#18760)
* feat(native-chat): label Codex tool rows by what the command actually did

Codex's app-server `commandExecution` item carries `commandActions`, which
already classifies each command as a read, a search, or a directory listing
with the target path, name, or query extracted. Orca ignored the field, so
every shell call rendered as an undifferentiated row of raw argv.

Read it and name the row by its class, keeping the raw command and cwd for the
expanded view. Unclassified commands are untouched: absent, null, or malformed
`commandActions` produces byte-identical output to before.

Rank the search term above the command in the shared label keys so a classified
search row reads by what it looked for rather than the shell text that ran it.
No first-party tool input carries both keys today, so this only reaches the new
rows; an MCP tool supplying both would prefer its search term.

Note `commandActions` is the app-server spelling. `parsedCmd` is the rollout-file
shape and never arrives on this lane; a test pins that it stays ignored.

* feat(native-chat): give tool rows a category glyph beside their word

A row named only by a word makes the reader parse text to tell a read from
a search. Pair the word with an icon: icon for category, word for action,
argument for target.

Name the full eight-category vocabulary in `src/shared/native-chat-tool-icon.ts`
now — read/search/listFiles/unknown/fileChange/webSearch/mcpToolCall/
subAgentActivity — even though only the classified shell categories reach a row
today, so the MCP and web-search rows landing separately inherit these names
rather than coining their own. Glyph ids are the lucide spelling shared by
`lucide-react` and `lucide-react-native`, so mobile can resolve one name to its
own component when it adopts this; mobile rows stay text-only for now.

The glyph is decorative and `aria-hidden`: the word is the accessible name, and
never renders without it. One glyph per category, fixed across running,
completed, and failed — a row that swapped icons on completion would read as
changing identity — so the run header's active row also takes its category glyph
instead of the generic wrench it fell back to once these rows stopped being
called `shell`. A word outside the vocabulary gets the terminal glyph rather
than a blank slot, so rows stay left-aligned.

Also stand `.` in for a `listFiles` action whose `path` is null, which is what a
bare `ls` sends. The row named the action and then showed the raw argv as its
target; now it names the directory it listed.

* fix(native-chat): hold the tool run header's glyph fixed and size its slot to 16/14

The header swapped its leading glyph on settle: the active tool's icon while
running, a check once done. That is the identity swap a fixed per-category glyph
exists to prevent — the row appeared to become a different thing when it
finished. Name the header by the run's latest tool in both states and move the
completion check to the trailing edge, where the rest of the state signal already
lives.

Size both header slots to the mock's 16px slot with a 14px glyph, matching the
tool rows beneath them and the subagent summary row landing separately. They were
24/16, so the icon columns sat 8px apart and broke the left alignment the icon
treatment depends on.

The fixity test walks running, completed, and failed and pins the leading glyph
of every row by lucide's own class name, so a swap shows up as a different name
rather than a still-present icon.

* fix(codex): stop a classified shell row from asserting facts the command doesn't support

Three claims the `commandActions` row model was making on its own:

- `listFiles` with a null path was given `path: '.'`. Codex sends null for a
  recursive walk and for the repo root, and the invented path flows into
  `createToolInputDisplay().filePath`, which mobile turns into a tappable
  "open file" link onto a directory — an affordance that can only fail. The row
  now keeps the raw command, which is what the label logic already falls back to.
- A command whose actions classify as two different things (`cat a.txt && ls src`)
  was named after the first one, silently dropping the rest. Recognized actions
  must now agree on one class; a repeat of one class keeps the class and only a
  target every entry names.
- `read` lifted `name` into the journal payload, where no label ever reads it —
  `path` always wins — so it was bounded weight carrying nothing.

* fix(native-chat): give an unmodelled tool row a generic glyph, not a terminal

The row-word vocabulary named seven words, and everything else fell through to
the terminal glyph — which reads as "a shell ran here" for rows where nothing
says one did. Codex's own `apply_patch` row, `Grep`/`Glob`/`Task`/`WebFetch`/
`TodoWrite`, and every `mcp__*` tool all rendered a terminal, leaving the
declared `mcpToolCall` and `subAgentActivity` categories unreachable.

- Split the vocabulary: `unknown` stays the shell command Codex could not
  classify and keeps the terminal, while a new `other` carries the generic
  wrench that unmodelled words now fall back to.
- Read the edit family from `EDIT_TOOL_NAMES` and the command tools from
  `isCommandToolName` rather than restating either. Command tools resolve first:
  `isEditToolName` counts `shell`/`exec` as possible patch carriers, and a shell
  row is not an edit.
- Result rows get no category glyph. Their word is `translate(…, 'Result')`, so
  keying a category off it resolved a different glyph per locale; an empty slot
  keeps the rows aligned.
- The header and the row now resolve through `NativeChatToolIcon`, so one `Grep`
  run can no longer show a wrench in the header and a terminal on its line. The
  glyph map and the unused `category` prop go with the duplication.

* fix(native-chat): give the projected Diff row the file-change glyph

Every Codex fileChange item projects to a tool call named `Diff`, which the
edit set does not name — it names the tools that carry the edit in their own
input. So a run whose body renders an edited-file card was headed by the
generic wrench.

* fix(codex): stop a classified shell row offering a folder as a file to open

A listFiles action's path is a directory, and a search action's path is the
root it scanned. Lifted under `path`, both became the row's file target, which
mobile renders as a tappable open-file link that can only fail — the same dead
link the removed `{ path: '.' }` stand-in would have produced. They lift to
`directory` instead, which still labels the row but is never a file target.

* fix(mobile): keep the terminal glyph on a classified Codex shell row

Mobile's run header picks between a terminal and a generic glyph by tool
name. Now that the host publishes `read`/`search`/`list` for the same
commands it used to publish as `shell`, that name check answers false and
a command that really ran heads its run with a wrench.

Ask the shared category vocabulary instead. Mobile keeps its two icons —
porting the full glyph set is a separate lane.

* fix(native-chat): say what the run header's glyph actually guarantees

The comment claimed the header names the same tool in both states, so its
glyph cannot change on settle. It can: the live header names the running
call while the settled one names the run's last tool call, and with
out-of-order completion those differ. The glyph is fixed for whichever
tool the header names — say that, and drop the never-taken running branch
from the settled header's call.

Also pin the other half of the file-target rule: `read` keeps `path`, so
its row stays tappable, where `list`/`search` lift a folder to
`directory` and offer no target at all.

* fix(native-chat): give a rollout-transcript shell row the terminal glyph

`exec` and `local_shell` are what the Codex rollout transcript names a
shell call — `native-chat-edit-normalize` already treats those three
words as the command tools — but the activity set the glyph vocabulary
reuses carries neither, so both rows headed a real command with the
generic-tool wrench.

Named in the vocabulary rather than in that activity set, because that
set also picks the running row's copy and this is only about the glyph.

* fix(mobile): pick the run-header glyph from the call's input, not its word

Codex now names a classified shell row `read` / `search` / `list`, which
lowercase to Claude's own `Read` / `Grep` / `Glob`. Mobile has only a terminal
and a wrench, so keying that choice on the row word gave Claude's filesystem
tools a terminal for a shell that never ran.

The input separates them: Codex keeps the raw command on a classified row,
while Claude's `Read` carries only a file path. `isShellActivityToolCall`
replaces `isShellActivityToolRow` and asks the command tool names first, then
the call's input.

* fix(native-chat): give the projected diff fixture its required digest

* fix(native-chat): head a settled run with a glyph the whole run shares

The settled run header drew the glyph of the run's last tool call while the
text beside it summarizes the run's first three, so a ten-call run ending in a
`read` showed an eye above "shell npm test · shell git status · …" — a category
the summary never described.

Resolve the header's glyph from every call in the run instead: the shared
category's glyph when all agree, the generic tool glyph when the run spans
categories, and no glyph when there are no tool calls. The running header still
names the active call, whose glyph is true of it.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-05 14:03:45 -07:00
Neil 0a821e5bc8 fix(crash-reporting): make the own-Chromium gate a real choke point, and stop a refusal leaking the root (#18459)
* fix(crash-reporting): make the own-Chromium gate a real choke point

Round-3 review found the guard was not the choke point its own comments
claimed: six pid-addressed `taskkill /pid <pid> /t /f` families in main were
ungated and uninstrumented, so the stale-pid shape stayed producible and a
`selfInitiatedTreeKillCount: 0` could read as exculpatory when it was not.

- Gate the remaining main-process families: the git command-runner abort, the
  notebook-cell and automation-precheck timeouts.
- Turn the `src/shared` seam into the gate itself (`process-tree-kill-gate`), so
  the runProcess choke point, the codex app-server deadline kill and the
  ephemeral-VM recipe kill ask the same decision. Those three are compiled into
  the CLI/relay too and cannot import main; main installs the guard at preflight.
- Ratchet (`main-process-tree-kill-gate.test.ts`): a new pid-addressed taskkill
  in main that skips the gate fails, and the allowlist entries must still exist.
- Give pid-addressed kills eviction priority in the 32-entry ring: 32 routine
  `win-pty-job` teardowns from a window-close burst no longer evict the one
  entry that discriminates a self-kill from an external one.
- Correct the coverage doc, which described the uninstrumented Windows sites as
  POSIX `process.kill(-pid)` group kills and omitted the git and codex paths.

* fix(crash-reporting): keep a refused tree-kill from leaking the root it owns

A refusal must block the pid-addressed tree walk, not the termination. Five of
the six gated sites returned on refusal with no fallback, so a refused
`taskkill /pid /t /f` left git.exe, a timed-out notebook cell, an automation
precheck or an ephemeral-VM recipe running while the caller reported it stopped.
The root kill is addressed by the child handle, which cannot reach the recycled
pid the refusal is about, so it stays correct and required on that path.

Also fixes the ring eviction the scope preference introduced: with the ring
saturated by pid-addressed kills, the only non-pid-addressed entry is the one
just pushed, so the splice evicted itself and the detail came back `{}` --
byte-identical to the external-kill arm, in the window-close case the guard
exists for. Eviction now excludes the newest entry and falls back to FIFO.

Tests: refusal now asserts the root kill at all six sites, and the ring covers
the saturated-pid ordering as well as round 3's group-burst ordering.

* fix(crash-reporting): stop a refused tree-kill leaking the commit-message agent, and count call sites

Two round-5 blocking findings, both open on main and on both branches.

`killSourceControlAgentProcess` had no root-kill fallback on its win32 arm: the
taskkill was the only termination, so once the own-Chromium gate could refuse it
the promise resolved having killed nothing. Both callers do
`terminationComplete ??= killSourceControlAgentProcess(child)` and then release
the managed-home lock on that promise, so a refusal left the local Codex/Claude
commit-message agent running while the caller reported it stopped -- the
lock-contention failure the taskkill was added for. Same fix as the six sibling
sites: the handle-addressed root kill cannot reach the recycled pid the refusal
is about, so it stays correct and required on that path.

The ratchet was file-granular, not call-site granular: one gate mention anywhere
in a file exempted every taskkill in it, which left the six files that now ask
the gate ratchet-blind -- the inverse of what it is for. It now counts `/pid`
call sites against gate admissions per file, so a second ungated kill inside an
existing family fails. Keying on the `/pid` argument rather than a quoted
`taskkill` also catches a kill whose program name comes from a constant. The
three comments that claimed more than the old scan enforced now state the rule
and its two remaining blind spots.

Also: the recording in `admitSelfInitiatedTreeKill` is now wrapped the way the
`admitProcessTreeKill` seam already wraps it, with the refusal decision taken
before anything that can throw so a diagnostics failure cannot flip it; and
`orca-chromium-process-pids` documents the false-positive direction (a stale
`getAppMetrics()` entry plus pid reuse refuses a live unrelated child), which is
the mechanism the root-kill fallback exists to bound.

Tests: refusal now asserts the root kill at all seven sites; the ratchet asserts
call-site counting and the constant-program form.

* test(crash-reporting): run the own-Chromium gate against real Windows trees

Nothing on this branch had ever executed on Windows. The unit tests pin the
gate's decision against a mocked taskkill, which cannot show that the decision
does anything to a real process: that `/T /F` reaps a detached grandchild, that
a refusal leaves that tree standing, or that the handle-addressed root kill the
refusal path falls back to reaps the root while orphaning descendants.

Adds a win32-gated live test covering all four, registered in both the
`package_windows` CI lane and `WINDOWS_PACKAGE_TESTS` as
`win32-test-lane-registration` requires.

Also completes the coverage doc's "never instrumented" list, which omitted the
macOS keyboard-input-source probe's POSIX group kill in `ipc/app.ts`.

* fix(crash-reporting): pin the commit-message root kill on the Windows arm

The first Windows run of this branch found nine failures the macOS suite
cannot see: `commit-message-text-generation-test-harness` asserts
`expect(child.kill).not.toHaveBeenCalled()` on `process.platform === 'win32'`,
which is the contract the previous commit deliberately replaced — and it
branches on the real platform, so it is dead code everywhere CI runs today.

The harness now asserts the handle-addressed root kill on every platform. On
win32 it lands after the tree walk, so the expectation waits rather than reading
one tick early, and its ten call sites await it. Red against the pre-fix arm at
all seven sites; the production code is unchanged.

* test(crash-reporting): remove the Windows lane marker tree through the retrying helper

The new win32 spec teardown used a raw rmSync, which the windows-lane-tree-removal
boundary ratchet rejects — and which is exactly the EPERM the ratchet exists to
prevent, since this spec's marker directory is written by processes it has just
force-killed.

* fix(crash-reporting): only refuse pid-addressed tree walks, disclose the handle-less codex site

The own-Chromium gate refused the POSIX process-group arm of
signalProcessTree as well, which was new macOS/Linux behaviour: a stale
getAppMetrics() entry plus pid reuse would orphan a group that main reaps
today. A POSIX group only holds what Orca put in it, so the refusal is now
scoped to win-taskkill-tree and the POSIX arm is recorded and admitted like
the other group kills in main. That also drops the synchronous
getAppMetrics() read from every POSIX termination.

codex-turn-added-roots kills roots found by a table walk, so a refusal has
no handle to fall back to. Pin that the refusal is visible - crumb written,
turn reported as not cancelled - rather than fixing what cannot be fixed.

* test(crash-reporting): detach the Windows survival fixture and observe real spawns
2026-09-04 16:42:42 -07:00
a65332a8bd feat(claude): move structured native chat onto the Claude Agent SDK and enable it on macOS and Linux (#18560)
* Join structured attach teardown through journal bind

* fix: restore structured chat parity

* feat: add Claude structured session adapter

* fix: harden Claude structured adapter

* fix: close Claude adapter edge cases

* fix: start Claude init deadline after launch

* feat: wire Claude structured sessions

* fix: harden Claude structured runtime

* fix: fence Claude structured compatibility

* fix: preserve Claude free-text prompt answers

* fix: decode addressed Claude prompt text

* feat: enable Claude structured chat on mobile

* fix(mobile): keep structured chat provider-aware

* fix(mobile): negotiate Claude structured tabs

* fix: keep scoped RPC tests native-free

* fix: secure mobile structured image delivery

* fix: close structured session data-loss gaps

* fix: prove real Claude structured startup

* fix: consume pre-spawn proof before retry

* feat(native-chat): add desktop structured sessions

* fix(native-chat): satisfy structured session cleanup gates

* fix(native-chat): keep structured renders pure

* fix(native-chat): open composer pickers upward

* fix(native-chat): use existing view for structured sessions

* fix: harden structured desktop status projection

* fix: close structured desktop lifecycle gaps

* fix: fence structured AI Vault resumes

* fix: fence structured AI Vault resumes

* fix: preserve structured tabs during activation

* feat: toggle structured sessions between chat and TUI

* fix: harden structured session handoffs

* fix: bind structured TUI before rollout proof

* fix: complete structured chat round trips

* fix: align structured TUI return readiness

* fix(native-chat): make reverse handoff transactional

* Add Claude structured TUI handoff seams

* fix(native-chat): clear sticky handoff recovery

* fix(native-chat): complete mobile reverse after TUI exit

* fix(native-chat): keep TUI transcripts readable

* fix(native-chat): recover TUI transcript gaps

* fix(native-chat): recover claimed TUI owners

* fix(native-chat): retain cold TUI proof authority

* fix(native-chat): preserve Claude handoff authority

* fix(native-chat): recover TUI transcripts read-only

* fix(native-chat): harden Claude handoff recovery

* fix(native-chat): serialize structured handoff recovery

* fix(native-chat): close handoff admission races

* fix(native-chat): validate pinned launch environment

* fix(native-chat): revalidate restored and retried owners

* fix(native-chat): gate restart recovery publications

* fix(i18n): catalog Claude session controls

* fix(native-chat): wait for structured TUI process proof

* fix(native-chat): queue stale idle TUI handoffs

* fix(native-chat): route structured Codex options directly

* fix(native-chat): persist structured session options

* fix(native-chat): hydrate resumed structured options

* fix(native-chat): preserve options across structured handoffs

* fix(native-chat): replay pending option mutations

* fix(native-chat): rotate settled handoff operations

* fix(native-chat): rotate refused send operations

* test(native-chat): derive refusal retry state from host

* test(native-chat): give the host-oracle matrix test an explicit timeout

* fix(native-chat): keep Claude option controls idle

* fix mobile structured first-send hydration race

* fix(native-chat): preserve handoff launch authority

* fix(native-chat): harden shared handoff recovery

* fix(native-chat): serialize structured handoff recovery

* fix(native-chat): close handoff admission races

* fix(native-chat): validate pinned launch environment

* fix(native-chat): revalidate restored and retried owners

* fix(native-chat): gate restart recovery publications

* fix(i18n): catalog structured session recovery control

* fix(native-chat): wait for structured TUI process proof

* fix(native-chat): queue stale idle TUI handoffs

* fix(native-chat): keep structured recovery provider-neutral

* fix(native-chat): drop local terminal topology from structured sync

* fix structured outbox and tab restore races

* fix(native-chat): preserve Claude question groups

* fix structured provider visibility and request handling

* fix structured session TUI handoff recovery

* fix reverse structured session handoff

* fix(native-chat): recover Claude outbox and resume state

* chore(mobile): preserve the working-tree lockfile state before the main merge

Carries the pre-existing uncommitted mobile/pnpm-lock.yaml modification into history so the
main merge cannot overwrite it. Verified benign pnpm drift (babel 7.29.7->7.29.8 transitives
plus deprecation metadata); drops no patchedDependencies (the mobile lockfile declares none).

* test(native-chat): drop orphaned Claude handoff-auth test left by the main merge

'pins Claude handoff auth through the terminal provider boundary' is absent from main and its
production counterpart preserveClaudeAuthEnv no longer exists outside this test - orphaned residue
of the terminal/native handoff work this PR excludes by scope.

Removed rather than repaired: the failure was a renamed field (providerHome -> providerRoot), and
renaming it would have carried out-of-scope handoff code into the merge. Body preserved as evidence
and logged in CLAUDE-STRUCTURED-DISPOSITION-TABLE.md.

* Fix mobile structured turn state

* fix Claude structured session blockers

* fix claude structured lane blockers

* fix Claude acquisition exit proof

* fix(claude): route stream-json launch through process wrapper

* fix(claude): gate structured chat support

* Fix Claude structured launch gating

* fix(claude): split session acquisition and prune mobile scope

* test(claude): align structured session fixtures

* fix(agent-session): preserve handoff launch arguments

* fix(claude): open journals through the factory after origin/main split

The journal opener moved to journal-store-factory on main; retarget the
Claude structured tests that still imported the old path.

* fix(claude): resolve Claude structured launch args, auth, and win32 proof

The origin/main merge re-expressed the lane's Claude wiring onto main's split
orca-runtime facade and dropped three wires past green typecheck and lint.

- resolveLaunchArgs discarded its provider parameter, so structured Claude
  sessions were launched with Codex app-server flags; Claude exits on
  --dangerously-bypass-approvals-and-sandbox, and a Codex arg-parse throw
  could block Claude session creation outright.
- resolveClaudeLaunchEnv was no longer supplied, so the launch resolver fell
  back to the whole process env as configuredEnv and
  buildClaudeChildProcessEnv re-applied every auth var it had just stripped.
  The resolver now merges the Claude overlay onto a strip-applied copy of the
  inherited env, which also keeps PATH intact for withCliRuntimeOnPath.
- The windowsProcessStartTimeAvailable producer was gone while the contract
  field and both consumers survived, so the renderer gate fail-closed and
  structured native chat was unreachable on every win32 host.

Separately, structured Claude pinned CLAUDE_CONFIG_DIR unconditionally. An
explicit pin makes the CLI abandon the macOS Keychain even when it names the
CLI's own default, so a default claude.ai account could not authenticate where
the legacy Claude terminal could. Pin only a home the CLI would not resolve on
its own, matching ClaudeRuntimePathResolver, and compare against the env the
child would otherwise inherit so a diverging overlay cannot outrank the
record's account home.

Also await the now-async revealNativeSession in its regression test, and set
the native status before revealing so a rejecting reveal cannot leave a
session released but never marked native.

Claude-Session: https://claude.ai/code/session_013UqKCRB6k5e8UaYhXUHeWY

* fix(claude): scrub case-insensitive Windows auth env

* fix(native-chat): settle handoff outcome-write failures instead of leaking them

A store write failure while recording a handoff outcome escaped the flow
runner's catch handler, so the client never received the failure and the
flow surfaced as an unhandled rejection (seen as an intermittent
agent_session_store_corrupt error in the proven-dead-retry suite, whose
teardown raced the flow's trailing outcome write). Record the failed
outcome best-effort, and drain the coordinator before that test's
teardown removes the store root.

Claude-Session: https://claude.ai/code/session_011aXkcHyeiRJuezupQdjZaM

* fix(native-chat): make the structured close-failure toast provider-neutral

The structuredSessionCloseFailed toast fires for any structured session,
but its copy said 'Codex chat', so a Claude structured session that fails
to close showed the wrong provider name. The launch-failure toast is only
reachable behind the agent === 'codex' gate, so its copy stays as is.

Claude-Session: https://claude.ai/code/session_013ugSpCx4AWkySaJb69BQax

* fix(native-chat): wire structured handoff proof recovery

* fix(native-chat): wire structured handoff proof recovery

* fix(native-chat): correct the structured chat opt-in copy

The one `experimentalStructuredNativeChat` toggle gates both providers —
`useStructuredAgentSessionCreate` runs `canUseStructuredNativeChat` for
`'claude'` as well as `'codex'` — but its description named only Codex.

Its scope line also said Windows keeps using terminal chat, while the gate
refuses win32 only until the host proves it can read a process start time.
`structured-native-chat-availability.test.ts` already pins that Windows is
allowed once the proof is cached, so the two contradicted each other.

Claude-Session: https://claude.ai/code/session_01RJFsidQWmKYFmeoUuVu4Tp

* test(claude): pin @anthropic-ai/claude-agent-sdk 0.3.251 contracts against a scripted CLI

PR 1 of the SDK migration: dependency + test-only harness, no product wiring.

- Pin @anthropic-ai/claude-agent-sdk to exactly 0.3.251 — not the newest
  release — because 0.3.251 (published 2026-08-28) clears the repo's 3-day
  minimumReleaseAge supply-chain gate with no exclusion, while the newest
  release was minutes old and would have required excluding a brand-new
  publish from the exact control built to catch brand-new malicious
  publishes. Every contract this design depends on was verified identical
  on 0.3.251: the full option surface, no pid on SpawnedProcess (custom
  spawner stays mandatory), env defaulting to process.env when omitted, and
  --replay-user-messages appearing only via extraArgs.
- Exclude all eight bundled CLI platform binaries via
  ignoredOptionalDependencies. The setting lives in pnpm-workspace.yaml
  because pnpm 12 no longer reads the package.json "pnpm" field (it warns
  and ignores it; verified by install ablation). Excluding the binaries is
  what makes Orca's pathToClaudeCodeExecutable override mandatory rather
  than merely preferred. Note: pnpm 12.0.0 honors the ignore list when
  reconciling an existing lockfile but not on fresh resolution of a new
  dependency, so the lockfile's SDK entry was pinned surgically; both
  'pnpm install' and 'pnpm install --frozen-lockfile' verify clean and
  stable against the committed lockfile.
- Contract-pin suite drives the real SDK against a scripted fake CLI and pins:
  unknown type/field/content-block pass-through (and keep_alive interception),
  spawner env fidelity plus the omitted-env process.env inheritance sharp edge,
  extraArgs producing --replay-user-messages, argument parity for every
  CLAUDE_STRUCTURED_BASE_ARGS entry plus --session-id/--resume/
  --resume-session-at, canUseTool wire request_id stability and abort on
  control_cancel_request, one spawn per query, pathToClaudeCodeExecutable
  honored by the default spawner, the exact SDK version, and the eight platform
  binaries staying uninstalled.

Claude-Session: https://claude.ai/code/session_01FGCRfYUnb4hbvfTAHGtJKQ

* feat(claude): drive the structured transport through the agent SDK

Replaces the hand-rolled `claude -p --input-format stream-json` transport with
@anthropic-ai/claude-agent-sdk 0.3.251, keeping the existing connection
interface for this commit so the acquisition path changes minimally. The
control-plane rewrite is a separate change.

Orca still supplies the process. `spawnClaudeCodeProcess` routes through
`spawnProcess`, retains the child and its pid — the triple the durable lease
adjudicates on — drains stderr so exit errors keep their tail, and hands `.cmd`
shims to Orca's Windows argument encoder rather than the SDK's plain spawn.
`close()` keeps Orca's own bounded tree-kill and exit deadline, so it still
resolves true only after an observed exit.

Launch resolution emits an SDK options object instead of argv; durable
`launchArgs` translate to a typed option where one exists and to `extraArgs`
otherwise, refusing a token neither can carry rather than dropping it. The
child env is always passed explicitly — omitting it would let the SDK inherit
`process.env` and reintroduce the ambient `ANTHROPIC_*` leak. The stdout line
parser is deleted; the SDK owns framing, and unknown frames still reach the
translator verbatim.

Claude-Session: https://claude.ai/code/session_01JMhFjh9HEnkcJ5YTfCdgD3

* fix(claude): settle the frame the SDK pulled but never wrote

The SDK's input pump is `for await (frame of prompt) { await transport.write(frame) }`.
When that write rejects — the child dies between Orca's liveness guard and the
write — the for-await ends abruptly and calls the generator's `return()`, so the
code after `yield` never runs. The frame was already shift()ed out of `queued`,
so the later `fail()` from the exit path could not reach it and `send()` never
settled: `dispatchClaudeTurn` awaits that send before it can return `unknown`,
wedging the caller and the durable outbox. The pre-SDK transport rejected on the
stdin write callback instead.

Retain the in-flight entry and settle it from the generator's cleanup, and let
fail() reach it too for the pump that never resumes at all.

Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum

* fix(claude): keep the agent SDK behind the structured-Claude boundary

The ordinary OrcaRuntimeService graph statically reaches the Claude adapter and
so the transport module, whose first line imported @anthropic-ai/claude-agent-sdk.
The SDK is evaluated whenever the regular runtime loads, before any structured
Claude session is chosen: it sets process.env.NoDefaultCurrentDirectoryInExePath,
changing Windows executable resolution for later subprocesses, and a missing or
incompatible install would break normal runtime startup — for a user who never
leaves the terminal/TUI path.

Defer the SDK to the connection, memoized so it loads once per process, and add
the import-graph ratchet: a walk from the Electron main entry that fails on any
static import of the package, plus a clean-fork check that loading the runtime
leaves the Windows search variable untouched and a child-process pin that the
side effect is still real.

Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum

* fix(claude): answer list_models so the picker stops serving the seed

sendControlRequest had no list_models case, so every request hit the default
reject; readClaudeStructuredSessionOptions swallows that with .catch(() => null)
and falls back to the static catalog. Every structured session therefore served a
hardcoded model list with no per-model effort levels, no resolvedModel and no
default detection, and nothing surfaced the failure. The pre-SDK transport got the
live catalog from the CLI.

Route it through the SDK's supportedModels(), wrapped in the { models } envelope
the existing parser reads.

Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum

* fix(claude): reap the child's descendants before killing it

The forced step of the exit ladder went through the Codex helper, which spawns
`pkill -KILL -P <pid>` and SIGKILLs the parent in the same tick: the parent
usually dies first, the descendants reparent to pid 1, and `-P` matches nothing.
An MCP or launcher descendant of a stubborn Claude child was left running. The
test named for that requirement declined to assert it and killed the survivor by
hand instead, so it could not fail for the thing it was named after.

Route the Claude reap through Orca's existing sweep, which snapshots descendants
while their parent link still exists and signals them before the root goes, and
on Windows uses the identity-gated `taskkill /T /F`. The test now asserts the
descendant is dead; the manual kill stays only as a failure-safe. close() still
returns true only on an observed exit.

Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum

* fix(native-chat): merge the duplicated handoff type import

CI's static-analysis lint (`oxlint --config
config/oxlint-code-quality-native-plugins.json src config tests mobile
--deny-warnings`) exits 1 on the two separate `import type` statements from the
same module.

Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum

* fix(claude): answer a permission callback whose signal already aborted

settleFrom registered the abort listener and then delivered the request. A
callback that arrives already aborted never fires that event, so the promise
stayed pending behind a durable prompt with no cancel path. Check the signal
first, emit the cancel, and resolve the SDK's null sentinel without registering.

Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum

* test(claude): wait for the child to record the frame, not just for its report

The scripted CLI writes its report at startup, so `until(readReport)` returned a
report with no user messages whenever the child had not yet read the line. The
assertion then failed under parallel load. Poll for the frame instead of for the
file.

Claude-Session: https://claude.ai/code/session_01AobxxokqQ3qcxS7sy7ckum

* fix(claude): coalesce partial deltas onto one assistant item and stop painting result frames

Under --include-partial-messages every stream_event frame carries its own
uuid, and the final assistant frame for a block carries yet another; only
message.id ties them. The translator keyed each delta by its frame uuid, so a
reply painted as one bubble per delta chunk followed by a complete duplicate
under the final frame's uuid. The block's first stream frame now mints the
claude:(sessionId, uuid) identity, deltas coalesce onto it through the shared
60ms seam, and the final frame reconciles onto that same item.

Known SDK bookkeeping no longer reaches the provider-fallback row: result
subtypes are catalogued and settled by the turn lifecycle, an empty thinking
block (redacted thinking) is a modeled kind, a string-content user replay is a
text block, and an empty user frame paints nothing. An unmodeled result
subtype or content kind still lands on the bounded fallback row.

Claude-Session: https://claude.ai/code/session_01GaP5HpYQbvy2hYehVhwfEW

* fix(claude): prove descendant exit at the close boundary instead of on an unref'd timer

close() reported proven=true as soon as the direct child exited while the
descendant sweep's SIGKILL sat on an unref'd 2 s timer, so a SIGTERM-resistant
MCP server outlived the lease release. The reaper now composes the same shared
primitives the Codex structured provider uses: snapshot, verified bounded
descendant termination on POSIX, taskkill /T /F on Windows. The proof is false
whenever descendants outlive the deadline, a retried close re-verifies the
retained snapshot rather than trusting the dead root, and the raw pipe child no
longer goes through the PTY job sweep it never owned a job for.

Measured on macOS: a killed child of a SIGSTOPped parent stays a matching zombie
row in ps, so the root is killed while verification runs rather than stopped
first as the Codex non-group path does.

Claude-Session: https://claude.ai/code/session_0161QFm3KVRNJKfdzWVGVNWk

* feat(claude): replace the hand-rolled control plane with the SDK's native surface

PR 3 of the Claude structured SDK migration removes the wire-frame scaffolding
PR 2 kept, so Orca drives the SDK's typed control surface directly.

Inbound permissions move from a rebuilt control_request dispatch to the SDK's
canUseTool / onUserDialog callbacks. The prompt registry now carries the
callback's own resolver: a decodable can_use_tool becomes a durable prompt whose
answer settles the callback; a malformed one is denied without registering; the
SDK's abort signal (fired on control_cancel_request, which the SDK matches and
dedups itself) forgets the prompt and settles it null, and a late answer after
abort finds no prompt and is refused. Closing settles every in-flight callback so
no promise dangles. The claude-agent-sdk-control-bridge that rebuilt the wire
frame is deleted.

Outbound control maps to Query methods: interrupt() for cancel, setModel /
setPermissionMode / applyFlagSettings for options, supportedModels for the model
list, initializationResult() for init proof, each under Orca's own request
deadline and error classification. Cancel is interrupt-receipt aware: a CLI
advertising interrupt_cancel_queued_v1 gets cancel_queued in one round trip,
otherwise the receipt's still_queued uuids are swept with cancel_async_message so
a cancelled turn cannot spawn a later unexpected turn; older CLIs resolve no
receipt. Init keeps the 10s deadline and the unauthenticated-startup guidance.

Every behavior is failing-first and ablation-proven; the toggle-off import
boundary and the accepted loss of unknown-control visibility rows are unchanged.

Claude-Session: https://claude.ai/code/session_01Pqjduxt5G4rr9aYvtp7rNm

* fix(claude): arm the descendant snapshot before stdin closes and make the tree verdict unproven by default

A healthy Claude root leaves within the graceful window, and the close ladder
only snapshotted descendants when the root was still alive after that window.
So the common close never looked at the tree: `treeExited` stayed null,
`!== false` passed it, and close() reported a proven exit with an MCP child
still running. A root that died before the walk made the snapshot vacuous too.

The proof is now unproven by default. The reaper holds one verdict in Orca's
vocabulary (exited / live / unverifiable), assigned in exactly one place from
the bounded verification, and close() returns true only on `exited`. The
snapshot is armed before stdin closes, while the root can still be walked, and
is verified after the root exits; a root that left before any snapshot could
be armed stays unverifiable rather than vouching for descendants it never
showed us. The shared verifier gains the three-way verdict behind its boolean
face, and the connection reports the root and tree verdicts separately along
with the child's exit status.

One verification per close attempt: the retried close re-verifies, so the
intra-attempt re-reap is gone from the teardown budget.

Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9

* fix(claude): verify the Windows tree after taskkill instead of trusting that it ran

`terminateWindowsProcessTree` resolves from taskkill's callback whatever the
error says, so a timeout, an access denial, a recycled root and a surviving
descendant all looked identical to the reaper — which then returned a proven
exit unconditionally. close() reported true and the lease was released with an
MCP descendant potentially still live.

The Windows branch now snapshots the root's descendants while it is alive and,
after taskkill, polls a fresh process table to a bounded deadline: a row still
matching by pid AND creation time is `live`, an unreadable table is
`unverifiable`, and only a table with no match is `exited`. Creation time is
the PID-reuse guard the POSIX path gets from ps lstart, so a descendant that
denied a creation-time query is omitted rather than signalled on a bare pid.
A root already observed exited is never taskkilled: `/T /F` on a recycled pid
would take an unrelated tree down with it.

The captured tree is tagged by platform so neither verifier can be handed the
other's rows.

Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9

* fix(claude): release a reservation on a first-hand root exit instead of latching it into manual recovery

Making close() strict about the descendant tree exposed a second defect at the
same boundary. A create-time acquisition has no ownerProcess until publication,
so an unproven cleanup mapped to handoffStage `manual-recovery`, and
adjudication then refuses every later attach with agent_session_ownership_unknown.
A user who was merely signed out, or whose --resume the CLI rejected, wedged the
session id permanently.

Each question now answers from its own evidence. close() is unchanged and stays
strict about the tree. Separately, the lease is keyed on the root's pid and
start time, so when Orca's own child handle observed that root exit and no
descendant snapshot was ever admissible, the reservation is released and the
CLI's exit code and stderr reach the user. A descendant observed still alive,
or a root Orca never saw leave, stays unproven and keeps the reservation.

The settlement records only what was observed: the released lease says the
provider process exited and its descendants were not verifiable, rather than
reusing the wording that claims cleanup proved no child remains.

Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9

* fix(claude): surface an API error a result frame reports instead of settling the turn on it

The SDK models an API failure as a SUCCESS-subtype result whose `result` string
is the user-facing error text, with no assistant frame behind it. The translator
suppressed every catalogued result subtype as turn bookkeeping, so that turn
tombstoned its lifecycle and showed the user a completed, empty reply with no
sign anything had failed.

Suppression is now by meaning. A result reporting a failure routes to the
bounded provider-error surface, leading with the provider's own sentence and
keeping the raw frame behind the row's disclosure; ordinary successful results
stay off the timeline as before. A turn the user aborted also stays suppressed:
its interrupt frame already says so, and its execution diagnostic would only be
noise on every stop.

Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9

* fix(claude): drop the stream state of turns that never received their final frame

Every streamed delta recorded its block's identity, latest text and checkpoint
length. Only the final assistant frame removed them, so an interrupted turn left
its whole accumulated reply reachable until the session was disposed, and a long
session with repeated interruptions grew those maps without bound. The partial
text was already journaled by the flush that precedes settlement, so the live
copy was pure retention.

That state now lives in its own module, named for what it does — grow a streamed
block's journal row between its deltas and its final frame — and turn settlement
drops every block still awaiting a final. The translator reports how many remain,
which is the invariant: a settled turn leaves none.

Also makes a timed-out process-table read retryable while the root is still
alive. A loaded host can miss the table's one-second deadline, and latching that
as "no descendants" both lost the descendant sweep and, on a busy machine, made
the close ladder report unproven for a tree it never actually looked at. Only
the root's death still makes a missing snapshot final.

Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9

* perf(claude): capture the Windows descendant tree from one process-table read

The capture walked the descendant tree and then read the table again for the
creation times the walk's projection drops. Each read is bounded in seconds and
both run inside the close ladder's budget, so the second one cost the worst-case
teardown three seconds for data the first read already held.

The walk is now exported from the module that owns it and runs over rows the
caller has already read, which is also what lets the snapshot keep the
PID-reuse guard the projection cannot carry.

Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9

* fix(pty): spend the descendant verification window instead of surrendering on one slow table read

The verification abandoned the whole check the first time a process-table read
missed its own one-second deadline, with seconds of its window still unspent.
On a loaded host that reported a tree unverifiable without ever having looked at
it, which the Claude close ladder then turned into an unproven close and a
retried teardown. It also made the descendant-exit tests flake under a parallel
suite run, for the same reason and with the same honest-but-premature verdict.

A read that missed its deadline is now simply not an answer: the loop waits and
reads again until its own deadline, and only a window that ends without a
readable table reports unverifiable. This can only turn a premature verdict into
one backed by evidence; it never manufactures a proof.

Claude-Session: https://claude.ai/code/session_01BSmXgkWSsNHft8jFkdBFG9

* fix(claude): never let a later failed look collapse an observed live descendant into unverifiable

The reaper's single assignment site latched only 'exited', so a second reap
whose table reads all missed their deadline overwrote an earlier completed
verification's 'live' with 'unverifiable'. The acquisition release gate
discriminates on exactly that pair, so a root exit after such a decay released
the lease over a descendant that had been observed alive. The latch is now
monotone in trust order: exited is final, and live is only ever raised to exited.

Claude-Session: https://claude.ai/code/session_01HfdhsvSJucLw4cTZxzg2CP

* fix(claude): never prove a Windows tree gone while a descendant denied identification

The Windows snapshot dropped rows that denied the creation-time query, and an
emptied snapshot was judged exited without any table read: a descendant Orca was
refused information about was treated as one that had left. The snapshot now
counts the unidentified rows it saw, and verification caps its verdict at
unverifiable while any exist. Nothing is ever signalled on a bare pid, as before.

Claude-Session: https://claude.ai/code/session_01HfdhsvSJucLw4cTZxzg2CP

* fix(claude): classify cleanup after a first-hand exit as a root exit instead of a proven tree

When the CLI died between a successful acquire and the host's commit or proof
of the lease, handleExit had already removed the session, so releaseAcquisition
found nothing and reported true. The attach flow then settled exit-proven with
deathEvidence claiming cleanup proved no provider child remains, though the
tree was never verified. The adapter now keeps the exit that removed a
published session until the session is acquired again; acquisition cleanup runs
that connection's close ladder and classifies its verdict exactly as a
start-time failure would be, so the record reads root-exit-observed. The wire
helper keeps that typed classification and its provider diagnostic instead of
wrapping it as unproven, and the router gives up its owner even when the
release throws.

Claude-Session: https://claude.ai/code/session_01HfdhsvSJucLw4cTZxzg2CP

* fix(claude): integrate SDK teardown and picker lifecycle fixes

* fix(claude): preserve resume leaf and settle processless spawns

* fix(claude): reacquire from persisted resume leaf

* fix(native-chat): restore Claude grouped question handling

* fix(claude): persist only resumable transcript leaves

* fix(claude): recover structured session exits safely

* fix(claude): close remaining structured session P1s

* fix(claude): harden transcript branch proof

* Remove superseded root fix reports

* fix(windows): restore indexed descendant row walk

* fix(router): forward force-close lifecycle

* fix(claude): fence stale turn cancellations

* fix(claude): fence cancellation after unknown dispatch

* fix(claude): fence replay and option recovery races

* fix(claude): block replay fallback after waiter eviction

* fix(claude): fence evicted slash results

* fix(claude): fence ambiguous results and restore options safely

* fix(claude): scrub SDK child env and localize pending launch

* fix(claude): pin transcript roots and exit recovery proofs

* fix(claude): retain unproven SDK exits

* fix(claude): settle retained exit before reacquire

* fix(claude): resume from settled retained cursor

* chore: remove tracked review artifact

* fix: harden Claude SDK transport session cleanup

* fix: close Claude sessions safely

* fix(claude): close races with fresh child snapshots

* fix(claude): fail closed on recycled child identities

* fix(claude): gate root cleanup on process identity

* fix(claude): fence same-second root identity reuse

* fix(claude): restore the root SIGKILL fallback the identity gate took away

The direct root kill goes through the handle Node owns, not through a pid:
libuv drops that handle in the same turn it reaps, so the signal either
reaches the process Orca spawned or reaches nothing at all. Gating it on a
process-table probe therefore bought no safety and cost the tree its only
fallback whenever the probe declined -- a first capture landing in the fork's
own second, a recycled descendant pid voiding the snapshot, or a process table
that could not be read on either platform.

Identity verification stays where a bare pid is genuinely addressed: Windows
`taskkill /T /F`, and the descendant sweep's own revalidation before it signals.

Also stops a declined root probe from collapsing an observed `live` or `exited`
descendant verdict into `unverifiable`, and stops a successful taskkill from
reporting `unverifiable` because a later probe found the root correctly dead.

* docs(claude): rewrap the root-kill ordering comment

* Match the Claude structured launch to the terminal path's managed-account auth rules

The SDK path stripped ambient Anthropic auth unconditionally, let an explicit
agentDefaultEnv override beat a pinned managed account, and had no account-switch
guard. Reuse the terminal preflight's own predicate and messages so both transports
strip, refuse, and report identically, and cover the CLI transcript location that
mobile native chat depends on.

* Reach the Claude structured chat lane from the desktop UI

The main process has had a complete, correctly gated Claude Agent SDK lane for
a while, but no renderer ever asked for it: the launch route accepted only
`codex`, and the create path was typed `agent: 'codex'` end to end.

Widen both to the structured provider union that already exists
(`AgentSessionHandleProvider`), and generalize the codex-named create path
instead of adding a Claude twin beside it. The pending-launch registry is now
keyed by agent as well as workspace — a shared key handed a second caller the
first agent's intent, so a Claude and a Codex launch in one worktree collided.

Windows, per agent. Codex's client-side win32 refusal is deliberate and settled
elsewhere, so it stays exactly as it was. Claude's answer is no longer guessed
from the client's platform: a structured session fences its provider child on
that child's process start time, and only the executing host knows whether it
can read one. `agentSession.createSupport` already answers precisely that, per
agent, and had no renderer caller — so the Claude create path asks it before
creating and turns a "no", or a probe it cannot get answered, into the
definitive refusal the launch fallback already handles. Fail closed either way.

That refusal mapping also closes a real gap: the host reports an unsupported
location by throwing `structured_agent_session_unsupported`, which reaches the
client as a transport rejection rather than a refusal envelope, so
`StructuredAgentSessionCreateRefusalError` never fired. The launch would retry
the create, strand itself in `visibilityUnknown`, run no legacy fallback, and
show an error toast.

Close a fail-open hole while Claude and win32 become reachable: `create` with a
client-supplied location, and `ensure`, both skip the worktree-resolving support
check. They now ask the executing host the same question directly, so a host
that cannot fence a provider child no longer creates one on a client's say-so.

Also deletes `structured-agent-session-provider-routing.ts`, a duplicate of
`structured-agent-session-provider-support.ts` with no importers.

WSL, SSH and paired hosts, floating workspaces, draft prompt delivery, explicit
TUI customization and initial session options all keep refusing; folder
workspaces keep working.

* P1-1: make the structured Claude auth policy required and testable

The optional dep plus a {stripAuthEnv:false} fallback meant a dropped wiring
under-stripped silently. Required at all three hops, asserted at install time for
the @ts-nocheck caller, and the settings-to-policy mapping is now a named tested
function.

* P2-3: mobile's default Claude transcript root must follow CLAUDE_CONFIG_DIR

session-file-resolver's default ignored the variable the pinned account home
follows, so a CLAUDE_CONFIG_DIR launch wrote one tree and mobile read another. The
Task-4 test now resolves with no root override (mobile's own call) and checks the
answer against the root the CLI itself reports, instead of mirroring the code under
test's own expression.

* P2-1/P2-2/P3: close the teardown window, join the live-auth gate, align the refusal

P2-1: a switch beginning inside the acquire teardown left a dead chat and no
replacement. Past that point the launch waits the swap out and refuses only if it
never settles; the entry guard still refuses outright, because nothing is torn down
there yet.
P2-2: structured children now hold the same OAuth-refresh gate a Claude PTY does,
so a managed refresh cannot rotate the token out from under a live turn.
P3: the refusal now matches the strip it guards (case-folded on win32, presence not
truthiness), and the dead structured-to-TUI builder states its auth policy instead
of silently signing a system-auth user out.

* Make the live-auth gate tests independent of sibling connection teardown order

* Do not offer structured Claude under a WSL-only managed account

Structured Claude launches against the ambient Claude config, which the account
service keeps in sync with the selected HOST account. A WSL-bound managed
account lives inside the distro and is never synced there, so on Windows a
structured session would authenticate as whatever the ambient identity happens
to be while the UI names the WSL account — the user is told one identity and
given another.

That was unreachable only because nothing offered structured Claude on win32.
Enabling it makes it reachable, so gate it here rather than patching the auth
layer: refuse the structured path when the active managed Claude account is
WSL-bound, and let the terminal-backed path — which resolves the account per
runtime — handle that account shape.

The answer rides the agentSession.createSupport seam the renderer already
consumes, so no new capability and no renderer knowledge of account internals.
A create the host declines becomes the definitive refusal the launch fallback
already turns into a legacy native chat tab, with no error toast.

Unknown answers refuse. An install with no managed accounts claims no identity
and is fine, but an active selection that cannot be resolved — or account state
that cannot be read at all — is not evidence that the ambient identity is right.

Claude only. Codex resolves its account through a different path and its
createSupport answer is untouched, as is every Codex routing decision.

* Read the structured Claude account gate through the auth policy's accessor

The gate resolved the active account from the account-service snapshot's
runtime map; the auth policy resolves it with
getSelectedClaudeAccountIdForTarget(settings, { runtime: 'host' }). Those are
two sources and two resolution rules, and they disagree on a legacy settings
blob that carries the selection only in the flat activeClaudeManagedAccountId:
the accessor falls through to it, a direct read of the runtime map does not. The
gate would then refuse a launch the policy would have run under host-1 — and in
the mirror case a session could be admitted under a policy computed from a
different account than the gate approved.

Read the same settings through the same accessor so agreement is structural
rather than coincidental, and drop the controller accessor that existed only to
reach the snapshot.

No behaviour change for any state both already agreed on; Codex is untouched.

* Round-3 review fixes: N-1 empty-value regression, N-2 gate leak window, N-4 lost history

N-1: my presence-based conflict predicate refused a terminal launch that works
today. 'ANTHROPIC_API_KEY=' is how a user blanks a variable and the settings
pipeline preserves that empty value; an empty override cannot beat the pinned
account and the strip removes the name anyway. Back to truthiness for the value,
keeping the win32 case folding.
N-2: enter the live-auth gate only after the exit/close handlers that release it,
so no throw in between can leave an entry nothing reconciles.
N-4: the Claude transcript resolver searches config-dir-then-default and de-dupes,
matching the Codex sibling in the same file, so adopting CLAUDE_CONFIG_DIR no
longer hides history written before it.

* Run the managed-account gate on every Claude acquisition, not just create

createSupport gates the create path, but a session's account state can change
while it lives. A reacquire after an unexpected child exit re-resolves the
launch and re-derives auth, with nothing re-checking the gate — so a session
created while supported could come back up in the refused shape. With the strip
predicate keyed on there being an active non-WSL account, the WSL-only user's
normalized steady state (accounts exist, none active) does not strip, and that
reacquire reaches the child with ambient auth while the UI names the account.

Gate at resolveLaunch, the one choke point every acquisition passes through,
refusing with the pre-spawn error the caller already handles. Same predicate as
create-time, now sharing one settings reader so the two cannot drift.

Claude only; Codex resolves its account on a different path and is untouched.

The runtime class that wires this does not typecheck its own `this` calls — a
missing hookup compiles clean — so the wiring is pinned behaviourally rather
than trusted to the compiler.

* Move the structured Claude gate out of the @ts-nocheck runtime files

Both call sites of the managed-account gate sat in files whose first line is
`// @ts-nocheck`, so neither was typechecked: three arguments to a one-argument
function plus an undeclared identifier compiled clean. New auth-identity
decision logic had no compiler behind it.

Move the verdict into a checked module that takes the two facts the runtime
owns — the adapter's answer and a settings getter — and decides. The runtime
class now only forwards. Move the gate reader's construction into the checked
installer too, so the nocheck file passes a plain settings closure and never
names a gate symbol.

Every reference to the gate predicate and its reader now lives in a checked
file, so the ablation that used to pass silently is a compile error at both the
create-support and reacquire sites.

Removing the file-level @ts-nocheck is a separate, larger job and is not
attempted here.

* Derive the gate test's auth policy from the settings under test

A hardcoded stripAuthEnv asserts a gate/policy pairing production cannot
produce, and false additionally lets launch.env inherit the runner's real
process.env. Derive via claudeStructuredAuthPolicyForSettings instead: the
gate settings type is the same Pick the policy takes, and both resolve the
account through getSelectedClaudeAccountIdForTarget.

* Pin the absent-vs-empty distinction in the managed-account gate

An empty claudeManagedAccounts array is a real answer: the user has no managed
accounts, nothing claims an identity, and the ambient path is legitimate. A
readable settings object with no such field is settings we failed to parse —
the same unknown as unreadable — so it refuses.

The two are one character apart in the code and the difference is invisible
without the reasoning, so record it at the branch and pin both sides. The test
fails under the obvious "consistency fix" of treating a missing field as empty.

* fix(claude): keep command queue bookkeeping out of the transcript

Claude Code 2.1.258 emits a `command_lifecycle` frame for every uuid-stamped
command it starts, completes or cancels. The frame carries a command uuid and a
state and no content, and the CLI keeps it out of its own transcript -- but it
is absent from the SDK's SDKMessage union and so from Orca's frame catalogue,
where an uncatalogued kind defaults to a substantive row. Every structured turn
therefore painted raw JSON rows into the user-visible transcript.

Catalogue it and disposition it as status chrome. The unknown-kind default stays
`timeline-substantive`: a kind we have never seen is likelier to carry content
than to be chrome, and a visible row we can catalogue later beats content we
silently dropped. A lifecycle state that reads as a failure still surfaces,
because the payload error check in `classifyProviderFrame` outranks the
catalogue.

* fix(claude): let a re-walked descendant become eligible for the forced sweep

A descendant first observed by a capture inside its own birth second could never
be SIGKILLed: `ps lstart` is second-resolution, so that capture cannot rule out
a pid recycled later in the same second, and the merge pinned each retained row
to the boundary of the walk that first saw it. SIGTERM-resistant children forked
in that window were signalled and then never escalated -- they survived close,
quit and restart, reparented to init, and had to be killed by hand.

Advancing that boundary on any later capture would be unsound: a later capture
matching pid, pgid and start-second is exactly what an impostor would also show.
But a capture is not a match -- it is a fresh ppid walk from a root Node pins
through its own handle, so a row it re-derives is proved ours at that instant
without appealing to its start time. Chain the fence from there instead, and
take that walk at the close boundary while the root certainly still lives: the
root may leave inside the grace window, and the post-timeout refresh never runs.

A row absent from the later walk still keeps its earlier boundary, and a row no
walk has ever re-derived in a later second is still never escalated.

* Treat an absent managed-account list as empty, not as unreadable

An empty claudeManagedAccounts array and a missing one are the same answer:
this user has no managed Claude accounts, so nothing claims an identity and
ambient auth is the truth. Refusing on absence strands any profile that simply
never wrote the key, and it disagrees with the auth policy, whose own predicate
takes `(accounts ?? [])` for exactly this reason.

Only settings that cannot be READ stay unknown, and those still refuse — as do
a WSL-bound active account and a selection naming an account the list does not
explain.

The earlier reasoning treated a missing field as settings we failed to parse.
That conflated "not present" with "not readable"; only the second is unknown.

* Support structured Claude when accounts are registered but none is selected

Registered-but-deselected Claude accounts were refused, which is behaviourally
identical to having no accounts at all: the auth policy does not strip, ambient
auth is the truth, and the UI names no host identity. A user who deselected
their accounts silently got legacy chat with nothing explaining why.

Nothing selected for the host runtime is two states the settings cannot tell
apart after the fact, because pruneInvalidClaudeRuntimeSelection empties the
host slot and persists null in the second one:

  honest deselection      -> ambient auth, UI names nothing   -> SUPPORTED
  the WSL-only steady state -> ambient auth, UI names the WSL account -> REFUSED

The presence of any WSL-bound account in the list decides. Simplifying this to
"none active -> supported" re-opens the auth-identity misrepresentation, so the
tests fail loudly on exactly that: five of them, across the unit rule and the
createSupport path.

* Stop treating an unanswerable create-support probe as a refusal

A worktree is not resolvable for a beat after createWorktree resolves, so a
probe fired immediately after creation fails the RPC with selector_not_found
instead of answering. The catch collapsed that into `supported = false`, so the
composer refused and quietly built a terminal session — the gate never said no,
it was never asked successfully. Elapsed time was the only input that decided
whether a Claude launch went structured.

"Could not answer" and "answered no" are different states and only the second
is a verdict. Retry while the host cannot yet resolve the selector, with a
bounded backoff that covers the measured window with margin, and keep refusing
on the first ask for everything else. Fail-closed is unchanged: a probe that
still cannot be answered when the budget is spent refuses.

The retry is narrowed with the shared error-code matcher, which classifies a
token that transports re-wrap into a longer message without matching prose that
merely mentions it.

Codex never probes, so this race has never been able to refuse a Codex launch —
the race itself is identical for it. Recorded at the early return, because
whoever gives Codex a probe inherits the bug.

* fix(claude): fence the forced sweep on re-derivation, not on lstart's second

A descendant forked in the same wall-clock second as every walk that sees it was
signalled with SIGTERM and then never escalated, so a SIGTERM-resistant child
survived tab close, app quit and a full relaunch. Two children of one parent
96ms apart across a second boundary took opposite paths. The leak predates this
branch: it reproduces with the change reverted.

`ps lstart` has one-second resolution, so a walk landing inside a row's birth
second can never rule out a pid recycled later in that same second. But a walk
is not a match: a ppid walk only reaches what the root actually parents, and the
root is pinned by Node's own handle, so a row the walk re-derived is ours
whatever second it was born in -- a stranger would have to have been forked into
our tree, and then it is not a stranger. Fence the escalation on that.

Rows a merge retained from an earlier walk are not re-derived and still answer
to the start-time fence, which remains correct for them.

Scoped to callers that revalidate identity before signalling, which is the
Claude close path. Codex teardown reaches this same verifier and is unchanged;
the argument holds there too, but widening it is its own deliberate change.

Also reverts two changes from the previous attempt at this leak. Advancing the
capture boundary on a later walk is inert once the sweep fences on re-derivation
-- both key on the same set of rows, so the new term short-circuits for exactly
the rows whose boundary it advanced. The extra ladder refresh was a duplicate
full process-table read: close() already awaits tree.refresh() immediately
before proveClaudeChildExit, on the only path that reaches it.

Known property: the kill lands roughly a grace window after the walk that proved
membership, so a pid recycled inside that gap could in principle be signalled.
It is bounded -- matchingSnapshotRows already requires the live row to carry the
same start-second and pgid, so an impostor must be born in the remainder of that
one second, land on that exact pid, and sit in the same process group, and it
has already received the unfenced SIGTERM from the same loop.

* Run the Claude structured integration suite as a runtime client

The suite exercises agentSession.* for Claude, not the mobile surface: nothing
in it asserts anything mobile-specific and its sibling integration suites use
'runtime'. Mobile now additionally requires the experimental structured-chat
setting, which structured-agent-session.test.ts pins in both states, so the
stale 'mobile' fixture was claiming coverage it never had.

* fix(claude): report effort from get_settings, which is the only frame that has it

The composer's Effort pill rendered blank in every structured session. This is
not a missing source: the publication reads `effortLevel` off the `system/init`
frame, and that frame has never carried an effort of any kind, while the correct
value is already fetched at acquisition and thrown away on the auth diagnostic.
Verified two ways -- a live get_settings probe against Claude Code 2.1.258, and
the shipped binary's own init frame construction, which lists `model` and no
effort. So `reportedOptions.effort` was always empty, the options reader dropped
the key, and the pill had no value. Model survived only because
`currentModelId()` has a fallback chain.

The get_settings call acquisition already makes reports the session's current
effort as `effective.effortLevel`; pass that into the publication instead.
Selecting an effort already worked, so this is the arrival value only.

The legacy PTY path is unaffected and must not be "fixed" to match: it reads its
effort by parsing the startup banner (`CLAUDE_MODEL_EFFORT` in
src/renderer/src/components/native-chat/claude-terminal-session-options.ts),
which is why it shows a value where the structured path does not.

Also removes the fixture that hid this: the fake init frame invented
`effortLevel: 'high'`, a field the CLI does not send, which is why every gate
stayed green over a value that is always empty in production. The fixture's
get_settings now returns the real {applied, effective, sources} shape instead of
a bare `{env: {}}`, so the two adapter tests that asserted an effort keep
asserting it through the path production actually uses.

The reader returns null rather than defaulting: an effort nothing measured would
repeat the fixture's mistake, and a blank pill is the honest degradation if the
provider ever renames the key.

* fix(claude): only record an effort the child confirms it adopted

apply_flag_settings answers `success` for an effort it then ignores. Measured
against Claude Code 2.1.258: applying `bogus-effort-xyz` returns
subtype "success" with no error while `applied.effort` stays at its previous
value, and a valid `low` moves it. The option write treated the absence of a
throw as adoption and recorded the requested value unconditionally, so Orca
would show and persist an effort the child was not using, with nothing anywhere
reporting a problem.

Read the effort back after applying it, through the same reader the arrival
value uses, and reject when the child reports a different one. A readback that
could not be taken is not evidence of a refusal -- the apply itself succeeded --
so it still records; only a readback that disagrees rejects.

Not reachable from today's picker, which offers catalog values only, but the
CLI's effort catalog is server-delivered and has changed before, so a retired id
would otherwise become a pill confidently displaying a setting that never took.

* test(claude): assert the effort contract against the real binary

The blank pill survived every gate because the only tests that touched it were
fixture-backed, and the fixture invented the field. A test that pins the shape
we read cannot catch the provider renaming the key, which is the failure mode
that produced this defect.

Asserts both halves against a live authenticated CLI: that no frame it publishes
carries an effort at all, and that the session's current effort arrives through
get_settings. Which frame proves the session varies by host -- this machine
proves it with a SessionStart hook rather than a system/init frame -- so the
negative half asserts over every published frame rather than picking one.

Skips with the rest of the file when no authenticated CLI is present.

* fix(claude): stop the synthesised content-part kinds leaking into the transcript

Sending an image put a bare `claude · message:user:content:image` row between
the user's bubble and the answer. Two causes, and only the second is a family.

An image part counted as modelled only when `source.type === 'url'`, but
claudeDispatchMessageContent sends a local attachment as a base64 source and the
CLI replays that shape back, so every attached image was classified unmodelled.
Accept the base64 and file sources Orca itself sends.

The family is the real defect. `message:<role>:content:<type>` kinds are
synthesised at runtime from whatever `part.type` arrives, so unlike the
top-level frame catalogue they can never be enumerated ahead of time -- the
`?? 'timeline-substantive'` default then prints the synthesised name at a user
who cannot act on it. That default is right for top-level frames, where
"substantive" means show the frame; here it meant show our own vocabulary, which
drops the content AND leaks the opcode.

So an unrenderable part now renders a sentence saying exactly that, with the
kind and payload still on the row's disclosure. A part that carries its own
readable sentence keeps it -- the placeholder is a fallback, not an override.

An unknown future part type is therefore visible, never silently dropped and
never printed as a kind: the same principle as the effort readback, which
records only what the provider confirms.

* Declare agentSession.requestHandoff on the cross-version wire surface

The manifest is a ratchet for cross-version reachability, so the method is
declared with real HandoffParams rather than counted. requestHandoff is
capability-gated through requireStructuredHost and has no client caller, so
declaring it is the whole of the change.

Also model two host capabilities the harness omitted: the stub host's
supportsCreate, and the fake adapter's, without which adapterSupportsCreate
falls through to a supportsLocation the fake also lacks. Every ensure was
refused for the harness's silence rather than for its location.

* Gate structured Claude session tabs on the client capability that names them

The Claude structured lane deleted the projection's `agent !== 'codex'`
filter and added CLAUDE_STRUCTURED_AGENT_SESSION_RUNTIME_CAPABILITY in the
same commit, but never wired the constant to anything. Paired clients then
received agent-session tabs for Claude, which no shipped client renders --
mobile's resolveMobileNativeChat returns null for every agent but codex, so
the row listed and selected into a pane with neither chat nor terminal.

Restore the filter behind the declared capability instead of the bare agent
name. No client advertises it yet, so this matches main's behaviour today
and becomes a negotiation a future client can opt into.

* Confirm the structured Claude model against the model the CLI reports

set_model answers success for any string, including a model it cannot
resolve — the failure only surfaces when the turn runs — and get_settings
reports the settings-file model, not the session's. The init frame that
opens each turn is the only channel carrying the adopted model, so keep
the session's reported model current from it instead of reading it once
at acquisition.

Also stop rejecting an effort the readback cannot represent: max is
session-scoped and excluded from the persisted effortLevel, so a readback
reporting the level underneath it is an absence of evidence, not a refusal.

* Clear the session-option hedge when the provider confirms the value

The pill claimed every option was unconfirmed for the life of the session:
the renderer recorded each write as dispatched and nothing ever moved it,
so a model the CLI had already reported back still read as unconfirmed.

Carry the provider's own confirmation to the surface. Main reports which
option ids the provider named rather than merely accepted, and the client
re-reads options as a turn changes, because the frame that opens a turn is
where the adopted model arrives. A value the provider has not reported
stays hedged, including an effort whose readback could not be taken.

The confirmed list is optional on the wire: a host that predates it sends
nothing and the client keeps hedging, which is the behaviour it had.

* Keep the model report current across an acquisition fence bump

* Show the picked session-option value and let the provider report correct it

The pill showed a "not confirmed" second tooltip line for any value we had sent
but not yet seen reported back. Nothing acts on it, and for the PTY lane it was
permanent — that transport has no report channel. The pill now shows the picked
value immediately and the provider's per-turn report corrects it when the two
disagree; a newer local write still outranks a report that precedes it.

`dispatched` stays as a provenance member rather than collapsing into `applied`:
it is produced independently by the PTY lane, and it is where the `confirmed`
wire field lands, which would otherwise be unobservable.

Effort keeps its readback and its rejection path. That matters more now, not
less: with the hedge gone the rejection is the only user-visible failure signal
on this surface, so a spurious one would be the loudest bug here. Skipping the
readback for an effort the settings response structurally cannot echo is what
prevents it — the response carries the persisted level, so reading it back for a
session-scoped value would report the level underneath and fail a valid write.

* Hedge a session-option value only when the terminal transport sent it

Both lanes emit `dispatched`, so it could never say which one produced a value.
The descriptor now carries the transport that built it, set once in the shared
snapshot builder from a parameter that is required rather than defaulted — the
builder is the only place a descriptor is constructed, so a new producer has to
name its lane or fail to compile.

The structured lane confirms every value from the provider's own per-turn report,
which makes the hedge transient noise there. The terminal lane can only learn an
outcome by parsing the screen back, and only for Claude: every other agent's
`dispatched` value stays unconfirmed for the life of the session, so the line is
the only signal that we sent something we never saw land.

* Refuse an effort the session's model advertises no control for

* Refuse tab mutations on a Claude row the client never negotiated

The branch added a case asserting a client advertising only
agent-session.structured.v1 may mutate a claude row. That is the same
ungated behaviour the projection gate removes, encoded a second time —
mutation authorization reads the projection, so hiding the row refuses
the write. Assert that contract instead, and add the positive case for a
client that does negotiate Claude rows.

* Resolve the Claude session's current model in one place so the effort guard and the pill agree

* Record an effort the child did not adopt instead of refusing the write

apply_flag_settings answers success for an effort it then ignores, so the
readback exists to detect that. Refusing on it made the detection a veto,
and a veto is only correct if the readback can never be wrong about which
model is current -- which it was, twice. The pre-flight guard already
refuses a level the model advertises no control for, so the veto guarded a
door that is now locked upstream.

Keep the detection, drop the refusal: a disagreement records the child's
own answer and omits the option from confirmed, so main stops vouching for
a value the provider rejected without blocking the user's write.

* Stop a slow whole-machine ps from being read as an absent process

`ps -axo ...command=` pays a per-pid argv read: measured 1.15s for 1,948
processes (0.03s without `command=`), and CPU contention stretched the same
capture to 6.0s. Two budgets sized for a cheap look then misreport a readable
machine.

The reader's 3s ceiling killed 6 of 20 consecutive captures at load 27, so
every consumer answered "unverifiable" about a table it could read. Raise it
to 15s, and stamp the capture instant at ps START so `capturedAgeMs` is the
upper bound its contract promises -- a 6s capture used to report itself as
freshly taken, understating staleness against a 5s kill gate. The TTL keys on
completion so a slow capture still coalesces instead of forking ps per caller.

`readStructuredTuiProcessIdentity` then spent its whole 5s wait inside one
capture and concluded "no exact child" after a single look taken before the
child existed (observed landing at ~3.5s). Absence needs a look that did not
race the spawn, so require two captures before the deadline can end the loop.

Both surfaced by the real-binary Claude TUI resume test, which failed ~1 in 5
under load; 14/14 now, 8 of those runs containing a capture the old 3s budget
would have killed.

* Let the desktop renderer negotiate Claude structured tabs

The paired-client gate hides agent-session rows an agent the client cannot
render. The desktop renderer's own IPC dispatches as clientKind 'runtime'
advertising only agent-session.structured.v1, so the gate hid Claude rows
from the surface this feature ships on. It renders them; it should say so.

* Stop a slow process table from silently blinding every freshness gate

Stamping `capturedAgeMs` at ps START made the number honest, and honest broke
both consumers that read it. `ps -axo ...command=` measured 2.5-9.0s on an idle
2,002-process laptop and 4.0-18.6s at load 46, so the age it now reports lands
past every budget: `planRelayPtySweep` refuses the stop as "too old", and the
renderer's `admitRemoteForegroundEvidence` refuses the record outright. That
second one is the expensive half and was outside the diff -- a refusal bumps
`consecutiveInspectionErrors`, the poll scheduler backs off to its 10s floor,
and agent-completion detection stops for the pane. The subsystem went blind on
exactly the loaded hosts the honest stamp was meant to serve.

The evidence-publishing read now gives up at 1,200ms instead of waiting out
`PS_TIMEOUT_MS`. It is one budget for one question: these consumers ask whether
an observation describes NOW, and past this it does not -- a late answer is
refused by the age gate anyway, having first blocked a polled path for the whole
capture, so a prompt `unverifiable` is both the truthful verdict and the cheap
one. Both relay call sites already produce it from a rejection, and an admitted
`unverifiable` costs a poll where a refusal costs the cadence. Identity proof
keeps the full 15s through `getFreshProcessTableSnapshot`, because it asks
whether a process EXISTS and must never read slow as absent. The budget bounds
the wait, never the capture: the reader coalesces, so an abandoned wait leaves
its capture running to fill the cache rather than forking a second whole-machine
`ps` on the host that can least afford one.

1,200ms is bracketed rather than picked. The floor is the capture's own cost --
`command=` measured 1.15s for 1,948 processes on an idle host, and a budget
under that answers `unverifiable` about a machine nobody is straining. The
ceiling is the consumer's: 2,000ms, less the 500ms a TTL-shared capture may
already have aged, leaves 1,500ms, and transit takes the rest.

That ceiling only fits once the capture stops being charged twice. `ps` runs
inside the RPC round trip, so its duration is already in `receiveDelay`, and
`capturedAgeMs` is that same duration on the host's clock; summing them halved
the budget this gate grants a host from ~2.0s of `ps` to ~1.0s, which is why a
1.2s capture arriving at 1.3s read as 2.5s old and was refused. Admission now
takes the larger of the two. The sweep's gate keeps its sum, which is correct
there: `evidenceAgeSinceListingMs` is stamped after the listing ARRIVES, so it
measures planning time and overlaps nothing.

A stated limit rather than an assumed one: 15s is not proven sufficient for
identity proof. The same capture reached 18.6s at load 46, so that path can
still time out and answer "no exact child" about a host it simply could not read
in time. Narrowing it needs a cheaper question than a whole-machine argv read,
not a larger number.

The one test guarding this field could not fail. `beginPtyHandlerTest` installs
fake timers, so `Date.now()` is frozen, the real reader reports exactly +0, and
`0 <= 500` held identically for a hardcoded zero, for completion-stamping and
for start-stamping -- while the real reader on that host returns thousands of
ms. It now drives a measured age in and asserts the handler publishes it rather
than restamping; that the reader MEASURES it correctly stays pinned separately,
against a controllable clock. Both consumers get boundary coverage either side,
and each new gate was ablated red before it went green.

* Keep the compatibility fields off the capture the budget just abandoned

inspectProcess falls back to processHasChildren and listProcesses to
getForegroundProcessName, and both read the same TTL-shared capture with
no budget of their own. On a slow host they joined the in-flight capture
the budgeted evidence read had just given up on, so the call still blocked
for the full 6-18s and the budget bought nothing -- once for inspectProcess
and once per managed PTY for listProcesses.

Use the degraded answers those helpers already give for an unreadable
table, reached promptly. pty.hasChildProcesses keeps its unbudgeted fresh
probe: it is a one-shot destructive gate that can afford to wait.

---------

Co-authored-by: Merge Sim <merge-sim@local>
Co-authored-by: Merge Sim <sim@local>
2026-09-04 15:55:20 -07:00
Brennan BensonandMerge Sim f4c2821167 refactor(agent-session-journal): move the session journal onto SQLite (#18652)
* refactor(agent-session-journal): move the session journal onto SQLite

The agent-session journal kept its state in three hand-rolled file formats: an
append-only `log.jsonl` with torn-tail repair, a `snapshot.json` holding folded
state plus a retained tail, and byte-quarantine files for anything unreadable.
This replaces all of it with one SQLite database per session — `journal.db`
beside the existing `blobs/` store — using the in-house adapter and the
open/pragma/migrate/harden pattern the orchestration database already follows.

Two tables: `journal_rows` (the append-only log, keyed by
`(session_id, epoch, seq)`) and `journal_sessions` (the derived projection,
upserted in the SAME transaction as every insert). Rows stay JSON in one
column, so the row schema, the version upcast chain, and the reducer survive
byte for byte — `journal-reducer.test.ts` and four other suites pass unchanged
and are the regression proof.

Deleted: `journal-log-file.ts`, `journal-compaction.ts`,
`journal-corruption-quarantine.ts`, and the public `compact()` /
`compactionBoundary` / `autoCompact` members, none of which had a non-test
caller.

Existing `log.jsonl` / `snapshot.json` journals are deliberately abandoned. No
importer: a session created on the old path stops working, which is acceptable
because the feature is off by default.

## The physical quota is repriced, because SQLite does not charge like a file

The 256 MiB per-session bound is unchanged, but the arithmetic under it could
not survive: SQLite grows the database in pages and the WAL in frames, and the
checkpoint that copies the WAL forward holds the same pages in both files at
once, so a transaction's peak is about twice its content. Admission now charges
the candidate transaction's own measured page cost, validated against a sweep
that runs as a regression test (`journal-database-space.test.ts`) rather than
derived from reasoning about the allocator.

Four things are load-bearing rather than tuning, each measured:

- `auto_vacuum = INCREMENTAL` must be set BEFORE `journal_mode = WAL`. Set it
  after and it is ignored with no error, reclamation silently becomes a no-op,
  and the file never shrinks again. Both halves are asserted.
- `wal_autocheckpoint = 0` plus an explicit `wal_checkpoint(TRUNCATE)` at the
  end of every write path, so the one moment the same pages live in two files is
  a moment the charge accounts for.
- Reclamation runs in bounded chunks. A single unbounded `incremental_vacuum`
  took a 252 MB directory to 504 MB — the reclamation added to defend the bound
  would have breached it. `PRAGMA incremental_vacuum(N)` also frees exactly one
  page unless it is stepped to completion, which no size assertion catches, so
  the freed page count is asserted directly.
- A blocked checkpoint leaves the WAL on disk together with the database growth
  it already copied, so admission charges that deferred copy explicitly. The
  term is zero whenever the last checkpoint succeeded, so the uncontended path
  admits and refuses an identical set.

The epoch discard is `DELETE FROM journal_rows` with no WHERE clause, which
takes SQLite's truncate optimization: measured at ~0.26% of the database in WAL
bytes where the `WHERE session_id = ?` form rewrote every emptied leaf at up to
99%. One database per session is what makes the unqualified form correct.

An open, empty journal costs 57,344 bytes before a single row exists, so a
configured quota below `JOURNAL_MIN_SESSION_BYTES` now fails loudly at open with
the existing `journal_bound_exceeded` instead of as a run of identical append
failures. No production caller configures one; the affected surface is test
fixtures, rescaled to the smallest value that restores what each case proves.

## One deliberate behaviour change

Compaction was the only mechanism that shed bytes inside an epoch, and the write
path called it precisely so an append at the bound was not refused. The
SQLite-shaped replacement — a bounded prefix delete — cannot be used: with the
snapshot gone the surviving rows ARE the state, so dropping the oldest of them
loses the oldest transcript silently at the next reopen. So no row is ever shed
inside an epoch, and a session whose row bytes alone reach the bound now refuses
every append where it previously compacted and continued. A loud typed refusal
beats silent data loss.

What still sheds is unreferenced BLOB bytes — the dominant and unbounded byte
source — on the same write-path hook. The escape from the hard stop is the fold
that already exists, `replaceEpochItems`, which now actually returns bytes to
the filesystem instead of leaving them on the freelist.

The prune's protected set is a union of live reducer digests AND the candidate
row's own digests, including those cited only by a nested lifecycle-batch
mutation. Content addressing never rewrites a digest already on disk, so
protecting live state alone deletes the blob the append is about to cite — a
dangling reference that surfaces one reopen later as an empty expansion on an
item the user can see. `journal-store-blob-budget.test.ts` pins it, and it goes
red when the set is narrowed back.

## Handle ownership

A file handle used to be opened and closed per append; a SQLite handle is held
for the session's lifetime. Every path that can open a connection now has one
owner: the open function owns its raw connection until it returns, the store
owns its retained one and releases it in a new `close()`, and every other
connection is closed by the call that opened it. The attach, recovery,
eviction, map-overwrite and host-teardown paths close what they drop, and host
teardown is failure-complete — the sink-barrier flush throws by design, so a
trailing close statement would be skipped on exactly the path that leaks.

`close()` has a stated contract: admission at enqueue and permanent, the close
step on the same queue past that gate, one shared in-flight attempt, fulfilment
terminal, and the release last and deliberately unguarded so a retry re-enters
it. Guarding the release would skip it on retry, guaranteeing a permanent leak
in exactly the case where it did not release.

`journal_closed` joins the error union for a write after `close()`; no file
outside the directory references any of these codes.

* fix(agent-session-journal): make a COMMIT final, stop repairs deleting valid rows, and keep rejected closes retryable

Six review findings on the SQLite journal migration.

1. A successful COMMIT is now the point of no return. The ordinary append,
   the epoch roll and the epoch replacement each adopt the committed row or
   epoch BEFORE any post-commit filesystem work; checkpoint, reclaim, blob
   prune and directory measurement run through `runJournalPostCommit`, which
   is best-effort by design and falls back to the transaction's own charge as
   a conservative footprint. Previously a post-COMMIT scan failure rejected a
   durable append and the next one reused its sequence, and a failed epoch
   housekeeping step left the store writing into a prefix already deleted.

2. Corruption repair preserves instead of destroying. A rejected suffix is
   copied into a new `journal_quarantine` table and removed from the live
   epoch in ONE transaction per chunk, charged against the session bound
   before a byte is written; a journal that cannot afford the copy refuses to
   open rather than falling back to deletion. The repair state is exposed as
   `journal.repair` and the rows are readable through
   `recoverQuarantinedRows()`, so Orca-owned submission, receipt and
   lifecycle identity survives a gap or a malformed row.

3. The physical charge covers the B-tree key payload. `session_id` and
   `epoch` are stored in both tables and both primary-key indexes and appear
   nowhere in `row_json`, so the journal boundary now bounds them and
   `journalTxnPhysicalCost` charges those bounds plus the projection upsert.
   The charge sweep runs the exact production transaction at maximum admitted
   key sizes.

4. A rejected `close()` no longer orphans its handle. Callers hand the
   journal to `agentSessionJournalCloseRetries` instead of swallowing the
   rejection, the attach map replacement is ABORTED when the previous
   journal will not close, host teardown retries what the registry holds, and
   a failed runtime teardown is retained so the next stop is a real retry.

5. `journalWalBytes()` returns zero only for ENOENT and propagates every
   other stat error, so admission and reclamation fail closed.

6. The WAL contention test closes the writer before removing its temp root
   and asserts the directory is removable once handles close.

Regression coverage: post-commit divergence (4), corruption repair (5),
key bounds (5), WAL stat (8), close retry (5), plus a runtime stop-retry
case. Each fix was ablated on this head and the matching tests go red.

* fix(agent-session-journal): anchor replay at sequence 1, make quarantine append-only, and charge it in bytes

Three ways the corruption quarantine still lost rows it was written to keep.

Replay validated contiguity from the first row that HAPPENED to remain, so an
epoch missing only its sequence-1 row declared the leftovers contiguous and set
no `truncateFrom`. The load was still corrupt, so recovery imported provider
history and `replaceEpochItems` deleted every live row — including Orca-minted
submission, receipt and lifecycle identity that no transcript can reconstruct,
and that nothing had quarantined. Replay now anchors at sequence 1, so a missing
epoch row rejects the whole surviving range before any replacement runs.

`journal_quarantine` was keyed on `(session_id, epoch, seq)` and copied with
`INSERT OR REPLACE`. A repair frees the sequences it removed and the live epoch
reuses them, so a second repair in the same epoch silently deleted what the
first preserved. The table is now keyed on a surrogate `quarantine_id`, the copy
is a plain append, and `(epoch, seq)` is metadata; existing v1 databases are
rekeyed in the migration that already bumps `user_version`.

The admission charge read `length(row_json)`, which counts CHARACTERS for a TEXT
value where `journalTxnPhysicalCost` expects physical UTF-8 bytes. A multibyte
suffix was charged at up to a third of what it writes, which defeats the
pre-write physical bound — over a megabyte on a maximum-size lifecycle batch.

* fix(agent-session-journal): keep a repaired epoch anchored and stop the v1 quarantine migration doubling the file

Replay validated numeric contiguity from sequence 1 but never that sequence 1
IS the epoch row. When the anchor was missing the repair set aside every
surviving row, and if provider-history import then failed — a transcript that
is temporarily gone is enough — the journal reopened as a clean, row-less
epoch: an ordinary append took sequence 1, replay accepted it, read-restore
published it as history, and automatic recovery never ran again while the
user's real messages sat in quarantine.

Replay now rejects an unanchored prefix, the open publishes an
`unreconcilable_prefix` anchor for an epoch its repair emptied, and that anchor
keeps reporting corrupt — so provider history is retried on every attach —
until the timeline is rebuilt or the session writes content of its own. A
repair also discloses rows it set aside when no line was unreadable at all,
which is the case that removes the most.

The v1 quarantine rekey copied every legacy row into the new table inside one
transaction and dropped the old one. A quarantine holds whole rejected rows: a
single 8 MiB row nearly doubled the database past the physical bound the open
had already checked, the dropped pages only reached the freelist, and the next
open refused the session it had just migrated. The v1 table is renamed and
frozen instead, and reads take both generations. Table creation also moves
inside the migration transaction, so a crash can no longer leave a v2-shaped
database still reporting version 0 for an older build to write into.

* fix(agent-session-journal): stop an empty provider transcript retiring the repair marker

A transcript that exists but decodes to zero messages was imported as a
success: the import published an empty `legacy_import` replacement that
deleted the `unreconcilable_prefix` anchor and its disclosure, so the next
probe read the session as clean and every later attach skipped provider
recovery while the user's rows sat in quarantine for good.

The import now leaves the epoch untouched when nothing decodes, reporting
`replaced: false`, and recovery treats that like a transcript it could not
read — the marker stands and a later attach with real history rebuilds the
timeline.

* style(agent-session-journal): merge the duplicate journal-database-space import

* refactor(agent-session-journal): drop quarantine, byte bound, blob spill and rate limit

Match what comparable implementations do: the journal is an unbounded
append-only SQLite log with no side tables and no admission control.

Corruption: the rejected suffix is DELETED rather than copied into a
quarantine table. The load still reports `corrupt` and recovery still
rebuilds the epoch from provider history, so the observable outcome is
unchanged — only the preservation half is gone. The schema is back to one
version with two tables; no v1 database exists outside unmerged commits of
this branch, so the rekey migration and the two-generation read path go with
it. Sequence-1 epoch anchoring and the empty-provider-transcript retry are
kept: both are about the corrupt signal being correct.

Size: no `maxSessionBytes`, so no page-cost arithmetic, reclaim band,
incremental vacuum, lifecycle byte reservations or `journal_bound_exceeded`.
`auto_vacuum` and `wal_autocheckpoint = 0` existed only to make a
transaction's physical cost predictable for that charge; with the charge gone
SQLite's default checkpointing is what the journal wants, and the explicit
pre-close checkpoint is redundant with the one `db.close()` performs. WAL,
`synchronous = FULL` and `busy_timeout` stay.

Payloads: an oversized body is truncated at the existing inline cap with the
existing marker and the remainder is discarded, bounded at the translation
layer that already calls these helpers. The truncation point and message do
not change; the content-addressed blob directory and all digest tracking do.

Rate: no `maxAppendsPerWindow` and no `journal_rate_exceeded`.

`JournalPayloadLimits` is now just the inline cap.

* fix(agent-session-journal): mark a partial repair pending and bound multi-block tool input

A repair that keeps its prefix had nothing durable to show for the suffix it
deleted: a sequence gap costs no malformed row, so no disclosure is appended,
and the surviving rows keep their epoch anchor. The next probe read a
contiguous anchored prefix, called it clean, and the deleted stretch of
timeline was never asked for again — silent loss, with the deletion already
committed. The deletion now writes a `journal_repairs` marker in the SAME
transaction, and replay keeps reporting corrupt while it stands. It retires
under exactly the rule the emptied-epoch anchor takes: a fresh epoch carries
the rebuild, or the session writes content of its own past the sequence the
repair left free. The repair's own disclosure is not that content.

Legacy import bounded a tool call's input only when it was the message's sole
block; the multi-block path returned `tool-call` unchanged, so a mixed message
from Claude, Grok or an omp execution cell persisted the whole input despite
`inlineHeadBytes`. `boundBlock` now routes it through `boundToolInput`.

Also drops canonical comments describing quarantine, snapshot files, blob
storage and blob compaction — none of which exist any more.

* fix(agent-session-wire): stop awaiting the synchronous journal probe

loadJournal runs on a sync-database connection and returns JournalLoad | null, so both wire call sites were awaiting a non-Promise. The type-aware code-quality gate flags it; the native gate does not.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-04 15:09:23 -07:00
Neil 9acfba401a fix(crash-reporting): stop claiming kills that never landed, and leave proof when the own-Chromium pid set is unreadable (#18578)
* fix(crash-reporting): stop the codex POSIX teardown claiming a group that was already gone

terminatePosixTree's default group signal swallowed every process.kill error and
then recorded a self_tree_kill unconditionally, so an ESRCH — proof the group was
already gone and this teardown killed nothing — still put a suspect in the
five-second render-process-gone attribution window.

Every sibling group-kill in the tree already records only on a proven signal:
terminateDedicatedPosixGroup in this same file, forceKillPosixPtyProcessGroups,
and the claude account-login teardown. This makes the outlier match them.

* fix(crash-reporting): leave proof when the own-Chromium pid set cannot be read

`readOrcaChromiumProcessPids` returns an empty set when `getAppMetrics()`
throws, which is the right decision — refusing every kill would orphan every
PTY, git, codex and notebook tree main tears down, and on main a refusal from
`killSourceControlAgentProcess` releases the managed-home lock with the agent
still alive. But the empty set was byte-identical to "no Chromium on this
host", so the fail-open was invisible in a field bundle.

Keeps the decision, adds a coalesced durable `own_chromium_pids_unreadable`
crumb so the two cases are distinguishable. Coalesced because the gate reads
this set on every tree kill.

* style(crash-reporting): tighten the group-signal comments to the WHY
2026-09-04 00:53:31 -07:00
Neil 7b530f1eb5 fix(crash-reporting): record Orca-initiated tree kills so a killed renderer is decidable (#18367)
* fix(windows): refuse tree-kills of Orca's own Chromium pids and record the rest

G2 is 20 field reports that share only a symptom. It is at least four
fingerprints: ~15 Windows `reason=killed exitCode=1`, 3 POSIX SIGKILL under
memory pressure (G4-oom), 2 duplicate reports of one macOS V8 Proxy Resolver
SIGKILL, and 1 `0x80000003` install-dir ACL crash (G1; #17740 ships in
v1.4.196 only, not 1.4.195). Nothing here claims to fix all of them.

Two changes:

1. Behaviour. `classifyWindowsTreeKillTarget` returns `own` for any direct
   child of the main process — which our renderer, GPU and network-service
   utility all are — so PTY teardown could `taskkill /T /F` Orca's own UI
   (#10680). Both that classifier and `terminateWindowsProcessTree` now refuse
   any pid Electron is currently accounting for in `getAppMetrics()`.

2. Diagnosis. An Orca-issued kill and an external one are byte-identical in
   every field the crash report records today, so the cluster is undecidable.
   Every main-process force-kill choke point now records a durable
   `self_tree_kill` breadcrumb, and `process_gone` reports carry
   `selfInitiatedTreeKills` naming the pid and its offset from the death.
   A refused kill records `self_tree_kill_refused_own_chromium`, which is
   falsifiable: if it ever shows up in the field, we were the killer.

* fix(crash-reporting): coalesce self-kill breadcrumbs and scope the discriminator

Round-1 review remediation. Three blocking findings, all accepted.

1. Breadcrumb flood (accepted). recordSelfInitiatedTreeKill wrote an
   uncoalesced durable crumb from two routine teardown paths, and the
   reviewer reproduced 12 terminal closes x 3 process groups completely
   evicting the 30-slot ring — including this PR's own refusal crumb — plus a
   forced writeSync per killed group. It now uses the existing
   recordCoalescedDurableCrashBreadcrumb (5s window for pid-addressed
   taskkills, 60s for routine group/job teardown), so a burst costs one ring
   slot and one flush. The refusal crumb is coalesced per victim pid, so a
   retry loop cannot flood while a distinct pid always gets its own crumb.
   Regression test replays the reviewer's exact 12x3 reproduction and asserts
   the refusal crumb and a pre-existing gpu_process_crashed both survive.

2. Undifferentiated count (accepted). posix-process-group and win-pty-job are
   structurally incapable of reaching a Chromium process, and scope was absent
   from the persisted string. Scope is now in every entry
   (`<scope>/<site>/pid<N> +Nms`), and the count is split:
   selfInitiatedTreeKillCount now counts only pid-addressed taskkills — the
   kills that can land on a recycled pid that is now our renderer — with
   pty-scoped sweeps in selfInitiatedGroupKillCount. The list is renamed
   selfInitiatedKills because it carries both, and sorts pid-addressed kills
   first so truncation never drops the discriminating ones for teardown noise.
   The reviewer's repro (routine macOS terminal close + unrelated exit-133
   crash) now yields selfInitiatedTreeKillCount undefined.

3. Recording gaps and a false comment (accepted). New
   admitSelfInitiatedTreeKill gate: it refuses own-Chromium pids and records
   the rest, and all three main-process taskkill families now go through it —
   terminateWindowsProcessTree plus codex-accounts/service.ts and
   claude-accounts (which keep their own spawn lifetimes). The false "single
   taskkill choke point" comment is gone. The runProcess choke point the
   investigation asked for is instrumented via a
   setProcessTreeKillObserver seam in src/shared/child-process — shared code
   runs in the CLI and relay so it cannot import the main breadcrumb store —
   registered in main preflight. The codex app-server POSIX group teardowns
   and the claude POSIX branch record too. The module doc no longer claims
   absence is discriminating: it enumerates what is instrumented and names the
   direct process.kill(-pid) sites that are not.

Non-blocking, also fixed:
- Breadcrumb calls moved out of the try blocks whose catch is the ESRCH
  contract (posix-pty-process-groups, codex teardown, claude POSIX), so a
  throw from the diagnostic path can never be reported as a failed kill.
- Detail truncation now bounds the first entry too, matching its comment.
- own-chromium-tree-kill-refusal.test.ts renamed to
  own-chromium-tree-kill-guard.test.ts, colocated with the module it tests.

Not changed, with reasons:
- Date.now() vs performance.now(): kept. Offsets are computed against
  goneAt = Date.now() in process-gone-recorder; a monotonic clock here would
  make the offsets meaningless. The reviewer verified this and agreed it is
  not a defect.
- app.getAppMetrics() per force-kill remains unbenchmarked. It reads
  in-process browser state rather than enumerating the OS process table, and a
  TTL cache would let a recycled pid slip past the refusal, so it stays
  uncached.
- The ~15 remaining direct process.kill(-pid) sites (browser routes,
  notebooks, automation prechecks, ephemeral VM recipes) are not instrumented.
  Rather than claim coverage this PR does not have, the module doc names them.

claude-command-process.ts crossed the 300-line cap, so terminateClaudeProcess
moved to claude-login-process-termination.ts. No max-lines suppression added.

* fix(crash-reporting): scope the self-kill guard to its real host topology

Round-2 review findings on the own-Chromium tree-kill guard.

BLOCKING 1 — "the own-Chromium refusal is a no-op in the process that issues
the pty-descendant-sweep taskkill". Correct on the mechanism, wrong on the
consequence; REBUTTED in part and documented in full.

Confirmed: the only non-test `setAppEnvironment` installs are
main-process-preflight.ts:177 (Electron) and orcad-entry.ts:84 (Node, whose
`getAppMetrics()` is `[]`); daemon-init-fresh-import.ts is a test harness. So
in the standalone daemon `readOrcaChromiumProcessPids()` is empty and
`admitSelfInitiatedTreeKill` always admits.

But that is not a live hazard. `killWithDescendantSweep` reaches
`terminateWindowsProcessTree` only when `verifyWindowsTreeKillTarget` returns
`own`, and that walks ancestry back to `deps.ownerPid ?? process.pid` — the
KILLING process's pid. In the daemon that is the daemon's pid. Orca's Chromium
processes are children of Electron main, a sibling of the daemon, so their
chain never reaches it: hop 0 lands on main, and within MAX_ANCESTOR_HOPS the
walk dead-ends and returns `foreign`. The reviewer's probe passes
`ownerPid: 1000` with the renderer as a direct child of 1000 — that is the
Electron-main topology, where the AppEnvironment IS installed and the guard DOES
fire, not the daemon's. On an orcad/SSH host there is no Chromium on the box at
all, so `[]` is accurate rather than degraded.

Locked in as tests rather than prose (own-chromium-tree-kill-guard.test.ts):
a renderer classifies `foreign` from a daemon ownerPid with an empty pid set,
and `own` from main's ownerPid with an empty set — the falsifiable pair showing
the pid set is load-bearing in main and nowhere else. Documented the host
coverage in orca-chromium-process-pids.ts and own-chromium-tree-kill-guard.ts.

One genuine hole the finding exposes: `signalProcessTree`'s `taskkillTree` is a
fourth pid-addressed taskkill family (non-blocking item 2), it runs in the
daemon/relay/CLI where the guard cannot run, and it guarded only on
`!child.pid`. Reusing the predicate the codex login teardown already uses, the
win32 branch now refuses a reaped child and falls back to `killRoot` — the same
shape as the existing `!child.pid` branch. That closes the reaped-then-recycled
pid path in every host.

BLOCKING 2 — module doc overstates coverage. Rewritten: the ring is per-process
and its only reader lives in Electron main, so a count on a `render-process-gone`
covers main-issued kills only. Sites are now split into main-only, main-and-
other-hosts (runProcess choke point, POSIX PTY group sweep, Windows Job Object —
which record into a ring nothing reads when they run in the daemon or relay),
and never-instrumented, with the note that a daemon/relay omission is a
diagnostics gap, not a missed suspect, per the topology argument above.

BLOCKING 3 — the three out-of-main instrumentation sites were untested. Added
regression coverage: the runProcess seam on both branches plus the reaped-child
refusal (process-tree-termination.test.ts), the group sweep recording only
groups it actually signalled and skipping an ESRCH group
(posix-pty-process-groups.test.ts), and the Job Object recording the shell pid
only on `terminated` (windows-pty-job.test.ts). Verified red: reverting the
three production files to origin/main fails 7 of the new tests.

BLOCKING 4 — the Windows evidence validates a single-process model. Accepted.
The main2.js arms exercise `pty-descendant-sweep` inside one Electron process;
that models the in-process/degraded daemon and the local PTY provider, not the
standalone daemon. Arm C's "the 449351d6 shape is not producible with the guard"
holds for main-issued kills only. In the daemon the shape is blocked one layer
earlier, by the ancestry check, which the arms do not exercise.

NON-BLOCKING taken: `recordSelfInitiatedTreeKill` moved outside the native
`terminateJob` try in windows-pty-job.ts, so a diagnostics throw can no longer
downgrade a real termination to `unavailable` and escalate callers to a broader
kill; covered by a test. The "all three families" parenthetical is gone with the
doc rewrite. `pnpm build:relay` run: exit 0, all seven targets built.

NON-BLOCKING declined: codex-accounts/service.ts records before the spawn
because a refusal must prevent the spawn — the crumb means "we were about to
kill this pid", which is the artifact worth having; the existing comment already
says so. `app.getAppMetrics()` perf is unbenchmarked and unchanged by this round.

Verification: pnpm tc clean; oxlint clean on touched paths;
check:code-quality:changed 0 new findings; oxfmt applied. 730 tests pass across
shared/child-process, main/crash-reporting, main/pty, main/windows and the guard
and descendant-sweep suites. The 4 failures in providers/git/codex-integration
reproduce on HEAD without these changes.

* fix(crash-reporting): keep the reaped-pid skip from flipping the termination barrier

The win32 hasExited short-circuit correctly avoids taskkill on a pid Windows
may have reissued, but it resolved `true` — verified tree termination. A
taskkill against a reaped pid already resolved `false`, and run-process turns
`true` into barrierTerminationVerified + terminationReporter.report(), which
releases the git admission grant on root exit instead of on `close`. That
admits the next git command while a descendant holding the inherited pipes is
still writing the repo. Resolve `false` so the skip changes only which process
we refuse to signal, not what the barrier claims.
2026-09-03 02:40:58 -07:00
djegor315-sketch f27a30d1b6 fix(ai-vault): avoid large Codex scan timeouts (#17889)
Agent Session History exceeded its 130-second deadline on large local Codex
histories. Three costs combined: excluded worker transcripts were recognized on
their first line but still drained to EOF (1,011 files / ~18.3 GiB on the
reported corpus), large ignored records were fully decoded and JSON.parsed, and
the persisted parse cache was discarded on every app update.

- Stop resumable reads the moment `session_meta` marks a worker transcript.
- Skip decode + `JSON.parse` for records the parser only feeds to the timeline.
  The skip set is the complement of what `consumeCodexRecordLine` reads, and
  applies only above the bounded prefix limit, so a long opening prompt (which
  is the session title) still takes the exact parser.
- Prove cross-volume rollout aliases from a bounded `session_meta` read routed
  through the WSL transcript FS gate, carrying the scan's AbortSignal, fanned
  out across contested candidates with bounded concurrency.
- Make parse-cache schema 2 the semantic compatibility boundary so an update no
  longer forces a multi-gigabyte cold scan, fenced by a build-time ratchet on
  the persisted session shape.
- Report early-stopped transcripts as their own `aiVault.scan` attribute.

Verified on macOS, Ubuntu over SSH, and Windows: read volume drops 336 -> 49.5
MiB identically on all three; the Windows failing-test set is byte-identical to
main. Reported corpus: 130s timeout -> 58.6s, 244 sessions, 0 issues.

Fixes #17888.
2026-09-01 04:21:43 -07:00
Brennan BensonandMerge Sim d4db524ba7 fix(native-chat): stop unjournaled provider frames from killing the session (#17813)
A frame the classifier declines (status-chrome, suppressed-benign, stream-into-item)
is deliberately not journaled. #17720 turned that null translation into
`{accepted: false, reason: 'untranslated'}`, which is not `backpressure`, so the
notification retry queue treated it as unreplayable and escalated through fail() ->
forceCloseUnexpected -> connection.close(). The app-server latched `closing` and the
create path's next model/list rejected with "codex app-server connection is closed
(model/list)".

The provider emits `remoteControl/status/changed` right after initialize, so every
structured Codex session died on its first chrome frame. Admit the null translation
instead, before any bookkeeping or publish.

Also restore the error-frame exemption from the generic row cap, dropped by the same
PR: the cap now runs after the classification check, so a noisy turn can no longer
reduce provider errors to a suppression count. The test that pinned the capped
behavior is inverted to assert the exemption.

Co-authored-by: Merge Sim <sim@local>
2026-08-31 23:31:27 -07:00
Neil 20a12a6a46 perf(codex): share one launch-prep hook install across a spawn burst (#17669)
* perf(codex): share one launch-prep hook install across a spawn burst

Codex launch prep runs a full managed-hook install on every local PTY
spawn, and both install lanes serialize globally per Codex home. Opening
a multi-pane worktree therefore paid N full installs back to back, and a
resumed Codex pane prepares twice. Concurrent spawns for the same runtime
home now share one run; the promise is dropped as soon as it settles, so
the next launch still re-reads hooks.json and the user's trust state.

Also split the `host_env` spawn-timing phase, which spanned the entire
Codex preamble and pinned that cost on the env builder that ran last.

* refactor(codex): unify the two hook-install single-flight lanes

Both the WSL and launch-prep lanes now share one generic in-flight helper
instead of duplicating the map bookkeeping. Also routes the WSL launch-prep
install through the serialized variant, which closes the same per-spawn
serialization gap on WSL that the native lane just got.

* refactor: extract the shared in-flight run dedupe

The codex hook service and the GitHub conflict-summary cache had grown
near-identical private copies of the same single-flight helper. Both now
use one module, which also keeps the hook service clear of the 300-line
budget. The shared copy keeps the identity check on clear so a late settle
cannot evict a newer entry for the same key.
2026-08-31 15:16:38 -07:00
Brennan BensonandMerge Sim 872bd51d47 fix(native-chat): reland large structured command results (#17720)
* fix(native-chat): preserve large structured command results (#17707)

* fix(native-chat): preserve large structured command results

* chore: place native chat validation artifacts under docs

* chore: drop stale root package config

* fix(native-chat): enforce rebuilt lifecycle append slots

---------

Co-authored-by: Merge Sim <sim@local>

* chore: omit native-chat reland planning docs

* fix(native-chat): remove journal store import cycle

* fix(native-chat): keep journal factory acyclic

---------

Co-authored-by: Merge Sim <sim@local>
2026-08-31 13:12:19 -07:00
Brennan Benson 894ed75abb Revert "fix(native-chat): preserve large structured command results (#17707)" (#17719)
This reverts commit 5fe37729ea.
2026-08-31 12:34:49 -07:00
Brennan BensonandMerge Sim 5fe37729ea fix(native-chat): preserve large structured command results (#17707)
* fix(native-chat): preserve large structured command results

* chore: place native chat validation artifacts under docs

* chore: drop stale root package config

* fix(native-chat): enforce rebuilt lifecycle append slots

---------

Co-authored-by: Merge Sim <sim@local>
2026-08-31 12:34:03 -07:00
Brennan BensonandMerge Sim 5ff1aa540e fix(codex): re-land WSL direct-home cutover with counsel findings fixed (#16854)
* fix(codex): safely re-land WSL direct homes

* fix(codex): finish WSL direct-home cutover

* fix(codex): coalesce WSL launch hook installs

* perf(codex): avoid duplicate retired WSL session scan

* fix(codex): retain canonical WSL retired-home path

* fix(codex): fail closed before retiring WSL auth

* fix(codex): reopen WSL drain after rollback

* fix(codex): preserve WSL source on unknown panes

* fix(codex): harden repeated WSL runtime drains

* perf(codex): bound pending WSL session scans

* fix(codex): recover invalid WSL session watermarks

* fix(codex): validate retained WSL scan state

* fix(codex): accept durable WSL scan state

* test(codex): cover the drain's inode-identity guard against destination replacement

Removing the four `target_auth -ef temporary_destination_auth` assertions left
all 33 apply-script tests passing, so a regression deleting them would have
shipped silently. Reproduced before writing this.

A hash check cannot catch the case. The pinned hard link keeps the original
inode, so it still hashes correctly after another writer atomically renames a
different file over the destination path; only inode identity sees it. Without
the guard the script exits 0 and retires the source, leaving the user holding
bytes nothing validated. The new case asserts the source survives.

The harness is split by responsibility so no file exceeds its max-lines budget:
fixtures, the coreutils interference shims, the run types, the apply runner, and
the recovery/absent runners. The atomic-rename hook is deliberately separate
from the in-place rewrite shim because different guards catch them.

* fix(codex): keep the split drain harness inside the child-process boundaries

Extracting the harness into non-test modules moved it out of the exemptions the
single test file had: three new files import child_process, and two spawned
without windowsHide.

Adds the three to the import allowlist, and sets windowsHide on the spawns
rather than exempting them - the flag is correct for these calls regardless of
the ratchet, and they are skipped on win32 anyway.

---------

Co-authored-by: Merge Sim <sim@local>
2026-08-31 11:20:22 -07:00
Brennan BensonandMerge Sim f2e9ba453c fix(agent-hooks): route reminted pane keys to canonical identity (STA-3993) (#15714)
* fix(agent-hooks): route reminted pane keys to canonical identity (STA-3993)

Spawn was stripping $$<base32>:L$$ ORCA_PANE_KEY values (and the launch
token) instead of rewriting them to the metadata-proven tab:leaf key, so
OMP hooks never entered last-status.json and sleeping rows stayed working.

Alias that exact remint form onto the canonical pane so later posts still
route, and keep unmatched tokens from stamping another pane.

* fix(agent-hooks): keep reminted pane-key aliases first-pane-wins

Remint tokens have no embedded tab identity, so a later spawn that reused
the same $$ token with a different tab/leaf was overwriting the alias and
routing leftover hook posts onto the new pane. Refuse destination changes
for that form while still allowing same-pane pty id updates.

* fix(agent-hooks): keep pane alias limit import valid after refactor

* fix(agent-hooks): bound pane alias destination keys

* fix(ssh): keep pane identity env stripped when hooks disabled

---------

Co-authored-by: Merge Sim <sim@local>
2026-08-30 17:54:13 -07:00
Brennan BensonandMerge Sim 585b4086d3 test(codex): pin Codex read-repair with a real-binary contract check (#17300)
* test(codex): pin Codex read-repair with a real-binary contract check

Orca's session index-heal depends on a Codex behavior: a `thread/read` of an
unindexed rollout performs a read-repair that inserts the `threads` row. All 55
existing heal tests drive a stub app-server and assert "healed" as "the call did
not error", so if Codex ever dropped the repair they would all stay green while
the subsystem went silently inert.

Adds a real-binary contract check built to the same shape as the Git binary
compatibility contract (src/shared/git-binary-compatibility.test.ts): env-gated
test file, version asserted against the binary, dedicated path-filtered PR job.

Pins only the four arms ablation established Orca relies on:
  - a read of an unindexed rollout inserts the state row
  - a session with no read inserts nothing (the negative control that makes the
    insert causal rather than incidental)
  - re-reading an indexed thread inserts nothing
  - an archived thread stays archived rather than being resurrected

Written against codex-cli 0.150.1. The job sets ORCA_CODEX_CONTRACT_REQUIRED=1
so a missing or failed CLI install fails red instead of silently skipping.

Existing heal tests are unchanged.

* test(codex): register the contract job in the verify aggregate contract

`pr-workflow-parallelism.test.mjs` pins `verify.needs` exactly, so adding the
job to pr.yml without updating that list failed the shard. Adds the entry, and
adds a workflow contract test mirroring `git-binary-compatibility-workflow.test.mjs`:

  - the pinned CODEX_CLI_VERSION is the single source for both the npm install
    and the runtime version assertion, so the two cannot drift apart
  - the install prefix and the binary path the test is pointed at are the same tree
  - ORCA_CODEX_CONTRACT_REQUIRED=1 is set, so a failed install fails red rather
    than turning the job into a green no-op

Removing the REQUIRED env from pr.yml reddens the new test, confirming it is live.

* test(codex): make binary version guard exact and bounded

* ci(codex): cover index-heal transport dependencies

* test(ci): pin Codex contract dependency coverage

* test(codex): align contract watchdog with child deadlines

* test(codex): cover three-session contract watchdog

* fix(codex): add sqlite sync-database to index-heal scope

---------

Co-authored-by: Merge Sim <sim@local>
2026-08-30 14:39:46 -07:00
Neil 1369821bad Split Codex hook service responsibilities (#17260)
* Split speech session lifecycle

* Split terminal output scheduler pipeline

* Split mobile browser pane modules

* Prune resolved max-lines suppressions

* Split pane tree equalization logic

* Extract mobile troubleshoot screen styles

* Split external automation manager

* Split main window service attachments

* Split hosted review creation checks

* Split automation dispatch event handling

* Split settings navigation metadata

* Split daemon initialization lifecycle

* Split GitLab item dialog

* Split relay dispatcher layers

* Split mobile host screen

* Retarget mobile view settings source test

* Split runtime file client layers

* Split ports panel layers

* Split runtime environments pane layers

* Split local PTY provider responsibilities

* Split CDP bridge responsibilities

* Split relay Git handler responsibilities

* Track moved relay Git fetch audit

* Split Linear item drawer responsibilities

* Split telemetry event schema responsibilities

* Split resource usage status responsibilities

* Split remote terminal multiplexer responsibilities

* Split Git worktree responsibilities

* Split Codex hook service responsibilities

* Keep mirrored hook trust type private

* Fix F3-speech for #17123

* Fix F1-cycle for #17131

* Fix F4-navtest for #17157

* Fix F2-allowlist for #17161
2026-08-29 20:21:14 -07:00
Neil 2dfaa676d8 chore: update oxlint and oxfmt (#17150) 2026-08-29 14:13:35 -07:00
Brennan Benson fd9125ea8c feat(native-chat): Codex structured native chat restructure (#16729)
* feat(native-chat): port structured Codex sessions from restructure-recovery

Rebuilds the desktop structured native-chat implementation from
brennanb2025/native-chat-restructure-recovery (tip 4e31c08db3) on top of
current main as a single commit, scoped to the local Codex path.

Ported:
- Structured agent-session core: durable record store + single-writer lease,
  canonical journal, agent-session wire host/attach/eviction/subscribers,
  `agentSession.*` RPC surface (registered via ALL_RPC_METHODS; host-side
  mobile allowlist included for wire compat), pty write gate, transcript
  additions, and the Codex app-server adapter/launch resolution.
- Renderer: NativeChatStructuredSession view/composer stack, structured
  launch path with the single-flight guard, local structured session tabs
  sync, activation gate + structured inventory (read-only
  `agentSession.handoffStatus` probe), agent-session tabs in the tab strip,
  AI-vault structured session activation, and the settings pane with the
  parent Experimental Chat UI toggle plus the nested "Use updated structured
  native chat" toggle. New sessions require both flags, agent codex, no
  prompt, and a local non-WSL, non-Windows-host execution host
  (structured-native-chat-availability).
- Fixes 72c013cea6 (verified Codex launch recovery), 8ddbaf5e3d (defer
  native terminal view switching affordances), and 4e31c08db3 (release the
  launch gate after a visibility retry) with their regression tests,
  including the third-launch-after-retry guard case.
- Cross-version agent-session wire test + CI lane, packaging entries
  (proper-lockfile, agent-tooling asar excludes), and the wire-compat doc
  section.

Deliberately not ported: mobile/ changes, the Claude structured runtime
(only the claude-transcript-branch-proof and claude-structured-owner-identity
leaf modules remain, backing the kept TUI-recovery arms), the terminal↔chat
adoption/handoff flow (`agentSession.adoptTerminal`/`requestHandoff`, the
handoff request engine, TUI adoption machinery, orca-runtime adoption
methods), renderer switching affordances and their dead leftovers, the
hook/subagent-status refactor cluster, and unrelated branch changes. The
crash-during-acquisition recovery path (restart handoff adjudication,
restore/reverse re-acquire, lease schema handoff keys) is kept because every
plain direct launch depends on it; a trimmed handoff coordinator exposes
only status/restore/close.

Branch edits that targeted files main has since split (ipc/pty.ts,
worktrees.ts, rpc/methods/terminal.ts, useIpcEvents, pty-connection,
store/slices/terminals.ts, runtime-types, web preload) were re-applied to
the split modules, preserving main's newer logic (Windows CIM fallback,
browser tab close rework, cold-restore resume flow, dispatcher threading).

Known seam: the mobile clipboard image-provenance CONSUMER gate ships
(agentSession.send refuses unproven mobile image refs with
agent_session_image_untrusted) but the producer hunk in
rpc/methods/clipboard.ts stays with the unported mobile cluster, so mobile
image sends into structured chat fail closed until that side ports.

* fix(native-chat): trust only authenticated local image uploads

* fix(build): preserve Windows process-tree patch application

* test(windows): include process creation time in addon fixture

* fix(build): run windows-process-tree node-gyp from the physical package dir

gyp expands the node-addon-api dependency by probing node, whose cwd
resolves to the package's physical directory in the store, so the emitted
target is a store-relative ../../../../node-addon-api@... hop. gyp then
resolves that hop against the rebuild cwd; from the node_modules
symlink/junction it escapes the store and configure fails with
"node_addon_api.gyp not found" (run 32999886072).

Rebuild from realpath(package dir) so both bases agree, matching how the
package manager itself runs native install scripts. The regression test
replays gyp's expansion+resolution against the planned cwd and fails
without the fix.

* fix(native-chat): keep chat tabs visible through terminal closes and empty-worktree launches

Two proven blockers in the native Codex tab contract:

closeTerminalTab pre-empted the canonical unified close. With one terminal
left it deactivated the worktree on a terminal/editor/browser-only check,
blanking a workspace that still held a renderable agent-session tab; with
two or more it pre-picked a successor from terminal entities only,
re-stamping the group active before closeUnifiedTab's MRU/neighbor repair
could land on the chat tab. Successor choice now defers to the unified
contract whenever the terminal has a unified row, and deactivation is
gated on the unified renderable count (matching leaveWorktreeIfEmpty),
with the legacy pre-pick kept only for terminals without a unified row.

A structured session created on an empty worktree was published into the
host's headless group while preserveLocalLayout froze the local layout,
leaving the tab in store but permanently off screen. A preserveLocalLayout
owner now always takes client-owned placement — repairing a rendered
leaf whose group record is missing, or materializing a rendered group on a
truly empty worktree — and applies the client-derived layout repair while
still rejecting host-authored layout.

Regression tests drive the real store through closeTerminalTab (git
worktree and folder workspace) and the real snapshot applier for the
empty-worktree adoption states; all fail without the fixes.

* fix(native-chat): close stale turns and retry rejected sends

* fix(native-chat): retire hosted rows on structured tab activation

* fix(native-chat): preserve rpc defaults across main merge

* chore: format remote wire compatibility guide

* test(native-chat): cover retry after unconfirmed send

* fix(native-chat): reload outbox on session switch

* docs(settings): disclose structured chat platform limits

* fix(native-chat): await Codex launch-home preparation

* fix(codex): align child-process allowlist with async trust bridge

* test(identity): update inventory for tab surface refactor

* fix(windows): preserve process-tree CRLF patch sources

* fix(native-chat): anchor an unmatched chat echo where it was sent (#16117)

* fix(native-chat): anchor an unmatched chat echo where it was sent

The reported symptom was old user messages replaying below every new turn, so the
conversation read as scrambled. The cause was not that the echo failed to match a
transcript row. Claude consumes a mid-turn send through a `queued_command`
attachment and writes no `type:"user"` record for it, so some echoes can never
match, and no amount of matching will change that. The cause was WHERE an
unmatched echo rendered: buildMobileNativeChatTransientData appended every pending
item after the entire transcript, so it re-read below each turn that landed
afterwards.

Render each echo directly after the transcript row it was sent against, using the
baseline the send already captures. An unmatched echo is then at worst a duplicate
in the right position rather than a scrambled one, and it stays visible. Echoes
sharing an anchor keep send order; a send with no baseline, or one whose anchor
folding dropped, still falls back to the tail.

Deliberately NOT fixed by deleting the echo. Inferring from send ordering that an
echo can never match, then removing it, loses the user's own text for a message
the agent did receive, and it cannot fire in the common case anyway - measured
drain groups are 1,017 of size 1 against 55 larger. It also escalates an existing
gap: the count pass has no baseline-tail guard, unlike the glue pass, while
`messages` is a 40-row window that head-trims, resets on reconnect and grows at
the front on loadEarlier, so a false landing there would license deleting a
DIFFERENT outstanding message.

That count-pass gap is real and left for a separate change; anchoring makes its
worst case a duplicate in place rather than a scrambled conversation.

* fix(native-chat): preserve folded echo anchors

* fix(native-chat): preserve forward-folded echo anchors

* fix(native-chat): keep leading folded echoes in place

* fix(workspace-cleanup): show git status for every row (#16690)

* fix(native-chat): refuse structured chat on every Windows execution path

canUseStructuredNativeChat only refused win32 when a project runtime
resolved, so folder-workspace keys (and other keys with no project
runtime) failed open into structured chat on Windows. Fail closed on
win32 unconditionally after the host check, matching the settings copy:
local macOS/Linux only; Windows/WSL/SSH stay on terminal chat.

* fix(native-chat): restore runtime refusals behind the win32 gate

506d375de3 replaced the project-runtime checks with a bare platform test,
so a WSL or repair-required runtime resolution would no longer refuse
structured chat off-win32. Keep the unconditional win32 refusal and
re-run the runtime resolution after it, so the gate does not depend on
the resolver's own platform guard. Tests inject WSL and repair-required
resolutions on darwin/linux and fail against the regressed gate.

* fix structured session journal durability

* fix structured tab active pointer after restart

* fix(native-chat): await optional lease renewal callbacks

* refactor(skills): extract install error messages

* fix(agent-session): harden recovery ownership

* fix(native-chat): retain panes across tab activation

* fix(native-chat): address round-one review findings

* test(native-chat): align integration coverage after main merge

* fix(native-chat): harden round-two reliability

* fix(native-chat): harden round-three reliability

* fix(native-chat): close round-four recovery gaps

* fix(native-chat): separate bounded journal key forms

* fix(native-chat): reset outbox error in render on session switch

The switch effect adjusted error state after the sessionId prop changed,
tripping react-doctor's no-adjust-state-on-prop-change on the changed-code
gate and flashing the old session's banner for a frame. Reset it with the
render-time previous-value guard instead.

* fix(native-chat): invalidate stale outbox settlements

* test(native-chat): restore settled-error session-switch regression

a6e2379bd1 replaced this test with the in-flight settlement race test,
leaving the render-time error reset unpinned: deleting the reset block
still passed the whole native-chat suite. Keep both scenarios pinned;
they are distinct (settled error clears on switch vs stale settlement
invalidated in the commit-to-passive window).

* test(wire): make release checkouts race safe

* test(wire): pin cross-process checkout single-flight and importer specifier contract

* test(wire): harden release checkout lifecycle

* fix(build): drop CR-byte residue from windows-process-tree patch

The two trailing CR bytes on the patch's deletion lines are a proven
no-op: pnpm hashes patches CRLF-normalized (both forms hash to the
lockfile's 946ffb2b) and materializes this package without applying the
patch in either form, so the load-bearing build edits come solely from
applyWindowsProcessTreeBuildFixes() (#16947), which handles both source
EOL forms. Restore byte-identity with main and repin the contract test
to the post-#16947 reality: LF-only patch bytes plus lockfile hash sync.

* fix(native-chat): skip empty startup recovery
2026-08-28 16:45:58 -07:00
Brennan Benson 3558cf943f fix(codex): heal WSL hooks before typed launches (#16535)
* fix(codex): heal WSL hooks before typed launches

* test(codex): keep launcher fixture type-safe on Windows

* fix(build): list codex-home-wsl-env in the CLI typecheck project

`managed-home-shell-preflight.ts` is already in the CLI project's include list and now imports
`wslCodexRuntimeHomeForGuestHome` from `src/main/pty/codex-home-wsl-env.ts`, which the list did not
cover — TS6307, so the CLI typecheck failed on every push.

Added the single module rather than a `src/main/pty/**` glob: it is a 31-line leaf with no imports
of its own, so it does not widen what the CLI bundle can reach.

* fix(codex): converge the two WSL hook install lanes onto one writer

Two independent readiness reviews agreed the Orca-terminal boundary holds, but Codex Sol found a
P1 the other rated P2: the new just-in-time repair raced the existing relay installer and the two
produced DIFFERENT hook and trust representations for the same managed home. Two unserialized
writers emitting different formats is worse than the bug this PR fixes, because it fails
intermittently rather than cleanly — a pane works or does not depending on which lane won.

- Relay Codex installs now delegate to the runtime-home writer, so there is one canonical
  representation instead of two. Redirected scripts use the runtime path, the readable wrapper,
  and the prepended group.
- `installForRuntimeHomeSerialized` puts every asynchronous WSL caller for a given home on one
  queue (`wslInstallQueues`), so concurrent panes cannot interleave writes.

Also rewrites the stale pin test the new `-x` guard broke. It asserted the defect —
"would run the impostor if the preflight carried an unqualified command name", expecting the
hijack marker to exist. The guard is a security improvement, so the test now asserts the contract:
an unqualified preflight is skipped and the marker is never written. Rewritten to the new
behavior, not loosened or deleted.

818 tests pass across the affected suites; typecheck clean. The changed-file quality gate could
not run locally — its pnpm engine-warning JSON parser fails under Node 26 — so CI covers it.

The boundary both reviews verified is untouched: paired/relay/mobile clients stay hard-blocked
from the RPC, params remain shape-locked to the managed home suffix with traversal rejection,
nothing is written outside the managed home, and macOS/Linux stay inert.

* fix(codex): serialize resolved WSL hook homes

* fix(codex): recover managed WSL homes after restart

* fix(wsl): translate Codex preflight through WSLENV

* fix(cli): cover bounded WSL Codex repair

* fix(codex): coalesce duplicate WSL hook repairs

* fix(codex): verify reconstructed WSL homes
2026-08-27 13:05:12 -07:00
Mark XianandBrennan Benson cc384c5a3d fix(agent-hooks): post posix payloads as json (#11292)
* fix(agent-hooks): post posix payloads as json

* fix(agent-hooks): mark header merged envelopes

* docs(agent-hooks): describe header merge envelope

* fix(agent-hooks): encode posix metadata headers

* test(agent-hooks): update WSL JSON hook assertions

* fix(agent-hooks): negotiate raw JSON transport

* fix(agent-hooks): preserve packed metadata in POSIX shells

* test(agent-hooks): include hook envelope in relay boundary inventory

---------

Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
2026-08-27 12:32:29 -07:00
Brennan Benson 928d306b53 fix(agent-hooks): stop the Windows hook launcher spelling the AV-denied flag pair (STA-5237) (#16739)
* fix(agent-hooks): stop the Windows hook launcher spelling the AV-denied flag pair (STA-5237)

`-WindowStyle Hidden` + `-EncodedCommand` is denied at CreateProcess by
Kaspersky on Windows 11, whatever the payload decodes to. Bash reports it as
`Permission denied` and every managed hook event fails, so agent status never
arrives; the parent shell also briefly cannot spawn anything afterwards, so a
denied hook can take the user's next command down with it.

Measured on the reporting host (#16003), with a harmless `exit 0` payload:

  -NoProfile -ExecutionPolicy Bypass -WindowStyle Hidden -EncodedCommand  126
  -NoProfile -WindowStyle Hidden -EncodedCommand                          126
  -WindowStyle Hidden -EncodedCommand                                     126
  -NoProfile -EncodedCommand                                              0 (5/5)

#16576 removed `-ExecutionPolicy Bypass`, which is the one flag of the three
NOT in the signature, so hooks kept failing after that fix. The pair that has
to stop being spelled is `-WindowStyle Hidden` + `-EncodedCommand`.

Because the change is to the shared switch constant, it covers every site that
spells the denied pair in one edit: Claude via `wrapWindowsPowerShellEncodedCommand`,
gemini/cursor/droid/command-code/copilot via `wrapWindowsHookCommand`, the
`runtime-home-hook-command` unsafe-HOME fallback, and the spaced-path fallback
for codex/grok/devin/antigravity. Only a flag is removed, so parser and payload
compatibility is unchanged for every executor: the string is still a PowerShell
command line, still one self-contained token, still base64-shielded.

The tradeoff, recorded rather than hidden: `-WindowStyle Hidden` was the shipped
fix for #14815 (+#14828, #15117, #15447, #15767), and this removes it. Its
suppression was never measured — #14825 confirmed it visually, #16576's author
stated it "remains unverified on a real box", and #15506's author argued it
cannot help a `.cmd` child with no console to inherit. The console is allocated
by the parent chain, not by this command line. A live window measurement is
still outstanding and is called out in the PR.

Also adds `windows-hook-payload-delivery.test.ts` to the PR CI Windows leg,
which had never run it.

* test(agent-hooks): keep launcher token out of source grep
2026-08-27 10:48:53 -07:00
Neil 9062494f9b fix(ai-vault): stop a whole opencode.db failure reading as one skipped transcript (#16587)
* fix(ai-vault): stop a whole opencode.db failure reading as one skipped transcript

#15036 reported "1 transcript skipped / database is locked" with both Agent
Session History scopes empty. Two separate defects.

The panel counts every unkinded scan issue as a skipped transcript, so a
failure that lost an entire *source* was reported as one lost *file*. The
whole-database failure is now kinded `scope`, and an unknown `kind` from a
newer host degrades to `scope` instead of failing validation and coming back
unkinded — a mixed-version remote host previously turned a source-level
failure into a phantom skipped transcript.

The read also inherited sqlite3's 0 ms busy timeout, so a genuinely contended
open failed in ~1 ms. It now opens once with a bounded timeout. No retry loop:
sqlite's own busy handler already blocks and retries internally for the whole
timeout, and WAL readers do not block on a writer at all (measured: 547/547
cross-process reads at timeout=0 while a writer held open transactions).

Measured against a real Ubuntu-24.04 distro, Windows cannot take SQLite's file
locks over \\wsl.localhost at all: an idle, never-WAL, nothing-attached
database still answers SQLITE_BUSY, a 5 s busy timeout does not change it, and
the identical bytes open fine once copied to local disk. So a lock-family error
on that share never means "a writer holds it" and no timeout can help. The copy
says so rather than sending the user after a write-ahead log that is not the
problem. Restoring those sessions needs an in-distro read; that is a follow-up,
and this PR no longer pretends a timeout will do it.

immutable=1 is deliberately not used as a workaround: over the same share it
opens and returns 100 of 150 rows, silently dropping everything still in the
uncheckpointed -wal — in a history panel, exactly the newest sessions.

* skip the provably futile busy wait on \\wsl.localhost paths
2026-08-26 23:37:16 -07:00
Brennan Benson aaef5e8c9f fix(agent-hooks): deliver hook events that fire while Orca is restarting (STA-5329) (#16685)
* fix(agent-hooks): correct durable spool delivery

* fix(agent-hooks): spool curl failures after retries

* fix(agent-hooks): keep replay out of runtime observations

* test(agent-hooks): pin managed hooks inert outside an Orca terminal

* fix(agent-hooks): address review findings on the durable spool

- claude: pass the literal source; options.agent does not exist (typecheck)
- kimi: the windows-local ordering runs its guard pre-stdin and before the
  function exists, so it no longer spools there (printed command-not-found)
- writer: require a readable endpoint file before creating a spool tree
- antigravity: carry its out-of-band event name into the record and filter on it
- drain: truncate only the bytes consumed, preserving concurrent appends and a
  torn trailing line

* fix(agent-hooks): ignore spool events without pane attribution

* fix(agent-hooks): make spool replay and appends robust

* test(agent-hooks): type spool replay records

* fix(agent-hooks): defer unterminated spool records

* fix(agent-hooks): replay spool events through relays

* fix(agent-hooks): preserve Codex prompt across child replay

* fix(relay): keep startup alive when spool replay fails

* fix(relay): simplify spool replay startup guard
2026-08-26 23:17:34 -07:00
Jinjing 9c01e09ecc Revert "fix(codex): launch WSL accounts from direct homes" and "refactor(codex): remove WSL runtime mirror machinery" (#16722)
This reverts commit ebcd637db9 (#16504) and dependent commit 673842db35 (#16505).

Launch-blocker rationale (findings-counsel validated):
- P0 Data Loss (F06-1): The WSL legacy auth drain deleteSource=1 path deletes intact source auth without re-validating the destination after concurrent destination rewrites, permanently corrupting auth credentials on upgrade.
- P1 Workflow Regression (FC-01): Pre-upgrade sessions under ~/.local/share/orca/codex-runtime-home/home are not linked into direct homes, breaking /resume in the Codex CLI for upgrading WSL users.
2026-08-26 22:42:37 -07:00
Brennan Benson ebcd637db9 fix(codex): launch WSL accounts from direct homes (#16504)
* fix(codex): launch WSL accounts from direct homes

* fix(codex): coalesce WSL auth drains and validate distro homes

* fix(codex): preserve legacy WSL account home metadata

* fix(codex): retain marked WSL home compatibility

* fix(codex): verify the bytes the WSL drain promotes, not an earlier read

The apply script validated the source hash and then re-read it with cp, so a
legacy pane rotating in that window put bytes freshness never judged over a
valid account home. Codex rewrites auth.json in place, so that read can be torn.

Covers it by running the real guest script under sh with a sha256sum shim that
rotates the source between the two reads; without the guard it exits 0.

* fix(codex): harden WSL auth drain races
2026-08-26 21:22:05 -07:00
Brennan Benson 07f2e14c08 refactor(codex): make WSL account surfaces direct-home aware (#16499)
* refactor(codex): make WSL account surfaces direct-home aware

* fix(wsl): keep Codex relay hooks on managed runtime home

* test(wsl): assert relay hooks use managed Codex home
2026-08-26 17:00:21 -07:00
Neil 26721bd632 fix(codex): stop blocking the main thread on trust grants (#16441) (#16594)
* fix(codex): stop blocking the main thread on trust grants (#16441)

Codex hook trust was granted by blocking the Electron main thread on
`spawnSync` of a bundled ELECTRON_RUN_AS_NODE entry for the whole
app-server deadline: 15s native, 35s WSL, ~45s on the real-home path
(rebase inspect + repair + grant). Cold start and every Codex pane
launch showed "Not Responding"; the reported event-loop gap was
15,049 ms.

The subprocess only ever existed to donate an event loop to a
deliberately blocked parent — `runCodexHookTrustGrantSession` was
already the real async implementation. Make the callers async and the
fork is unnecessary, so the bridge, the forked entry and its envelope
are deleted along with their build/knip/tsconfig registrations. The CLI
`agent hooks prepare-codex` handler is already async, so it awaits the
in-process session and saves a process spawn per managed-home shell.

`resolveCodexTrustGrantHost` is async too; the WSL identity probe moves
from `execFileSync` to `runProcess`, dropping that file from the
child-process import allowlist. Status reads keep a synchronous
native-only stamp path.

Two invariants that held only because the lane blocked:

- Overlapping capability probes were impossible by construction.
  `GitCapabilityCache`'s dedupe engine is extracted to a shared
  `CapabilityProbeCache` and `CodexAppServerCapabilityCache` now
  inherits it, so concurrent launches against a cold host share one
  app-server session instead of one each.
- Two grants on one `config.toml` could not interleave capture and
  restore. A reentrant per-file lane now serializes the whole install
  sequence (managed, WSL runtime, real-home ensure, legacy sweep) and
  the grant and rebase inside it.

Cold-start work moves off the critical path: retained-home
reconciliation (N sequential sessions) is fire-and-forget behind the
daemon provider, and the startup real-home ensure chains into managed
hook reconciliation instead of blocking app init.

Every preserved semantic is unchanged: never throws, the
ORCA_DISABLE_CODEX_TRUST_RPC kill switch, ledger hits, backfill-pending
and cooldown fallbacks, config rollback on every failure path,
pre-grant self-computed trust removal, the verify-failure taxonomy,
diagnostics and telemetry.

* fix(codex): widen the trust-config lane to every config.toml writer

Review follow-ups on #16441's async trust grant:

- `markCodexProjectTrusted` now runs inside the runtime+system config.toml
  lanes, so a project-trust write can no longer land inside a hook grant's
  capture->restore window and be silently reverted. Its callers await it.
- `install`/`refreshRuntimeUserHooks`/`remove` hold the system config.toml
  lane as well as the runtime one — they promote approvals into
  ~/.codex/config.toml and mirror it back. Lock order is runtime-before-system
  everywhere.
- The real-home ensure chain resumes after a rejection instead of returning
  the same rejected promise to every later pane launch, and resolving the real
  home is now inside the module's never-throws boundary.
- `buildSpawnEnv` awaits inside a cancelable pending-spawn registration, so
  shutdown during the (now long) env build stops the PTY from launching.
  `prepareLocalPtySpawn` generalizes into `awaitCancelableLocalPtySpawn`.
- CapabilityProbeCache drops the test-only `nowMs` passthrough; its probe
  backstop comment now describes what it actually guards.
- Preflight is a plain async function; the trust dispatch in orca-runtime
  collapses into one `markWorkspaceTrustedForAgent`.

* test(codex): exercise the trust-config lane under real concurrency

The async grant makes two pane launches overlap for the first time. These
drive the real modules end to end on real files: a rollback swallowing a
sibling's grant, a markCodexProjectTrusted write landing inside a capture
-> restore window, shared capability-probe dedupe on a cold host, the
host-scoped transient cooldown, and reentrancy from inside an installer.

Each was verified to fail against a deliberately broken implementation
(lane removed, dedupe disabled, cooldown made global, reentrancy pass-
through disabled).

* test(codex): stop hook-service suites spawning the developer's real codex

The forked grant bundle never existed under vitest, so the RPC lane was
unreachable in tests on main. Running it in-process makes these suites
spawn a real `codex app-server` when one is installed: 38 spawns and two
failures in hook-service-runtime-trust-repair on a machine with codex,
green in CI where there is none. Stand in for the missing binary so both
environments exercise the same fallback lane.

* docs(codex): scope the trust-RPC kill switch comment to what it actually gates

The comment read as though the flag forces the fallback lane everywhere. It
gates the managed grant only: the real-home rebase still runs its own
inspect/repair app-server sessions when Orca's insertion shifts a user's hook
positions, and never reads the flag.

Verified by exercise, not by reading — with the flag set, both
inspect-user-hook-trust and repair-user-hook-trust still ran. Pre-existing:
main has no check there either, it just blocked the main thread while doing it.

Widening the flag to cover the rebase is a follow-up; this only stops the
comment promising something the constant does not do.
2026-08-26 16:44:55 -07:00
Neil 015f904fca fix(codex): stop re-scanning all Codex session history on every launch (#16251) (#16593)
* fix(codex): stop re-scanning all Codex session history on every launch (#16251)

A launch deleted the backfill completion marker, and a marker could never
be written while a Codex pane was open, so every launch re-derived
"needs full scan" and walked the entire .codex/sessions tree — on Windows
with a large history that read as a hung window.

- v4 marker keeps a durable full-history baseline plus a bounded set of
  pending dates. v3 is read as a baseline, so upgrades pay no full scan.
- A launch now marks dates pending instead of deleting the marker, and a
  full pass certifies the baseline even while a pane is still running;
  the live pane's own date just stays pending.
- Pending dates are persisted, so an abnormal exit or a cross-midnight
  pane recovers a bounded window instead of a full walk.
- A date-limited pass can only extend an existing baseline, never create
  one, so it can no longer certify history it never looked at.
- Marker and index-heal target roots compare through
  normalizeRuntimePathForComparison, so Windows spellings of one
  directory stop invalidating each other.
- Both append-only ledgers stream instead of readFileSync + whole-file
  JSON.parse, keeping the main thread responsive on large histories.

* fix(codex): keep the backfill marker's full-scan demand durable

Review follow-ups on the v4 backfill marker:

- markCodexSessionBackfillMarkerPending no longer erases a persisted
  needsFullScan; the demand survives until a generation-current full walk
  retires it, and the function now reports it so the launch path folds it
  into its own in-memory flag (as @rumoii's #16252 does).
- A full pass settles the whole pending set instead of subtracting the
  empty set, so a date a full walk provably covered stops forcing an extra
  bounded pass on every startup.
- isCodexSessionBackfillDate does a real calendar check, so a corrupted
  marker cannot carry 2026/99/99. No age or future bound: the same guard
  gates rollout publication and a clock-skewed directory holds real
  sessions.
- 'scans only the current date once a baseline exists' now has a second
  date directory, so it fails on a full walk instead of passing either way.
2026-08-26 15:42:54 -07:00
Brennan Benson e8005c3325 fix(codex): preserve WSL account home trust (#16496)
* fix(codex): preserve WSL account home trust

* fix(codex): preserve WSL drive path semantics

* fix(codex): preserve mounted-drive WSL config paths

* test(codex): preserve WSL path helpers in mock
2026-08-26 14:33:34 -07:00
Neil c1a7748267 fix(agent-hooks): stop the Windows hook launcher spelling the AV-denied flag triple (#16576)
* fix(agent-hooks): stop the hook launcher spelling the AV-denied flag triple

Orca's Windows agent-hook launcher ran

  powershell.exe -NoProfile -ExecutionPolicy Bypass -WindowStyle Hidden \
    -EncodedCommand <base64>

That exact combination is the textbook "hidden encoded PowerShell" malware
shape, and endpoint security denies it at process creation whatever the
payload decodes to -- even `exit 0`. Every injected hook then failed with
`powershell.exe: Permission denied` (exit 126 from bash's execve, EACCES)
on every turn, for both Claude Code and Codex, with no AV exclusion that
re-enabled it.

Dropping any one of the three flags clears the signature. `-ExecutionPolicy
Bypass` is the one that can move: it sets the Process scope, and so does
`Set-ExecutionPolicy -Scope Process`, which now rides inside the encoded
payload. `-EncodedCommand` is never policy-gated, so the bypass always gets
to run before the managed script does -- which is what keeps Copilot's .ps1
hook working under a Restricted or AllSigned machine policy.

The hidden window and the encoding are unchanged, so nothing regresses for
#14815, #14818 or #6078.

Closes #16003

* fix(agent-hooks): ship the launcher shape #16003 actually measured as allowed

The previous revision of this branch dropped only `-ExecutionPolicy Bypass`
and kept `-WindowStyle Hidden -EncodedCommand`, on the reasoning that
"dropping any one of the three flags clears the signature". That sentence is
not in the bisect. The reporter ran exactly four command lines on the affected
Kaspersky/Windows 11 host:

  -NoProfile -WindowStyle Hidden -Command 'exit 0'                       -> 0
  -NoProfile -EncodedCommand <b64>                                       -> 0
  -NoProfile -ExecutionPolicy Bypass -Command 'exit 0'                   -> 0
  -NoProfile -ExecutionPolicy Bypass -WindowStyle Hidden -EncodedCommand -> 126

Every passing row drops two flags. No row drops exactly one, so the shape the
branch was about to ship had never been executed on the machine that reports
the bug -- and it is `-WindowStyle Hidden -EncodedCommand`, which is the
"hidden encoded PowerShell" pair the denial is named for in our own comment.
Shipping it would have closed #16003 while leaving every hook on that host
dying at CreateProcess, with no tracking left open.

So emit the measured-passing encoded row instead: `-NoProfile -EncodedCommand`.
Of the two flags there was a choice between, `-EncodedCommand` is the one that
carries correctness -- it is what keeps paths and switches intact across
cmd.exe and MSYS (#6078, #14815). `-WindowStyle Hidden` costs at most a console
flash, and only where the parent has no console to inherit.

Second, the relocated bypass now runs inside try/catch. Under a MachinePolicy
or UserPolicy GPO scope, `Set-ExecutionPolicy -Scope Process` reports that the
process scope did not take. `-ErrorAction SilentlyContinue` covers only the
non-terminating half of that; the command-line switch it replaces was silent
either way. This file already documents that non-stdout PowerShell streams
corrupt consumers merging our output into JSON stdout, so a per-invocation
ErrorRecord on stderr is a regression we should not trade for the switch.

Refs #16003

* fix(agent-hooks): keep the hook console hidden while dropping the AV-denied flag

Round 2 of this PR widened the fix from "stop spelling -ExecutionPolicy Bypass"
to "stop spelling it and -WindowStyle Hidden", on the reasoning that the #16003
reporter never measured a shape that drops exactly one flag, so keeping the
hidden+encoded pair would be extrapolation.

That trades a reproduced regression for an unmeasured one. Window suppression is
the shipped fix for #14815 and its four duplicates (#14828, #15117, #15447,
#15767): a hook launched from a parent with no console gets a fresh console per
event, which takes foreground and eats whatever the user is typing into Orca,
and never closes at all on the stdin-blocking path hook-stdin-contract.ts exists
to guard. That fires on every prompt, tool call and stop of every managed agent.
The AV denial, by contrast, is measured only for the full triple; that the
remaining pair still trips it is a hypothesis. Between a certain regression and
a possible one, keep the certainty.

So the flag that leaves the command line is the policy bypass alone — the only
one of the three with an exact in-payload equivalent, hence the only one that
can move without losing behaviour. If the pair turns out to be denied too, the
answer is a different shape that still hides the window.

* fix(agent-hooks): silence progress before the policy bypass can autoload (#16621)

Hardware-measured on Windows 11 while exercising #16576.

Set-ExecutionPolicy autoloads Microsoft.PowerShell.Security, and that module's
"Preparing modules for first use." progress record is written before any later
assignment can suppress it. Running the bypass first therefore defeated the
silencer that runs immediately after it:

  bypass-first    stderr = 616 bytes, first merged line '#< CLIXML'
  silencer-first  stderr = 0 bytes,   first merged line '{"decision":"approve"}'

That is precisely the corruption HOOK_PROGRESS_SILENCER's own comment warns
about -- redirected progress becoming CLIXML that can corrupt merged JSON -- so
the PR reintroduced the hazard it documents, one line below documenting it.

Both existing tests asserted the broken order, so they enforced the bug rather
than catching it. Reordered them and added one that pins the ordering itself
rather than the literal string, since the string will drift again.
2026-08-26 14:08:58 -07:00
Neil 48e63c015f refactor agent config and auth services (#16195)
* refactor: split agent config and auth services

* chore: repoint wsl and global-fetch guards at split module paths

* fix: restore merge-base Claude CLI error propagation

Drop the secret-redaction rewriting added to Claude CLI error paths in the
refactor: spawn errors again reject with the original Error (preserving
.code/.errno/.syscall/.stack) and command output/auth-status logs are no
longer rewritten.
2026-08-24 23:15:01 -07:00
Neil 2b1b094aa8 fix(cli): pair every resolved CLI with its runtime, and ratchet it (#16383)
Follow-up to #16365, which paired 8 spawn sites by hand. Hand-pairing is how
the class got introduced, so close it structurally instead.

cliPath is now required on CodexAppServerInvocation, `null` only for the
guest-side wsl.exe launcher where a host path pairs nothing. Optional let a
native builder omit it and silently fall back to pairing against a cmd.exe
wrapper with no type error. Every production site already passed it; only
test fixtures needed updating, which is the type doing its job.

Four more sites now pair. codex-state-db-backfill-recovery spawns the same
`codex app-server` subcommand #16365 fixed elsewhere. cli/handlers/account
was the worst case: addAgentNodePaths prepends the *newest* version-manager
bin, which is not necessarily where the CLI being launched lives, so it
actively created the mismatch — pairing now runs last so the CLI's own node
wins. commit-message-text-generation and skills/skill-update-run spawn
resolved binaries with inherited env.

cli/handlers/skills had grown its own buildNpxPath: a weaker local copy that
prepended unconditionally, ignored the Windows `Path` key, and special-cased
a '.' dirname. Deleted in favor of the shared helper, which checks the
sibling node actually exists — the behavior change one test had pinned.

The ratchet is the point: any file that resolves a CLI and spawns must
reference withCliRuntimeOnPath, with a shrink-only allowlist. It caught
skill-update-run, which I had missed. Its first draft required a call paren
and so let dependency-injected resolvers (`resolveCommand: resolveCodexCommand`)
through — verified by removing a pairing and watching it stay green, then
widened until it failed. A second assertion fails on a stale allowlist entry
so an exemption cannot outlive its reason.

external-editor-launch stays allowlisted: it launches a GUI editor, not a
Node CLI whose ABI matters.
2026-08-24 23:12:37 -07:00