Commit Graph
2288 Commits
Author SHA1 Message Date
Brennan BensonandClaude 1069bb053f fix(agent-launch): a cwd at the workspace root no longer forces a terminal (#22729)
"Continue in New Session…" always names a cwd, and both the renderer route
input and the host launch-mode decision read any cwd as a custom start
directory, so the continuation opened a terminal agent even when chat was
the user's default. Both now share one rule: only a cwd outside the
workspace root (after normalising slashes, Windows case, WSL aliases and
the distro's Linux spelling) requires a terminal. A subdirectory still
does, because a structured session cannot start there.

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-24 16:57:38 -07:00
Brennan Benson 5e6fcca0b3 fix(native-chat): date a session by its own lifecycle, not its subagents' work (#22520)
* fix(native-chat): date a session by its own agent's rows, not its subagents'

The journal reducer's lastActivityAt is the structured status summary's
updatedAt, which the status row uses as its completion stamp and
acknowledgement clock. It took the max over every journal row, and a
session's subagents write into the same journal after its own agent has
settled, so an idle parent was re-dated and marked unread on child work.

A row now dates the session only when the session's own agent produced it:
not a row whose producer linkage names a subagent, and not a subagent
roster row (a subagent-group block), which the session writes but revises
on every child transition. The roster rule is derived from the row body;
no new persisted field. Replay folds through the same rule, so existing
journals are re-dated to their own last row on reopen.

Claude: a backgrounded subagent emits no child frames, so its re-dating
came entirely from roster revisions (task_updated, task_notification) and
from the stale-roster revision written when a journal reopens. Codex: the
roster is revised on every child token-usage report; child-thread rows
carry no producer linkage yet, and read as the session's own until they do.

* fix(native-chat): a reopened journal's verdict on stale work does not date the session

Reopening a journal settles rows the previous host left live (a working
subagent roster, a live background task) to unverifiable. Those revisions
were appended at the reopen moment, and a background-task row is the
session's own non-roster row, so a crash-restarted session with a live
shell was re-dated to the restart although no agent acted.

The reconciler now writes each verdict revision with the row's own observed
time. That is one rule at the one writer, covering both settle shapes; the
render item's observedAt was already pinned to the row's first write, so
nothing the transcript shows changes. The live-transition roster exclusion
stays: live roster revisions are written by the providers, not here.

The `recovered` row flag is not used as the discriminator: the live
unexpected-exit settlement also writes recovered rows, and a clock rule
keyed on it would stop dating a provider crash the host just observed.

* fix(native-chat): date a session by the reducer's attribution of what a row wrote

The clock read producer linkage off the raw row. A lifecycle batch names no
row-level producer, a tombstone names none, and a revision may name none while
the reducer still attributes the item to a subagent, so each of those dated an
idle parent. The clock now asks the reducer: after a row applies, whether any
item it wrote is the session's own work; before a removal, whether the item it
removes was.

* fix(native-chat): date a session's status by its own lifecycle edges

A subagent writes into its parent's journal and keeps going after the parent
settles. Every one of its rows advanced the summary's updatedAt, and the status
row re-dated a done parent to it, so an idle parent read as newly finished and
unread on each child step.

The host now publishes statusStartedAt beside updatedAt: when the session's own
agent entered its status, read off edges only it writes. Idle is when its newest
turn ended; working is when the running turn was requested, or the earliest
send still unanswered; attention is its own oldest pending ask, or a subagent's
when that alone holds it. A turn that recovery settled after its host went away
ended when that settle was written, so it reads as a completion the user has not
seen; it carries no outcome, so no completion event or notification calls it a
success. Render items carry recoveredAt, the recovered row's own write time, so
nothing new is persisted.

The sidebar bridge and the host ingest date the row and the main agent's clock
from it whenever the row shows the main agent's own state, and keep their
existing rules for a row child work holds open or a summary from an older host.
The status feed republishes when the clock moves instead of on every idle row.

* revert(native-chat): keep the journal clock over every row

The row filter this branch put on the reducer's lastActivityAt decided which
rows could date a session: a list of exclusions that each new row kind could
slip past. The session's state is now dated by its own lifecycle edges, so the
filter, its attribution helper and the backdated reopen verdicts go back to
main. lastActivityAt, and the summary's updatedAt it feeds, is again the
evidence clock over every row, including a subagent's.

* fix(native-chat): keep republishing an idle session its live child work holds open

A row held open by live child work is dated by when each reader saw the
publish, and mobile decays a working row whose evidence is older than the
staleness window. Suppressing row-activity republishes for every dated idle
session froze that evidence, so a subagent running more than 30 minutes past
its parent's turn made the row read idle on mobile. Only a session nothing
holds open stays quiet on row activity now; its state clock is unchanged.

* fix(activity): order an agent's timeline by when each state was seen

An answered ask returns a settled parent to its own turn's end, so its done
repeats the time of the done before the ask. Activity keyed and ordered
events by that time: the new done collided with the old one and was
dropped, and the row took its state from the newest-dated event, the
blocked ask, so a done parent read Blocked and needed attention.

Each state switch now records the `updatedAt` it was seen at. Events are
keyed and ordered by that, while unread and "Clear completed" still compare
the state's own time, so the answer neither re-lights unread nor revives a
cleared done. The row's state comes from the pane's own status entry, so a
clear that hid the done cannot leave it reading Blocked either.

* test(activity): pin the timeline across repeated asks, a clear, and a stale turn

Three parts of ordering the timeline by when each state was seen had no test
that failed without them:

- A second ask moves the answered done into history. Both dones share the
  turn's end, so only the history entry's own seen time keeps them apart;
  without it one done collided with the other and the timeline showed two
  Blocked events in a row. Three asks also exceed the per-pane cap, which must
  keep the most recently seen events, not the most recently started.
- "Clear completed" on an answered row must cut off past the ask, which is
  dated after the done, or the cleared row stays listed. A done that the user
  cleared must also stay hidden once a later ask moves it into history.
- A stale working row must not read as running just because the pane's own
  status says working.
2026-09-24 16:17:08 -07:00
Brennan BensonandClaude 6ae6ed08bb fix(claude): open structured chat without a startup deadline, and make Retry start fresh (#22364)
* fix(claude): open structured chat without a startup deadline, and make Retry start fresh

Publish the Claude session as soon as its process is spawned instead of racing
initialize against a fixed 10s deadline. Prompts sent before startup lands are
held and written in order once it does. An exit or sign-in failure before startup
ends the session with the reason and the CLI's stderr.

A create that failed because the process provably exited now carries
ownerVerdict 'exited', so the client marks the launch failed and Retry mints a
new operation instead of replaying the stored failure.

* fix(native-chat): sending into a chat that failed to start restarts it

* fix(native-chat): a send with no live owner restarts it once

A provider child that timed out or exited hands its lease back, and every
later send was refused agent_session_ownership_unknown. Clients read that
code as "not admitted yet" and resend forever, while only a surface hold
could make a new child, once per mount, with its failure swallowed.

The send now routes to a live owner, otherwise restarts one from the
persisted resume state where resume eligibility allows it (single-flight
per session), otherwise refuses with the new settled
agent_session_owner_unrecoverable. Unverifiable, reserved and handed-off
leases are left alone. The desktop hold now logs its failure.

* test(native-chat): pin the unrecoverable refusal as settled in the outbox

* test(native-chat): pin the release clock after a send restarts an unheld owner

* test: read the sent operation id without a cast

* fix(native-chat): type the send-recovery record lookup as the store returns it

* fix(native-chat): a send ensures its owner before admission, and an unheld owner idles for 30 minutes

* fix(native-chat): a create that throws releases its event sink

A child that dies between spawn and journal attach can still write through
the host's event sink, which attach unbound in onAcquiring and never re-bound
because onAttached never ran. The orchestration released that sink only when
performAttach returned a refusal; a thrown failure (the root-exit path) kept
the sink cached with its queued write, so the next attach's drain barrier and
runtime shutdown's flush waited forever.

Also pins the publish-on-root-exit clause for a start that never proved:
deleting it reddened nothing before.

* fix(native-chat): a resend the journal answers restarts nothing, and a send joining a restart rebases from the fence it replaced

* fix(native-chat): the host learns a Claude start positively, and persists only proven options

A publish-first create used to read the session's options before Claude had
answered initialize. With startup pending that read fell back to the built-in
catalog's default, so `record.options.model` was persisted as `sonnet` for
every user whose CLI default is something else; an owner handoff or a reopen
then replayed `set_model('sonnet')` and silently switched their model.

The adapter now reports `started` once startup facts are applied and saved
options restored. The host keeps a `providerChildPhase` on the session it
owns: a starting child hands over nothing but the saved options as intent,
and the `started` event re-reads the options as fact and persists them through
the same record write a user's option change takes. The status summary carries
`hostExecutionPhase` (optional, wire-safe), and the chat pane says the agent is
still starting instead of showing nothing.

A child whose exit already reached the adapter before acquire returns is no
longer handed over as live; the create fails with the CLI's diagnostic.

* fix(native-chat): a hold and a send that find the owner gone share one restart, and a send the ledger already holds restarts nothing

* fix(native-chat): a failed create answers one refusal shape, stamped once at the boundary

A create whose Claude process was seen to exit answered twice in two shapes:
the first call threw a generic runtime error, and only the replay of the same
operation carried the `ownerVerdict: 'exited'` refusal that lets a client
retry under a new operation. Three sites stamped the verdict and the store
failure path stamped nothing.

The first-hand root exit is now returned as the refusal on the first call,
with the provider's own diagnostic as its message. The verdict is stamped in
one place, at the boundary of the attach, from the durable row the operation
settled to, so every refusal shape answers the same fact and no site can
forget it. The per-site stamps are gone.

* fix(native-chat): a send into a session whose child ended restarts it before admission

A session that published and then lost its Claude child before startup (not
signed in, for one) keeps a released lease and a chat the user can still type
into. The send was refused as ownership-unknown, the outbox parked it as
pending admission, and nothing ever restarted the child: the message sat
there until the user closed and reopened the tab.

A send reaching a session with no provider child now runs the same resume a
surface's first hold runs, before the write is admitted. The resume reserves
a new fence, so that send is answered stale with the published fence and the
client's outbox re-drives under it, as after any fence change. A resume that
fails is not this send's answer; admission reports the lease as it stands.

* chore: restore pnpm-lock.yaml to origin/main (local pnpm rewrote it)

* test(native-chat): pin the pre-handover exit as a failed acquire; stub the status feed in the delivery test

An exit the adapter observes before acquire returns now fails the acquire
with the CLI's diagnostic instead of handing over a dead child; the
published-then-ended path stays pinned by the slow-init startup case. The
delivery test renders the pane, which now activates the host status feed.

* test(native-chat): a same-ID re-hold over the wire joins the one resume, and a replay reopen goes on the idle clock

* test(native-chat): a re-hold that joins a failing resume proves one resume ran

* fix(native-chat): a create whose child was proven gone answers the refusal on the first call

The previous change answered a first-hand root exit as the exited refusal on the
first call, but the common failed start never took that path: when the close
ladder proves the whole tree dead the acquisition error is a plain one, the
store-failure classifier rethrows it, and the client still saw a runtime error
first and the refusal only on replay.

The cleanup that proves the child gone now names such a failure
`AgentSessionAcquisitionExitProvenError`, carrying the provider's diagnostic,
unless it already names its own verdict (a refusal, a typed exit proof, a host
store code). The attach answers both proven-exit kinds as the refusal its replay
gives. How a failed acquisition settles and how it is first answered now live
beside the verdict stamp, in the failed-create module.

* test(native-chat): pin the outbox re-driving a stale-refused send under the resumed fence

A send into a session whose child ended is answered stale once the host has
restarted the child. The outbox keeps that operation queued and blocked, and the
fence change the resume publishes re-drives the same operation under the new
fence; the host admits it.

* fix(native-chat): a child restarted for a send nobody holds is still released

The restart a send runs for a childless session takes no holder, on the premise
that the sending surface already holds one. A one-shot writer holds nothing, so
the child it restarted had no release clock and lived until the app quit. The
write resume now arms the clock when no holder is present, as the first-hold
resume already does. The send-after-failed-start cases also pin that the stale
answer's operation is admitted when re-sent under the new fence, and that two
racing sends restart the child once.

* test(native-chat): pin the picked Claude model across a resume whose child starts on its own default

The started event re-reads and persists what the child reports. A resumed child
answers initialize with its CLI default before the saved pick is restored over
it; the record must hold the pick while starting and after started.

* Revert "fix(native-chat): a child restarted for a send nobody holds is still released"

This reverts commit e52c4a6f08.

* Revert "test(native-chat): pin the outbox re-driving a stale-refused send under the resumed fence"

This reverts commit 136a39deb0.

* Revert "fix(native-chat): a send into a session whose child ended restarts it before admission"

This reverts commit 39234e44bf.

* refactor(native-chat): make ensure-owner a step of the serialized send

A send that found the owner gone restarted it OUTSIDE the host's per-session
serialize, through a single-flight resume map shared with the surface hold, then
rebased its fence by heuristic. The attach body is now callable from inside
`serialize` (`attachStructuredAgentSessionUnderSerialize`), and every restart
runs there: a hold, a send's ensure-owner step, provider-exit recovery and the
rewind owner replacement take turns on one queue, so the first to run attaches
and the next finds its child. The single-flight map and `isResuming` are gone.

Admission is two-phase for a send: the ledger's answer comes first and places
nothing; a send it will admit gives the session an owner, and only then are the
row placed and the lease and fence checked. A send it will replay into a closed
session makes the journal readable and spawns nothing. The session entry
records the released fence the child replaced (`resumedFromFence`), so a writer
current as of that owner is admitted at the new fence by bookkeeping, whether it
ran the restart or arrived behind it.

The resume reads its record only after this host has reconciled it and exited
any recovery stage a failed attempt latched, so a hold behind a failed attempt
makes its own attempt against the lease as it now stands.

* fix(native-chat): a Claude start proving itself no longer waits on the CLI

The host handles a Claude child's `started` on the recovery chain every
session's unexpected-exit handling shares, under that session's serialized
step. It then asked the CLI for the model list and settings again, so one slow
CLI held every other session's exit recovery, and its own close, behind up to
two request timeouts.

The adapter already holds those answers when startup proves: the settings read
at startup, the restore's confirmations, and the initialize result the SDK
answers the model list from. `started` now carries that snapshot, and the host
turns it into one record write without any provider I/O.

* test(native-chat): a hold reads its lease only after this host has reconciled it and exited a latched recovery stage

* fix(claude): a chat whose first start failed resumes as the same conversation

A Claude start that dies before initialize writes no transcript, so the next
start launches the chain head's provider id fresh instead of `--resume`. The
launch flag that chose that mode also chose the provider-handle link's origin,
so the fresh launch published a second `created` link onto a chain that
already had a head. The store refused it, the healthy child was closed, and
every later reopen, hold or send spawned and killed another Claude.

The launch now carries the two facts separately: `resumesTranscript` (launch
mode, from whether Claude wrote a transcript) and `continuesChain` (lineage,
from the record's chain head). The link origin reads lineage; rewind and the
Fast opt-in carry-over read launch mode.

* refactor(native-chat): a resume answers with a typed refusal the send classifies

`resumeHeldStructuredAgentSession` and the holds' `ensureProviderChild` answer
`{ ok: true } | { ok: false, refusal }` instead of throwing the refusal code.
The refusal is the attach's own, with its message and, when the failed attach
proved its child gone, its `ownerVerdict`. An attach that settles a failed
acquisition in the ledger and then rethrows the cause is read back off that row,
so a durably failed restart is a refusal and only an unrecorded error is a fault.

The send classifies the refusal through a `Record` over every wire code — a new
code does not compile until it is placed — into transient (the lease is someone
else's to settle; the send runs as the lease stands) or terminal. A terminal
one answers `agent_session_owner_unrecoverable` carrying the cause, forwards the
verdict, and writes the same status row into the chat that a start that failed
leaves, so the user sees why after the error strip is gone. Nothing about the
failure is remembered; a Retry is a fresh attempt. A fault thrown by the restart
itself is reported and the send runs as the lease stands, since bookkeeping
never gates a user's action.

`hold()` still raises the refusal code for its RPC caller.

* fix(native-chat): a child's event sink belongs to the attach attempt that spawned it

The runtime kept one event sink per session id and handed it to every attach.
An attach that acquired a new child unbound that sink first, so when the
acquire then failed its dead child's queued frames stayed in the cached,
unbound sink. The earlier guard only discarded it when no session entry was
left, which a resume of a still-indexed session never satisfies: the next
attach's drain and shutdown's flush waited on it forever. A TUI-to-native
handoff acquire had the same shape.

Each acquiring attempt now mints its own sink. Only a successful attach (or a
proven handoff owner) adopts it as the session's, closing the one it
replaces; any other exit closes it with whatever its child queued. A re-attach
to a live child keeps the sink that child already writes through. A sink that
is not the session's own can no longer force the session's provider down.

The native handoff acquisition moves to its own module, which keeps the
handoff file under its line budget.

* test(native-chat): pin that only the adopted child's event sink still takes writes

Closing a failed attempt's sink and closing the sink a resume replaces were both
unpinned: removing either left every suite green, because neither sink is in the
map that drains and flushes read. The resume test now asserts the failed
attempt's sink and the exited generation's sink refuse writes, and the adopted
one accepts them; deleting either close reddens its own assertion.

* perf(native-chat): the chat reads only the host's startup phase from the status feed

The chat took the whole status summary to read one field, so every status change
for its session (prompt, update time, background tasks) re-rendered the chat
view. It now subscribes with the phase itself as the snapshot, so it re-renders
only when the phase changes.

* fix(native-chat): the startup-phase hook answers a phase or null, never undefined

* fix(native-chat): every restart is counted from the moment it is asked for, and a handoff clears the restart fence

Provider-exit recovery now restarts through the holds' `ensureProviderChild`
like a hold and a send do, so a child whose only surface left while the attach
ran goes on the idle clock instead of living until quit. A hold's resume and a
client attach are tracked as in flight from enqueue, not from their turn on the
queue, so a quit's drain waits for one queued behind a close before it decides
what to evict. A handoff back to native moves the fence in place and now clears
`resumedFromFence`: only a restart may rebase a writer. The failed-restart
status row is keyed by the send's operation id, not the clock, so a resend of
the same id that fails again adds no second row.

* fix(native-chat): a Claude start no longer waits behind another session's exit recovery

The runtime delivered every Claude lifecycle event on the single chain
exit recovery uses so teardown can drain it. That chain orders nothing
across sessions, and an exit recovery on it can run a full reacquisition,
so one chat's `started` waited on an unrelated chat's respawn and kept
its 'still starting' line up. `started` now takes only its own session's
serialized step, is queued the moment it is emitted (ahead of any later
exit of that child), and is tracked in a set the same teardown drain waits
on.

* fix(native-chat): a Claude create that dies at spawn is refused with the CLI's own diagnostic

A CLI that exited before its acquisition handed the child over was
refused with 'claude stream-json for session … exited while being
acquired', or with an unreadable start time, and the stderr the exit
carried (for example 'not signed in') appeared nowhere. The acquisition
now keeps the error its connection ended with and answers with it at
both sites; the generic message is only a fallback when none exists.

* test(native-chat): pin that a reopened Claude chat dying before initialize says why

A resumed start is published at spawn, so a CLI that exits before it
answers initialize fails a chat the user is looking at. Pin that the
open chat is sent the 'stopped before it finished starting' row with the
CLI's diagnostic even when the child's tree cannot be proven gone, and
that a message held for that start is refused rather than left in doubt.

* fix(native-chat): Stop while a Claude start drains its held prompts withdraws the rest

Stop withdrew held prompts only while startup was pending. Once startup
landed and the gate began writing them one by one, a Stop interrupted the
CLI and the prompts still waiting were written straight after it. Stop
now withdraws whatever the gate still holds in both states; the drain
takes each prompt off the queue immediately before writing it, so a
withdrawn prompt can never be written. The one already written still
gets the interrupt.

* fix(native-chat): the release clock keeps a session that still owes a sent message

When the last surface stops holding a session, the release clock evicts
it after the grace unless a turn is running. A message sent while Claude
is still starting is held, not running, so switching away from that chat
for the grace evicted the session and refused a message the user had
already sent. The clock now asks whether the session owes work: a running
turn, or a submission the provider has not taken yet (still pending in
the journal). Both are read from the journal; nothing new is stored. A
starting session that owes nothing is still released, and an explicit
close still ends everything.

* fix(native-chat): only a start that holds a sent message keeps a released session

The release clock kept any session with a pending submission. A Codex
send is admitted and stays pending until its echo, which may never come,
and only an eviction retires it, so such a session was never released
while the app ran. A pending send now keeps the session only while its
child is still starting, which is when the send is held for that start.
Pins that a ready session with an unechoed send is evicted, and that a
Claude chat whose turn finished is released after the grace.

* fix(native-chat): a send waits for the owner it met to prove its start before it is admitted

A Claude child is published before the CLI has answered initialize, so a send admitted right
behind a restart — or right behind the first start — was dispatched into a child that could die
milliseconds later, and learned of the death only as a delivery nobody could confirm. The terminal
refusal the send was written to give was unreachable on the real adapter for exactly the failure
it was written for.

The send's serialized step now admits nothing against a `starting` child. It registers for the
child's startup verdict and returns having placed nothing; the send waits off the session's queue
(the `started` and `ended` settlements run on it) and admits again once the child is `ready`, or is
refused `agent_session_owner_unrecoverable` with the child's own exit reason when it exits first.
The exit settlement writes the one status row, decided by the host's own phase rather than only
the provider's flag. A close, an eviction or a replacement answers the wait too, and quit releases
whatever is left; there is no timer. One spawn per user action holds across re-entries.

* fix(native-chat): restart the release grace when a start writes its held prompts

A prompt held while Claude starts is written when the start lands, but
its turn opens only when Claude echoes it. The release clock stopped
counting it once the child read ready, so a tick landing in that gap
stopped the child before it ran the user's first message. The start
landing now restarts a pending release's full grace, the same grace a
message sent to a ready chat gets before it is released.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): a starting child owns the send; the adapter holds the message for its start

A send that meets a child still proving its start is admitted against it, as it was before the
off-queue startup wait: the adapter holds the message until startup lands and rejects it with the
child's own diagnostic when the child dies first, the exit settlement writes that cause into the
chat, and the release clock keeps a starting session that holds a sent message. The startup watch,
the off-queue wait loop and their teardown phase are gone; the exit settlement still reads a start
that failed off the host's own phase when the provider omits the flag.

The scripted-CLI test now pins that contract end to end: a restart a send asked for whose CLI dies
at initialize leaves the message rejected with the diagnostic, one row naming it, and the fence
moved by two; a healthy CLI is restarted once and written to; a send during the first start is
held and written once initialize answers, or rejected with the diagnostic when the CLI dies.

* fix(native-chat): a failed send restart says why, and offers a new chat only when nothing can restart it

The refusal a send gets when the host cannot restart the chat's agent is renamed
agent_session_owner_restart_failed and now reads "<Agent> couldn't restart: <reason>." with the
restart's own cause. "Start a new chat to continue." is added only when the resume was refused
because this host has no record to restart from or cannot run the one it has. Any other failure,
such as a CLI that is not signed in, leaves the chat retryable: the outbox stops auto-retrying, and
a manual Retry or a new send tries the restart again, since a refusal before admission leaves no
ledger row.

* fix(native-chat): a Claude start skips an option write the CLI never answers instead of faulting at the request deadline

* test(native-chat): wait for the recovery's reserved lease, not the released one it replaces at once

* fix(native-chat): a send whose restarted child dies before starting is rejected with the child's diagnostic

A child that never proved its start has accepted nothing: input is written
only after it initializes. A send admitted against such a child, whose
dispatch then found no session, settled unknown, twice, and the outbox took
Retry away. It now settles rejected with the child's own diagnostic, both
when the dispatch throws and when the exit settles the sends it left
unanswered, so the chat says why and offers Retry. A proven child's
unanswered sends stay in doubt, as before.

* test(native-chat): expect a send held for a start that never proved itself to settle rejected

* fix(claude): write a prompt held after startup already drained, instead of stranding it pending

* test(native-chat): pin that a send to a child that died before starting is answered rejected

* test(native-chat): leave the cast exit-session fixtures as they were, since a proven exit never rejects

* fix(claude): a saved option the CLI never answered stays saved instead of being replaced by the CLI's value

A start skips an option write the CLI does not answer within the request
deadline, and then persisted what the CLI reported in its place, so a slow
answer silently replaced the user's saved model or dropped their saved
permission mode. Silence is not a refusal: the unanswered option is now
recorded apart from a rejected one, the live child keeps running on the
CLI's value, and the saved choice stays on the record for the next start to
retry. An option the CLI rejects is still dropped as before.

* fix(native-chat): a rejected send opens no turn, so the row naming why it failed is not folded away

A send whose restarted child died before starting is rejected, and the
exit writes a row naming the cause. The chat's local clock had watched the
send go pending and stop, so it gave the message "Worked for 0s"; that
settled a turn that never ran, and the fold hid every non-prose row after
the message behind it, including the one naming the cause. The row only
appeared when a later send moved the turn anchor, which read as two rows
for one Retry. The host's journal already says the send was rejected; it
now answers that such a message opened no turn, which outranks the local
clock on desktop and mobile alike. A rejected send whose journal does
record a turn keeps its duration.

* fix(native-chat): a send whose restart died starting leaves the same row as any start that died

One failed attempt already leaves one row, but which row depended on when
the child died. A child that died after the send was admitted left "The
provider stopped before it finished starting: <cause>."; one that died
before the send was admitted left "Claude couldn't restart: <cause>." So the
same failure read two ways from one Retry to the next. When the refused
restart proved its child exited, the send now writes the startup-failure row
itself, as its comment always said it did. The refusal under the composer
still says the restart failed; a restart that failed for a reason other than
a child exiting keeps its own wording.

* fix(native-chat): a send rejected because the agent never started names the cause under the composer

When the child a send was admitted against died before starting, the host
rejected the send with the child's diagnostic behind the internal transport
marker. The client rightly hides that marker's detail, so the red line read
"Couldn't reach the agent" while the cause sat in the record. A startup
death is not a failed write: the host now words that rejection the way the
chat row does, "The provider stopped before it finished starting: <cause>.",
at every site that rejects for it. Desktop and mobile show a reason in words
verbatim already, and older clients do too, so no client change is needed.
Real write failures keep the marker and the generic copy.

* fix(claude): a saved option the CLI never answered survives a later change to a different option

The saved choice a start could not apply was kept on the record, but the next
option the user set persisted only what the child had applied, so changing the
permission mode or effort, or clearing the chat, silently dropped the saved
model. The adapter now reports which saved options are still unanswered, every
option write keeps those saved values, and a write the child accepts for that
option retires it.

* fix(claude): a send that meets a child whose exit already settled names that exit's cause

When the child a send was admitted against died starting and its exit
finished settling before the send reached it, the send was rejected with
"no live claude stream-json session for <id>", now shown under the composer
as the cause. The adapter keeps a settled exit's diagnostic until the chat is
acquired or closed again, so that send names what the CLI said. A refused
restart whose child died at spawn or while its start time was read is pinned
to leave one row in the words any failed start uses.

* test(native-chat): pin the words an exit settlement rejects a never-started send with

The startup gate and the dispatch reject a send first in every existing
scenario, so the exit settlement's own rejection had no test of its wording.

* fix(claude): derive which saved options are still unanswered from what the child applied

A write that lands already puts its option in the session's applied set, so
the unanswered list is that list minus what has since been applied, rather
than a second copy every option write must remember to edit. Session
fixtures built without the new set no longer throw on an ordinary write.

* fix(native-chat): a cleared chat starts from a saved choice the child never answered

Clearing a chat seeded the replacement from the values the child reported,
so a saved model or effort whose restore write the CLI never answered was
replaced by the CLI's own value in the new chat, even though the retired
record kept it. The replacement now keeps those saved values too, and its
start retries them.

* fix(claude): closing a chat forgets its exit's diagnostic even when the exit settles during the close

The diagnostic was dropped when the close began, but closing over an exit
that was still settling finishes that settlement, which kept it again, so a
closed or deleted chat held it until its next acquire. It is now dropped once
the close finishes. Pins that an acquire and a close each retire it.

* refactor(claude): keep a saved option the CLI never answered as the wanted value, not a list beside it

A restore cleared the session's wanted options and added back only the writes
the CLI answered, so an unanswered one lost the user's value and every later
writer had to be told to put it back: the start report, each option change and
/clear each carried a list of unanswered keys. The restore now keeps the saved
value as wanted and unconfirmed, so what the session reports and persists
already carries it, and the list, its adapter method and the started-event
field are gone. A refused option is still dropped.

/clear now starts the replacement from the record's options instead of reading
the child's live values, which can be a model the CLI fell back to.

* test(claude): wait for the start to finish before changing the saved model

The record holds the saved model from creation, so waiting for it returned at
once and the option write could reach Claude while it was still starting,
which refuses it. Wait for the effort the finished start reports instead.

* fix(i18n): translate the still-starting chat notice

The notice that a structured chat is still starting was only in English.

* test(claude): pin the failed acquisition's own reading-control release

The merge re-pointed this test at a child that exits after publish, where the
exit path also releases the binding, so it passed with the acquisition's release
deleted. A child that exits before publish leaves only that release. Also drops
the create 'init' phase, which lost its last producer when rewind stopped
proving before publish.

* refactor(claude): move unexpected-exit handling into the exit lifecycle module

The adapter crossed the 300-line limit once main's context-usage change
landed beside this branch's growth. The two methods that turn a Claude
process exit into an ended event now live next to the existing exit
helpers; behavior is unchanged.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-24 15:23:15 -07:00
Brennan Benson f0b3f44f10 feat(agent-session): let the host own a chat's tab id and let a create reserve it (#22616)
* feat(agent-session): let the host own a chat's tab id and let a create reserve it

A structured chat's tab id was derived from its session id by every layer that
needed one: the renderer, the host snapshot and the status address each built
their own spelling. The join between a conversation and the tab that shows it
must be a pointer the host owns, not a derivation each client repeats.

The session record now carries surfaceTabId. A create pins it: the tab half of
the pane agent.launch reserved, an optional tabId on agentSession.create, or a
host-minted UUID. Records written before the field existed are backfilled at
open with the string clients derived, in memory at once and on disk with the
store's first transaction, so nothing keyed by it (read state, notification
ids, worker rows) moves on upgrade. A second record under a held id is refused.

Only the record and the two create wires change here. The snapshot still
publishes agent-session:<sid> and the renderer still derives its local id;
those move in the next two changes. agentSession.create is a strict object, so
the field is advertised as a capability a client checks before sending it.

* fix(agent-session): record the derived tab id for an unreserved create

A create that reserved no tab minted a random UUID that no reader uses: the
renderer, status address, worker rows and host-shared read state all still key
by structured-agent-session-<sid>. Persisted, that id would move every chat
created before readers switch to the recorded one, orphaning its read state
and worker rows the way the backfill exists to prevent. An unreserved create
now records the derived id, the same rule the backfill applies, so the record
always matches the prefix every existing key uses; an opaque mint belongs with
the change that moves the last reader.

Also:
- a chat tab id must be a host tab id on the record, the create wire and in
  admission, matching what agent.launch already requires of paneKey; a
  web-surface id would decode as another tab
- the stored launch-result guard checks the structured outcome's tabId
- comments no longer claim a retry naming another tab conflicts; replay keys
  on the attach fingerprint and answers with the recorded id (now pinned)
- the wire refusal test used a non-hex digest, so the schema refused it for
  that reason; it now reaches the tab id rule
- pin that the reload path refills the id without forcing a save

* test(agent-session): correct the tab-id fingerprint comment to match replay
2026-09-24 14:42:12 -07:00
Brennan Benson 98584332a3 fix(native-chat): record which Codex agent produced each journal row (#22532)
* fix(journal): a batch revision restates the producer of each row it revises

The reducer rebuilds a row's producer linkage from its NEWEST revision, and
absence is a positive claim: no agent id means the session's own agent wrote
the row. So any revision written without the stamp hands a subagent's row
back to its parent, permanently.

Three host paths revise rows they did not write, from the render item they
already hold, and all three dropped the stamp:
- answering a prompt re-appended the asker's row with the fence only;
- dead-generation settlement failed running tool calls and cancelled pending
  prompts through a lifecycle batch;
- stale-session settlement on acquire cancelled lost prompts the same way.

The batch path could not carry a producer at all: linkage was removed from
the batch row because one row covers N mutations, with a note that a mixed
batch would have to stamp per mutation. Dead-generation settlement is such a
batch already, and Codex settlement is about to become one. So each item
mutation now names its own producer, inline like the row base. A mutation
that names none falls back to the row's linkage, which is what a batch read
before. Parse sanitizes a bad per-mutation id the same way it does a row's:
the field is dropped and the mutation kept.

No schema version bump. An older host's mutation validator ignores unknown
keys, so it reads a stamped mutation as the session's own, which is exactly
what it shows today. Old journals carry no stamp and read as before.

Turn revisions still carry nothing: a turn is the session's unit of work,
and the live-turn scans rely on a turn row never carrying linkage. The note
recording that invariant is updated to the new write sites.

* fix(native-chat): attribute a Codex subagent's journal rows to the subagent

Codex journals every thread on its app-server connection into the session's
journal, and a spawned subagent's items arrive on the child's own thread.
None of those rows carried producer linkage, so under the journal's rule that
absence means the session's own agent wrote a row, every child's command,
message, reasoning, prompt and status row read as the PARENT's: the parent
could show its child's running command, its child's reasoning as "thinking",
and its child's prose as its own latest line.

The Claude lane's model is reused, not reinvented: the same fields and the
same absence rule. What differs is how the producer is known. Orca opens
exactly one thread per app-server, so any other thread is one Codex spawned.
That decides WHETHER a row is a child's from its first frame, announced or
not, and the thread id is final at once: it is never re-minted the way a
tool-call reference is, so no correction ledger is needed for identity.

- agentId: the child thread id, the same id the status side keys a Codex
  child on.
- parentAgentId: the thread whose stream carried the child's `started`
  activity. Codex emits that item on the spawning agent's own session, so a
  child that spawned a grandchild is named; the session's own thread is not.
  Other activity kinds ride whichever agent acted and are not used.
- producerKind: 'agent'.
- attempt: which run of the child the row's own turn was, counted from the
  child turns the roster already observes; absent on the first run. Taken
  from the row's turn rather than the child's latest, so a persistent shell
  that outlives its turn keeps its run across revisions.
- providerParentRef is omitted: a Codex child's frames carry no parent
  reference of their own beyond the thread id, which is already agentId.

One resolver, owned by the roster (which already owns what is known about
each child thread), is handed to every writer: items, streams, generic and
summary rows, prompts, compactions, goals, and the three settlement batches.
The session-end settlement mixes every thread's rows in one batch, so each
mutation names its own producer. Turn rows stay unstamped: Codex writes them
only for the primary thread.

The spawn-group roster row stays unstamped on purpose: a child's frame can
trigger its write, but it is the parent's list of its children.

The translator's construction moves to a parts module so the translator
stays a router under the line cap, and the item streams reuse one
append-and-publish helper instead of two copies. Children are never swept
at turn end; nothing here changes that.

* test(native-chat): pin Codex subagent attribution at every writer and every parent reader

Two layers, so a stamp that is correct in the store and never read, or read
and never persisted, cannot pass.

The readers, through the real path: translator, deferred sink, on-disk
journal, snapshot. Each is a defect on main: the parent named its child's
running command as its own tool, read its child's reasoning as itself
thinking, showed its child's compaction as its activity line, and quoted its
child's prose as its latest line (checked after closing and reopening the
journal, so the stamp is read back from disk). The transcript still renders
the child's rows.

The writers, through a sink that records the linkage of every plain append,
batch mutation and lifecycle transition: start, streamed checkpoint and
completion of one command all restate the child; a row that beats the spawn
announcement is still the child's; a grandchild names the child that
announced it, while an `interacted` activity names no parent; a follow-up
turn is the child's second run, and a shell that outlives its turn keeps its
own; the exit batch settles each thread's rows under its own producer and
the turn row under none; a child's provider frames, approval and goal rows
are its own; nothing is stamped while the session thread is still opening;
and the spawn-group row stays the parent's.

* test(native-chat): pin linkage forwarding on the sink's lifecycle-transition path

A Codex child's goal row is written through a lifecycle transition, so a sink
that forwarded only the fence there would file the child's goal as the
session's own.

* test(native-chat): type the Codex item fixtures as thread items

* refactor(journal): keep a row's producer across revisions that name none

The reducer took a row's producer linkage from its newest revision, so every
writer that revised a row it did not write - a prompt answer, a dead-generation
or stale-session settlement, the reopen sweep of stale subagent rosters - had to
restate the producer or silently hand a subagent's row to the session's own
agent. Three of those writers had been patched to restate it; the next one to
forget would reintroduce the bug.

Attribution is now fixed by a row's first write. A revision that names no
producer keeps the row's existing linkage; one that names any replaces the
whole bundle, which is how a provisional stamp is still corrected in place. A
row re-created after a tombstone starts with nothing. The reducer runs the same
fold on replay, so the kept producer survives a reopen.

The three restatements are removed. Per-mutation linkage on lifecycle batches
stays: a batch can create a row (a Codex child's prompt, or a child's item
settled before any checkpoint landed) and one batch can mix producers.

* test(journal): pin producer inheritance in the reducer and across a reopen

A revision naming no producer keeps the row's, on the plain item path and in
a batch settling a child's row beside the session's own; one naming any
replaces the bundle wholesale; a tombstone clears it; a stale revision cannot
touch it; and a reopened journal replays it exactly as it was folded live.

* refactor(codex): name the translator's writer factory for what it builds

* docs(codex): say why a settled row names its producer

* test(journal): drop a producer test the stale-revision guards make unreachable

The stale revision is dropped whole by two independent guards before the
inheritance rule runs, so its producer assertion could never fail; the
reducer's own stale-revision tests already cover the drop. Also say what
the batch sink does forward: each mutation's own producer.
2026-09-24 13:33:56 -07:00
Jinwoo Hong 5610b11703 feat(ipynb): create a .venv when pip is locked out, and show ipykernel setup progress (#22710)
* feat(ipynb): set up ipykernel in a new .venv when pip is locked out, and show install progress in the dialog

* fix(ipynb): drop the retired installFailed string from translated catalogs

* refactor(ipynb): drive the setup dialog from one setup state; fix review findings

- Kernel status now only describes the kernel; a single `setup` object (base, offer, phase, error) drives the dialog, replacing the extra statuses and the externallyManaged/setupError fields.
- The picker's 'Create virtual environment…' opens the same dialog; success switches through selectEnvironment, so a running kernel is only replaced once the venv exists.
- Async setup results are dropped when the dialog they belong to is gone (tab closed/reopened).
- main verifies ipykernel imports after pip, reuses an existing .venv instead of re-running venv over it, and explains failures that printed nothing.
- Windows copy command guards the install with if ($?); notebooks at a filesystem root get a correct .venv parent; 'Try again' shows for both retry paths.

* fix(ipynb): close the setup prompt when another Python is picked
2026-09-24 16:22:34 -04:00
Brennan Benson 85642d0d88 fix(agent-status): count only agent work in stats, and read a Grok background subagent as working (#22474)
* fix(agent-status): renderer and recorder consumers read the question they mean

Since #22295 a row's combined `state` reads `working` both while the lead's
turn runs and while a subagent or background shell outlives a settled lead.
The lead's own state now rides beside it (`lead`); each consumer in this slice
reads the question it actually asks.

- Smart sort and the Activity unread badge keep reading the combined state:
  their classes and rows are what the sidebar shows. Pinned with tests,
  including a restored `lead.state: 'working'` row that must never read live.
- The stats recorder asks "was an agent executing" and now reads a shared
  derivation (`isAgentExecutionOwed`): the lead's turn, or live agent child
  work holding a settled lead's row open. A settled lead's background shell
  no longer accrues "Time agents worked". Old hosts without `lead` fall back
  to today's read; restored and replayed rows still never open a session.
- The `agent.status.changed` plugin event gains `lead` as an optional field
  through one tested projection; `state` keeps its meaning and restored rows
  still project to nothing.
- A Codex root Stop that follows an inferred interrupt keeps the
  `cancellation` verdict, as the Claude lane already does at its turn
  boundary, on both the hook and relay paths.

* test(agent-status): pin that a child's approval wait no longer splits the recorded span

The recorder's move to the lead fact quietly changed one more story: a Codex
child's PermissionRequest turns the combined row waiting while the root's own
turn keeps running. The old state read closed the span there and minted a
second spawn on resume; the new read keeps one span, because the lead never
stopped. Pin it at both boundaries (the shared derivation and the recorder)
so the change is deliberate, not incidental.

* fix(agent-status): date stats edges by this host's clocks and scope the accrual predicate to stats

The recorder dated a start by the producer's mainAgent.stateStartedAt. An SSH
host stamps that with its own clock while every stop is stamped locally, so each
span gained or lost the clock skew. The same clock also survives a row that
briefly lost the fact (an OSC repaint to another state), dating the reopen
before the close already sent, and a subagent reopening a monitoring row took
the row clock the hook lane pins to the main agent's turn start, re-billing the
whole watch-loop window. Edges now use the row clock when the row settles or
pauses and the evidence clock otherwise.

Rename isAgentExecutionOwed to isAgentTimeAccruing and state that it is the
stats question, not a liveness gate: it excludes watch loops, which lifecycle
gates must keep treating as live. Note on the plugin schema that
mainAgent.stateStartedAt is the execution host's clock.

* fix(agent-status): pause agent time while the row waits on the user, whoever asked

Time agents worked now accrues only while the combined row reads working. A
child's approval or question wait pauses the clock exactly like the main
agent's own prompt, and the pause edge is dated by the row's own clock.

* refactor(agent-status): read agent time from the combined row and date edges by the row's own state

Time agents worked now accrues while the combined row reads working and is not
a watch loop. The shared fold emits monitoring only for a settled main agent, and
hosts that predate the main agent fact did the same, so this is the same answer
on every new-host row without reading mainAgent, and it applies the watch-loop
rule to older hosts too instead of billing their monitoring windows.

An edge that leaves working is dated by the row's state clock; an edge inside
working is dated by the evidence clock. This also stops a live repeat of a
hydrated working row (any row without the main agent fact, such as an OSC row)
from dating its start at the persisted state clock from the earlier runtime.

* fix(grok): read a background subagent as agent work, not a watch loop

Grok's end-of-turn Stop lists each in-flight background task with its type
(shell, monitor or subagent). Orca filed a running subagent with the shells,
so a Grok subagent that outlived the main agent read "Monitoring background
tasks" and, with the stats recorder now skipping watch loops, stopped the
"Time agents worked" clock. Map shell and subagent entries to the shared
child-work kinds and let the shared liveness classifier decide: any live
subagent keeps the pane working, a shell alone or an active stop hook stays
monitoring, monitors stay excluded.

* test(agent-status): cover a waiting child in the fold's every-input accrual check

Since the shared fold learned a child's human wait, a waiting child makes the row wait, so it
must not accrue agent time whatever the main agent is doing. The exhaustive check now includes
that input.
2026-09-24 10:12:16 -07:00
Brennan BensonandClaude ad6cb0e05c fix(worktrees): version every catalog publication so a stale listing cannot undo a create (#22507)
* fix(worktrees): version every catalog publication so a stale listing cannot undo a create

A worktree listing is a snapshot from when its scan began. The renderer treated any
authoritative listing that lacked a known worktree as proof of deletion, judged at apply
time against the live store, so a listing delayed past a create reply purged the new
workspace: tabs wiped, selection cleared to the landing, pending structured launch
tombstoned so the host session was closed the moment it published. #22311 re-runs a scan a
mutation overtakes, which covers a bump during the scan but not a reply that is simply
applied late, on the host or in the renderer, or a refresh that joined an older one.

The host now stamps every listing with the catalog version its scan began at (the existing
per-repo scan generation, scoped by a per-process epoch) and every create and remove reply
with the version the mutation produced. Clients keep the newest version applied per repo
and host; a listing older than that is not applied at all, not its rows, not its purge, not
the pre-merge terminal teardown. Coalesced joiners inherit the reply and therefore the rule.
Fields are optional on the wire; older hosts and clients keep today's behavior.

* fix(worktrees): relist after a refused stale listing and stop version churn

- A refused listing can be the only answer a caller gets (a change-event
  refresh that joined an older in-flight listing), so fetchWorktrees lists
  once more; that listing scans at or past the applied version.
- An equal catalog version keeps the held object, so a no-op listing no
  longer writes store state on every refresh.
- Versions the client cannot order are treated as unstamped at ingest.
- worktree.rm takes the repo from its id selector instead of resolving the
  worktree a second time, which also stamped nothing for an id two hosts share.
- A removal on one of two hosts sharing a worktree id records its version.

* test(worktrees): pin the scan-generation bump right after git worktree add on every create path

A listing is stamped with the generation its scan began at, so a create must
advance it before any post-add work can yield. Pins the local desktop, SSH and
runtime local create paths.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(worktrees): bump the scan generation right after git worktree remove

A listing is stamped with the generation its scan began at. Removals bumped it
only at the end, after watcher release, push-target cleanup and the metadata
purge, so a listing scanned before the git removal and one scanned after it
could carry the same sequence. Applied out of order, the older one restored the
removed row until the removal reply. Bump right after the git removal on the
desktop local, desktop SSH and runtime paths, as creates already do.

The ordering pins now witness the generation at the first step after the git
mutation rather than at the re-list, so moving a bump past any intervening
await fails them.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(runtime): stop leftover worker-recovery retries from firing into later tests

The legacy worker terminal recovery retry timer re-arms itself and had no
way to end, so a runtime from one aggregator test kept rescanning repo-1
through the shared listing mock during later tests, consuming the listing a
lineage create expected ("Worktree created but not found in listing").

Give the controller a dispose() that cancels pending retries and refuses to
re-arm, and have the runtime test harness dispose every controller it
constructed after each test.

* Revert "fix(runtime): stop leftover worker-recovery retries from firing into later tests"

This reverts commit 5879110163. The leaked
recovery retry timer is a pre-existing test-harness flake that also hits
main; it belongs in its own change, not in the catalog-version fix.

* test(worktrees): pin the relist bound and the teardown gate on fetch-all and paired runtime listings

The relist after a refused stale listing had no test holding it to one retry, and the
pre-merge terminal teardown gate was pinned only on the direct fetchWorktrees path: removing
it from fetchAllWorktrees or from the paired-runtime listing path left every suite green.

* fix(worktrees): keep an SSH reconnect going when its listing is older than an applied create

A direct SSH listing refused because a newer catalog is already applied for
that host reported 'stale', the same result as a moved connection. The
reconnect preparation ends on any 'stale' repo and skips the post-connect
workspace sync and terminal correction, and nothing retries that while the
connection holds. The refused listing now reports the host's answer, since
the store already holds a newer catalog; 'stale' stays for a moved
connection or owner.

* test(worktrees): pin the listing teardown gate on the fetchAllWorktrees startup hydration pass

The hydration pass lists through its own call site, and removing its gate left every suite green.

* fix(worktrees): decide a listing's refusal reason inside the merge, and defer the startup purge behind a newer create

The listing merge now returns 'applied', 'superseded' (a newer catalog is already
applied) or 'not-current' (its connection or repo owner moved), decided against live
state inside the store update. Callers switch on it instead of re-deriving the reason
afterwards, which misreported an owner that went away during an older listing as current.

The one-shot startup purge keeps only ids from scanned rows. A create applied after a
repo's listing but before the purge wrote only the visible rows, so the purge closed the
new workspace's tabs, chat tab included. It now defers when any repo's listing is older
than that repo's applied catalog, like a refused listing, and runs on the next pass.

* test(worktrees): pin the startup purge deferral on a listing refused for a changed repo owner

* test(worktrees): pin a direct-authority fetch reporting an older listing as current

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-24 00:46:29 -07:00
Brennan Benson 25d7c21fcb feat(native-chat): show context window usage in the composer (#22301)
* refactor(native-chat): move the composer's stop action into its own hook

The composer is at its line budget; lifting the stop action out makes room
for the context usage ring without changing what Stop does.

* feat(native-chat): record the Claude CLI's context window facts on the structured journal

A structured Claude session now keeps what the CLI says about its context
window on journal rows, so every client reads the same answer and a restart
replays it:

- Each main-thread assistant response keeps its API usage. A subagent's
  response measures its own window, so it carries none.
- The turn a result settles records the session's window: the largest
  contextWindow across the result's per-model usage, since side calls to a
  smaller model report their own smaller window.
- After a result and after a compaction boundary the host asks the CLI for
  its /context breakdown (5s bound) and records the answer on the current or
  last turn. An answer is dropped when the main conversation moved, a send was
  accepted, a newer request was issued, or the session was released while it
  was in flight; a failure or an older CLI leaves the row unchanged.
- A compaction boundary or conversation reset records that the used count is
  unknown until the next response or report, so the pre-compaction size is
  never shown as current.

Every part carries its own host clock, and the reader takes the newest, since
a revised turn row keeps its place in the transcript. The persisted validator
admits every value the writer can write, including a zero auto-compact
threshold: a row replay rejects truncates the journal from that row.

* feat(native-chat): show context window usage in the composer

A structured Claude chat shows a ring beside send once the journal can state
the session's context usage. Hovering shows used/window with a bar and, when
the CLI has reported its breakdown, one row per CLI category as a share of the
window, largest first. Between reports the ring shows the newest response's
usage against the newest window the CLI reported, marked as estimated. Before
the CLI has reported any window, and after a compaction or reset until the next
response, there is no ring. A terminal-backed chat shows none.

* fix(native-chat): measure the context ring against the main thread's model window

The result's per-model usage is cumulative across the session and includes
subagents and side calls, so the largest window was often not the one the
main conversation runs in: after switching from a 1M model to a 200k one, or
when a subagent ran on a larger-window model, the ring read against the
wrong window after every turn. Pick the entry named by the model that served
the newest main-thread response, and among its [1m]/non-[1m] entries the one
the result moved; fall back to the largest only when nothing names it.

* fix(native-chat): keep the context ring moving through tool-only responses

The live estimate lived on assistant message rows, and a response with only
tool calls or thinking writes no message row, so the ring froze through long
tool loops and stayed hidden after a mid-turn auto-compaction until the next
text reply. Record every main-thread response's usage on its turn row
instead, once per response, so the selector sees each one.

* fix(native-chat): read the ring's window from the model the turn's init names

The CLI keys per-model usage by the main loop's model string, [1m] included,
and every turn's system/init frame carries that exact string, while a
response drops the suffix. Match the init's key first, so a session that
switched between the 1M and 200k variants of one model reads the right
window; fall back to the newest response's model, then the largest entry.

* fix(native-chat): correct the composer's control-order note for the context ring

* fix(native-chat): keep a running turn's context facts when the host settles it

A turn row now carries the live context estimate while it runs. When the host
settles a running row itself (a crashed or stale generation, a close the
translator never saw), it rebuilt the record field by field and dropped those
facts, so after a crash the ring fell back to an older turn's size, or to a
pre-compaction size the dropped reset had superseded.

* refactor(native-chat): revise Claude turn rows from the journal so the ring survives a restart

The context ring's facts were written to turn rows through an in-memory list
of recent turns. A new translator is built on every acquisition, so after a
restart or reattach that list was empty and every fact for a turn that was
not open was dropped: a /compact as the first action after a restart never
cleared the ring and never showed the fresh breakdown.

Every Claude turn-row write is now a revision of the row as the bound journal
holds it when the write runs. The sink gains a resolved revision that reads
the target row and its body at execution; the queue runs one operation at a
time, so the read-modify-write cannot interleave, and revisions are never
coalesced. Lifecycle writes own the lifecycle fields and context writes own
contextUsage; each keeps every other field. Only the open turn is kept in
memory. A fact with no open turn lands on the newest turn row, and a report
lands on the turn it was requested for.

The persisted facts are simplified to a window, which now names the model it
was measured for, and a single used part (report, estimate or unknown) that
each write replaces. The ring reads the newest turn row carrying each part,
and hides an estimate whose model the window was not measured for instead of
dividing by another model's window. A turn opening, and a reset, count as
activity, so a late report can never land behind a newer turn.

Host settlement of a stale running turn now drops only the fields its verdict
owns, so context facts and any field a newer build wrote survive it.

* fix(native-chat): keep a turn row whose context facts this build cannot read

Context facts are validated deeply, so one malformed or future-shaped fact
made the whole turn row malformed, and replay truncates the journal from that
row on. Replay now drops unreadable facts from a turn row, in item rows and in
settlement batches, and keeps the row, the same way it already drops producer
linkage it cannot trust. The ring shows nothing for that turn instead of the
session losing its history.

* test(native-chat): pin that a child exit mid-turn keeps the ring's last size

A lifecycle-only revision, the end a turn gets when its child exits without a
result, must keep the context facts the row already carries.

* perf(native-chat): revise a named Claude turn row by key instead of scanning the journal

Every Claude turn-row write walked every reduced journal item to find its
row, even when it already knew the row's identity, so a long session paid
O(items) per write on the main process. The journal now answers a keyed read,
and a context report names its turn by row identity rather than turn id, so
only a write made while no turn is open still scans.

* fix(native-chat): tell a 1M window from a 200k one of the same model

Responses drop the [1m] suffix, so after a switch between the 1M and 200k
windows of one model the running turn was measured against the previous
turn's window until its result arrived. An estimate now records the turn's
init model, which keys the window exactly, and the reader requires the full
model id to match.

* fix(native-chat): show no ring for a context kind a newer host writes

A paired client reads turn rows from the host unvalidated, so a used-count
kind this build does not know fell through to the estimate branch and threw
reading its missing usage. Only the kinds this build can measure now produce
a ring.

* fix(native-chat): keep the context ring through plan-mode turns on another model

Plan mode can run a turn on a model the turn's init does not name (opusplan
upgrades to Opus's 1M window). The estimate then carried only the response's
id, which drops [1m], and the exact comparison against the window hid the ring
for every plan-mode turn. The estimate now records the response's id beside
the init's exact key, and the reader matches the base model only when no exact
key was recorded.

* refactor(native-chat): pair the context ring's window by model change, not by model id

The ring divided the newest response's size by the newest window only when
their model ids matched, which meant comparing ids from the init frame, the
response, per-model usage keys and canonical ids. Those disagree in plan mode
and across 1M and 200k windows of one model.

The writer now knows when the model may have changed: after a model or
permission-mode write that changes the value, when a restore cannot put the
stored model back, and when a main-thread response comes from a different
model than the one the window serves (an approved plan). It then marks the
size unknown, holds estimates, and asks the CLI for its context report, which
states the new model's window. Any new window, from a report or a turn
result, releases the hold. The reader compares nothing: a report, or the
newest estimate over the newest window.

Turn rows no longer store window.model, window.canonicalModel,
estimate.model or estimate.responseModel.

* fix(native-chat): keep a late context report's window when only its count went stale

* test(native-chat): pin that each turn's init lets its result restate the context window

* fix(native-chat): open the context card on click and tap

* fix(native-chat): wait a beat before a mouse hover opens the context card

* fix(native-chat): publish each context write in the operation that makes it

A context report answers after the turn's last frame, so a revision that waited
for the next frame's publish reached live clients only on the next turn. Context
writes now queue their revision and its publication as one operation.

* fix(native-chat): write million-token counts with a capital M

A lowercase m read as minutes on the context card.

* fix(native-chat): keep the context card open while the pointer crosses into it

The card closed the moment a mouse left the ring, so the pointer could not
cross the gap into the card. Leaving now waits a beat, and entering the card
cancels the close.

* fix(native-chat): let Escape close the context card without stopping the agent

The card keeps focus in the composer, so the Escape that closed it also
reached the composer and interrupted the running turn. The composer now
skips an Escape an open layer already handled.

* fix(native-chat): show the context ring when the chat has not loaded the turn it belongs to

The ring read context facts only from the rows the chat had loaded, so a
reopened chat whose recent page started after the last turn row, or a live
turn longer than the retained window, showed no ring until the next turn.

The host now derives the newest context facts from its whole journal with
the same selector the chat uses, and returns them on agentSession.options
for sessions that write them. The chat prefers each fact its loaded rows
carry and takes the host answer for a fact they lack. When a live batch
revises a turn row older than the loaded window, the chat asks for options
again so that answer stays current.

* fix(native-chat): bound context refresh reads and refresh when the turn row is trimmed

Each turn-row revision the loaded window missed started its own options
read. Those reads share the session's host queue with sends and interrupts,
and each asks the CLI for its settings, so a burst could pile reads in front
of a user action and discard every answer before it landed. The chat now
keeps one options read in flight and at most one behind it.

A live turn longer than the retained window also lost its turn row to the
trim without asking for a fresh host answer, so the ring fell back to the
answer read at turn start until the next response. Trimming a turn row now
asks again, like a dropped revision does.

* fix(native-chat): show the context ring from the first response, sized from the session's model

A new session has no measured window until its first result, so the ring
stayed hidden for the whole first turn. The host now keeps the window the
applied model's name implies (1M for a [1m] name, unknown for default, 200k
otherwise) and writes it beside an estimate when the journal holds no window,
or after a model write, until the result or the CLI's report replaces it.

* fix(native-chat): imply a context window only from a [1m] model name

A bare model name does not fix the window: first-party runs today's opus,
sonnet and fable models natively at 1M while a gateway or cloud provider runs
them at 200k, and opusplan and haiku run another model in plan mode. Sizing
their first response at 200k read the ring about five times too full, so only
a [1m] name implies a window now; any other name waits for the result.

* fix(native-chat): size the first response from a report taken before any turn

A model picked in a chat with no turn yet asks the CLI for its context
report, but with no turn row the report's write lands nowhere. Recording it
still marked the journal as holding a window, so the first response wrote
none and the ring stayed hidden until the turn's result.

The report's window now serves as the fallback a response writes while the
journal holds no window, and recording a report no longer assumes its write
landed.

* test(native-chat): move the fake Claude connection out of the structured integration suite

The context-report delivery case pushed the suite past the 800-line limit,
failing repo-wide lint. The fake child now lives in its own fixture.
2026-09-24 00:12:08 -07:00
Jinwoo Hong ca75bc4c8d fix(orchestration): type a request ahead of pasted dispatch briefs so Claude workers follow them (#22582)
* fix(orchestration): type a request ahead of pasted dispatch briefs so Claude workers follow them

Claude Code wraps a bracketed paste in <pasted_content> and tells the model to
follow instructions inside it only where the user's own message asks. Orca sent
the whole dispatch brief as a bare paste, so Claude workers (Opus 5.5, Sonnet 5)
refused it as suspected prompt injection. Every dispatch path now types a short
lead line in the same PTY write as the paste frame, the preamble drops shouted
rules, and dispatch detection accepts the lead line and pasted_content wrapper.

Fixes STA-8200

* refactor(orchestration): tidy dispatch lead-line delivery after review

- Share one dispatchPreambleSendOptions() across the four dispatch paths.
- Fold every C0 control and DEL out of the typed lead line.
- Bound the <pasted_content> tag scan and let compaction return null for
  non-dispatch prompts, removing the separate detector.
- Restore the stay-off-other-channels rule in plain wording.
- Test through the real status normalizer and trim duplicated assertions.

Refs STA-8200

* test(orchestration): guard coordinator auto-dispatch lead line

- Capture send options in the coordinator runtime fake and assert the
  auto-dispatch send uses dispatchPreambleSendOptions.
- Drop the helper test that only restated its literal.
- Share DispatchPreambleSendOptions with the coordinator runtime contract.

Refs STA-8200

* docs(orchestration): fit the pasted-spec note inside the kernel line budget

Refs STA-8200

* docs(orchestration): drop the pasted-spec note from the coordinator guide

The typed lead line is the fix; the advisory note cost always-loaded context.

Refs STA-8200

* fix(orchestration): type the dispatch lead line only for Claude agents

Codex discards typed text that shares a PTY write with a bracketed paste,
so the lead line never reached it. Only Claude Code needs the lead to follow
a pasted brief, so known non-Claude agents now get the pre-lead bytes and
unidentified agents keep the lead in case they are Claude.

Refs STA-8200
2026-09-24 01:33:16 -04:00
Brennan Benson b4d732685c feat(agent-status): combine Codex child work through the shared main-agent status fold (#22475)
* feat(agent-status): combine Codex child work through the shared main-agent status fold

* docs(agent-status): correct two comments the waiting child-work arm made stale

A child failure reported in place as `blocked` now pins the row `waiting`, not
`working`; and no relay ever sent an unfolded `working` beside a waiting child.

* fix(agent-status): only a waiting child asks for a human

A child's `blocked` state means its task failed (the only producer maps a
failed background task to it, and the background-task view labels it
"failed"), not that a human must act. Folding it into the waiting arm would
surface a failed child as needs-you. It stays live work, as before this
series.

* docs(agent-status): say a waiting child, not a blocked one, makes the row wait

A child's blocked state means it failed; only its waiting state feeds the
waiting arm. Two fold comments, a test describe and two parity story names
still called the waiting child blocked.

* docs(agent-status): name where a child's wait is still lost, and pin the structured lane's real input

The doc said the Claude hook lane's rows match Codex and that every lane feeds a
child's wait into the fold. Neither holds: Claude keeps the wait in one slot the
next main agent event overwrites, the structured lane turns a child's prompt
into the main agent's own attention, and Codex drops its roster on a root Stop
when it tracks no child transcripts. The parity story now drives the structured
lane with the input it actually receives.
2026-09-23 22:09:07 -07:00
Brennan Benson 7a4f080086 revert: #18790 (orchestration incarnation reap fallback and bundled Freebuff agent) (#22601)
This reverts commit 0677271709.

#18790 was merged as one squash commit that carried two unrelated changes:
a process-incarnation fallback for reaping leaked orchestration worker
terminals, and an unannounced "Freebuff" third-party agent (catalog entry,
icon, locale strings, README rows). The Freebuff agent was never meant to
ship, so the whole PR is reverted; the reap fix should be re-submitted on
its own.

Until that re-land, a worker whose durable terminal handle goes stale is
again reported missing on release/stop instead of being re-found through
its process incarnation, so its terminal can leak on Remote Server.

The mobile session page closure pin moves 4218 -> 4219: the revert drops
the freebuff icon #22119 pinned (-1), and #22452 had already added two
src/shared modules without re-pinning (+2).
2026-09-23 21:39:28 -07:00
Brennan BensonandClaude 3ea15dd0a2 fix(native-chat): keep chats that failed to resume in the status bar and say what to do (#22448)
* fix(native-chat): keep chats that failed to resume in the status bar and say what to do

After a restart, a chat whose resume did not carry on was reported only by a
four-second toast that named nothing, and the status bar entry vanished because
the reattach had already spent the offer.

The host now files the outcome as a durable `failed` entry in the recovery
capsule, with the refusal code and the prompt the offer quoted, and returns it
from the restart-resume RPCs. The renderer shows a "N chats failed to resume"
status bar entry, a count-only toast with Show and Dismiss, and keeps the resume
dialog open with a status icon per row and a "To resume" line whose action is
chosen from the reason. A failure dies on dismiss, on a successful retry, on the
user's own send in that chat, or with the marker's 24h expiry.

* fix(native-chat): keep the resume dialog unchanged and add failed chats as rows

The failure view had replaced the resume dialog's title, checkboxes, preference
box, and footer. The dialog is back to main's layout. A chat an earlier resume
could not carry on is now an ordinary selectable row there, plus a status icon
whose tooltip carries the reason, a dismiss control, and a "To resume" line.
Selecting it and pressing Resume retries it; it is pre-selected only when a
retry can succeed. Resumed chats leave the list as before.

Also stubs the new failure listing on the cross-version wire host fixture, which
the restart-resume RPC now reaches.

* refactor(native-chat): release a failed-resume record where a send is admitted

Keeps the host file at main's size, and only releases the record for a send the
controller actually lets through.

* fix(native-chat): note a failed restart resume in the chat and derive when it is settled

The chat itself now says when Orca could not continue it after a restart,
with an error (refused) or warning (unconfirmed) status row, so the failure
survives the toast, a dismissed record, and another restart.

A recorded failure is current only while the chat's newest user message is
the one it had when the failure was filed. Listing re-derives that from the
journal and prunes superseded records, replacing the in-memory set and the
hook on every send.

Failures move to their own optional top-level key in the recovery file, so an
older build that rejects unknown entry states keeps reading its offers. A
retried failure stays a failure through a rollback or a lapsed lease instead
of returning as a pending offer, and the toast's Dismiss names only the chats
the host listed as failed.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): check the older reader against a filed failure before any retry

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): file a failed restart resume against the chat as its attempt ended

A failure was filed against the chat's newest user message read at settlement, after every chat in
the action had finished. The chat's note asks the user to send a message, and one sent while other
chats were still being continued became part of the filed state, so the failure stayed listed after
the user had done what it asked. Each chat's newest message is now observed as its own attempt ends.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): let a reply that made a chat ineligible retire its failure

When a restart resume reached a chat the user had already replied in, the
attempt was refused as no longer eligible, but the failure was filed against
that very reply. It then stayed listed as "finished on its own" until the user
sent yet another message. An ineligible chat is no longer observed at the
attempt, so its failure falls back to the reserved marker and the reply that
made it ineligible supersedes it at the next listing.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): offer Retry on a failed resume only when a retry would run

When the agent refused Orca's "continue" message with its own reason, the
failed row fell to the generic guidance, which offers Retry and pre-selects
the chat in the resume dialog. The refused message is already the chat's
newest user message, so a retry is never eligible: it did nothing and the
same "couldn't be resumed" toast came back.

The host now reports whether a retry would run, derived at list time from
the same predicate the retry applies to the failure's marker. Where it would
not, the row offers Open chat and Dismiss and is not pre-selected. An older
host omits the flag and the reason alone decides, as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): file only restart failures the user must act on

A chat that stopped being resumable between listing and acting (it
finished on its own or is waiting on the user) was filed as a failure
with no note in the chat. It now just spends the offer.

A failed reattach now writes the same in-chat note as a refused
continuation, so every filed failure explains itself in the chat.

An unconfirmed continuation's failure retires once the chat shows the
continuation's own message opened the newest turn.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): clear a failed resume from the status bar once its chat moves on

While a failed resume is listed, the restart store watches the host's
status feed; when a failed chat's status or latest prompt changes after
the list was read, it re-reads the host once. The host still decides
whether the failure stands. Nothing is watched while nothing failed.

The failure toast now counts only the requested chats the host still
lists as failed, keeping the old count for a host that sends no list.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that an unjournaled continuation never retires its unconfirmed failure

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop telling the user to send a message into a chat Orca couldn't reconnect

A failed reattach, or a continuation refused because another window or terminal owns the
session, now leaves a note saying Orca couldn't reconnect the chat instead of advising a send
that would meet the same refusal. The restart list keeps the reason-specific advice.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the failed-chat re-read from undoing an action or missing a reply

The status-bar re-read no longer runs while a resume or dismiss is in flight, and its answer is
dropped if one settled meanwhile, so a dismissed failure cannot come back. A change to a failed
chat already seen always triggers it, whatever the host's timestamp says. A chat whose unconfirmed
continuation the host already retired is now reported as resumed instead of saying nothing.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that no failed-chat re-read runs under a resume in flight

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): typecheck the unjournaled-continuation case against a nullable marker

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): drop the stale "no arguments" note on the restart-offer params

The dismiss call now names sessions, so the older comment contradicted the schema below it.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): count unconfirmed resumes apart from refused ones in the action toast

The post-action toast said "N chats couldn't be resumed" for chats whose
continuation may well have gone out, while the list and the chat itself say
Orca couldn't confirm it. That wording invites a duplicate "continue" send.
Unconfirmed chats now get their own count, classified by the outcome the
host filed, so the toast matches the row it points to.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop offering Resume on a failure the host says cannot retry

A failed row the host marks unretryable could still be ticked, sending a resume
that could only fail again; its checkbox is now disabled and it never joins the
action. An older host that omits the flag keeps today's selectable row.

The status bar no longer calls a chat "failed to resume" when the host only
couldn't confirm the resume, matching the dialog's own wording, and the mixed
toast's second line now says "other" so it cannot read as the same chat.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): skip unreadable failure records instead of rejecting the capsule

A failure record this build cannot parse (a newer outcome, say) made the
whole recovery file unreadable, so a downgraded build listed no restart
offers and could not record new teardowns. Failures are advisory: parse
each one on its own and drop what does not parse.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a failure record with an unreadable marker is skipped

The skip-unreadable-failure test only covered an unknown outcome, so going back to the throwing
marker parser for failure records still passed. A failure record usually outlives its offer
entry, so a newer marker shape can appear only there.

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-23 20:45:07 -07:00
Jinwoo Hong 8ba7f829ac feat(ipynb): run notebook cells in a persistent Jupyter kernel (#22581)
* feat(ipynb): run notebook cells in a persistent Jupyter kernel

Replaces the fresh-process runner (which silently re-ran every earlier cell)
with the user's own ipykernel, driven by a small bundled Python bridge over
line-framed JSON. One kernel per open notebook: started on first Run, shut
down when its tab closes or Orca exits (stdin EOF), and the kernel's own
parent poller reaps it if the bridge dies.

The header gains a kernel pill (workspace .venv/.conda recommended, PATH
interpreters, Browse), Interrupt/Restart/Run all/Clear all, and a one-time
missing-ipykernel dialog with Install. Outputs stream live per cell and are
written into the document when the run finishes.

* test(ipynb): cover the stalled-interrupt restart offer

* refactor(ipynb): disable the kernel pill while settling; merge its classes with cn

* refactor(ipynb): tie kernels to their renderer document and simplify the run flow

- Main keys kernels per renderer, so a reload, renderer crash or closed
  window shuts them down, and two windows never share one notebook kernel.
  One close-driven cleanup replaces the separate exit and start-failure
  deletions.
- The first run uses the nearest workspace env, else the first Python on
  PATH; the picker no longer opens itself, so its open state stays in the
  toolbar. Closing the picker brings the missing-ipykernel dialog back
  instead of dropping the queue, which also keeps a Browse pick's run.
- Discovery marks the kernel starting, so a second run during it queues
  instead of starting a second kernel, and a tab closed mid-discovery no
  longer leaks one.
- Running an nbformat 4.4 notebook gives its cells ids (upgrading to 4.5),
  so moving a cell mid-run cannot misroute its output.
- The death notice drops stderr from before the kernel was ready (the
  unencrypted-TCP warning).
- SSH and non-Python runs toast instead of writing a notice into the cell.
- Windows conda envs are named after their folder.

* fix(ipynb): install ipykernel into envs without pip

uv-created venvs ship without pip, so Install failed with 'No module named
pip' there. When pip is missing, bootstrap it with the stdlib's ensurepip
and retry. The install moves beside the other interpreter probes, and the
copyable install command comes from one helper.

* fix(ipynb): address PR review comments on stream errors, old jupyter_client and the Windows venv hint

- Swallow stdout/stderr stream errors on the bridge child, as spawnProcess
  requires, so a broken pipe cannot crash main.
- The bridge exits (reporting the death) even when cleanup_resources is
  missing (jupyter_client < 6.1.5) or raises.
- The install-failure hint suggests `py -m venv .venv` on Windows.

* feat(ipynb): add Cancel to the missing-ipykernel dialog

It does what Esc does: drops the cells waiting on the kernel.

* fix(ipynb): recover from a rejected kernel start; quote the install command per shell

- Discovery moves into start, so one catch turns a rejected
  listPythonEnvironments or startKernel into the usual failed start: the
  session returns to off with the error in the cell, instead of sticking
  at starting.
- The copyable ipykernel command quotes the interpreter only when its path
  has whitespace, prefixing PowerShell's call operator on Windows. Install
  itself still spawns without a shell.

* fix(ipynb): always shell-quote the copyable ipykernel install command

Quote the interpreter path for every path, not only ones with whitespace,
so paths with shell metacharacters like & copy as a working command.
Single quotes are literal in POSIX shells and PowerShell; embedded quotes
are escaped per shell, and PowerShell keeps its & call operator.
2026-09-23 23:38:33 -04:00
Neil 795b64b9a6 docs(tui-agent-config): correct the OpenCode readiness-budget rationale (#22593)
The comment merged with #22546 claimed ConPTY never forwards DECSET 2004 and
that the signal therefore cannot fire on Windows. Verification on two real
Windows hosts refuted that: the sequence arrives in order on both ConPTY
backends, and the readiness signal fired in every run.

The budget was the actual problem — opencode does not enable bracketed paste
until ~4.8s and its composer is not ready until ~10s, so the 8s default expired
first and the draft was pasted blind. Same fix, accurate reason.

Note the claim that seeded this: terminal-agent-paste-bracketing.ts says 2004
"can be lost by remote replay or ConPTY", which is careful and not contradicted
here; the absolutism was mine.
2026-09-23 20:29:00 -07:00
Brennan Benson 5e3effc32f fix(native-chat): show every user message on the message rail, not just loaded ones (#22558)
* fix(native-chat): show every user message on the message rail, not just loaded ones

The message rail was built only from transcript rows the renderer had
loaded, so any prompt above the loaded page had no tick, and the rail lost
ticks when a long live session trimmed its retained window.

The host now answers `agentSession.conversationOutline`: every user message
in a structured session's journal (item id, creation sequence, a preview
cut to 200 characters, image count) plus the journal position it is current
through. It is derived from the reduced journal on each request with the
same projection the transcript runs, so an entry's id and preview are what
the loaded row shows. The reply is bounded like a history page: previews
shorten, then drop, and only then do the oldest entries go.

The renderer asks only while the pane is visible and older history is
unloaded, uses outline entries only for messages older than its loaded
window (the window is authoritative for the rest), and falls back to loaded
messages while the outline is stale (epoch change, or the window trimmed
past what it covers). Selecting a tick with no row pages older history in
until the row exists, then uses the existing rail jump.

The method is negotiated with `agent-session.conversation-outline.v1`; a
client never calls a host that does not advertise it, and any failure
leaves the rail on loaded messages.

* fix(native-chat): keep a rail jump from the bottom from re-arming follow and cancelling itself

A rail jump started by a reader following the end stopped a few pixels
above the bottom instead of reaching the message. Paging older history in
for a jump always leaves the reader following at the very end, so jumps to
unloaded messages hit it every time; a jump to a loaded message from the
end did too.

The jump scrolls smoothly, and only its landing is marked as the
application's own scroll. Its first frames move a pixel or two, still
inside the band where a reader event re-arms follow, so the list read the
jump leaving the end as the reader arriving at it. The next frame, just
outside the band, then read as the reader taking over and rebased the view
with an instant write, which cancels the smooth scroll.

Re-arming follow now needs the reader to be arriving at the end: an
unmarked event that moved the view up never reattaches a detached reader.

* test(native-chat): check a rail jump left the end before reading where it landed

* fix(native-chat): keep the rail's message list still while a press selects an item

With the whole conversation in the rail, its hover list overflows and opens
scrolled to the message being read. Clicking an older item did nothing:
pressing it focuses it, which turns the hover preview interactive, and that
switch re-ran the effect that scrolls the lit row into view and focuses it.
The list moved under the pointer between press and release, so the click
landed on the list instead of the item, and focus jumped to the lit row.

Revealing the lit row now follows the list opening (and its rows shifting),
not the switch between hover and interactive. Entering interactive moves
focus into the list only when focus is not already on one of its items.

* perf(native-chat): keep the rail's outline entries stable while a trimmed window slides

A long live session holds a head-trimmed window, so every new row moved the
oldest-loaded edge and rebuilt the outline view even when no user message
crossed it. The rail then re-merged, re-rendered and re-read the scroll
geometry on each new row. The view is now reused while the set of entries
older than the edge is unchanged.

* fix(native-chat): let a rail jump wait out an older page already loading

Scrolling to the top of the loaded window asks for the next older page. A rail
jump made while that page was in flight asked again, got the lane's immediate
no-op return, read it as a page with no progress, and dropped the click. The
jump now waits for the in-flight page to land before deciding.

* fix(native-chat): reattach follow when content shrinking clamps a reader onto the end

The rule that stops a smooth scroll leaving the end from re-arming follow
compared offsets, so it also refused a reader whose offset dropped because
settled content folded away beneath them and the browser clamped them onto
the end. They sat at the bottom without following, and the next reply grew
out of view. Re-arming now requires closing on the end rather than moving
down, which still rejects a scroll leaving it.

* fix(native-chat): keep the rail hooks' ref writes out of render

Both hooks wrote a ref while rendering, which React may replay or discard.
The rail's structural-sharing baseline is now recorded after commit, and the
history jump calls the lane's page loader from its effect instead of through
a render-updated ref.

* fix(native-chat): let the latest rail pick win over a jump still paging

A jump to an unloaded message keeps paging older history until it lands. A
pick made meanwhile lost to it: a loaded message scrolled into view, then the
earlier jump finished and pulled the reader away; another unloaded message was
ignored. Picking a loaded message now cancels the paging jump, and picking an
unloaded one retargets it without asking for a second page.

* fix(native-chat): step a rail history jump with a functional update

The step that requests the next page wrote the pending jump from the value
its effect closed over, so a pick or cancel queued since that commit would
be overwritten.

* refactor(native-chat): run a rail history jump as one abortable awaited loop

The jump through unloaded history was an effect-driven state machine that
guessed "no progress" from the message list's identity and could not be
cancelled by anything but another rail pick. A diff reveal, "Jump to latest"
or the reader scrolling left it paging, and when its page landed it pulled
the reader away; a history read that kept failing during a live turn could
repeat back to back.

Loading an older page now reports how it ended, and a second request while
a page is in flight joins it instead of being refused. The jump is an
awaited loop that reads the rail from a commit after each page, stops on
anything but a page that moved the window, and is aborted by any other
navigation, reader input (wheel, touch, scroll keys, scrollbar), a session
switch or unmount. The latest pick wins.

* fix(native-chat): derive the rail outline from the transcript's own projection

The host built the outline from user items alone, while the transcript
orders every message by when it was observed, folds tool results into the
turn above and then drops harness turns. An imported user row carrying a
tool result beside harness text therefore got a rail tick previewing the
harness text, and clicking it paged history for a row that never draws.

The transcript's order-fold-strip projection now lives in one shared
function. The renderer's list projection wraps it with its own tail-row
order, and the host runs it over the whole journal and keeps the user rows
that draw content, so the outline lists the same messages in the same order.

* fix(native-chat): retry a failed rail outline read a few times

A failed outline read left the rail on loaded messages until a new gap
opened, the pane was shown again or the epoch changed. The client now
rejects a failed read (a host without the outline still resolves to
nothing, without being called), and the rail retries up to three times
with doubling backoff.

* test(native-chat): give the rendered-transcript fixture the older-page result contract

* perf(native-chat): sort the shared transcript projection without a spread copy

* fix(native-chat): let a wheel over the rail cancel a rail jump still paging

The rail forwards its wheel to the transcript, so a reader scrolling there is
scrolling the transcript. That wheel never reached the scroller's reader-input
handlers, so the jump kept paging and later pulled the reader to its target.

* fix(native-chat): keep the rail's list open while a picked message pages in

Picking a message that is not loaded yet can take several pages of older
history. The list closed on the pick, so its busy item was never seen and
the click looked ignored. The list now stays open with that item pulsing
until the jump lands or is abandoned, and stops revealing the lit row
meanwhile so the rows do not move under the pointer.

* fix(native-chat): keep the shared transcript projection loadable on mobile

The projection moved to src/shared, which mobile's Hermes engine also loads,
and switching its sort to toSorted broke the Hermes compatibility guard.
Sort a copy made with Array.from instead, as other shared code does.
2026-09-23 20:26:54 -07:00
5802b54579 fix(rate-limits): read OpenCode Go usage with the Go API key (#22551)
* fix(rate-limits): read OpenCode Go usage with the account API key

Since OpenCode's console migration (upstream fe51b0b19a, "fix(console):
restrict legacy access to Black"), an account with no Black subscription
is redirected from the legacy console to /console/login, so Orca's
cookie-based workspace lookup returns nothing and the Go bar stays empty.

Fetch usage from GET https://opencode.ai/zen/go/v1/usage instead, which
authenticates with `Authorization: Bearer <key>` and needs no console
session. The key resolves in order: Orca settings override,
OPENCODE_API_KEY, then whatever OpenCode itself stored on /connect --
auth.json for 1.x, the credential table for 2.x. The cookie path stays
as the fallback so Black/legacy accounts keep working.

A 403 EntitlementError now reads as "no OpenCode Go subscription" in the
status bar instead of a generic refresh failure (#22257's reporter was
misled by exactly that).

* fix(rate-limits): prefer OpenCode's stored Go key over OPENCODE_API_KEY

OpenCode applies the key saved on /connect after the environment, so the
stored key is the one its own Go requests use. OPENCODE_API_KEY is also
the Zen provider's variable, so ranking it first could read a key that
OpenCode itself is not using for Go.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(rate-limits): name the API key when OpenCode Go usage lands on sign-in

A redirected usage request arrives as a 200 sign-in page because Electron
follows redirects; report it as a rejected key instead of a parse failure.
The cookie path's empty workspace lookup is what non-Black accounts now
hit after the console migration, so its message points at the API key
rather than only the workspace override.

Co-Authored-By: Claude <noreply@anthropic.com>

* chore(i18n): add the OpenCode Go API key strings to the English catalog

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(rate-limits): stop calling the credential table an OpenCode 2 marker

Verified on two real Windows hosts running OpenCode 1.18.16: the `credential`
table exists there too (empty, same columns), so its presence does not identify
a 2.x install. Neither host had an `auth.json` at all.

The resolution already probes both stores on every version, so only the comments
were wrong. Says so now, and records that a 2.x install which never ran the
legacy import has no `auth.json` either — which is why both tiers exist.

* refactor(shared): move GhosttyImportPreview out of global-settings-types

Adding `opencodeGoApiKey` pushed global-settings-types.ts one line past the
300-line ceiling, failing `oxlint` in CI. AGENTS.md forbids a max-lines
suppression, so split instead: the Ghostty import preview is a distinct concern
that never belonged in the settings-shape file.

Re-exported from the original module so no importer changes. 293 code lines now.

---------

Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
2026-09-23 20:09:02 -07:00
Neil b0ae7d18a0 fix(opencode2): resolve subagent session lineage so child work stops taking over the pane (#22444)
OpenCode 2's plugin adapter unwraps a single-property `{ data }` success schema,
so `ctx.session.get` resolves to the bare session record. The shared lineage
lookup only accepts `result?.data?.id === sessionID`, and OpenCode 2 has no
`session.list` fallback, so `resolveRootSessionID` returned null for every
session and `childState` was permanently null.

With unknown lineage `canFailOpen` is true for attention events, so a subagent's
`permission.asked`/`question.asked` fell through and pinned an un-evictable
blocker keyed to the child's own session id — publishing a subagent as if it
were a root. Observed in hook posts: SessionBusy for a child session id whose
`session_v2` row carries a parent.

Envelope the result in the OC2 client shim so the shared lineage module works
unchanged; OpenCode 1 already receives enveloped results and is untouched.

Also adds `opencode2` to the double-Escape interrupt list, extracted into one
shared helper so the server inference and renderer gate cannot drift. A single
Escape was inferring an interrupt, and Escape is how the Subagents dock closes.

7 of 11 new lineage tests fail without the shim.
2026-09-23 20:08:01 -07:00
Neil b7a4fee700 fix(agents): stop claiming an unconfirmed OpenCode handoff succeeded (#22546)
* fix(opencode): stop claiming a handoff prompt was delivered when it was written blind

"Continue in New Session…" to OpenCode reported success even when the prompt
never reached the TUI (#22479). The paste-after-ready helper falls back to a
blind write when the composer-ready signal never arrives and only the agent
process is known to exist; that write was indistinguishable from a real
delivery, so the continuation showed its success toast.

- pasteDraftWhenAgentReady / pasteDraftToAgentPtyWhenReady report the blind
  fallback via onUnconfirmedDelivery, plumbed to launchAgentInNewTab as
  onPromptDeliveryUnconfirmed.
- The session continuation hedges instead of claiming success, and both the
  failure and the hedged notice offer "Copy prompt".
- OpenCode gets Codex's 20s composer budget. Both are quiet-window-less
  signals anchored on DECSET 2004, which ConPTY never forwards, so on Windows
  that budget is the settle delay before the blind paste.

* test(runtime): retarget the 8s startup budget test off OpenCode

The main-runtime startup-draft budget test used `opencode` as its stand-in for
"an agent without an override", which this branch invalidates by giving OpenCode
20s. It failed with "expected vi.fn() to not be called at all, but actually been
called 1 times" — the readiness signal now legitimately arrives inside budget.

Point it at `claude`, which still takes the 8s default, and add a companion
pinning OpenCode's 20s: a readiness signal at t+10s, past the old default, must
now deliver the draft. Removing the override makes that companion fail.
2026-09-23 19:47:47 -07:00
Brennan Benson 80f5aae0f9 feat(agent-status): publish the main agent's own state beside the combined row state (#22452)
* feat(agent-status): publish the lead agent's own state beside the combined row state

Every status producer folded the main agent's state together with live child
work into one `state`, so a lead that had finished while a subagent still ran
read `working` and its own state was lost. The row now also carries
`lead: { state, outcome?, stateStartedAt }`, admitted by the one payload
normalizer on the relay wire, IPC and disk, and published from the Claude hook
lane, the structured host ingest and renderer bridge, Grok (now on the shared
fold) and Codex (own combine kept). The persisted child-only boundary flag is
derived from `lead` plus child evidence and no longer written; old rows map
onto `lead` at hydrate. Combined `state` and `workingMode` are unchanged for
every reader; a cross-lane parity table pins that, with the cancelled-turn
watch-loop story recorded as a known divergence.

* fix(agent-status): make Orca's inferred interrupt the primary source of a Claude lead cancellation

Current Claude Code sends no hook at all on a cancel and no is_interrupt on
Stop, so the cancellation enters the lead record from the server's inferred
interrupt and rides into the next real Stop; is_interrupt on a turn boundary
stays as the secondary source for builds that send it. Comments, the store
reference and the parity table say so; no suppression changes.

* docs(agent-status): the child-only boundary comment now describes the persisted shell fact

The old sentence said a hydrated row no longer carries the shell fact, which is
the opposite of the mechanism: claudeRunningNonAgentTask is persisted precisely
so hydration can read it, and only a pre-lead row lacks it — reading as
shell-free, the same assertion its legacy flag made at write time.

* rename the lead fact to mainAgent: the main agent's own state

* docs(agent-status): the inferred cancel comes from Ctrl+C, not Esc

* fix(agent-status): an inferred interrupt keeps an already settled main agent, and the row verdict docs name its inferred source

* fix(agent-status): a child-induced wait publishes the main agent state it displaced

* fix(agent-status): decide child-held Claude rows from the saved main agent fact

Restart seeds the Claude main agent from the row's saved mainAgent whenever it
settled and no shell held the row, instead of re-deriving a child-only shape.
OSC cannot settle or repaint a row child agents hold open, including a row
waiting on a child's permission prompt. A sticky child permission prompt still
records the main agent's own progress, and OSC repaints and inferred answers
keep the shell fact beside the main agent they preserve.

* fix(agent-status): keep a finished turn's main agent verdict and clock with that turn

A Claude SessionStart restarts the main agent's clock instead of inheriting the
previous session's last Stop. A Grok idle prompt or session end, and a late
Codex root Stop after an inferred cancel, restate the same finished turn, so
they keep its recorded verdict; only a new turn clears it.

* test(agent-status): publish the Grok verdict restatement past the late-event window

* docs(agent-status): describe hydrate seeding and the OSC refusal from the saved main agent fact

* fix(agent-status): push a held child permission row when its main agent changes

* fix(agent-status): keep the shell fact on a held child permission row so restart does not settle it

* docs(agent-status): note the held child permission row carries the shell fact and is pushed

* fix(agent-status): pair the Claude shell fact with the main agent at the one row-build point

Every non-hook rewrite (terminal-title repaint, inferred answer, held child
permission) had to re-carry the shell fact beside `mainAgent`, and each one that
forgot let a restart settle a row while a shell still ran. The row builder now
pairs the fact once: a listener event restates it, any other write keeps it only
while `mainAgent` is unchanged. Restart seeds a settled main agent only when the
row says no shell ran, and legacy child-only rows map to that explicitly.

A held child permission now also accepts the main agent event's background
evidence, as it already accepts its `mainAgent`, so the child's drain no longer
settles a row a shell still holds. The renderer keeps a previous `mainAgent`
only for writers that never carry one, so a hook row without it matches the
host snapshot.

* test(agent-status): pin that restart never seeds a main agent from a row silent about its shell

* docs(agent-status): the row builder pairs the shell fact with the main agent, and restart seeds only on an explicit no-shell

* test(agent-status): name the legacy-row case parameter for what it holds

* docs(agent-status): name which rows carry the main agent fact
2026-09-23 17:45:50 -07:00
Brennan Benson 800d33e5c9 feat: name runtime machines (#22094)
* feat: name runtime machines

* fix: preserve pairing address optionality

* fix(cli): keep host and environment listings local

Listing paired servers read each one's machine name by dialing it, so both listings made a network
round trip per server and waited out a timeout on any that were offline. They answer from this
machine's own pairing store; `orca host name --environment <name>` reads one server's name.

* fix(settings): caption the machine name paired devices actually receive

The caption read the runtime's published name once, when the pane opened, so saving an override
left it naming the old computer while phones already showed the new one. It now re-reads whenever
the saved override changes; the settings write lands in the main process before the store publishes
it, so that read already sees the new name. The name is interpolated rather than baked into the
fallback, and the caption, label and placeholder are in the English catalog.

* refactor(settings): normalize the machine name in one place

The trim and length rules for `machineName` were spelled out separately at the
renderer IPC (trim + 255), the settings load path (trim only, no cap), the RPC
schema (zod trim + 255) and the runtime reader (trim). A hand-edited or legacy
profile could therefore load a longer name than any writer accepts.

`src/shared/machine-name.ts` now owns `MACHINE_NAME_MAX_LENGTH` and
`normalizeMachineName`, and every writer and the load path use it. The RPC
schema keeps rejecting over-long names but derives its cap from the constant,
and the runtime settings controller normalizes an RPC write before storing it.

* fix(runtime): detect the machine name once and label handoffs with it

Every runtime constructed in a process (the app, plus each one a test builds)
ran its own `scutil` lookup. The friendly name is a property of the host, so the
lookup is now a single shared promise; construction still never blocks on it,
and a rejected lookup can no longer surface as an unhandled rejection.

The structured-chat handoff banner ("Agent is open in terminal on X") named this
host with the bare `os.hostname()` while paired devices saw the published name.
The transport now reads the same `RuntimeMachineName`, through a getter so a
rename in Settings is reflected without rebuilding the transport.

* fix(cli): print the name the runtime publishes and keep its envelope

`orca host name --name X` printed `undefined`: `settings.update` replies with
`{ settings }`, but the handler read a bare `machineName` off the reply, and the
test fixture mirrored the wrong shape so it passed. After a write the command
now re-reads `status.get` and prints what the runtime publishes, so a blank
`--name` prints the detected name it returned to rather than an empty string.

The read path wrapped a possibly routed answer in a local envelope, stamping
`_meta.runtimeId: "local"` on a reply from another server. It now returns the
`status.get` envelope itself, and an unreachable runtime is reported as the
usual error instead of an invented "unknown" name.

`environment list` had gained machine-name and platform columns that no caller
populated, so every row printed "platform unknown"; the columns are removed.

* refactor(settings): give the machine name field its own component

The caption under the field re-read runtime status every time the saved value
changed, relying on a comment about write ordering to show the new name. A saved
override already is what paired devices see, so the hook now derives the caption
from it and asks the runtime only for the detected name; a stale status read can
no longer show the previous name.

`MobileMachineNameField` owns the store read, the published-name hook and the
debounced input, so `MobilePairingSetupSection` returns to its prop shape and
the pass-through `MobilePanePairingOutput` wrapper is gone. Paired-device
revocation moves into `useMobilePairedDeviceRevocation`, which keeps
`MobilePane` within its line budget with an extraction that carries behavior.

The web client mounts this pane too, but its settings store kept the name
locally where nothing published it. `machineName` now rides the existing
runtime-backed settings sync so the field renames the paired runtime.

* refactor(settings): normalize the machine name at the store boundary

Every writer (desktop IPC, web RPC, CLI) reaches the store through
updateSettings, which already normalizes the other free-text settings
there. Trim and bound the machine name in that one place instead of at
two upstream edges, so a future main-process writer is covered too.

* test(settings): pin machine-name routing and detection, and make the field searchable

The shared machine-name lookup test spawned the real `scutil` twice and compared the answers, so a
slow runner could time one spawn out to the hostname and fail. It now mocks the subprocess, proves
the hostname answers until the one shared lookup lands, and that a second runtime does not spawn
again.

`host name` is no longer pinned local, but only the explicit `--environment` route was covered; an
ambient `ORCA_ENVIRONMENT` now has its own test so the pin cannot silently grow back.

The Machine name field is added to the Mobile pane's search catalog at the tail, keeping every
existing row's tie-break index.

* fix(runtime): wait for the machine-name lookup before publishing status

A status read answered in the first few milliseconds after launch published the bare
hostname because the friendly-name lookup had not landed yet, and a caption fetched in
that window never corrected itself. RuntimeMachineName now exposes the settled lookup
as a promise, and both status publishers (the status.get RPC and the desktop
runtime:getStatus IPC) await it before reading. Construction, listen, and every other
method stay unblocked; the worst case is one wait of at most a second on the first read.

* fix(cli): refuse to rename a runtime that does not publish a machine name

An older Orca runtime rejects the unknown settings field with a bare invalid_params, so
'orca host name --name' routed at one failed with no explanation. The runtime that does
not publish machineName on status cannot store one either, so the CLI reads status first
and refuses with incompatible_runtime and a message that says to update that host,
before writing anything.

* fix(ipc): introduce this desktop to remote hosts by its machine name

When this desktop connected to a remote workspace host it announced itself under a
hostname captured once at module load, so a renamed machine kept its old name on every
other device's connected-clients list. The client name is now read at send time from
the runtime's machine name (the configured override, else the detected one), passed in
where the remote workspace handlers are registered, so a rename reaches the next
presence frame without a relaunch.

* fix(runtime): keep the machine-name lookup under the status probe budget

Status publishers now wait for the one-time name lookup, and `orca status`
probes them with a one-second budget. scutil answers in milliseconds, so a
half-second cap keeps a stalled lookup from making a healthy runtime read as
"starting" while still preferring the friendly name.

* refactor(web): drop the unreachable machine-name write path

The Mobile settings section is desktop-only, so the paired web client can
never render the field. Forwarding the name through the web settings sync was
dead code, and against an older host the strict update contract would have
rejected it while the local mirror kept the value. Remove it until a web
surface exists.

* chore(i18n): translate the machine-name strings and document paired-server rows

Add the Machine name field and its Settings search entry to the five non-English
catalogs, explain in the host list spec why paired-server rows report an unknown
platform, and drop a stale timeout figure from a test comment.

* refactor(settings): make the machine name a machine-wide setting with a General home

The name other devices and hosts list this computer under is not a mobile
setting. Rename MobileMachineNameField to MachineNameField, give it a per-mount
id, and put its primary home in Settings > General under "This computer". The
Mobile pane keeps the same field. One shared search entry feeds General, the
Mobile pane, and the copy now says "other devices and hosts" in all six locales.

The web client has no machine of its own to name and its settings mirror cannot
persist one, so the field renders nothing there and General omits the section.

* feat(mobile): name this computer in the Orca Mobile pairing step

The "Pair this computer" step now shows the same machine name field above the
connection choice and code, so a user pairing a phone from the sidebar page can
name the computer right there.

* feat(settings): name this host when sharing it with other devices

Share this host produces the access link other devices use to reach this
machine, so it mounts the machine name field first. The pane's search entry
takes the shared machine-name keywords so a search lands there.

* feat(sidebar): name this desktop when adding a remote host

This desktop introduces itself to a new SSH host or remote server under its
machine name, so the Add Remote Host dialog mounts the field once, between the
header and the host fields, in both modes. Submit logic is unchanged.

* feat(settings): name this computer in the SSH pane add form

The SSH pane's add form mounts the machine name field above the host fields.
Editing a saved host leaves it out; that host already met this computer.

* fix(mobile): drop the empty machine-name grid row on the web client

The pairing step wrapped MachineNameField in its own grid-area div. On the
web client the field renders nothing, so the wrapper left an empty row and
an extra row gap between the copy and the connection options. The field now
takes a className for its root, so the grid slot disappears with it.

* fix(settings): let Enter in the machine name field submit its form like sibling inputs

The field intercepted Enter to blur and commit instead of submitting the enclosing
SSH add form. The draft is already flushed on blur and on unmount, and the name is
read from the store whenever a peer asks, so nothing is lost when the form submits
first. Enter now behaves like the neighbouring inputs; the test proves the submit
fires and the name still commits when the form closes.

* fix(mobile): keep the machine name inside the pairing copy cell

A dedicated grid row stayed in the template on the web client, where the field
renders nothing, adding an empty track and a second row gap between the copy and
the connection options. The field now sits at the end of the copy cell with the
same 18px rhythm, so an absent field leaves nothing behind.

* fix(runtime): retry a failed machine-name lookup instead of latching the hostname

On a loaded Mac the scutil lookup missed its 500 ms cap during app boot, and
because the fallback was memoized for the process, every status read and the
Settings caption showed the bare hostname for the rest of the session.

The lookup now gets a 5 s timeout, a failed attempt (timeout, spawn error,
non-zero exit, empty output) clears the shared memo so a later ready() retries
after a 30 s interval, and status publishers wait only up to a 750 ms publish
budget before answering with what read() has now. A friendly name and the
non-darwin hostname stay final.

* refactor(settings): show the machine name only where other devices join this computer

The Add Remote Host dialog, the SSH pane add form, and General all describe
another machine, so a field about this computer's own name read as a third
kind of label there. The field now mounts only where other devices pair with
or connect to this computer: the Mobile pane, the Orca Mobile pairing step,
and Remote Servers > Share this host.
2026-09-23 17:30:14 -07:00
Brennan Benson 98e5ea3d5f refactor(tabs): delete the terminal tab's dead adopted-session field (#22557)
* refactor(tabs): delete the terminal tab's dead adopted-session field and every branch that read it

* refactor(tabs): drop the stale adopted-agent comment and pin legacy load on a chat terminal

The comment above the terminal chat-eligibility agent fallback described the
removed adopted-session agent fallback. The legacy load test now puts the
retired key on a chat-mode terminal tab, the only shape that ever carried it.
2026-09-23 17:29:11 -07:00
Brennan Benson 069dc8a1d8 feat(agent-launch): let a caller reserve the chat session, and start terminal launches with the session picks (#22523)
* feat(agent-launch): let a caller reserve the chat session and carry session picks to a terminal launch

* fix(agent-launch): keep a caller-minted session id named for its agent, and mint the fallback the same way

* test(mobile): model the older host from the launch fields, not the refined schema

* docs(agent-launch): describe the reserved session id as conversation identity, not placement

The caller mints the session id so it knows which conversation it
started; tab placement is not keyed on it. Also puts the terminal
surface's doc comment back on createTerminalSurface.

* docs(agent-launch): say a terminal launch reads the session picks on the wire contract

The `sessionOptions` field doc still said a terminal launch ignores them, which this branch changed.

* fix(agent-launch): check a reserved session id's token after the agent name, not the whole id

A hyphenated agent name failed the one-token check, so any session id for such an agent was
refused at the wire, while every other agent without a chat has its id ignored on the terminal.
2026-09-23 16:45:28 -07:00
Brennan Benson 60bd1dfdea feat(native-chat): one shell-environment setting for every structured chat (#22387)
* feat(native-chat): one shell-environment setting for every structured chat

Structured Codex chats started from the login-shell environment, while
structured Claude chats started from Orca's own process environment, so a
variable exported in .zshrc reached one and not the other. Both now start
from the same base, chosen by a new setting:

- on (default): the whole login-shell environment, as a terminal gets
- off: Orca's environment plus PATH, locale, SSH_AUTH_SOCK, and the
  variable names the user lists

The setting is re-read each time a chat starts or resumes. It is shown
only when Chat UI, the Chat UI default view, and structured native chat
are all on. Terminal-backed chat is unchanged.

* fix(native-chat): normalize the shell-environment settings when a profile loads

A hand-edited settings file could store the variable list as something other
than an array, and the structured runtime called `.filter` on it per launch, so
a malformed value failed every structured chat create and resume, and the
settings pane render. Normalize both keys where the profile loads, the same way
the other array settings are, through one shared normalizer the runtime policy
also uses. Also pin that an uncommitted name draft survives an unrelated
settings re-render.

* fix(native-chat): keep the pinned account as the only source of a structured chat's Claude home

The session record owns which Claude home a structured chat uses, and the
acquisition pin (claudeConfigDirEnvPatch) is the only emitter of
CLAUDE_CONFIG_DIR, compared against what the child would otherwise inherit.
With the login-shell snapshot as the inherited base, a CLAUDE_CONFIG_DIR
exported only in a shell rc flipped that comparison and produced an explicit
pin to the CLI default home, which moves the CLI off its default Keychain item.

Drop the inherited CLAUDE_CONFIG_DIR in the Claude launch resolver before the
pin runs, as Codex already does for an inherited CODEX_HOME. A configured
per-agent overlay still passes through, since the record already honors it.

* fix(native-chat): drop Orca's own CLAUDE_CONFIG_DIR from a structured Claude child too

The process spawner merges Orca's process env under the launch env, so a
CLAUDE_CONFIG_DIR exported to Orca itself reached the child around the launch
resolver's drop and unseen by the account pin. One helper now strips it from
both inherited bases. Also declare the two shell-environment settings on the
runtime store contract and add the six new strings to every locale catalog.

* feat(native-chat): add shell variables one at a time with a removable list

* fix(native-chat): return focus to the name input after removing a shell variable

* fix(native-chat): use a neutral placeholder for the shell variable input

The empty input showed a grey HTTPS_PROXY as its placeholder, which reads as a
saved value, especially right after that exact entry is removed from the list.
Use "Variable name" instead, in every locale catalog.
2026-09-23 14:29:20 -07:00
Brennan Benson 845db9e5e2 fix(native-chat): underline only file links a click can act on (#22370)
* fix(native-chat): underline only file links a click can act on

A chat message could underline a bare file name such as `deck.md` that
resolved nowhere, and clicking it did nothing, so it read as a broken link.

- Inline code and quoted text become file links only when they name a path
  (contain a `/` or `\`), matching plain prose; a bare file name stays plain code.
- Every file link click now answers: it opens, or says the file was not found,
  that the host could not be checked, or that the path could not be resolved.
- Explicit links like [x](README.md:5) route as files, and linked text keeps
  `#`, `?` and `%XX` literally instead of re-parsing them as URL syntax.

* fix(native-chat): wrap the parsed file location so file URIs in chat text still open

Linkified prose, quoted text and inline code wrapped their display text, which the
literal wrapped-href route no longer URL-parses, so file:///... resolved as a relative
path under the worktree. Wrap pathText[:line[:col]] from the parsed link instead.
2026-09-23 13:29:24 -07:00
Brennan Benson 641a7f36d9 fix(native-chat): keep one live tool-run header from a call's start to the turn's end (#22432)
* fix(native-chat): keep one live tool-run header from a call's start to the turn's end

The collapsed tool run's header was two elements, one for "a call is running"
and one for "nothing is", chosen call by call. Every call start and end
remounted it, the count disappeared while a call ran and came back one
higher, and a call that finished inside a frame still bought the whole swap.
That is the 42→43 flicker in the report.

The header is now one element whose live state belongs to the turn, not to
any call: it stays live from the run's first call until the agent moves past
it (prose, a further run, or the turn's end), and settles in place. While
live the sentence speaks in the present tense and counts the call in flight
("Running 3 commands"), with the latest call's command beside it as a muted
preview; once settled it reads as before ("Ran 3 commands ✓"). The category
glyph is the run's in both states, and the completion mark only appears once
settled, so nothing pops between calls.

Which run is live is derived where the transcript is sliced into rows: the
last row that speaks or acts is the trailing one. A reasoning aside after it
leaves it live; an answer or a further run settles it.

Present-tense forms for the ten sentence categories are added to the shared
copy and the English catalog. The transcript-file lane, which renders with
the structured activity UI off, is unchanged.

* fix(native-chat): settle a run blocked on the reader, keep it live past an approval

- A run whose question is awaiting the reader's answer no longer pulses
  "Reading 1 file" while the agent is blocked; it falls back to its calls.
- An approval's receipt no longer moves past the run above it, so the call
  it just approved reads as running while it runs.
- The header button is the live region, so the count is announced too.
- Drop the unused live option and record from the shared English sentence;
  nothing renders it yet.

* fix(native-chat): stop the settled run's check from fading in on every mount

Windowing remounts settled rows as the reader scrolls, and a restored transcript
mounts them all at once, so the fade replayed where nothing had changed. Also
pin that the live header counts the next call on the same element.
2026-09-23 10:49:04 -07:00
Brennan Benson 563dd5487f feat(native-chat): show a Codex chat's goal above the composer, and set it from goal mode (#22377)
* feat(native-chat): show a Codex chat's goal above the composer and set it from goal mode

Structured Codex chat now treats the thread goal as session state: a banner above the
composer shows the current goal (pursuing / paused) with clear, pause/resume and expand;
/goal enters a goal mode whose send calls thread/goal/set; the objective is journaled as a
user message marked as sent as a goal. The banner is derived from the journaled goal rows,
which Codex's resume snapshot refreshes, so a reopened or adopted chat shows its goal.

Fixes STA-8159

* fix(native-chat): replace a recorded goal by clearing first, and recover a lost goal-change response

- A set while the journal records a goal (any status) clears it before setting,
  so the new goal starts with its own time and token counters instead of
  rewriting the old goal's objective in place.
- The threadGoal plan answers an unknown outcome from the goal the journal
  records and reruns otherwise, so one request timeout no longer refuses every
  later Clear/Pause/Resume as unknown for the mounted session.
- The goal-mode chip says "Exit goal mode"; "Clear goal" stays the banner's
  action on the provider goal.
- A typed bare /goal on Enter enters goal mode, the same as picking it.
- The renderer reads the goal off the tail of its ordered snapshot; the host's
  unordered map keeps the by-sequence reader.
- Drop the composer's duplicate in-flight guard; the goal controller already
  serializes changes.
- Pin that a counter-only revision reaches a subscriber's live page under its
  original sequence.

* fix(native-chat): keep a bare /goal inside goal mode as the entrance, and pin goal delivery and serialization

- A bare `/goal` submitted while already in goal mode re-enters the mode instead
  of setting a goal whose objective is the literal text "/goal".
- The counter-only revision pin now drives the host's own event sink bound to a
  real journal, so it goes red when the publish after a lifecycle transition is
  dropped; the previous fake sink never published.
- Pin that a set which threw after journaling its objective puts that objective
  back exactly once when the ledger reruns the same operation id.
- Cover the goal controller hook: absent without host support, the loaded window
  wins over the host's answer, a second change while one is unsettled answers
  false without a request, and a refused change frees the next one.

* fix(native-chat): resume a blocked or usage-limited goal, and keep goal-mode drafts honest

- The goal bar offers Resume on a blocked or usage-limited goal, which the
  provider resumes exactly as it resumes a paused one; a goal whose token budget
  is spent still offers only Clear. The rule lives beside the other goal facts
  in shared code so every reader answers it the same way.
- A `/goal <text>` typed inside goal mode sets the objective `<text>`, as it
  does outside goal mode, instead of a goal whose objective is the literal
  command.
- Setting a goal is a host round trip; a draft edited while it was in flight is
  no longer wiped when the goal lands, matching every other host command.
- Pin that a lost status-change response is read as applied only when the
  recorded goal is in that status, that a cleared row in the loaded window
  outranks the host's earlier answer, and that the PTY lane is untouched.

* fix(native-chat): keep the load-older anchor on the loaded window when a live revision lands below it

A live revision of a row keeps that row's original sequence. When the row is
older than the client's loaded window, the shared reducer merged it in and it
became the load-older anchor, so paging `before` it skipped every row between.
A goal's counter-only revisions during a long goal turn reach any client that
attached after the goal row left its window, so a reopened chat lost rows on
scroll-back.

The reducer now admits live rows only at or above the window's oldest row
while older rows remain on the host; the journal keeps the revision and the
page reader serves it once the window reaches the row. With nothing older on
the host the window is the whole journal, so a row below the head is admitted
as before.

Also drain accepted provider events before a goal set reads the journal to
decide whether it replaces a recorded goal.
2026-09-23 10:34:06 -07:00
Brennan Benson a375936c04 feat(agent-launch): let a caller reserve the pane its terminal launch creates (#22291)
* feat(agent-launch): let a caller reserve the pane its terminal launch creates

* fix(agent-launch): refuse a launch whose reserved pane is already live

* fix(agent-launch): refuse a live reserved pane before it is revealed

The live-pane refusal used to fire in the executor, after createTerminal
had already issued a handle, published the mobile snapshot and revealed
the tab. The reveal re-registered a fresh launch config over the running
agent's. agent.launch now passes requireFreshPane with a reserved pane,
and createTerminal throws AgentLaunchPaneAlreadyLiveError as soon as
spawn reports it attached to a live pane. That is before any handle,
snapshot or reveal. The spawn reattach itself is the one terminal.create
already uses, so the live PTY is never killed, and the stable-pane
create claim is still released in finally. The isReattach plumbing
added to the launch factory for the old check is gone.

A replay-safe launch refused this way on an existing workspace now
records a failed ledger row, the same way a name collision does.
Before, the row stayed claimed, so every retry got
agent_session_operation_unknown. agent.launchReplay passes the code
through. On create-worktree the workspace already exists when the
terminal is refused, so the row stays unknown. The code is added to
the runtime passthrough list so callers can branch on it.

The pane key is now in the replay fingerprint, deliberately. It is not
placement: group, anchor and focus still stay out of the request and
out of the ledger. It is identity. It is written into the pane's PTY
environment and names the tab the caller has placed. A retry that
reserved a different pane is therefore a different request. Replaying
the first answer would return a key the new reservation can never
find. This matches terminal.createAgentSession, which also fingerprints
its tab and leaf ids. The key is only folded in when present, so every
existing digest is unchanged, and a test pins that.

The wire schema now refuses a tab id the runtime would not adopt as
sent: one with surrounding whitespace, which the runtime trims, and one
longer than 512 characters, which the spawn reservation does not key
on. It reuses the tab-id schema that Placement uses.

* test(agent-launch): pin that a refused live pane issues no handle

The refusal test named handle issuance but only asserted the reveal, so a
throw moved to just before the reveal would still pass. Assert no terminal
is registered, with the attach test as the positive control.
2026-09-23 10:18:16 -07:00
Neil eb18eaf2b6 feat(usage): add Muse Code local usage provider (#22379)
* feat(usage): add Muse Code local usage provider

Scan Muse session logs (including subagent logs, which hold usage the parent
log does not) for model_completed token events and surface them as a fourth
local usage provider: shared scan worker, persisted per-file cache reused by
mtime/size, cross-log dedupe, Stats tab, and Usage Overview integration.
Muse logs carry no price, so the provider reports tokens only.

* fix(usage): name Muse in Stats & Usage copy; skip partial-cost warning when nothing is priced

* fix(usage): surface unreadable Muse sessions root; name Muse in remaining Stats & Usage copy

* fix(usage): count distinct same-content Muse records within one log
2026-09-22 22:44:10 -07:00
Brennan Benson 8757e40063 fix(native-chat): keep a structured agent's tool line between tool calls (#22349)
* fix(native-chat): keep a structured agent's tool line between tool calls

A structured session's status named a tool only while the call was still
running, so the sidebar's tool line went blank whenever the agent was
thinking or writing between calls. Terminal agents keep naming the finished
tool until the next one starts, and clear it after a failure. The structured
status projection now does the same: a running call wins, otherwise the
turn's newest root call if it completed.

* fix(native-chat): bound the structured tool line by the running turn, not the user row

A send made while a turn is running writes its user row into the journal
straight away, and the turn keeps going. Stopping the scan at that row
blanked the tool line while a tool was still running. The scan now runs to
the turn record and names a call only when that record is still running, so
a turn that already ended never lends its last tool to a pending follow-up.

This lookup was the running-only lookup's only production caller, so it
replaces that lookup instead of sitting beside it.

* fix(native-chat): keep naming a structured agent's failed tool until the next one

Clearing the tool line after a failed call brought the blank gap back for
much of a turn: Codex marks any nonzero exit as failed, so a search with no
match or a red test run is enough. The failure already shows on the tool's
own row in the transcript. The running turn's newest running call still
wins; otherwise its newest root call is named whatever it settled to.

* fix(native-chat): name a structured Codex edit on the tool line as the chat draws it

Once a Codex edit's changes exist, its apply_patch call becomes a diff row,
which the status lookup skipped, so the row named the command before the edit.
The chat's tool-call block for a journal row now comes from one shared builder,
and the status lookup reads the same definition: a diff is named as Diff with
its path, and counts as settled since it carries no lifecycle.

* docs(native-chat): describe the structured tool field as running-or-latest

The status summary's toolName/toolInput now name the running turn's
latest tool between calls, not only a running one. Update the wire type
and status bridge comments that still said "the running tool".
2026-09-22 22:42:31 -07:00
Jinwoo Hong 996f9cc306 feat(mobile-web-bundle): gzipped 384 KiB ranges over a capability-negotiated mobileWeb.bundle.range (OTA phase C follow-up) (#22381)
* feat(mobile-web): serve gzipped 384 KiB bundle ranges behind a capability

Adds mobileWeb.bundle.range with its own strict params and result, so
shipped chunk readers see no reply change. The host gzips each range at
level 6 and sends identity when gzip does not shrink it, sharing the chunk
method's read-slot budget and per-asset verification. status.get
advertises mobileWeb.bundle.range.v1 beside mobileWeb.bundle.v1, and the
method is allowlisted for paired phones.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): read the bundle range capability and range replies

Adds the range reply reader and operation, and picks range or chunk from
the status.get capabilities the connection already proved, so an older
desktop keeps being paged in chunks with no probe round trip.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* perf(mobile): keep four bundle chunk reads in flight across the whole manifest

The fetch ran one worker per asset and paged inside an asset sequentially, so
the largest script's 71 chunks were 71 serial round trips while the other
readers idled. One window of four chunk reads now covers every (asset, offset)
on the host's chunk grid, largest asset first. A read_limited refusal narrows
the window and retries the read; eof is still read from the reply.

Synthetic manifest (one 71-chunk asset, five small): 72 round trips -> 19.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* chore(mobile): add fflate 0.8.2 for gzip bundle ranges

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): decode gzip bundle ranges into a bounded buffer

Inflates each range into a buffer one byte past its window, so a gzip
bomb costs at most that allocation and an overlong body is visible. A
corrupt, truncated or unknown-encoding body refuses as range-undecodable;
a body of the wrong decoded length refuses as range-length-mismatch.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): pass the bundle read method from the session to the fetch

The download reads the capabilities of the gates the reducer decided
under and hands the fetch range or chunk. The fetch does not act on it
yet; the range read lands on the pipelined window.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-record the bundle-fetch family under pipelined reads

Baseline moves to a1ee317368, the pipelined fetch.
781 goldens change only their `baseline` header line. Six bodies move:
mobile-web-bundle-fetch-paged, mobile-web-bundle-build-changed, and the four
matrix-mobileweb.bundle-fetch-* goldens.

The two bundle-fetch scenarios now bind requests in pipelined order, largest
asset first, with every chunk sent before any reply: index.html@0 (#1),
index.html@16 (#2), assets/app.js@0 (#3).
- fetch-paged: the request set is identical, only reordered. The chunk
  sender names/ordinals and the scenarioSha256 moved; the replies and the
  fetched bytes did not.
- build-changed: the same reorder, plus one request that is new because
  pipelining puts it in flight before the refusal lands (index.html@16).
  The refusal and the checkpoint are unchanged.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): the bundle chunk comment no longer says a reply picks the next offset

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): read gzipped bundle ranges on the pipelined window

A host that advertised mobileWeb.bundle.range.v1 is paged in 384 KiB
ranges on the range grid, through the same four-read window and queue as
chunks; any other host keeps the chunk grid. Each range is decoded to its
exact window length before the fill checks and the asset hash.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-pin the recording corpus after the bundle comment fix

Baseline moves from a1ee317368 to e94bde327d,
the comment-only commit on a fenced path. All 787 goldens and the scenarios file
change only their `baseline` line. No golden body moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-record the corpus at the range-read pin

Repins baseline to the range-read commit and re-records every golden.
Only the baseline and lockfileSha256 headers move: the lockfile gained
fflate, and the bundle-fetch adapter pages the chunk path, whose
requests and replies are unchanged, so no golden body moves.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): state the on-settle reason that holds for pipelined bundle reads

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-pin the recording corpus after the on-settle comment fix

Baseline moves from 252c592b52 to fe41226ef5,
the comment-only commit on a fenced path. Re-recorded: all 787 goldens and
the scenarios file change only their baseline line. No golden body moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* chore(mobile): keep the lockfile's patch block in main's form

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-pin the recording corpus after the lockfile patch-block restore

Baseline moves from fe41226ef5 to 84d6fca6e7,
which restores main's patchedDependencies form in mobile/pnpm-lock.yaml.
Re-recorded: all 787 goldens change only baseline and lockfileSha256, and
the scenarios file only baseline. No golden body moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): pool four workers over planned bundle chunks, report progress per chunk

Design-review fix round, sketch C: four workers take reads from one planned
chunk queue, largest asset first. They replace the central pump and the
read_limited narrowing. A host frees its read slot before it replies, so a lone
fetch capped at four cannot trip the limit. A refusal now fails the fetch, as
it did on base, and stops the other reads.

- Each asset's buffer is allocated when the plan is built. That removes the
  nullable buffer and its guard. The per-asset byte count is gone, and the
  hash is the oracle (S1, S2).
- The caller's signal is checked before each read and once after the pool
  drains, so an abort during the final window rejects with fetch-stopped
  (S3). The stopped check now covers only the caller's abort. The internal
  stop only makes late replies skip checks, hashing and progress (S4).
- Progress is reported per accepted chunk. completedAssets still counts on
  completion (S6).
- Renames: MAX_CONCURRENT_CHUNK_READS, and `reply` for the RPC reply (N1).
- The slot check states exact-slot acceptance once, then classifies the
  refusal (N3).
- Stale test titles and comments are renamed (N4).

The synthetic manifest still takes 19 round trips.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): say why bundle reads settle at on-settle under pipelined chunks

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile-web): announce bundle ranges on the manifest reply

The manifest reply now names the range grid in an optional rangeBytes,
beside chunkBytes, and the status capability is gone. The range method
takes exactly the chunk params on that grid instead of a caller length.
Both methods share one verified read that returns the six-field header,
and the range handler checks the connection again before deflating.
Range schemas move into the bundle RPC contract; SHA256_PATTERN is shared
from the manifest contract. Shared refusals are tested once over both
methods.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): drop the range capability read and its session threading

The phone will read rangeBytes off the loose manifest reply instead, so
the read-method module goes and the session effects and hook return to
the pipeline branch's version. Range imports move to the bundle RPC
contract, the reply reader reuses the shared SHA256_PATTERN, and a new
test pins that node's level-6 gzip from the host encoder inflates with
fflate to the same bytes. The fetch and window-read modules still import
the deleted names until part 2.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-record the bundle-fetch family under per-chunk progress

Baseline moves to 34fda6f62e. 782 goldens and
the scenarios file change only their `baseline` line. Five bodies move:
mobile-web-bundle-fetch-paged and the four matrix-mobileweb.bundle-fetch-*
goldens. The only change is bundle-progress effects. One report now lands
after the first accepted chunk of index.html (0 assets, 16 bytes), and the
later progress ordinals shift by one. Requests, replies and fetched bytes
are identical. mobile-web-bundle-build-changed keeps its body, because its
one accepted chunk is the whole of assets/app.js.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): page bundle ranges through one window reader chosen by the manifest

The fetch keeps the pipeline's four-worker pool and builds one window
reader from the manifest reply: ranges on rangeBytes when the host names
it, chunks on chunkBytes otherwise. The reader returns the six-field
header and a lazy bytes() so the stop and misroute checks run before any
decode. A range that inflates to the wrong length now falls to the slot
checks, with the one-byte-over buffer as the memory bound, so
range-length-mismatch is gone. A rangeBytes this build cannot page reads
as absent. Fetch names say window, not chunk.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): drain the fake host after the fetch settles so the sibling-stop bounds can fail

The wave host stopped releasing replies once the fetch settled. Reads that
should have been stopped were never answered, so the read_limited bound (7)
and the chunk-failure bound (5) held even with no sibling stop at all. It now
drains until nothing waits. With the worker's stopped.abort() removed, both
bounds fail at 76 requests. assertChunkDescribesAsset's parameter is renamed
to `reply`.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-pin the recording corpus at the window-reader commit

Baseline moves to a25a355546. Re-recorded:
every golden and the scenarios file move only baseline, and the five
bundle-fetch goldens also move lockfileSha256 to this branch's lockfile.
Every golden body is byte-identical to the pipeline branch's. The
recording adapter's scripted host names no rangeBytes, so the bundle
family still records the chunk path.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-pin the recording corpus after the sibling-stop test fix

Baseline moves from 34fda6f62e to 97b13ec7f0.
All 787 goldens and the scenarios file change only their `baseline` line. No
golden body moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-pin the recording corpus at the second pipeline merge

Baseline moves from a25a355546 to 4d2cab31e5,
the merge of the pipeline's sibling-stop test fix. Re-recorded: every golden
and the scenarios file move only baseline. Against the pipeline branch, only
baseline and lockfileSha256 differ; every golden body is identical.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): read the bomb inflation without depending on call order

With the fetch's sibling stop removed, a read left over from the previous
test inflated into the bomb test's record first, and indexOf(601) picked
it. The test now asserts some inflation stopped at 601 and none exceeded
its buffer, whatever else ran.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-pin the recording corpus at the bomb-test fix

Baseline moves from 4d2cab31e5 to 73fde15487,
the test-only commit on a fenced path. All 787 goldens and the scenarios
file change only their baseline line. No golden body moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile-web): tighten the bundle window contract and pin the range sibling stop

The range bomb test reads only its own host's inflations, keyed by the
gzip bodies that host sent, and plans twenty reads so a missing sibling
stop is visible: with stopped.abort() removed it sends all twenty.
The range params are an alias of the chunk params, and the chunk data
bound is the exact base64 length of a full chunk. The phone's chunk and
range replies share one header shape. The window reader closes over the
client and bytes() takes the slot length the fetch computes. The host's
positional read is readMobileWebBundleAssetWindow.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-pin the recording corpus at the window-contract commit

Baseline moves from 73fde15487 to 7346e005a3.
All 787 goldens and the scenarios file change only their baseline line.
No golden body moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-pin the recording corpus at the main merge

Baseline moves from 7346e005a3 to 9c0fe1a546,
the merge of main at 98a6a5325c. Recorded with --record: all 787 goldens
and the scenarios file change only their baseline line. No golden body
moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile-web): drop a test cast and shape-named field maps

The bomb test's inflation log is typed by its hoisted factory's return
instead of an assertion, and the zod field maps shared by the bundle
window schemas are windowParamsFields and windowHeaderFields.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-pin the recording corpus at the lint fix

Baseline moves from 9c0fe1a546 to f4f0915e70.
Recorded with --record: all 787 goldens and the scenarios file change only
their baseline line. No golden body moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): page the range fetch fixture on a small advertised grid

The fake host names a 4 KiB range grid and a 1 KiB chunk grid on its
manifest reply, which the phone pages as it would the real ones, so the
fixtures shrink to a few KiB with the same shapes and each asset is
hashed once. The file runs in about 360 ms instead of 3.5 s, which a
loaded CI runner pushed past the 5 s test timeout. The desktop range
suite still pins that the real host names MOBILE_WEB_BUNDLE_RANGE_BYTES.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-pin the recording corpus at the range fixture fix

Baseline moves from f4f0915e70 to d4d6aadea2.
Recorded with --record: all 787 goldens and the scenarios file change only
their baseline line. No golden body moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-23 01:34:49 -04:00
Neil 51d3cafc4f feat(editor): add a setting to turn off preview tabs (#22398)
Single-clicking a file in the Explorer, or following a link in Markdown
source, opens it as a preview tab that the next preview open replaces.
There was no way to turn that off, so browsing files kept swapping one tab.

Adds `editorPreviewTabsEnabled` (General -> Navigation, on by default).
A caller's `preview` flag is now an intent that `resolveEditorPreviewIntent`
resolves against the setting, covering every open path - files, diffs,
history diffs, conflicts - in one place.

Preview-ness is derived rather than reconciled: readers treat a tab as a
preview only when the stored flag and the setting agree, so a flag left
over from a saved session, another window, or a host switch is inert while
previews are off. Nothing rewrites stored flags when the setting changes,
so no settings-landing path has to remember to clean up.

Fixes #22397
2026-09-22 22:33:06 -07:00
Neil 52a1e2875b feat(orchestration): accept Muse model and effort for supervised workers (#22383)
* feat(orchestration): accept Muse model and effort for supervised workers

`worker-start --agent muse` already launched, but `--model` was refused because
Muse had no session-option catalog. Add one that maps worker preferences to
`muse --model <id>` and `--reasoning-effort <level>`; it seeds no models, so
native-chat surfaces show no picker.

opencode stays without `--model`: the opencode 2 TUI (now shipped as
`opencode`) rejects the flag, so the refusal now tells callers to rely on the
agent's own config. Help, skill guide, and docs list valid `--agent` ids and
the agents that accept `--model`.

Refs #19823

* test(mobile): repin session route closure for the Muse option catalog
2026-09-22 22:20:35 -07:00
Neil 83dd047fd9 fix(explorer): find files by name in large local workspaces (#22369)
* fix(explorer): search local workspaces by file name across every file

The Explorer name filter only searched remote workspaces directly; local
workspaces still filtered the first 20,001 listed files, so files beyond
that cap never matched in large repos. Local name queries now rank the
whole workspace on the host, and fall back to an uncapped git listing
when ripgrep is not installed.

* fix(explorer): filter capped local listings on the host with the Explorer word rule

Replaces the Quick Open fuzzy top-32 routing, which dropped multi-word
matches and capped visible results. The Explorer keeps its instant
renderer-side filter; only when the local listing hits its cap does it
re-list on the host with the same word rule applied before the cap.

* fix(explorer): keep capped matches when host name filtering fails

- Fall back to the capped listing (and stop re-listing) if the host scan fails
- Keep primary matches when the ignored-file pass fails during a filtered scan
- Key host scans on normalized filter words; reset capped state per filter session
- Bound nameFilter size at the IPC boundary; drop the double readdir walk

* fix(explorer): match name filters without locale-dependent lowercasing

* fix(explorer): avoid render-time ref writes in the host name filter fallback
2026-09-22 19:39:52 -07:00
NeilandAdrien De oliveira ebed0964a2 feat(agents): add first-class Muse Code harness (#22216)
* feat(agents): add first-class Muse Code harness

Add Muse as a supervised Orca agent across desktop, mobile, session history, source control, local hooks, SSH, WSL, and native Windows. Preserve user settings, support Muse 1.3 hook environment allowlists, and recognize versioned foreground processes. Include question, waiting, completion, resume, and readiness coverage.

Co-authored-by: homesh-dev <300847526+homesh-dev@users.noreply.github.com>

Co-authored-by: jeffhuen <32542276+jeffhuen@users.noreply.github.com>

Co-authored-by: John Cusack <johncusackccm@gmail.com>

Co-authored-by: Adrien De oliveira <75085839+adriendeoliveira@users.noreply.github.com>

* test(agents): cover Muse remote hook registration

* test(agents): cover Muse hook and source-control contracts

* test(agents): exclude Muse hook metadata from script mode check

* test(agents): keep Muse skill picker coverage stable

* test(ai-vault): include Muse in every-agent fixture

* test(mobile): repin Muse agent icon closure

* fix(muse): detect questions and approvals from structured Muse signals

Muse 1.3 fires no hook for request_user_input, so a pending question left
the pane "working". Its internal reminder subagents also post hooks with
their own session ids (even after Stop), which surfaced "tool failed" rows
and flipped finished panes back to working.

- Read pending questions from Muse's session log
  (user_input_prompt_requested/settled) via the existing transcript poll,
  now generalized from Codex subagents to Muse on main and relay.
- Drop child-session hooks (SubagentStart ids, or turn_id === session_id).
- Treat Notification permission_prompt as the approval wait; PermissionRequest
  also fires for auto-approved calls, so it only caches the approval card.
- Ignore Notification copy as the prompt; poll replays are not new prompts
  or turn boundaries.
- Allowlist USERPROFILE so Windows cmd AutoRun doesn't fail every hook.

* perf(muse): parse only question events from the session log

Most Muse session-log lines are large model/tool records. Filter raw lines
by the user_input_prompt_ marker before JSON.parse via an optional
readJsonlCursor line filter.

* fix(muse): unwrap batched log records and scope questions to the live turn

Review follow-ups: question events inside retained_frame batches were
skipped, and a question left open by a crash or interrupt stayed pending
for the pane's life. Share the history scanner's retained_frame unwrapper,
and only report a pending question whose run_id matches the hook turn_id.

* refactor(muse): drop type assertion in retained_frame unwrap

* fix(agent-hooks): satisfy exhaustive-switch lint in transcript poll policy

---------

Co-authored-by: Adrien De oliveira <75085839+adriendeoliveira@users.noreply.github.com>
2026-09-22 19:13:11 -07:00
Jinwoo HongandDavid Bebawy 7c46a69049 feat(telemetry): report the macOS daemon's code identity on adoption and folder-denial events (#22171)
* feat(daemon): import the macOS process code-identity probe from PR #21826

Takes `daemon-mac-code-identity.ts` and its test verbatim from David Bebawy's
community PR #21826 (stablyai/orca). The probe asks Security.framework, via
`codesign --display --verbose=1 +<pid>`, where a live process's code lives on
disk — the question Node cannot answer, and the one that decides whether tccd
can still resolve a running daemon's code identity after an app update.

Imported unchanged here so the adaptation that follows is reviewable as a diff
against the author's original.

Co-authored-by: David Bebawy <david.ayad2@gmail.com>

* feat(telemetry): report the daemon pid's macOS code identity on the two adoption events

Community PR #21826 argues that macOS terminal daemons lose Documents/Desktop/
Downloads access after an update because the daemon's own executable is
unlinked — Squirrel parks the outgoing bundle under a ShipIt staging directory
and later deletes it — so tccd can no longer map the daemon pid to on-disk
code. Today's `spawner_path_class` and `tcc_attribution` read the binary that
forked the daemon, which an in-place update deletes and recreates, so neither
can see that state.

This adds the detector as a measurement only. `code_identity` rides on
`daemon_adopted` and `daemon_pty_cwd_denied`, the two events that already
describe an adopted daemon, so denied daemons can be cross-tabbed against
healthy ones. Nothing reads the verdict: no replacement, no notice, no UI.

The probe is David Bebawy's, narrowed from a path-carrying union to the closed
enum the wire allows, and memoised per pid so one codesign spawn answers for a
whole daemon generation. Off macOS, or with no pid, it reports `probe-failed`,
which keeps both schemas strict and non-optional.

Co-authored-by: David Bebawy <david.ayad2@gmail.com>

* fix(telemetry): read the daemon's code identity fresh on every adoption event

The probe memoised its verdict per pid and never expired it, so
`daemon_pty_cwd_denied` reported whatever the probe saw at adoption rather than
what was true at the denial. That breaks the measurement in both directions: a
transient codesign failure during startup pinned `probe-failed` for the rest of
the run, and the `parked` to `unresolvable` transition became invisible.
Squirrel leaves the parked bundle in place until the next update, which can be
days, so a daemon adopted as `parked` and denied as `unresolvable` is the exact
crossover this study exists to catch, and the cache hid it.

Now every ask runs its own codesign. Only concurrent asks about the same pid
share a probe, and that entry is cleared as soon as it settles, so nothing
survives to be reported later. Both events are rare enough that one spawn each
is not worth a cache.

* fix(telemetry): drop the dead existence check from the code-identity probe

The classifier stat'd the path codesign displayed and called a missing one
unresolvable. That path is unreachable: once the executable is unlinked,
`codesign --display` prints no `Executable=` line at all and exits 1 with
"No such file or directory", which the fallback below already classifies as
unresolvable. Verified directly on Darwin 25.5 against a signed binary deleted
out from under a running pid.

All the branch actually covered was the window between codesign reading the
path and this process stat'ing it, and it paid for that with a synchronous
stat on the main thread.

* fix(telemetry): never classify a timed-out codesign probe as a verdict

`runProcess` kills the child at the deadline and reports `timedOut`, but the
runner type dropped that field, so a codesign killed mid-display could still
have printed an `Executable=` line and been read as `resolved` or `parked`.
A half-written display proves nothing about where the daemon's code lives.

The runner result now carries `timedOut`, and a timed-out probe returns
`probe-failed` before the output is looked at.

* docs(telemetry): state what each code-identity verdict actually asserts

A reviewer read `resolved` as a claim that the executable sits inside the
installed app and asked for that to be validated. It is not that claim, and we
are not making it: proving containment needs the pid record's spawner path, and
deciding anything from where the code lives is #21826's proposed behaviour
rather than this measurement.

The enum doc now spells out all four verdicts in the terms the probe can
actually support, and says plainly why `resolved` stops at "exists and is not
parked". A matching note sits beside the parked-path pattern.

* docs(telemetry): stop asserting how long a parked bundle survives

The probe's rationale claimed Squirrel keeps the parked bundle "until the next
update". A reviewer claimed the opposite, that it is deleted at the end of the
same install. Neither holds up against this Mac's ShipIt log: the install moves
the outgoing bundle to a TMPDIR ShipIt directory and logs no removal of it at
all, and the one "Couldn't remove owned bundle" line names the incoming
download staging copy, not the parked one. Every parked bundle from the last
two days is nevertheless gone now.

So the rationale in the probe doc, the enum doc, and the reprobe test comment
now assert only what is established: the outgoing bundle is moved aside at
install and disappears later on a schedule we have not pinned down. That is
already enough to justify the design, since one pid's verdict can change
within an app run, which is exactly why every ask reads fresh.

* feat(telemetry): report readable TCC-gated spawns as the code-identity control

`daemon_pty_cwd_denied` gives code_identity's hit rate on denials, but a
readable spawn emitted nothing, so an `unresolvable` adoption with no denial
could not be told apart from a user who never opened a terminal in Documents,
Desktop, or Downloads. The false-positive rate that gates #21826's
auto-replacement was unmeasurable.

`daemon_pty_cwd_readable` now fires when a daemon reads a TCC-gated cwd, once
per daemon and folder class per app run, with the same origin properties as
the denial event. The read-out becomes a 2x2 of code_identity against
readable/denied on protected-folder spawns. Fire-and-forget on the spawn path
like the denial emit, and no app-side directory read.

* refactor(telemetry): one emitter and schema for both cwd verdicts, no dedupe state

The once-per-daemon dedupe on `daemon_pty_cwd_readable` was keyed before the
probe ran, so a daemon first seen readable while `parked` never reported again
once it turned `unresolvable` — the one cell that would count most against
#21826. It also counted per daemon while denials count per spawn, so the 2x2
mixed units.

Readable now reports every spawn, like denied, and both events share one
emitter (`trackDaemonPtyCwdVerdict`) and one schema. The TCC-folder gate lives
in the verdict branch. The origin fields are one shape spread into both
schemas. The codesign probe calls `runProcess` directly and tests mock it,
replacing a test-only runner parameter. The repeated "never cached" rationale
is now said once.

* fix(telemetry): rename the shared origin schema fields for the anti-slop gate

no-shape-in-symbol-names rejects daemonOriginShape; the fields are event props.

---------

Co-authored-by: David Bebawy <david.ayad2@gmail.com>
2026-09-22 22:10:51 -04:00
Brennan Benson 1b85be67d8 feat(native-chat): notify on every settled structured turn (#22105)
* feat(native-chat): notify on every settled structured turn

A structured chat that finished while you were elsewhere lit the sidebar
but never raised an OS notification, and a notification that did arrive
for one could not open the chat it came from.

Unread and delivery now come out of the single resolveAgentAttention
decision the terminal lane already uses: the structured dispatcher calls
applyAgentAttention instead of applyAgentAttentionUnread, so the same
policy that decides what to light also decides what to deliver, through
the same sound and blocked-permission tail.

Every settled turn notifies, as the CLI lane does. Success says
"finished"; failure and cancellation say "stopped" through the shipped
agentInterrupted flag rather than a second vocabulary. A turn whose
outcome the host never stated stays unknown and lights nothing.

The host now dedupes mobile fan-out by event identity (scope, session,
turn) beside the existing per-workspace burst cooldown, so a completion
two windows both saw reaches the phone once while each window still
decides its own banner. Clicking a structured notification reveals the
chat tab: its pane key's leaf is synthetic, so focusTerminal would hunt
a split-layout leaf that does not exist.

* fix(notifications): spend each mobile gate only when it actually notifies

Two review findings on the structured-chat notification lane, both real.

The mobile event gate consumed its reservation before the per-workspace
burst cooldown ran. Two chats in one workspace share that cooldown key,
so the second chat's completion could burn its event key and then lose
the cooldown to the first chat — never announced, yet permanently marked
as announced, so a later window dispatching it could no longer reach the
phone. The gate now peeks first and records the event at dispatch, which
also keeps a known duplicate from burning the cooldown slot.

A notification id is minted from the status row's stateStartedAt, and the
row re-projects that field as the turn settles: the working episode's
start moves into stateHistory and the settled start takes its place. A
banner raised in the window before that re-projection therefore carried
an id acknowledgement never rebuilt, leaving it on screen for good.
Acknowledgement now collects ids for the row's left episodes too — the
same episodes the unread check beside it already scanned, so the two
halves finally read the same turns. Lane-neutral: the terminal lane
mints its ids the same way and had the same gap.

* fix(notifications): drop the mobile event gate and reveal chats in folder workspaces

The per-event mobile dedupe defended against one completion being
dispatched by several Orca windows. Only one renderer mounts the
structured attention bridge, the completion feed is live-only with no
replay, and any in-process duplicate lands inside the existing 5s
per-workspace burst cooldown, which already collapses mobile and
desktop alike. The gate never acted on a real sequence, so the wire
field, the shared ledger and its tests go; mobile delivery is back to
main's behavior.

A folder workspace id ("folder:<id>") has no "repoId::" prefix, so the
click binding was skipped and clicking a chat notification there did
nothing. The chat route selects its workspace itself through
ui:focusEditorTab, so it now binds without a repoId; the terminal
route is unchanged.

* fix(notifications): retire the banner ids actually dispatched, not ids rebuilt from a moved row

A banner's id is minted from the status row's stateStartedAt at dispatch, and that field moves
afterwards: a completion can outrun the settled re-projection, and a settled structured row is
re-stamped with no history entry by any later journal row (a cancel appends a status note after
the turn settles). Rebuilding ids from the row's episodes at acknowledgement missed the second
case and fanned out up to 21 mobile dismissals per pane for ids never raised.

The shared delivery tail now records each dispatched id per subject; acknowledgement retires
those plus the current-row rebuild it always had. The acknowledgement collector is back to
main's single-field form.

* refactor(notifications): retire announced notifications by subject in main

Main now records, per pane, the ids it actually announced (a desktop banner
shown or a phone alert sent) and an acknowledgement passes the acknowledged
pane keys so main retires all of them. This replaces the renderer-side record
of dispatched ids: main is where the announcement happens, so it records only
real announcements, including phone alerts whose desktop banner focus
suppressed. The id rebuilt from the current row stays as the fallback after
a restart empties the in-memory record.
2026-09-22 18:35:18 -07:00
Jinwoo Hong dfff3915c4 fix(browser): scope back/forward/reload/zoom/grab shortcuts to the originating split (#22340)
* fix(browser): scope back/forward/reload/zoom/grab shortcuts to the originating split

With two browser panes visible in a split, Back, Forward, Reload, Hard
Reload, page zoom and Focus Address Bar fired in every visible pane. Main
forwarded these guest chords without the page id, and each split's active
pane subscribed. The renderer-side listeners for the same chords were also
window-wide per pane, so a key pressed in the toolbar (or in a terminal in
another split) reached every active browser pane.

Guest-forwarded chords now carry the originating browserPageId; preload
admits only well-formed payloads and each pane ignores ids that aren't its
own. Toolbar-path listeners use the same focused-split scope Find already
uses. The streamed remote pane's history chord moves onto that scoped hook.

Cmd/Ctrl+C grab (STA-3319) gets the same scope and no longer arms while a
text selection exists outside the browser pane, so copying from the native
chat transcript works again.

* refactor(browser): simplify split shortcut scoping per review

Drop the preload payload admission (main and preload ship together), fold
the three inline scope checks into browserChromeShortcutOwnsEvent, and
replace the outside-overlay selection check with a plain live-selection
rule so Cmd+C copies from surfaces that do not move split focus.

* refactor(browser): share one zoom command type and tidy shortcut comments

BrowserPageZoomEventDetail and BrowserPageZoomCommand were the same shape;
keep one in shared/browser-page-zoom.ts and route guest and local zoom
through a single handler.

* refactor(browser): narrow the zoom event with instanceof instead of a cast

* test(e2e): pin split-scoped browser shortcuts

Two browser splits (and a terminal beside a browser) now prove that Back,
Forward, Reload, Hard Reload, page zoom, Focus Address Bar, and the element
grab chord act only on the split that sent them, from both the guest page and
the browser toolbar. A native chat selection proves Cmd/Ctrl+C copies instead
of arming grab. Split fixtures move to a shared helper so both specs reuse them.
2026-09-22 20:14:19 -04:00
Brennan Benson 9ece273056 fix(native-chat): journal rows name the agent that produced them (#22299)
* feat(native-chat): carry producer linkage on every journal row

A journal is the durable record of one agent SESSION, and a session that runs
subagents journals their rows into the same timeline with nothing on the row
saying which agent wrote it. Add that: a per-row linkage bundle naming the
producing agent, its parent, the provider's raw parent reference as provenance,
the kind of work, and which run of the agent produced the row.

The bundle rides the row BASE, not the body: two nested prompt shapes are
strict, so an unknown key on a body makes the whole row parse as malformed. It
is deliberately not a schema-version bump either — an unknown `v` makes a row
unreadable and latches the host read-only, while an unknown key is ignored, so
an older host reads a stamped row and behaves exactly as it does today.

One reader predicate interprets absence, by presence and not by truthiness: an
id that failed to resolve is still an id, and a truthy test would read it as
root and put the child's content back on the parent. The parent-facing status
scans — thinking, the running tool call, the latest assistant line and the
quoted prompt — now skip rows a subagent produced. The transcript is left
unscoped on purpose: it shows every agent's output.

No producer stamps anything yet; this is the carrier and the reader.

* fix(claude): attribute a subagent's journal rows to the subagent

The Claude translator already parsed `parent_tool_use_id` on every envelope and
threw it away. It now resolves that reference to the producing agent's canonical
task id — never to the reference itself, which names the tool CALL and is
re-minted on every resume, so a row stamped with it would split one child into
two the moment it resumed. The raw reference is kept beside it as provenance.

Resolution is its own module rather than more roster: the roster maintains the
spawn-group row a user reads, while this answers, for one frame's parent
reference, whether the rows it produces are the session's own agent's, some
child's, or nobody's yet. It reads the alias table directly to tell "an
announcement named this spawn call" from "this id is simply unknown", which
comparing the canonical id against the raw one cannot do when the two match.

Where the identity is not final the row waits rather than guessing. A top-level
spawn whose `task_started` has not landed is the one case that can still
resolve, so its rows are held — bounded at 64, oldest written first — and
released when the announcement arrives or when nothing can name the producer any
more. Nothing is dropped and nothing is written as the parent's.

A release that announces no tasks at all is decided immediately instead of held:
nothing stable is ever reachable for its children, and an id that rotates is
worse than no id because it is silently wrong rather than visibly absent. Those
rows read as the session's own, exactly as they do today. Holding them instead
would strand every row that REVISES an earlier one — a tool result would leave
its tool row reading "running" for the rest of the turn.

The attempt counter moves in exactly one place, the existing reactivation branch
where a new spawn alias reopens an entry, and is gated on that observed alias
change rather than on the counter, so a late duplicate cannot advance a settled
run. The first run carries no attempt at all.

The group row keeps its own module's write path, now named there, because it is
the one row written from a child's frame that is deliberately the parent's.

* test(native-chat): pin producer linkage end to end, and fix the harnesses first

Four harnesses in this area silently discarded the append options they were
handed, so every assertion about attribution would have passed against
`undefined`. Two are fixed here — the journal double behind the real deferred
sink, and the Claude subagent translator's sink — and each records the options
beside its existing call log rather than on it, so the assertions about call
order stay about call order.

The read side is pinned first, because linkage correct in the store and never
read by the projection is the way this ships looking finished and fixing
nothing. The three defects are asserted through the live-turn and projection
readers: a parent no longer reads as thinking because its child is reasoning, no
longer shows its child's running tool, and no longer quotes its child's prose or
prompt. The opposite direction is pinned too — the transcript still renders the
child's output, and a parent's own line is never suppressed.

Also covered: a resumed child keeps one identity while its spawn call id
rotates; a child row arriving before its announcement is held and then written
linked rather than dropped; a held row is written under the raw reference when
no announcement ever comes; the buffer's bound writes the oldest row rather than
losing it; an unresolvable id reads as a child rather than as the parent; a row
with no linkage reads as root; the schema version is unchanged; and a strict
prompt shape still parses, with a positive control proving that strictness is
real and is why linkage rides the row rather than a body.

* test(native-chat): give the journal double's cast its SAFETY rationale

Editing inside the object literal re-attributes the pre-existing assertion to
changed lines, and the changed-code gate requires a line-specific rationale.

* fix(claude): name the agent that spawned a nested subagent

A grandchild's rows carried an agent id but no parent, and under this
journal's semantics an absent parent is not silence — it is the claim that
the session's own agent spawned the row's producer. For a task spawned from
inside another subagent's sidechain that claim was simply false.

The frames from such a task carry exactly one handle: the nested call's tool
id. That id was journaled once already, as a tool-use block on the row of the
child that made the call, so the child is recoverable from it — but only if
something remembers which row carried it. The registry that already tracks
which tool calls reached the top-level transcript now records the sidechain
ones too, against the reference naming their owner, and the resolver follows
that reference to name a row's parent.

The reference is recorded, not an identity, and it is resolved through the
same path the owner's own rows resolve through, so a parent id always matches
the agent id the parent's rows carry however either was settled. A row now
persists only once BOTH its producer and its parent are final; a grandchild
whose child has not been announced yet waits in the same buffer, and leaves
it through the same three doors.

* test(claude): pin the streamed-text lane's attribution instead of only capturing it

The checkpoint harness was fixed to record the append options, and then nothing
asserted them: all five of its tests passed unchanged against an implementation
that resolves no producer at all, so the lane's attribution was covered by a
capture and no claim.

These assert it: a block streamed inside a child carries that child's linkage,
the session's own carries no keys at all, a checkpoint is held rather than
written while the producing agent is provisional, the flush before settlement
writes a held block under the raw reference rather than losing it or filing it
as the parent's, and a block's producer is resolved once and kept — every
checkpoint rewrites the same row, so a producer that moved would file one
agent's prose under two identities.

* refactor(journal): name the linkage row fields for their role, not their shape

The anti-slop gate rejects "Shape" in a symbol name. These are the linkage
fields a render item carries, and the name now matches the sibling helper
that builds them.

* fix(claude): hold the session's first subagent instead of filing it as the parent

A child's first frames can arrive before the `task_started` that names it, and
the resolver treated "this release has announced no task" as a settled fact
about the CLI. Before its own first announcement every session looks exactly
like that, so the FIRST subagent's pre-announcement rows were written straight
out as the session's own — the whole defect, for the first child of every
session, persisted with no backfill to repair it.

A spawn call the session forwarded at top level is positive evidence that an
announcement is still expected, so it now outranks the release check. A release
that genuinely announces nothing is unchanged: its rows reach the release
verdict at settle and still read as root, just written a little later.

Also stops a malformed owner chain that loops back from naming an agent its own
parent; the depth guard bounded that walk but could not make its answer mean
anything, and absence is the truthful claim.

* fix(claude): read a tool result as its caller's row, not a child's

Every top-level tool call is a forwarded tool id, not just a spawn, so a result
frame naming its own call as parent resolved as a child awaiting an
announcement that is never coming. The row was parked until the turn settled
and the tool sat `running` in the meantime.

A frame delivering the result of the very call it names is the caller consuming
its own output; only a spawn call ever gets a sidechain. Pins added for that and
for the first-subagent hold, and the never-announced case now asserts the row
was WRITTEN as root rather than that it carries no agent id, which an absent row
also satisfied.

* fix(journal): refuse an empty producer id, and drop a bad one without losing the row

The reader that scopes a parent's surfaces tests PRESENCE, so `agentId: ''` is
present: a row carrying it reads as a subagent's and disappears from its own
author's surfaces for good. Neither validator caught it — the wire schema
accepted any string, and the persisted-row guard type-checked nothing in the
bundle at all, against that file's own stated policy.

The wire schema now requires a non-empty id. The persisted side sanitises
instead: a bad linkage field is DROPPED and the row is kept. Rejecting there
would turn a tightened validator into a whole-store kill switch, and degrading
a row to the session's own agent is what every row said before linkage existed.

* fix(native-chat): answer the turn activity line for the session's own agent

`selectStructuredAgentTurnActivity` is a "what is this agent doing right now"
reader and was not scoped by producer. It builds a label set from every
tool-call row in the turn — a subagent's included — and both readers below it
use that set to suppress a line that repeats it. So a CHILD's tool label could
blank the PARENT's activity line: child data deciding the parent's surface.

Live, not latent: the provider-activity branch is populated in this lane, and
it consults the label set without ever consulting `providerFrame`, which is
what the status fallback loop relies on. Scoped once at the top, so both
readers share one interpretation point. Renderer and mobile share this
function, so both are covered.

* fix(journal): stop a lifecycle batch stamping one producer onto N mutations

A lifecycle-batch row carries N mutations but stamped linkage at ROW level, so
a future mixed-producer batch would silently attribute every mutation to
whoever opened it. Both callers are single-producer today, so this was latent.

The write path no longer accepts linkage for a batch, which removes the failure
mode by construction rather than guarding it. The reducer still READS linkage
off a batch row — a row may arrive from a host that writes one — and a genuinely
mixed batch would have to stamp per mutation, which nothing needs yet. Chosen
over adding a per-mutation field because that would persist a new key forever
with no writer and no reader.

* docs(journal): say why each unlinked write site is unlinked, and drop two false claims

Completes the write-site audit the PR claims. Prompt rows carry no linkage and
CANNOT: a prompt arrives through the SDK's permission callback, whose options
carry a request id and the tool awaiting approval and no parent reference of
any kind — unattributable at that site, not deliberately root. Turn rows are
deliberately root and now say so.

Two comments justified decisions by mechanisms this store does not have. The
linkage docblock cited compaction dropping a start row and a pagination
boundary; there is no compaction, and pagination is complete-or-reset. Per-row
repetition is still right, for the reason that is actually true: every reader
scans back from the tail and stops at the turn. A test carried the same false
framing. `claudeFrameParentRef` claimed to read the field by the same rule as
`isRootClaudeFrame`; it is deliberately stricter on the empty string.

* refactor(claude): write a child's rows through, then correct the attribution

Four misattribution paths shared one cause: the lane committed to an
attribution verdict at write time and could never revise it. That followed from
"the journal has no backfill", which is false — re-appending an `itemId` bumps
its revision, the reducer rebuilds linkage from the newest row, and it pins
`sequence`/`observedAt` so a correction does not move the bubble. This lane
already relied on that twice.

So the order inverts. A row whose producer is still provisional is written
immediately, stamped with the spawn call's own id, and re-attributed in place
when the announcement names it. Bookkeeping no longer gates a user's view of
what an agent said.

The hold buffer is deleted rather than left as a pass-through. Corrections are
bounded and die four ways: the announcement, turn settle, teardown, or the
bound. Passing the bound gives up on that producer WHOLESALE — correcting some
of a child's rows and not the rest splits one child across two ids, which is
worse than correcting none. A correction that would change nothing is dropped
rather than burning a revision.

The streamed lane loses its producer latch, which pinned the first verdict
permanently and is why an announcement one frame later could never reach the
row. Every checkpoint rewrites the same identity, so there is one row per block
and re-resolving can only revise it; the latch was guarding against a split
that cannot happen on this path. A block that stops streaming before its
announcement is re-attributed explicitly, since nothing else revisits it, and
the announcement is now observed BEFORE the forced flush that would otherwise
stamp it a line too early.

Also narrows the no-announcements-at-all escape so it no longer swallows a
forwarded spawn call. That escape now applies only to a sidechain id no spawn
call ever forwarded, where there is genuinely no handle to stamp.

* fix(journal): move two test doubles onto the signatures they pin

Both failed typecheck while passing at runtime, which is what a test double
gets to do: vitest never typechecks them.

The sink's lifecycle-batch double still read producer linkage off the batch
input after that input stopped carrying any, so it had no property in common
with the linkage type. The fence is now all it records, which is what the
narrowed contract actually forwards — and what the test beside it already
asserts.

The row-schema helper returned the whole six-arm `JournalRow` union while every
caller reads `body`. It now narrows to the item arm it always builds, so the
assertions read it directly rather than through a cast.

* fix(claude): resolve a tool result to its real caller, not to the session root

A nested tool row could end stuck `running` with its result content dropped.

Cause was in the result-frame attribution, not in the correction ledger. A frame
delivering the result of the call it names as parent is the CALLER consuming its
own output — but the code read "the caller" as "the session's own agent", which
is only true when the caller is the root. A call a child made is owned by that
child. Collapsing it to root both misattributed the row and made the result's
write resolve through a different reference than the call's, so the correction
owed to that row was left holding the body it had BEFORE the result landed, and
re-attribution then reverted the row.

The caller is now resolved through the registry that already records which agent
journaled a tool call, so both writes to one row resolve through the same
reference and the newest body wins.

A settled write also supersedes any correction owed to its row. One `itemId` is
legitimately written under two references — `claudeToolIdentity` is keyed on the
tool id alone — and a settled write already carries a final verdict, so an
outstanding correction could only restamp it from a reference that write did not
use. Dropped rather than re-bodied for that reason.

Adds the ledger's first unit tests, including the invariant this defect broke: a
correction changes a row's attribution and never its content.

* fix(claude): keep a correction owed when the sink refuses it

A correction went out through the plain append, which discards the queue's
admission. Under backpressure the write was refused and `retry` had already
dropped the entry, so the obligation died with nothing re-deriving it — the
failure class this work exists to refuse. It is self-feeding too: a correction
costs a commit on the same serialized writer that carries live rows, so the
burst that generates many corrections is what builds the backlog that drops
them.

It now uses the admission-returning path the sink already exposes, keeps the
entry outstanding on a refusal, and lets `abandon` try once more. A refusal
there ends it: the row keeps the spawn call's own id, which is usable, and an
obligation with no exit is worse than one that settles for less. `settle` also
reports what actually happened instead of always claiming it wrote, so publish
no longer fires for a write nobody accepted.

Also records why the live-turn scans may read the turn record before checking
the producer. A turn is the session's unit of work and no producer of a
turn-bearing body stamps linkage: Claude's turn rows carry none, Codex has no
linkage concept, the compact row passes only a fence, and the stale-turn sweep
goes through the lifecycle-batch path, which cannot carry linkage by type. The
ordering is safe by construction rather than by accident, and the comment says
so, so a future producer knows what it would break.

* fix(claude): never read a row naming a parent as the session's own

A non-null `parent_tool_use_id` names a child, always. The resolver still had
one branch that read such rows as the session's own agent's — a release that
had announced no task, where the comment claimed "there is no handle to stamp".
There is one: the reference itself. The branch was buying a false attribution to
avoid an id nothing joins on, which is the trade already reversed once for
forwarded spawn calls, and every reader of this field is a presence test.

So the branch goes, and with it the `root` arm of the verdict and the resolver's
whole dependency on whether the release announces tasks. Two states remain:
linked now, or linked now and owed a correction. The type deleted a stale test
double on sight, which is the argument for removing the arm rather than the
branch alone.

This also closes the severe half of the tool-origin eviction exposure. A spawn
id evicted from the bounded top-level set used to flip the release check on and
stamp a child's rows as the parent's; with nothing returning root that cannot
happen. What remains is a missed correction, which splits one child across two
ids — the same end state as passing the correction bound, benign in kind and
disclosed.

`isForwardedParentTool` stays where it gates PENDINGNESS. It now decides only
whether a correction is owed, never whether a row is a child's, so a stale
answer costs precision rather than correctness.
2026-09-22 13:57:46 -07:00
Brennan Benson 4c696a1e2a fix(agent-status): a structured session with live child work reads as working (#22295)
* fix(agent-status): a structured session with live child work reads as working

An idle native-chat session whose subagent was still running showed a green
check in the sidebar, the collapsed worktree pill, and worktree ps, while a
terminal Claude session in the same situation showed working. The two lanes
folded child work into the parent's status with different code: the hook
listener did, the structured lane did not.

Both lanes now share one child-work liveness vocabulary and one lead-status
fold. Live agent work makes a settled lead working; shells and monitors alone
make it monitoring. The structured lane derives liveness from the background
task list already on the wire, in both its readers, so the sidebar, the CLI,
the dashboard and mobile agree. The Claude task-kind table is one shared file
covering both the hook inventory and SDK stream names, and the renderer bridge
reuses the shared child-work projection instead of carrying its own copy.

* fix(agent-status): a blocked or out-of-contact subagent still holds its session working

Child-work liveness retired an agent-kind child on any state but working/monitoring,
while the shell beside it stayed live on everything except done/idle. A subagent
waiting on a permission prompt, or one whose host lost contact, therefore counted
for less than a backgrounded sleep and let the session read done. Both kinds now
share the settlement rule `resolveAgentChildWorkFreshness` already reads rows by:
only an explicit done/idle retires child work.

Also keep empty task labels out of the shared background-task projection candidate,
so a host that publishes `name: ''` cannot beat the child-row fallbacks.

* test(agent-status): pin the widened hook-inventory agent names, and correct two stale claims

The hook inventory now classifies through the shared kind table, which also maps the
SDK stream's `local_agent` / `local_subagent`. Nothing pinned that widening, so add
cases for all four agent names — including `teammate`, whose pane state stays `done`
under the #8825 idle-squat rule.

Two comments the fold made false:
- the teardown marker rule's comment claimed it could not disagree with what the UI
  calls working; it is deliberately lead-only, so now it says that and why;
- the agent-status store reference still described the structured row's `state` as the
  deleted `structuredAgentSessionStatusState`, and omitted the `workingMode` the ingest
  now writes.

* fix(agent-status): the state clock restarts when monitoring becomes a real turn

`stateStartedAt` carried forward whenever the prior `state` matched, which was sound
while `state` meant "a turn is running". Now that it folds in child work, an idle lead
watching a `sleep 3600` publishes `working`/`monitoring`; the user's prompt 45 minutes
later keeps `state: 'working'`, so the row inherited the watch loop's clock and read
"Working for 45m" the instant the turn began. Monitoring is its own displayed label
(`worktree-card-compact-agent-row.tsx:40`), so the continuity key is now the whole
published work identity — state AND workingMode — in both writers.

Also record two facts the code stated wrongly: the structured lane's `interrupted: false`
is inert (a projected session status has no interrupted member) rather than a decision,
and the child-work liveness rule's escape hatch is the roster's session lifetime, not a
settled state.

* fix(agent-status): a workflow is watch work, and child work dates itself

Two defects the fold introduced.

`isAgentChildWorkKind` counted `workflow` as agent work, so a structured session
whose only live task was a backgrounded `local_workflow` published a full working
spinner while the children projection — which admits `kind === 'agent'` only —
rendered nothing to expand, and the same workflow in a terminal pane showed the
monitoring badge instead. The repo already decides this: `isClaudeSubagentTask`
excludes workflows by name, and MATERIALIZED_TASK_KINDS leaves "the backgrounded
shell command and the workflow" to the non-agent owner. The predicate is now
`kind === 'agent'`, and the three sites that restated the same test route through
it, so a new kind is decided in one place instead of three that merely agree.

`evidenceObservedAt` dated every row by `summary.updatedAt`, the journal's last
activity. The journal cannot date child work: its clock stopped when the lead's
turn did, so a genuinely live roster aged past the 30-minute staleness window and
mobile's dot decayed a running session to idle. The fold now reports whether child
work alone holds the row open, and only then does the host's observation clock
stand in — keeping "a restart's republish is not new evidence" for lead turns.

* fix(agent-status): the sidebar dates child work the same way the host does

`fromChildWork` reached the host ingest but not the renderer bridge, so after ~30
minutes of live child work with no journal activity the sidebar's row aged into
staleness while `worktree ps` and mobile stayed fresh — two writers for one session
answering differently, which is the defect this PR exists to remove. For a remote
host the client's own receipt time is also the more honest clock, since the journal
stamp is the host's and is never comparable against this machine's now.
2026-09-22 13:48:07 -07:00
Brennan Benson 60c43695e5 feat(agent-launch): report the pane a terminal launch created (#22108)
* feat(agent-launch): report the pane a terminal launch created

A `term_*` handle is a main-side mapping the renderer cannot resolve
(terminal-handle-links.ts:309), so a client that draws its own tabs had no
way to name the tab it had just asked `agent.launch` to build. The runtime
already mints that pane, bakes it into the PTY's environment and hands it
to its own reveal; the surface factory then dropped it on the floor.

Carry it through as `paneKey` on the terminal outcome. Identity, not
placement: where the pane goes — which group, what order, whether it takes
focus — stays with whichever client is drawing, and nothing here rides the
wire for it. One field rather than a tabId/leafId pair, because the key
already holds both and two copies of one fact can disagree.

Absent when this launch minted no pane: a reused terminal was already
running, and a worktree-create startup terminal is built by the create,
which reports only a handle. Naming the wrong surface is worse than naming
none.

Optional on the wire and optional on the read side. Mobile parses the
receipt with a loose object and is deliberately mode-blind, so it ignores
the field; the persisted-row guard checks it when present and accepts a row
written before it existed, because a read rule stricter than the write side
turns one odd row into a refused replay.

* fix(agent-launch): retain startup terminal pane identity
2026-09-22 09:35:11 -07:00
OrcaWinandm4air ba742a86bb fix(linux): release orphaned processes when their owner exits (#22247)
* fix(linux): release orphaned processes when their owner exits

* fix(linux): handle inhibitor errors until streams close

---------

Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
2026-09-22 05:10:03 -07:00
90d363afc9 fix(renderer): contain Monaco initialization failures (#21555)
* fix(renderer): add defensive error handling for Monaco editor crashes

Analyzed 34 crash reports for v1.4.205 released 2026-09-17. Identified and
added defensive fixes for React error boundary crashes in Monaco editor setup.

- Error: ReferenceError: thũs is not defined
- Location: Monaco editor initialization (editor.api2 bundle)
- Platforms: Linux, Windows, macOS
- Root cause: Undefined variable in Monaco setup or language registration
- Status: Added try-catch to prevent cascade crash

- Pattern: Cascading process deaths (network service + GPU service)
- Platforms: Primarily Windows
- Root cause: Infrastructure/concurrent process failure (not code defect)
- Status: Documented, requires Electron/Chrome infrastructure review

- Pattern: Renderer memory grows to 851MB on low-RAM Windows systems
- Root cause: Memory exhaustion on systems with <2GB free RAM
- Status: Existing memory monitoring detected; needs leak investigation

- Status: Requires minidump analysis with source maps

1. Added try-catch to Monaco editor mount callback (use-monaco-editor-mount.ts)
   - Catches errors during editor initialization
   - Logs file path and error for better diagnostics
   - Prevents crash cascade to React error boundary

2. Added try-catch to Monaco language registration (monaco-setup.ts)
   - Catches errors during Vue/Svelte/Astro/Nim language registration
   - Logs failures without crashing Monaco setup
   - Allows app to continue even if optional features fail

- Analyzed 34 crash reports across 3 categories
- Examined crash dumps, diagnostics, and memory profiles
- Reviewed Monaco setup and editor component code
- Checked git history for recent changes

- Crash breadcrumbs (memory, process state, user actions)
- Process metrics (heap, private memory, system memory)
- Component stacks (React error boundaries)
- Exit codes and system signals

- Error silently continues instead of crashing: Users get degraded experience
  instead of app crash, can still use editor in most cases
- May hide underlying issues: Errors are logged for crash reports, but won't
  be surfaced as prominently

- Type checking: pnpm tc:renderer (passed)
- Changes preserve existing error reporting through crash breadcrumbs
- Defensive coding only adds try-catch, no behavior change for success path

* fix(renderer): keep Monaco mount failures inside error boundary

* fix(renderer): isolate Monaco setup failures

* fix(renderer): contain Monaco mount failures at the editor surface

The try/catch around the onMount body did the opposite of containment: React
already routed that throw to the page boundary, so swallowing it left a
half-wired editor and hid the crash from the reporting pipeline. It also never
saw the reported failure, which is raised inside @monaco-editor/react's own
create effect before onMount runs.

Revert the hook to main and wrap the editor element in
RecoverableRenderErrorBoundary instead, so either throw degrades the file pane
only, still files a crash report, and retries by remounting on the existing
pane+path key.

* refactor(renderer): drive Monaco setup steps from one guarded table

Ten near-identical guarded calls, each repeating its own function name as a
label, become one [label, step] table run by a single loop. Same behaviour: an
optional registration that throws is logged and the rest still run.

loader.config and the editor model registry stay unguarded — they are
load-bearing, so catching there would only move the failure later.

* fix(renderer): breadcrumb swallowed Monaco setup-step failures

A guarded registration that throws was console-only, so a lost language or
behaviour guard never reached crash reports. Record a breadcrumb so the
containment stays visible in the field.

Claude-Session: ab8ff806-4870-4ea8-bbf5-bbd123b1166e

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
2026-09-22 02:57:18 -04:00
Jinwoo Hong e47ef8cc28 feat(mobile): the shell tells a page which optional capabilities it has (OTA phase D, C8.1) (#22141)
* fix(mobile): publish page-route pairs the strict host schema accepts (OTA phase D, C8.1)

`routeViewOf` handed the manifest's own route entries to the host as
`pageRouteGrants`. The phone reads a manifest route loosely, so an entry
arrives carrying whatever field the desktop that wrote it knew about, and
`BridgePageRouteGrantsSchema` is `.strict()`: one unread key refuses the
pairs, `createBridgeHost` refuses the route with them, and the page gets no
`init` at all rather than losing one field.

Fixed before any route carries an optional grant (ruling 37.4), so the
manifest field the next commits add costs an installed shell nothing.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* chore: drop the closure and bundle probe scripts from the tree

Scratch measurements for C8.1 (which route closures reach the HTML preview,
and what the preview render rig costs to bundle with a client provider). They
belong outside the repository and were swept in by the previous commit's
`git add -A`.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): a manifest route may declare optional grants (OTA phase D, C8.1)

Design B of design-ota-c8-1.md, ruling 37. `MobileWebBundleRouteSchema` grows
`optionalGrants` under the required lane's own grammar, with the 16-name
ceiling applied over the union of the two lists rather than to each. Serving a
route still reads `grants` alone, so a capability a screen cannot work without
stays required and takes the route native; a session's granted list is
`[...grants, ...optionalGrants]` narrowed to what this shell implements, from
one helper that both `grantsForRoute` and the `pageRouteGrants` publish read.

The ruling's compatibility rationale is corrected in place. `z.looseObject`
passes unknown members through rather than dropping them (measured, zod
4.4.3), so a shell older than the field still receives the key; what it lacks
is a policy that reads one. What makes the lane safe against such a shell is
therefore the previous commit's publish fix, not the reader.

BRIDGE_PROTOCOL_VERSION stays 1. No new notify, verb or frame field.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): name the shell's cancelled-navigation behaviour as a grant (OTA phase D, C8.1)

`externalNavigation` joins `MOBILE_WEB_SHELL_GRANTS` beside `screencastBinary`
and `haptics`, declared in `cancelled-navigation-target.ts` because that is
the module holding the rule which acts on it. A third token that is neither a
verb nor a notify: the page posts nothing to make a cancelled top-frame
navigation happen, so this list is the only thing that can tell a page whether
a tap inside the sealed HTML-preview frame escapes at all.

A constant and not a platform read (ruling 37.1): both engines dispatch the
event, `ios/MobileWebShellView.swift:481` and Android's
`MobileWebShellView.kt:382`, so an app build carries the behaviour on both or
on neither.

The policy census grows the half that was only pinned by the verb table: the
implemented set is that table plus exactly three non-verb tokens, each read
off the module that declares it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): the bundle builder carries a route's optional grants (OTA phase D, C8.1)

`resolveMobileWebPageRoutes` maps each declaration member by member, so a
field the declaration grows reaches a phone only once the map names it: until
now `optionalGrants` would have been dropped in silence and every route would
have declared nothing optional. Omitted when the route declares none, because
absent and empty are the same answer to a shell.

The declaration suite grows the rule rather than a row: the map carries the
lane through and writes no key without one, and the lane is held to the
manifest's own grammar and to the ceiling over the union of the two lists.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): the HTML preview hides its links on a shell that cannot open one (OTA phase D, C8.1)

The session route declares `externalNavigation` on the optional lane, and the
preview asks for it before it renders an artifact's links as links. Ruling
37.2's three readings are what "hide" means here, and removing `href` is what
delivers all three at once: `a:any-link` stops matching, so the UA stylesheet
stops underlining, the element leaves the tab order, and there is no dead
anchor a tap does nothing on. The text the author wrote stays where it was,
the artifact paints, and the Preview/Source toggle is untouched.

Done with the browser's own parser rather than over the source text: an `href`
inside a comment or a `<template>` is text to a browser, and a pass that
rewrote either would be editing the artifact instead of its links. The frame
also loses `allow-top-navigation-by-user-activation` on that path, so a link
the pass somehow missed is refused by the browsing context as well.

One route, measured rather than assumed: the design said two, and the file
preview route's closure does not reach the HTML preview at all - it renders
`MobileFilePreviewScreen`. The new closure census derives that list from the
hook's callers.

The render rig grows the case on both engines and the readings it needs, and
`mobile-web-app-preview-frame-readings.mjs` is split out of it at the
readings/arms boundary, because the two were over the 600-line cap together.
Two engine findings are recorded in the rig: an `<a>` with no `href` still
answers `tabIndex` 0 on both, so focusability is asked by focusing; and WebKit
computes `cursor: auto` for a real link, so that reading is pinned where it
discriminates and its blindness pinned where it does not.

The hop-coverage census now reads the effective set, because that is what the
running rule compares. Inert today: the session route is the only declarer and
an opener into every other route.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): pin the preview's hidden link path where the unit suite can reach it (OTA phase D, C8.1)

The mobile suite runs in a `node` environment whose resolver has no `.web`
precedence, so `MobileHtmlPreview.web.tsx`'s import of the grant hook lands on
the native sibling, which answers yes unconditionally. That is why the
existing component suite still measured the granted frame without knowing a
grant exists, and it means the hidden path had no coverage in the sharded
`test` job, where the render rig is skipped for want of the bundler's
dependencies.

So the wiring gets its own file with the module replaced: that the component
asks, and that both the frame's sandbox and the document it is handed follow
the one answer. happy-dom rather than the suite default, because the inerting
pass parses with the browser's own `DOMParser`.

`String(node.type)` rather than a literal comparison: `node.type` is
`ElementType`, which overlaps a real intrinsic tag and not the host strings
these mocks render, so `=== 'Pressable'` is a no-overlap error under
`tsconfig.test.json` and the tests-typecheck ratchet reds on it.

Also replaces a `Reflect.get` the anti-slop gate refuses with an `in` check.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the session page closure at 4,362 for C8.1's three modules

Measured on both sides with `mobileWebAppRouteClosure(SESSION_ROUTE)` at base
`841d06a969` with all five postinstall generators run first, and the two
`local` lists diffed rather than the total inferred: 4,359 -> 4,362 modules,
1,017 -> 1,020 local.

All three are local source modules and none is vendored: the page's read of
`init.grants.native`, the pass that turns an artifact's links back into text
without the grant, and the module declaring the token beside the rule that
acts on it - reached both by that hook and by `page-route-policy.ts`. The
`bridge-caps.ts` it imports was already in this closure, and the hook's native
sibling is replaced rather than joined.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): allowlist the preview's grant sibling among the .web.* overrides

`mobile-web-app-web-overrides.test.mjs` pins the allowlist against the `.web.*`
files on disk, so a new web sibling reds it until the file says why the page
needs one. Red before: `expected [ …(36) ] to deeply equal [ …(37) ]`, naming
`src/components/use-html-preview-link-grant.web.ts`.

The preview's own entry is corrected with it: its reason said
`allow-top-navigation-by-user-activation` is granted, and that token is now
conditional on the shell answering that it can open such a navigation.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): the hidden-link render case waits on the frame's own reading (round 1)

CI's chromium arm timed out at the full 240 s on this case alone while the
WebKit sibling passed in 1.5 s and it passed 26/26 locally. The cause is the
third arm: it tapped the granted link and waited through
`expectNavigation: 'main-frame'`, and `waitForRecordedNavigation` has no bound
but the case's own timeout. Under CI load the click missed its 2 s
actionability window, no navigation was ever recorded, and the arm sat in that
wait until vitest gave up - `recorded []`, with the frame attached only at
38.9 s. Three arms sharing one budget is what made this the case to find it.

The arm is dropped rather than its wait lengthened or retried. Every verdict
left is a reading the frame itself publishes: the anchors its document holds,
the style the engine computed for one, whether focus lands on it, and now
whether the tap this arm made landed at all - `actError` is asserted null, so
a click that never reached its target is no longer the same three zeros as a
tap that did nothing.

Nothing is lost. The tap's outcome on a granted shell is the next case, on
these same counters from this same rig and with a budget of its own, which is
the presence precondition this file already uses elsewhere for the same
reason. The WebKit sibling's discriminating reads are untouched.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): the inert-link pass changes nothing an engine renders but the links (round 2)

Round 2's ruling: the hidden-link path may change nothing about the artifact's
rendering except that links are not links. A parse and a reserialise is not
free of that by default, and all four findings reproduced on Chromium 147 and
WebKit 26.4.

A same-document fragment link is kept. It starts no navigation at all, so it
goes on working inside the sealed frame whatever the shell can do, and taking
it away would be degradation over a capability it never needed - an artifact's
own table of contents is the case. Its `target` still goes, because a fragment
aimed at another frame is a navigation rather than a scroll, and `href=""` is
not a fragment: it resolves to the frame's own URL.

Links inside `template.content` are reached, recursively. `<template
shadowrootmode>` is a declarative shadow root the frame's parser attaches and
renders, and `querySelectorAll` does not walk into template content, so those
links arrived live inside a sandbox that refuses their navigation - the dead
anchor ruling 37.2 forbids. Measured: `parseFromString` attaches no such root
on either engine or in happy-dom, so the pass can reach them.

The leading newline of a `pre`, `listing` or `textarea` is written back. A
parser drops one after the start tag and the serialiser is specified to put it
back; measured, neither engine's does, so a round trip lost a blank line from
every such block.

The doctype is carried whole, and the reason is corrected from the one the
finding gave. It cannot move this frame between layout modes: a `srcdoc`
document takes its mode from its embedder, and measured, a quirks doctype, the
bare name and no doctype at all all read `CSS1Compat` inside the frame. What
rewriting it does is change the document the author wrote for no reason, with
`document.doctype` observable beside a Source tab showing the original. The
render case pins `compatMode` as the blind reading it is and reads the frame's
own doctype identifiers as the one that discriminates.

Option B was not available: the frame has no `allow-scripts` and inherits
`script-src 'self'`, so nothing runs inside it and there is no injection to
carry the work.

Also drops a vacuous half of the affordance test. `renderSource()` is called
with no argument, so the markup a Source view shows is the caller's own
closure and asserting it equals the fixture passed whatever the component did.
What the component decides is whether the rewritten frame stays mounted
underneath, and that is what is read now.

`mobile-web-app-preview-arm-driver.mjs` is split out of the render rig at the
boundary the readings module already names - the rig holds what each case
claims, the driver how an arm is driven, the readings what it reports - since
the three were over the 600-line cap together. No max-lines disable or bump.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): a fragment link is a frame navigation in this preview, so it is inerted too (round 3)

pullfrog is right, reproduced on both engines before believing it. Round 2
kept `#`-prefixed hrefs on the theory that they are same-document scrolls. In
this frame they are not: the document's URL is `about:srcdoc` while its base
URL is inherited from the embedder, so `#section` resolves against the shell's
own URL and the destination differs from the document's by more than a
fragment - which makes activating it a frame navigation, and the shipped
`frame-src 'none'` refuses it.

Measured under the shipped policy, one tap, with something to scroll:

  Chromium 147   scrollY 0, frame becomes chrome-error://chromewebdata/,
                 artifact gone, embedder reports frame-src <origin>/preview
  WebKit 26.4    scrollY 0, frame stays about:srcdoc and intact, same report

So the destruction is Chromium-only but the absence of a scroll is not: there
was no working affordance to carve out for, and the carve-out left a live link
that destroys the preview - worse than the inert text it was meant to avoid.
Both sandbox values behave the same, so this is the base URL and the policy
rather than the sandbox.

The same tap does the same thing on the granted path, where this pass does not
run, so an artifact's internal links have never worked in the preview. That is
not this change's to fix; it is recorded in
`followup-html-preview-fragment-links.md`, and the render case reads the
granted arm's violation as its presence precondition so the behaviour is
pinned rather than merely known.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-22 02:11:40 -04:00
Jinwoo Hong 7650abe224 fix(macos): tell the user when Orca's terminal service can't read their folder, and walk them through the fix (#21923)
* fix(macos): tell the user when Orca's terminal service can't read their folder

On macOS, a terminal daemon that survived an app update can be refused access to
a workspace under Documents, Desktop, or Downloads while the Orca app itself can
still read it. Terminals opened there die with "Operation not permitted" and
nothing on screen explains why. The daemon has reported `cwdReadableByDaemon` on
every create since #18043 and main has emitted `daemon_pty_cwd_denied` on proven
divergence since then; the field data says 1,438 users hit it in 21 days. What
was missing was the notice.

The verdict itself moves off `access()`. A grant-less probe on an affected
machine showed a TCC mode where `access(R_OK|X_OK)` passes on `~/Documents` and
`opendir` still fails, so the check now does what a shell listing its cwd does:
`opendirSync`, one `readSync`, `closeSync`. Only EPERM/EACCES reads as denial —
a missing path, a non-directory, or an unexpected error still reads as readable,
so a non-permission failure can never masquerade as one. The same probe is what
the app side compares with, through one oracle shared by the telemetry emitter
and the notice, so the spawn path reads the directory once.

Proven divergence now also records evidence in main: one entry, keyed by the
daemon's pid, start time and launch nonce, carrying an opaque digest of that
identity and the folder class. No path leaves main. The existing focus-time
`macTccAttribution` poll carries it to the renderer, which raises a second toast
latched per daemon scope: dismissed stays dismissed, and a restart mints a new
identity so the poll returns null and the toast clears with no post-restart
probe. If the replacement daemon is denied too, about 31% of cases, the next
spawn re-records under the new scope and the notice returns, now with the
re-allow sentence doing the work.

No new IPC channel, no daemon protocol field, no polling change, and nothing new
on the spawn path beyond one `opendir`. `daemon_folder_access_notice` counts
shown, dismissed and open_manage_sessions against `daemon_pty_cwd_denied` as the
denominator; `shown` is emitted from main the first time a scope leaves the IPC
handler, so the renderer carries no telemetry plumbing for it.

* fix(macos): clear folder-access evidence only when the same folder class reads back

A readable spawn in ~/code said nothing about a Documents denial but was
hiding the notice; retire the evidence only when the daemon reads a folder
of the class it was denied on.

* fix(macos): say what a terminal-service restart actually does

The Manage Sessions restart confirmation still described the product as it was
before agents resumed themselves: it promised panes showing "Process exited"
that the user reopens by hand, and mentioned legacy-protocol sessions nobody
outside the daemon code can act on. Open terminals and agents come back on
their own now, so the old copy made a routine remedy sound like data loss.

It also called the thing a "daemon". The same restart is about to be offered
from a user-facing fix dialog, so both surfaces now say "terminal service", and
the confirm button is just "Restart".

The new body adds the one fact the old one never stated: terminals on remote
hosts are not affected. Translations of the two changed strings are dropped so
the five non-English locales fall back to English rather than keep showing copy
that is now wrong.

* feat(macos): give the denied-folder notice a fix the user can follow

The folder-access toast told the user their terminal service could not read
Documents and then handed them a paragraph: restart from Manage Sessions, and
if that does not work, re-allow Orca in System Settings. Both halves were
guesses. Roughly a third of restarts do not fix it, and the user had no way to
know which case they were in before spending every open terminal on finding
out.

Main can now answer that. `daemon-folder-access-probe.ts` forks a short-lived
child of the app binary the same way the daemon itself is forked, runs one
opendir/readdir/closedir against the denied path, and prints a single JSON
line. macOS attributes a TCC grant to the process that forked the child, so a
child of the app running now answers exactly the question the running daemon
cannot: would a replacement daemon get in? The child goes through the shared
child-process wrapper, never a shell, with a 3s deadline, a 1KB output cap and
an environment scrubbed to PATH/HOME/TMPDIR. Every failure — timeout, bad
output, spawn error — reads as `unknown`, never as a verdict.

That answer rides out as `restartWillHelp` on the evidence the existing
focus-time poll already carries, and the toast becomes a title and two buttons:
Fix… and Not now. Fix opens a dialog with the two real steps. When the grant is
already in place, step one is shown as done and Restart is live. When it is
not, step one is open and Restart is disabled until it completes — which it
does by itself, because the poll re-probes while the answer is still no, and
returning from System Settings is the moment that lands. An unanswered probe
never accuses the user of a missing grant; it leaves both steps open.

Restart calls the management API directly rather than stacking the Manage
Sessions confirmation on top, since the dialog already states the consequence.
Success replaces the steps with a done line and takes the toast down; failure
says so inline and leaves the button usable.

System Settings opens through the existing developer-permissions pane opener,
which takes an id rather than a URL, with Files and Folders added to it. The
event's action enum now also counts fix_opened, settings_opened,
restart_clicked and — emitted from main when a replacement daemon's first spawn
lands in the folder class the previous one was denied on — whether the restart
actually worked.

* fix(macos): let the folder-access notice return after a poll that read no daemon

A daemon identity reads as null during any reconnect blip, and the poll reports that as
"no mismatch". The notice dismissed itself and then never showed again for that daemon,
because the once-per-daemon latch still held its scope. Only "Not now" should latch.

* fix(macos): say what the folder-access notice costs the user

One line read like a stray warning. The toast now says who is blocked and what fails,
and still leaves the steps to the fix dialog.

* fix(macos): give the folder-access toast one action and the X, like every other toast

"Fix" is the only button; the X dismisses. Sonner fires onDismiss for programmatic
dismissals too, so the post-restart takedown now goes through the store and the hook,
and only a user's X is counted as dismissed.

* fix(macos): keep the fix dialog's steps a checklist and put the one action in the footer

Buttons inside each step made the list look like a form, and a footer Close duplicated
the X. The footer now carries the active step's action, with a ghost Cancel; a probe
that could not answer says so under step 1 instead of showing a check.

* fix(macos): let the checklist show the fix landed instead of saying so

A hedged sentence addressed to the user read like chat. On success both steps check
off and the footer offers Done; the unanswered-probe helper is a status, not advice.

* chore(i18n): drop the fix dialog's unused close key

* Revert "chore(i18n): drop the fix dialog's unused close key"

This reverts commit 365915df48.

* chore(i18n): drop the fix dialog's unused close key

* fix(macos): tell step 1 what to do when the folder toggle is already on

Users who need step 1 usually find Orca already allowed in System Settings; the grant
is recorded but not honoured for the daemon. Re-toggling re-records it.

* fix(macos): drop the unverified toggle instruction from step 1

Nothing has been confirmed to fix a grant that is already on, so the step says only
what the probe knows.

* refactor(macos): share the tccutil reset and bundle-id read behind one module

Clearing a macOS TCC row is about to have a second caller: the daemon
folder-access fix (STA-7948) needs the exact `tccutil reset` the computer-use
helper already issues. Extract both it and the PlistBuddy bundle-id read into
src/main/macos-tcc-reset.ts so the two remedies cannot drift apart.

The extracted calls go through runProcessSync rather than a fresh
node:child_process import: the spawn chokepoint's ratchet holds the direct
importer count at a pin, and a new module with its own spawnSync would raise it.
Behaviour is unchanged except that both calls now carry a 10s bound, and the
computer-use test asserts the same argv against the chokepoint's options.

* feat(macos): offer a permission reset when restarting the terminal service cannot help

About a third of the users who see the folder-access notice are still denied by
a freshly forked daemon even though Orca itself is allowed under Files and
Folders, so the restart the dialog offers cannot fix anything for them. That
state previously had one action: open System Settings, where the toggle they
would look for is already on.

The denied state now offers "Reset permission". Main clears Orca's TCC row for
that folder class with tccutil, then reads the folder from the app itself so
macOS raises its prompt against Orca rather than the daemon, then forces a
fresh-daemon re-probe that bypasses the poll's reuse interval. The dialog
re-renders from that verdict: allowed turns step one green and offers Restart,
still denied says so, and a refused reset points back at System Settings.

Nobody has confirmed this remedy on an affected machine, which is why main emits
the re-probe's verdict as reset_outcome_allowed/still_denied/unknown. Those
three, plus reset_clicked, are the evidence that decides whether the feature
stays.

* fix(macos): say what the permission reset does, and keep System Settings as the fallback

Step 1 was labelled like a Settings task while the button did something else, with two routes
in the footer for one step. The denied state now names the step for what the reset does,
explains it under the step, and shows System Settings only after a reset fails or leaves
things blocked.

* fix(macos): count a folder-access restart only against evidence that survived

The stored denial is the prior denial, so a second copy of it outlived the
one event that retires it: a daemon that read its own folder back cleared the
entry but left the copy, and the next daemon's first denial was then reported
as a restart that had never happened.

Track the outcome on the entry itself, drop the spawn-path probe (ten denied
terminals forked ten probe children the focus-time poll re-runs anyway), and
stop emitting `shown` from a getter the reset path calls for data. The
renderer's toast latch is what decides a scope is shown, so it emits it.

Both accessors now read one identity-matched entry.

* refactor(macos): name the folder-access verdict instead of encoding it as a tri-state

`restartWillHelp: boolean | null` re-encoded a verdict the probe already
returns as a named union, so every reader had to remember that `false` meant
"Orca itself must be re-allowed" and `null` meant "no answer".

`freshDaemonAccess: 'allowed' | 'denied' | 'unknown'` says it, end to end
through main, the IPC payload, the preload mirror and the dialog. The reset's
outcome event becomes a lookup. No user-visible string changes.

* refactor(macos): give the folder-access notice one latch instead of three

Two refs in the hook and a field in the store tracked the same fact, and the
dialog reached the hook through a store field plus an effect just to take its
own toast down before sonner echoed the dismissal back.

The store now holds the visible scope and the scopes the user closed, and
exposes the three things that happen to a notice: it is shown, someone else
retires it, or the user dismisses it. The dialog calls retire directly and the
effect is gone. `settingsIsFallback` loses an argument that was always true at
its only call site, so it becomes the local it always was.

* refactor(macos): stop blocking main on the tccutil reset

Two spawnSync calls with a ten-second timeout sat inside an async IPC handler,
so clearing a TCC row held main's event loop for as long as either binary took.

Both now run through runProcess. The computer-use caller that shared them was
already async, so it awaits them.

* test(macos): run the folder-access probe script against real paths

Every other test mocks the spawn away, so the minified child script — the one
piece that duplicates enumerateDirectoryOnce's errno mapping — had no oracle.
It now runs against a temp directory, an absent path, a file, and a directory
whose mode withholds it, which is skipped for root and on Windows.

* refactor(macos): read the folder-access entry through one identity match

All four callers that ask "is this evidence still this daemon's?" now go
through the same private accessor, so the rule the canonical path depends on
lives in one place.

* fix(macos): keep folder evidence through a failed health read, and make a forced re-probe always probe

A rejected attribution-health read nulled the folder evidence on the same poll, which the
renderer read as "cleared". A forced refresh after a reset returned early on an older
settled verdict. The dialog also closes when a reset finds the evidence gone, and stops
showing the unverified helper once the restart is done.

* fix(macos): name the folder in the access-notice scope

One daemon denied two protected folders kept one scope, so the toast, the
fix dialog, and the tccutil reset could each be about a different folder.

* refactor(macos): derive the folder-access dialog from the latest verdict

The store held an `open` flag and a mismatch frozen at the moment the toast
was raised, so the dialog could open on a stale verdict and its remedy state
could survive a close. It now keeps the latest verdict and the scope the user
opened, and the dialog is shown only while the two agree.

* fix(macos): offer the permission reset only where there is a row to reset

A workspace symlinked out of Documents or on an external volume can be denied
too, and the dialog offered a reset that main refuses. One shared list of the
TCC-backed folder classes now decides both.

* fix(macos): give the permission prompt's read a deadline

An unanswered macOS sheet blocks the app's folder read for as long as the user
ignores it, and the fix dialog is modal and busy until that read returns. The
wait now ends after a minute and reports an unknown outcome rather than
probing under the sheet.

* fix(macos): count the folder-access notice once per scope

A reconnect blip reports no daemon, which takes the toast down and lets the
same scope raise it again. Both raises counted as separate notices, inflating
the denominator behind the affected-user rate. The two latches are now one
map from scope to phase, and the count follows first insertion.

* fix(macos): drop the restart warning once the restart is done

Step two ticked green while its helper still warned that open terminals and
agents would restart, which had already happened.

* fix(macos): keep the folder-access toast up when the fix dialog opens

Sonner deletes a toast after its action button runs unless the handler
prevents the event, and it does so without calling onDismiss. Clicking Fix
therefore took the notice off screen while the scope stayed latched as
visible, so cancelling the dialog left no way back to it.

* refactor(preload): reuse the shared daemon cwd class instead of copying it

The five folder classes were hand-mirrored in preload behind a comment saying
preload cannot depend on main-only modules. The enum lives in src/shared,
which preload already imports from elsewhere, so the copy could drift.

* refactor(macos): close the fix dialog when its evidence disappears

A null verdict left the opened scope set, so the same scope coming back
remounted a checklist nobody had opened. Clearing it on a null verdict also
makes the dialog's scope key redundant, so it goes.

* fix(macos): let each fix-dialog button report its own work

The footer swaps the reset for a restart as soon as a poll says the grant
landed, which can happen while the reset is still running. Both buttons read
their label off the dialog being busy at all, so the restart button appeared
spinning as "Restarting…" for a restart nobody had started.

* fix(macos): clear the reset failure once the permission is granted

"Couldn't reset the permission" stayed on screen after the user granted it in
System Settings and the probe read allowed, contradicting the ticked step
above it. Its sibling line was already gated on the same verdict.

* fix(macos): end the folder-access remedy with the evidence it is about

Two ways out were missing. A reset that cleared the evidence closed the dialog
but left the toast on screen, because only the poll retired it; the store now
retires the notice whenever a verdict comes back null, so both callers get it
and the hook's own branch goes. And the opened scope survived a verdict for a
different scope, so the original one returning later reopened the dialog with
nobody having asked for it.

* fix(macos): keep the folder prompt off main's spawn path

The app-side readability check moved from accessSync to opendir when the
notice was added. TCC lets accessSync through but gates opendir, so on a
machine that has never granted Orca the folder, spawning a terminal there
raised the macOS sheet and froze main until the user answered it. The read is
async now and the spawn no longer waits for it. The blocking variant keeps a
name that says so, and the reset module's own copy of the read is gone.

* fix(macos): only say a folder is still blocked when something re-read it

Two paths reached "Still blocked after the reset." with no verdict behind it:
an unanswered prompt, where the reset returns the verdict stored before it
ran, and a re-probe that could not answer. The reset now returns the same
access it reports to telemetry, and the line waits for a real denial.

* refactor(macos): let the folder-access refresh decide when to skip itself

The poll handler re-implemented the refresh's own two guards, a null entry
and a settled allowed verdict, so each had to be kept in step by hand.

* fix(macos): stop the daemon blocking on its own folder read

The daemon reads the requested cwd before forking a shell to report whether
it can list it. That read is the one macOS gates, so on a folder the daemon
is refused it could hold the daemon's event loop behind a prompt. It is
awaited now, which leaves the blocking enumerator with no callers.
2026-09-22 01:09:26 -04:00
eb92222e7f feat: support Antigravity as supervised worker (#21705)
* feat: add supervised Antigravity worker support

* fix: address Antigravity worker review findings

* fix: stabilize Antigravity readiness detection

* fix: allow Antigravity resume footer after readiness

* fix(antigravity): make agy reach worker_done as a supervised worker

Three defects each blocked `orchestration worker-start --agent antigravity
--worktree new-child` at the agent_readiness stage.

1. Readiness never fired. The composer check required the trimmed line to be
   exactly one character, but agy 1.2.7 launches in accept-edits mode and paints
   it into the caret row (`> Accept-edits mode: ...`). Widened narrowly to a bare
   `>` or `> <name> mode:`; matching any `> <text>` would make every menu dialog
   read as ready, since they all prefix their highlighted row the same way.

2. No trust artifact for agy. Added markAntigravityWorkspaceTrusted, writing
   ~/.gemini/antigravity-cli/settings.json under `trustedWorkspaces` — verified
   empirically against agy 1.2.7, and distinct from the Gemini CLI's
   trustedFolders.json, which agy does not consult. Trust is exact-path and not
   inherited by subdirectories, so each child worktree needs its own entry.

3. The orchestration path skipped the preset. Orca has two trust dispatch
   chains: the renderer's preflightAgentTrust and the main-process
   markLocalWorktreeTrusted. worker-start only takes the second, which matched
   cursor/copilot/codex and fell through for antigravity, so the trust write
   never happened while renderer-side tests passed.

Verified live end to end: the dispatch settles `succeeded` with worker_done
carrying the right task and dispatch ids, and the worktree is appended to agy's
settings with sibling keys untouched.

Known gap: remote-agent-trust-presets.ts has no antigravity branch. The SSH
artifact path is unverified, so agy over SSH still stalls at agent_readiness.
Recorded in a comment there rather than guessed at.

* fix(antigravity): wire trust preset through preload safely

* fix: preserve Antigravity readiness across transcript tails

---------

Co-authored-by: Neil <neil@stably.ai>
Co-authored-by: LielinaH <lielinah@gmail.com>
2026-09-21 20:22:08 -07:00
0677271709 fix(orchestration): reap leaked worker terminals via process-incarnation fallback — stops an unbounded PTY/process leak on Remote Server (OOM / cgroup PID exhaustion) (#18790)
* fix(orchestration): remint live handle from process incarnation on worker release

When a durable terminal handle goes stale (rendererGraphEpoch fence),
inspectWorkerTerminal re-mints a live handle via
resolveTerminalHandleByProcessIncarnation + matchesProcessIncarnation so
release/stop/read act on the still-running PTY instead of reporting
missing and leaking the agent process tree.

- keep main shared host-scope re-exports; add matchesProcessIncarnation
- wire observation.terminalHandle through control/stop/release
- rebuild release-completion on main structured paths
- on missing/unattached + provably exited: settleDead fence first, then
  same-incarnation settleWorker fall back (archive may block settleDead
  mid-request); settle before recovery defer

* fix(orchestration): derive SSH host scope from the reminted handle; reuse fresh-request recovery guidance for structured workers

Addresses two open CodeRabbit review comments on PR #18790.

inspectWorkerTerminal read the dispatch authority with the stale durable
terminalHandle, so after a remint the lookup resolved nowhere and
currentHostScope was always undefined — an SSH worker with no liveness
verdict and no persisted host_scope got classified from terminal.connected
instead of unverifiable. It now reads the same effectiveHandle every other
observation in the function uses.

stopStructuredWorkerForRelease told the caller to repeat the release with
the same --retry-request, which only replays the stale release_unknown
receipt and made a structured-worker close failure permanently unretryable.
It now sources releaseUnknownRecovery from worker-release-completion so the
fresh-request-ID guidance lives in one place.

Pre-commit lint-staged (oxlint + oxfmt) run manually: clean.

* test(orchestration): exercise incarnation recovery through runtime paths

* test(orchestration): pin the incarnation read scenario to the reminted terminal

The read scenario only asserted that the call resolved, so it documented
nothing about which handle the read reached. Assert that the handle
readTerminal received resolves to the registered pane and incarnation, so
the scenario proves the read went through the reminted terminal instead of
passing on the incarnation fence's throw.

* refactor(orchestration): drop redundant incarnation prefix check; require liveTerminalHandle

* feat: add freebuff as a first-class TUI agent (#42)

<!-- orca-pr-loc -->
<!-- Programmatic LoC summary. Do not edit by hand; rewritten on every
commit. -->

| | Files | Added | Deleted | Net |
| :--- | ---: | ---: | ---: | ---: |
| Test | 0 | 0 | 0 | 0 |
| Prod | 28 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$​37 | 0 |
$\color{#1a7f37}{\Huge{\mathbf{+}}}$​37 |

<!-- /orca-pr-loc -->

## ELI5

Add Freebuff (`freebuff`) as a recognized first-class TUI coding agent
in Orca alongside Codebuff and other supported agents.

## What Changed

- Registered `freebuff` across shared TUI agent definitions,
configuration catalogs, display names, and telemetry schemas.
- Added agent icons, favicons, status mappings, and mobile asset
references for Freebuff.
- Added localization strings across supported language packs (`en`,
`es`, `fr`, `ja`, `ko`, `zh`) and updated locale translation policy.
- Documented Freebuff CLI in README agent table (`npm i -g freebuff`).

## Why

Freebuff is a CLI coding agent twin of Codebuff (`npm i -g freebuff`).
Adding it to the catalog enables users to launch worktrees, run
automated sessions, and pick Freebuff directly within Orca.

## Linked Issue

N/A

## Visual Proof

`N/A` - Catalog registration and metadata definition for CLI agent
launch; UI rendering uses existing TUI agent picker and status
components.

## Testing

- Verified TypeScript contracts, schemas, and catalog configurations.
- Tested CLI detection / agent picker integration locally on Linux
(`worktree create --agent freebuff`).

## AI Disclosure

Assisted by AI coding tooling.

## Checklist

- [x] This PR is small and focused
- [x] I explained what changed and why (including ELI5)
- [x] Before/after screenshots or videos attached for UI changes, or
`N/A` with reason
- [x] Self-reviewed for correctness, security, and performance
- [x] Cross-platform, SSH/remote, and path/shortcut impact considered
(or N/A)

---------

Co-authored-by: Lesley Murfin <lesley@revivebusiness.ca>

* test(orchestration): erase method overloads in worker reap fixtures

* test: document worker fixture type boundaries

* test: simplify worker fixture typing

---------

Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
Co-authored-by: svc-orca[bot] <313947298+svc-orca[bot]@users.noreply.github.com>
Co-authored-by: m4air <m4air@Mac.localdomain>
2026-09-21 17:23:33 -07:00
88f2f01061 fix(daemon): escape the terminal daemon into its own systemd scope so a service restart no longer kills every live PTY (#19430)
* fix(daemon): escape the terminal daemon into its own systemd scope so a service restart no longer kills every live PTY

Root cause: daemon-launched-child.ts forks the detached terminal daemon with
detached: true, which escapes the POSIX process group (setsid) but never the
systemd cgroup. Every PTY the daemon owns is itself an undetached direct
child of the daemon (native-pty-spawn.ts). Under a combined systemd unit
(Type=simple, KillMode=mixed, per docs/reference/headless-linux-server.md),
a systemctl restart/stop SIGKILLs every process still in the cgroup at the
stop timeout -- the daemon and every live terminal -- even though the
codebase already has a fully-built adoption/reattachment path for a
surviving daemon (orcad-entry.ts's refreshRestoredOrchestrationAuthority +
reconcileLegacyWorkerTerminals, gated on daemonOwnsFreshPersistentPtys()).
That path never fires today because the daemon never survives long enough.

Fix: when systemd is actually supervising the process and the OS user has a
reachable systemd --user manager (isDurableDaemonScopeSupported(), Linux
only), launch the daemon via systemd-run --user --scope so it lands in a
cgroup that is a sibling of the service unit's cgroup, not a descendant of
it. A systemctl restart of the combined unit then never reaches it. Any
failure of the scoped launch (no reachable bus, D-Bus policy rejection,
etc.) falls back transparently to the existing plain fork() launch, so
every platform/environment without this capability is unaffected.

The daemon self-detects its own resulting cgroup scope via /proc/self/cgroup
(detectOwnCgroupScopeUnit()) rather than trusting the launcher's intent, and
publishes it as cgroupUnit in its pid record and orcad's health/readiness
payload (health.terminalDaemon.cgroupUnit), so a running deployment can be
observed to confirm the fix actually engaged.

No new session registry is added: the existing daemon pid-record + adoption
protocol (publishDaemonPidFile, daemon-pid-record-quarantine.ts's
dead-record reclaim, refreshRestoredOrchestrationAuthority) already
implements durable, crash-safe reattachment for a surviving daemon -- it
was simply never exercised against a full unit restart before now.

Proven via a systemd-in-Docker recovery test: a live PTY session's shell
process, its daemon, and the daemon's cgroup scope were all confirmed
unchanged across a real systemctl restart of a Type=simple/KillMode=mixed
unit, while the main process pid changed (confirming the unit actually
restarted) and the new process's health payload recognized the surviving
daemon as adopted and live. A fresh write into the same PTY post-restart
reached the same running shell. Ordinary terminal create/work/release and
the #18789/#18790 worker-release reap-fix regression tests are unaffected.

Fixes stablyai/orca#19408

* fix(daemon): probe the real per-UID XDG_RUNTIME_DIR before trusting the process's own env

isDurableDaemonScopeSupported()/buildDurableDaemonScopeCommand() trusted the current
process's own XDG_RUNTIME_DIR env var first, falling back to /run/user/<uid> only when
that var was unset entirely. On mtl-02, orca-serve@factory.service's RuntimeDirectory=
hardening directive makes systemd export XDG_RUNTIME_DIR=/run/orca_serve/factory into the
unit's process -- a private scratch dir that shares the env var's name but has nothing to
do with the user session bus. /proc/<pid>/environ on that host confirmed exactly that path
plus DBUS_SESSION_BUS_ADDRESS=disabled:, while the real bus was reachable the whole time at
/run/user/985 (confirmed via systemctl --user is-system-running with that dir exported by
hand). The probe treated the hardened override as authoritative, found no bus socket there,
and reported unsupported on every launch -- so the cgroup-escape fix from #19408/#19430
never actually engaged on real hardware, even though tonight's factory deployment picked it
up.

Fix: resolveUserRuntimeDir() now always tries the conventional /run/user/<uid> path first
(computed independently via getuid(), never trusted from env), checking for a genuinely
connectable bus socket via statSync(...).isSocket() rather than a bare existsSync. It falls
back to the process's own XDG_RUNTIME_DIR only when that canonical path has no reachable
bus -- covering hosts that legitimately have no /run/user/<uid> at all but do have a
working bus wherever their own environment points. buildDurableDaemonScopeCommand() now
explicitly sets XDG_RUNTIME_DIR to whichever path this resolution picked, rather than
inheriting the spread env's (possibly hardened-wrong) value.

Both isDurableDaemonScopeSupported() and buildDurableDaemonScopeCommand() gained an
injectable canonicalRuntimeDir parameter (defaulting to the real computed path) so tests
can exercise the hardened-override scenario deterministically with a real, connectable
AF_UNIX socket fixture instead of the live host's actual runtime directory.

Docker's stock jrei/systemd-ubuntu test container never had this hardening directive, so
this gap was structurally invisible to the container-based verification in #19430 -- only
caught against real mtl-02 hardware.

* fix(daemon): report the daemon's own pid over the ready handshake, not systemd-run's

The launcher used to infer the daemon's identity pid from the immediate
spawned child (`child.pid`). On the durable-scope path that child is
`systemd-run --user --scope`, not the daemon, so the launcher was asserting
an identity it had no authority over.

`DaemonReadyIdentity` now carries a required `pid` populated from
`process.pid` inside the daemon itself, and `daemon-launched-child.ts` takes
`launchedIdentity.pid` from that self-report. Both sides of the
`holdDaemonAdoptionLease` pid comparison therefore originate inside the
daemon process, which is the idiom this branch already uses for cgroup
membership (`detectOwnCgroupScopeUnit` reads `/proc/self/cgroup` rather than
trusting what the launcher intended).

Note on the reported consequence: `systemd-run --scope` registers its *own*
pid on the transient scope unit and then `execvpe()`s the target command --
same pid, no intermediate process -- so adoption did not in fact fail on
systemd >= 206 (verified against systemd 255.4-1ubuntu8.17 and current main,
`src/run/run.c` `start_transient_scope()`). The fix stands on its own merits:
it removes a silent dependency on that exec-vs-fork implementation detail,
which a `systemd-run` shim earlier in PATH or any future systemd change would
have broken with no diagnostic.

`terminateLaunchedDaemonChild` was audited and deliberately left on
`child.pid`: for the same execve-preserves-pid reason that pid is either
still systemd-run mid-scope-setup (killing it correctly aborts the launch) or
already the daemon, so it targets the right process either way.

Regression coverage: `daemon-launched-child-identity.test.ts` pins the
identity source, and `daemon-ready-identity.test.ts` gains pid-validation
cases. Ready-message fixtures across the `daemon-init-*` suites were updated
for the now-mandatory field.

Addresses:
https://github.com/stablyai/orca/pull/19430#discussion_r3953722704
https://github.com/stablyai/orca/pull/19430#discussion_r3954346518

* test(daemon): assert cgroupUnit in the pid-file parse contract

`parseDaemonPidFile` returns `cgroupUnit` on every branch as of the
durable-scope commit on this branch, but five exhaustive `toEqual`
assertions in daemon-health.test.ts still described the pre-scope shape, so
they failed on the branch independently of any later change.

Adds the field to those expectations. Deliberately not relaxed to
`toMatchObject`: asserting the full parsed shape is what makes these tests
catch a field silently dropped from the pid-file contract.

* refactor(daemon): resolve the canonical user runtime dir at one point

The per-UID path cannot change for a live process, so compute it once into a module
const instead of threading the same default call through three signatures, and drop
the try/catch around a getuid() that cannot throw once it exists. Trims the module
prose to the non-obvious facts and corrects the pid-file record comment: an unscoped
daemon writes null; only records no daemon wrote are absent.

* test(daemon): clean up the cgroup-scope fixtures and assert a verdict

The cgroup fixture tracked only the file it wrote, leaking one temp dir per case.
Drains both fixture lists with splice so the pop-may-be-undefined guards go away,
and replaces a not-throw/typeof-boolean pair with the verdict it was circling:
no resolvable runtime dir means unsupported.

* refactor(daemon): share the detached child options across both launch paths

cwd, detached and stdio were repeated in the fork and systemd-run branches, which
left the two comments explaining them hovering over the env block instead. Names
them once so each branch carries only its own delta.

* refactor(daemon): validate the ready pid like every other field

typeof-first narrows the value, so the two 'as number' casts the isSafeInteger check
needed disappear and the pid guard reads like the startedAtMs guard below it.

* fix(daemon): don't retry the launch unscoped after losing the endpoint race

A scoped attempt that lost the endpoint to another daemon was retried unscoped: a
second doomed fork, a misleading 'cgroup-scope launch failed' warning, and the same
DaemonEndpointUnavailableError the caller was already going to adopt on. Rethrows it
instead, since no launch mode can win a race that is already lost.

Also drops a private alias for DaemonChildSpawnOptions and the two 'as number' casts
on child.pid in the startup-failure cleanup.

* fix(daemon): unlink the pid record by the pid the daemon published

The record holds the daemon's self-reported pid, so match on that rather than on the
immediate child's, which is the systemd-run wrapper's until it execs.

* fix(daemon): route the scope launch through the child-process chokepoint

The two files this PR added imported `node:child_process` directly, which
`child-process-import-boundary.test.ts` fails on deterministically: the
offender count went 155 -> 157 against a pin of exactly 155. Raising the pin
or listing the files is what that test explicitly forbids, and the allowlist's
own note says a split "moved the import, it did not add one" -- so the fix is
to get both new files off the module and put the count back at 155.

- `daemon-cgroup-scope.ts`: the `systemd-run --version` probe now uses
  `runProcessSync` instead of `execFileSync`, so it gets the shared spawn
  decisions. Kept synchronous deliberately: `launchDaemonChild` attaches the
  readiness listener in the same tick it is called, and an await before the
  spawn moves the child past that tick. A non-zero exit is data rather than a
  throw here, so the verdict now checks `code === 0 && !timedOut`.
- `daemon-launched-child-spawn.ts`: the scoped launch uses `spawnProcess`, and
  the long-standing unscoped launch keeps `fork` semantics through a new
  `forkProcess`.
- `src/shared/child-process/fork-process.ts`: the fork arm of the chokepoint.
  `spawnProcess` cannot express a Node child with an IPC channel started from
  a module path under an overridden `execPath`, and the existing launch tests
  are written against `fork`'s contract, so a spawn rewrite would have changed
  module resolution, `execPath` and `execArgv` at once. It passes
  `windowsHide: true` -- the flag every other call site in that directory
  sets, reachable via an assertion because `ForkOptions` omits it -- which
  keeps `windows-console-visibility.test.ts` at its pin of 65 too.

Both ratchets pass with both pins and both allowlists untouched.

Docs: `orcad-operations.md` and `headless-linux-server.md` still described the
limitation this PR removes as permanent. Both now describe the durable-scope
survival path and its preconditions (systemd as PID 1, a reachable user bus /
`loginctl enable-linger`, `systemd-run` on PATH), and scope the old text to
the unscoped-fallback case, pointing at `health.terminalDaemon.cgroupUnit` as
the way to tell the two apart on a running host.

* fix(daemon): seal the cgroup capability probe from the host and correct KillMode=mixed docs

The capability probe consulted the host's own /run/systemd/system marker and
spawned the real systemd-run binary, so the hermetic unit tests could only pass
on a systemd host (and fail closed otherwise, even with faked bus sockets).

- Thread systemdBootPath and runVersionProbe as test seams through
  isDurableDaemonScopeSupported, defaulting to the real boot marker and
  systemd-run --version probe in production.
- Narrow the injected probe to the ProcessResult slice it consumes.
- Cover: no-systemd-boot, non-zero probe exit, and probe-timeout cases.
- Correct KillMode=mixed semantics in the docs: the cgroup-wide SIGKILL fires
  the instant the main process exits, not after TimeoutStopSec; document the
  Docker-container caveat and add KillMode=mixed to the multi-service template.

* fix(daemon): satisfy assertion checks in scoped launch

* fix(daemon): satisfy anti-slop and console guards

* test(serve): update shutdown docs assertions for daemon scope

* fix(daemon): migrate adopted legacy scopes

* docs: qualify restart safety by daemon scope

* docs(daemon): qualify Upgrade restart prose with durable scope caveat

Align the Upgrade section in docs/reference/headless-linux-server.md with
the earlier preservation section and docs/reference/orcad-operations.md:
a service restart terminates live processes only when running under the
unscoped fallback, and stops should be treated as destructive unless
health.terminalDaemon.cgroupUnit names an orca-daemon-*.scope.

Update the shutdown workflow test assertion in
config/scripts/headless-serve-shutdown-workflow.test.mjs to match.

* fix(daemon): harden legacy scope migration

---------

Co-authored-by: Lesley Murfin <260182349+LesleyMurfin@users.noreply.github.com>
Co-authored-by: m4air <m4air@Mac.localdomain>
2026-09-21 17:23:30 -07:00
Brennan Benson 15472cd4c6 feat(native-chat): keep restart recovery available in status bar (#21397)
* feat(native-chat): keep restart recovery available in status bar

* fix(native-chat): source the restart offer from the host and retire it on recovery

Closing the reconnect dialog spent the durable recovery offer, so looking around
before deciding lost the recovery for good. The offer now survives a close, and
the status bar carries it — but a durable offer needs a way to die, and it only
had a reconnect, an explicit dismiss, and a 24h expiry.

The claim's launch-scoped lifecycle moves into its own collaborator, which splits
what the host ADVERTISES from the evidence it holds. A resume-capable hold that
hands a marked chat its provider child back is the recovery the offer existed to
perform, so it stops being advertised and stops being written back at quit, while
the marker stays valid evidence — a user who reopened a chat can still ask the
agent to carry on. Teardown re-derives the snoozed offer rather than round-tripping
raw markers, and this teardown's own witness now outranks the stale claim for the
same chat instead of being overwritten by it, which was silently persisting an old
turn id and making the next launch refuse the chat that was actually mid-turn.

On the renderer the candidate list gets its own producer against
agentSession.restartResumable, so the status entry and the dialog read one
host-owned answer instead of the dialog pushing its local state at a sibling. The
entry re-reads the host before reopening, so a reopened list can never name a chat
the host would now refuse; dialog open becomes the external one-shot request
rather than a flag mirrored into render state, which is what let a reopen replay
the launch answer and re-offer chats already reconnected. Dismiss all is quiet
rather than destructive, saves the preference like every other exit, and reports a
write the host never confirmed instead of trapping the dialog open.

* fix(native-chat): keep a durable offer a launch never read, and settle the one a continuation spent

Teardown replaced the recovery capsule with whatever this launch still owed,
and a launch that never read the offer owes nothing — so a quit after a
failed first read, a disabled flag, or a window that never mounted deleted a
recovery the user was never shown. The write-back now distinguishes "claimed
and still owed" from "never claimed": the first is re-derived as before, the
second carries forward verbatim, because nothing revealed those sessions and
the predicate would refuse every one for want of a journal nobody opened.

Reconnect and continue spent the same claims Reconnect does but never shrank
the offer, leaving the status bar counting chats the host had already handed
back and sending the user to an entry that re-reads, finds nothing and does
nothing.

* fix(native-chat): stop a teardown answering for an offer it could not read

Two ways the write-back deleted a durable recovery offer nobody had seen.

A take that FAILED left the claim holding an empty list and reporting that
this launch had answered for the offer. The markers were still on disk,
unread and unknowable, and teardown then overwrote them with its own empty
list. It now writes nothing at all unless it has a witness of its own.

`owed()` read "has the capsule been touched" where it meant "did anything
here LOOK at the offer" — and its own write-back read counted. Teardown is
retried when a phase fails, so the second attempt re-derived carried markers
against a session map eviction had already emptied, refused every one, and
wiped what the first attempt had just carried forward. The flag is now set
only by the paths that actually read or act on the offer.

The mock guard for the carry could not fail: it indexed the session it
claimed nothing had revealed, so re-deriving passed and the verbatim carry
was never the reason it went green. It now runs against no indexed session,
which is what an unread offer looks like.

Also drops the `Not now` row from the preference table, where it was paired
with a dismiss method it no longer calls, and asserts the same thing where
the snooze is already covered. Splits the marker predicate's journal reader
out of the resume host, which was at its line ceiling.

* fix(native-chat): clear the corrupt recovery capsule the take refused

A capsule whose contents no longer parse made take() throw before it ever
reached the clear, so the bad file survived every launch. Nothing else
rewrites it now that a teardown owing nothing readable declines to write, and
the freshness filter runs after the parse, so the 24h window could not release
it either: one corrupt file refused recovery forever.

Clear it inside the same transaction that failed to read it, then rethrow, so
the poison dies on the next launch while callers still see why the take failed.
A clear that fails is swallowed rather than allowed to mask the parse error.
Refusing to expose partial candidates is unchanged, and a read that fails for
any other reason still writes nothing.

* feat(native-chat): make resume the one restart action, and make it actually resume

The restart prompt offered two actions: "Reconnect all", which reattached
and sent nothing — exactly what opening the chat already does — and
"Reconnect and continue", which reattached and asked the agent to carry
on. The vacuous one is gone, the "Not now" button and the info popover
with it, and the feature is now called resume throughout.

"Don't ask again (resume automatically)" now runs the action the button
runs: the launch calls agentSession.restartContinue instead of
agentSession.restartResume, so the preference means what it says. Several
comments asserted the opposite as a structural guarantee and are
corrected. agentSession.restartResume stays: no in-app caller is left,
but it is a published wire method a non-desktop or older client can call.

* fix(native-chat): label the resume button with the number of chats selected

The button read "Resume all" whenever every chat happened to be ticked,
which described the selection rather than the action. It always acted on
the selected chats only. Now it always names that count, with a singular
variant so one chat does not read "1 chats".

* refactor(native-chat): drop the reconnect vocabulary the resume action left behind

Resuming became one action — reattach and ask the agent to carry on — so the
notification helpers no longer need to be told which action they are reporting.
Every caller passed `continue`; the `reconnect` branch, its helper and its
catalog keys are gone.

The dialog and the launch path had grown two copies of the same call: same RPC,
same response shape, same announce-and-settle. That now lives once in the store
module that owns the offer, which also takes the dismiss call, leaving the modal
presentational. The two copies had drifted — only the dialog's caught a
malformed payload — and the unified one keeps the defensive reading.

No behaviour change. `agentSession.restartResume` stays: it is a published wire
method even though nothing in the app calls it.

* refactor(native-chat): derive the resume selection instead of intersecting it

The modal's selection was intersected back against the host's candidate list
before every action, as a guard against naming a chat the host never offered.
That guard could never fire: the selection was already derived from that same
list, so the intersection was the identity. The array of chosen ids is now the
derived value and the lookup set falls out of it, which makes the property
structural rather than checked. The helper had no other caller and is gone,
along with its three tests.

Three tests mocked the resume response in the shape the old API returned. Two
never reached that branch at all; the third only passed because the unreadable
shape happened to exercise the malformed-payload path. All three now use the
real shape, and the malformed-payload behaviour — report an unconfirmed
delivery, leave the offer standing — gets a test that says so.

Also: the candidate reader took two trailing optional parameters, so one caller
passed a placeholder `false` to reach the second; they are an options object
now. `isFolderWorkspaceId` had no caller outside its own module and is no
longer exported. `RestartActionOutcome` only ever describes a continuation row,
so it is named for that. `dismissAll` set a busy flag that nothing could
render, since it closes the dialog first. Several comments repeated an argument
already made in the module they point at.

Settings: the automatic-resume description is one sentence again.

No behaviour change.

* fix(native-chat): make restart recovery explicitly durable

* fix(native-chat): preserve dismissal fence across new interruptions
2026-09-21 16:38:05 -07:00
Brennan Benson cd59678394 refactor(agent-launch): assemble host startup-plan inputs in one resolver (#22082)
* refactor(agent-launch): assemble host startup-plan inputs in one resolver

buildAgentStartupPlan was already one shared implementation, but every host
re-derived its argument object by hand from the same four settings
(agentCmdOverrides, agentDefaultArgs, agentDefaultEnv, terminalWindowsShell),
and the copies had drifted.

resolveAgentStartupPlanInputs owns that assembly. What genuinely varies per
launch stays a parameter: the host (platform, isRemote), a requested shell, the
per-launch agentArgs override, and the picked session options.

Fixes a live divergence on the agent.launch path: orca-runtime-create-agent-session
passed sessionOptions without sessionOptionsOverrideAgentArgs, so a configured
`--model` in agentDefaultArgs reached argv alongside the picked model and won on
argv order, while the same launch through worktree.create honored the pick.
The plan also reported no applied sessionOptions, so the chat surface could not
name the model the user chose.

Migrates the four host sites; the eleven renderer sites are unmigrated and still
assemble their own inputs.

* fix(agent-launch): preserve picked options in draft launches

* test(agent-launch): assert draft option precedence
2026-09-21 16:15:44 -07:00