Commit Graph
2319 Commits
Author SHA1 Message Date
Neilandcarlosbaraza 15baf86660 fix(agents): honor environment prefixes in generation commands (#22427)
* fix(agents): preserve environment prefixes in generation commands

Adapted from the proposal by @carlosbaraza.

Co-authored-by: carlosbaraza <carlosbaraza@users.noreply.github.com>

* test(agents): respect Windows environment key normalization

---------

Co-authored-by: carlosbaraza <carlosbaraza@users.noreply.github.com>
2026-09-25 20:50:39 -07:00
Neilandbbingz 14087c8e32 Accept repeated leading BOMs in agent hooks (#22414)
Adapted from the investigation and proposal by @bbingz.

Co-authored-by: bbingz <bbingz@users.noreply.github.com>
2026-09-25 20:50:15 -07:00
OrcaWinandm4air 2ed6505a41 fix(native-chat): preserve current pane ownership through toggles and restore (#23049)
* fix(native-chat): persist current pane ownership across lifecycle events

* fix(native-chat): retain ownership when client chat rendering is disabled

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-25 20:46:24 -07:00
Brennan Benson c5f33bd139 fix(ipynb): run no workspace interpreter until the notebook is trusted (#22962) 2026-09-25 20:08:24 -07:00
OrcaWinandm4air ff74506c0b fix(native-chat): keep terminal pane chat ownership stable (#22984)
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-25 19:40:32 -07:00
8846987c99 feat(rate-limits): add Cursor usage tracking (#22633)
* feat(rate-limits): add Cursor usage tracking

## ELI5

If you use Cursor, Orca now shows how much of your monthly Cursor plan you
have used, next to the Claude, Codex and Grok meters, and in Settings →
Accounts. It reads the sign-in Cursor already saved on this computer and never
changes it.

## What changed

Cursor becomes a rate-limit provider like Grok: a status-bar meter (default-on,
with its own toggle), a row in the usage roster, and a Settings → Accounts
section naming the signed-in account.

The credential is read from whichever of three stores has it, first match wins,
all read-only:

- the macOS login keychain item `cursor-access-token` / `cursor-user`, which is
  where `cursor-agent` 2026.06+ keeps the session;
- `~/.cursor/auth.json` and its platform variants, used by older CLIs;
- the Cursor IDE's `state.vscdb` (`cursorAuth/accessToken`), for people who
  never run the CLI.

The keychain entry is the one current CLIs use, and reading only `auth.json`
finds nothing on an up-to-date macOS install. A locked keychain cannot mask a
readable `auth.json`, and a locked `state.vscdb` cannot mask either.
`~/.cursor/cli-config.json` supplies the account's email and display name; it
never holds a token.

Usage comes from the dashboard route the Cursor web dashboard itself reads,
because Cursor documents no individual-user usage API — every documented API is
team- or Enterprise-scoped. Per Cursor's pricing docs an individual plan has two
pools, Cursor Models and Other Models, both resetting with the billing cycle,
plus optional on-demand spend; each becomes a named bucket. The headline
percentage prefers `used / limit` over the sibling percentage fields, which are
pre-rounded for the dashboard's own copy. Because the route is undocumented the
mapping is defensive: an unrecognised payload resolves to `unavailable` and
hides the bar rather than publishing a zero that reads as "no usage".

Orca never runs `cursor-agent login` and never writes, refreshes or rotates a
Cursor credential. An expired token short-circuits to an actionable
"run cursor-agent login" instead of spending a request that can only 401 — not a
rare case, since `cursor-agent status` still reports `isAuthenticated: true`
against a token that expired months ago.

## Why this shape

Six open PRs implement this feature and none reads the keychain, so each finds
nothing for a large share of users; this takes the auth layer further and keeps
what those PRs verified live. The bar is not gated on `cursor-agent` being on
PATH, unlike other CLI providers, because an IDE-only session is real usage with
no CLI to detect.

`readKeychainPassword` moved out of the Claude keychain reader into
`src/main/macos-keychain/generic-password.ts` so both providers share one
`security(1)` wrapper. It is a byte-for-byte relocation, so Claude's credential
path is unchanged; the two child_process allowlists move the entry with it and
neither ratchet count changes.

Co-authored-by: Preschian Febryantara <preschian@users.noreply.github.com>
Co-authored-by: Qwesdy <qwezdi@proton.me>
Co-authored-by: ivo922 <github.concur614@passmail.net>
Co-authored-by: Mihail Vratchanski <mivrkiki@gmail.com>
Co-authored-by: Tauri-EPO <enrico.pin@gmail.com>
Co-authored-by: Raajik <44516546+Raajik@users.noreply.github.com>

* test(rate-limits): name the JWT helper's segment type in the Cursor tests

The anti-slop gate rejects a bare `object` parameter; the fixtures build a
claims record, so say that.

* fix(rate-limits): render Cursor's pools and keep its plan total visible

Review of the first commit found the meter effectively blank for a healthy
account, which the screenshots missed because the only Cursor session on hand
had expired and never reached the success path.

- The verbose status-bar segment filtered buckets through an allowlist written
  for Gemini's experimental models, so both Cursor pools were dropped and the
  fallback needed a `session` window Cursor never reports. A signed-in account
  rendered an icon and no number. The allowlist now admits Cursor's pools, and
  the fallback accepts a monthly window.
- `getWindowSections` dropped `monthly` whenever buckets existed. Cursor puts
  the plan total there and its sub-pools in buckets, so a plan at 92% showed as
  50% in the roster, the tooltip, and the tightest-usage pick.
- A plan reporting `enabled: false` still published its 0% pools, painting a
  healthy meter for a pool the account does not own and skipping the
  request-quota fallback.
- `redirect: 'error'` turned the dashboard's bounce to /login into a generic
  network failure, hiding the actionable sign-in message.
- A busy `state.vscdb` (the IDE holds it open) surfaced as a provider error,
  which would pin an alert bar on Cursor IDE users who never set Cursor up in
  Orca. It falls through to "no credential" instead.
- Refreshing the Accounts section read the keychain twice for one update.

* fix(rate-limits): pin the platform in the Cursor keychain tests

Review caught three cases that assumed macOS: the keychain source is behind an
explicit `process.platform` check, so on the Linux CI runner the mocked read was
never reached and the tests read the CLI file instead. They now set the platform
they mean, and two new cases assert the off-macOS fall-through.

Also track the credentials reference doc (docs/** is ignored by default, so a
new reference needs its own allowlist entry) and give the visibility fixtures
their own provider id instead of Grok's.

* fix(rate-limits): prefer a live Cursor session and report a failed refresh

Review round two, from CodeRabbit and Pullfrog.

- Credential precedence returned the first token that parsed, so an expired
  keychain token in front of a fresh Cursor IDE session reported "sign-in
  expired" on every poll while a usable session sat one source below. A live
  session now wins; the expired one is returned only when nothing live exists,
  so the actionable message still appears in that case.
- The usage schema took `.optional()` where the route sends `null` for an absent
  sub-object, so one null pool failed the parse for the whole body and threw
  away valid pools and the billing cycle with it.
- Cursor usage could survive an account switch: a failed refresh for account B
  kept account A's figures beside B's name in Accounts. The snapshot now carries
  a hashed account fingerprint, and a known-and-changed identity clears the
  previous reading. A refresh that names no account still keeps its own.
- The Accounts section rendered nothing at all when a signed-in account's fetch
  failed, and could repaint an older account when two status reads overlapped.
  It now states the failure — beside the numbers when a stale snapshot remains —
  and ignores superseded reads.
- A web client claimed "not signed in" for a host it cannot read, contradicting
  the meter beside it; it now says the detail is host-only.
- Signed-out copy named `cursor-agent login` as the only way in, though an IDE
  sign-in works just as well.
- The census comment ended at 4219 after the pacer squash without naming the two
  modules #22616 added; recorded them, re-measured on a clean origin/main.
- Narrowed the docs claim: Cursor documents all-plan APIs, but no individual
  usage endpoint.

* fix(i18n): localize the web client's Cursor host-only notice

It reaches the Accounts pane like any other string, so the coverage gate is
right to want it in the catalog rather than allowlisted.

* fix(rate-limits): name the Cursor account on failed refreshes, and ship the reworded copy

Review round three. Both findings say an earlier fix did not actually take.

- The account-switch guard reads `authProvenance` off the fresh result, but the
  fetcher stamped it only on success and network failures. The `stale-token`,
  429, 5xx and parse results omitted it, and so did the expired-session branch —
  so a switch whose first refresh failed, which is precisely the case the guard
  exists for, still rendered the previous account's figures under the new name.
  Every failure holding a readable session now names its account; a missing or
  unreadable credential still names none. The service test also fed a result
  shape the fetcher never produces, so it proved nothing; it now uses the real
  stale-token shape, and the fetcher test asserts provenance across 401/429/5xx
  and expiry.
- The reworded signed-out copy never rendered: a present catalog value beats the
  `translate()` fallback, and `sync:localization-catalog` only adds missing keys
  rather than updating changed defaults. Updated both strings in en.json, which
  also prunes them from the runtime-required catalog now that they match.

* docs: keep the Cursor credentials reference out of the tree

Its content lives in the PR description instead; docs/** stays ignored rather
than gaining an allowlist entry for this branch.

* test(mobile): drop the census note main no longer pins

main removed `SESSION_ROUTE_MODULES` and re-pinned this lane on a different
count, so the paragraph this branch added documents a number series that is
gone. The branch touches nothing in this file now.

---------

Co-authored-by: Preschian Febryantara <preschian@users.noreply.github.com>
Co-authored-by: Qwesdy <qwezdi@proton.me>
Co-authored-by: ivo922 <github.concur614@passmail.net>
Co-authored-by: Mihail Vratchanski <mivrkiki@gmail.com>
Co-authored-by: Tauri-EPO <enrico.pin@gmail.com>
Co-authored-by: Raajik <44516546+Raajik@users.noreply.github.com>
2026-09-25 18:54:11 -07:00
1c2cf120e3 fix: stop process-tree loops (#22411)
Based on the report and proposal by @brynnclaw.

Co-authored-by: brynnclaw <brynnclaw@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-25 18:35:48 -07:00
Brennan BensonandClaude 16784c1a67 fix(native-chat): name a chat write by its target, not the owner generation (#22812)
* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* fix(native-chat): name a chat write by its target, not the owner generation

A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.

Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.

Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.

* docs(native-chat): say mutation admission checks only the writer lease

* docs(native-chat): drop the send rebase from comments that still described it

* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent

* docs(native-chat): drop the fence from the admission the send effects run behind

* docs(native-chat): give the fence move on release the reason that still holds

* docs(native-chat): stop citing a write fence check in launch and mailbox comments

Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-25 18:31:29 -07:00
Brennan Benson 18bbf6f209 refactor(orchestration): resolve every caller and target to one orchestration party, keyed by the Orca session id (#22555)
* refactor(orchestration): resolve a session caller at the dispatch entry and bind it by actor

WIP: entry resolver on both dispatchers, caller identity through run scope,
actor-keyed Run binding and unbind, actor writes on bind/create/assign, and
actor-aware mail ownership exclusions.

* test(orchestration): pin actor-keyed Run binding, stale-actor precedence and actor mail ownership

Keeps the non-session dispatch path synchronous so terminal and session-tab
streams reach their handler without an extra async hop.

* feat(orchestration): resolve session callers before params parse and pin every verb on both routes

A session caller need not name itself in a param that requires a caller: the
entry binds the declared caller to the session before the schema runs. Session
refusal codes pass through the RPC error map, and DB row reads added here carry
their SAFETY rationale.

* test(orchestration): pin the SSH check's pane through the caller-identity lookup

* test(orchestration): pin a Run-less session's direct check and receipt binding without a caller param

Drops the actor clause from self-dispatch detection: a creator and assignee can
only share an actor when they already share a handle or pane.

* test(orchestration): pin a session's own Run for plain and group sends and the assignee-only mail sweep

* test(orchestration): name the party-naming field population for its role

* fix(orchestration): clear a worker actor an older binary's unbind leaves, and refuse a worker without its identity

An older binary unbinds a structured worker's Run by clearing handle and pane, which
leaves the actor looking like a handle-less chat binding. The every-open repair
clears that shape for actors recorded as structured workers only, and the resolver
refuses a worker session whose worker identity is gone, so this binary never
writes the shape itself.

* refactor(orchestration): read a Run's coordinator actor through its generation, and give a party's addresses one owner

The coordinator actor now counts only at the consumer generation it was written at, so a
Run binding matches a session by that rule alone. It replaces two mechanisms for the same
fact: the rule that an actor beside a handle it did not bind with never matches, and the
open-time repair that cleared a structured worker's actor an older binary's unbind left.
Every write of an older binary that rebinds or unbinds bumps the generation, so both
shapes stop counting by themselves, including a chat's actor after a rebind then an
unbind, which neither old mechanism caught. createRun and the same-coordinator actor
correction write the generation in the statement that writes the actor. The resolver
still refuses a structured worker whose worker identity is gone; its predicate moves
next to the worker identity lookup.

addressSpellingsOf is the one owner of the addresses a party is reachable at (a
structured worker's handle and session actor). createRun, bindRun, the coordinator
unbind and the declared-caller check use it instead of hand-built sets, and each has a
test at both of a worker's addresses.

* fix(orchestration): say a released session is not running, and scope the pane-key credential claim to requests without a session

A released lease is evicted, not ended: a user turn resumes the session, so the refusal
now says it is not running right now instead of that it has ended.

The worker pane-key comment claimed the random leaf is what stops anyone who learns a
session id from acting as the worker. On the same-host socket route the session id now
names the worker with no token by design; the pane key still matters where a request
names no session (a PTY agent's, or the paired-client route, which refuses session ids).

* test(orchestration): pin the same-coordinator actor correction's generation write on a row an older binary wrote

* refactor(orchestration): resolve session callers by the bare Orca session id, typed apart from its address

Carries the Orca session id rename into caller resolution, Run binding and Dispatch
creation. The caller identity holds the bare `orcaSessionId`; the `session:<id>`
spelling is derived by formatOrcaSessionAddress wherever mail needs it.

`OrcaSessionId` and `OrcaSessionAddress` are distinct branded strings. Only
isOrcaSessionId and parseOrcaSessionAddress produce an id, and only
formatOrcaSessionAddress produces an address, so comparing the two is a type error.
The Orca session id columns on the row types carry the id type.

Every reader that compares a stored id with a mail address now compares like with
like: the active-Dispatch ownership check formats the stored id
(orcaSessionAddressSql), and the stray-mail sweep, the creator nesting lookup and the
recorded-worker check bind a parsed or typed bare id. The Run-mailbox ownership
check keeps its existing handle comparison beside the session one.

* fix(orchestration): accept a worker's ask to every address its coordinator is reachable at

A worker's preamble names its coordinator as `session:<id>` when the coordinator is a structured session, but ask only accepted `run:<id>` or the coordinator's terminal handle. A chat coordinator has no handle, and a coordinating structured worker has two addresses, so ask --to the session address was refused as dispatch_run_mismatch. The check now takes the Run's current coordinator addresses from addressSpellingsOf(runCoordinatorKey(run)).

* refactor(orchestration): require every caller-identity entry point to be handed the resolved session

The resolved session parameter was optional on resolveRunScope, resolveOrchestrationCaller, orchestrationCallerIdentity, resolveDispatchCreator and resolveDispatchCallerWorktreeId, so a method that forgot to pass it would compile and silently treat a chat's session address as a terminal handle. It is now required and typed `OrchestrationSessionCaller | undefined`, so leaving it out is a type error. Every call site already passed it; no behavior changes.

* fix(orchestration): deliver mail sent to a session address to the mailbox that session reads

A send to session:<id> was resolved like a terminal handle: no live pane, so a chat's
current Run was missed (two Runs read as ambiguous), a Run-less chat was refused though
it reads its direct mailbox, and a structured worker's session address never reached
its Dispatch. Resolve a worker's session address as its handle, a chat's by its bound
Run, and fall back to the chat's durable direct mailbox while it runs on this host.

* refactor(orchestration): resolve every party through one resolver with one mailbox address

A structured worker is reachable at its handle and at its session address, and a chat only at its
session address. Callers and targets were each compared or resolved at their own site, some against
one spelling and some against every spelling, and the sites that did neither refused or misrouted.

Add orchestration-party: resolveOrchestrationParty (and resolveOrcaSessionParty for a bare id) is
now the only place an address becomes a party. The session caller, the declared-caller check,
ask's target and inbox's filter all resolve through it, so caller and recipient resolution cannot
disagree, including on a recorded worker whose identity this host lost. Every session-to-party step
passes through canonicalOrcaSessionId, the seam later lineage canonicalization plugs into.

mailboxAddressOf replaces addressSpellingsOf: a party has one mailbox address (a worker's handle
today), and Run binding remembers and reroutes only that.

A request with no session id that declares a session address as its caller now gets the party it
names: a worker's handle, as if it had named it, or a session_caller_chat_not_declarable refusal
for a chat, which is identified only by the session id its own environment sends.

* fix(orchestration): route dispatch and mail targets through the party they name

dispatch --to a worker's session address stored that address as the assignee handle, so the worker,
which reads by its handle, never saw the Dispatch. The assignee now resolves through the party
resolver: a worker's either spelling assigns its handle and records its Orca session id. A chat
cannot receive a dispatch yet, so dispatch --to a chat is refused with
session_chat_not_dispatchable before any row is written.

Recipient routing resolves the party once and, for any session-backed party, finds its Run by its
durable binding rather than by a live pane. A structured worker that coordinates a child Run while
assigned in its parent got mail at its parent Dispatch mailbox whenever its session was evicted,
which it never reads while bound to the child Run. Terminal handles keep the same live-pane lookup.

* refactor(orchestration): delete the SQL that matched a worker's second spelling

With every caller and target resolved to one mailbox address, no writer stores mail under a
structured worker's session address: sends resolve it to the handle, a declared session caller is
rewritten to the handle or refused, replies answer a stored canonical sender, and questions,
answers, escalations, federation and legacy mail write run:, dispatch: or handle addresses (legacy
rows are legacy_direct, which these queries never read). The branches that matched that spelling
were unreachable and are removed:

- the session-address OR in activeDispatchOwnsAddressSql, back to the assignee handle alone;
- the session branches in routeForeignDirectMessagesToOwnedMailboxes and
  findActiveDispatchForDirectMessageOwner, which return to the base-branch form.

Tests that inserted such rows directly now pin the behaviour through the verbs: both spellings of a
worker land in the one mailbox it reads.

* test(orchestration): pin one canonical address per party across every caller and target

- every target param in ORCHESTRATION_TARGET_PARAM, for a worker's handle and session address and
  a chat's session address: send, ask, dispatch (a chat refused with no row) and inbox;
- a declared session-address caller on a request with no session id: a worker acts as its handle,
  a chat is refused;
- a worker that coordinates a child Run while assigned in its parent gets mail in the child Run
  with no live pane;
- no mail writer stores to_handle or from_handle as a worker's session address;
- the caller a session id resolves to and the recipient its address resolves to agree in every
  party state, and every session-to-party step goes through canonicalOrcaSessionId.

* refactor(orchestration): drop the session caller's chat-to-terminal-view handoff remnants

The structured-chat terminal handoff is gone from the base branch: a session is owned natively only.

- the lease rule no longer speaks of either owner or a handoff keeping identity;
- an in-progress owner change (new-owner-proving, recovering, manual-recovery, which the base branch
  still produces) is refused as changing owners, not as switching between chat and terminal view;
- the native to terminal-view to native identity test is deleted, and the terminal-evidence test
  no longer describes that evidence as a terminal view's.

* fix(orchestration): read a declared caller param by typed access, not Reflect.get
2026-09-25 18:01:40 -07:00
Brennan BensonandClaude 6627c6503c refactor(agent-session): keep which conversation each chat tab shows in one host table (#22709)
* refactor(agent-session): keep which conversation each chat tab shows in one host table

The host now persists one table in the agent-session store, from chat tab id to
the conversation that tab currently shows. It replaces both the visible-session
list and the per-record surfaceTabId, so there is one answer to "which chats have
a tab, and under what id".

- /clear moves the tab's entry from the old conversation to its replacement in
  the same transaction that commits the clear. No record carries a copied tab
  id, so a tab id names one conversation by construction, and the lookup from a
  conversation to its tab returns at most one id.
- A create that reserved a tab id claims it in its reservation transaction;
  uniqueness is a key check on the table. A create that fails releases its
  claim, and hiding a chat frees its id.
- Showing a chat with no entry gives it today's id, structured-agent-session-<sid>,
  unless a cleared chat's tab kept that id; then it gets a fresh one.
- A close that does not land puts the tab back under the id it had.
- Stores written by older builds are seeded on read from the visible list and,
  for chats that have one, the record's surfaceTabId. That field is no longer
  part of the record type, is never written, and is read only by this seed when
  the file has no table. The visible list is still written, derived from the
  table, for older builds.

Tab snapshots, status keys and worker pane keys are unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(agent-session): take a reserved chat tab id when its tab is published

A create that reserved a tab id put it in the tab table at reservation, and table membership is
visibility. A create that stopped before its tab was published (a crash, or a failure after the
provider started) left an entry that the next launch either restored as a tab nobody asked to see
again or kept forever with no way to close it. Reservation now only refuses an id another chat's
tab holds; the id is taken when the tab is published, so there is nothing to release on failure.

The create reply now reads the tab id after publishing, so an unreserved create answers with the
id its tab was given, as it did before the table, and agrees with a replay of the same create.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(agent-session): seed a chat cleared before the upgrade under its first tab id

A chat cleared on an older build shows a later conversation of its /clear chain in the tab that was
opened for the first one, and clients key that tab, its read state and its status by the first
conversation's id. Seeding gave it the latest conversation's derived id instead, so the stored id
disagreed with the one clients hold and with what a /clear on this build leaves behind.

Seeding now follows the source records' committed clears back to the chain's first conversation
and gives the chat showing the chain that conversation's id. Those chats seed first, so a cleared
conversation reopened from history takes a fresh id when its own is held. A seeded chat is no
longer dropped when every candidate id is taken, and the table keeps the visible list's order.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(agent-session): give a reopened cleared chat the same tab id at runtime and at seeding

A cleared conversation reopened from history got a random tab id at runtime but a
deterministic `-reopened` id when the table is seeded from an older store. One rule now
serves both, so a table dropped by an older build and seeded again gives that chat the
id it already had.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-25 17:30:27 -07:00
Brennan BensonandClaude add99c908b fix(native-chat): send typed question answers as structured answers, not an option id (#22793)
* fix(native-chat): send typed question answers as structured answers, not an option id

A typed "Other" answer was packed into the `optionId` of
agentSession.respondToQuestion, a field capped at 1024 characters, so a
long answer failed with "Invalid option id" and never reached the agent.

respondToQuestion now carries per-question `answers` in their own field,
bounded like a typed answer, and a host advertises
agent-session.question-answers.v1 when it takes them. Clients fall back to
the packed option id for older hosts. The host reads either form once into
a typed response, records the structured answers on the resolution (and
keeps the packed form older clients read), and the Claude and Codex
adapters build their reply from the typed answers before the journal
commits, so an answer the agent cannot take is refused rather than
recorded unanswered.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): hold one answer per single-select question card

Typing an answer deselects a picked option, and picking an option leaves the typed
text in the field without sending it, so the card never shows two answers while
sending one. Multi-select still sends picked options and typed text together.

* fix(native-chat): keep keyboard tabbing from re-choosing a typed answer; accept untrimmed question ids

Clicking or typing in the answer field chooses the typed answer; focus alone
no longer does, so tabbing to Submit keeps the option the user picked.
A question id is matched exactly by the host, so the wire no longer rejects
agent-written ids with edge spaces, which older builds accepted.

* fix(native-chat): choose the typed answer on click so a disabled or scrolled field cannot

* test(native-chat): cover pointer events on a disabled answer field

* refactor(native-chat): record the typed answer as a choice in the question card

Choosing the typed answer is now an entry in the question's selection, set by typing
or clicking the field and replaced by picking an option, instead of being inferred
from an empty selection. Unpicking an option no longer silently chooses kept text.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-25 15:54:46 -07:00
mmarabelandNeil 7d2c399329 fix(web): keep Remote Web loading over plain HTTP without crypto.randomUUID (#22516)
* fix(web): keep Remote Web loading over plain HTTP without crypto.randomUUID

Browsers hide crypto.randomUUID outside secure contexts, so Remote Web over
http://<lan-or-tailnet-ip> threw while importing the store and never painted.
createAgentStatusAuthorityId now takes its UUID source (renderer passes
createBrowserUuid, main passes node:crypto randomUUID), and the other
unguarded renderer calls go through createBrowserUuid.

* refactor(renderer): route remaining randomUUID fallbacks through createBrowserUuid

Replaces five hand-rolled crypto?.randomUUID?.() fallbacks (including a copy
of the browser-uuid fallback in mint-stable-pane-id) with createBrowserUuid,
and adds an oxlint no-restricted-properties rule so renderer code cannot call
randomUUID directly again.

* refactor(shared): move the non-secure-context UUID generator into src/shared

The white screen came from src/shared, so the fix belongs there. src/shared had
three hand-rolled copies of the same randomUUID-then-getRandomValues-then-Math.random
ladder (nested-repo-telemetry, project-groups, setup-agent-sequencing) because there
was nothing in that layer to import; createBrowserUuid lived one directory over in
the renderer.

createNonSecureContextUuid() now holds the single implementation, @/lib/browser-uuid
re-exports it under the renderer's existing name so no renderer import site changes,
and the three duplicates call it.

That also lets createAgentStatusAuthorityId go back to one argument. The injected
randomUuid source was justified as keeping browser APIs out of shared code, but this
generator is runtime-agnostic — it works unchanged in Node. Injecting it bought no
layering and made the safe choice a parameter every future caller had to get right,
unguarded: a caller could pass () => globalThis.crypto.randomUUID() and restore the
white screen with lint and tests green.

* fix(lint): ban crypto.randomUUID in src/shared and scope the escape hatch

vite.web.config.ts compiles src/shared straight into the web bundle, but the new
randomUUID ban only covered src/renderer/src — so the exact module that white-screened
the app sat outside the guard it shipped with, and the regression could come back with
a green lint. The override now covers src/shared/**/*.ts too; it costs zero diagnostics
because the duplicates it would have flagged are gone. `import { randomUUID } from
'node:crypto'` is untouched, so main-only shared modules keep working.

Both blanket "off" overrides are gone. no-restricted-properties is keyed by property
name, so the moment a second property joins the renderer block those overrides would
have silently exempted it — in the one file that is the escape hatch, and in every test
in the repo. Tests are where people copy patterns from, so they stay covered; the four
real uses carry line-scoped disables with a reason.

* fix(terminal): keep render-desync capture ids inside main's 120-char cap

createCaptureId builds `${Date.now()}-${panePart}-${nonce}`. A real paneKey is
`${tabId}:${leafId}` — two UUIDs, 73 chars after sanitizing — so with a 36-char UUID
nonce the id is 124 chars and main rejects it with 'Invalid render-desync capture id'.
persistHealedReference swallows that into console.error, so it shows up as diagnostics
that silently never appear.

This was already broken on the desktop app, where randomUUID is available; routing the
non-secure path through the same generator would have made it unconditional, including
on the plain-HTTP web client this branch exists to repair.

Bound the pane part rather than the nonce: keep the trailing 40 chars, which is the
whole leaf id (the identifying half, unique on its own) and drop the tab-id prefix, so
ids stay unique and traceable at 91 chars. The 120-char contract now lives in
src/shared next to the IPC args, imported by both sides, so the renderer cannot mint an
id main will reject without the test noticing.

* test(web): cover the whole store graph and the Vault token without randomUUID

The reported stack was the store chunk, not two named modules, so the repro test now
evaluates the store root under the stubbed non-secure crypto. Any new import-time
secure-context call anywhere in that graph fails here, not just the one this branch
removed.

Also ports the request-token regression from #20465, the one piece of coverage the
competing branches for this bug contributed that this one lacked. Both cases fail with
"randomUUID is not a function" when their production change is reverted.

* test(web): restore the real crypto.randomUUID after the non-secure Vault case

randomUUID lives on Crypto.prototype, so stubbing it as an own property of
globalThis.crypto left the restore branch with an undefined descriptor and a
leaked own `randomUUID: undefined`. Swap the whole crypto own property instead,
through one shared stub the repro suite already needed.

---------

Co-authored-by: Neil <neil@stably.ai>
2026-09-25 15:01:30 -07:00
841503152c fix(runtime-environments): don't crash when a server removed via the CLI still responds (#22517)
* fix(runtime-environments): don't crash when a server removed via the CLI still responds

orca environment rm edits the environment store behind the running app, so the
next ok response on a live socket called markEnvironmentUsed, which threw
'Unknown environment' out of an unguarded socket callback. Main-process callers
now use markEnvironmentUsedIfPresent, which skips a missing environment and
keeps every other store error; the status owner pauses shared control instead
of re-establishing it for a removed server.

* fix(runtime-environments): guard usage bookkeeping at the main-process boundary

Keep one strict store contract and move the leniency to the caller that cannot
report a failure to anyone.

- Revert markEnvironmentUsedIfPresent: drawing the line around one error string
  left corrupt, unreadable and oversized store files still fatal on the same
  unguarded socket callback.
- Add recordRuntimeEnvironmentUsage, a named main-process boundary that says
  lastUsedAt is advisory and swallows every store failure. Route only the three
  sites with no observer through it (subscription onResponse in transport- and
  support-routing, and the status owner's verified hook, where a throw skips
  settleWaiters and hangs refresh callers). Awaited request paths stay strict.
- Guard onResponse/onBinary in the subscription frame router the way the sibling
  request router already guards validateStatus, so no consumer throw can reach
  the ws 'message' emitter and become main_uncaught_exception.
- Drop the status-owner `capable && present` gate: pauseStandingRetry no-ops
  while subscriptions exist, and removal teardown belongs to #21048's watcher.

Co-authored-by: mmarabel <mmarabel@users.noreply.github.com>

* test(runtime-environments): cover the real socket path a consumer throw escapes

The existing tests invoke the captured onResponse directly, which never touches
the surface that actually kills the app. Drive a real WebSocket server through
subscribeRemoteRuntimeRequest so the throw travels ws 'message' -> handleFrame
-> consumer; without the frame-router guard vitest reports it as an unhandled
error, which is main_uncaught_exception in production.

---------

Co-authored-by: Neil <neil@stably.ai>
Co-authored-by: mmarabel <mmarabel@users.noreply.github.com>
2026-09-25 14:48:54 -07:00
c220d92c03 fix(codex): Codex 0.157+ starts in Orca-managed homes instead of failing with SUN_LEN (#22878)
* fix(codex): turn off Codex daemon auto-start in homes whose socket path exceeds sun_path

Codex >= 0.157 auto-starts a background app-server daemon and connects to
<CODEX_HOME>/app-server-control/app-server-control.sock. Orca's managed homes
under userData make that path longer than sun_path (104 bytes on macOS, 108 on
Linux/Windows), so every interactive codex in an Orca terminal failed with
'path must be shorter than SUN_LEN'. The config mirror now writes a marked
[features] daemon_auto_start = false into only those homes, removes it when the
home fits, and never promotes it into ~/.codex.

* fix(codex): address review of the daemon socket guard

- A runtime config.toml holding only Orca's daemon override no longer reads as a
  config-sync stall, so users without ~/.codex/config.toml get no false
  "missing" warning in the accounts pane.
- The legacy shared-home refresh re-applies the guard, so retained pre-rollout
  panes keep daemon auto-start off after a system-default launch.
- Warn once when an inline `features = {...}` or `[[features]]` blocks the
  override instead of failing silently.
- Rename the upsert's TUI-specific internals now that it serves any table.

* fix(codex): apply the daemon socket guard even when the settings mirror stalls

When the settings write-back or mirror refused (unreadable baseline, failed
write to ~/.codex, unreadable source), the whole pass returned before the
daemon guard was applied. A home whose config.toml predates the guard then
kept failing with SUN_LEN on every launch for as long as the stall lasted.
The guard now lands on those paths too; the mirror itself is unchanged.

* fix(codex): guard managed account homes when ~/.codex/config.toml is missing

* test(codex): keep reset-credit ownership checks scoped to the retry, not service construction

* test(codex): build the account mirror test without a type cast

* fix(codex): keep blocking WSL ownership checks off the no-config guard pass

Guarding account homes with no ~/.codex/config.toml ran the WSL ownership
check, a synchronous wsl.exe call per account, at startup before the window
opens and on every account switch. WSL homes are guarded by WSL launch prep,
so that pass now covers host homes only.

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
2026-09-25 14:29:08 -07:00
Jinwoo Hong 64704daac5 feat(feature-tips): one-time tip for agent session search (#22923)
* feat(feature-tips): one-time tip for agent session search

Session search is only discoverable from Settings. Add a feature tip that
existing users see once, which turns search on, shows the first index
build's progress in place, and opens the sidebar search once it is ready.
A toast says when a build left running in the background finishes.

Also share the two-column tip layout across the voice, Cmd+J and session
search dialogs, and move the CLI tip dialog into its own file.

* refactor(feature-tips): simplify the session search tip after review

- Watch for a background finish only when this tip closes mid-index; closing
  any other tip no longer arms a stray "ready" toast.
- The hook detects the close itself, so dialogClosed/reset and the onStatus
  callback on useSessionSearchStatus are gone.
- Table-driven dialog copy, reuse FeatureTipActions, and share the eyebrow
  badge and settings link across the voice, Cmd+J and session search tips.
- One getPendingFeatureTips for the startup gate and the modal.
- Demo: a phase timing table and hoisted header props.
- e2e helpers mark the new tip seen too.

* fix(feature-tips): retire the session search tip once the user has switched search

Turning session search on or off in Settings, or enabling it from the
sidebar, now marks the tip seen, the way the Voice switch does, so a user
who turned search on and later off is not pitched it again.

* test(ai-vault): give the legacy-filter store mock markFeatureTipsSeen

Also mark the tip seen only after the sidebar's enable actually saves.
2026-09-25 17:14:15 -04:00
Brennan Benson acf8e679ea feat(native-chat): Claude sessions write their subagents into the host status store (#22536)
* refactor(native-chat): the host hands out client delivery's status subscriptions as they are

subscribeStatus and subscribeTurnCompletions wrapped client delivery's bound
methods in forwarding lambdas; they are now the same members, the way
waitForSendSettlement already is. The host is at its size limit, and the
next channel it hands out needs the line.

* feat(native-chat): Claude sessions write their subagents into the host status store

The Claude background-task tracker queues child-work evidence at each decision it
already makes (start, update, progress, terminal frame, roster replacement, turn
end, session end), plus the two facts its legacy row ignores: a foreground child's
progress and a foreground spawn call's result. The adapter drains that evidence
after the journal handled the frame and the parent row was republished, and the
host folds it into one record per child in its canonical store.

Nothing reads the records yet; the strip and sidebar keep their current sources.

* test(native-chat): pin the Claude child-work evidence and the host reduction of it

* test(native-chat): prove every hop from a Claude frame to the host's child record

The adapter delivers evidence after the frame's journal rows and the parent's
republished row; the frame script keeps the parent state today reads while the
records add outcome and activity; the runtime hands the evidence to the status sink
under the session's own address; both entry points wire the sink to the ingest.

* test(native-chat): read an optional task list as optional in the producer script

* test(native-chat): an address whose publish threw carries no child work

* test(agent-status): a foreign record differs from ours by producer alone

* feat(native-chat): a foreground Claude child's own tool call is what its record says it is doing

A child's tool traffic reaches the parent stream only for a foreground child. Read
after the journal handled the frame, the child's newest call still awaiting a result
becomes its open operation, previewed the way a hook-reported row previews its own
tool; the result closes it. The open call is derived from the journal's own
bookkeeping, not held a second time.

* fix(native-chat): a Claude child restarted under a new spawn call keeps reporting to its record

A task that ended and starts again stays hidden from the legacy row until a roster
lists it, so the tracker held no run for it: the new run's progress reached nothing
and a foreground re-run's own spawn result settled nothing. The run is now held
beside the live map, where the legacy row never reads it, until a roster hands it
back or it ends. A parity test pins the record's run count to the journal roster's
attempt on a new spawn call, the one event both count.

* refactor(native-chat): the Claude child-tool queries and translator contract get their own homes

The translator's child-tool queries move into claude-child-tool-queries.ts and its contract
type into claude-journal-translator-contract.ts. Brings the translator back under the size
limit.

* refactor(native-chat): Claude child evidence carries only its own edge's facts

Admission now keeps what a child's record already knows: labels, model,
owner, residency, the last message within an invocation, and a token count
that never shrinks. The evidence side copied all of those forward itself, a
second owner of the same rule. It now sends only what this edge observed,
and a task's token count comes from the frame that reported it.

* refactor(native-chat): Claude child evidence hands admission its raw labels

Admission now folds provider text to one line and drops a malformed fact
instead of refusing the record, so the evidence side no longer folds labels
itself. The description keeps admission's longer bound.

* fix(agent-status): admission alone decides a settled child's second ending

The reconciliation returned before admission whenever a record had already
settled with a definite outcome. That dropped the evidence an `unknown` ending
carries (its last message and tokens), which admission's refine-only rule keeps,
so that rule never ran for the structured producers.

The latch goes. Admission keeps the definite outcome, lands the late evidence,
and refuses a conflicting definite ending as `stale-invocation`, which the host
ingest already counts as the fence doing its job, not a fault.

Pinned through the real Claude producer and the host's own ingest.

* perf(agent-status): keep child records off the status hot paths

Child records made every store write and every status notification scale with the
whole store. Each Claude child progress frame cost about 2 ms with 5 chats holding
~200 child records (about 14 ms at ~1,400), and every status change on any lane
re-parsed every child record just to list parent rows.

- The store derives each frozen record's key once instead of re-parsing it on every
  mutation's validation and every alias lookup.
- Settled history is trimmed only when a batch settles something.
- Parent listing and the structured row's revision stamp read the parents and the
  revision directly instead of building a full snapshot.

A progress frame now costs about 0.3 ms at the same size, and listing parent rows no
longer depends on how many child records the store holds.

* fix(native-chat): an errored Claude spawn result no longer decides how the child ended

Interrupting a foreground Claude agent while it runs a tool delivers the spawn call's errored
result before the child's own killed/stopped frames. The spawn result settled the record
`failed` first, and admission then refused the later `cancelled` as a conflicting ending, so an
interrupted child read as a failure.

An errored spawn result now settles the child `unknown`; the child's own terminal frame refines
it to `cancelled` or `failed`. A successful spawn result still settles `succeeded`. The test
replays both frame orders the real CLI produced when interrupted.

* test(native-chat): pin a Claude foreground child's real finishing order

The real CLI ends a foreground agent with its own completed update, then a notification
carrying the final summary and usage, and only then the spawn call's result. Existing tests
modeled the spawn result arriving first, so nothing checked that the notification's summary
and tokens still land on a record the update already settled.

* perf(agent-status): a store write costs what it touches, not the whole store

With child records on the host, every mutation copied all five store maps and re-validated
every record, and reads scanned every child and alias. A parent status publish cost about
10 ms with 4,000 child records in the store, and a child update about 13 ms.

- A mutation writes into drafts over the committed maps and lands in place; a refused one
  is dropped with nothing to undo. The drafts keep the exact map order a copy would have.
- Only what a mutation touched is re-validated: touched parents, children, aliases, facts
  and tombstones, plus every alias of a touched child and whatever a removed parent owned.
  The full validation stays for snapshot restore.
- The snapshot byte budget is a running total instead of a re-measure.
- Children by parent, facts by parent, aliases by child, aliases by identity and retired
  aliases are indexed, so reads return stored records without scanning or re-parsing.
- The memoized alias identity and tombstone-key checks are gone: indexes derive them once.

A parent publish now costs about 0.015 ms and a child update about 0.06 ms at 40, 1,000 and
4,000 children alike. A seeded fuzz holds the store to the copy-and-validate-everything
path decision for decision, snapshot for snapshot and read for read, and a replica fed the
envelopes ends identical.

* fix(native-chat): a Claude child ends only on its own terminal frame

The child records were fed from the legacy background-task tracker's display decisions, so
they inherited rules that are not truth: a turn ending swept foreground children, a roster
omitting a background child settled it, a foreground spawn call's result ended the child,
and a new background start after any roster produced no record. Captured from the real CLI,
an agent moved to the background keeps its own shell running for 40 s after the parent's
turn ends, and that shell was settled `unknown` at the parent's `result`. Replayed with the
spawn result ahead of the roster, the same agent settled as a false success and was then
revived as a spurious second run.

A new decoder reads the task frames directly. `task_started` opens a child (a start for an
ended task id is a restart, the way messaging a finished agent resumes it), progress and a
live `task_updated` update it, and a terminal `task_updated` or `task_notification` ends it.
Rosters, turn ends and spawn results say nothing about a child. Every child in every capture
gets its own terminal frame, so no evidenced ending is lost. The notification's `tool_use_id`
names the run that ended (captured on a resumed agent's second run), so an ending from a run
that is already over no longer ends the current one; a run id the record never saw still
ends it, so nothing strands.

The tracker, its settled-task retention and the frame readers are back to exactly what main
has: the aggregate-roster split and the restart holding map are deleted, and the legacy row is
unchanged by construction.

* fix(agent-status): a session's end settles its live children instead of erasing them

When a structured session ended, the reducer removed every child record it held, finished
or not, so a reader could no longer tell how the session's work had ended. Now a child still
live when its session ends settles `unknown` (nothing reported how it ended), and a child that
had already ended keeps its outcome. The records still die with their parent: closing or
releasing the session drops the parent row, and the store drops its children with it. A
child's own outcome arriving after the session ended still refines the `unknown`.

The `inventory` and `turn-ended` edges, and the rules that settled children on a roster
omission or at a turn boundary, are deleted: no producer sends them any more. A restart is
now its own flag on a live edge, which is what a producer reports when a finished child
starts again under the same run handle.

* test(native-chat): replay the real Claude CLI's frame orders into a real host

Scrubbed cuts of five Claude CLI 2.1.280 stream-json captures (ids, paths and prompts replaced,
frame order and relative clock kept), replayed through the adapter into a hook server:

- an agent moved to the background keeps its own shell live past the parent's turn, and the
  shell settles at its own notification's time;
- the same capture with the spawn result ahead of the move ends the agent once, from its own
  notification, with no second run;
- a roster that omits a background child without its own ending leaves it live;
- a session that ends settles what still runs `unknown` and keeps every record;
- messaging a finished background agent opens its second run, which ends from its own frame;
- interrupts in both captured orders end `cancelled`, and a finished foreground agent keeps
  its summary and usage.

* test(agent-status): hold the store's running indexes and byte total to a rebuild

The copying-store fuzz never reaches the snapshot byte budget, so a drift in
the running byte total (or any index the public reads do not surface) passed
it. After every fuzzed step, including refusals, compare every index with one
rebuilt from the committed maps.
2026-09-25 12:50:50 -07:00
Brennan BensonandClaude d443320af2 refactor(native-chat): remove the unused terminal handoff (#22783)
* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-25 10:17:36 -07:00
745cde69cd fix(feedback): send text-only report when screenshots exceed the upload limit (#22508)
Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com>
Co-authored-by: Neil <neil@stably.ai>
2026-09-25 02:49:52 -07:00
Neil 69839c253e feat(zcode): explain a ZCode build that has no terminal UI (#22730)
* feat(zcode): explain a ZCode build that has no terminal UI

ZCode ships one agent runtime behind two front ends. The desktop app bundles it
without `@zcode/tui`, because it draws its own window in Electron. Put that
bundle on PATH as `zcode` and it answers `--version`, runs `-p` headlessly, and
passes `zcode doctor` — so Orca detects it, launches it, and installs hooks
against it, all successfully. Only the interactive session fails, leaving a bare
Node stack trace in the pane that reads as a broken Orca integration.

Watch a freshly launched ZCode pane's first output and replace that with an
explanation: Orca's hooks are fine, this `zcode` just cannot open a session,
install one that ships the TUI.

The rule keys on Node's own module-resolution error rather than on the healthy
build's "TUI requires an interactive terminal." message, because ZCode localizes
the latter (`TUI 需要交互式终端。` in zh-CN) and matching it would miss every
non-English user. Node's error is not translated and names the package.

Scoped so it costs a healthy pane nothing: it runs only for a pane Orca launched
as `zcode`, and only over the first 8 KiB, because a module-resolution failure
happens before the runtime renders anything.

Evidence: `src/main/runtime/__fixtures__/zcode-missing-tui.txt`, a recorded PTY
capture of the desktop bundle refusing to start, per
docs/reference/agent-pty-transcript-capture.md.

Reported-by: JWu527

* refactor(zcode): ask the CLI if it can open a session instead of watching for the failure

The stream watcher this replaces never fired. Before/after screenshots were
identical and instrumentation showed the hook never ran, so the sidecar was
both misplaced and racing a failure that lands ~440ms after spawn.

Replace it with a direct question, answered once per run and cached.

Reading zai-org/ZCode shows why running it is the only way to ask, and why the
answer is unambiguous. `--version` and `doctor` are byte-identical in shape
between a build that has the terminal UI and one that does not, because the TUI
is only ever touched on the `tui` command path. There, `runTuiCommand` calls
`loadTuiRuntime()` before anything else, and `runTui` checks for a TTY only
after that module is already loaded. So with stdin at EOF:

  - no TUI  -> fails in the loader  -> Node's module-resolution error
  - has TUI -> loads, then declines -> "TUI requires an interactive terminal."

The module error is therefore present exactly when the terminal UI is absent.
All three shipping shapes land correctly: an npm/node-bundle install resolves
`@zcode/tui` as a real package (esbuild marks it external, so it is never
inlined), a SEA build always carries it as embedded assets, and the desktop
app's bundled runtime carries neither.

Verified against both real builds on this machine rather than a mock: the
desktop bundle answers `missing-tui`, a CLI built from source answers
`interactive`, and a command that does not exist answers `unknown` — the probe
fails open so an unrelated spawn failure never accuses a working CLI.

* feat(zcode): warn at launch when the installed zcode cannot open a session

Wires the capability probe to the one place a ZCode launch is first known:
terminal tab creation, which runs before the pane connects, so the explanation
reaches the screen alongside the failure rather than after it.

- main exposes the cached probe over `preflight:zcodeInteractiveCapability`,
  beside the other "what can the installed CLIs do" answers
- the web preload stub answers `unknown`, because a paired client has no
  business deciding anything about the host's CLI install
- the renderer notice is advisory: a probe that cannot run never blocks a launch

Verified in the running app against the real desktop bundle: creating a ZCode
workspace now shows "This ZCode build has no terminal UI" next to the stack
trace, where before the trace stood alone.
2026-09-25 02:28:40 -07:00
Neil 90801e2deb feat(agents): add first-class ZCode harness (#22464)
* feat(agents): add first-class ZCode harness

Add ZCode (Z.ai's `zcode` CLI) as a supervised Orca agent: managed lifecycle
hooks on local, SSH and Windows hosts; status, question and approval reporting;
synthetic status titles; session resume; orchestration worker launch options;
and desktop + mobile agent-picker registration.

Written against the newly open-sourced `zai-org/ZCode` (agent CLI 0.16.9), not
against a remembered screen:

- ZCode's hook runner writes a Claude-compatible stdin alias set, so it routes
  through the existing Claude-compatible vendor path while keeping its own
  identity in the sidebar.
- `PermissionRequest` fires only once the approval card is on screen and racing
  the user's answer, so it is proof the pane is blocked, not an auto-approval.
- ZCode's clarification tool is literally `AskUserQuestion` with Claude's
  questions/options shape, so Orca's question card renders it unchanged.
- ZCode's `hooks.enabled` defaults to false, which is why configured hooks were
  reported as never firing; the installer sets it.
- ZCode renames its own process to `zcode-cli`, so the expected foreground
  process cannot be the launch command or dispatch refuses the pane.
- ZCode emits no OSC title in any state and repaints its ASCII banner forever,
  so readiness comes from Orca's synthetic hook title and launch drafts wait on
  the composer box rather than on a quiet render window.

Three files crossed their max-lines limit, so each is split along a real seam:
command-line entrypoint parsing out of agent process recognition, skill
classification out of skill root discovery, and registry coverage out of the
remote hook installer tests.

Refs #10564

* fix(zcode): drop the session-option catalog and pin the orchestration contract

ZCode's CLI exposes no `--model` flag at all, and the session-option launch path
refuses to apply any option until a model id is chosen. A catalog therefore could
not deliver `--mode` per worker, and would have accepted `--model` only to drop
it silently. Take opencode's position instead: no catalog, so `worker-start
--model` is refused with a clear message and ZCode launches with the model from
its own config. `--mode` stays reachable through agent args, which is also how
the yolo default is applied.

Add a contract test covering the parts that make ZCode a usable worker:
dispatchable foreground process, stdin prompt delivery, the prompt staying out
of the launch command, and the composer-gated draft paste.

* refactor(zcode): reuse shared helpers and cut the harness down

No behaviour change; every ZCode test still passes.

- Use installer-utils' own `hookDefinitionHasManagedCommand` instead of
  re-walking a hook definition by hand, which also drops a local string reader.
- Share one `readZCodeEventMap` instead of keeping the same narrowing in both
  hook-settings and hook-config-json.
- Collapse five identical error returns into one `zcodeHookError` builder, and
  return early from the status branches instead of assigning through `let`.
- Split the event-to-status decision out of `normalizeZCodeEvent` into a pure
  `readZCodeTurn`, so the normalizer reads as decide-then-build and stops
  computing the tool name for events that never look at it.
- Take a script file name in `readManagedZCodeHookEvents` like its siblings,
  which removes a `Parameters<typeof …>` indirection at the call site.
- Drop the unused `ZCodeHookEvent` export and inline a single-use path helper.
- Correct a stale comment: ZCode's loader is a strict `JSON.parse`, so the
  in-place edit preserves key order and indentation, not comments.

* fix(zcode): address review — keep unmanaged event keys, correct comment, de-dupe README

- `removeZCodeManagedHooks` deleted any event key whose list ended up empty, so an
  unrelated `"Notification": []` the user wrote was removed as collateral whenever a
  managed hook elsewhere made the write happen. Only touch an event Orca actually
  owned something in; covered by a new regression test.
- The `isNewTurnEvent` comment claimed UserPromptSubmit was ZCode's only turn
  boundary while the expression below it also returned true for SessionStart. Say
  what the code does: SessionStart lands the idle boundary, UserPromptSubmit is the
  turn boundary (the Codex/Claude shape).
- ZCode appeared twice in the README's single agent-badge block; keep the
  local-icon entry the link checker validates and drop the favicon duplicate.

* docs(zcode): call out that the desktop bundle's CLI cannot open a session

From live testing on #22464: pointing `zcode` at the desktop app's bundled
`glm/zcode.cjs` installs Orca's hooks fine but then fails with
`Cannot find package '@zcode/tui'`, so the pane never opens a session. The
symptom reads as a broken harness when the CLI simply has no TUI. Say which
build to use and how to check before reporting a problem.

Reported-by: JWu527
2026-09-25 02:17:51 -07:00
Brennan Benson 67fc894c8e refactor(orchestration): give structured sessions an orchestration actor column (#22522)
* refactor(orchestration): give structured sessions an orchestration actor column

Adds nullable session:<id> actor columns to runs (coordinator) and
dispatch_contexts (assignee, creator) at schema v42, a shared codec, a
fill for rows that provably belong to a structured worker, and a
coordinator mail-address cache that remembers a handle-less coordinator
by its actor address.

* test(orchestration): pin the actor columns, their fill, cache and v40/v41 upgrade paths

* refactor(orchestration): fill structured-worker actors from one open-time call site

* fix(orchestration): refuse terminal handles as session actors and clear the assignee actor on reassignment

* test(orchestration): read the current schema version from its constant in the delivery downgrade contract

The contract asserted user_version 41 after old code reopens a database
current code wrote, so the v42 bump failed it. Assert SCHEMA_VERSION so
the next bump cannot strand it; the pre-v41 pin and its v40 stamp stay.

* fix(orchestration): count a Run's coordinator actor only at the generation it was written at

A binary without the actor column rebinds and unbinds a Run by rewriting its handle and pane,
which it cannot clear the actor beside. A rebind followed by an unbind leaves a row identical
to a live chat binding. Both writes bump consumer_generation, which every binary already
maintains, so the actor now carries the generation it was written at
(coordinator_actor_generation, set in the same statement) and counts only while the two are
equal. The coordinator cache, its triggers and the open-time fill read the actor through one
rule in run-coordinator-actor; the fill also replaces an actor an older generation left behind.

Still schema v42 (unreleased): the column joins migrate-v42 and the v42 skew-probe entries, so
a database stamped v42 without it replays the chain.

* refactor(orchestration): drop the unused coordinator-actor index and bare-id normalizer

Nothing in this stack looks a Run up by coordinator_actor: callers load the Run and compare its
current actor, so idx_runs_coordinator_actor would ship in every database with no reader. v42 is
unreleased, so it leaves the migration rather than needing a later drop. normalizeOrchestrationActor
had no caller outside its tests; bare session ids enter through sessionOrchestrationActor, and the
handle-refusal cases stay covered there and in parseOrchestrationActor.

* fix(orchestration): keep the coordinator-actor index the caller lookup needs

The next step finds a caller's Runs with one statement that ORs a pane-leaf
match with `coordinator_actor = ?`. SQLite splits that OR across two indexes
only when both sides have one; without idx_runs_coordinator_actor the plan
falls back to scanning every Run on each lookup. v42 is unreleased, so the
index returns to migrate-v42 rather than needing a later schema step.

* refactor(orchestration): store the Orca session id instead of an "actor"

"Actor" read as a new concept when the columns only ever named a structured
session. Rename them to what they hold: coordinator_orca_session_id (with its
_generation), assignee_orca_session_id and creator_orca_session_id, plus the
matching indexes, still added by migrate-v42 since v42 has not shipped.

The columns now store the bare Orca session id rather than session:<id>. The
session:<id> mail address is derived from ORCA_SESSION_ADDRESS_PREFIX where
mail needs it: the coordinator address triggers and the cache refill share
one SQL builder. isOrcaSessionId keeps refusing terminal-handle-shaped ids, and
the generation rule and backfill evidence rules are unchanged. A dev database
stamped v42 with the earlier *_actor columns replays the chain and gains the
new ones.

* fix(orchestration): remember every address a Run coordinator has, not the handle first

The v42 coordinator triggers and the on-open refill stored one address,
COALESCE(handle, session address), so a structured worker coordinator was
remembered by its handle only. Remember each address the coordinator has,
its handle and its current session address, each where present, so this
cache follows the same rule as bindRun and no precedence is persisted.

* docs(orchestration): define the Orca session id without a variable this change does not add

The shared codec's comment named ORCA_AGENT_SESSION_ID, which nothing in this
change defines, and ran one line past the wrap. It now says the stored id is
the one the agent is addressed by (a /clear'd chat's lineage root), as the
column comments do, and that PTY agents have none today rather than never.
migrate-v42's note stated the lineage rule twice; it is folded into one sentence.
2026-09-25 00:35:06 -07:00
Brennan BensonandClaude fec8fe4822 fix(native-chat): offer a resume for every chat that was working, and say what it was doing (#22560)
* fix(native-chat): offer a resume for every chat that was working, and say what it was doing

* fix(native-chat): keep a child-work resume offer after the chat is reopened

The offer's cut-off work was read from the items' current state, so once the
chat was opened (reattaching the provider, which rewrites the rows it lost),
a chat offered for its stopped subagents or background tasks dropped out of
the dialog and its retry. Judge the revision rows written after the marker's
cursor instead: they only accumulate, so the reading is the same before and
after reattach, and the per-run pre-reattach copy is no longer needed.

Also pass the lease fence into the shared "shows work" check, as the status
feed does, so a send stranded under an older fence cannot hold an unheld
session's provider alive while the sidebar shows it idle.

* refactor(native-chat): offer is the working check taken right before each child stops

Every review loop found the same bug class in the "is this offer still owed?" re-derivation that
ran when a child was stopped and again after restart. It re-read a journal the provider had already
rewritten on reattach (a notice turn of its own, restated subagent rows) and kept misreading it:
a send made after an earlier completed turn was dropped at shutdown, and Claude's notice turn
refused the resume after restart.

- Teardown now snapshots each session with the sidebar's working check in a new eviction step
  right before its provider child is stopped (after draining events Orca already accepted), and
  keeps it once the stop is proven. The capture-first/confirm-on-stop re-judgment is gone.
- After restart an offer is withdrawn only by a newer user message, dismissal or expiry, beside
  the structural checks (record, support, lease, no fork). Listing, retryable and the pre-send
  check no longer re-derive work state.
- The pre-send barrier drains again while a provider keeps streaming instead of refusing; both
  providers queue a message sent mid-turn.
- Row activity stays a display-only read of the rows after the marker's journal cursor.

* test(native-chat): pin the re-drained admission barrier and the proven-stop gate on offers

* test(native-chat): cover a send-shaped offer across a closed provider notice turn

* refactor(native-chat): drop the pre-stop drain nothing depended on

* docs(native-chat): align marker and retry comments with the simplified offer rule

* fix(native-chat): an unfinished admission drain no longer refuses the restart continuation

The drain before the continuation's pre-send check refused the send when provider events were
still arriving at its 2 s bound or the barrier failed. That refusal journaled the continuation as
rejected, so its own message then read as the user moving on and the offer could never be
retried. The drain is now best effort and the pre-send check judges what the journal holds.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin a stalled or failed admission barrier dispatching once

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): name a cut-off reply from the settlement's row, not the turn as restated

A reattached provider may restate the offer's turn, so the row's mid-reply label read off the
turn's current state could vanish and count the cut-off reply's own tool calls as background
commands. Also corrects the ineligible-offer comment to match the user-moved-on rule.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): the offer is a stop-time snapshot on the marker

The marker records the main agent's own state, the pending prompts and the
live child roster at the moment before its child stops; the dialog row,
status bar and candidate read only that snapshot. Deletes the journal
cut-off reader, journalCursor, the offer TTL and the pre-send admission
drain. Superseded offers and chat closes now delete their records.

* test(native-chat): every offer ending deletes the durable record

* fix(native-chat): preserve restart offers on unreadable journals

* fix(native-chat): release idle sessions with pending sends

* fix(native-chat): keep a message held while the CLI starts from being released

Idle release had been switched to the sidebar's working check minus pending sends, which dropped
the rule that a send held while the provider CLI is still starting keeps the session. An unheld
chat left during startup was then evicted and its message refused. Restore the release rule this
branch never needed to change: an open turn, or a pending send while the child is starting.
Also keeps the host file within its line budget.

* fix(native-chat): delete a restart offer whose conversation forked

A forked conversation can never become the marked one again, but its offer was only skipped, so
with no expiry it sat unseen in the recovery file forever. Report it for deletion on the same path
as a newer user message. Also corrects comments that still called listing read-only.

* fix(i18n): keep Agent untranslated in the Japanese subagent activity rows

The catalog keeps Agent in English for Japanese, and the localization
gate rejects the translated form.

* fix(native-chat): don't call a mid-reply command monitoring in the resume dialog

The live task roster also lists the foreground command a reply is
running, so a chat stopped mid-command read "Was mid-reply · Monitoring:
<command>". Speak monitoring only for an idle lead, as the sidebar does;
the row's tooltip still names every task.

* chore(native-chat): state the mid-reply label rule exactly

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-24 23:19:21 -07:00
Brennan Benson e0144a9eb6 fix(native-chat): show the Codex and Claude model picker the moment a chat opens (#22756)
* fix(native-chat): show the Codex and Claude model picker the moment a chat opens

A new structured chat showed no model picker until its session had been
created, spawned, initialized and had answered a model listing — and the
picker then listed the models a second time. Codex's listing often goes to
the network, so the picker took 0.6-2 s to appear.

- Keep a host-owned model catalog per agent and account home, persisted on
  success only and refreshed in the background once it ages out. A new
  read-only agentSession.modelCatalog RPC answers from it without a live
  session; sessions reuse it instead of listing again.
- Render the picker while the launch is still provisional, showing the saved
  default. A pick made before the session exists is held and applied once it
  publishes; only the host's acceptance saves it as the default.
- Mark the model and effort set in the user's Codex config as the listing's
  default (config/read), so the first frame names what the chat will run.
- Resolve the account a record-less read would use without running launch
  preparation, which writes and syncs account state.

* fix(codex): disable plugins in the model catalog probe app-server

* fix(native-chat): read the host model catalog only for panes on this machine

* fix(native-chat): name a pre-report model only for a chat this view launched

* test(native-chat): pin the launch latch across publish

* test(native-chat): pin a held pick reaching the host before the first send

* fix(claude): pin the catalog probe's config dir by the session spawn's rule

* fix(native-chat): read the host model catalog only for a visible chat

* fix(native-chat): rewrite the model catalog file only when a listing changes

* test(native-chat): type-check the first-send order fixture

* refactor(native-chat): keep the structured options hook under the line cap

* fix(native-chat): send the first turn only after every pick held during launch settles

* fix(claude): name no default effort from the catalog probe listing

* fix(native-chat): name no listed default model for a chat resumed from history

* refactor(native-chat): let the launch own picks made before it publishes

A pick made while a chat launches had no fence to go to, so the pane held it
and flushed it after publish; every other sender (the outbox, the launch
prompt) then needed its own gate to wait for that flush. The launch now keeps
those picks in its own state, applies them against the create receipt's fence
before it counts as published, and every sender follows publish by
construction. The pane flush, the outbox gate and the module-wide held-pick
registry are gone.

The launch also snapshots the saved selection its create seeds when the intent
is built, so a pick in another chat no longer relabels one still launching, and
a pick the host refuses is reported the way a refused mid-session pick is.

* fix(native-chat): name no default model a workspace's own config can replace

The catalog's default is the account's, read without a working directory, but
a chat runs in its worktree, where a project config (Codex's .codex/config.toml
between the project root and the worktree, or a Claude .claude settings file
that sets a model) picks the model instead. The picker named the account
default there while the chat ran the project's model.

A new chat's catalog read now names its worktree. The host checks that
workspace for such config (existence only for Codex, the model key for
Claude) and, when any is present or the workspace is not a local directory,
serves the listing with no default, so the picker names nothing until the chat
reports its model.

* fix(native-chat): name the listed default model before the report only for Codex

* fix(claude): let an option pick made while Claude starts wait for it instead of being refused

* fix(codex): name no listed default when the configured model is not in the listing

* fix(native-chat): show the picker as unavailable until a published chat attaches

* fix(codex): keep the catalog probe's listing when config/read stalls

* fix(native-chat): write the pending model catalog save before quit

* chore: drop an unrelated lockfile rewrite

* fix(native-chat): rename the catalog store's listing parameter off the global fetch name

* chore: drop an unrelated lockfile rewrite

* fix(native-chat): name the model Claude will run before its first turn

* fix(native-chat): keep Claude's pre-turn applied effort out of the saved session options
2026-09-24 23:04:07 -07:00
Brennan Benson 8f7cbad07b feat(mobile): name the machine after pairing (#22104)
* feat(mobile): confirm host identity after pairing

* refactor(mobile): unify host descriptor state

* chore(i18n): translate the last-known host descriptor label

The remote-host row's "Last known" label shipped in English only. Every other locale now carries it, worded as each catalog already words "last known".

* fix(mobile): make the pairing naming step safe to abandon and show the machine live

Pairing:
- An unreadable status.get reply no longer strands the pairing race: the
  descriptor read ran inside the race's success handler and threw, so the
  candidate was never counted and a direct-only pairing sat on
  "Connecting..." until the timeout. The race now reads the status through
  a reader that returns null instead of throwing.
- A pending pairing is now a small state machine: a save in flight owns the
  outcome (Cancel, back, or unmount no longer clear the journal under it),
  a failed save stays pending and the naming screen shows the error with
  Save still available, and Cancel never rejects.
- When the desktop refuses relay provisioning, the journal is cleared before
  the naming step instead of at save, so an app kill on that screen no longer
  blocks every later scan with "recovery pending".
- A pairing that resolves after the screen went away is cancelled rather
  than left with its journal.
- The screens keep their root ref callbacks stable (the latest pending
  pairing is read from a ref) instead of re-creating them per pairing.
- The label field's placeholder shows the name that an empty label saves.
- "Is this an existing host" is derived from the id identity resolution
  hands back, so host-store and its tests go back to main's shape.

Machine descriptor:
- Mobile keeps the host-reported machine name and OS in memory only, filled
  by the status reads that already happen, and labels it "Last known" from the
  row's own connection state. This drops the per-host AsyncStorage copy, its
  web sibling and web-overrides entry, and the removal cleanup. It also fixes
  a latch: freshness used to stay true for the whole process once a host
  had answered.
- Desktop reads the descriptor through lastVerifiedRuntimeStatus and marks
  it "Last known" using the same reachability verdict as the row's dot.
- Drops the unused hostname/previousPlatform resolver inputs and moves the
  "OS · machine" formatting into the shared resolver.

Also restores main's page-only Reconnect gate in the host header (#22326),
undoes the no-op toStoredHostProfile reformat, and re-measures the web
session route at 4222 modules (one fewer: the dropped web persistence file).

* test(mobile): re-record RPC goldens for the deferred pairing save

Repins baseline to the pairing fix commit and re-records every golden.
Against main, 781 goldens move only in the header (baseline everywhere,
adapterSha256 for the pairing adapter family). Six pairing goldens move in
the body, and only in effect order: the pairing now resolves (closing its
candidate sockets) before the naming step saves the host, and a refused
relay provision clears its journal before the save rather than after. The
set of effects and every outcome match main.

This also removes the unhandled-rejection effects and the failed
result-absent/result-null cells that the previous recording captured from
the pairing race's throwing status read.

* fix(mobile): keep the machine name out of the host header title while the label loads

The host screen starts with an empty saved label and loads it asynchronously. With a descriptor
already in memory, the resolver fell back to the machine name as the title for that window, so
opening "Windows-Low Spec" briefly titled the header with the Mac's name. Read the descriptor
only once the label is known.

* test(mobile): build the pending-pairing status through its schema so the test typechecks

The mobile tests typecheck ratchet rejected a partial status literal: the reply type keeps every
optional field as a required key. Parsing the literal through the status schema yields that shape
without a type assertion.

* test(mobile-web): re-pin the session route closure after merging main

#22301 added two src/shared modules this route reaches; its own CI never ran this suite, so the
merged branch read 4224 against the 4222 pin. Measured on the merge.

* feat(mobile): name each host by what its desktop reports

Pairing saves the host immediately again and names it after the machine name the
desktop publishes; every connection refreshes that name and OS, so a rename on the
desktop reaches the phone. A name typed on the phone's Edit screen is kept as a
phone-local override that wins; clearing it returns to the desktop's name.

- Stored host profiles gain optional personalName, lastKnownMachineName and
  lastKnownHostPlatform; `name` stays the resolved value older builds read.
  Legacy records classify a generated "Host N" as desktop-named and any other
  name as a phone override.
- The connection layer runs one retrying status probe per connected host and
  records the descriptor; the capability probe becomes a projection of it.
- Rows and the header always show the OS, add the machine name under a phone
  override that hides it, and mark it "Last known" while the host is offline.
- Removes the deferred-save naming screen and the one-shot home status fetch.
- Name rules move to host-name-identity.ts and the host-list mutation queue to
  host-list-mutation-queue.ts; an unchanged mutation no longer rewrites storage.

* test(mobile): restore the RPC recordings to main's

Pairing saves the host back to back again and the recording adapter is main's, so
every recording matches main byte for byte; the earlier deferred-save re-record and
its baseline repin no longer apply.

* fix(mobile): keep stored host name identity across snapshot saves and on the web page

A connection re-saves its host profile snapshot on relay credential rotation or relay
re-resolution. The re-pair merge let that snapshot's name identity win, so a phone rename
cleared or changed since connect came back, and a newer desktop machine name rolled back.
The stored record now keeps the name identity on every save; the save supplies the rest.

The web page receives only the app's resolved host name, with no identity fields, so the
docked host header treated a phone rename as "no override" and titled the host with the
live machine name. The display hook now reads such a source the way storage reads a legacy
record: a non-generated name is the user's label.

* fix(mobile): hand the web page the host's stored name identity

The page received only the app's resolved host name, so it had to guess whether that
name was the phone's override or the desktop's name. It guessed "override" for any
non-generated name, which froze a desktop-adopted name as the title after the desktop
was renamed, and left an offline page header without the last-known OS and machine name.

The shell now puts the stored personalName and last-known descriptor on the init host as
optional fields. The page reads them exactly as the app does; a page handed its host by an
older shell still falls back to classifying the name, and an older page ignores the fields.

* fix(mobile): drop an unreadable host name field, not the whole host

The stored host record and the page's init host checked the platform against a closed list
and required non-empty names. A value this build does not know, such as a platform added in a
later build, failed the whole record: the host list dropped the paired host and the next write
persisted the list without it, and the page refused its init message. Those three optional
fields are now salvaged, so an unreadable value drops only that field.

Also states the one exception to the stored name rule (an OS reported without a machine name
keeps the adopted name), and brings four mobile test files in line with the branch: the edit
screen now saves `personalName`, and host opens now start a descriptor status probe.

* refactor(mobile): name the shared host name fields for their role

* fix(mobile): drop the "Last known" prefix from the host machine line

The OS and machine name line under a host's name reads the same whether or not the
host is connected; the connection status already says when it is offline, and a
prefix that users could read as applying to the name added nothing. Removes the
resolver's liveness input and the desktop row's translation key.
2026-09-24 22:35:22 -07:00
Jinwoo Hong f559c0588a fix(terminal): ground a program that dies with input modes armed on the normal screen (#22739)
* fix(daemon): rebase durable checkpoints on the live terminal

A durable checkpoint was folded from the previous checkpoint plus recorded
output, so it inherited that checkpoint's modes forever. After a daemon
restart killed a full-screen TUI and a new process started inline, the chain
kept the dead TUI's alt screen and mouse tracking (?1049h ?1003h ?1006h)
while the live emulator was clean. Every reattach and getBufferSnapshot
served the stale chain, the renderer re-armed mouse tracking, and wheel
scrolling went to a program that never asked for it: scrolling froze.

Each full checkpoint is now the live snapshot verbatim (screen, layout,
alt frame, modes, owner) with only the normal-buffer rows live has evicted
taken from the durable replay. A checkpoint can no longer carry a dead
process's modes, and checkpoints already poisoned on disk heal on the next
compaction.

- The first fold after a cold restore replays the same seed segments live
  was given, so rows line up even over a dead TUI's alt screen.
- Idle zero-record folds keep the disk copy when it already agrees with
  live, so quit and relaunch bursts don't replay every session.
- Held teardown bytes are already in the drained records and the live
  snapshot, so they are no longer replayed twice or appended as a tail.
- The bounded getBufferSnapshot path honors the requested depth even when
  the live window is deeper, without phantom link rows.
- The fold's ownership scanner and frame merge are removed; owner and
  frame come from live.

* fix(terminal): one process-boundary ground for every known or proven boundary

Three copies of the "the process that armed these modes is gone" reset had
drifted: the cold-restore seed cleared only pen and mouse, the recovery
barrier used the renderer's dead-TUI profile, and the cold-restore payload
had none. A cold restore therefore left the dead process's focus reporting,
bracketed paste, application cursor and keypad modes armed in the live
emulator, the first checkpoint, and main's mirror. And the seed wrote the
dead process's torn escape after the reset, so the new shell's first bytes
could complete it (for example retitling the pane).

PROCESS_BOUNDARY_GROUND replaces them: CAN, leave the alt screen without
moving the normal-buffer cursor, every mouse protocol and encoding off,
focus/paste/app-cursor/keypad off, cursor shown and style reset, kitty
popped, SGR reset, grounded DECSC. It stays inert for the lifecycle
scanner. The seed, the recovery barrier, and the cold-restore payload all
use it, and the seed no longer carries the torn tail.

The first fold after a cold restore now always rebases on live, because
focus and keypad are not in TerminalModes and the zero-record shortcut
could not see them differ.

* fix(terminal): ground a program that dies with input modes armed on the normal screen

The daemon's in-stream crash detector only fired when a program died with
the alternate screen up. A normal-buffer program that armed mouse tracking,
focus reporting, keypad or kitty keyboard flags and exited without
disabling them was cleaned up only in the renderer, so the daemon kept the
modes and re-armed them on the next reattach, mobile included (#13077's
garbage-at-the-prompt family).

The lifecycle scanner now tracks armed input modes (mouse protocols and
encodings, ?1004, ?66, and kitty flags as per-screen stacks that mirror
xterm's main/alt swap and its 16-entry cap). ?2004 and ?1 are excluded:
shells arm them at their own prompts. Modes armed when a command starts
(OSC 133;C) count as the shell's, so a prompt that leaves modes on never
triggers. At OSC 133;D the trigger is now "alt screen or a program-armed
input mode", still one-shot and still gated by the shell proof, and the
existing PROCESS_BOUNDARY_GROUND is recorded through the stream so live
and durable history change together.

WSL panes spawn wsl.exe, which the shell proof does not recognise, so the
detector never grounds them; a test pins that and the renderer keeps
covering them. The mouse-leak e2e now keeps its arming process alive until
the live pane is checked, because the daemon grounds a proven exit.

* fix(terminal): keep shell- and host-armed input modes through the process-boundary ground

ConPTY arms focus reporting (?1004h) before the first prompt, and the live
recovery ground cleared it for the rest of the pane. The barrier now re-arms
the modes that were on at OSC 133;C right after the ground, so only the dead
program's modes are reset.

* fix(terminal): re-assert only modes the shell or host armed outside a command

A mode a program leaked past a refuted proof was still on at the next OSC
133;C, so the baseline snapshot re-armed it after a later ground. Record who
armed each mode instead: only enables outside a command (before any marker,
or between 133;A/D and C) form the baseline.

* fix(daemon): keep OSC links and kitty flags through durable checkpoint folds and trims

Stop seeding persisted OSC link ranges into the fold replay: they index the
base buffer, so rows evicted by pending output left a link on the wrong text.
The serializer already writes OSC 8 into the ANSI the fold replays.

Re-apply kitty keyboard flags when replaying a snapshot for trimming, since
rehydrateSequences omits them.

Bound a smaller restore request by trimming the committed checkpoint instead
of re-reading disk and rebasing the live window at a smaller depth.

* test(daemon): follow the isFirstTake rename in the process-boundary ground suite

* refactor(daemon): drop the unreachable deep-live branch from the durable fold

The live window's override cap now derives from the restore depth, so live can
never be deeper than the fold. pendingRecords and isFirstTake are required.

* test(daemon): pass pendingRecords to the process-boundary ground fold

* fix(terminal): reset alt-screen kitty flags in the process boundary ground

Kitty keyboard stacks are per screen, so resetting only after ?1049l left
a dead TUI's alt-screen flags for the next alt-screen app. Also drop the
inert CAN from the ground (every site grounds after complete bytes) and
correct two stale comments.

* fix(terminal): track input-mode ownership in one map

Each armed mode now has one owner: host (before any marker, or a prompt a
133;C proved), prompt (unproven until C), command, or stale (left past a D).
Host arming is sticky, 133;D demotes command modes (the one-shot), and the
ground re-asserts only host modes. Fixes a D without C triggering on host
modes, an ESC c mid-command turning later enables into host modes, and a
program's repeated host enable dropping host ownership. The reattach e2e now
keeps the arming program alive so only the reattach reset can disarm it.

* refactor(terminal): stop treating kitty flags as host state

fish, the one shell that pushes kitty flags at its prompt, pops them before
running a command and re-pushes at the next prompt, so the ground never needs
to restore them. Only host private modes are re-asserted now.

* fix(terminal): keep host input-mode ownership across RIS

ConPTY answers a mid-command ESC c by re-sending ?1004h, which reset() had
recorded as the command's, so the ground turned host focus reporting off for
the rest of the pane. RIS now drops only non-host ownership.

* fix(terminal): let only the host own focus reporting and leave it in the ground

Host ownership covered every mode armed before the first marker, so a tmux
that died with mouse on had it re-armed by the ground. And the re-assert's
?1004h enable made the runtime's ownership mirror revoke, so remote owners
never settled on Windows. Only ?1004 can be host-owned now, and the ground
skips its ?1004l instead of turning it off and back on, so injected bytes
carry no enables.
2026-09-25 00:50:04 -04:00
Brennan Benson 58ba75b5a5 feat(agent-status): child work records say what the child is doing, how it ended, and when (#22521)
* test(agent-status): pin the legacy child-work projection of published background tasks

* feat(agent-status): child work records say what the child is doing, how it ended, and when

A child-work record gains the facts every surface needs from one host-owned
record: the child that owns it (parentChildWorkId), whether the provider said
it may outlive its launch turn (residency, host-only), what it is doing now
(operation, with an open/reported basis), the newest thing it said
(lastMessage), and when its current invocation settled (settledAt, stamped by
admission, never by a producer).

The codec enforces one membership x state legality matrix: live work is never
done and carries no outcome or settle time; only a shell or monitor stores
monitoring; settled work is done with an outcome and a settle time inside its
own evidence window; an operation exists only while live and working, waiting
or blocked. A settled record written without an outcome reads as unknown, never
success. Malformed descriptive fields drop and keep the record.

A new read-only view (AgentChildWorkView) is the one projection surfaces read;
the legacy subagent and background-task shapes are derived from it with
today's output unchanged for today's inputs. deriveAgentChildDisplayState
folds a child's own state and the liveness of the work it owns through the
same fold a parent row uses, so a child whose own work is idle or done reads
monitoring while a shell it launched runs.

Codex children get a thread_id alias kind.

* fix(agent-status): an unknown child ending can gain its real outcome; operation clock clamped

A settled child whose ending was first recorded as unknown (a roster omission
can land a tick before the frame naming the outcome) now accepts the definite
outcome for the same invocation and keeps its original settle time. A definite
ending still never changes, and a later unknown ending is ignored rather than
downgrading it.

Admission clamps operation.observedAt into the child's evidence window, so an
operation stamped in provider time is kept instead of silently dropped.

The record codec is pinned as host-internal: it rejects a whole record over one
unknown key, so a ratchet test fails if anything outside the host store and
admission path imports it.

* test(agent-status): pass the fold-parity input as a value; name the hook lane's alias kinds

* fix(agent-status): a sparse child observation never erases what the record already knows

Admission merged a later observation by replacing the whole record, so an
ending that knew only that the child was gone dropped its name, model and
token count, and an outcome refinement dropped the recorded last message.
Labels now fill or replace but never clear, tokens never shrink, and a
settled ending keeps its last message unless new evidence carries one.
The provider-id preference is keyed by alias kind so a new kind cannot
compile without a rank.

* fix(agent-status): a child's owner, residency and last message outlive a sparse observation

A settle that knows only that the child is gone dropped who owned it and
whether it ran in the background, and the last thing the child said while
live. They now survive like the labels do: the last message for its
invocation, owner and residency for the child.

* fix(agent-status): an unknown ending keeps a definite outcome and still lands its evidence

A settled child's later `unknown` (or omitted) ending was acknowledged without a write, so a
late last message, token count, alias or reclassification it carried was dropped while the
caller was told it was accepted. The outcome now merges like every other sparse fact: an
`unknown` claims nothing and keeps the stored definite outcome, and only a different definite
ending conflicts.

* fix(agent-status): group child aliases and owned work in one pass

Appending by spread copied each bucket on every insert, quadratic in a bucket's size on the
projection and per-row liveness paths.

* fix(agent-status): group child aliases and owned work without Map.groupBy

The relay runs this core on Node 18, which lacks Map.groupBy; a plain loop into a Map is
equally linear and portable.

* fix(agent-status): a child's activity and last message survive the codec

Admission folded raw provider text with the status-row normalizer, which can leave a tab or
other control character and can end a truncation on a space. The record codec drops such a
field, so a long command cut at a space, a tab in a command, or an escape in a message
silently erased the child's current operation or last message. Admission now folds control
characters to spaces and trims the cut, with the codec's own control-character predicate.

* refactor(agent-status): parse child facts, merge, then check the record

Admission merged provider values before anything knew they were valid, and
the codec then either rejected the whole record or silently dropped the
field depending on how old the field was. A malformed owner erased the
stored one, a label cut on a space rejected the announce, and a bad token
count blocked a settle.

Admission now parses every descriptive fact into a value the codec accepts
or "not said", merges it over the stored record with one rule per fact (a
typed map, so a new request field without a rule fails to compile), and the
codec checks the result. Text goes through one normalizer and the codec
accepts exactly its image; any value outside it is a writer bug and
rejects. The owner is now a fact of the invocation, like the last message.

* refactor(agent-status): name each erasure row by what makes its value malformed

* refactor(agent-status): pin the token parse where the max-merge cannot hide it

* fix(agent-status): provider timing lasts only for its own run

providerTiming records the provider's start and end of one run. Keeping it
across a resume left a live restarted child claiming the previous run's
completion time. It now follows the owner and last message: kept within an
invocation, reset by a new one.
2026-09-24 21:44:35 -07:00
Brennan Benson eb746a6d32 fix(codex): a Codex native chat that never sent a message reopens after restart (#22639)
* fix(codex): start a new thread when a chat's thread was never saved

A structured Codex chat records its thread at create time, but Codex writes
no rollout until the first input. After a restart, launch resumed that
thread, Codex answered "no rollout found for thread id", and the chat could
never run again.

When the head of the handle chain is the session's own creation and Codex
answers that exact error for that exact thread, start a new thread instead.
The new link supersedes the unsaved creation in place and names it, so the
chain keeps one live identity and does not grow across restarts. A thread a
resume, fork or adoption proved is never superseded, and no other resume
error starts fresh.

* test(codex): build launch-resolution chains without a type assertion

* fix(codex): match only Codex's own no-rollout text, pinned through the real connection

The fallback matched Orca's own error-wrapper prefix too, and every test built
that string itself, so rewording the wrapper would have disabled the fallback
with the suite green. Match the method, code -32600 and Codex's exact detail
as the message suffix, and drive Codex's raw error frame through the real
connection in a test.

The link builder now refuses, at the type level, a supersession on an adopted
or resumed link, which the chain would reject downstream anyway.

* docs(codex): note why the no-rollout text is safe on the resume path
2026-09-24 21:43:04 -07:00
Brennan Benson a0e24905f6 fix(agent-status): a cancel never hides live work (#22476)
* fix(agent-status): a cancel never hides live work

After the user cancels a turn, a background shell, scheduled check or
subagent that is still running keeps reading as it truly is in both
lanes. The fold no longer takes a verdict input; the cancellation
survives only as lead.outcome, restated as the row's interrupted flag on
a settled row for readers that predate lead.

* fix(agent-status): keep a cancel's verdict and clock on every settle path

A Grok turn cancelled while a task ran now reads monitoring, and the
idle_prompt backstop that later settles it restated done without the
row's `interrupted` flag, so notification readers announced the
cancelled turn as a clean finish. Derive `interrupted` from the main
agent's outcome, as the Claude builder already does.

The inferred Claude cancel now folds through the host's local main
agent record, which a relayed pane never refreshes, so a second cancel
on an SSH pane inherited the first cancel's clock. The caller admits
only a working main agent, so the cancel always starts a new done clock.

* fix(agent-status): keep the shell fact on an inferred cancel so restart can seed it

An inferred Ctrl+C cancel beside a working subagent publishes a row held
open by child work, but the synthesized event dropped the row's paired
claudeRunningNonAgentTask fact because mainAgent changed. Hydration seeds
a settled main agent only when that fact says no shell ran, so after a
restart the child's drain left the row working with no mainAgent. Carry
the fact forward: a cancel does not change what the shell inventory said.

* fix(agent-status): a Ctrl+C at an idle main agent's prompt cancels nothing

Every row that publishes the main agent fact now admits an inferred cancel
only while that main agent is working. Grok's Ctrl+C at the idle prompt
leaves its background task running, so settling the monitoring row to
done hid live work. Rows without the fact keep the evidence guard, and
Codex keeps it too because its synthesized row is a plain done.

* fix(agent-status): fold a relayed pane's cancel from its row, not the desktop's records

The inferred Claude cancel read and wrote the desktop's own listener
records for every pane. For an SSH pane those records are not the relay's:
hydration seeds them from the saved row and nothing reaps them, so a
subagent that finished on the remote after a desktop restart kept a
cancelled row spinning with nothing running. A local pane still records
the verdict on its listener and folds its own roster; a relayed pane
folds only the child work its row carries. The relayed-pane parameter and
forced clock the shared record path grew for this are gone.

* fix(agent-status): hold a cancel verdict in the store until a new turn or the provider's own

A relay never learns of the cancel the desktop infers from Ctrl+C, so its
next child hook or reconnect replay restated the main agent as working and
flipped the row back. The late-hook suppression that guarded this keyed on
a done row flagged interrupted, which a cancel held open by a shell or
subagent no longer is; it also dropped Grok's own stop_cancelled when the
inference won the settle race, hiding the task that hook reported.

The suppression is replaced by a latch derived from the row: its main
agent reads cancelled (or, from an older host, a done row flagged
interrupted). A settled incoming main agent, another prompt, an explicit
prompt or a session start releases it. Child and replayed events keep the
latched main agent and are re-folded with their own child evidence; late
main agent work is held as before, and Codex keeps its record re-mark.

* test(agent-status): pin Codex's evidence guard beside the main agent fact

* fix(agent-status): a prompt submission ends the cancel verdict latch

The task notification Claude starts when background work ends is a real
turn, but it keeps the cached prompt and carries no explicit prompt, so
within 15 s of a cancel the latch held its prompt submission and every
tool event after it: the turn read as monitoring under a cancelled main
agent until its Stop. The captured shell cancel has exactly this: the
notification lands 0.17 s after the cancel key.

* fix(agent-status): derive a Codex row's interrupted flag from its main agent

The cancel verdict latch lets any settled mainAgent through, so a late
root Stop after an inferred Codex cancel now applies where the old
same-prompt window held it. It restates the cancellation on mainAgent
but, unlike Claude and Grok rows, carried no interrupted flag, so mobile,
the dashboard and notification text read the cancelled turn as finished.
Codex rows (local and relayed) now derive the flag from the main agent
record, like the other providers that publish one.

* docs(agent-status): describe cancel admission for every provider and the store's cancel-verdict hold

* docs(agent-status): correct the idle-prompt Ctrl+C claim to the measured CLI behavior

* fix(agent-status): preserve waiting relay children on cancel

* fix(agent-status): resolve the cancel hold before a child's permission card adopts a relayed main agent

The permission-card hold took the incoming event's mainAgent before the cancel hold ran,
so on an SSH pane a child's next tool under a sticky card restated the relay's stale
working main agent and dropped the cancellation the desktop had inferred.

* test(agent-status): pin that a cancelled turn's drained subagent settles as stopped, not completed

* fix(agent-status): keep a cancel through a restarted relay's child hook and a teammate's idle

A relay that restarts after a desktop-inferred cancel has lost its prompt cache,
so the child's next hook arrived with an empty prompt, read as a new turn, and
replaced the cancelled main agent with none; the row then stayed working after
every child stopped. A child's empty prompt is now unknown, not another turn; a
non-empty different one still releases, since it is the listener's newer prompt.

TeammateIdle names its child by teammate_name and carries no agent id, so the
latch treated it as the main agent's and let the late-hook window apply it after
15 s, reviving the cancelled turn. It is now re-folded as child work.
2026-09-24 21:24:59 -07:00
Jinwoo Hong fe46138716 fix(terminal): one process-boundary ground for every known or proven boundary (#22735)
* fix(daemon): rebase durable checkpoints on the live terminal

A durable checkpoint was folded from the previous checkpoint plus recorded
output, so it inherited that checkpoint's modes forever. After a daemon
restart killed a full-screen TUI and a new process started inline, the chain
kept the dead TUI's alt screen and mouse tracking (?1049h ?1003h ?1006h)
while the live emulator was clean. Every reattach and getBufferSnapshot
served the stale chain, the renderer re-armed mouse tracking, and wheel
scrolling went to a program that never asked for it: scrolling froze.

Each full checkpoint is now the live snapshot verbatim (screen, layout,
alt frame, modes, owner) with only the normal-buffer rows live has evicted
taken from the durable replay. A checkpoint can no longer carry a dead
process's modes, and checkpoints already poisoned on disk heal on the next
compaction.

- The first fold after a cold restore replays the same seed segments live
  was given, so rows line up even over a dead TUI's alt screen.
- Idle zero-record folds keep the disk copy when it already agrees with
  live, so quit and relaunch bursts don't replay every session.
- Held teardown bytes are already in the drained records and the live
  snapshot, so they are no longer replayed twice or appended as a tail.
- The bounded getBufferSnapshot path honors the requested depth even when
  the live window is deeper, without phantom link rows.
- The fold's ownership scanner and frame merge are removed; owner and
  frame come from live.

* fix(terminal): one process-boundary ground for every known or proven boundary

Three copies of the "the process that armed these modes is gone" reset had
drifted: the cold-restore seed cleared only pen and mouse, the recovery
barrier used the renderer's dead-TUI profile, and the cold-restore payload
had none. A cold restore therefore left the dead process's focus reporting,
bracketed paste, application cursor and keypad modes armed in the live
emulator, the first checkpoint, and main's mirror. And the seed wrote the
dead process's torn escape after the reset, so the new shell's first bytes
could complete it (for example retitling the pane).

PROCESS_BOUNDARY_GROUND replaces them: CAN, leave the alt screen without
moving the normal-buffer cursor, every mouse protocol and encoding off,
focus/paste/app-cursor/keypad off, cursor shown and style reset, kitty
popped, SGR reset, grounded DECSC. It stays inert for the lifecycle
scanner. The seed, the recovery barrier, and the cold-restore payload all
use it, and the seed no longer carries the torn tail.

The first fold after a cold restore now always rebases on live, because
focus and keypad are not in TerminalModes and the zero-record shortcut
could not see them differ.

* fix(daemon): keep OSC links and kitty flags through durable checkpoint folds and trims

Stop seeding persisted OSC link ranges into the fold replay: they index the
base buffer, so rows evicted by pending output left a link on the wrong text.
The serializer already writes OSC 8 into the ANSI the fold replays.

Re-apply kitty keyboard flags when replaying a snapshot for trimming, since
rehydrateSequences omits them.

Bound a smaller restore request by trimming the committed checkpoint instead
of re-reading disk and rebasing the live window at a smaller depth.

* test(daemon): follow the isFirstTake rename in the process-boundary ground suite

* refactor(daemon): drop the unreachable deep-live branch from the durable fold

The live window's override cap now derives from the restore depth, so live can
never be deeper than the fold. pendingRecords and isFirstTake are required.

* test(daemon): pass pendingRecords to the process-boundary ground fold

* fix(terminal): reset alt-screen kitty flags in the process boundary ground

Kitty keyboard stacks are per screen, so resetting only after ?1049l left
a dead TUI's alt-screen flags for the next alt-screen app. Also drop the
inert CAN from the ground (every site grounds after complete bytes) and
correct two stale comments.
2026-09-24 23:38:53 -04:00
Brennan Benson 517ef56c66 fix(agent-status): end a Claude helper's turn when an API error stops it (#22745)
* fix(agent-status): end a Claude helper's turn when an API error stops it

When a Claude helper agent's request fails (for example a 429 rate limit),
Claude skips the helper's SubagentStop and TeammateIdle hooks and sends only
a StopFailure carrying the helper's agent_id. Orca treated that StopFailure as
ordinary helper activity, so the helper row stayed "working" and pinned the
pane "working" indefinitely, even after the lead agent finished.

Route a helper's StopFailure through the same child turn-end path as
SubagentStop, in both the hook listener (roster update) and the server's
sticky-permission rule (a failed helper no longer holds its permission prompt).

* test(agent-status): cover a failed background child that held a permission prompt

* test(agent-status): pin the captured order of a rate-limited teammate before the lead stops
2026-09-24 18:39:58 -07:00
Neil 7ea01279cd feat(search): bundle ripgrep for local, WSL, and SSH search (#22396)
* feat(search): bundle ripgrep for local, WSL, and SSH search

Ship @vscode/ripgrep-universal's prebuilt rg for all six relay platforms in
every desktop artifact. Local and WSL searches spawn the bundled binary and
drop the git ls-files / git grep fallbacks; SSH deploys upload the remote's
binary once per ripgrep version and the relay prefers it over PATH rg.

* fix(search): address bundled ripgrep review findings

- Key the SSH ripgrep cache on the binary's content hash; a package bump is the only update step
- glibc verifier: read arch tokens below the slice root and accept static ELFs (arm64 release blocker)
- Ship ripgrep/PCRE2/musl license notices; bundle rg with orcad
- Packaged builds never spawn a bare rg; report fd pressure as transient
- SSH: install rg before sweep/GC, size-validate installs, back off instead of disabling on launch failure
- Scope Dependabot to @vscode/ripgrep-universal; revert unrelated lockfile churn

* chore(search): drop bundled-ripgrep reference doc; assert full packaging layout parity

* refactor(search): one entry point for spawning the bundled ripgrep

Local Quick Open, Quick Open path search, the Explorer name filter, and
runtime text search each repeated the same three steps: resolve the bundled
command, spread in the WSL distro, spread in the WSL shell expression. Fold
that into spawnBundledRipgrep so one place owns the rule that a bare 'rg'
must never reach spawn, and simplify the resolver's command/packaged checks.

Restore the AGENTS.md ripgrep rule dropped alongside its reference doc in
63f4dac, and note why the relay's availability probe may spawn a bare 'rg'.

No behaviour change; verified by the existing suites plus a new test that
pins the local, WSL-routed, and distro-routed-but-Windows-output cases.

* refactor(search): drop the local install-ripgrep path; enforce the rg rule

Bundling rg removed the local git/readdir fallback, so nothing can produce
the "install ripgrep on the host running the Quick Open scan" guidance any
more -- only a remote host an upload never reached still reaches the capped
listing. Drop the host parameter, the renderer's local branch and its
translation key, and the relay wrapper that existed only to pass 'remote'.

Add a ratchet test for bare 'rg' spawns, since the AGENTS.md rule alone had
nothing enforcing it. Its one allowlist entry is the relay's PATH probe,
which asks about PATH by definition. Verified the guard catches a planted
offender rather than passing vacuously.

Also stop chaining the remote cleanup sweep behind the ripgrep upload: on a
cold host that is a multi-MB transfer, and stale upload stages and
superseded version dirs were left on the remote for its whole duration. The
two touch different trees, so they now run concurrently.

* test(ssh): pin that the cleanup sweep does not wait on the ripgrep upload

* fix(search): derive rg spawn types instead of importing node:child_process

A type-only import still counts against the child_process ratchet, whose pin
and allowlist only ever shrink. Derive both types from wslAwareSpawn instead.

* fix(search): surface an unreachable WSL workspace instead of an empty result

Inside `bash -c`, a failed `cd` exits 1 -- the same code ripgrep uses for "no
matches" -- so a WSL workspace whose directory had gone away reported an empty
listing as a successful scan. main did not have this hole: checkRgAvailable ran
the same `cd` wrapper first and settled on `code === 0`, diverting to the git
fallback that this PR deletes. The WSL wrapper now takes an optional
cwdFailureExitCode; rg passes 97, and all four close handlers reject with a
clear error before the unavailable check can blame the install.

Also from review:
- Bound the fire-and-forget ripgrep upload with deploySignal. The controller
  aborts only on the deploy timeout, never on success, so this cancels a
  still-running upload when the deploy gives up.
- Run the stale-stage sweep before the installed check rather than inside its
  else branch. Once rg was installed every later deploy took the PRESENT path,
  so a stage orphaned by a dropped connection was never collected again.
- Note in orcad-remote-deploy.ts why wiring it up needs ripgrep work first:
  build-orcad.mjs copies only the build host's rg, and orcad reports
  isPackaged() === true, so a remote of another platform would find nothing.

ssh-relay-deploy.test.ts sat at the max-lines cap, so any edit to it failed the
gate. Split the four Windows named-pipe deploys into their own file (926 -> 737
+ 333); both are now well clear of it.

* fix(search): name the unreachable root in every handler, not three of four

Round-two review caught that the missing-cwd branch in scanRipgrepPaths sat
AFTER isRipgrepUnavailableExit, which classifies any code above 2 as a broken
install -- so for exit 97 it was dead code and Quick Open still told the user to
reinstall Orca. Reordered; all four handlers now check it first.

Also from review:
- A vanished workspace makes spawn fail with ENOENT, which read as a damaged
  install on every local path. Confirm the cwd with isRipgrepSpawnCwdUsable --
  the guard the relay already applies -- before blaming the binary. The async
  continuation re-checks `resolved`, because finish() drops its argument once
  settled and the rejected promise would otherwise go unhandled.
- bundledRipgrepCommand returned a bare 'rg' for an arch outside the bundled
  set, bypassing the guard that exists so Windows cannot resolve a bare name
  against the repo cwd. A packaged app now always names an absolute path.

Drop ci-shards/unit-assignment.json, a 9,425-line CI artifact swept in from
reproducing a shard locally, and gitignore the directory that produced it.

The "rg genuinely cannot start" test pointed at a synthetic /repo, which the
new guard correctly reports as unreachable; it now resolves to a real root so
it still tests what its name says.

* fix(search): let the error handler own the spawn-failure verdict

A failed spawn emits 'error' and THEN 'close' with a negative code. The cwd
check added in the error handler did not settle, so the close handler settled
first -- synchronously, with the reinstall message -- and won the race every
time. The branch was not merely flaky, it was unreachable in all four handlers:
it is guarded by pid === undefined, which is exactly the case that always
produces a following close(code < 0). Verified against a real spawn: 3/3 runs
give error(ENOENT) -> close(-2). The error handler now detaches 'close' before
the probe, so it owns the outcome.

The probe also had no rejection handler, so a probe that rejected left the
search unsettled forever -- a hang, not just a wrong message. It now falls back
to the prior verdict rather than inventing one.

Tests: filesystem-search-rg-timeout and orca-runtime-files-search already cover
error-first and close-first, but against synthetic roots that the new guard
correctly calls unreachable; they now resolve to a real root, keeping each
test's stated intent. Added a Quick Open case for the vanished-workspace path
and confirmed it fails with the old ordering.

* test(search): cover exit code 97 in all four ripgrep close handlers

Round-four review found the missing-cwd branch had zero handler coverage: no
test anywhere emitted close(97), only -2/0/1/2/127. Ordering was correct, but
guarded by source-line order alone -- and that exact ordering was wrong in
three of four handlers two commits ago. Each suite now drives close(97) through
its real handler and expects the unreachable-root message.

Verified the tests earn their place: neutering the missing-cwd check fails
exactly four tests, one per handler.

Also drop a Reflect.get the anti-slop gate rejects, in favour of `in` narrowing.

* docs(search): stop claiming the close handler always wins the race

The previous commit asserted close "would beat this threadpool round-trip every
time", from an n=3 sample that measured event ordering -- which was never in
dispute -- rather than probe-vs-close. Two later measurements disagree with each
other: 50/50 close-first here, 30/50 probe-first in review. Either way it is a
race on a sub-millisecond margin, and the detach is what makes the verdict
deterministic.

Why this wording matters: "close wins every time" is an argument for deleting
the detach as a guard against an impossible race. No test would catch that --
the suites emit error and close in the same synchronous tick.

* chore(search): ship the jemalloc and libunwind notices the Linux rg needs

The statically linked Linux builds carry jemalloc (BSD-2-Clause) and LLVM
libunwind (Apache-2.0 WITH LLVM-exception) in addition to PCRE2 and musl, and
both require their notice on binary redistribution. Confirmed with `strings`:
their symbols are present in linux-x64 and linux-arm64 and absent from the
darwin and win32 builds. Texts taken from the upstream canonical sources.

extraResources already copies the whole licenses directory, so these ship
without a packaging change.

* fix(relay): stop spawning a bare rg, name unreachable roots, collect old builds

Three gaps the reviews surfaced on the remote side, all pre-existing on main.

Bare `rg` on Windows remotes. Both relay spawn sites pass the user's repo as
cwd, and CreateProcessW searches the cwd before PATH -- the same hijack the
desktop side already fixes. The relay now walks PATH itself and spawns an
absolute rg.exe, skipping relative PATH entries because those resolve against
the cwd. No rg on PATH yields null, which callers treat as "ripgrep
unavailable" rather than handing spawn a bare name. POSIX keeps the bare name:
execvp never consults the cwd, so there is nothing to resolve and nothing to
gain. With the last probe converted, the bare-spawn ratchet allowlist is empty.

Empty results for an unreachable root. settleLaunchFailure resolved an empty,
successful-looking scan when the root was gone but PATH rg existed, and the
git/readdir chain never engaged because it only triggers on
RipgrepUnavailableError. Both relay paths now reject naming the root, matching
local workspaces. Missing-rg keeps precedence over a missing root, because only
that verdict engages the fallback chain -- two tests pinned that deliberately
and it would have been wrong to flip it.

Unbounded ~/.orca-remote/ripgrep/. Nothing collected this tree; the relay's
version GC only matches `relay-*`, so every rg bump left another ~5 MB per host
forever. The probe command now also drops sibling builds older than two weeks,
sparing the current one and live upload stages, on POSIX and PowerShell alike.
Two weeks because a client pinned to an older build may still be using it; the
cost of collecting one early is that client re-uploading once.

* fix(relay): probe the rg that failed, and close the drive-relative PATH hole

Five review findings against the previous commit, all reproduced first.

The launch-failure classifier probed PATH rg, but the spawn that failed was the
bundled binary. On the normal remote setup -- no rg on PATH, which is why Orca
uploads one -- the probe failed and a moved workspace was reported as a missing
ripgrep, telling the user to install what Orca already ships. So the fix was
inert on exactly the hosts the uploader exists for. It now takes a candidate
list and asks the binary that actually failed first, then PATH.

path.win32.isAbsolute accepts `\tools` and `/tools`: rooted, but carrying no
drive, so they resolve against whatever drive the process is on. The probe
would have validated one against the relay's drive while the spawn, running
with the user's repo as cwd, resolved it against the repo's -- the same
cwd-dependence this lookup removes, narrowed from directory to drive. A real
drive letter or UNC root is now required.

probeRipgrepVersion had lost the timeout's kill in the rewrite, leaking a live
process and a ref'd handle per launch failure -- for a hang, which is the very
case the bundled-rg back-off exists for. It also spawned without windowsHide,
which would flash a console; fixing that made an allowlist entry stale, so the
entry is gone and the pin ratchets down 63 -> 62.

`windowsPathRipgrep ??= …` never memoised a miss, because null is nullish. The
caching was inverted against cost: a hit stops at the first directory, a miss
stats every one, and only the miss was repeated -- per spawn.

The bare-spawn ratchet claimed "nothing in production spawns a bare rg", which
is false on POSIX. It now also matches PATH_RIPGREP_COMMAND at a spawn site,
and the comment states plainly what a textual guard cannot see: the POSIX bare
name reaches spawn as a parameter, and is safe because execvp ignores the cwd.

The drive-rooted predicate is tested directly rather than through the
filesystem -- a temp dir on a POSIX CI host has no drive letter to exercise
win32 semantics with, so the filesystem test could never have caught this.

* test(mobile): repin the session closure past #22452's two shared modules

Merging main brought the closure to 4220 against a pin of 4218. The two extra
modules are `src/shared/agent-turn-outcome.ts` and `src/shared/main-agent-status.ts`
from #22452, which the status projection this route already reaches import.
That change was src/shared-only, so the mobile job never ran on it -- the same
way the structured tool line slipped past, as the ledger above already records.

Repinned here because this PR's file set is what next made the job run, not
because this PR reaches either module. Verified: of the 28 source files this
branch changes, none appear anywhere in the route's 4220-module closure.

* fix(search): preserve remote binaries and complete runtime packaging

* test(relay): pin the probe's env now that it inherits the relay's PATH

8d6759a threaded the relay env into probeRipgrepVersion -- correctly, since the
probe decides whether a launch failure was the binary or the root and so has to
resolve the same rg the failed spawn would have. It left the assertion that
pins the probe's spawn arguments behind, which is what CI caught.

Asserting buildRelayCommandEnv() rather than loosening the match to any object:
under process.env the probe could resolve a different rg, or none, which is the
regression the change exists to prevent.

* feat(ssh): collect remote ripgrep builds by reference, not by age

Nothing collected `~/.orca-remote/ripgrep/`: the version GC matches only
`relay-*`, so every change to the shipped bytes left another ~5 MB on every SSH
host, permanently. The age window this replaces was the wrong instrument --
a directory's mtime is when it was written, not when it was last used, so it
cannot tell a superseded build from the one a live relay was launched against.
Deleting the latter is not graceful degradation: without a PATH ripgrep remote
text search rejects outright, and listing drops to the capped walk this PR
exists to remove.

So the question is reference. Each relay directory now records the build it
runs against in `.ripgrep-ref`, written only once that binary is confirmed
present, and the GC collects a build only when no installation names it.

The discipline is ssh-relay-native-deps-cache-gc.ts': anything the pass cannot
account for blocks the whole pass. A relay directory with no readable marker is
an older Orca's, possibly running right now against a binary it never recorded,
so the pass declines rather than guessing. Those directories are removed by the
version GC in time, which is what makes their builds collectable -- hence
running after it, not beside it. Deletion is the same tombstone, recheck under
the rename, then remove, so a deploy that takes a reference mid-pass gets its
tree restored. Windows has no pass yet, matching the native-deps cache's gate.

One test note: the first version of the "unaccountable blocks the pass" test
passed against a deliberately broken guard, because the tombstone recheck
masked its absence. The test now puts a readable recheck behind an unreadable
first scan, which is the only shape that fails when that guard is removed.

Recording the reference lives inside ensureRemoteBundledRipgrep rather than at
the call site: it is the same concern, and it keeps the deploy's ripgrep
surface to one call for the tests that mock it to protect their exec queues.

* feat(ssh): collect Windows remotes too, and ship the Rust crate notices

Three items previously left documented-but-open.

Windows remote accumulation. The cache GC was POSIX-gated, so the leak did not
go away -- it moved to the platform with the larger binary (rg.exe is 5.43 MB on
win32-x64, against 4.77 MB for linux-arm64). The PowerShell dialect now does the
same reference scan: entries and references carry token prefixes, because
PowerShell writes every uncaptured value to stdout and an untokenised listing
would feed Remove-Item whatever a cmdlet happened to emit.

Verified on a real Windows host rather than a mock: the listing emits its
ENTRY/LIST_OK tokens, a relay directory carrying a marker yields REF <entry>,
and a relay directory without one yields REFS_ERR -- the safety path, on the
real interpreter.

Rust crate notices. The crate set was read out of the shipped binary's symbols
and the licence identifiers taken from crates.io rather than assumed. Where a
crate offers the Unlicense, Orca elects it: a public-domain dedication carries
no notice obligation, and that covers eight of them. The four that do not offer
it get their MIT text reproduced. encoding_rs carries a BSD-3-Clause notice for
its WHATWG-derived encoding data that is joined by AND, not OR, so electing MIT
does not discharge it.

Release-only validation, corrected rather than repeated. Linux AppImage/deb/rpm
already runs in CI's package job on every PR, and Windows signing was already
rehearsed on this branch. macOS notarization is the only item a release must
still exercise, and the exposure is narrow: notarization requires signatures on
Mach-O binaries, and of the six bundled builds only the two darwin ones are
Mach-O -- `file` reports ELF for linux and PE32+ for win32 -- so signIgnore
excludes only files the notary never asks about.

orcad-artifacts.test.ts caught the new notice file missing from the standalone
runtime's shipped list, which is exactly the gap that test exists to catch: a
notice committed to the repo but never actually shipped.

* fix(search): protect relay cache references and handle failed spawns

* fix(ripgrep): close review gaps and repair deployment fixtures

* test(mobile): refresh merged session module census

* fix(ssh): preserve ripgrep caches with empty legacy references

* test(mobile): assert bundle boundaries instead of global module count
2026-09-24 17:25:48 -07:00
Brennan BensonandClaude 1069bb053f fix(agent-launch): a cwd at the workspace root no longer forces a terminal (#22729)
"Continue in New Session…" always names a cwd, and both the renderer route
input and the host launch-mode decision read any cwd as a custom start
directory, so the continuation opened a terminal agent even when chat was
the user's default. Both now share one rule: only a cwd outside the
workspace root (after normalising slashes, Windows case, WSL aliases and
the distro's Linux spelling) requires a terminal. A subdirectory still
does, because a structured session cannot start there.

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-24 16:57:38 -07:00
Brennan Benson 5e6fcca0b3 fix(native-chat): date a session by its own lifecycle, not its subagents' work (#22520)
* fix(native-chat): date a session by its own agent's rows, not its subagents'

The journal reducer's lastActivityAt is the structured status summary's
updatedAt, which the status row uses as its completion stamp and
acknowledgement clock. It took the max over every journal row, and a
session's subagents write into the same journal after its own agent has
settled, so an idle parent was re-dated and marked unread on child work.

A row now dates the session only when the session's own agent produced it:
not a row whose producer linkage names a subagent, and not a subagent
roster row (a subagent-group block), which the session writes but revises
on every child transition. The roster rule is derived from the row body;
no new persisted field. Replay folds through the same rule, so existing
journals are re-dated to their own last row on reopen.

Claude: a backgrounded subagent emits no child frames, so its re-dating
came entirely from roster revisions (task_updated, task_notification) and
from the stale-roster revision written when a journal reopens. Codex: the
roster is revised on every child token-usage report; child-thread rows
carry no producer linkage yet, and read as the session's own until they do.

* fix(native-chat): a reopened journal's verdict on stale work does not date the session

Reopening a journal settles rows the previous host left live (a working
subagent roster, a live background task) to unverifiable. Those revisions
were appended at the reopen moment, and a background-task row is the
session's own non-roster row, so a crash-restarted session with a live
shell was re-dated to the restart although no agent acted.

The reconciler now writes each verdict revision with the row's own observed
time. That is one rule at the one writer, covering both settle shapes; the
render item's observedAt was already pinned to the row's first write, so
nothing the transcript shows changes. The live-transition roster exclusion
stays: live roster revisions are written by the providers, not here.

The `recovered` row flag is not used as the discriminator: the live
unexpected-exit settlement also writes recovered rows, and a clock rule
keyed on it would stop dating a provider crash the host just observed.

* fix(native-chat): date a session by the reducer's attribution of what a row wrote

The clock read producer linkage off the raw row. A lifecycle batch names no
row-level producer, a tombstone names none, and a revision may name none while
the reducer still attributes the item to a subagent, so each of those dated an
idle parent. The clock now asks the reducer: after a row applies, whether any
item it wrote is the session's own work; before a removal, whether the item it
removes was.

* fix(native-chat): date a session's status by its own lifecycle edges

A subagent writes into its parent's journal and keeps going after the parent
settles. Every one of its rows advanced the summary's updatedAt, and the status
row re-dated a done parent to it, so an idle parent read as newly finished and
unread on each child step.

The host now publishes statusStartedAt beside updatedAt: when the session's own
agent entered its status, read off edges only it writes. Idle is when its newest
turn ended; working is when the running turn was requested, or the earliest
send still unanswered; attention is its own oldest pending ask, or a subagent's
when that alone holds it. A turn that recovery settled after its host went away
ended when that settle was written, so it reads as a completion the user has not
seen; it carries no outcome, so no completion event or notification calls it a
success. Render items carry recoveredAt, the recovered row's own write time, so
nothing new is persisted.

The sidebar bridge and the host ingest date the row and the main agent's clock
from it whenever the row shows the main agent's own state, and keep their
existing rules for a row child work holds open or a summary from an older host.
The status feed republishes when the clock moves instead of on every idle row.

* revert(native-chat): keep the journal clock over every row

The row filter this branch put on the reducer's lastActivityAt decided which
rows could date a session: a list of exclusions that each new row kind could
slip past. The session's state is now dated by its own lifecycle edges, so the
filter, its attribution helper and the backdated reopen verdicts go back to
main. lastActivityAt, and the summary's updatedAt it feeds, is again the
evidence clock over every row, including a subagent's.

* fix(native-chat): keep republishing an idle session its live child work holds open

A row held open by live child work is dated by when each reader saw the
publish, and mobile decays a working row whose evidence is older than the
staleness window. Suppressing row-activity republishes for every dated idle
session froze that evidence, so a subagent running more than 30 minutes past
its parent's turn made the row read idle on mobile. Only a session nothing
holds open stays quiet on row activity now; its state clock is unchanged.

* fix(activity): order an agent's timeline by when each state was seen

An answered ask returns a settled parent to its own turn's end, so its done
repeats the time of the done before the ask. Activity keyed and ordered
events by that time: the new done collided with the old one and was
dropped, and the row took its state from the newest-dated event, the
blocked ask, so a done parent read Blocked and needed attention.

Each state switch now records the `updatedAt` it was seen at. Events are
keyed and ordered by that, while unread and "Clear completed" still compare
the state's own time, so the answer neither re-lights unread nor revives a
cleared done. The row's state comes from the pane's own status entry, so a
clear that hid the done cannot leave it reading Blocked either.

* test(activity): pin the timeline across repeated asks, a clear, and a stale turn

Three parts of ordering the timeline by when each state was seen had no test
that failed without them:

- A second ask moves the answered done into history. Both dones share the
  turn's end, so only the history entry's own seen time keeps them apart;
  without it one done collided with the other and the timeline showed two
  Blocked events in a row. Three asks also exceed the per-pane cap, which must
  keep the most recently seen events, not the most recently started.
- "Clear completed" on an answered row must cut off past the ask, which is
  dated after the done, or the cleared row stays listed. A done that the user
  cleared must also stay hidden once a later ask moves it into history.
- A stale working row must not read as running just because the pane's own
  status says working.
2026-09-24 16:17:08 -07:00
Brennan BensonandClaude 6ae6ed08bb fix(claude): open structured chat without a startup deadline, and make Retry start fresh (#22364)
* fix(claude): open structured chat without a startup deadline, and make Retry start fresh

Publish the Claude session as soon as its process is spawned instead of racing
initialize against a fixed 10s deadline. Prompts sent before startup lands are
held and written in order once it does. An exit or sign-in failure before startup
ends the session with the reason and the CLI's stderr.

A create that failed because the process provably exited now carries
ownerVerdict 'exited', so the client marks the launch failed and Retry mints a
new operation instead of replaying the stored failure.

* fix(native-chat): sending into a chat that failed to start restarts it

* fix(native-chat): a send with no live owner restarts it once

A provider child that timed out or exited hands its lease back, and every
later send was refused agent_session_ownership_unknown. Clients read that
code as "not admitted yet" and resend forever, while only a surface hold
could make a new child, once per mount, with its failure swallowed.

The send now routes to a live owner, otherwise restarts one from the
persisted resume state where resume eligibility allows it (single-flight
per session), otherwise refuses with the new settled
agent_session_owner_unrecoverable. Unverifiable, reserved and handed-off
leases are left alone. The desktop hold now logs its failure.

* test(native-chat): pin the unrecoverable refusal as settled in the outbox

* test(native-chat): pin the release clock after a send restarts an unheld owner

* test: read the sent operation id without a cast

* fix(native-chat): type the send-recovery record lookup as the store returns it

* fix(native-chat): a send ensures its owner before admission, and an unheld owner idles for 30 minutes

* fix(native-chat): a create that throws releases its event sink

A child that dies between spawn and journal attach can still write through
the host's event sink, which attach unbound in onAcquiring and never re-bound
because onAttached never ran. The orchestration released that sink only when
performAttach returned a refusal; a thrown failure (the root-exit path) kept
the sink cached with its queued write, so the next attach's drain barrier and
runtime shutdown's flush waited forever.

Also pins the publish-on-root-exit clause for a start that never proved:
deleting it reddened nothing before.

* fix(native-chat): a resend the journal answers restarts nothing, and a send joining a restart rebases from the fence it replaced

* fix(native-chat): the host learns a Claude start positively, and persists only proven options

A publish-first create used to read the session's options before Claude had
answered initialize. With startup pending that read fell back to the built-in
catalog's default, so `record.options.model` was persisted as `sonnet` for
every user whose CLI default is something else; an owner handoff or a reopen
then replayed `set_model('sonnet')` and silently switched their model.

The adapter now reports `started` once startup facts are applied and saved
options restored. The host keeps a `providerChildPhase` on the session it
owns: a starting child hands over nothing but the saved options as intent,
and the `started` event re-reads the options as fact and persists them through
the same record write a user's option change takes. The status summary carries
`hostExecutionPhase` (optional, wire-safe), and the chat pane says the agent is
still starting instead of showing nothing.

A child whose exit already reached the adapter before acquire returns is no
longer handed over as live; the create fails with the CLI's diagnostic.

* fix(native-chat): a hold and a send that find the owner gone share one restart, and a send the ledger already holds restarts nothing

* fix(native-chat): a failed create answers one refusal shape, stamped once at the boundary

A create whose Claude process was seen to exit answered twice in two shapes:
the first call threw a generic runtime error, and only the replay of the same
operation carried the `ownerVerdict: 'exited'` refusal that lets a client
retry under a new operation. Three sites stamped the verdict and the store
failure path stamped nothing.

The first-hand root exit is now returned as the refusal on the first call,
with the provider's own diagnostic as its message. The verdict is stamped in
one place, at the boundary of the attach, from the durable row the operation
settled to, so every refusal shape answers the same fact and no site can
forget it. The per-site stamps are gone.

* fix(native-chat): a send into a session whose child ended restarts it before admission

A session that published and then lost its Claude child before startup (not
signed in, for one) keeps a released lease and a chat the user can still type
into. The send was refused as ownership-unknown, the outbox parked it as
pending admission, and nothing ever restarted the child: the message sat
there until the user closed and reopened the tab.

A send reaching a session with no provider child now runs the same resume a
surface's first hold runs, before the write is admitted. The resume reserves
a new fence, so that send is answered stale with the published fence and the
client's outbox re-drives under it, as after any fence change. A resume that
fails is not this send's answer; admission reports the lease as it stands.

* chore: restore pnpm-lock.yaml to origin/main (local pnpm rewrote it)

* test(native-chat): pin the pre-handover exit as a failed acquire; stub the status feed in the delivery test

An exit the adapter observes before acquire returns now fails the acquire
with the CLI's diagnostic instead of handing over a dead child; the
published-then-ended path stays pinned by the slow-init startup case. The
delivery test renders the pane, which now activates the host status feed.

* test(native-chat): a same-ID re-hold over the wire joins the one resume, and a replay reopen goes on the idle clock

* test(native-chat): a re-hold that joins a failing resume proves one resume ran

* fix(native-chat): a create whose child was proven gone answers the refusal on the first call

The previous change answered a first-hand root exit as the exited refusal on the
first call, but the common failed start never took that path: when the close
ladder proves the whole tree dead the acquisition error is a plain one, the
store-failure classifier rethrows it, and the client still saw a runtime error
first and the refusal only on replay.

The cleanup that proves the child gone now names such a failure
`AgentSessionAcquisitionExitProvenError`, carrying the provider's diagnostic,
unless it already names its own verdict (a refusal, a typed exit proof, a host
store code). The attach answers both proven-exit kinds as the refusal its replay
gives. How a failed acquisition settles and how it is first answered now live
beside the verdict stamp, in the failed-create module.

* test(native-chat): pin the outbox re-driving a stale-refused send under the resumed fence

A send into a session whose child ended is answered stale once the host has
restarted the child. The outbox keeps that operation queued and blocked, and the
fence change the resume publishes re-drives the same operation under the new
fence; the host admits it.

* fix(native-chat): a child restarted for a send nobody holds is still released

The restart a send runs for a childless session takes no holder, on the premise
that the sending surface already holds one. A one-shot writer holds nothing, so
the child it restarted had no release clock and lived until the app quit. The
write resume now arms the clock when no holder is present, as the first-hold
resume already does. The send-after-failed-start cases also pin that the stale
answer's operation is admitted when re-sent under the new fence, and that two
racing sends restart the child once.

* test(native-chat): pin the picked Claude model across a resume whose child starts on its own default

The started event re-reads and persists what the child reports. A resumed child
answers initialize with its CLI default before the saved pick is restored over
it; the record must hold the pick while starting and after started.

* Revert "fix(native-chat): a child restarted for a send nobody holds is still released"

This reverts commit e52c4a6f08.

* Revert "test(native-chat): pin the outbox re-driving a stale-refused send under the resumed fence"

This reverts commit 136a39deb0.

* Revert "fix(native-chat): a send into a session whose child ended restarts it before admission"

This reverts commit 39234e44bf.

* refactor(native-chat): make ensure-owner a step of the serialized send

A send that found the owner gone restarted it OUTSIDE the host's per-session
serialize, through a single-flight resume map shared with the surface hold, then
rebased its fence by heuristic. The attach body is now callable from inside
`serialize` (`attachStructuredAgentSessionUnderSerialize`), and every restart
runs there: a hold, a send's ensure-owner step, provider-exit recovery and the
rewind owner replacement take turns on one queue, so the first to run attaches
and the next finds its child. The single-flight map and `isResuming` are gone.

Admission is two-phase for a send: the ledger's answer comes first and places
nothing; a send it will admit gives the session an owner, and only then are the
row placed and the lease and fence checked. A send it will replay into a closed
session makes the journal readable and spawns nothing. The session entry
records the released fence the child replaced (`resumedFromFence`), so a writer
current as of that owner is admitted at the new fence by bookkeeping, whether it
ran the restart or arrived behind it.

The resume reads its record only after this host has reconciled it and exited
any recovery stage a failed attempt latched, so a hold behind a failed attempt
makes its own attempt against the lease as it now stands.

* fix(native-chat): a Claude start proving itself no longer waits on the CLI

The host handles a Claude child's `started` on the recovery chain every
session's unexpected-exit handling shares, under that session's serialized
step. It then asked the CLI for the model list and settings again, so one slow
CLI held every other session's exit recovery, and its own close, behind up to
two request timeouts.

The adapter already holds those answers when startup proves: the settings read
at startup, the restore's confirmations, and the initialize result the SDK
answers the model list from. `started` now carries that snapshot, and the host
turns it into one record write without any provider I/O.

* test(native-chat): a hold reads its lease only after this host has reconciled it and exited a latched recovery stage

* fix(claude): a chat whose first start failed resumes as the same conversation

A Claude start that dies before initialize writes no transcript, so the next
start launches the chain head's provider id fresh instead of `--resume`. The
launch flag that chose that mode also chose the provider-handle link's origin,
so the fresh launch published a second `created` link onto a chain that
already had a head. The store refused it, the healthy child was closed, and
every later reopen, hold or send spawned and killed another Claude.

The launch now carries the two facts separately: `resumesTranscript` (launch
mode, from whether Claude wrote a transcript) and `continuesChain` (lineage,
from the record's chain head). The link origin reads lineage; rewind and the
Fast opt-in carry-over read launch mode.

* refactor(native-chat): a resume answers with a typed refusal the send classifies

`resumeHeldStructuredAgentSession` and the holds' `ensureProviderChild` answer
`{ ok: true } | { ok: false, refusal }` instead of throwing the refusal code.
The refusal is the attach's own, with its message and, when the failed attach
proved its child gone, its `ownerVerdict`. An attach that settles a failed
acquisition in the ledger and then rethrows the cause is read back off that row,
so a durably failed restart is a refusal and only an unrecorded error is a fault.

The send classifies the refusal through a `Record` over every wire code — a new
code does not compile until it is placed — into transient (the lease is someone
else's to settle; the send runs as the lease stands) or terminal. A terminal
one answers `agent_session_owner_unrecoverable` carrying the cause, forwards the
verdict, and writes the same status row into the chat that a start that failed
leaves, so the user sees why after the error strip is gone. Nothing about the
failure is remembered; a Retry is a fresh attempt. A fault thrown by the restart
itself is reported and the send runs as the lease stands, since bookkeeping
never gates a user's action.

`hold()` still raises the refusal code for its RPC caller.

* fix(native-chat): a child's event sink belongs to the attach attempt that spawned it

The runtime kept one event sink per session id and handed it to every attach.
An attach that acquired a new child unbound that sink first, so when the
acquire then failed its dead child's queued frames stayed in the cached,
unbound sink. The earlier guard only discarded it when no session entry was
left, which a resume of a still-indexed session never satisfies: the next
attach's drain and shutdown's flush waited on it forever. A TUI-to-native
handoff acquire had the same shape.

Each acquiring attempt now mints its own sink. Only a successful attach (or a
proven handoff owner) adopts it as the session's, closing the one it
replaces; any other exit closes it with whatever its child queued. A re-attach
to a live child keeps the sink that child already writes through. A sink that
is not the session's own can no longer force the session's provider down.

The native handoff acquisition moves to its own module, which keeps the
handoff file under its line budget.

* test(native-chat): pin that only the adopted child's event sink still takes writes

Closing a failed attempt's sink and closing the sink a resume replaces were both
unpinned: removing either left every suite green, because neither sink is in the
map that drains and flushes read. The resume test now asserts the failed
attempt's sink and the exited generation's sink refuse writes, and the adopted
one accepts them; deleting either close reddens its own assertion.

* perf(native-chat): the chat reads only the host's startup phase from the status feed

The chat took the whole status summary to read one field, so every status change
for its session (prompt, update time, background tasks) re-rendered the chat
view. It now subscribes with the phase itself as the snapshot, so it re-renders
only when the phase changes.

* fix(native-chat): the startup-phase hook answers a phase or null, never undefined

* fix(native-chat): every restart is counted from the moment it is asked for, and a handoff clears the restart fence

Provider-exit recovery now restarts through the holds' `ensureProviderChild`
like a hold and a send do, so a child whose only surface left while the attach
ran goes on the idle clock instead of living until quit. A hold's resume and a
client attach are tracked as in flight from enqueue, not from their turn on the
queue, so a quit's drain waits for one queued behind a close before it decides
what to evict. A handoff back to native moves the fence in place and now clears
`resumedFromFence`: only a restart may rebase a writer. The failed-restart
status row is keyed by the send's operation id, not the clock, so a resend of
the same id that fails again adds no second row.

* fix(native-chat): a Claude start no longer waits behind another session's exit recovery

The runtime delivered every Claude lifecycle event on the single chain
exit recovery uses so teardown can drain it. That chain orders nothing
across sessions, and an exit recovery on it can run a full reacquisition,
so one chat's `started` waited on an unrelated chat's respawn and kept
its 'still starting' line up. `started` now takes only its own session's
serialized step, is queued the moment it is emitted (ahead of any later
exit of that child), and is tracked in a set the same teardown drain waits
on.

* fix(native-chat): a Claude create that dies at spawn is refused with the CLI's own diagnostic

A CLI that exited before its acquisition handed the child over was
refused with 'claude stream-json for session … exited while being
acquired', or with an unreadable start time, and the stderr the exit
carried (for example 'not signed in') appeared nowhere. The acquisition
now keeps the error its connection ended with and answers with it at
both sites; the generic message is only a fallback when none exists.

* test(native-chat): pin that a reopened Claude chat dying before initialize says why

A resumed start is published at spawn, so a CLI that exits before it
answers initialize fails a chat the user is looking at. Pin that the
open chat is sent the 'stopped before it finished starting' row with the
CLI's diagnostic even when the child's tree cannot be proven gone, and
that a message held for that start is refused rather than left in doubt.

* fix(native-chat): Stop while a Claude start drains its held prompts withdraws the rest

Stop withdrew held prompts only while startup was pending. Once startup
landed and the gate began writing them one by one, a Stop interrupted the
CLI and the prompts still waiting were written straight after it. Stop
now withdraws whatever the gate still holds in both states; the drain
takes each prompt off the queue immediately before writing it, so a
withdrawn prompt can never be written. The one already written still
gets the interrupt.

* fix(native-chat): the release clock keeps a session that still owes a sent message

When the last surface stops holding a session, the release clock evicts
it after the grace unless a turn is running. A message sent while Claude
is still starting is held, not running, so switching away from that chat
for the grace evicted the session and refused a message the user had
already sent. The clock now asks whether the session owes work: a running
turn, or a submission the provider has not taken yet (still pending in
the journal). Both are read from the journal; nothing new is stored. A
starting session that owes nothing is still released, and an explicit
close still ends everything.

* fix(native-chat): only a start that holds a sent message keeps a released session

The release clock kept any session with a pending submission. A Codex
send is admitted and stays pending until its echo, which may never come,
and only an eviction retires it, so such a session was never released
while the app ran. A pending send now keeps the session only while its
child is still starting, which is when the send is held for that start.
Pins that a ready session with an unechoed send is evicted, and that a
Claude chat whose turn finished is released after the grace.

* fix(native-chat): a send waits for the owner it met to prove its start before it is admitted

A Claude child is published before the CLI has answered initialize, so a send admitted right
behind a restart — or right behind the first start — was dispatched into a child that could die
milliseconds later, and learned of the death only as a delivery nobody could confirm. The terminal
refusal the send was written to give was unreachable on the real adapter for exactly the failure
it was written for.

The send's serialized step now admits nothing against a `starting` child. It registers for the
child's startup verdict and returns having placed nothing; the send waits off the session's queue
(the `started` and `ended` settlements run on it) and admits again once the child is `ready`, or is
refused `agent_session_owner_unrecoverable` with the child's own exit reason when it exits first.
The exit settlement writes the one status row, decided by the host's own phase rather than only
the provider's flag. A close, an eviction or a replacement answers the wait too, and quit releases
whatever is left; there is no timer. One spawn per user action holds across re-entries.

* fix(native-chat): restart the release grace when a start writes its held prompts

A prompt held while Claude starts is written when the start lands, but
its turn opens only when Claude echoes it. The release clock stopped
counting it once the child read ready, so a tick landing in that gap
stopped the child before it ran the user's first message. The start
landing now restarts a pending release's full grace, the same grace a
message sent to a ready chat gets before it is released.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): a starting child owns the send; the adapter holds the message for its start

A send that meets a child still proving its start is admitted against it, as it was before the
off-queue startup wait: the adapter holds the message until startup lands and rejects it with the
child's own diagnostic when the child dies first, the exit settlement writes that cause into the
chat, and the release clock keeps a starting session that holds a sent message. The startup watch,
the off-queue wait loop and their teardown phase are gone; the exit settlement still reads a start
that failed off the host's own phase when the provider omits the flag.

The scripted-CLI test now pins that contract end to end: a restart a send asked for whose CLI dies
at initialize leaves the message rejected with the diagnostic, one row naming it, and the fence
moved by two; a healthy CLI is restarted once and written to; a send during the first start is
held and written once initialize answers, or rejected with the diagnostic when the CLI dies.

* fix(native-chat): a failed send restart says why, and offers a new chat only when nothing can restart it

The refusal a send gets when the host cannot restart the chat's agent is renamed
agent_session_owner_restart_failed and now reads "<Agent> couldn't restart: <reason>." with the
restart's own cause. "Start a new chat to continue." is added only when the resume was refused
because this host has no record to restart from or cannot run the one it has. Any other failure,
such as a CLI that is not signed in, leaves the chat retryable: the outbox stops auto-retrying, and
a manual Retry or a new send tries the restart again, since a refusal before admission leaves no
ledger row.

* fix(native-chat): a Claude start skips an option write the CLI never answers instead of faulting at the request deadline

* test(native-chat): wait for the recovery's reserved lease, not the released one it replaces at once

* fix(native-chat): a send whose restarted child dies before starting is rejected with the child's diagnostic

A child that never proved its start has accepted nothing: input is written
only after it initializes. A send admitted against such a child, whose
dispatch then found no session, settled unknown, twice, and the outbox took
Retry away. It now settles rejected with the child's own diagnostic, both
when the dispatch throws and when the exit settles the sends it left
unanswered, so the chat says why and offers Retry. A proven child's
unanswered sends stay in doubt, as before.

* test(native-chat): expect a send held for a start that never proved itself to settle rejected

* fix(claude): write a prompt held after startup already drained, instead of stranding it pending

* test(native-chat): pin that a send to a child that died before starting is answered rejected

* test(native-chat): leave the cast exit-session fixtures as they were, since a proven exit never rejects

* fix(claude): a saved option the CLI never answered stays saved instead of being replaced by the CLI's value

A start skips an option write the CLI does not answer within the request
deadline, and then persisted what the CLI reported in its place, so a slow
answer silently replaced the user's saved model or dropped their saved
permission mode. Silence is not a refusal: the unanswered option is now
recorded apart from a rejected one, the live child keeps running on the
CLI's value, and the saved choice stays on the record for the next start to
retry. An option the CLI rejects is still dropped as before.

* fix(native-chat): a rejected send opens no turn, so the row naming why it failed is not folded away

A send whose restarted child died before starting is rejected, and the
exit writes a row naming the cause. The chat's local clock had watched the
send go pending and stop, so it gave the message "Worked for 0s"; that
settled a turn that never ran, and the fold hid every non-prose row after
the message behind it, including the one naming the cause. The row only
appeared when a later send moved the turn anchor, which read as two rows
for one Retry. The host's journal already says the send was rejected; it
now answers that such a message opened no turn, which outranks the local
clock on desktop and mobile alike. A rejected send whose journal does
record a turn keeps its duration.

* fix(native-chat): a send whose restart died starting leaves the same row as any start that died

One failed attempt already leaves one row, but which row depended on when
the child died. A child that died after the send was admitted left "The
provider stopped before it finished starting: <cause>."; one that died
before the send was admitted left "Claude couldn't restart: <cause>." So the
same failure read two ways from one Retry to the next. When the refused
restart proved its child exited, the send now writes the startup-failure row
itself, as its comment always said it did. The refusal under the composer
still says the restart failed; a restart that failed for a reason other than
a child exiting keeps its own wording.

* fix(native-chat): a send rejected because the agent never started names the cause under the composer

When the child a send was admitted against died before starting, the host
rejected the send with the child's diagnostic behind the internal transport
marker. The client rightly hides that marker's detail, so the red line read
"Couldn't reach the agent" while the cause sat in the record. A startup
death is not a failed write: the host now words that rejection the way the
chat row does, "The provider stopped before it finished starting: <cause>.",
at every site that rejects for it. Desktop and mobile show a reason in words
verbatim already, and older clients do too, so no client change is needed.
Real write failures keep the marker and the generic copy.

* fix(claude): a saved option the CLI never answered survives a later change to a different option

The saved choice a start could not apply was kept on the record, but the next
option the user set persisted only what the child had applied, so changing the
permission mode or effort, or clearing the chat, silently dropped the saved
model. The adapter now reports which saved options are still unanswered, every
option write keeps those saved values, and a write the child accepts for that
option retires it.

* fix(claude): a send that meets a child whose exit already settled names that exit's cause

When the child a send was admitted against died starting and its exit
finished settling before the send reached it, the send was rejected with
"no live claude stream-json session for <id>", now shown under the composer
as the cause. The adapter keeps a settled exit's diagnostic until the chat is
acquired or closed again, so that send names what the CLI said. A refused
restart whose child died at spawn or while its start time was read is pinned
to leave one row in the words any failed start uses.

* test(native-chat): pin the words an exit settlement rejects a never-started send with

The startup gate and the dispatch reject a send first in every existing
scenario, so the exit settlement's own rejection had no test of its wording.

* fix(claude): derive which saved options are still unanswered from what the child applied

A write that lands already puts its option in the session's applied set, so
the unanswered list is that list minus what has since been applied, rather
than a second copy every option write must remember to edit. Session
fixtures built without the new set no longer throw on an ordinary write.

* fix(native-chat): a cleared chat starts from a saved choice the child never answered

Clearing a chat seeded the replacement from the values the child reported,
so a saved model or effort whose restore write the CLI never answered was
replaced by the CLI's own value in the new chat, even though the retired
record kept it. The replacement now keeps those saved values too, and its
start retries them.

* fix(claude): closing a chat forgets its exit's diagnostic even when the exit settles during the close

The diagnostic was dropped when the close began, but closing over an exit
that was still settling finishes that settlement, which kept it again, so a
closed or deleted chat held it until its next acquire. It is now dropped once
the close finishes. Pins that an acquire and a close each retire it.

* refactor(claude): keep a saved option the CLI never answered as the wanted value, not a list beside it

A restore cleared the session's wanted options and added back only the writes
the CLI answered, so an unanswered one lost the user's value and every later
writer had to be told to put it back: the start report, each option change and
/clear each carried a list of unanswered keys. The restore now keeps the saved
value as wanted and unconfirmed, so what the session reports and persists
already carries it, and the list, its adapter method and the started-event
field are gone. A refused option is still dropped.

/clear now starts the replacement from the record's options instead of reading
the child's live values, which can be a model the CLI fell back to.

* test(claude): wait for the start to finish before changing the saved model

The record holds the saved model from creation, so waiting for it returned at
once and the option write could reach Claude while it was still starting,
which refuses it. Wait for the effort the finished start reports instead.

* fix(i18n): translate the still-starting chat notice

The notice that a structured chat is still starting was only in English.

* test(claude): pin the failed acquisition's own reading-control release

The merge re-pointed this test at a child that exits after publish, where the
exit path also releases the binding, so it passed with the acquisition's release
deleted. A child that exits before publish leaves only that release. Also drops
the create 'init' phase, which lost its last producer when rewind stopped
proving before publish.

* refactor(claude): move unexpected-exit handling into the exit lifecycle module

The adapter crossed the 300-line limit once main's context-usage change
landed beside this branch's growth. The two methods that turn a Claude
process exit into an ended event now live next to the existing exit
helpers; behavior is unchanged.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-24 15:23:15 -07:00
Brennan Benson f0b3f44f10 feat(agent-session): let the host own a chat's tab id and let a create reserve it (#22616)
* feat(agent-session): let the host own a chat's tab id and let a create reserve it

A structured chat's tab id was derived from its session id by every layer that
needed one: the renderer, the host snapshot and the status address each built
their own spelling. The join between a conversation and the tab that shows it
must be a pointer the host owns, not a derivation each client repeats.

The session record now carries surfaceTabId. A create pins it: the tab half of
the pane agent.launch reserved, an optional tabId on agentSession.create, or a
host-minted UUID. Records written before the field existed are backfilled at
open with the string clients derived, in memory at once and on disk with the
store's first transaction, so nothing keyed by it (read state, notification
ids, worker rows) moves on upgrade. A second record under a held id is refused.

Only the record and the two create wires change here. The snapshot still
publishes agent-session:<sid> and the renderer still derives its local id;
those move in the next two changes. agentSession.create is a strict object, so
the field is advertised as a capability a client checks before sending it.

* fix(agent-session): record the derived tab id for an unreserved create

A create that reserved no tab minted a random UUID that no reader uses: the
renderer, status address, worker rows and host-shared read state all still key
by structured-agent-session-<sid>. Persisted, that id would move every chat
created before readers switch to the recorded one, orphaning its read state
and worker rows the way the backfill exists to prevent. An unreserved create
now records the derived id, the same rule the backfill applies, so the record
always matches the prefix every existing key uses; an opaque mint belongs with
the change that moves the last reader.

Also:
- a chat tab id must be a host tab id on the record, the create wire and in
  admission, matching what agent.launch already requires of paneKey; a
  web-surface id would decode as another tab
- the stored launch-result guard checks the structured outcome's tabId
- comments no longer claim a retry naming another tab conflicts; replay keys
  on the attach fingerprint and answers with the recorded id (now pinned)
- the wire refusal test used a non-hex digest, so the schema refused it for
  that reason; it now reaches the tab id rule
- pin that the reload path refills the id without forcing a save

* test(agent-session): correct the tab-id fingerprint comment to match replay
2026-09-24 14:42:12 -07:00
Brennan Benson 98584332a3 fix(native-chat): record which Codex agent produced each journal row (#22532)
* fix(journal): a batch revision restates the producer of each row it revises

The reducer rebuilds a row's producer linkage from its NEWEST revision, and
absence is a positive claim: no agent id means the session's own agent wrote
the row. So any revision written without the stamp hands a subagent's row
back to its parent, permanently.

Three host paths revise rows they did not write, from the render item they
already hold, and all three dropped the stamp:
- answering a prompt re-appended the asker's row with the fence only;
- dead-generation settlement failed running tool calls and cancelled pending
  prompts through a lifecycle batch;
- stale-session settlement on acquire cancelled lost prompts the same way.

The batch path could not carry a producer at all: linkage was removed from
the batch row because one row covers N mutations, with a note that a mixed
batch would have to stamp per mutation. Dead-generation settlement is such a
batch already, and Codex settlement is about to become one. So each item
mutation now names its own producer, inline like the row base. A mutation
that names none falls back to the row's linkage, which is what a batch read
before. Parse sanitizes a bad per-mutation id the same way it does a row's:
the field is dropped and the mutation kept.

No schema version bump. An older host's mutation validator ignores unknown
keys, so it reads a stamped mutation as the session's own, which is exactly
what it shows today. Old journals carry no stamp and read as before.

Turn revisions still carry nothing: a turn is the session's unit of work,
and the live-turn scans rely on a turn row never carrying linkage. The note
recording that invariant is updated to the new write sites.

* fix(native-chat): attribute a Codex subagent's journal rows to the subagent

Codex journals every thread on its app-server connection into the session's
journal, and a spawned subagent's items arrive on the child's own thread.
None of those rows carried producer linkage, so under the journal's rule that
absence means the session's own agent wrote a row, every child's command,
message, reasoning, prompt and status row read as the PARENT's: the parent
could show its child's running command, its child's reasoning as "thinking",
and its child's prose as its own latest line.

The Claude lane's model is reused, not reinvented: the same fields and the
same absence rule. What differs is how the producer is known. Orca opens
exactly one thread per app-server, so any other thread is one Codex spawned.
That decides WHETHER a row is a child's from its first frame, announced or
not, and the thread id is final at once: it is never re-minted the way a
tool-call reference is, so no correction ledger is needed for identity.

- agentId: the child thread id, the same id the status side keys a Codex
  child on.
- parentAgentId: the thread whose stream carried the child's `started`
  activity. Codex emits that item on the spawning agent's own session, so a
  child that spawned a grandchild is named; the session's own thread is not.
  Other activity kinds ride whichever agent acted and are not used.
- producerKind: 'agent'.
- attempt: which run of the child the row's own turn was, counted from the
  child turns the roster already observes; absent on the first run. Taken
  from the row's turn rather than the child's latest, so a persistent shell
  that outlives its turn keeps its run across revisions.
- providerParentRef is omitted: a Codex child's frames carry no parent
  reference of their own beyond the thread id, which is already agentId.

One resolver, owned by the roster (which already owns what is known about
each child thread), is handed to every writer: items, streams, generic and
summary rows, prompts, compactions, goals, and the three settlement batches.
The session-end settlement mixes every thread's rows in one batch, so each
mutation names its own producer. Turn rows stay unstamped: Codex writes them
only for the primary thread.

The spawn-group roster row stays unstamped on purpose: a child's frame can
trigger its write, but it is the parent's list of its children.

The translator's construction moves to a parts module so the translator
stays a router under the line cap, and the item streams reuse one
append-and-publish helper instead of two copies. Children are never swept
at turn end; nothing here changes that.

* test(native-chat): pin Codex subagent attribution at every writer and every parent reader

Two layers, so a stamp that is correct in the store and never read, or read
and never persisted, cannot pass.

The readers, through the real path: translator, deferred sink, on-disk
journal, snapshot. Each is a defect on main: the parent named its child's
running command as its own tool, read its child's reasoning as itself
thinking, showed its child's compaction as its activity line, and quoted its
child's prose as its latest line (checked after closing and reopening the
journal, so the stamp is read back from disk). The transcript still renders
the child's rows.

The writers, through a sink that records the linkage of every plain append,
batch mutation and lifecycle transition: start, streamed checkpoint and
completion of one command all restate the child; a row that beats the spawn
announcement is still the child's; a grandchild names the child that
announced it, while an `interacted` activity names no parent; a follow-up
turn is the child's second run, and a shell that outlives its turn keeps its
own; the exit batch settles each thread's rows under its own producer and
the turn row under none; a child's provider frames, approval and goal rows
are its own; nothing is stamped while the session thread is still opening;
and the spawn-group row stays the parent's.

* test(native-chat): pin linkage forwarding on the sink's lifecycle-transition path

A Codex child's goal row is written through a lifecycle transition, so a sink
that forwarded only the fence there would file the child's goal as the
session's own.

* test(native-chat): type the Codex item fixtures as thread items

* refactor(journal): keep a row's producer across revisions that name none

The reducer took a row's producer linkage from its newest revision, so every
writer that revised a row it did not write - a prompt answer, a dead-generation
or stale-session settlement, the reopen sweep of stale subagent rosters - had to
restate the producer or silently hand a subagent's row to the session's own
agent. Three of those writers had been patched to restate it; the next one to
forget would reintroduce the bug.

Attribution is now fixed by a row's first write. A revision that names no
producer keeps the row's existing linkage; one that names any replaces the
whole bundle, which is how a provisional stamp is still corrected in place. A
row re-created after a tombstone starts with nothing. The reducer runs the same
fold on replay, so the kept producer survives a reopen.

The three restatements are removed. Per-mutation linkage on lifecycle batches
stays: a batch can create a row (a Codex child's prompt, or a child's item
settled before any checkpoint landed) and one batch can mix producers.

* test(journal): pin producer inheritance in the reducer and across a reopen

A revision naming no producer keeps the row's, on the plain item path and in
a batch settling a child's row beside the session's own; one naming any
replaces the bundle wholesale; a tombstone clears it; a stale revision cannot
touch it; and a reopened journal replays it exactly as it was folded live.

* refactor(codex): name the translator's writer factory for what it builds

* docs(codex): say why a settled row names its producer

* test(journal): drop a producer test the stale-revision guards make unreachable

The stale revision is dropped whole by two independent guards before the
inheritance rule runs, so its producer assertion could never fail; the
reducer's own stale-revision tests already cover the drop. Also say what
the batch sink does forward: each mutation's own producer.
2026-09-24 13:33:56 -07:00
Jinwoo Hong 5610b11703 feat(ipynb): create a .venv when pip is locked out, and show ipykernel setup progress (#22710)
* feat(ipynb): set up ipykernel in a new .venv when pip is locked out, and show install progress in the dialog

* fix(ipynb): drop the retired installFailed string from translated catalogs

* refactor(ipynb): drive the setup dialog from one setup state; fix review findings

- Kernel status now only describes the kernel; a single `setup` object (base, offer, phase, error) drives the dialog, replacing the extra statuses and the externallyManaged/setupError fields.
- The picker's 'Create virtual environment…' opens the same dialog; success switches through selectEnvironment, so a running kernel is only replaced once the venv exists.
- Async setup results are dropped when the dialog they belong to is gone (tab closed/reopened).
- main verifies ipykernel imports after pip, reuses an existing .venv instead of re-running venv over it, and explains failures that printed nothing.
- Windows copy command guards the install with if ($?); notebooks at a filesystem root get a correct .venv parent; 'Try again' shows for both retry paths.

* fix(ipynb): close the setup prompt when another Python is picked
2026-09-24 16:22:34 -04:00
Brennan Benson 85642d0d88 fix(agent-status): count only agent work in stats, and read a Grok background subagent as working (#22474)
* fix(agent-status): renderer and recorder consumers read the question they mean

Since #22295 a row's combined `state` reads `working` both while the lead's
turn runs and while a subagent or background shell outlives a settled lead.
The lead's own state now rides beside it (`lead`); each consumer in this slice
reads the question it actually asks.

- Smart sort and the Activity unread badge keep reading the combined state:
  their classes and rows are what the sidebar shows. Pinned with tests,
  including a restored `lead.state: 'working'` row that must never read live.
- The stats recorder asks "was an agent executing" and now reads a shared
  derivation (`isAgentExecutionOwed`): the lead's turn, or live agent child
  work holding a settled lead's row open. A settled lead's background shell
  no longer accrues "Time agents worked". Old hosts without `lead` fall back
  to today's read; restored and replayed rows still never open a session.
- The `agent.status.changed` plugin event gains `lead` as an optional field
  through one tested projection; `state` keeps its meaning and restored rows
  still project to nothing.
- A Codex root Stop that follows an inferred interrupt keeps the
  `cancellation` verdict, as the Claude lane already does at its turn
  boundary, on both the hook and relay paths.

* test(agent-status): pin that a child's approval wait no longer splits the recorded span

The recorder's move to the lead fact quietly changed one more story: a Codex
child's PermissionRequest turns the combined row waiting while the root's own
turn keeps running. The old state read closed the span there and minted a
second spawn on resume; the new read keeps one span, because the lead never
stopped. Pin it at both boundaries (the shared derivation and the recorder)
so the change is deliberate, not incidental.

* fix(agent-status): date stats edges by this host's clocks and scope the accrual predicate to stats

The recorder dated a start by the producer's mainAgent.stateStartedAt. An SSH
host stamps that with its own clock while every stop is stamped locally, so each
span gained or lost the clock skew. The same clock also survives a row that
briefly lost the fact (an OSC repaint to another state), dating the reopen
before the close already sent, and a subagent reopening a monitoring row took
the row clock the hook lane pins to the main agent's turn start, re-billing the
whole watch-loop window. Edges now use the row clock when the row settles or
pauses and the evidence clock otherwise.

Rename isAgentExecutionOwed to isAgentTimeAccruing and state that it is the
stats question, not a liveness gate: it excludes watch loops, which lifecycle
gates must keep treating as live. Note on the plugin schema that
mainAgent.stateStartedAt is the execution host's clock.

* fix(agent-status): pause agent time while the row waits on the user, whoever asked

Time agents worked now accrues only while the combined row reads working. A
child's approval or question wait pauses the clock exactly like the main
agent's own prompt, and the pause edge is dated by the row's own clock.

* refactor(agent-status): read agent time from the combined row and date edges by the row's own state

Time agents worked now accrues while the combined row reads working and is not
a watch loop. The shared fold emits monitoring only for a settled main agent, and
hosts that predate the main agent fact did the same, so this is the same answer
on every new-host row without reading mainAgent, and it applies the watch-loop
rule to older hosts too instead of billing their monitoring windows.

An edge that leaves working is dated by the row's state clock; an edge inside
working is dated by the evidence clock. This also stops a live repeat of a
hydrated working row (any row without the main agent fact, such as an OSC row)
from dating its start at the persisted state clock from the earlier runtime.

* fix(grok): read a background subagent as agent work, not a watch loop

Grok's end-of-turn Stop lists each in-flight background task with its type
(shell, monitor or subagent). Orca filed a running subagent with the shells,
so a Grok subagent that outlived the main agent read "Monitoring background
tasks" and, with the stats recorder now skipping watch loops, stopped the
"Time agents worked" clock. Map shell and subagent entries to the shared
child-work kinds and let the shared liveness classifier decide: any live
subagent keeps the pane working, a shell alone or an active stop hook stays
monitoring, monitors stay excluded.

* test(agent-status): cover a waiting child in the fold's every-input accrual check

Since the shared fold learned a child's human wait, a waiting child makes the row wait, so it
must not accrue agent time whatever the main agent is doing. The exhaustive check now includes
that input.
2026-09-24 10:12:16 -07:00
Brennan BensonandClaude ad6cb0e05c fix(worktrees): version every catalog publication so a stale listing cannot undo a create (#22507)
* fix(worktrees): version every catalog publication so a stale listing cannot undo a create

A worktree listing is a snapshot from when its scan began. The renderer treated any
authoritative listing that lacked a known worktree as proof of deletion, judged at apply
time against the live store, so a listing delayed past a create reply purged the new
workspace: tabs wiped, selection cleared to the landing, pending structured launch
tombstoned so the host session was closed the moment it published. #22311 re-runs a scan a
mutation overtakes, which covers a bump during the scan but not a reply that is simply
applied late, on the host or in the renderer, or a refresh that joined an older one.

The host now stamps every listing with the catalog version its scan began at (the existing
per-repo scan generation, scoped by a per-process epoch) and every create and remove reply
with the version the mutation produced. Clients keep the newest version applied per repo
and host; a listing older than that is not applied at all, not its rows, not its purge, not
the pre-merge terminal teardown. Coalesced joiners inherit the reply and therefore the rule.
Fields are optional on the wire; older hosts and clients keep today's behavior.

* fix(worktrees): relist after a refused stale listing and stop version churn

- A refused listing can be the only answer a caller gets (a change-event
  refresh that joined an older in-flight listing), so fetchWorktrees lists
  once more; that listing scans at or past the applied version.
- An equal catalog version keeps the held object, so a no-op listing no
  longer writes store state on every refresh.
- Versions the client cannot order are treated as unstamped at ingest.
- worktree.rm takes the repo from its id selector instead of resolving the
  worktree a second time, which also stamped nothing for an id two hosts share.
- A removal on one of two hosts sharing a worktree id records its version.

* test(worktrees): pin the scan-generation bump right after git worktree add on every create path

A listing is stamped with the generation its scan began at, so a create must
advance it before any post-add work can yield. Pins the local desktop, SSH and
runtime local create paths.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(worktrees): bump the scan generation right after git worktree remove

A listing is stamped with the generation its scan began at. Removals bumped it
only at the end, after watcher release, push-target cleanup and the metadata
purge, so a listing scanned before the git removal and one scanned after it
could carry the same sequence. Applied out of order, the older one restored the
removed row until the removal reply. Bump right after the git removal on the
desktop local, desktop SSH and runtime paths, as creates already do.

The ordering pins now witness the generation at the first step after the git
mutation rather than at the re-list, so moving a bump past any intervening
await fails them.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(runtime): stop leftover worker-recovery retries from firing into later tests

The legacy worker terminal recovery retry timer re-arms itself and had no
way to end, so a runtime from one aggregator test kept rescanning repo-1
through the shared listing mock during later tests, consuming the listing a
lineage create expected ("Worktree created but not found in listing").

Give the controller a dispose() that cancels pending retries and refuses to
re-arm, and have the runtime test harness dispose every controller it
constructed after each test.

* Revert "fix(runtime): stop leftover worker-recovery retries from firing into later tests"

This reverts commit 5879110163. The leaked
recovery retry timer is a pre-existing test-harness flake that also hits
main; it belongs in its own change, not in the catalog-version fix.

* test(worktrees): pin the relist bound and the teardown gate on fetch-all and paired runtime listings

The relist after a refused stale listing had no test holding it to one retry, and the
pre-merge terminal teardown gate was pinned only on the direct fetchWorktrees path: removing
it from fetchAllWorktrees or from the paired-runtime listing path left every suite green.

* fix(worktrees): keep an SSH reconnect going when its listing is older than an applied create

A direct SSH listing refused because a newer catalog is already applied for
that host reported 'stale', the same result as a moved connection. The
reconnect preparation ends on any 'stale' repo and skips the post-connect
workspace sync and terminal correction, and nothing retries that while the
connection holds. The refused listing now reports the host's answer, since
the store already holds a newer catalog; 'stale' stays for a moved
connection or owner.

* test(worktrees): pin the listing teardown gate on the fetchAllWorktrees startup hydration pass

The hydration pass lists through its own call site, and removing its gate left every suite green.

* fix(worktrees): decide a listing's refusal reason inside the merge, and defer the startup purge behind a newer create

The listing merge now returns 'applied', 'superseded' (a newer catalog is already
applied) or 'not-current' (its connection or repo owner moved), decided against live
state inside the store update. Callers switch on it instead of re-deriving the reason
afterwards, which misreported an owner that went away during an older listing as current.

The one-shot startup purge keeps only ids from scanned rows. A create applied after a
repo's listing but before the purge wrote only the visible rows, so the purge closed the
new workspace's tabs, chat tab included. It now defers when any repo's listing is older
than that repo's applied catalog, like a refused listing, and runs on the next pass.

* test(worktrees): pin the startup purge deferral on a listing refused for a changed repo owner

* test(worktrees): pin a direct-authority fetch reporting an older listing as current

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-24 00:46:29 -07:00
Brennan Benson 25d7c21fcb feat(native-chat): show context window usage in the composer (#22301)
* refactor(native-chat): move the composer's stop action into its own hook

The composer is at its line budget; lifting the stop action out makes room
for the context usage ring without changing what Stop does.

* feat(native-chat): record the Claude CLI's context window facts on the structured journal

A structured Claude session now keeps what the CLI says about its context
window on journal rows, so every client reads the same answer and a restart
replays it:

- Each main-thread assistant response keeps its API usage. A subagent's
  response measures its own window, so it carries none.
- The turn a result settles records the session's window: the largest
  contextWindow across the result's per-model usage, since side calls to a
  smaller model report their own smaller window.
- After a result and after a compaction boundary the host asks the CLI for
  its /context breakdown (5s bound) and records the answer on the current or
  last turn. An answer is dropped when the main conversation moved, a send was
  accepted, a newer request was issued, or the session was released while it
  was in flight; a failure or an older CLI leaves the row unchanged.
- A compaction boundary or conversation reset records that the used count is
  unknown until the next response or report, so the pre-compaction size is
  never shown as current.

Every part carries its own host clock, and the reader takes the newest, since
a revised turn row keeps its place in the transcript. The persisted validator
admits every value the writer can write, including a zero auto-compact
threshold: a row replay rejects truncates the journal from that row.

* feat(native-chat): show context window usage in the composer

A structured Claude chat shows a ring beside send once the journal can state
the session's context usage. Hovering shows used/window with a bar and, when
the CLI has reported its breakdown, one row per CLI category as a share of the
window, largest first. Between reports the ring shows the newest response's
usage against the newest window the CLI reported, marked as estimated. Before
the CLI has reported any window, and after a compaction or reset until the next
response, there is no ring. A terminal-backed chat shows none.

* fix(native-chat): measure the context ring against the main thread's model window

The result's per-model usage is cumulative across the session and includes
subagents and side calls, so the largest window was often not the one the
main conversation runs in: after switching from a 1M model to a 200k one, or
when a subagent ran on a larger-window model, the ring read against the
wrong window after every turn. Pick the entry named by the model that served
the newest main-thread response, and among its [1m]/non-[1m] entries the one
the result moved; fall back to the largest only when nothing names it.

* fix(native-chat): keep the context ring moving through tool-only responses

The live estimate lived on assistant message rows, and a response with only
tool calls or thinking writes no message row, so the ring froze through long
tool loops and stayed hidden after a mid-turn auto-compaction until the next
text reply. Record every main-thread response's usage on its turn row
instead, once per response, so the selector sees each one.

* fix(native-chat): read the ring's window from the model the turn's init names

The CLI keys per-model usage by the main loop's model string, [1m] included,
and every turn's system/init frame carries that exact string, while a
response drops the suffix. Match the init's key first, so a session that
switched between the 1M and 200k variants of one model reads the right
window; fall back to the newest response's model, then the largest entry.

* fix(native-chat): correct the composer's control-order note for the context ring

* fix(native-chat): keep a running turn's context facts when the host settles it

A turn row now carries the live context estimate while it runs. When the host
settles a running row itself (a crashed or stale generation, a close the
translator never saw), it rebuilt the record field by field and dropped those
facts, so after a crash the ring fell back to an older turn's size, or to a
pre-compaction size the dropped reset had superseded.

* refactor(native-chat): revise Claude turn rows from the journal so the ring survives a restart

The context ring's facts were written to turn rows through an in-memory list
of recent turns. A new translator is built on every acquisition, so after a
restart or reattach that list was empty and every fact for a turn that was
not open was dropped: a /compact as the first action after a restart never
cleared the ring and never showed the fresh breakdown.

Every Claude turn-row write is now a revision of the row as the bound journal
holds it when the write runs. The sink gains a resolved revision that reads
the target row and its body at execution; the queue runs one operation at a
time, so the read-modify-write cannot interleave, and revisions are never
coalesced. Lifecycle writes own the lifecycle fields and context writes own
contextUsage; each keeps every other field. Only the open turn is kept in
memory. A fact with no open turn lands on the newest turn row, and a report
lands on the turn it was requested for.

The persisted facts are simplified to a window, which now names the model it
was measured for, and a single used part (report, estimate or unknown) that
each write replaces. The ring reads the newest turn row carrying each part,
and hides an estimate whose model the window was not measured for instead of
dividing by another model's window. A turn opening, and a reset, count as
activity, so a late report can never land behind a newer turn.

Host settlement of a stale running turn now drops only the fields its verdict
owns, so context facts and any field a newer build wrote survive it.

* fix(native-chat): keep a turn row whose context facts this build cannot read

Context facts are validated deeply, so one malformed or future-shaped fact
made the whole turn row malformed, and replay truncates the journal from that
row on. Replay now drops unreadable facts from a turn row, in item rows and in
settlement batches, and keeps the row, the same way it already drops producer
linkage it cannot trust. The ring shows nothing for that turn instead of the
session losing its history.

* test(native-chat): pin that a child exit mid-turn keeps the ring's last size

A lifecycle-only revision, the end a turn gets when its child exits without a
result, must keep the context facts the row already carries.

* perf(native-chat): revise a named Claude turn row by key instead of scanning the journal

Every Claude turn-row write walked every reduced journal item to find its
row, even when it already knew the row's identity, so a long session paid
O(items) per write on the main process. The journal now answers a keyed read,
and a context report names its turn by row identity rather than turn id, so
only a write made while no turn is open still scans.

* fix(native-chat): tell a 1M window from a 200k one of the same model

Responses drop the [1m] suffix, so after a switch between the 1M and 200k
windows of one model the running turn was measured against the previous
turn's window until its result arrived. An estimate now records the turn's
init model, which keys the window exactly, and the reader requires the full
model id to match.

* fix(native-chat): show no ring for a context kind a newer host writes

A paired client reads turn rows from the host unvalidated, so a used-count
kind this build does not know fell through to the estimate branch and threw
reading its missing usage. Only the kinds this build can measure now produce
a ring.

* fix(native-chat): keep the context ring through plan-mode turns on another model

Plan mode can run a turn on a model the turn's init does not name (opusplan
upgrades to Opus's 1M window). The estimate then carried only the response's
id, which drops [1m], and the exact comparison against the window hid the ring
for every plan-mode turn. The estimate now records the response's id beside
the init's exact key, and the reader matches the base model only when no exact
key was recorded.

* refactor(native-chat): pair the context ring's window by model change, not by model id

The ring divided the newest response's size by the newest window only when
their model ids matched, which meant comparing ids from the init frame, the
response, per-model usage keys and canonical ids. Those disagree in plan mode
and across 1M and 200k windows of one model.

The writer now knows when the model may have changed: after a model or
permission-mode write that changes the value, when a restore cannot put the
stored model back, and when a main-thread response comes from a different
model than the one the window serves (an approved plan). It then marks the
size unknown, holds estimates, and asks the CLI for its context report, which
states the new model's window. Any new window, from a report or a turn
result, releases the hold. The reader compares nothing: a report, or the
newest estimate over the newest window.

Turn rows no longer store window.model, window.canonicalModel,
estimate.model or estimate.responseModel.

* fix(native-chat): keep a late context report's window when only its count went stale

* test(native-chat): pin that each turn's init lets its result restate the context window

* fix(native-chat): open the context card on click and tap

* fix(native-chat): wait a beat before a mouse hover opens the context card

* fix(native-chat): publish each context write in the operation that makes it

A context report answers after the turn's last frame, so a revision that waited
for the next frame's publish reached live clients only on the next turn. Context
writes now queue their revision and its publication as one operation.

* fix(native-chat): write million-token counts with a capital M

A lowercase m read as minutes on the context card.

* fix(native-chat): keep the context card open while the pointer crosses into it

The card closed the moment a mouse left the ring, so the pointer could not
cross the gap into the card. Leaving now waits a beat, and entering the card
cancels the close.

* fix(native-chat): let Escape close the context card without stopping the agent

The card keeps focus in the composer, so the Escape that closed it also
reached the composer and interrupted the running turn. The composer now
skips an Escape an open layer already handled.

* fix(native-chat): show the context ring when the chat has not loaded the turn it belongs to

The ring read context facts only from the rows the chat had loaded, so a
reopened chat whose recent page started after the last turn row, or a live
turn longer than the retained window, showed no ring until the next turn.

The host now derives the newest context facts from its whole journal with
the same selector the chat uses, and returns them on agentSession.options
for sessions that write them. The chat prefers each fact its loaded rows
carry and takes the host answer for a fact they lack. When a live batch
revises a turn row older than the loaded window, the chat asks for options
again so that answer stays current.

* fix(native-chat): bound context refresh reads and refresh when the turn row is trimmed

Each turn-row revision the loaded window missed started its own options
read. Those reads share the session's host queue with sends and interrupts,
and each asks the CLI for its settings, so a burst could pile reads in front
of a user action and discard every answer before it landed. The chat now
keeps one options read in flight and at most one behind it.

A live turn longer than the retained window also lost its turn row to the
trim without asking for a fresh host answer, so the ring fell back to the
answer read at turn start until the next response. Trimming a turn row now
asks again, like a dropped revision does.

* fix(native-chat): show the context ring from the first response, sized from the session's model

A new session has no measured window until its first result, so the ring
stayed hidden for the whole first turn. The host now keeps the window the
applied model's name implies (1M for a [1m] name, unknown for default, 200k
otherwise) and writes it beside an estimate when the journal holds no window,
or after a model write, until the result or the CLI's report replaces it.

* fix(native-chat): imply a context window only from a [1m] model name

A bare model name does not fix the window: first-party runs today's opus,
sonnet and fable models natively at 1M while a gateway or cloud provider runs
them at 200k, and opusplan and haiku run another model in plan mode. Sizing
their first response at 200k read the ring about five times too full, so only
a [1m] name implies a window now; any other name waits for the result.

* fix(native-chat): size the first response from a report taken before any turn

A model picked in a chat with no turn yet asks the CLI for its context
report, but with no turn row the report's write lands nowhere. Recording it
still marked the journal as holding a window, so the first response wrote
none and the ring stayed hidden until the turn's result.

The report's window now serves as the fallback a response writes while the
journal holds no window, and recording a report no longer assumes its write
landed.

* test(native-chat): move the fake Claude connection out of the structured integration suite

The context-report delivery case pushed the suite past the 800-line limit,
failing repo-wide lint. The fake child now lives in its own fixture.
2026-09-24 00:12:08 -07:00
Jinwoo Hong ca75bc4c8d fix(orchestration): type a request ahead of pasted dispatch briefs so Claude workers follow them (#22582)
* fix(orchestration): type a request ahead of pasted dispatch briefs so Claude workers follow them

Claude Code wraps a bracketed paste in <pasted_content> and tells the model to
follow instructions inside it only where the user's own message asks. Orca sent
the whole dispatch brief as a bare paste, so Claude workers (Opus 5.5, Sonnet 5)
refused it as suspected prompt injection. Every dispatch path now types a short
lead line in the same PTY write as the paste frame, the preamble drops shouted
rules, and dispatch detection accepts the lead line and pasted_content wrapper.

Fixes STA-8200

* refactor(orchestration): tidy dispatch lead-line delivery after review

- Share one dispatchPreambleSendOptions() across the four dispatch paths.
- Fold every C0 control and DEL out of the typed lead line.
- Bound the <pasted_content> tag scan and let compaction return null for
  non-dispatch prompts, removing the separate detector.
- Restore the stay-off-other-channels rule in plain wording.
- Test through the real status normalizer and trim duplicated assertions.

Refs STA-8200

* test(orchestration): guard coordinator auto-dispatch lead line

- Capture send options in the coordinator runtime fake and assert the
  auto-dispatch send uses dispatchPreambleSendOptions.
- Drop the helper test that only restated its literal.
- Share DispatchPreambleSendOptions with the coordinator runtime contract.

Refs STA-8200

* docs(orchestration): fit the pasted-spec note inside the kernel line budget

Refs STA-8200

* docs(orchestration): drop the pasted-spec note from the coordinator guide

The typed lead line is the fix; the advisory note cost always-loaded context.

Refs STA-8200

* fix(orchestration): type the dispatch lead line only for Claude agents

Codex discards typed text that shares a PTY write with a bracketed paste,
so the lead line never reached it. Only Claude Code needs the lead to follow
a pasted brief, so known non-Claude agents now get the pre-lead bytes and
unidentified agents keep the lead in case they are Claude.

Refs STA-8200
2026-09-24 01:33:16 -04:00
Brennan Benson b4d732685c feat(agent-status): combine Codex child work through the shared main-agent status fold (#22475)
* feat(agent-status): combine Codex child work through the shared main-agent status fold

* docs(agent-status): correct two comments the waiting child-work arm made stale

A child failure reported in place as `blocked` now pins the row `waiting`, not
`working`; and no relay ever sent an unfolded `working` beside a waiting child.

* fix(agent-status): only a waiting child asks for a human

A child's `blocked` state means its task failed (the only producer maps a
failed background task to it, and the background-task view labels it
"failed"), not that a human must act. Folding it into the waiting arm would
surface a failed child as needs-you. It stays live work, as before this
series.

* docs(agent-status): say a waiting child, not a blocked one, makes the row wait

A child's blocked state means it failed; only its waiting state feeds the
waiting arm. Two fold comments, a test describe and two parity story names
still called the waiting child blocked.

* docs(agent-status): name where a child's wait is still lost, and pin the structured lane's real input

The doc said the Claude hook lane's rows match Codex and that every lane feeds a
child's wait into the fold. Neither holds: Claude keeps the wait in one slot the
next main agent event overwrites, the structured lane turns a child's prompt
into the main agent's own attention, and Codex drops its roster on a root Stop
when it tracks no child transcripts. The parity story now drives the structured
lane with the input it actually receives.
2026-09-23 22:09:07 -07:00
Brennan Benson 7a4f080086 revert: #18790 (orchestration incarnation reap fallback and bundled Freebuff agent) (#22601)
This reverts commit 0677271709.

#18790 was merged as one squash commit that carried two unrelated changes:
a process-incarnation fallback for reaping leaked orchestration worker
terminals, and an unannounced "Freebuff" third-party agent (catalog entry,
icon, locale strings, README rows). The Freebuff agent was never meant to
ship, so the whole PR is reverted; the reap fix should be re-submitted on
its own.

Until that re-land, a worker whose durable terminal handle goes stale is
again reported missing on release/stop instead of being re-found through
its process incarnation, so its terminal can leak on Remote Server.

The mobile session page closure pin moves 4218 -> 4219: the revert drops
the freebuff icon #22119 pinned (-1), and #22452 had already added two
src/shared modules without re-pinning (+2).
2026-09-23 21:39:28 -07:00
Brennan BensonandClaude 3ea15dd0a2 fix(native-chat): keep chats that failed to resume in the status bar and say what to do (#22448)
* fix(native-chat): keep chats that failed to resume in the status bar and say what to do

After a restart, a chat whose resume did not carry on was reported only by a
four-second toast that named nothing, and the status bar entry vanished because
the reattach had already spent the offer.

The host now files the outcome as a durable `failed` entry in the recovery
capsule, with the refusal code and the prompt the offer quoted, and returns it
from the restart-resume RPCs. The renderer shows a "N chats failed to resume"
status bar entry, a count-only toast with Show and Dismiss, and keeps the resume
dialog open with a status icon per row and a "To resume" line whose action is
chosen from the reason. A failure dies on dismiss, on a successful retry, on the
user's own send in that chat, or with the marker's 24h expiry.

* fix(native-chat): keep the resume dialog unchanged and add failed chats as rows

The failure view had replaced the resume dialog's title, checkboxes, preference
box, and footer. The dialog is back to main's layout. A chat an earlier resume
could not carry on is now an ordinary selectable row there, plus a status icon
whose tooltip carries the reason, a dismiss control, and a "To resume" line.
Selecting it and pressing Resume retries it; it is pre-selected only when a
retry can succeed. Resumed chats leave the list as before.

Also stubs the new failure listing on the cross-version wire host fixture, which
the restart-resume RPC now reaches.

* refactor(native-chat): release a failed-resume record where a send is admitted

Keeps the host file at main's size, and only releases the record for a send the
controller actually lets through.

* fix(native-chat): note a failed restart resume in the chat and derive when it is settled

The chat itself now says when Orca could not continue it after a restart,
with an error (refused) or warning (unconfirmed) status row, so the failure
survives the toast, a dismissed record, and another restart.

A recorded failure is current only while the chat's newest user message is
the one it had when the failure was filed. Listing re-derives that from the
journal and prunes superseded records, replacing the in-memory set and the
hook on every send.

Failures move to their own optional top-level key in the recovery file, so an
older build that rejects unknown entry states keeps reading its offers. A
retried failure stays a failure through a rollback or a lapsed lease instead
of returning as a pending offer, and the toast's Dismiss names only the chats
the host listed as failed.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): check the older reader against a filed failure before any retry

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): file a failed restart resume against the chat as its attempt ended

A failure was filed against the chat's newest user message read at settlement, after every chat in
the action had finished. The chat's note asks the user to send a message, and one sent while other
chats were still being continued became part of the filed state, so the failure stayed listed after
the user had done what it asked. Each chat's newest message is now observed as its own attempt ends.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): let a reply that made a chat ineligible retire its failure

When a restart resume reached a chat the user had already replied in, the
attempt was refused as no longer eligible, but the failure was filed against
that very reply. It then stayed listed as "finished on its own" until the user
sent yet another message. An ineligible chat is no longer observed at the
attempt, so its failure falls back to the reserved marker and the reply that
made it ineligible supersedes it at the next listing.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): offer Retry on a failed resume only when a retry would run

When the agent refused Orca's "continue" message with its own reason, the
failed row fell to the generic guidance, which offers Retry and pre-selects
the chat in the resume dialog. The refused message is already the chat's
newest user message, so a retry is never eligible: it did nothing and the
same "couldn't be resumed" toast came back.

The host now reports whether a retry would run, derived at list time from
the same predicate the retry applies to the failure's marker. Where it would
not, the row offers Open chat and Dismiss and is not pre-selected. An older
host omits the flag and the reason alone decides, as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): file only restart failures the user must act on

A chat that stopped being resumable between listing and acting (it
finished on its own or is waiting on the user) was filed as a failure
with no note in the chat. It now just spends the offer.

A failed reattach now writes the same in-chat note as a refused
continuation, so every filed failure explains itself in the chat.

An unconfirmed continuation's failure retires once the chat shows the
continuation's own message opened the newest turn.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): clear a failed resume from the status bar once its chat moves on

While a failed resume is listed, the restart store watches the host's
status feed; when a failed chat's status or latest prompt changes after
the list was read, it re-reads the host once. The host still decides
whether the failure stands. Nothing is watched while nothing failed.

The failure toast now counts only the requested chats the host still
lists as failed, keeping the old count for a host that sends no list.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that an unjournaled continuation never retires its unconfirmed failure

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop telling the user to send a message into a chat Orca couldn't reconnect

A failed reattach, or a continuation refused because another window or terminal owns the
session, now leaves a note saying Orca couldn't reconnect the chat instead of advising a send
that would meet the same refusal. The restart list keeps the reason-specific advice.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the failed-chat re-read from undoing an action or missing a reply

The status-bar re-read no longer runs while a resume or dismiss is in flight, and its answer is
dropped if one settled meanwhile, so a dismissed failure cannot come back. A change to a failed
chat already seen always triggers it, whatever the host's timestamp says. A chat whose unconfirmed
continuation the host already retired is now reported as resumed instead of saying nothing.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that no failed-chat re-read runs under a resume in flight

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): typecheck the unjournaled-continuation case against a nullable marker

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): drop the stale "no arguments" note on the restart-offer params

The dismiss call now names sessions, so the older comment contradicted the schema below it.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): count unconfirmed resumes apart from refused ones in the action toast

The post-action toast said "N chats couldn't be resumed" for chats whose
continuation may well have gone out, while the list and the chat itself say
Orca couldn't confirm it. That wording invites a duplicate "continue" send.
Unconfirmed chats now get their own count, classified by the outcome the
host filed, so the toast matches the row it points to.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop offering Resume on a failure the host says cannot retry

A failed row the host marks unretryable could still be ticked, sending a resume
that could only fail again; its checkbox is now disabled and it never joins the
action. An older host that omits the flag keeps today's selectable row.

The status bar no longer calls a chat "failed to resume" when the host only
couldn't confirm the resume, matching the dialog's own wording, and the mixed
toast's second line now says "other" so it cannot read as the same chat.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): skip unreadable failure records instead of rejecting the capsule

A failure record this build cannot parse (a newer outcome, say) made the
whole recovery file unreadable, so a downgraded build listed no restart
offers and could not record new teardowns. Failures are advisory: parse
each one on its own and drop what does not parse.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a failure record with an unreadable marker is skipped

The skip-unreadable-failure test only covered an unknown outcome, so going back to the throwing
marker parser for failure records still passed. A failure record usually outlives its offer
entry, so a newer marker shape can appear only there.

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-23 20:45:07 -07:00
Jinwoo Hong 8ba7f829ac feat(ipynb): run notebook cells in a persistent Jupyter kernel (#22581)
* feat(ipynb): run notebook cells in a persistent Jupyter kernel

Replaces the fresh-process runner (which silently re-ran every earlier cell)
with the user's own ipykernel, driven by a small bundled Python bridge over
line-framed JSON. One kernel per open notebook: started on first Run, shut
down when its tab closes or Orca exits (stdin EOF), and the kernel's own
parent poller reaps it if the bridge dies.

The header gains a kernel pill (workspace .venv/.conda recommended, PATH
interpreters, Browse), Interrupt/Restart/Run all/Clear all, and a one-time
missing-ipykernel dialog with Install. Outputs stream live per cell and are
written into the document when the run finishes.

* test(ipynb): cover the stalled-interrupt restart offer

* refactor(ipynb): disable the kernel pill while settling; merge its classes with cn

* refactor(ipynb): tie kernels to their renderer document and simplify the run flow

- Main keys kernels per renderer, so a reload, renderer crash or closed
  window shuts them down, and two windows never share one notebook kernel.
  One close-driven cleanup replaces the separate exit and start-failure
  deletions.
- The first run uses the nearest workspace env, else the first Python on
  PATH; the picker no longer opens itself, so its open state stays in the
  toolbar. Closing the picker brings the missing-ipykernel dialog back
  instead of dropping the queue, which also keeps a Browse pick's run.
- Discovery marks the kernel starting, so a second run during it queues
  instead of starting a second kernel, and a tab closed mid-discovery no
  longer leaks one.
- Running an nbformat 4.4 notebook gives its cells ids (upgrading to 4.5),
  so moving a cell mid-run cannot misroute its output.
- The death notice drops stderr from before the kernel was ready (the
  unencrypted-TCP warning).
- SSH and non-Python runs toast instead of writing a notice into the cell.
- Windows conda envs are named after their folder.

* fix(ipynb): install ipykernel into envs without pip

uv-created venvs ship without pip, so Install failed with 'No module named
pip' there. When pip is missing, bootstrap it with the stdlib's ensurepip
and retry. The install moves beside the other interpreter probes, and the
copyable install command comes from one helper.

* fix(ipynb): address PR review comments on stream errors, old jupyter_client and the Windows venv hint

- Swallow stdout/stderr stream errors on the bridge child, as spawnProcess
  requires, so a broken pipe cannot crash main.
- The bridge exits (reporting the death) even when cleanup_resources is
  missing (jupyter_client < 6.1.5) or raises.
- The install-failure hint suggests `py -m venv .venv` on Windows.

* feat(ipynb): add Cancel to the missing-ipykernel dialog

It does what Esc does: drops the cells waiting on the kernel.

* fix(ipynb): recover from a rejected kernel start; quote the install command per shell

- Discovery moves into start, so one catch turns a rejected
  listPythonEnvironments or startKernel into the usual failed start: the
  session returns to off with the error in the cell, instead of sticking
  at starting.
- The copyable ipykernel command quotes the interpreter only when its path
  has whitespace, prefixing PowerShell's call operator on Windows. Install
  itself still spawns without a shell.

* fix(ipynb): always shell-quote the copyable ipykernel install command

Quote the interpreter path for every path, not only ones with whitespace,
so paths with shell metacharacters like & copy as a working command.
Single quotes are literal in POSIX shells and PowerShell; embedded quotes
are escaped per shell, and PowerShell keeps its & call operator.
2026-09-23 23:38:33 -04:00
Neil 795b64b9a6 docs(tui-agent-config): correct the OpenCode readiness-budget rationale (#22593)
The comment merged with #22546 claimed ConPTY never forwards DECSET 2004 and
that the signal therefore cannot fire on Windows. Verification on two real
Windows hosts refuted that: the sequence arrives in order on both ConPTY
backends, and the readiness signal fired in every run.

The budget was the actual problem — opencode does not enable bracketed paste
until ~4.8s and its composer is not ready until ~10s, so the 8s default expired
first and the draft was pasted blind. Same fix, accurate reason.

Note the claim that seeded this: terminal-agent-paste-bracketing.ts says 2004
"can be lost by remote replay or ConPTY", which is careful and not contradicted
here; the absolutism was mine.
2026-09-23 20:29:00 -07:00
Brennan Benson 5e3effc32f fix(native-chat): show every user message on the message rail, not just loaded ones (#22558)
* fix(native-chat): show every user message on the message rail, not just loaded ones

The message rail was built only from transcript rows the renderer had
loaded, so any prompt above the loaded page had no tick, and the rail lost
ticks when a long live session trimmed its retained window.

The host now answers `agentSession.conversationOutline`: every user message
in a structured session's journal (item id, creation sequence, a preview
cut to 200 characters, image count) plus the journal position it is current
through. It is derived from the reduced journal on each request with the
same projection the transcript runs, so an entry's id and preview are what
the loaded row shows. The reply is bounded like a history page: previews
shorten, then drop, and only then do the oldest entries go.

The renderer asks only while the pane is visible and older history is
unloaded, uses outline entries only for messages older than its loaded
window (the window is authoritative for the rest), and falls back to loaded
messages while the outline is stale (epoch change, or the window trimmed
past what it covers). Selecting a tick with no row pages older history in
until the row exists, then uses the existing rail jump.

The method is negotiated with `agent-session.conversation-outline.v1`; a
client never calls a host that does not advertise it, and any failure
leaves the rail on loaded messages.

* fix(native-chat): keep a rail jump from the bottom from re-arming follow and cancelling itself

A rail jump started by a reader following the end stopped a few pixels
above the bottom instead of reaching the message. Paging older history in
for a jump always leaves the reader following at the very end, so jumps to
unloaded messages hit it every time; a jump to a loaded message from the
end did too.

The jump scrolls smoothly, and only its landing is marked as the
application's own scroll. Its first frames move a pixel or two, still
inside the band where a reader event re-arms follow, so the list read the
jump leaving the end as the reader arriving at it. The next frame, just
outside the band, then read as the reader taking over and rebased the view
with an instant write, which cancels the smooth scroll.

Re-arming follow now needs the reader to be arriving at the end: an
unmarked event that moved the view up never reattaches a detached reader.

* test(native-chat): check a rail jump left the end before reading where it landed

* fix(native-chat): keep the rail's message list still while a press selects an item

With the whole conversation in the rail, its hover list overflows and opens
scrolled to the message being read. Clicking an older item did nothing:
pressing it focuses it, which turns the hover preview interactive, and that
switch re-ran the effect that scrolls the lit row into view and focuses it.
The list moved under the pointer between press and release, so the click
landed on the list instead of the item, and focus jumped to the lit row.

Revealing the lit row now follows the list opening (and its rows shifting),
not the switch between hover and interactive. Entering interactive moves
focus into the list only when focus is not already on one of its items.

* perf(native-chat): keep the rail's outline entries stable while a trimmed window slides

A long live session holds a head-trimmed window, so every new row moved the
oldest-loaded edge and rebuilt the outline view even when no user message
crossed it. The rail then re-merged, re-rendered and re-read the scroll
geometry on each new row. The view is now reused while the set of entries
older than the edge is unchanged.

* fix(native-chat): let a rail jump wait out an older page already loading

Scrolling to the top of the loaded window asks for the next older page. A rail
jump made while that page was in flight asked again, got the lane's immediate
no-op return, read it as a page with no progress, and dropped the click. The
jump now waits for the in-flight page to land before deciding.

* fix(native-chat): reattach follow when content shrinking clamps a reader onto the end

The rule that stops a smooth scroll leaving the end from re-arming follow
compared offsets, so it also refused a reader whose offset dropped because
settled content folded away beneath them and the browser clamped them onto
the end. They sat at the bottom without following, and the next reply grew
out of view. Re-arming now requires closing on the end rather than moving
down, which still rejects a scroll leaving it.

* fix(native-chat): keep the rail hooks' ref writes out of render

Both hooks wrote a ref while rendering, which React may replay or discard.
The rail's structural-sharing baseline is now recorded after commit, and the
history jump calls the lane's page loader from its effect instead of through
a render-updated ref.

* fix(native-chat): let the latest rail pick win over a jump still paging

A jump to an unloaded message keeps paging older history until it lands. A
pick made meanwhile lost to it: a loaded message scrolled into view, then the
earlier jump finished and pulled the reader away; another unloaded message was
ignored. Picking a loaded message now cancels the paging jump, and picking an
unloaded one retargets it without asking for a second page.

* fix(native-chat): step a rail history jump with a functional update

The step that requests the next page wrote the pending jump from the value
its effect closed over, so a pick or cancel queued since that commit would
be overwritten.

* refactor(native-chat): run a rail history jump as one abortable awaited loop

The jump through unloaded history was an effect-driven state machine that
guessed "no progress" from the message list's identity and could not be
cancelled by anything but another rail pick. A diff reveal, "Jump to latest"
or the reader scrolling left it paging, and when its page landed it pulled
the reader away; a history read that kept failing during a live turn could
repeat back to back.

Loading an older page now reports how it ended, and a second request while
a page is in flight joins it instead of being refused. The jump is an
awaited loop that reads the rail from a commit after each page, stops on
anything but a page that moved the window, and is aborted by any other
navigation, reader input (wheel, touch, scroll keys, scrollbar), a session
switch or unmount. The latest pick wins.

* fix(native-chat): derive the rail outline from the transcript's own projection

The host built the outline from user items alone, while the transcript
orders every message by when it was observed, folds tool results into the
turn above and then drops harness turns. An imported user row carrying a
tool result beside harness text therefore got a rail tick previewing the
harness text, and clicking it paged history for a row that never draws.

The transcript's order-fold-strip projection now lives in one shared
function. The renderer's list projection wraps it with its own tail-row
order, and the host runs it over the whole journal and keeps the user rows
that draw content, so the outline lists the same messages in the same order.

* fix(native-chat): retry a failed rail outline read a few times

A failed outline read left the rail on loaded messages until a new gap
opened, the pane was shown again or the epoch changed. The client now
rejects a failed read (a host without the outline still resolves to
nothing, without being called), and the rail retries up to three times
with doubling backoff.

* test(native-chat): give the rendered-transcript fixture the older-page result contract

* perf(native-chat): sort the shared transcript projection without a spread copy

* fix(native-chat): let a wheel over the rail cancel a rail jump still paging

The rail forwards its wheel to the transcript, so a reader scrolling there is
scrolling the transcript. That wheel never reached the scroller's reader-input
handlers, so the jump kept paging and later pulled the reader to its target.

* fix(native-chat): keep the rail's list open while a picked message pages in

Picking a message that is not loaded yet can take several pages of older
history. The list closed on the pick, so its busy item was never seen and
the click looked ignored. The list now stays open with that item pulsing
until the jump lands or is abandoned, and stops revealing the lit row
meanwhile so the rows do not move under the pointer.

* fix(native-chat): keep the shared transcript projection loadable on mobile

The projection moved to src/shared, which mobile's Hermes engine also loads,
and switching its sort to toSorted broke the Hermes compatibility guard.
Sort a copy made with Array.from instead, as other shared code does.
2026-09-23 20:26:54 -07:00
5802b54579 fix(rate-limits): read OpenCode Go usage with the Go API key (#22551)
* fix(rate-limits): read OpenCode Go usage with the account API key

Since OpenCode's console migration (upstream fe51b0b19a, "fix(console):
restrict legacy access to Black"), an account with no Black subscription
is redirected from the legacy console to /console/login, so Orca's
cookie-based workspace lookup returns nothing and the Go bar stays empty.

Fetch usage from GET https://opencode.ai/zen/go/v1/usage instead, which
authenticates with `Authorization: Bearer <key>` and needs no console
session. The key resolves in order: Orca settings override,
OPENCODE_API_KEY, then whatever OpenCode itself stored on /connect --
auth.json for 1.x, the credential table for 2.x. The cookie path stays
as the fallback so Black/legacy accounts keep working.

A 403 EntitlementError now reads as "no OpenCode Go subscription" in the
status bar instead of a generic refresh failure (#22257's reporter was
misled by exactly that).

* fix(rate-limits): prefer OpenCode's stored Go key over OPENCODE_API_KEY

OpenCode applies the key saved on /connect after the environment, so the
stored key is the one its own Go requests use. OPENCODE_API_KEY is also
the Zen provider's variable, so ranking it first could read a key that
OpenCode itself is not using for Go.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(rate-limits): name the API key when OpenCode Go usage lands on sign-in

A redirected usage request arrives as a 200 sign-in page because Electron
follows redirects; report it as a rejected key instead of a parse failure.
The cookie path's empty workspace lookup is what non-Black accounts now
hit after the console migration, so its message points at the API key
rather than only the workspace override.

Co-Authored-By: Claude <noreply@anthropic.com>

* chore(i18n): add the OpenCode Go API key strings to the English catalog

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(rate-limits): stop calling the credential table an OpenCode 2 marker

Verified on two real Windows hosts running OpenCode 1.18.16: the `credential`
table exists there too (empty, same columns), so its presence does not identify
a 2.x install. Neither host had an `auth.json` at all.

The resolution already probes both stores on every version, so only the comments
were wrong. Says so now, and records that a 2.x install which never ran the
legacy import has no `auth.json` either — which is why both tiers exist.

* refactor(shared): move GhosttyImportPreview out of global-settings-types

Adding `opencodeGoApiKey` pushed global-settings-types.ts one line past the
300-line ceiling, failing `oxlint` in CI. AGENTS.md forbids a max-lines
suppression, so split instead: the Ghostty import preview is a distinct concern
that never belonged in the settings-shape file.

Re-exported from the original module so no importer changes. 293 code lines now.

---------

Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
2026-09-23 20:09:02 -07:00
Neil b0ae7d18a0 fix(opencode2): resolve subagent session lineage so child work stops taking over the pane (#22444)
OpenCode 2's plugin adapter unwraps a single-property `{ data }` success schema,
so `ctx.session.get` resolves to the bare session record. The shared lineage
lookup only accepts `result?.data?.id === sessionID`, and OpenCode 2 has no
`session.list` fallback, so `resolveRootSessionID` returned null for every
session and `childState` was permanently null.

With unknown lineage `canFailOpen` is true for attention events, so a subagent's
`permission.asked`/`question.asked` fell through and pinned an un-evictable
blocker keyed to the child's own session id — publishing a subagent as if it
were a root. Observed in hook posts: SessionBusy for a child session id whose
`session_v2` row carries a parent.

Envelope the result in the OC2 client shim so the shared lineage module works
unchanged; OpenCode 1 already receives enveloped results and is untouched.

Also adds `opencode2` to the double-Escape interrupt list, extracted into one
shared helper so the server inference and renderer gate cannot drift. A single
Escape was inferring an interrupt, and Escape is how the Subagents dock closes.

7 of 11 new lineage tests fail without the shim.
2026-09-23 20:08:01 -07:00
Neil b7a4fee700 fix(agents): stop claiming an unconfirmed OpenCode handoff succeeded (#22546)
* fix(opencode): stop claiming a handoff prompt was delivered when it was written blind

"Continue in New Session…" to OpenCode reported success even when the prompt
never reached the TUI (#22479). The paste-after-ready helper falls back to a
blind write when the composer-ready signal never arrives and only the agent
process is known to exist; that write was indistinguishable from a real
delivery, so the continuation showed its success toast.

- pasteDraftWhenAgentReady / pasteDraftToAgentPtyWhenReady report the blind
  fallback via onUnconfirmedDelivery, plumbed to launchAgentInNewTab as
  onPromptDeliveryUnconfirmed.
- The session continuation hedges instead of claiming success, and both the
  failure and the hedged notice offer "Copy prompt".
- OpenCode gets Codex's 20s composer budget. Both are quiet-window-less
  signals anchored on DECSET 2004, which ConPTY never forwards, so on Windows
  that budget is the settle delay before the blind paste.

* test(runtime): retarget the 8s startup budget test off OpenCode

The main-runtime startup-draft budget test used `opencode` as its stand-in for
"an agent without an override", which this branch invalidates by giving OpenCode
20s. It failed with "expected vi.fn() to not be called at all, but actually been
called 1 times" — the readiness signal now legitimately arrives inside budget.

Point it at `claude`, which still takes the 8s default, and add a companion
pinning OpenCode's 20s: a readiness signal at t+10s, past the old default, must
now deliver the draft. Removing the override makes that companion fail.
2026-09-23 19:47:47 -07:00