Commit Graph
2457 Commits
Author SHA1 Message Date
OrcaWinandm4air 6d1a97ef98 fix(ssh): launch the Windows relay outside sshd's job so standard users work (#24224)
* fix(ssh): launch the Windows relay outside sshd's job without WMI

Win32-OpenSSH kills a session's job on close but allows breakaway. relay.js
gains a one-shot launcher mode that starts the detached relay with
CREATE_BREAKAWAY_FROM_JOB through the staged process-tree addon, so a standard
user no longer needs a WMI Remote Enable grant. WMI stays as the fallback for a
relay without the addon, and a refusal there is named. The Windows SSH-host
lanes drop their WMI grant and assert the breakaway route and adoption.

* fix(ssh): find runtime holds without WMI on a standard-user Windows host

The store GC read held runtimes through Get-CimInstance Win32_Process, which
WMI refuses to a standard user's SSH logon, so the pass kept every runtime.
On a refusal it now reads this account's own process image paths through
Get-Process.

* build(relay): ship the Windows relay launcher addon in every desktop package

macOS and Linux packages carried Windows relays without windows-process-tree.node,
so a legacy-runtime relay they uploaded to a Windows SSH host could not launch
outside sshd's job and fell back to WMI, which a standard user is refused.

A reusable Windows job now compiles the x64 and arm64 addons once and uploads
them; release-cut, release-mac-build, and the hourly/daily/adhoc mac builds
download them before build:release and require both arches. Staging now rejects
a binary with the wrong PE machine, the ReadProcessMemory import, or no
spawnOutsideJob export, so a stale pre-launcher build cannot ship.

* ci(ssh): run the Windows SSH-host lanes when the relay process-tree build scripts change

The staging and gyp-rebuild scripts decide which windows-process-tree addon the
relay ships, so a change to either must re-prove the Windows host cells.

* test(ci): find the mac orcad-template download by artifact name

The release mac job now also downloads the relay Windows process-tree addons, so
the first download-artifact step is no longer the template's.

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 05:32:09 -07:00
8afa1db50c feat(ssh): rung B glibc 2.17 compat runtime; gate remote vault on host node:sqlite (#24148)
* feat(ssh): wire rung B to the glibc 2.17 compat runtime; gate rung C vault on full node:sqlite

- COMPAT_RELAY_RUNTIMES lists linux-x64-glibc217; rung B plans the compat slot and compat
  pinned Node when glibc is below 2.28 or rung A refused with libc_floor/missing_lib.
- The relay version folds the compat runtime's executable hash; refusals are cached per runtime.
- The orcad template stages an optional linux-x64-glibc217 target (base package + compat
  node-pty slot + compat runtime marker); the verifier and materializer accept it.
- node-pty slot loader falls back to the compat slot when the default slot is missing or
  needs a newer glibc.
- Runtime store GC keeps the compat pin beside the default one on every relay connect.
- hasNodeSqliteReaderApi (DatabaseSync + backup) gates relay session search and the relay
  OpenCode reader, which now names the host Node version in its unavailable reason; the SSH
  vault reader installs the compat Node on old-glibc hosts and uploads nothing when no
  pinned Node can run.
- Rung D: a remembered noexec reports home_noexec and never advises installing Node.

* fix(ssh): re-prove a replayed noexec after rung D so allowing exec recovers the host

* fix(ssh): keep the rung B compat runtime pinned in the relay-connect store GC

* test(ssh): mock deployment-target facts in the Windows OpenCode runtime tests

* ci(ssh): build the glibc 2.17 compat slot for the hostile-host matrix; CentOS 7 lands on rung B

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>
2026-10-01 05:32:05 -07:00
dfdcfcf61f feat(ssh): plain SSH terminals and SFTP browsing when no Orca runtime can run (#24147)
* feat(ssh): connect in plain SSH mode when no Orca runtime can run on the host

Runtime ladder rung D (design D6): instead of failing the connect, register an
ssh2 shell-channel PTY provider and an SFTP-only filesystem provider and publish
the classified reason on the SSH connection state.

* fix(ssh): harden plain SSH mode against stale reconnects, host sleep and tilde cwd

- Only the current connect or reconnect attempt may enter plain SSH mode; a superseded
  reconnect whose ladder ends at rung D no longer registers a second provider set.
- Host-sleep resume probes a plain session over SFTP instead of always reconnecting,
  which ended every open plain shell.
- A home-relative cwd keeps its tilde outside the quotes so the shell expands it.
- The SFTP provider implements folder download, which the connect state advertises.

* docs(ssh): rung D now means plain SSH mode, not a failed connect

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>
2026-10-01 03:24:26 -07:00
OrcaWinandm4air 67014c8c60 feat(ssh): pinned-Node relay on Windows SSH hosts (#24135)
* feat(ssh): pinned-Node relay on Windows SSH hosts (D5 Windows, D2)

Windows hosts opted into remoteRuntime 'pinned-node' now get the same rung A
relay POSIX hosts do, instead of an early host-Node fallback.

- Runtime store: the official node-v24.21.0-win-<arch>.zip is uploaded to a
  stage under %USERPROFILE%\.orca-remote\runtimes, verified against the pinned
  archive hash, node.exe extracted with System32 tar.exe (Expand-Archive
  fallback), hashed with Get-FileHash, run once, and published with
  node.exe + .verified by one Directory.Move. One powershell.exe per phase via
  the existing powerShellCommand helper; the probe also creates the stage. No
  new -EncodedCommand site, no -ExecutionPolicy, no Add-Type. node.exe keeps
  its real name at runtimes\node-<sha>\node.exe.
- Bytes that change or vanish after Orca wrote and verified them are reported
  as ORCA_NODE_RUNTIME_SECURITY_MODIFIED and become a remembered
  'security_software' refusal (fallback to the host-Node relay); application
  control blocks classify as 'noexec'.
- Addons: the win32 slot's conpty.node, conpty_console_list.node,
  conpty\conpty.dll + OpenConsole.exe, watcher and windows-process-tree.node
  ride with the relay; the orcad template now carries the win32 targets.
- Self-test on Windows is one powershell.exe running relay.js on node.exe; the
  report must name the pinned Node. The relay self-test loads conpty.node and
  opens a PTY with useConptyDll, and reports a missing bundled ConPTY file as a
  load failure. A pinned relay's terminals use the bundled ConPTY too; host-Node
  relays are unchanged.
- describeRelayRuntime recognizes the Windows store layout.

* fix(ssh): skip the redundant stage-cleanup powershell.exe after a Windows runtime promote

The promote script already removes its stage on every path, so the client-side
cleanup only runs when promote never returned (upload failure, abort, timeout).

* test(ssh): expect the ladder's remembered flag and pin check on Windows pinned relays

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 02:34:58 -07:00
OrcaWinandm4air 53fd2dea0b feat(ssh): relay runtime fallback ladder, telemetry and host runtime setting (#24133)
* feat(ssh): complete the relay runtime fallback ladder (D6 rungs B slot, C, D)

Rung C runs the relay on the host's Node >= 18 with Orca's prebuilt N-API
addons and no npm (addon-only probe mode). Rung B is a data-driven slot chosen
only when a compat runtime is listed. Rung D fails the connect with a
classified reason carried as a TerminalUnavailableCause. The ladder steps
down only on classified refusals; unanswered probes throw. The rung decision
is persisted per host keyed by (glibc, runtime hash, Orca major), and
ssh_remote_runtime_resolved reports it once per host per session.

* feat(settings): SSH host runtime choice (Auto | Orca-managed Node | Host Node)

* docs(telemetry): describe ssh_remote_runtime_resolved

* fix(ssh): let a passing rung C disprove a remembered noexec; allow glibc-less compat runtimes

A remembered rung A noexec was re-persisted even after rung C self-tested addons from the same
~/.orca-remote tree, so rung A stayed skipped until the key changed. Rung B's evaluator also could
never match a musl compat runtime.

* test(ssh): import node:fs once in the host-node addon test

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 02:08:26 -07:00
OrcaWinandm4air a5601375d4 feat(ssh): opt-in SSH relay on the pinned Node with prebuilt addons (#24129)
* feat(relay): runtime self-test flag and informational runtime on handshake-ok

relay.js --orca-runtime-selftest <nonce> dlopens pty.node, opens and closes a
PTY, and prints one JSON line (nonce, node, napi, glibcVersionRuntime) for the
client to classify before it launches a daemon on a runtime (design D5).

handshake-ok gains an optional runtime {kind, version}; bridge and daemon
already match exactly on version, so it is informational only (D8.1).

* feat(ssh): opt-in pinned-Node relay with prebuilt addons (D5, D6 rung A, D8.1)

SshTarget.remoteRuntime (legacy | pinned-node, default legacy; env
ORCA_SSH_REMOTE_RUNTIME for development) selects the runtime. On POSIX hosts
the pinned path resolves the target with its glibc major.minor, ensures
~/.orca-remote/runtimes/node-<sha>/bin/node, uploads the relay bundle plus
the target's node-pty slot and @parcel/watcher from the orcad artifact
(no npm or node-gyp on the host), writes .runtime-ref-node-<sha>, and folds
the runtime and addon digests into the relay version so pinned and host-Node
builds never share a dir or socket.

A 30 s self-test (node --version, then the relay self-test) gates
.install-complete. Timeouts and lost channels are unverifiable and never step
down; noexec, missing_lib, libc_floor, illegal_instruction and wrong_libc
refusals fall back to the untouched host-Node path with a logged reason,
remembered for the session.

* fix(ssh): only an answered libc probe steps the pinned relay down

A lost channel during target detection says nothing about the host; descending
would launch a host-Node daemon beside a running pinned one and strand its sessions.

* test(ssh): mark the mocked SSH connection casts in the pinned relay tests

* test(ssh): resolve the pinned runtime mock to an executable path

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 01:41:14 -07:00
OrcaWinandm4air ddd4927a0b build(orcad): server node-pty slots at glibc 2.28, plus a glibc 2.17 compat slot (#24134)
* build(orcad): build server glibc slots on glibc 2.28 and add the glibc 2.17 compat slot

Design D6: the default linux-{x64,arm64}-glibc node-pty slots now build in
manylinux_2_28 (digest-pinned) and are gated at glibc 2.28 / GLIBCXX_3.4.25
through a floor profile on verify-linux-glibc-floor.cjs; the desktop keeps
its Ubuntu 20.04 (2.31) default.

Adds the opt-in linux-x64-glibc217 compat target: NODE_RUNTIME_COMPAT_ASSETS
pins the unofficial glibc-217 Node (update/check pin scripts cover it, outside
SERVER_TARGETS), and a new CI lane builds the compat slot in manylinux2014
with static libstdc++, gates it at glibc 2.17 with no shared C++ runtime in
DT_NEEDED, and smokes it under the glibc-217 Node.

* refactor(node-runtime-pin): route compat lookups through isCompatServerTarget; keep the glibc doc's slot-name paragraph intact

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 01:10:57 -07:00
OrcaWinandm4air 9ddcc9b9f0 fix(ssh): collect relay versions only when provably exited; runtimes/ store GC (#24130)
* fix(ssh): relay version GC deletes only on an exited verdict and keeps the previous build

The relay records .relay-pid in its version dir once it owns its socket. GC calls a
relay version dir exited only when that PID is provably dead and every relay-*.sock
refuses a connection; a dir without a PID file keeps the test -S rule. The most
recently completed other relay build is pinned like orcad's rollback target.
Design D5 GC liveness.

* feat(ssh): collect the shared runtimes/ Node store and give it its own owner

runtimes/ gets its own owner in the install model, so no version-dir GC (new or old
clients, whose listings are prefix-scoped) can list or delete it. A store pass
removes node-<sha> only when no retained dir references it, it is neither a current
pin nor the newest other verified runtime, and a ps or /proc check ran and found no
process using it. Legacy relay-*/orcad-* dirs are read for references and reported
as diagnostics only (design D10 two-step). Wired behind orcad GC's nodeRuntimePins.

* fix(ssh): runtime store process check holds runtimes reached through a symlinked home

/proc exe resolves symlinks and argv keeps whatever spelling launched the runtime, so
filtering on the exact $root path missed in-use runtimes on hosts like /home -> /var/home.
Filter on the store segment instead; the parser already attributes holds root-agnostically.

* test(ssh): wait for the holder process to spawn instead of a fixed delay

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 01:10:47 -07:00
OrcaWinandm4air 8c2cd7d331 feat(ai-vault): read remote OpenCode history with the pinned Node; remove Bun (#24128)
* feat(ai-vault): read OpenCode history with the pinned Node instead of Bun

SSH hosts whose Node lacks node:sqlite (or its backup(), which 22.13-22.15
omit) now get the pinned Node in the shared ~/.orca-remote/runtimes/node-<sha>
store orcad uses: POSIX hosts receive the official archive and extract and
hash-verify it on the host; Windows hosts receive the verified node.exe the
client extracted, promoted by host Node with the same hash check. WSL distros
use the same layout and checks under ~/.cache/orca/runtimes/.

The Bun release pin table and its materializer are deleted. Old relays keep
reading their vault-sqlite/<sha>/bun references; nothing deletes those files.
An unconfirmed runtime upload now keeps its stage instead of removing it.

* refactor(sqlite): drop the Bun SQLite adapter; node:sqlite is the only backend

Nothing outside Electron runs on Bun any more (design D4), so SyncDatabase
loses its Bun branch, and bun-sqlite-database, bun-sqlite-statement and
bun-readonly-wal go, with the relay's bun:sqlite external. The profile-state
backup worker admits Electron or an entry that exists, and startup errors
name the pinned Node. The D7 cross-runtime gate still runs Bun 1.4.2, now
reaching Bun's SQLite through its node:sqlite.

* test(native-chat): drop the Bun SQLite driver case now that node:sqlite is the only backend

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 01:10:31 -07:00
OrcaWinandm4air 6593d7d194 feat(orcad): run orcad on the pinned Node instead of Bun (#24110)
* ci(daemon): gate PRs on daemon protocol crossing from the newest release

Lands daemon-protocol-facts.mjs from the Windows update diagnostic branch with a
stricter parser, and adds check-daemon-protocol-crossing.mjs (rule R1): the working
tree must attach the newest release tag's daemon. Rollback crossing is reported only.
Runs in the cross-version-wire job, which already has full tags; tag selection moves
to config/scripts/stable-release-tags.mjs so both use one rule.

* feat(persistence): run profile backups in the worker whenever its entry is bundled

* refactor(orcad): make profile and native preflight runtime-neutral

The profile preflight parser now takes the expected runtime identity from the
caller (shipped callers pass the pinned Bun identity), and the native
preflight is renamed to orcad-runtime-native-preflight with neutral wording.

* feat(runtime): pin the Node 24.21.0 server runtime with an offline CI check

Add src/shared/node-runtime-pin.ts (NODE_RUNTIME_PIN, SERVER_TARGETS,
NODE_RUNTIME_ASSETS for all 8 server targets plus the headers tarball),
generated by config/scripts/update-node-runtime-pin.mjs from the nodejs.org
and unofficial-builds SHASUMS. check-node-runtime-pin.mjs verifies, with no
network, that the pin tracks the locked Electron, matches engines.node's
major, and covers exactly SERVER_TARGETS; it runs in the static analysis job.

ORCAD_BUN_TARGETS consumers now read SERVER_TARGETS so there is one target
list; orcad's Bun runtime and build output are unchanged.

* test(persistence): skip plain-Node backup selection tests in the Bun profile suite

* fix(runtime): reject a pinned archive that belongs to another target

* ci(daemon): fail PRs that swap a runtime launcher and bump the daemon protocol

D7.1 R3: hosting orcad or the daemon on another runtime is not a protocol change,
so one PR must not do both. The launcher file list lives in the check script; the
allow-runtime-launcher-protocol-bump label overrides it.

* feat(orcad): select pinned-Node slots by a .runtime-node marker

D7.1 R5: a Node slot names its shared runtimes/node-<sha256>/node through
.runtime-node instead of .build-target, so Bun-era clients read it as a legacy
slot rather than exiting 78 on a missing bun-runtime. Nothing builds the marker yet.

* fix(runtime): load the Node pin without the typeless-module warning

check-node-runtime-pin.mjs now requires the pin and takes nodeDistArchiveName from
its own module, so it no longer loads the update script's build graph.

* fix(orcad): resolve Node slots to the design's runtimes/node-<sha>/bin/node layout

* feat(orcad): 8-slot node-pty prebuilds against the pinned Node headers at N-API 8

- build-orcad-prebuilds.mjs adds win32-x64/arm64 (conpty.node, the vendored
  conpty.dll/OpenConsole.exe, upstream's N-API conpty_console_list.node), compiles
  in a scratch copy against the hash-verified pinned headers (node.lib pinned per
  Windows arch) with NAPI_VERSION=8, rejects post-8 node_api_* imports, and writes
  a schema 2 manifest with per-file sha256, N-API level and the glibc need.
- --require-slots [slots] verifies files against hashes; --smoke loads the slot
  under the pinned Node and spawns a PTY; --print-slot names the host slot.
- The slot installer gates on N-API, libc, arch, glibc and file hashes instead of
  the exact NODE_MODULE_VERSION, and installs nested files (conpty/).
- bun-profile-tests.yml builds, verifies and smokes each runner's slot.

* fix(orcad): scope node-pty's glibc .symver pins to glibc on musl prebuild slots

musl's unversioned libc cannot satisfy openpty@GLIBC_* references at link
time, so the Alpine slot compile would fail. Pin the staged pty.cc guard to
__GLIBC__ and assert both musl transforms against the installed patch.

* feat(orcad): run orcad on the pinned Node instead of Bun

A packaged orcad slot now references the pinned Node 24.21.0 by its
executableSha256 (`.runtime-node`, `.server-target`) instead of carrying
bun-runtime, and ships node-pty from the slot's prebuild, only its own
ripgrep, and no Windows Bun PTY gate. The runtime lives beside the slots
at runtimes/node-<sha>/bin/node (node.exe on Windows, upstream name).

- build:orcad (build-orcad-node.mjs) builds the host slot's prebuild when
  missing and places the pinned runtime; the template is schema 3 with
  per-target files.
- handoffToBundledOrcad() resolves the slot's runtime reference and checks
  process.versions.node against the pin; a host Node >= 18 still hands off.
  Startup preflight keys on running as that runtime; callers expect 'node'.
- orcad and its daemon use node-pty (ConPTY + windows-pty-job on Windows);
  the Bun PTY sources, gate entry and canUseBunPty branches are removed.
- SSH deploy uploads the official archive once per pin, extracts and
  hash-checks it on the host, and self-tests it before publishing. Bun
  slots stay launchable for rollback; Node slots never use host Node.
- The runtime materializer is generic over pinned assets; the Bun wrapper
  remains only for the OpenCode vault reader (design Phase 2).
- Cross-runtime test: a profile DB written by Bun 1.4.2 (WAL left by
  SIGKILL) opens and backs up under the pinned Node, and the reverse.

No daemon PROTOCOL_VERSION change (design D7.1 R3).

* docs(ci): name the headless lanes after the pinned Node, drop Bun shard timings

Design D10: ci-demand-rollout.md and ci-runner-efficiency.md follow the
bun-profile-tests.yml -> node-server-tests.yml rename; shard timings drop the
deleted Bun PTY tests and follow the renamed ones.

* chore(ci): count the runtime archive download as a runtime launcher path

* fix(orcad): pin the macOS C++ standard for node-pty prebuilds

The official Node headers' config.gypi sets clang: 0, so common.gypi skips its
gnu++20 xcode_settings and Apple clang 15 (macos-14 runners) compiles
node-addon-api as C++98.

* fix(orcad): resolve the preflight's slot through realpath, as the handoff does

A symlinked orcad.js handed off to its real slot's pinned Node, but the
startup and profile preflights read the symlink's directory, found no
runtime marker there, and silently skipped the readiness check.

* refactor(ssh): drop materializeCachedNodeRuntime, which nothing calls

Deploys upload the verified official archive (design D5); no client path
needs an extracted Node executable cached by digest.

* test(orcad): gate the Bun-to-Node upgrade and Node-to-Bun rollback with live terminals

Design D7.1 R1/R3/R4 and D7.2. The last Bun orcad and this checkout's Node
slot are installed side by side under ~/.orca-remote, launched and stopped
with the client's own deploy commands, and share one data root. Each
direction proves the incoming orcad adopts the outgoing runtime's daemon
(same PID, same shell, output continues), opens its profile database and
backs it up with its own shipped worker, and that GC keeps the slot the
live daemon was forked from.

The node-server Linux lanes provide Bun 1.4.2 and build that Bun orcad
from main, and run with --cross-runtime. --artifact and --cross-runtime
now make their tests fail on a missing input instead of skipping.

* ci(node-server): pin node:24.21.0-alpine by its multi-arch index digest

* test(ssh): name the runtime archive fixture after its role

* test(node-server): load node-pty from the packaged slot in artifact runs

The node-server lane installs dependencies without building node-pty, and
Linux has no upstream prebuild, so the real-PTY failed-I/O teardown test
(picked up by the pty-subprocess selector) could not load pty.node. In
--artifact runs, alias node-pty to out/orcad's shipped slot so the test
exercises the addon orcad actually runs under the pinned Node.

* fix(orcad): let the Windows profile preflight exit after its PTY probe

On Windows, node-pty keeps the conout worker thread and pseudoconsole alive
until kill(), even after the shell exits. The PTY health probe never killed a
cleanly exited probe, so the packaged preflight printed its readiness line
and then hung until the build's 30s timeout, reported with an empty stderr.

- The probe kills its PTY on Windows after exit and uses the bundled ConPTY
  the daemon spawns with.
- The preflight exits once stdout is flushed; its owner reads to EOF.
- Preflight failures now report code, signal, timeout, stdout and stderr.

* test(node-server): load the slot's node-pty in the real-PTY test, not by alias

A vite alias redirected only ESM imports of node-pty; windows-pty-job and
local-pty-utils resolve it through require, so Windows loaded two conpty.node
copies and the Git Bash job-membership proof read an empty job. The failed-I/O
teardown test now loads node-pty through a fixture that picks the packaged slot
in artifact lanes.

The pty-subprocess selector was a prefix that also pulled in its POSIX-host
sibling unit tests, which pr.yml runs and which were never qualified on
Windows. Select the directory plus the two sibling files that belong here.

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 00:39:00 -07:00
OrcaWinandm4air d2dfc79764 ci(daemon): runtime-launcher protocol ratchet and Node slot marker (#24108)
* ci(daemon): gate PRs on daemon protocol crossing from the newest release

Lands daemon-protocol-facts.mjs from the Windows update diagnostic branch with a
stricter parser, and adds check-daemon-protocol-crossing.mjs (rule R1): the working
tree must attach the newest release tag's daemon. Rollback crossing is reported only.
Runs in the cross-version-wire job, which already has full tags; tag selection moves
to config/scripts/stable-release-tags.mjs so both use one rule.

* feat(persistence): run profile backups in the worker whenever its entry is bundled

* refactor(orcad): make profile and native preflight runtime-neutral

The profile preflight parser now takes the expected runtime identity from the
caller (shipped callers pass the pinned Bun identity), and the native
preflight is renamed to orcad-runtime-native-preflight with neutral wording.

* feat(runtime): pin the Node 24.21.0 server runtime with an offline CI check

Add src/shared/node-runtime-pin.ts (NODE_RUNTIME_PIN, SERVER_TARGETS,
NODE_RUNTIME_ASSETS for all 8 server targets plus the headers tarball),
generated by config/scripts/update-node-runtime-pin.mjs from the nodejs.org
and unofficial-builds SHASUMS. check-node-runtime-pin.mjs verifies, with no
network, that the pin tracks the locked Electron, matches engines.node's
major, and covers exactly SERVER_TARGETS; it runs in the static analysis job.

ORCAD_BUN_TARGETS consumers now read SERVER_TARGETS so there is one target
list; orcad's Bun runtime and build output are unchanged.

* test(persistence): skip plain-Node backup selection tests in the Bun profile suite

* fix(runtime): reject a pinned archive that belongs to another target

* ci(daemon): fail PRs that swap a runtime launcher and bump the daemon protocol

D7.1 R3: hosting orcad or the daemon on another runtime is not a protocol change,
so one PR must not do both. The launcher file list lives in the check script; the
allow-runtime-launcher-protocol-bump label overrides it.

* feat(orcad): select pinned-Node slots by a .runtime-node marker

D7.1 R5: a Node slot names its shared runtimes/node-<sha256>/node through
.runtime-node instead of .build-target, so Bun-era clients read it as a legacy
slot rather than exiting 78 on a missing bun-runtime. Nothing builds the marker yet.

* fix(runtime): load the Node pin without the typeless-module warning

check-node-runtime-pin.mjs now requires the pin and takes nodeDistArchiveName from
its own module, so it no longer loads the update script's build graph.

* fix(orcad): resolve Node slots to the design's runtimes/node-<sha>/bin/node layout

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 00:10:57 -07:00
Neil b245b3e019 feat(sidebar): let the host pill be turned off per card (#24299) 2026-09-30 23:37:15 -07:00
OrcaWinandm4air 49a83deaef refactor(orcad): make profile backup and preflight runtime-neutral (#24088)
* feat(persistence): run profile backups in the worker whenever its entry is bundled

* refactor(orcad): make profile and native preflight runtime-neutral

The profile preflight parser now takes the expected runtime identity from the
caller (shipped callers pass the pinned Bun identity), and the native
preflight is renamed to orcad-runtime-native-preflight with neutral wording.

* test(persistence): skip plain-Node backup selection tests in the Bun profile suite

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-30 23:23:23 -07:00
Brennan Benson ebc479b9c0 fix(native-chat): /clear starts nothing; the new chat's first message starts its agent (#23935)
* fix(native-chat): every journal append reaches the chats that are open

A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.

A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.

* test(native-chat): an epoch replacement reaches the open chat

* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

* perf(native-chat): a publish behind a delivered commit reads nothing

Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.

* test(native-chat): state why the teardown test's fake journal is safe to cast

* docs(native-chat): say mutation admission checks only the writer lease

* docs(native-chat): drop the send rebase from comments that still described it

* fix(native-chat): a message is accepted, then delivered

A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".

A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.

Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.

A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.

Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.

* fix(native-chat): settle queued messages only for the child that ended

A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.

A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.

The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.

* fix(native-chat): an adoption that fails to import keeps the conversation open

The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.

* perf(native-chat): the recovering open reads the journal once

Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.

* fix(native-chat): an attach that fails after indexing its child leaves no child behind

A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.

* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer

The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.

A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.

* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down

The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.

* fix(native-chat): a message rejected while its chat was closed reads as not sent

A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.

* test(orchestration): name why the readiness settlement fakes are cast

* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent

* docs(native-chat): drop the fence from the admission the send effects run behind

* docs(native-chat): give the fence move on release the reason that still holds

* docs(native-chat): stop citing a write fence check in launch and mailbox comments

Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.

* refactor(native-chat): the provider child is its own record

A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.

- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
  patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
  own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
  the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
  dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
  is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
  the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.

* fix(native-chat): the delivery loop alone settles a message its start or child failed

A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.

- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
  reads how it ended: a Stop continues; anything else writes one failure row and rejects every
  queued message with the same words, then stops. A child still starting whose start the adapter
  says did not land fails the same way. The exit, eviction and the settlement retry only settle
  the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
  nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
  failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
  closed, with or without a child, and a start the loop already has in flight is waited for so the
  child it produces is stopped rather than left behind.

* refactor(native-chat): a stopped child ends on the one reading of its stop

The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.

* feat(native-chat): the host says it accepts a send before any agent has it

The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.

* refactor(native-chat): an attach never opens a journal of its own

The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.

* fix(native-chat): a moved fence resends nothing on a host that accepts first

The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.

The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.

* refactor(native-chat): a child's end says whether the user or the host stopped it

The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.

* fix(native-chat): a chat whose only work is a queued message is not offered for resume

A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.

* fix(native-chat): the conversation outlives its agent

Opening a chat no longer starts its agent. A conversation is reached through one host
accessor that opens its journal at rest, and a send is what starts the agent, through
the delivery loop. One idle sweep, every five minutes, stops an agent that has been
quiet for thirty minutes and owes no work, then drops an open journal handle that is
only a cache. Its record, tab, status row and readers stay.

- hold and release are no-ops; hold still builds the host for shipped mobile builds.
- The holders, the holds, the release clock and the exit respawn are deleted.
- Options, the model list, the goal and the context meter answer at rest; a model pick
  at rest is recorded as intent for the next start.
- Compact, rewind, clear and goal changes start the agent first. A send does too when
  a rewind is still in doubt after the conversation opens.
- Orchestration routes mail and group addresses on ownership (the record plus the chat
  tab), not on whether the process runs. An open dispatch keeps its worker running.
- The restart continuation is a send; Resume all holds each slot until the message is
  handed over or rejected.
- A read error never replaces a loaded transcript, and shows the host's own words.

* test(native-chat): type the queued-message fixtures in the resume-offer tests

* fix(native-chat): a start that dies while a message waits on it is that message's failed start

Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.

* fix(native-chat): a request that failed reads as failed

A structured chat whose only message the agent's start refused read as a
green finish, and a cancelled structured turn did too: the host published a
verdict only for turn records, and structured rows carried no `interrupted`.

The host projection now reads the session's latest request: its turn's
outcome, or `failure` for a send the agent or its start refused. A send
that was withdrawn, or left undelivered by a restart or a close, fails
nobody and makes nothing listable. The ingest publishes `interrupted` as the
hook lanes do, and every reader decodes the verdict through one accessor, so
a failure reads Failed on the dot, the rollups, history and `worktree ps`,
behaves like a cancellation in every clean-finish policy, and notifies as
"failed".

* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now

* test(native-chat): a verdict change republishes the mobile status projection

* refactor(native-chat): the store's retention trigger keeps its flag compare

A verdict change always moves the completion clock the same check already
reads, so a second verdict compare there caught nothing new.

* test(native-chat): a user message the provider journaled keeps its session listed

* test(native-chat): pin what a failed start settles, and what a resume offer names

A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.

* test(native-chat): the failed-start pins fail on what the message became, not on a timeout

* fix(native-chat): a restart offer ends when the chat's agent starts again

The offer used to end only when the chat's newest user message changed,
because opening a chat started its agent and that start could not be told
apart from real activity. Opening a chat starts nothing now, so the host
reads the fact it already publishes: a chat's status row goes from not
host-owned to host-owned exactly when its agent is started. At that edge the
offer and any failure record for the chat are withdrawn, unless the start is
a resume action's own (its continuation is the oldest undelivered message).

A continuation and a message racing to be first are decided at acceptance:
the continuation is refused, quietly and with nothing filed, when any other
message was accepted since the restart. A failed continuation start leaves
the offer retryable, and each resume action sends its own message id.

Deleted: the newest-user-message comparison, its journal reader, the
continuation filter, and the failure ledger's own "answered by the chat"
check. The marker still carries its message id for one release, so the
previous build can read it.

* fix(runtime): end a transcript stream when its client unsubscribes

Desktop: the IPC subscription controller was dropped as soon as the streaming
handler returned, which for most streams is right after it binds. A later
runtime:unsubscribe then found nothing to abort, so the host kept the subscriber
and derived and sent every publish to a channel no one listened to. The controller
now lives until the renderer unsubscribes, resubscribes the same id, or goes away.

Mobile: disposing an agentSession.subscribe stream now sends agentSession.unsubscribe
with the stream's frame id, so the host ends that subscriber and leaves a sibling
stream on the same socket running. The direct path now passes the frame id the relay
path already passed.

* fix(native-chat): a late provider-session update keeps a failed recovery record failed

A provider-session heartbeat that rewrites a completed recovery record kept
its interrupted flag but dropped the outcome it was copied with, so a live
failed checkpoint read as a clean finish until the next status write.

* test(orchestration): the preamble's host stub is typed, not cast

The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.

* test(native-chat): the terminal-bell check asserts the renamed verdict field

The bell notification test still checked for agentInterrupted, which no
longer exists, so it could not catch a verdict leaking into a bell dispatch.

* fix(native-chat): a failed turn ranks like a completion for attention

Attention readers (completion time, Smart Sort, sticky retention, Cmd+J
Recent) now demote only a turn the user stopped. A failure is news the
user has not seen, so it keeps its completion time, ranks in the Done
class, stays retained after its pane goes away, and a retained failure
reads failed in the worktree rollup instead of done. Clean-finish
policy (hibernation, pane ownership, the value moment) still treats a
failure like a stop.

The retention trigger compares verdicts again: success -> failure no
longer moves the completion clock.

* fix(native-chat): one fact ends a restart offer: the chat moved on since the restart

The offer is live while no other message has been accepted in the chat since the
restart and its agent has not proved a start since. The offer list, the resume's
reservation check and the continuation's acceptance check all read that one fact,
so a message whose start then failed withdraws the offer too, and a stale click
finds nothing to act on.

The fact is read off the conversation's open handle, which the restart closed, so
it is retired durably whenever it may have changed: a message accepted, a start
proven. A close and reopen within the same run therefore cannot bring the offer
back. A continuation rejected before it reached the agent does not count, so a
retry after a failed start still runs.

Deleted: the quit-time gate on withdrawal, which changed nothing because the
withdrawal and the quit's own offer write share one queue; the per-action
"withdrawn" flag and the separate acceptance check it paired with.

* test(native-chat): an older build reads the restart offer this build records

The offer lives in a file the previous release reads after a downgrade. Pin that
against the pinned release's own capsule, and run the lane when the marker or the
capsule changes.

* fix(native-chat): read a restart offer against where the journal stood when it was taken

"Since the restart" was read off the conversation's open handle, which the idle
sweep closes: after a reopen, a message the user had already sent looked older
than the handle and the withdrawn offer came back.

The offer now records the journal position (epoch and sequence) at the moment
it is taken, and a message accepted after that position, or a journal on another
epoch, means the chat moved on. That is derived from the journal, so it holds
across any number of closes and reopens. An older build's offer has no position;
only a start withdraws it. Because the message half is now durable, the offer is
no longer rewritten in the recovery file on every accepted message; a proven
start still writes it, since only the host that saw the start knows of it.

* test(native-chat): wait for the listing's retire write before reading the recovery file

* refactor(native-chat): every journal row states which turn it belongs to

Rows gain a turn scope stated by the write that creates them: the open root
turn, or the conversation. A queued message takes its scope from its handover.
Rows stored before scopes existed are placed on replay by the root turn open
when they were created, so no persisted state is needed for them. Rewind keeps
each retained row's scope and producer, so a subagent's row stays its own.

* fix(native-chat): keep the terminal-backed chat's read error over its local echoes

Messages winning over a read error is right for the structured chat, whose read retries and whose
messages came from the transcript. The terminal-backed view assembles its list from local echoes
too (a launch prompt, a pending send), so a failed read there showed only those bubbles and no
error. Only the structured pane now keeps messages over an error.

* fix(native-chat): a start retries the exit settlement a failed journal write left owed

An agent exit whose journal settlement write failed releases the lease latched until a retry lands.
Reopening the chat used to be that retry; with reveal now only opening the journal, nothing retried
it before the next app launch, and every send was refused. The start the send needs now runs the
retry first, where the attach would.

* fix(native-chat): a failed main agent reads failed while its subagents still work

The verdict is now read from the main agent's own state, not the folded
row: a main agent that is done and failed has a verdict even while its
subagents keep the row working. Without mainAgent (history, worktree ps,
older hosts) the old combined-done rule stands.

Display marks the verdict through agentVerdictDisplayMark: a failure
outranks every combined state on the agent's dot, label, tab badge,
dashboard and activity rows; a stop marks only a done row, so a
successful or stopped main agent with live subagents still reads
working. Subagent rows keep their own state. The worktree card, terminal
tab and Cmd+J rollups share one pane fold and rank a pending question,
then failed, then working, monitoring, interrupted and done.

worktree ps publishes the main agent's outcome on a working row, and the
mobile mirror reads it. The store's change check, the paired-client
mirror's equality and its epoch now see a verdict change on a working
row, which otherwise moves no state or clock and left the worktree card
reading working. Clean-finish policy is unchanged: a working row is never
hibernated and has no completion time.

* perf(native-chat): answer the owner check without opening the chat

Worktree activation calls agentSession.handoffStatus for every chat tab in the worktree, and the
answer comes from the session record alone. Reaching it through the accessor opened each resting
chat's journal (a full read, the crash-boundary write and a restored status publish), then kept it
open for the idle window. It now checks the record and the adapter's support, as before this series,
and opens nothing.

* fix(native-chat): a read waiting on the session lock opens nothing once quit began

The accessor checked for quit before queueing the open, so a read queued behind a session task ran
its open after teardown had begun and indexed a journal no teardown step would close. The check now
runs at the open itself.

* test(native-chat): pin stated turn scopes, the upcast of unscoped rows, and rewind attribution

* fix(native-chat): /compact is a message the chat sends, run as a turn of its own

The conversation command RPC now accepts /compact into the queue like any
send and answers once it is handed over. The delivery loop opens the command's
own turn, starts the provider on it, and waits for the provider's end off the
session's queue, so messages typed meanwhile are held and delivered after it,
even when it fails. It settles by re-reading the journal: a child that died
meanwhile already wrote the verdict. Stop ends the command at once. The 180 s
completion window, the unconfirmed row and the recovery of an older build's
compaction record are gone; that record no longer gates anything. On Codex the
provider turn the command opens is claimed into the command's turn.

* fix(native-chat): read a failed resume's chat before calling it retryable

Whether a failed resume is retryable is the offer's own rule: the chat has not moved on since the
restart, read from its journal. The failure list read it only for a chat already open, so once the
idle sweep closed a chat the user had moved on in, its failure showed Retry again, and the click did
nothing. The list now opens the failed chats first, as the offer list does.

* test(native-chat): type the provider event sink the settlement test reaches for

* docs(native-chat): the worktree ps outcome comment no longer claims old hosts send it

The field is new: an old host sends no outcome at all, so a reader falls
back to interrupted. The removed clause said old hosts send it on done
rows, which never shipped.

* fix(native-chat): say the structured read keeps trying only where it does

The structured pane's "Orca keeps trying to load it" line never showed: the view state filled in an
untranslated fallback whenever the read error had no text, and the empty state prefers any message.
The view state now leaves the message out, so the structured pane shows that line and the
terminal-backed pane its own translated one. Mobile's structured lane does not resubscribe after an
error frame, so it no longer makes the claim.

* fix(native-chat): rows group under the turn their record names, not the one above them

Each row's turn is the turn its stated scope names, anchored on the entry
that opened it, or on the turn itself when the provider opened it unasked.
So /compact groups its own rows and the previous turn is untouched, a message
typed into a running turn joins it, and a provider-resumed turn folds under
its own Worked-for. A row reporting how a turn ended, an error or the
compaction separator, never folds. Desktop and mobile read the same keys; a
host that states no scope keeps today's positional grouping.

* test(native-chat): await the send's settlement instead of polling for the start

The at-rest send tests polled for the provider start with vi.waitFor's one-second default, which a
loaded machine outran. They now await the host's own settlement of the message.

* docs(native-chat): the status-store listing rule names provider-journaled user messages

* fix(native-chat): a restart offer resumes any time after the quit, and knows its own continuations

The continuation's message id was dated by the quit, and the ledger refuses a new id dated more than
a day back, so Resume or Retry a day after quitting was always refused (on main too). It is now
dated by the resume action.

Telling a rejected continuation from the user's own message read the operation ledger, whose rows
expire after about a day; after that a failed resume stopped being retryable. The offer now
records the continuation each action sends on its own capsule entry, bounded to the newest 16, so
the ids end with the offer. The ledger read is deleted.

* fix(native-chat): a /compact is not a request the sidebar, notifications or restart resume report

The sidebar's prompt, preview, verdict and instant, the turn-completion feed,
and the restart-resume marker read past a conversation command and its turn to
the last real request, so a /compact neither notifies nor re-dates the row,
and a command in flight is never offered as work to resume. An older client
shown a command's turn in the legacy form names the session's own agent.

* fix(orchestration): route no mail to a structured worker its orchestration released

A structured worker is routed on ownership, and a resting worker's lease is released, so ownership
held while its chat tab stayed listed. A worker the coordinator abandoned and then released, found
at rest by the release, therefore still took peer mail and @worktree: broadcasts, and each one
restarted its agent. Routing now also reads the orchestration's own resource row: once it is
released, direct mail, group addressing and worker-show's addressable answer drop the worker, as
they would a terminal worker whose terminal closed. The chat tab stays, and nothing new is stored.

* fix(native-chat): a failed retry names the user's prompt, not Orca's continuation

A resume's continuation is written to the chat before its start, so after a failed attempt the chat's
newest user message is that rejected continuation. A second failure then showed Orca's own restart
text as the chat's prompt. A retry now keeps the prompt its first failure named.

* test(native-chat): pin what a conversation command's admission refuses at rest and at handover

* test(native-chat): tests merged from the base state which turn their rows belong to

* fix(native-chat): a refused send notifies failed through the completion feed

The host's completion feed followed only the newest turn, so a send the
agent or its start refused, which creates no turn, read Failed on its row
but sent no notification. The feed now follows the session's latest
request, read from the projection the status feed already makes for the
commit: a turn keeps its id, a refused send is named by its journal item
key. It announces only while the session is idle, as the row reports a
verdict, so queued sends refused one commit at a time notify once, and a
withdrawn send falls back to a request already announced.

* fix(orchestration): read the released row optionally, as the authority does

worker-show's observation called the row lookup directly, which a runtime double without it threw on
and failed the structured tab-retirement release.

* chore(native-chat): one import per module and no unexplained casts in the turn-scope changes

* test(claude): pin which turn a Claude row joins, including a subagent's after the turn ends

* fix(native-chat): the status bar drops a restart offer the chat moved on from

The renderer re-read the host's restart offer only when a failed chat showed activity, so after a
message withdrew a pending offer the host answered no chats while the status bar kept counting one,
and clicking it opened nothing. The same watch now covers pending offers: a status change in an
offered chat asks the host again, once.

* fix(native-chat): a refused steer is read from the turn its handover named

The latest-request reader decided whether a refused send had joined a running turn by comparing
host clocks: its handover time against the previous turn's end. The handover row now states the
turn it delivered into, so the reader reads that instead and the clock comparison goes. A journal
written before handover rows stated a turn is scoped on replay from the turn open when each row
was written, which can differ from the clock reading only when a send and a turn's end share a
millisecond.

* fix(mobile): the native-chat controller contract carries the turn journal

The controller and overlay already pass nativeChatTurnJournal, but the
contract type never declared it, so mobile failed to typecheck.

* fix(native-chat): the live turn is the running turn, not the newest user row

A turn the provider opened on its own (a background wake, a resumed turn)
anchors on its own record, but the list still treated the newest user row
as the live turn. While such a turn ran, the settled user turn before it
lost its duration and the running turn's own rows were drawn as settled,
so its tool calls lost their live state.

nativeChatTurnMembership now answers both questions from the turn record:
each row's turn, and the live turn (the running root turn's anchor, else
the newest user row, which is also all an unscoped host has). Desktop and
mobile key liveness, the timing clock and the live status's row on it.

* test(native-chat): a turn the provider opened keeps its own clock

Pins that the local turn clock follows the live turn, so a wake after a
settled turn does not restart that turn's clock when no host durations
are recorded.

* fix(native-chat): a running turn no message opened draws its status on no row

Its live status belongs to the transcript-tail indicator alone. Once it
settles, its duration draws at its first row as before; a running turn a
message opened still draws on that message.

* fix(native-chat): every copy of a row carries the main agent's own status

History entries, sleep records and `worktree ps` rows carried a flattened
top-level `outcome`, copied under different gates and without the main agent's
clock. They now carry `mainAgent` (state, outcome, stateStartedAt), the type
the live row already persists and sends, and every copy site takes it with
`interrupted` through one function, `agentVerdictFields`.

- The accessor reads `mainAgent` then the legacy flag; the mobile mirror
  matches it line for line.
- Sleep records admit `mainAgent` with `normalizeMainAgentStatusField`, so a
  malformed value drops the field, never the record.
- Mobile dates a main agent that failed under live subagents by its own clock,
  as desktop does, and its row equality compares `mainAgent`.
- The activity feed reads a history entry's own `mainAgent` instead of
  rebuilding one; the sync key and history equality compare it.

* test(native-chat): pin the worktree ps verdict across host and phone versions

Pairs the real v1.4.212 host and phone row reader with this build: an old phone
reads a new host's rows by `interrupted`, a new phone reads an old host's rows
(no `mainAgent`) the same way, and a new phone reads a failure under live
subagents as Failed, dated by `mainAgent.stateStartedAt`. The release checkout
now carries the phone's self-contained row reader, and the lane runs when the
`worktree ps` row producers change.

* test(mobile): name the parity table's row for its role

* test(native-chat): a roster of idle or finished children does not keep an agent awake

The sweep reads owed background work through the shared child-work liveness that upstream's
release clock adopted; a child that went idle or finished is not work the agent still owes.

* fix(native-chat): a request that settles while the user is asked something notifies once

The completion edge waited for an idle session, and a pending prompt (including a
subagent's approval) is not idle. Structured chat has no other attention producer,
so a main turn that finished while a subagent waited on the user sent nothing
until the prompt was answered.

The edge now waits only on owed work (a running turn or an unanswered send), which
the projection reports even beneath a pending prompt. A request that settles with
a prompt pending announces once; the renderer words it "needs input" from the
host status mirror's `attention`, and answering the prompt keeps the same request
identity, so it does not announce again. The wire shape is unchanged.

* fix(orchestration): a task dispatched into a resting structured worker keeps it running

The sweep's open-dispatch check read only the worker-start dispatch that owns the worker's terminal
resource, so a task later dispatched to the same worker (orchestration dispatch --to, which writes a
dispatch with no worker row) did not count: after thirty quiet minutes the worker was stopped while
that task was open, and its coordinator read exited. Any unsettled dispatch addressed to the worker's
process incarnation now counts, derived from the existing rows.

* fix(native-chat): a command's wait ends when its child does

The delivery loop waited for a /compact only on the adapter's compaction
tracker, which learns of the child's end only on some exit paths: a Codex
exit or close, and a Claude close, never reach it. The wait then never
ended, so nothing queued behind the command was delivered again, Stop had
no child to answer through, and the tracker's leftover entry refused the
next /compact.

Every way a child ends passes endProviderChild, so the host now offers a
per-child end signal there. The loop races the tracker against it (the
dead-generation settlement has already written the command's verdict),
and on that end asks every adapter to release the command, so a later
command runs and no later provider turn is claimed into the dead one.
The adapters' own exit-time releases were unreachable (Codex) or covered
one path of several (Claude), and are removed.

The Codex RPC test harness moves to its own module so the exit can be
driven through the real adapter's connection callback.

* fix(native-chat): keep refusing sends during a command on an older host

An older host's controller still refuses a send while a conversation
command runs, so dropping the client's block turned every message typed
during /compact into a 'not sent' row with Retry there. The block stays
for hosts that do not run the command as a send-path turn, and goes only
for those that do.

The signal is one the client already holds: a host that runs /compact on
the send path states a turn scope on every journal row it writes, the
same fact turn membership uses to tell it from an older host. Both now
read it from one predicate. On an empty conversation, or one whose rows
all predate the upgrade, the signal is absent until the command's own
entry streams in, so that brief window keeps the old local refusal; no
capability or wire field is added.

* docs(native-chat): comments stop describing the hold this PR removed

Eight comments still justified orderings and teardown choices by a viewer or dispatch hold that
pinned the provider child. Nothing holds any more; the orderings stand for the binding's redrive
subscription and parked mail, and a chat's agent runs from a send until the idle sweep rests it.
Comment-only.

* fix(native-chat): the completion says when the user is being asked

A request that settles while a prompt waits on the user was worded "needs input"
from the renderer's status-feed mirror. Remote clients receive the status and
completion streams over separate sockets, so they can arrive in either order and
the wording could be wrong both ways.

The host already knows at emit time, so the completion now carries an optional
`awaitingUser: true` in that case and omits it otherwise. The renderer words the
notification from that field alone and no longer reads the status mirror. Old
clients ignore the field and word by outcome; old hosts never send it.

* fix(native-chat): a restart offer keeps the start its own continuation made

Whose start ended an offer was decided at read time, from whether the offer's continuation was
still the queued message. Once the provider refused that continuation, the child it had started
read as someone else's start, so the offer ended and its failure showed no Retry. The delivery
loop now records which queued message a start is for on the in-memory child, and the child's end
carries it; the offer counts a start as its own when that message is one of its continuations.

* fix(native-chat): a rewound turn still names the message that opened it

A Codex rewind rebuilds the epoch without submissions, so each sent message survives only under
its provider key. The kept turn records still named the submission key, so each turn anchored on
itself and its rows grouped apart from the message that opened it. The rewind now renames the
turn's opener along with the message.

* fix(native-chat): Stop ends only the command it names

Stop on a command turn abandoned whatever compaction the session had pending, so a late Stop for
an earlier /compact cancelled the one running now. The tracker now ends a command only when the
Stop names its turn, and the cancel reply reports whether it did.

* fix(native-chat): an agent gets a full idle window after its owed work ends

The sweep measured quiet only from the last journal row, so once a subagent, command, monitor or
dispatch that had outlived the window ended, the agent was stopped at the next tick. A child can
read done before the lead's wake-up turn writes anything, and stopping in that gap loses the
wake-up. The sweep now counts owed work it observes as activity, which gives the agent the full
window afterwards, as the release clock it replaced did.

* test(claude): the options-read fixture runs a live child

The fixture marked its conversation running with a hasProviderChild field the
session type does not have, so the read took the at-rest path and refused a
session with no record. It now carries a child, which is what the read checks.

* test(native-chat): host tests reach its collaborators through a typed seam

The rest-test rig and three test files read the host's private members with
Reflect.get and cast the result. The host now exposes one test-only accessor,
collaboratorsForTests(), and the subscribers class a subscriberCountForTests()
beside its existing retainedActivityCountForTests(), so the tests are checked
against the real types and the casts are gone.

* fix(worktree-status): a departed agent's failure yields to live work on the worktree card

A retained failed agent has no expiry, so ranking it with a live failure pinned the card to Failed over other panes' live work. It now ranks below working, monitoring and permission, and above every finished outcome.

* refactor(orchestration): one owner answers a structured worker's custody

Routing, group addressing, worker-show and the idle sweep each composed their own reading of
whether orchestration still holds a structured worker, so each new obligation or retirement state
had to be added to every reader. structured-worker-custody now derives both answers from the
worker-terminal list state coordinators see in worker-list: addressable is owned and not released,
and owed work is an active custody or an unsettled task dispatched to the same incarnation. The
owner's state is read through the remote dispatch attachment too, as the terminal transfer lookup
already does. Behaviour is unchanged; a settled worker awaiting its coordinator still rests.

* refactor(orchestration): owed work is an open dispatch on the worker's incarnation

A supervised worker's own dispatch context stays open exactly while the worker is active, so the
separate active-custody branch only repeated it. Owed work is now one fact, which also states the
policy that a worker awaiting its coordinator's decision may rest, and both custody decisions are
written once at the top of the module.

* docs(agent-status): a departed agent's failure ranks below live work on the worktree card

* fix(native-chat): a restart offer knows its continuations by a tag in their id

The offer recorded each continuation id in a list on its capsule entry, capped at 16, and a running
action's id in memory. Both could disagree with the journal: past the cap an old rejected
continuation read as the chat moving on, and a crash during a retry restored the failure's older
entry, which lacked the retry's id. Each continuation id now carries a tag derived from the offer
(its teardown and chat), then the action's own part, so any continuation of this offer, queued or
rejected, is recognised from the journal row and the marker alone. The persisted list, its cap and
the in-memory action map are deleted; the agent-start withdrawal keeps an offer whose own
continuation the start was for, read against the stored marker.

* test(runtime): the legacy-worker reveal test judges its stale snapshot inside the wait

The tui-idle probe reads through readTerminal, which now awaits the structured
worker check before the PTY read, so the probe's snapshot request starts a
microtask later. vi.waitFor missed it on its first check and polled again at
50 ms, the same moment the wait's own 50 ms timeout fired. The stale snapshot
then resolved after the wait had already timed out, so the test passed without
judging it, and the rejection landed before any handler was attached. Vitest
reported that as an unhandled error and failed the shard.

Polling every 1 ms sees the request within a few ms, so the snapshot is judged
while the wait is still pending.

* fix(native-chat): a message held behind /compact is drawn where it was handed over

A message typed while /compact runs was drawn above the compaction's result, between
itself and its own answer. The reducer kept every item at the sequence and timestamp of
the row that created it, and a queued message is created at acceptance, long before the
command it waits behind writes its result. The phone orders by that sequence and the
desktop by that timestamp, so both put the message first.

A queued message now takes its position from its handover row, the same row that already
states its turn scope. Everything the agent did before the handover, a command it waited
behind included, draws above it. This holds for every held message, not only /compact's,
and needs no client change: every client, older builds included, reads the position the
host publishes. A live batch already carries the item when its dispatch row lands, and
history pages cut the reduced timeline by sequence, so paging stays contiguous.

* fix(native-chat): a phone's send during /compact answers without waiting out the compaction

A client that predates accepted-send replies, which is every phone build, has its send
reply held until the host hands the message over. A message sent during /compact is not
handed over until the compaction ends, so the phone's 15 s request timeout fired first
and showed the message as unconfirmed.

That wait now also ends once the message is queued behind a running command. This is
read from the journal's running turn and needs no new state. Every other wait still
ends at the handover: behind a starting child or an ordinary turn, and for restart
resume, the command front door and orchestration, which keep the plain handover point.

* perf(native-chat): a rewind places provider items with one pass over the merged rows

A Codex rewind gives each provider item the old epoch never held the turn record for its
provider turn. It found that record by scanning every merged row, restoring each row's
body, once per provider item. That is quadratic, and it runs on the host's main thread
up to the journal's 10,000-row cap, twice per rewind. A rewind record written before
rows carried their scope holds no scope for any provider item, so it paid the full cost.

The merge now indexes turn records by provider turn id once, keeping the first match as
the scan did, and each provider item looks its record up.

* fix(native-chat): a view never restarts a chat whose last start failed

A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.

* fix(native-chat): a message waiting behind /compact is drawn after it until it is sent

A message sent while /compact runs is placed where it was handed over. It was still
drawn where it was accepted until then. /compact writes its result one step before the
handover, so for that step the waiting message sat above the compaction's separator.

A message the host accepted but has not handed over is not part of the conversation
yet, so both clients now draw it after everything the agent has done. The shared
projection moves it to the end, which is the order the phone draws. The desktop ranks
it with the other not-yet-sent rows, after the streaming preview. At handover it takes
its place from its handover row, which is also after the separator, so it never
appears above the compaction it waited for.

* fix(native-chat): the idle sweep reads owed work every tick

Owed work counted as activity, but the sweep read it only once the idle window had elapsed, so it
refreshed the clock at most once a window. Work that ended just before the next read left the
agent to be stopped at that read, moments after the work ended, which is the gap the refresh was
meant to cover. The sweep now reads owed work on every tick for a started agent, so the window
always runs from the last tick that saw work owed.

* fix(native-chat): a continuation handed to the agent stays sent

The offer read its own continuation as not reaching the agent while its dispatch was pending, which
also covered one already handed over and still unanswered. When the wait for that answer ended first,
the failure it filed read as retryable, and a retry sent a second continuation to an agent that may
have acted on the first. Only a continuation still queued, or rejected, is now read as unsent.

* test(native-chat): start the child the loop waits on with an attach, not a second view

A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.

* fix(native-chat): settle a gone generation's turn wherever a conversation opens

A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.

* test(native-chat): prove the next child's start settles the turn an earlier child left

The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.

* test(native-chat): count a failed start's rows by row, not by text

Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.

* test(cross-version): load the phone row readers without mobile's toolchain

Vite transforms a file against its nearest tsconfig, and mobile/tsconfig.json
extends expo/tsconfig.base.json, which the root-only cross-version lane never
installs. The worktree ps verdict suite imported the current phone row reader
from mobile/ directly, so CI failed with TSConfckParseError before any test ran.

The harness now imports a copy of the working-tree reader placed under the
checkout cache, where the root tsconfig applies, as it already does for the
release checkout's copy. Both readers are still the real files.

* test(cross-version): keep the checkout path-guard message and justify the copy import's cast

* fix(native-chat): a command ends only by its own provider answer or its child's end

Stop no longer settles a conversation command. It interrupts it like any turn,
and when the provider cannot take that (Codex has not opened the command's turn
yet, or Claude refuses the interrupt) it stops the child, whose dead-generation
settlement writes the verdict.

The pending command now lives on the provider child's own session instead of an
adapter-wide map keyed by session, so it dies with the child and nothing has to
release it. Claude's /compact is sent under a uuid the slot records, and only a
root result naming that input (or naming none) ends it; its outcome is read with
the ordinary result reading, so a stopped /compact is a cancellation.

* fix(native-chat): a command's settle answers its message before ending its turn

The two writes are not one batch. Writing the message's answer first means a
crash between them leaves a running command turn, which the stale-turn sweep
already settles, instead of an ended turn whose message reads as in flight
forever. The settle now writes only while the command turn is still running.

* fix(native-chat): "Worked for" counts from the handover, not the send

A message held behind /compact, or behind a cold start, used to count the wait
as the agent's work, although its row is drawn at the handover. Every handed-over
submission's turn, the command's own included, now starts at the handover row's
instant, falling back to the send time for a host that recorded none.

* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget

* test(native-chat): the interrupted create's own retry continues again

The merge of main's lease-latch fix replaced that test's retry of the interrupted create, under its
own operation id, with a fresh start whose result nothing read. That fresh start passes with the
released-reservation continuation deleted, so the case the fix exists for went untested. The retry
and its assertion are main's again.

* docs(native-chat): three comments that still had views starting agents

A start with nothing queued now comes from a command, goal change or rewind; an interrupted compaction
left alone would refuse every send, so no agent would ever start to finish it; and a current host
raises the unattached read refusal only once quit began, with the attach window belonging to an older
host.

* test(native-chat): pin the open's and the send's start and row counts, however the view binds

Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.

* fix(native-chat): a second Stop on a command ends its child; one compaction verdict for every provider

A Stop's note now names itself in its key, so a later Stop on a command still
running reads, from the journal, that the provider was already asked and never
answered, and stops the child instead of interrupting again. Nothing is held in
memory for it.

Adds the rule both translators will read a compaction's end by: only a
compaction the provider reported is a success; none after Orca's interrupt is a
cancellation; anything else is a failure. A real Claude capture, pinned as a
fixture, is why: a stopped /compact ends in the same success result as a
finished one.

* test(native-chat): a reader's open settles the turn a failed exit settlement left running

An exit whose settlement write failed leaves its turn running in the open journal. PR 1's open now
settles it, and this pins the two reads that reach it here: a reader reopening a chat the idle
sweep closed, and a read that opens the chat before the restart restore reaches it.

* test(native-chat): the view-start test's starting window outlasts two subscriptions on a loaded runner

A subscription reads the conversation before it returns, so under load the two views took longer
than the create child's 300 ms start, which then exited before the test checked that it had not.
The child now takes a second to fail.

* fix(native-chat): settle a gone generation's turn at every open but an acquisition's

The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.

* test(native-chat): hold the create's start open until the views bind

The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.

* refactor(native-chat): the provider's translator ends a command's turn; the loop holds no command state

A conversation command is now a turn of the provider child's own journal
pipeline. The adapter-wide tracker, its promise and the loop's settle step are
gone.

- Codex: the translator claims the provider turn that carries the command, scopes
  its rows to the command's turn, and writes the command's end in the same batch
  that settles that turn. Codex's own compaction marker is the success row.
- Claude: the command's turn is the translator's open turn until the result that
  answers the /compact input ends it. The command's own frames, such as the
  continuation summary, its echo and "Compaction canceled.", draw nothing.
- Both read the end with the one compaction rule: success needs the provider's
  report of the compaction; none after Orca's interrupt is a cancellation.
- The message resolves at the provider's receipt, as any send does: the Codex
  ack, or the Claude slash-command waiter on its result. The host writes a
  command's end only when the provider never took it.
- The delivery loop stops while a command's turn runs, and every journal commit
  re-wakes it through the session's serialize, so an end that lands while a step
  decides to stop is never lost. A child that ends first is settled with it.

* test(native-chat): pin a command's end to real /compact frames and to each path it threads

The captured /compact frames drive the Claude translator's command turn: a
finished compaction ends as a success with only the separator drawn; a stopped
one ends as a cancellation with no failure row, and the next send answers in its
own turn; a result naming another input ends nothing. The command's end is
checked at each point the ordinary result path threads through: the reopen latch
after a failure, the settling of a child still working, the context facts the
result reports, and the provider's own error row.

On the host: a message held behind a command is handed over when the command
ends just as the loop stops for it, a refused command settles as a failure and
the loop moves on, and a Claude child that exits mid-command settles the command
and hands what waited to a fresh child.

* test(native-chat): tests merged from the base state which turn their rows belong to

* refactor(native-chat): drop the child-end waiter nothing waits on

A command no longer waits for its child here: its turn ends from the provider's frames or from
that child's settlement, and the delivery loop is woken by the commit. The waiter and its test
were left from the earlier shape.

* fix(native-chat): a command holds the queue only while its child runs it

The delivery loop stopped whenever the journal showed a command's turn running. When the
command's child ended and its settlement could not be written, that turn stayed running with
no child to end it, and the loop's gate kept it from ever starting the next child, which is
what settles a gone generation's leftovers. Every later send was held for good, and Stop had
no child to end.

The gate now holds only while the conversation has a child: with none, the command belongs to
a gone generation, and the loop's start settles it like any turn a dead child left running.

* fix(native-chat): a Claude /compact succeeds only on its compaction boundary

The command's evidence counted Claude's `compact_result: 'success'` status as the compaction
done. That status comes before the boundary that replaces the history, so a Stop landing
between the two read as a finished compaction even though no boundary was ever written. Only
the boundary now counts, as the rule for both providers states; the capture's finished
compaction carries one, so it still reads as a success.

* fix(native-chat): a Claude child's exit says why the turn it ended stopped

When a Claude child exited mid-/compact, the command showed "Worked for 0s" and no reason. The
child's translator ends its open turn the moment the exit is reported, stamped with the exit's
instant, so by the time the exit settlement ran nothing was running. The settlement recognises a
turn the exit already ended by that same instant, but the Claude lifecycle event dropped it on the
way to the host, which then used its own clock, matched nothing, and wrote no row. When the clocks
did agree, the row was scoped to the running turn, of which there was none, so it landed outside
the turn it explained.

The exit's instant now reaches the host, and the exit row belongs to the turn the exit ended:
still running, or ended by the translator at that instant.

* fix(native-chat): a message waiting behind /compact draws below its live activity

A message sent while /compact runs waits on the host until the command ends. Both clients moved
it to the end of the transcript rows, but the running turn's live activity line ("Compacting the
conversation") draws after every row, so the waiting message sat between the command and its own
live status.

A row that is queued, and not what the live turn is for, now draws after that live activity: on
desktop outside the transcript window, below the activity line; on the phone in the list footer,
below the live status. A message whose own start is pending still draws above the activity that
start reports.

* fix(native-chat): only a running command holds a message below its live activity

A message is accepted, then handed over a moment later, and in between it reads as waiting. Every
message waiting behind a live turn drew below that turn's activity line, so an ordinary message
sent while the agent was working crossed below "Thinking" and jumped back up once it was handed
over, on desktop and phone. Only a conversation command's turn holds the queue on the host.

A message now waits below the live activity only while the running turn is one a command opened,
read from the entry that opened it. The phone test also typechecks, which the mobile test ratchet
requires.

* test(codex): the claim test names its notification params as a record

* test(native-chat): a read that reaches a crashed chat before the startup reconcile settles its turn

On desktop the chat on screen at relaunch reads before startup reconciles the leases, while the
dead process's lease still reads live. The open settles the turn it left running anyway, and the
restore that follows finds it settled.

* refactor(native-chat): drop the composer's second error formatter

After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.

* test(native-chat): pin the reason on a message rejected while its chat was closed

The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.

* docs(native-chat): drop the removed dispatch hold from six comments

A worker's session no longer takes a dispatch hold, and no release clock
rests a chat by visibility; the agent-launch comments, the abandon test,
the teardown test and the refusal census still said so.

* test(native-chat): rest the owner-status chat through the idle sweep, not a hold

The activation-gate test from #22808 put its chat at rest by holding and
releasing it, and passed the release-clock grace. This branch deleted both,
so the case threw before it reached its assertions. It now moves the host's
clock past the idle window and lets the sweep stop the agent and close the
conversation, then asserts the same owner answer and activation gate.

* fix(native-chat): show the structured pane's retrying line when a read fails

The read transport always hands the pane the host's words, so the error
state's "Orca keeps trying to load it" line, which showed only when there
were none, was never seen: the pane showed the host's text twice, as its
subtitle and on the status line under it. The structured pane now always
says its read keeps retrying, and the host's text stays on the status line.
The terminal-backed chat is unchanged.

* fix(native-chat): a send the provider never received after a restart has no verdict

Restart reconciliation rejects a crash-stranded send that is absent from a
trustworthy provider history with reason 'not_delivered'. Nobody failed that
send, but the verdict allowlist did not name it, so after a crash the chat
read Failed, was listed, and could notify "failed". Give the reason a shared
constant (persisted value unchanged), add it to the no-verdict set, and treat
it as an internal marker so the Retry row no longer shows the raw string.

* fix(native-chat): a failed Codex compaction's late completion writes no turn of its own

Codex ends a failed turn with an error and then still completes it as failed.
The error settled the compaction and released its claim on the provider turn,
so the completion read that turn as an ordinary one and wrote a stray record.
The claim now lasts until the completion, which adds nothing to a command the
error already ended.

* test(native-chat): the mid-command exit case resumes its next child as a real one does

The case's fake started every child as a newly created thread with the same generation. The
store refuses a created link once the conversation has a thread, so the next child's start
failed and wrote its own error row, which landed before or after the case read the journal.
The next child now resumes the thread under its own generation, and the case reads the
journal once the waiting message is delivered, which also proves the loop moved on.

* fix(native-chat): a /clear that never committed no longer locks the chat

A /clear wrote a durable "prepared, outcome unknown" record before starting
the replacement conversation. When that start was refused without a definite
answer (or Orca died), the record stayed forever, and while it did the chat
refused every send, /compact, a new /clear and rewind. Its only exit was a
rerun under the same operation id, which only the renderer held.

The record guarded nothing the process does not already know: a clear in
flight holds the session's serialize for its whole run and the command
controller refuses sends meanwhile, and the replacement's id and start
operation are pure functions of the clear's operation id. So the clear now
writes nothing durable before its commit, the gates refuse only a committed
clear (an older build's prepared record is inert), and a clear with no
committed answer reruns: a same-op retry re-attaches the same replacement,
a new op id runs a fresh clear.

A crash between the replacement's start and the commit leaves a replacement
record nothing points at. Verified: it has no tab, is not in the
replacement list, and a restart opens and starts nothing for it (restore
reads only the visible tab index); restart reconciliation releases its lease
like any dead owner's. In a live process its agent is stopped by the idle
sweep like any quiet agent. Session History lists provider transcripts and
only annotates them with an owner, so it can list this only if the provider
wrote a transcript for a thread that never got a message. Its record stays
on disk, as every closed chat's does; the store deletes none.

* fix(native-chat): a Codex rewind the provider did not keep no longer fails every attach

When Codex acknowledged a revert and Orca stopped before proving it, the
rewind stayed prepared with providerApplied set. On the next attach,
recovery read the provider's history, found the target turn still there
(provider-refused), and threw, because that settlement was limited to
reverts never sent. The throw ran inside the attach, so every attach, and
every send that needs one, failed for good.

The journal is replaced only once the provider proves the revert, so both
the provider and the journal still hold the target turn: settling the
rewind refused is consistent whether or not the provider acknowledged it.

* test(native-chat): a clear retried after a crash starts no second replacement

The replacement's id is the only thing that keeps a retried clear from leaving a second one, and no test held it across a restart.

* chore(native-chat): the clear rerun comment claims only the stable replacement id

* test(native-chat): wait for a send's background start before the refusal oracle removes its store

An accepted send wakes the delivery loop, which starts the agent in the background. The oracle's teardown disposed the loop but did not wait for that start, so its lease write could create a temp file in the store directory while the directory was being removed, failing the test with ENOTEMPTY about one run in four. The teardown now drains tracked starts before it closes the journals.

* fix(native-chat): a start a message waited on gets one failure row, the delivery loop's

When a queued message's start failed, two writers could report it under the same row: the delivery loop, when the adapter settled the start without proving it, and the exit settlement, when the child's exit landed. The last one won, so the chat's row could name a different cause than the one the message was rejected with, or be written twice.

The exit settlement now writes the start's row only when no message is queued and the loop has not already recorded that start. A start for a command, goal change or rewind, with nothing queued, still gets its row from the exit.

* fix(native-chat): a /compact whose start failed says to run /compact again

The failure-words context named only /clear as a command to retry, so a
/compact whose agent failed to start read "Send your message to try again."
on its row, its rejected message and the command reply. The context now
carries any conversation command; the host derives it from the oldest
message still waiting on the provider, which is the one a failed start
fails first, and the /compact reply names it directly.

* fix(native-chat): a Codex /compact ends only on its turn's completion, below Codex's own error row

Since only turn/completed ends a Codex turn, Codex's turn-ending `error` is a row
inside the still-open command turn, and the failed completion that follows it is
the command's end: completed, outcome failure, at the completion's receipt time.
The command's own "Compaction failed" row was written on that completion too, so a
failed /compact read its reason twice.

The command turn now notes when Codex's turn-ending error for the turn it carries
was written as a row, and its end then adds no second row. A retried stream error
ends nothing and is not counted. The flag that let the error end the command and
kept the claim until the completion is gone with the error-driven end.

A test replays the captured failed compaction from the real app-server through a
claimed command turn.

* test(native-chat): main's crash-turn test states its row's turn, and a dead /compact settles on its recorded exit

Two tests the main merge brought together:
- The crash-turn test from #23456 writes a turn record through the event sink
  without options; every row here states its turn scope, and a turn record's is
  the thread.
- The /compact whose exit settlement could not be written no longer stays running
  until the next start: main now settles an open chat from the exit it recorded, so
  the command reads interrupted before the next message, which is then delivered.

* test(native-chat): main's new journal tests state each row's turn

The crash-turn, stale-turn and sink-queue tests main added wrote rows without a
turn scope, which every item write now states. Rows written inside a running
turn name that turn; the sink-queue batch and a send handed over with no live
turn name the thread.

* fix(native-chat): draw a queued turn's message after the earlier turn's rows

A message sent while A runs is written to the journal when it is sent.
When the provider queues it (Claude answers it after A), A's remaining
rows - its last tool run and its answer - are written after that
message, and the message's own turn opens only after them. Grouping put
those rows in A's turn, but the transcript still drew them in journal
order, below B's bubble and bar, where A's answer read as B's reply. This
is the residual #23671 left open.

A message that opened a turn now draws after the earlier turns' rows the
journal wrote after it, just before its own turn's rows
(nativeChatTurnDrawOrder, returned by nativeChatTurnMembership as
drawOrder). Desktop and mobile both draw in that order. A steer, and a
message that has opened no turn yet, stay where they were written. It
applies on hosts that state turn scopes and, through journal order, on
older ones.

* test(native-chat): run #23026's Stop tests against #23059's command turns

Two of #23026's tests call APIs #23059 changed, and failed after the
merge:

- codex-structured-conversation-stop: a compaction now goes through
  adapter.compact with the command run the host wrote (#23059), not a
  bare turn id, and answers with the provider's receipt. With the command
  claimed, a Stop that names no turn while the compaction's provider turn
  has not opened still interrupts nothing.
- main-agent-working-agreement: a provider row states its turn scope
  (#23059's appendItem contract); the retry and subagent rows are
  conversation-scoped.

* fix(native-chat): typecheck main's Stop and restore-grouping code against #23059

A Stop's compaction interrupt reads the narrowed requested turn, and the
restore-grouping test states whether each row reports its turn's outcome.

* fix(native-chat): say a /clear cut off by a restart left the chat unchanged

A /clear retried under the same operation after Orca restarted could not reuse the new conversation its first try started, and its row said "Codex couldn't start. Run /clear again." The agent did not fail to start: the earlier try was cut off. The row now reads "This /clear didn't finish, so the chat is unchanged. Run /clear again to start fresh.", from a new clearUnfinished failure fact written through agentSessionFailureWords.

The clearUnconfirmed and conversationCommandUnconfirmed reasons stay, with their words, for older hosts that still send them.

* fix(native-chat): a retried /clear finishes onto the conversation its earlier try started

When an earlier try of the same /clear started its replacement conversation and a restart or the
idle sweep has since stopped it, the retry could not replay that settled start and reported the
chat unchanged. That replacement is a fresh conversation at rest, so the retry now commits onto it
and its first message starts its agent. A replacement whose start definitely failed still reads
that failure, and one Orca can't prove stopped still commits nothing. The clearUnfinished failure
kind this made unnecessary is removed.

* refactor(native-chat): stop recording that Codex acknowledged a rewind

A refused rewind recovery now settles as refused whether or not Codex acknowledged the revert,
so nothing reads providerApplied any more. Stop writing it and drop the hook that wrote it.
Records that still carry the field load as before; the schema ignores the extra key.

* fix(native-chat): a /clear retried under a new operation id finishes the same replacement

A /clear's replacement id came from the client's operation id, so a retry the client sent
under a fresh id started a second replacement and orphaned the first. The host now derives
it from this caller's oldest /clear since its last commit whose replacement start reached
the operation ledger, so any retry from that caller finishes the same replacement, including
after a restart. A /clear after a committed one starts a new replacement. Another caller's
/clear is refused only while such a replacement is running or not proven stopped. An older
client that resends the same operation id still lands on the same replacement.

* fix(native-chat): a /clear retry never repeats a failed start or waits on an unproven stop

A retry under a new operation id could pick an earlier try whose replacement start had already
failed, replay that failure and commit it again, so a user who had since signed in was told
they were still signed out. Such a try is now skipped, and the retry starts afresh.

Another window's /clear was refused while the first window's leftover replacement was merely
not proven stopped. Nothing but the first window's own retry would settle that, so the refusal
could last until its ledger row expired a day later. It now waits only on a replacement whose
agent is running.

* fix(native-chat): a /clear retry finishes only a replacement that started

A retry picked an earlier try whose replacement start never answered, because a crash left
that start unsettled. Replaying it could only repeat "couldn't start" or, with the old agent
unproven, refuse every /clear from that window. Only a start that succeeded left a
conversation to finish; any other try is skipped and the retry starts afresh.

* refactor(native-chat): a record's identity fields are built in one place

A created record and a founded one (a conversation no agent has run yet, at
rest) share who and where the agent is and how it launches. The founding
builder is used by the /clear commit that follows.

* feat(native-chat): the store commits a /clear and its new conversation in one write

commitConversationClear founds the at-rest replacement from the cleared
record's identity and writes the committed marker and tab move in the same
transaction, so neither can land without the other. It refuses to overwrite
an existing record under the replacement id.

* fix(native-chat): /clear starts nothing; the new chat's first message starts its agent

/clear used to start the new conversation's agent before it committed, so it
could fail on that start ("Run /clear again"), and a crash between the start
and the commit left a running conversation nothing pointed at. #23524 then
needed a ledger scan to find an earlier try's replacement, a nonce half of the
derived ids, a refusal of another window's /clear while a leftover agent ran,
and a check for a start that had already finished.

Now /clear opens the chat for writing (it no longer starts an at-rest chat's
agent either) and makes one store write: the at-rest replacement under a
random id, the committed marker, and the tab move. The first message in the
new chat starts its agent through the existing send and delivery path, fresh
because its handle chain is empty. A failed start shows on that message with
the typed failure and a Retry, and a conversation no agent ever ran now reads
"couldn't start" rather than "couldn't restart".

Deletes clearTryToFinish, otherCallersClearIsLive, the attach block and the
committed start-failure branch, and the tests of that retry machinery.

* test(native-chat): drop the /clear retry wording test; no start runs for a /clear now

* test(native-chat): another window and a phone read a /clear's replacement from the host

Both list the replacement the committed marker names, under the chat's tab,
and each one's session list shows it with nothing unread until its first
message runs. A reader that recomputed the id from the operation turns this
red.

* test(native-chat): a never-started replacement closes as settled

A worktree delete closes every chat in it and asks the user to force any it
cannot prove stopped. A replacement no agent has run is released, so its
close settles like any at-rest chat's.

* fix(native-chat): /clear settles an interrupted Codex rewind the way a send does

/clear moved from starting the chat's agent to only opening the
conversation. A Codex rewind cut off mid-way on a chat at rest can only be
settled by its agent, so /clear was refused as "rewind unconfirmed" every
time until the user happened to send a message. It now prepares like a send
or /compact: the agent starts only when such a rewind is in doubt.

* fix(native-chat): a chat whose agent is not running keeps its `/` commands

Claude reports its skills and project commands only from a running process,
and the host served the `/` menu only from the running agent. Now that
/clear starts nothing, the new chat's menu lost those entries until its
first message; a chat stopped by the idle sweep already did.

The host now keeps, in memory, the list a running agent last reported for
its launch (provider, host, workspace, account and launch arguments) and
serves it to a chat of the same launch whose agent is not running. A new
report replaces it; nothing is stored on disk, so a relaunch still shows
the short menu until the agent reports again, and no list is ever served
across accounts, workspaces or hosts.

* test(native-chat): queued drafts around /clear follow what a /clear now is

Three queue tests from #23726 are red on main 29c49aec31 itself:
- Two expected an older build's unconfirmed ("prepared") /clear to hold
  sends and drafts back. #23524 made such a clear inert, because it
  changed nothing; Send-now and the drain now go past it, like any send.
- One held /clear in flight by holding the new conversation's start,
  which /clear no longer makes. It now holds the one store write, and
  still sees a send refused and no draft left on either conversation.

* fix(native-chat): /clear opens the new conversation under its own lock

Carrying queued drafts after a /clear opened the new conversation while
holding only the old conversation's lock. Every other open runs under the
opened conversation's own lock, so that a concurrent reader cannot open a
second handle on the same transcript. The carry now opens it through the
same entry point everything else uses. A clear with no drafts still opens
nothing.

* test(native-chat): queued drafts reach a /clear replacement that never started

Under lazy start the new conversation has no agent when /clear answers.
Pin that the drafts have already moved there by then, on a queue paused
"cleared", and that the user's first message starts the agent and goes
ahead of them. Red with the carry removed.

* fix(native-chat): /clear stops the old agent before it records the clear

/clear wrote its marker first and the RPC handler stopped the old
conversation's agent afterwards. It now stops the agent first, under the
same lock, then writes the marker, so nothing the old agent does can land
after the clear. A write that fails leaves the chat usable: its next
message starts the agent again, as after an idle stop.

The stop releases the lease, which moves its fence, so the marker is
written at the fence the record holds after the stop.

* fix(native-chat): a Claude chat at rest reads its `/` menu from Claude's folders

Claude reports its skills and custom commands only while it runs, so a
chat whose Claude was not running (right after /clear, after the idle
stop, or after a relaunch) offered only the built-in `/` menu. The last
commit on this branch kept the last reported list in memory, which could
not survive a relaunch and could only repeat what a running Claude had
said.

The host now reads that surface where Claude itself reads it, on the
host that runs the chat, without starting anything: the workspace's and
the account's custom command folders (every `*.md` below them, named by
path with `:` between folders), the skills Orca's existing skill
discovery finds for Claude there, and the built-in commands Orca knows
Claude has. A running Claude's own report still wins. One scan answers
for a workspace and account for 10 seconds; a scan that finds something
new is pushed to the panes showing those chats. A chat run by another
host is never answered from this host's folders. The in-memory list is
removed.

* fix(native-chat): a Claude chat at rest keeps /model, /effort, /clear and /compact in its menu

Once the host sends a `/` list for a chat, the menu shows that list
instead of its own. The at-rest list started from Claude's text-driven
commands, which is empty, so a Claude chat whose Claude was not running
lost /model, /effort, /clear and /compact from its menu. It now starts
from exactly the menu a chat at rest showed before, then adds what the
folders hold.

A scan that never answers also held the next one off for good, freezing
the menu until relaunch; one unanswered for 30 seconds is now given up
on and the next read scans again.

* fix(native-chat): an at-rest `/` scan keeps any newer answer and never piles up

Giving up on a scan after 30 seconds threw away every scan that took
longer than that, so a slow folder never updated the menu, and a folder
that stayed hung started another stuck walk every 30 seconds, each
holding one of Node's few file-system threads.

Each scan is now numbered and its answer is kept whenever it is newer
than the one already kept, however long it took; an older answer that
lands after a newer one changes nothing. A new scan starts beside an
overdue one, but never more than two run at once.

* fix(native-chat): a failed at-rest `/` scan no longer discards an older answer

A failed scan marked itself as the newest answer, so an older scan that
answered after it was ignored and the menu stayed without custom
commands until the next scan. A failure now only starts the 10-second
wait. Test pins that the wait runs from when a scan lands, a failed one
included.

* test(native-chat): open the at-rest command host on main's shared journal database

* refactor(agent-session): the transaction queue opens the session store file

AgentSessionRecordStore.open hardened permissions, loaded the file, marked every
lease unreconciled, built the transaction queue and persisted a pending rewrite.
That is the queue's load lifecycle, and the queue already applies the same
unreconciled rule when it reloads an externally changed file. Move it to
AgentSessionStoreTransactionQueue.open beside fromLoadedStore, and drop the
exported wrapper that existed only for the store's open.

No behavior change; the store's public API is unchanged. Brings the record
store back under max-lines after commitConversationClear.

* test: open the store on the journal database and pass the close cause, where main's tests still used the old calls

#24006 moved the store into the journal database and #23684 added a test on the
old open call; main's close now takes a cause. Four tests catch up.

* refactor(native-chat): a provider's at-rest commands are one adapter member

The at-rest `/` surface's read and change listener travel together as
`atRestCommands`, which the Claude catalog already is, so the adapter types
stay within their line limit after main's growth.

* test: the startup-reconcile tab close passes the close cause main now requires
2026-09-30 23:21:59 -07:00
OrcaWinandm4air 3135fbbf49 feat(runtime): pin the Node 24.21.0 server runtime with an offline CI check (#24087)
* feat(runtime): pin the Node 24.21.0 server runtime with an offline CI check

Add src/shared/node-runtime-pin.ts (NODE_RUNTIME_PIN, SERVER_TARGETS,
NODE_RUNTIME_ASSETS for all 8 server targets plus the headers tarball),
generated by config/scripts/update-node-runtime-pin.mjs from the nodejs.org
and unofficial-builds SHASUMS. check-node-runtime-pin.mjs verifies, with no
network, that the pin tracks the locked Electron, matches engines.node's
major, and covers exactly SERVER_TARGETS; it runs in the static analysis job.

ORCAD_BUN_TARGETS consumers now read SERVER_TARGETS so there is one target
list; orcad's Bun runtime and build output are unchanged.

* fix(runtime): reject a pinned archive that belongs to another target

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-30 22:57:10 -07:00
Brennan Benson 07e9fdfd13 fix(native-chat): a message the chat said was not sent is never sent later on its own (#24232)
* fix(native-chat): keep a message the chat said was not sent held until its Retry

A native-chat send the host refused (for example "Chats were saved by a newer
Orca. Your message was not sent.") or that never reached the host showed "not
sent" with a Retry button, but the hold that stopped it lived only in the
outbox hook's memory. The message itself was saved in the outbox, so the next
launch lifted the hold and sent it with no Retry; a copy the user retyped in
the meantime was held behind it and went out as well.

The hold is now read from the failure the message already saves: a queued
entry that carries its last failure waits for the user's Retry, on this launch
and every later one, and an entry saved by an earlier build in that shape is
held too. The drain passes over a held message instead of stopping behind it,
so what the user sends next goes out as they send it. Retry clears the saved
failure. On a host from before accepted-send, a new agent owner observed while
the chat is open still sends the refused message again, as before; a relaunch
does not.

A held message keeps its operation id, and a host refuses an id older than a
day as expired for good, so its Retry could never go through; a new id could
deliver a message an earlier attempt already delivered. Such a message now
goes back to the composer with a notice to check the chat before sending it
again, and leaves the outbox.

* fix(native-chat): keep refused messages as rows until Retry, and only release them for an older host's new owner

- A message the host refuses as expired under an id it kept stays a saved row
  reading "Orca couldn't confirm what happened. Check the chat.", and its Retry
  sends it under a new id. It no longer moves into the message box, where an
  automatic resend after a relaunch could put text the user never asked for,
  held only in memory.
- An owner change releases a refused message only on a host known to predate
  accepted sends, and only for the refusals such a host gives while it restarts
  the chat's agent. Those rows say Orca will send it again when the agent
  restarts, beside their Retry. A host whose capability check has not answered,
  or failed, no longer releases anything.
- Every failed message ahead of the one the queue stopped on keeps its Retry,
  since that Retry sends it at once.
- A journal row saying the host cannot tell whether a message landed, and the
  unconfirmed probe's resend, replace an earlier attempt's saved failure, so
  the message is probed rather than held.
- The drain stages from the hook's own outbox, so a hold kept only in memory
  after a failed save survives the next send; the hold is written once more
  after that failed save.

* fix(native-chat): a refused message waits for its Retry on every host, and a send is staged from the latest outbox

A message the chat showed as not sent no longer goes out on its own when an
older host's chat gets a new agent owner. Resending it on the owner change
sent it after messages typed later, still delivered a retyped copy twice,
and its "Orca will send it again when the agent restarts" row promised a
resend that often never came. It now waits for the user's Retry, as it does
on every current host. A send still in flight when the owner changes is
still sent again under its id; it was never shown as failed.

The drain admitted and staged the next send from the render's outbox. An
owner change requeues the send it interrupted in an effect earlier in the
same commit, and staging from the render's list wrote the old list back,
leaving that send stuck as sending. The drain now reads the latest list,
which still carries a hold kept only in memory after a failed save.

Tests pass the view's target as one stable object, as the view does: a new
object each render re-ran the owner-change requeue, which hid the drain bug.

* fix(native-chat): keep a not-sent message out of newer turns, and word its saved cause only when seen

A message shown as not sent stays in the outbox and draws below every
newer turn. It was an ordinary user row there, so it counted as the newest
user row: while a new send waited for its turn to open, that turn's
"Working for" clock drew under the old message, and once the turn ended an
empty "Worked for" divider was left under it. The projection now marks
such a bubble (held for its Retry, or rejected) as unsent; turn membership
gives it no turn and never makes it the live one, on hosts that state turn
scopes and on those that do not; and the transcript draws it after the
live activity, as it draws a message waiting behind /compact. Mobile has no
outbox, so its rows never carry the mark and its grouping is unchanged.

A held message read back from storage repeated the cause it was saved with,
which may no longer hold: "Update Orca to keep using them" after the user
updated Orca. The outbox hook now remembers, in memory only, which messages
failed while the chat was open; only those word their cause. Any other held
message reads "Your message was not sent." with its Retry, and a Retry the
cause still stops brings the full words back. An expired id keeps its words,
since that cause cannot clear.

* fix(native-chat): follow the bottom and light a tick for a chat whose only rows are not sent

A message shown as not sent draws after the windowed transcript. When it was
the only row, the windowed list was empty, and following the bottom or "Jump
to latest" asked the virtualizer for an end it computes from its own rows:
the top. It now scrolls to the container's own bottom when no row is
windowed.

A rejected send the journal recorded keeps its place but opens no turn, so
the rail lit no tick when it was the row being read. A user row in no turn
now lights its own tick.

Also pins that a refusal seen while the chat is open reaches the rendered
notice in full, and reads only "not sent" after the chat is reopened until a
Retry is refused again.

* test(native-chat): name the relaunch test parameter for how the refusal arrives

The low-evidence lint rejects "shape" as a symbol name.
2026-09-30 22:49:26 -07:00
Kelvin Amoabaandfruit c0ac2b1fcc feat(browser): add a rebindable shortcut for Annotate page element (#23879)
Fixes #23470

Co-authored-by: fruit <200041037+guozi-lab@users.noreply.github.com>
2026-09-30 22:09:26 -07:00
Brennan Benson a4606ccae3 fix(cli): orca file open no longer moves your view unless you pass --focus (#24244)
* docs(cli): file open/diff/open-changed say they switch the user's view and are for user requests only

Refs #9944

* fix(cli): file open/diff/open-changed leave the user's view alone unless --focus

`orca file open`, `file diff` and `file open-changed` always switched the
desktop to the target worktree, selected the tab and revealed it in the
sidebar. An agent skill that opens its answer pulled the user out of whatever
they were typing in (#9944), and a phone opening a file moved the desktop too.

The commands now add the tab in its worktree without changing anything on
screen, including when that worktree is the one being viewed: the new tab is
added to the tab bar but the active tab, tab type and focus stay put. In a
worktree the user is not viewing, the tab becomes that worktree's selection so
it is in front when they go there. `--focus` keeps today's behavior.

files.open / files.openDiff take an optional `navigation` target (the existing
RUNTIME_NAVIGATION_TARGETS vocabulary); the CLI sends 'all' for --focus, like
`worktree create --activate`, and nothing otherwise. The renderer moves the
host view only when the target reaches the host; a missing field (phones,
older CLIs) leaves it still. Editor opens for a worktree other than the
on-screen one no longer write the global activeFileId/activeTabType.

Refs #9944

* test(cli): justify the window and runtime stubs in the file-open notification test

* fix(cli): keep phone file opens switching the desktop; the CLI asks for 'caller'

Phone opens send no `navigation` field, and the phone's diff-review "Open in
session" relies on the desktop selecting the diff it opened. A missing field
now keeps the original switch exactly; the CLI says what it wants instead:
'caller' (no host move) by default and 'all' for --focus. Older CLIs, which
send nothing, keep switching as they always have.

Refs #9944

* fix(cli): background file opens select the tab without counting as a visit

A CLI open into a worktree the user is not viewing selected the new tab with
the same activation a user click uses, which stamps lastFocusedAt and the
group's recency list. The worktree jump palette sorts recent tabs by that
time, so every agent `orca file open` into another worktree jumped to the top
of the user's recent tabs.

Editor opens now take a selection mode: 'focus' (default, unchanged),
'background' (select within its worktree without recording focus or recency)
and 'none' (add only). createUnifiedTab and activateTab gain recordFocus:false
for the background case.

Also: tests for reopening an already-open file or diff without --focus, a
comment that file opens move only the host window ('all' acts as 'host'),
root help lines back under 100 columns, and an accurate remote test title.

Refs #9944

* fix(tabs): a background-selected tab still joins its group's tab history

recordFocus:false skipped both the focus-time stamp and the group's
recentTabIds append while still making the tab the group's active tab. Ctrl+Tab
looks the active tab up in that history, so after a background CLI open it
did nothing (or went to the wrong tab) once the user switched to that
worktree, and hydrate kept the broken history across a restart.

Only the focus-time stamp is skipped now; the jump palette's recent rows sort
by that alone, so the palette fix stands.

Refs #9944

* fix(cli): file open/diff/open-changed --focus help says it brings the user to the file

The three commands borrowed the shared --focus line written for terminal
create ("Reveal the created terminal session in Orca"). They now use the
per-command flag help table; terminal create's line is unchanged.

Refs #9944
2026-09-30 20:46:47 -07:00
Brennan Benson 12b8ef8c0b fix(worktree): update local main safely, once per branch, alongside the checkout (#23698)
* fix(worktree): retry local main refresh through git lock contention and skip false alarms

* fix(worktree): overlap the local main refresh with the checkout and run one refresh per repo at a time

* fix(worktree): skip the local base refresh when the create makes that branch itself

Creating a workspace named feature-x from origin/feature-x runs `worktree add -b feature-x`,
which now overlaps the refresh. The refresh's drift probe could see refs/heads/feature-x
missing and its presence probe then see it (the add just wrote it), which reported
"not fast-forward" and showed a sticky "Local feature-x was not refreshed" warning.
`-b` refuses an existing branch, so there is nothing to refresh in that case: skip it on
the local, prepared-checkout and SSH create paths.

The SSH overlap tests move to their own file so the existing suite stays under the line limit.

* fix(worktree): say plainly what happens after the local base refresh queue wait expires

* test(worktree): prove SSH local base refreshes of one repo run one at a time

* test(worktree): drop type assertions from the SSH refresh overlap test mocks

* fix(worktree): fast-forward local main with one host-owned merge --ff-only per branch

Moves the whole local base refresh into one shared routine that runs on the
execution host (main process for local and WSL repos, the relay for SSH), so
the app no longer keeps a second copy of the checks, queue and retry.

A checked-out branch now moves with merge --ff-only (hooks, auto-gc and
autostash off) instead of status-then-reset --hard, which silently overwrote an
untracked file the new commit adds and could discard an edit or a commit made
after the check. A free branch moves with a compare-and-swap update-ref that
writes a reflog message. Status reads no longer take index.lock.

Creates of one branch share one run plus at most one trailing run; a create
waits at most 30 s and never starts a competing mutation. The failure toast is
keyed by repo and branch because every create that joined a run reports the
same fact.

* fix(worktree): fast-forward local main even when the repo requires signed merges

With merge.verifySignatures=true, the owner-checkout fast-forward refused an
unsigned origin/main tip, so every create warned "Local main was not
refreshed" where the old reset moved main. The new workspace is already
created from that same unsigned commit, and a branch that is not checked out
moves without a signature check, so the refusal protected nothing. Turn the
setting off for this one merge, like the hooks, gc and autostash overrides.

* fix(worktree): clear git read caches when a shared local main update lands late

The update of local main can finish after a create stopped waiting for it, so
the shared run now invalidates git read caches itself. The index.lock real-git
test also no longer reads the developer's global git config.

* fix(worktree): never overwrite an ignored file when fast-forwarding local main

A plain `git merge --ff-only` silently replaces an ignored file (for example a
local `.env`) at a path the new commit starts tracking. Pass
`--no-overwrite-ignore` so git refuses instead and the create reports the
checkout as having local changes. Supported on the fast-forward path since
well before Git 2.25.

Also make the relay test for one-refresh-per-branch hold the first merge until
the second request has reached the relay, so it fails without the coalescing.

* fix(worktree): keep the local main update a plain fast-forward whatever the user's merge settings say

A per-branch mergeOptions such as '-s ours' or '--squash', or pull.twohead=ours,
made the update create a merge commit that dropped upstream, or stage upstream
without moving main, while reporting success. The command now clears the
branch's mergeOptions and passes the strategy and signature choice on the
command line, which beats any config. After the move Orca confirms local main
is exactly the target before reporting it updated. The exact command also runs
in the Git 2.25 compatibility suite.

* fix(worktree): make the Git 2.25 fast-forward contract pass in CI and rerun on every change to it

The new real-Git contract for the local main fast-forward wrote a post-merge hook into .git/hooks, which does not exist when the repo is created by the uninstalled Git 2.25.5 build CI uses (no templates), so the Git compatibility check failed. Create the directory first.

The Git compatibility check also did not run when only the fast-forward module changed, so a later edit to its merge arguments (for example a flag Git 2.25 lacks) would skip the one check that tests them. Add the module to the check's paths.

* fix(worktree): answer every create from a local main update toward its own base

Creates from different remotes' main (origin/main and upstream/main) shared one queued update
per repo and branch, which ran only the latest caller's target: a create could get no result
for its own base, or a false "not refreshed" warning computed for another remote's main.

The per-branch runner now queues one run per distinct target, still one at a time per branch,
and only callers toward the same target share a queued run. Applied in the app and the relay.

* test(worktree): record the third create's result in the mixed-remote burst tests and update the toast id rationale

* chore(worktree): correct the toast id rationale
2026-09-30 18:30:22 -07:00
Brennan Benson 0b79720c2e feat(native-chat): the chat strip and the sidebar read the host's child records (#22614)
* feat(native-chat): publish the host's child records to the status summary and the chat strip

The status summary and the background-task channel now read a session's child
records from the host's canonical store, through the status sink its row
landed in, and derive the legacy task and subagent shapes from the same views.
The parent row folds its child-work liveness from those records at ingest,
not from the summary's task list. The adapters no longer push their task DTO
to clients: the onBackgroundTasksChanged path is gone, and a child-work ingest
is what republishes both the summary and the strip. Finished children stay
listed until the session's own next turn starts. A reader that predates child
views never receives a roster whose rows are all settled.

* feat(native-chat): the chat strip reads the host's child records with its parent's verdict

The strip's roster now renders from the child views its channel carries, and
passes the verdict the session's own status row gives its children, built the
way the sidebar builds it (the row's freshness and the status feed's
observation). So one child reads the same in the strip and the sidebar, live,
after the transport drops, and once the row goes stale. A roster of finished
children stays shown until the next turn but no longer animates the monitoring
indicator or blocks conversation commands.

Tests: an end-to-end run on a host with no renderer (a real hook server as the
status sink) shows the summary and the strip channel carrying the same records
at every step, the parent row folded from them, retention, and an older
client's task list holding live work only; a wired renderer test shows both
surfaces agree when live, lost and stale.

* test(native-chat): a newer host's view degrades, an older reader keeps its live roster, a finished roster holds nothing open

- The view decoder ignores unknown keys, degrades unknown kinds, states,
  outcomes and memberships, and drops only rows it cannot identify.
- At the RPC boundary a reader that predates child views gets no strip for a
  roster of finished children and never the views themselves; a stop-only
  reader keeps rows whose host offers no targeted stop.
- The strip shows finished children without reading them as live work.
- The row keeps its child list's identity when a summary repeats it.

* refactor(native-chat): the summary's task list is the legacy projection's live rows, unfiltered

* test(native-chat): type the switch tests' mocks instead of asserting them

* test: remote clients advertise reading child views

* docs(agent-status): the structured row folds the store's child records

* fix(agent-status): keep the view reader's header from reading as a value import to the renderer boundary

The renderer node-builtin boundary test scans raw text, so a header comment
that said "imports" ahead of the import block turned the type-only import of
agent-status-child-work into a value edge that reaches node:crypto.

* refactor(native-chat): the status summary's broadcast equality gets its own module

The status feed crossed the file-size limit once the summary gained the main agent's turn
outcome beside the child views. Which summary changes reach every session list now lives in
structured-agent-session-status-summary-equality.ts.

* fix(native-chat): command admission reads the strip's child records

A conversation command was refused on the provider tracker's own roster
while the strip read the host's child records, so a drift between the two
rule sets could refuse /clear with a stop instruction the strip had no
button for. Admission now reads the same records through the same read as
the strip, uses the strip's liveness fold, and asks for a stop only when
the strip renders one. The adapter contract no longer exposes the tracker
roster, so no host decision can read it.

Also records when the legacy child shapes die, every earlier death of a
settled child, and the display-precision invariant behind the summary's
clock tolerance.

* refactor(native-chat): command admission takes only what it reads of a turn

* fix(native-chat): the session list drops a session's children when the store does

A session's end no longer removes its child records: a child still running
settles with an outcome nobody reported, and a finished one stays listed.
Records now leave only at the session's own next turn, at the cap on settled
records, or when the host lets go of the session and its row leaves the store.

The summary kept after the host lets go used to strip its children on close,
a rule of its own. It now re-reads them from the store when the row leaves,
through the same read every live summary uses, so the session list and the
chat strip list the same children at each step, including a forget with no
close. Closing only revokes ownership, as before the child records existed.

* test(native-chat): write the Codex frame script's parent row out step by step

Once every surface reads the child records, the provider tracker's roster is
no oracle: it and the records read the same child executions, so agreeing
with it cannot catch a defect in either. Each frame now states the child
liveness and the parent row it must fold to.

* fix(native-chat): the idle sweep and the restart snapshot read the host's child records

The idle sweep (keep an agent running while its subagents or commands run) and
the restart-resume snapshot (what a chat was doing when Orca stopped it) both
read the provider tracker's roster through the adapter interface, which no
longer carries it. Both now take the host's one child-record read, the same one
the status summary, the chat strip and command admission use.

The snapshot's working test also folded that roster through the shared fold's
old `backgroundTasks` input, which the fold no longer reads, so a settled lead
whose subagent was still running would have been offered nothing. It now hands
the fold the records.

* test(native-chat): the child-record tests follow the merged command lifecycle

A command is a live child record from its start and is removed, not settled,
when it stops, whatever Codex tagged it. The end-to-end switch now shows the
child's `npm test` as a live row beside its dev server, and both are gone once
they exit; only the finished subagent stays listed until the next turn. Letting
go of the session is its tab closing, since a closed conversation whose tab
remains keeps its row.

Command admission's finished row is a subagent, the one kind that settles, and
the failed-verdict row test admits its live subagent as a host record, the only
thing the row folds.

* refactor(native-chat): the status feed's journal projection cache gets its own module

The status feed crossed the file-size limit once the child records joined the
agent-start signal and the completion feed's status read. The per-journal
projection, cached per commit, now lives in
structured-agent-session-status-journal-projection.ts.

* test(native-chat): the admission test's compaction resolves with a real outcome

Main's compaction result is a tagged outcome; the host-level admission test
resolved its mock compaction with an empty object.

* feat(native-chat): the sidebar lists running subagents; the strip, running then the newest finished

The host keeps every child record; what each surface lists is picked from them on every read, so
nothing is stored twice. The status summary, which every session list reads, now carries only
running children (and a finished one whose shell still runs, which reads monitoring): a finished
or failed subagent leaves the sidebar and stays in the chat's strip. The strip lists every running
child, then the newest finished ones, 100 rows in all; more than 100 running all show.

This matches common practice: sidebars show live subagents, and finished ones stay in the chat's
panel, newest first. No wire field is added. An older client reads fewer rows: its legacy task
lists were already live-only in the summary, and the strip's settled tasks come from the same
bounded roster.

* fix(sidebar): one rule for what the worktree sidebar lists: running children, from every source

`worktreeSidebarListsChild` is the one definition: a child that runs, counting a finished one
whose own shell still runs (it reads monitoring). The sidebar's row builder applies it to every
child source it reads, a terminal agent's hook roster and a chat session's records alike, and the
host's status summary applies the same predicate, so the sidebar's payload stays small. The chat's
strip keeps finished children, newest first.

A terminal agent's hook roster already drops a child on its own stop, so nothing changes there:
a teammate between turns and a child gone quiet still run, and still show. The selection module
moves to `agent-child-work-listing.ts`, since it now covers every source, not only chat sessions.

* test(native-chat): the switch test passes the startup child key main's status bar takes

* fix(native-chat): a finished child stays until the user's next send, not a turn Claude opens on its own

Claude wakes the agent on its own when a background task ends, and that wake is a
new root turn. Keying retention on the newest root turn retired every finished
child about two seconds after a background agent or shell finished, so its
outcome never showed in the strip.

Retention now keys on the user's newest send the provider accepted (a message, a
steer or a command), with the journal epoch so a rewind still retires. A turn the
provider opens itself and a subagent's turn carry no send. Replayed captured wake
orders through the real adapter, hook server and status feed.

* fix(native-chat): a background Stop reaches the tasks the child records show

The strip draws a row's Stop, and /clear, /compact and rewind wait for background
work, from the host's child records, but the Claude adapter still resolved which
tasks a Stop reached from its own tracker's roster, and refused to stop at all
once that roster was empty. A task the records kept live after the roster dropped
it showed a Stop that sent nothing and blocked those commands until the chat tab
closed.

The host now resolves the provider ids a Stop sends from the records (the same
per-row rule the strip and admission use; every such row for stop-all), and the
adapter stops exactly those, with no tracker guard. An acknowledged stop ends the
record: a running task sends its own stopped frame first, and the CLI answers
success with no frame for a task it no longer knows. A refused stop leaves the
record live. No production code reads the tracker's roster any more.

* fix(native-chat): one rule for a finished child that still owns live work, at any depth

The listing kept a finished child whose work ran through any depth of ownership,
but retention at the user's next turn protected only the direct owner, so a
finished agent whose finished subagent still ran a shell was removed and that
subagent jumped to the top level. Both now read settledOwnersOfLiveWork.

* fix(native-chat): an older client sees a Codex child's shell as it did before views

Clients that predate child views read a flat task roster derived from the views.
It listed a Codex child agent's shell as an extra row beside the running agent,
then as a bare command once the agent finished. The derivation now hides a
running agent's commands and names a finished agent's as "<agent> — <command>",
as the Codex tracker did; the label rule moves to a shared module both use.

* fix(native-chat): the chat decodes a roster's child rows once, as the frame arrives

The client reducer compared raw wire rows, so a row shaped by a newer host could
throw there, and the strip decoded a new array on every render, which defeated its
grouped-rows memo while a turn streamed. Rows are now decoded where the frame
enters the reducer, an unchanged roster keeps its identity, and the strip's parent
context is rebuilt only when one of its values changes.

* fix(native-chat): the strip channel forgets a closed conversation's roster

It kept the last roster fingerprint of every conversation for the host's lifetime.
The conversation map now tells observers when one leaves it. Also corrects the
summary's children comment: it carries running children only.

* docs(native-chat): rewrap the retention comment

* fix(native-chat): a task's own ending replaces a Stop's, and a child finished after the user wrote stays

Two lifecycle gaps from the round-1 fixes.

A Stop acknowledged ahead of the task's own ending relabelled it. The SDK hands
Orca a control answer as soon as it reads it and queues other frames, so when a
task finished just as the user pressed Stop, the acknowledgement arrived before
the task's completion the CLI wrote first, and the task read "Stopped" with its
result lost. An acknowledged Stop now ends a record provisionally (outcome basis
`stop-acknowledged`); the task's own terminal frame replaces it, and nothing
replaces an ending the task reported itself.

A finished child still vanished with no new action from the user when the send
the provider took was written before the child finished: a steer Claude takes at
its next boundary, or a queued draft handed over at the end of the turn. The
user's next turn now carries when they acted (the send's written time, or the
draft's queued time, both on the host clock), and only children that finished at
or before it retire; a later one stays until the user's next send. A rewind
still retires every finished child.

Also: a Stop-all keeps stopping the remaining tasks after one request fails, then
reports the failure.

* fix(native-chat): the strip keeps one empty list for a roster that omits one

A roster with only running or only finished rows made a new empty array on every
render, so the strip regrouped its rows each time. One shared empty list keeps
its memo.

* fix(native-chat): a strip row whose owner the 100-row budget cut renders under the main agent

The budget can keep a finished child and cut its finished owner. The child still
named that owner, so it rendered nowhere. The selection now clears an owner it
did not keep, as the view contract says for an owner outside the projection.

* fix(native-chat): a stop-all that times out stops asking, and "no children" is sent once

A stop-all kept asking after a request timed out, so a Claude CLI that stopped
answering control requests cost one full deadline per task, while the chat's
sends, Stop and /compact waited behind it. A timeout now ends the loop; other
request failures still let the remaining tasks be stopped.

The chat strip channel never remembered that it had sent "no children", so
every change in a chat with none re-sent that frame to each subscriber. It now
remembers it, and forgets only when the conversation closes.

Also renames agent-child-work-stop.ts to agent-child-work-stop-targets.ts, which
says what it answers: the provider ids a background Stop reaches.

* fix(native-chat): the status feed reads a journal snapshot with no submissions, and e2e tests use the child-work reader

CI on dbd2439cd4 was red in three places:
- first-work-branch-rename and the agentSession.subscribeStatus RPC test feed the status feed a
  journal whose snapshot lists items only. The projection reads the user's newest accepted send
  from `snapshot.submissions`; it now tolerates their absence, as the status projection beside
  it already did.
- the cross-version downgrade test still passed `backgroundTasks` to the teardown's working
  marker, which now takes `childWork`.
- an e2e unit test still gave the status feed the removed `readBackgroundTasks` dependency
  (harmless at run time, a type error in the tests/ project).

* fix(native-chat): the chat strip lists running children only, by the sidebar's rule, and hides when none runs

A finished subagent's result is already in the transcript ("Ran N subagents ·
completed"), so the strip is for work that runs. It now lists exactly what the
sidebar lists, by one predicate (a running child, or a finished one whose own
shell still runs, which reads monitoring), and the host sends no roster once none
runs, so the strip hides.

Gone with it: the 100-row budget and the running-then-newest-finished
selection, the re-homing of a child whose owner the budget cut, and the RPC
gate's rule for a roster of finished rows only (no such roster exists now).
Older clients still get their derived task list, running work only.

Finished records still stay in the host's store until the user's next accepted
message: they refuse a late frame of their run, let a task's own ending replace
an acknowledged Stop's, and keep a running shell's owner. Dating that retention
by when the user wrote the message only kept finished rows visible longer, so it
is removed.

* fix(native-chat): the strip shows running work only from any host, and hides after a released session's last child

- A new app paired with an older host no longer shows that host's finished task
  rows: the strip lists running work only, whatever host sent it, and hides when
  an older host's roster has only finished rows left.
- A test for the path that hides the strip when a session's last running child
  settles after the provider let go of the session (Claude's release path): the
  channel sends `null` though no provider answers for the session any more.
- A test comment still described the strip keeping finished children.
2026-09-30 18:23:24 -07:00
Brennan Benson 24540300f0 fix(agent-status): preserve hook presence when process checks cannot answer (step 1 of 3) (#23947)
* fix(agent-status): admit hook process presence on the execution host

* fix(agent-status): restrict process checks to real hook ingress

* fix(agent-status): keep presence checks from causing false exits or losing real ones

- Pin the macOS process start time to UTC on both the hook and the host so a
  shell TZ or a time-zone change cannot turn a live Claude into an exit.
- An unanswered process check falls back to the foreground confirmation, so
  Codex, SSH and Windows panes still leave the agent state on a real exit and
  the Codex late-completion recovery still runs.
- A nested agent that inherits the pane key cannot take over the pane's
  presence; a retired session is replaced by the next session even when its
  SessionStart was lost.
- Drop the unused Windows process read (no Windows hook captures an identity
  yet) so shared code no longer imports main-process modules.
- Skip the capture outside Orca panes, gate relay re-checks to real title
  changes, and list the capture module in the CLI project.

* refactor(agent-status): own pane presence by the agent's process, not its session

Presence now exists only when a hook carries the agent's process identity,
and only that process's evidence changes it: its SessionEnd ends the pane,
its /clear and /resume keep it, and hooks from any other process (a nested
agent inheriting the pane key) update status without taking ownership.
Hooks without an identity (Windows, sessions started before the capture)
behave exactly as before, so a nested agent can no longer end a pane it
does not own. An ended owner stops answering 'exited', and the host probes
the owner only when another process reports in the pane.

* fix(agent-status): close round-3 review gaps in presence handling

- Relay retries and transcript polls schedule against the row the relay
  cached, and identity-less events pass through the transition unchanged,
  so SSH Grok replies and Codex transcript polls deliver again.
- A live process check restores the runtime's agent status and releases
  queued orchestration mail, like a foreground read that finds the agent.
- An answered foreground read naming a non-agent (wsl.exe, tmux) is still
  an exit; only silence is not.
- A suspended (Ctrl-Z) agent is unverifiable, not live, so the foreground
  read decides as before.
- An agent of another type started mid-turn cannot own or end the pane.
- Replayed spool hooks check each pane once, and SessionEnd ends presence
  only for reasons that end the process.

* refactor(agent-status): record the pane's owning agent even without a process id

Ownership is now decided only by comparing the recorded owner with the
sender, never from the row's agent type or turn state (which identity
resolution rewrites). The first agent hook in an empty pane claims it; a
live owner keeps it against any other agent (a nested claude -p, a Codex
started inside Claude or the reverse); only the owner's own proven process
can end it. An owner no hook identified never ends from a hook and is not
probed, so those panes behave as before.

* test(agent-status): read the optional process id in the relay presence test

* fix(agent-status): other agents' SessionEnd hooks settle their status again

Only an admitted exit (Claude's process-ending SessionEnd, or a host-proved
exit) is marked ended on the event, and ownership keys on that marker, not
the hook name. Devin, Qoder, CodeBuddy and Copilot SessionEnd hooks are
ordinary status updates again, locally and through the relay.

* test(agent-status): cover the owner's own SessionEnd through the relay

* fix(agent-status): rows without a process identity keep today's command-finished cleanup

The renderer's command-finished cleanup kept every row when its shell check
could not answer, which left Codex, hookless-agent and old-relay rows over
SSH showing done after a real exit. It now asks the host whether the pane's
agent process can be checked: only a pane with an identified, running owner
keeps its row on an unanswered check; every other pane drops it exactly as
before. The drop stays armed while the host answers, so a new command still
cancels it, and a missing or failing answer (web client, older host) keeps
today's behaviour.

* fix(agent-status): panes without an identified owner keep today's exit confirmation

The title-driven exit confirmation applied 'silence is never an exit' to
every pane. It now applies only when the pane has an identified owner whose
process cannot be checked right now; a pane with no process identity
(Codex, hookless agents, Windows, old relays, sessions started before the
update, or an owner that already ended) confirms exits exactly as before.
2026-09-30 18:21:46 -07:00
Brennan Benson 3ab3c9239f fix(native-chat): the working line shows only what the agent is doing now (#24218)
* fix(native-chat): the working line shows only what the agent is doing now

A chat's live "Working…" line could show an old notice, such as "Claude hit a
temporary problem and is retrying.", long after the agent had moved on and was
running new commands. When the host had no live activity for the turn, the line
fell back to the newest status row in the turn, and any status row qualified:
retry warnings, "Context compacted", "Cancellation requested.", and notification
summaries. Those rows record the past and already appear in the transcript.

The line now reads only the host's live, per-turn activity, which is never saved
and is cleared at turn boundaries. Without it, the line says Thinking or
Working…. Desktop and mobile share the selector, so both change.

* test(codex): guard that a subagent's compaction never becomes the parent's live activity
2026-09-30 16:44:29 -07:00
Brennan Benson 6f2a7d05c9 fix(worktrees): let git delete removed checkouts so chat sends never wait behind them (#23837)
* fix(worktrees): delete removed checkouts in git, not in Orca's file pool

Local worktree removal renamed the checkout into a sibling trash root and
deleted it in the background with a recursive fs.rm in the main process.
That queued one request per entry on libuv's shared 4-thread file pool, so
for minutes every other async fs call in the main process (the agent-session
store behind chat sends, file explorer reads) waited behind the delete.

`git worktree remove` now deletes the checkout inline in git's own process
again, so the card stays in its Deleting state for the length of the delete
while Orca's file pool stays free. No timeout applies to the call, so a
large delete is never killed halfway.

If git reports success but the path still exists (Git for Windows leaves
junctions and their parent directories in place), the leftover is deleted
with the existing removeHostTree; WSL checkouts stay with the distro.

Nothing creates trash any more: the scheduling queue, rename/restore
helpers and the trash_rename span are gone. The startup sweep stays to
drain entries older releases left behind, and now removes each emptied
trash root so the obligation ends.

* fix(worktrees): let Git delete Windows checkouts with long paths enabled

Removal now always runs Git's own recursive delete, and worktree creation
checks out with core.longpaths on Windows, so a deep checkout Orca created
could fail to delete with "Filename too long" (#6433). The Windows recovery
then finishes the delete but keeps the branch. Pass the same command-scoped
core.longpaths option to `git worktree remove` so Git can delete what it
created.

Also point the CI shard timing entry at the renamed real-git removal suite.

* fix(worktrees): keep an inherited GIT_ASK_YESNO out of the worktree delete

Git for Windows asks $GIT_ASK_YESNO whether to retry when a file stays
locked during a recursive delete. Orca's git env inherits the user's
environment, so an inherited value would run an arbitrary prompt program
in the middle of a removal. Drop it for the removal call only.

* perf(worktrees): run worktree deletes under their own limit, outside git admission

`git worktree remove` now deletes the whole checkout in Git's own process,
which takes 20-35 s on a large tree. It took a general git admission slot at
status tier for that whole time, and that cap is as small as two slots on a
machine with six or fewer cores, so two deletes blocked every status read.

Deletes now skip general admission and queue under their own limit of two
per host instead: two concurrent deletes already saturate one disk, and more
only slow each other down. Leftover cleanup runs inside the same slot.

* fix(worktrees): delete removed checkouts in the background and mark them removing

Since the checkout is deleted by `git worktree remove` in Git's own process,
a large delete takes 20-35 s. Answering the request only after that made web
and mobile (30 s), paired desktop (60/180 s) and the CLI (60 s) report a
failure for a delete that was still going, and mobile silently re-showed the
row.

The request now does everything that can refuse (lock, cleanliness, archive
hook, watcher/terminal gate, terminal stop, shared-link unlink), records the
removal in an in-memory table on the host and answers `removing: true`. The
delete, branch cleanup and metadata purge run after it in the same order as
before, and the watcher/terminal gate stays held until they finish.

- Listings mark rows in the table `removing` for clients that advertise
  `worktree.background-removal.v1` (the desktop renderer, paired desktop and
  web), and leave them out for everyone else (older clients, mobile, the
  CLI), which already dropped the row when the request answered.
- The outcome (removed, with any preserved branch, or the error) rides the
  existing worktrees-changed event as an optional field, sent after the row
  has left the table.
- A repeat delete while Git runs joins it. A create at the same path or with
  the same branch is refused with "Cleanup is pending; try again shortly";
  create's name search skips the path, so generated names move on.
- Nothing is persisted: after a quit or crash Git still lists the checkout
  and it can be deleted again. WSL checkouts still delete inline.
- `orca worktree rm` says the checkout is still being deleted.

* fix(worktrees): keep the existing Deleting card until the host's Git finishes

The host now answers a local worktree delete on acceptance and deletes in the
background. The renderer keeps the existing delete state set until the host
publishes how it ended:

- The delete that asked waits for the outcome on the worktrees-changed event
  (local IPC or the paired runtime's client event), then runs the same
  teardown, preserved-branch toast and card error an inline delete did. If
  that event is lost to a dropped connection, a listing that shows the row
  gone after it was marked removing finishes the wait, and one that shows it
  back without the marker fails it.
- Any other renderer (a reload, a paired desktop, web) sets the same delete
  state from the host's `removing` marker and clears it when the marker goes.
  A failure the host publishes lands on that card's existing error.
- Web advertises `worktree.background-removal.v1` so the host sends it the
  marker; paired desktop does through the Electron capability list.

No new component, style or state: the card reads the delete state it always
did. A host that predates this answers when done without `removing`, and the
renderer takes that as finished, as before.

* test(worktrees): type the removal harness and projection for the node typecheck

* fix(worktrees): don't fail a delete retry with an earlier attempt's buffered failure

A background removal's outcome that reached this renderer with no waiter (another client's
delete, a host-marked card, or one already settled from listings) was buffered for 60 s and
consumed by the next delete of the same workspace, so retrying a failed delete failed at once
with the old error while the host was deleting. Drop the buffered outcome before sending the
request; only an outcome that arrives after it can belong to it.

* fix(worktrees): let only a gap in host events settle a background delete from listings

Git unlists the checkout before the host deletes the branch, cleans the push target and purges
metadata, and the worktree-directory watcher refetches within 250 ms. The renderer read the
missing row as a finished delete, so the waiter resolved without the preserved branch (no
toast) and a failure in those last steps showed as success; the real outcome was then dropped.
The listing fallback exists only for a lost outcome event, so it now applies only after this
host's event stream had a gap: a new subscription or a replay after reconnect.

* perf(worktrees): let a bulk delete start each same-repo checkout delete once the host accepts the last

A bulk delete ran one worktree at a time per repo (#2259, for packed-refs and ref-lock races in
branch cleanup). With Git now deleting each checkout for 20-35 s before the request settles, N
worktrees in one repo took N times that. The renderer now queues same-repo deletes only until
the host accepts each one; a parent still waits for its nested children to finish. The host
serializes the branch cleanup step per repo itself, which also covers removals started by
different clients.

* test(worktrees): pin the host platform in the mocked removal suites so they pass on Windows

Removal now passes -c core.longpaths=true on Windows, so the exact-argv
assertions and command-keyed mocks never matched there (17 failures on a
Windows host). Pin darwin as the add-worktree suites already do, and drive
the one Windows-specific case through the same spy.

* test(worktrees): type the blocked git remove result instead of a broad object

The anti-slop static-analysis gate rejects `object` parameters.

* test(worktrees): clear the changed-code quality gate in the removal suites

Merge the duplicate node:fs import, build the mock child without a cast, read
worktrees:list rows through one typed helper, and give the remaining casts a SAFETY line.

* fix(worktrees): record each background delete durably and finish it after a quit or crash

A quit mid-delete left git to finish the checkout on its own while the branch
delete and metadata purge never ran; a crash left a normal-looking row. Each
accepted local removal now writes a record beside the profile state before git
starts, clears it on success or failure, and the host runs the same delete
again for any record left at startup, re-deriving what remains from git and
disk. An orderly quit stops the checkout delete without waiting for it.

* test(worktrees): type the interrupted-removal assertions for the node typecheck

* fix(worktrees): finish an interrupted delete that already removed the checkout's .git file

Quit stops git worktree remove mid-delete, and Git deletes the checkout's .git
file wherever it falls in directory order. Git then refuses the checkout
("validation failed ... .git does not exist") on every retry, so the startup
finish failed and the row could never be deleted from Orca. A registered
checkout this record owns that has lost its .git file now finishes like an
unregistered one: leftover files, prune, then the branch.

* fix(worktrees): let Git finish an interrupted delete, and never take a different checkout

A quit or crash that stops `git worktree remove` after it deleted the checkout's
.git file left a registered checkout Git refuses to remove. The previous fix
deleted that leftover inside Orca's process, which is the bulk delete this
change exists to avoid (and on Windows the leftover can be most of the
checkout). The startup finish now rewrites the missing .git file from Git's
own admin entry for that path and lets `git worktree remove --force` delete
it. `git worktree repair` is not used: it also re-points every other
registered path, including a checkout another repository now owns there.
Orca deletes the leftover itself only when no admin entry claims the path.

The startup finish forces, so it now leaves the path alone when the checkout
there is not the one recorded: a registered worktree on a different branch or
head, or a `.git` at a path Git already unregistered. The record is dropped and
the card shows why.

The record write before Git starts is now bounded (2 s, logged when exceeded)
so a stalled disk cannot hold the delete, and the outcome is published before
the record's clear reaches disk.

* test(worktrees): compare worktree paths by value and tear down with Windows lock retries

Git prints forward slashes in `git worktree list` on Windows, so the real-Git
removal suites never found a joined path there: positive checks failed and
negative ones passed without proving anything. They now compare Git's parsed
rows by value. Teardown uses the shared retrying removeTree, since Windows can
hold the deleted checkout busy for a moment after Git exits. Adds a
relative-path worktree case for the .git restore (skipped before Git 2.48).

* fix(worktrees): reply to a worktree delete when it has finished, not on a broadcast event

A current client's delete request now waits for the host's background delete and gets its real
result (removed, a preserved branch, or the error) as the reply, the way it did before the delete
moved off the request. A request that arrives while the delete runs joins it and gets the same
result. Every other view keeps reading the host's `removing` marker: the row leaving means the
delete finished, and the row listed again without the marker shows "The delete did not finish.
Try again." on a card that view had marked Deleting. A request whose reply is lost (a timeout or a
dropped connection) settles the same way from a fresh listing instead of reporting a failure.

Clients without the background-removal capability (mobile, the CLI, older desktops) are still
answered on acceptance and have rows under removal left out of their listings.

This removes the outcome on worktreesChanged and everything it needed: the renderer's outcome
waiters, early-outcome buffer and TTL, per-host event-gap generations, the request pre-registration,
and the accept callback bulk delete used. Bulk delete runs same-repo deletes in parallel only on
this machine, whose host serializes branch cleanup per repo; SSH and paired hosts stay serialized.

* test(worktrees): type the pending-removal host id in the background-removal suite

* fix(worktrees): answer a delete request even when a concurrent removal of the same worktree replaced its record

The desktop app's removal and the runtime removal (CLI, paired clients) coalesce separately, so
both can be accepted for one worktree. The second replaced the first's record, and the first
delete then finished without resolving the request waiting on it, leaving the desktop card on
Deleting indefinitely. Each delete now settles the request it was started for.

* fix(worktrees): run same-repo removal archive hooks and teardown one at a time on the host

Local bulk delete now sends same-repo removals in parallel, so their archive hooks, terminal
teardown and preflight ran at once; a hook that writes refs can race the repo's ref locks
(#2259). The host now serializes each local removal up to acceptance per repo, for every
client; Git's checkout delete still runs in parallel under the delete limit.

* fix(runtime): keep waiting worktree deletes out of a host's foreground call slots

worktree.rm now replies only after Git deletes the checkout (up to minutes), so on paired
desktop and web each waiting delete held one of the host's 8 foreground call slots, and a
bulk delete queued listing refreshes and every other foreground call behind it. Deletes now
run in their own lane with the same bound; the 2-slot background lane stays for status polls.

* fix(worktrees): join a same-worktree delete accepted while a removal waited its repo turn

The desktop app and the runtime (CLI, paired clients, web) check for a running delete before
they queue for the repo's acceptance turn. A delete of the same worktree from the other path,
accepted while this one queued, was missed: this request re-ran the archive hook, stopped the
terminals again and started a second `git worktree remove` on the directory Git was deleting.
The queued acceptance now re-checks and joins the running delete.

* fix(worktrees): fence a resumed delete's checkout from startup, and drop rows a listing read before the delete finished

A delete a quit or crash interrupted took its terminal and file-watcher gate only when the resume
job ran, after the first window was shown; session restore could open a shell or watcher inside the
half-deleted checkout first, and on Windows that handle can fail the resumed git delete. Loading the
records now fences each recorded path, and the resumed job takes the fence over in the same tick it
takes its own gate.

A listing that read git's registration before a delete finished, and replied after the removal
record cleared, returned the row unmarked, so other views briefly showed "The delete did not
finish". Listings now capture the pending removals before reading git and leave out a row whose
delete finished successfully since; a row whose delete failed stays listed as before.

* test(worktrees): keep git's auto-maintenance out of the real-git removal suite

CI's Git 2.55 failed the file-pool test in teardown with ENOTEMPTY on the scratch repo's
objects/pack after the test body passed: the 3,000-file commit's detached auto-maintenance was
still writing a pack. The scratch repo now disables auto-maintenance and auto-gc.

* fix(worktrees): one archive-hook approval covers a same-repo bulk delete again

Local same-repo deletes now start together, so each queued its trust prompt with a state snapshot
taken before the first prompt was answered; approving the first still showed the same prompt once
per remaining worktree. The queued check now reads the store when its turn comes.
2026-09-30 16:32:20 -07:00
Brennan Benson 3727100cc9 fix(opencode): report OpenCode 2 status from each pane's own TUI (#23722)
* fix(opencode): report OpenCode 2 status from each pane's TUI

OpenCode 2 serves every pane from one shared server whose env names only
the pane that started it, so every same-folder pane's work showed on that
pane, and the plugin's single aggregate swallowed the Idle of a pane whose
turn ended while another pane was busy.

Install the status plugin a second time as an OpenCode 2 TUI plugin
(plugins/<name>-tui/tui.js, written only when its bytes differ, locally,
in overlays and in the SSH/WSL relay installs). In a TUI process setup()
runs the same engine, fed through the same event translation as the
server, but only for root sessions this pane owns: the route's session
from when it starts (or when the route reaches it while running, hydrated
from the TUI's session status) until it settles, is deleted, or the TUI
exits. The server plugin stands down in any serve process when the TUI
copy is installed beside it; a relay that predates the TUI copy keeps the
old behavior. OpenCode 1 loads no plugin directories and is unchanged.

Delete the session-to-pane binder, registry, client sweep and ingest
reattribution: with every post stamped by its own pane nothing is left
for them to correct.

Known gap: OpenCode 2 `opencode run` in a pane has no TUI, so it reports
no pane status.

* fix(opencode): settle missed OpenCode 2 turn ends and order TUI installs

- The TUI copy now settles an owned root whose run the TUI's session data
  reports ended, with no pending permission or form, when the engine still
  holds it busy. An execution end missed across a service restart or
  reconnect no longer leaves the pane Working until the TUI exits.
- The TUI copy stays idle without ORCA_PANE_KEY. post() cannot report without
  it, so OpenCode 2 TUIs outside Orca no longer run the route poll and engine.
- Installers write the TUI copy before the server plugin file. The server
  decides at load whether to stand down, so a reload between the two writes
  now finds the TUI copy.

* fix(opencode): reconcile OpenCode 2 TUI status only once queued events drain

The TUI's session data applies each event before this plugin's queued handling
reaches it, so settling against it mid-backlog published a false Done before a
fast turn's later steps (Working, Done, Working, Done). Reconcile only when no
event is queued; a mismatch then is a start or end missed across a reconnect,
and both directions are now re-derived (a missed start left the pane on Done).

Also install the TUI copy beside the server plugin in the retired shared hooks
dir, so a TUI or service still loading that dir reports per pane instead of
leaving the service reporting under its starter pane.

* fix(opencode): keep OpenCode 2 TUI panes silent on plugin dispose

A TUI plugin hot reload disposes the plugin while the pane's turn keeps
running. The server path already passes sessionsOutliveDispose so dispose
publishes nothing; the TUI adapter now does the same, or every TUI reload
would still show a false Done. Also refreshes the generated-bytes digest
after rebasing onto the write-if-changed installers.

* fix(opencode): refresh installed OpenCode status plugins at app start

After an Orca upgrade, an OpenCode 2 service that was already running kept
the previous plugin, and with it the old wrong-pane status, until any new
terminal pane rewrote the file. OpenCode 2 reloads a plugin whose file
changes, so Orca now refreshes its existing installs once after the first
window shows: the global config dir, source overlays and the retired shared
dir, TUI copy first. It reuses the per-pane writers, which skip unchanged
files, never creates an install the user did not have, and honours the
status-hook and per-agent switches. An SSH relay does the same for its
canonical install when Orca connects and ships the plugin sources.

The plugin source assembly moves to its own module (re-exported unchanged)
to keep hook-service.ts under the line limit.

* fix(opencode): derive each OpenCode 2 pane's status from its TUI session data

The TUI copy of the status plugin translated OpenCode events into the
server-side engine and then patched the engine's latches back toward the
TUI's own session data: a settle on every tick, a queued-event counter, a
re-assert, synthetic Busy and Idle. Each review found another place where
the two copies disagreed: a missed end left the pane Working, a backlog of
slow posts flickered Done, a missed start showed Done mid-turn, a turn held
only by a subagent never settled, and a request answered while disconnected
pinned Needs input.

The TUI reporter now reads the pane's level straight from the session data
OpenCode keeps current (and re-hydrates on reconnect) on every event and on
a 100 ms tick: for each root session this pane owns, Needs input (an open
permission, else form, anywhere in its family while it runs) outranks
Working (any family member running) outranks Done. It posts one status per
level change through the plugin's existing delivery functions (retry,
dedupe, message-part throttle, ordering), so the wire and Orca's ingest are
unchanged. Ownership is kept in OpenCode's storage.memory, which survives a
plugin hot reload, so a turn that ends during a reload still shows Done;
dispose publishes nothing, and a level the old generation could not deliver
is re-posted by the next. On reconnect it re-syncs blockers for the owned
sessions OpenCode would not re-sync itself, ignoring requests already
answered. Events are handled synchronously, so a slow post can no longer
hold up event processing. The server path, OpenCode 1 and mimo are
unchanged.

* fix(opencode): keep an OpenCode 2 pane on Needs input while its root streams text

Orca treats every OpenCode MessagePart as Working. While a background
subagent waits on a permission or form, the root session can keep streaming
reply text, and the TUI reporter forwarded that text as a MessagePart. The
pane then flipped from Needs input to Working, and nothing restored it until
the level changed, so the user could miss the open request.

Skip reply text while the pane's level is Needs input, as the reporter
already does for queued prompts.

* fix(opencode): leave OpenCode 2 step events to the server path's own change

The shared-server translation re-derived Working from session.step.started.
The pane reporter no longer uses it, so it only changed the server path and
duplicated a separate open change. Drop it; server behaviour matches main.

* fix(opencode): reset an OpenCode 2 pane to idle when its TUI starts

Before this change, a pane could keep a status an earlier process left on
it. The case that matters is an upgrade mid-turn: the old shared-service
plugin posted pane B's turn under pane A's key, then stood down, and pane
A's TUI loaded the new reporter owning nothing, so it never posted and A
showed a wrong Working until its own next turn.

A freshly started reporter that owns no running turn now posts the host's
existing session-start boundary once, which the host shows as connected
idle: no completion, no notification, and an unseen Done it lands on stays
unread. A plugin hot reload keeps its memory and skips it, and where
OpenCode keeps no plugin memory it is never sent.

* fix(opencode): keep OpenCode 1 serve + attach sessions on the attaching pane

OpenCode 1 `opencode serve` in one pane plus `opencode attach` in others
reports every session from the serve process, whose environment names only
the serve pane. Before this branch the session-to-pane binder moved those
posts to the attaching pane; deleting it for OpenCode 2 put them back on the
serve pane.

Restore the binder, registry, correlation and client sweep for OpenCode 1
(and mimo-code, which it also served), gated so OpenCode 2 never uses it:

- The status plugin now sends `opencodeMajor: 2` on every post from a process
  whose loader called setup(), which only OpenCode 2 does. The host skips the
  binder (no rewrite, no kicked round) for any post that carries it. OpenCode 1
  and older plugins send nothing and keep the previous behaviour.
- The binder reads only OpenCode 1's `session` table. OpenCode 2 writes
  `session_v2`, so its sessions never bind; on a database both versions wrote,
  OpenCode 1 sessions are no longer hidden behind the v2 table.

OpenCode-1-only; it goes when OpenCode 1 support is removed.

* fix(opencode): show OpenCode 2 `opencode run` as Working, then Done, on its own pane

OpenCode 2's `opencode run` loads no plugin, and the shared service that
runs its session cannot tell which pane it belongs to, so a pane running it
showed nothing.

The runtime now reports it from the pane's own process lifetime. On the
pane's OSC 133 command start (main already parses 133 for every local PTY;
the start callback was never wired), after the pane tracker's 350 ms settle
it reads the foreground process name through the existing foreground
reader, and only when that name is OpenCode, the foreground command line
from one fresh shell-foreground process-table capture. An `opencode run`
posts Working through the existing terminal-status path into the hook
server's store; the pane's 133;D (or the daemon's background fact) posts
Done, marked interrupted on Ctrl-C. No polling.

Exclusive with plugin reporting: each write carries the command's start
time, and the store drops it once a hook has written the pane since then,
so an OpenCode 1 `run` (in-process plugin) or a server plugin without the
TUI copy owns its command alone and there is never a second Done.

Local macOS and Linux panes only; SSH, WSL and Windows panes stay silent
because their foreground cannot be read on this host.

* fixup! fix(opencode): show OpenCode 2 `opencode run` as Working, then Done, on its own pane

Type the selected descendants as process-table rows; ReturnType of the generic collector widened them to bare identity rows.

* fix(opencode): keep an `opencode run` pane's Done after the command exits

The run's Done was published inside the chunk that carried its OSC 133;D,
before the chunk's command-finished fact. The renderer drops an exited
agent's row on command-finished when the row has not changed since that
fact arrived, so it took the Done as the stale row and dropped it, in the
renderer and in main.

Publish the run's Done after the chunk's side-effect facts are emitted (and
after the daemon's background fact). The renderer then sees command-finished
while the row is still Working, and the Done that follows counts as a change,
so it stays, the same way a hook Done that lands after exit already does.

* fix(opencode): let an `opencode run` revive a pane Orca retired

When an agent Orca launched exits, command completion retires the pane, and
only a hook new-turn event revived it. A later `opencode run` in that pane
posts no hook, so its process-lifetime Working was refused and the pane
stayed silent.

A process-lifetime Working is posted only after a fresh OSC 133;C and the
pane's own foreground argv prove a new OpenCode run, which is at least as
strong as a hook new turn. The store now revives the retired pane on it the
same way: it clears the retirement, drops the launch-token fence, and rebinds
the observation. OSC status and a lone process Done still cannot write a
retired pane, and a closed tab stays closed.

* fix(opencode): report `opencode run` on local Windows panes too

The run producer skipped every Windows pane, although Orca already resolves
a Windows pane's foreground agent from the native process table. On main an
OpenCode 2 `run` showed there (on the service's pane); on this branch it was
silent.

The argv read now has a Windows branch: one fresh native table read
(windows-process-table, no interpreter spawn), the same foreground identity
the Windows resolver already computes, and the command line of the process
that identity names. It runs only after the name read says OpenCode. Local
Windows panes whose shell prints OSC 133 C/D (PowerShell with PSReadLine,
Git Bash with Orca's wrapper) now go Working, then Done; cmd.exe prints no
markers and stays silent. SSH and WSL panes stay silent.

* fix(opencode): re-read an `opencode run` foreground on the pane tracker's ladder

The producer read the pane's foreground once, 350 ms after the command
started. A wrapper, a shim or `sleep 1; opencode run` execs OpenCode later,
so those runs stayed silent. The renderer's pane tracker already re-reads a
command's foreground at 350, then 1200, then 6000 ms for exactly this.

Move those delays to one shared module and use it from both. The producer
re-reads only while the foreground is a non-shell process that is not
OpenCode, and at most on those three rungs; no new polling.

* fix(opencode): end an armed `opencode run` when the next command starts

A new OSC 133;C in a pane whose `opencode run` was still armed dropped the
armed state without a Done, so a run whose 133;D never arrived left the pane
on Working. A new command start proves the previous command ended, so it now
posts that run's Done (not marked interrupted: no exit code is known) before
the new command is inspected.

* fix(opencode): skip the session-start row inside an OpenCode 1 `run` process

OpenCode 1 `run` loads the status plugin in its own process, and the plugin
posts SessionStart when the run's session is created. The host lands that as
an idle session boundary, so a pane the run producer had already shown as
Working blinked idle before the plugin's Busy.

The plugin now skips SessionStart when its own process is a `run`, read from
its argv the same way isOpenCodeRunCommand reads a pane's foreground (first
positional after global options, past the compiled binary's entry path). A
`run` session goes Busy at once, and its first prompt part still revives a
retired pane and resets the turn caches. The TUI and `serve` still post it,
and OpenCode 2's `run` loads no plugin.

* chore(opencode): say where the plugin's opencodeMajor comes from

The field is set when OpenCode's loader calls setup(), which only OpenCode 2
does; the plugin never reads OpenCode's version. Say so at the field. Comment
only; the generated plugin is unchanged.

* test(opencode): type the late-Done pane fixture without bare casts

The new renderer test passed its fixtures with `as never`; use the checked
fixture-tuple cast with a SAFETY note, like the sibling pty-connection tests.

* fix(opencode): keep `opencode run` silent when OpenCode status is turned off

The run producer ignored the status-hooks switch, so a user who turned
status off globally or for OpenCode (#23667) still got a run's Working and
Done. It now checks isAgentStatusHooksEnabledForAgent for the run's agent
(opencode or opencode2), the same predicate the plugin install honours,
before reading argv or posting anything.

* test(opencode): read the late-Done row without an untyped property access

The mock store types agentStatusByPaneKey values as unknown, so reading
`.state` failed tc:web. Assert the row with toMatchObject instead.

* perf(opencode): read a command's foreground only while it could become `opencode run`

Two costs the run producer paid on every command in every local pane:

- The retry ladder re-read the foreground (a process-table capture on the
  local provider) for any non-shell program still running, including other
  agents, editors and dev servers, up to three times. It now re-reads only
  while the foreground is still the shell (the command has not exec'd yet)
  or an unrecognised launcher that may still exec OpenCode (node, bun, bunx,
  npx, npm, pnpm, pnpx, yarn). Any other program is read once.
- With OpenCode status turned off for both opencode and opencode2, it still
  set the timer and read the foreground only to discard the result. It now
  returns before any timer or read. The per-agent check after the name read
  stays for the case where only one of them is off.

* fix(opencode): let a new hook turn end a pane's process-exit completion

A confirmed agent process exit records a pane-wide completion identity that
names only the agent. Hook Dones are matched against it by agent alone, and
only a working title cleared it, so in a pane whose agent paints no working
title (OpenCode) every later hook-reported Done was treated as already
notified: no notification and no unread mark. An unreported exit (a status-off
`opencode run`, or quitting an idle client) was enough to set it.

A fresh hook Working now clears a process-exit identity, since a new turn
cannot be a duplicate of an earlier exit. Hook identities stay, so
same-turn and replay dedupe is unchanged.

* refactor(opencode): share the foreground read schedule as one value

The pane tracker's import of the two shared foreground-read delays took
four lines where its old local constants took three, which put the file
one line over max-lines once merged with main. Export the settle delay
and retry ladder as one object so each consumer imports a single name.
No behaviour change: same 350 ms settle and 1200/6000 ms retries.
2026-09-30 15:37:18 -07:00
Brennan Benson b4b708c2c4 fix(native-chat): a Codex stream retry is one warning row that updates in place (#23684)
* fix(native-chat): a provider's own retry progress is quoted in its retry row

A retry row whose fact carries a detail the provider wrote for a person now
quotes it, the same way a rejected message or failed compaction does, so the
row says how the retry is going. A log detail still stays out of the sentence.

* fix(native-chat): a Codex stream retry is one warning row that updates in place

An error Codex says it will retry used to fall through to the generic frame
row: red, and a new row for every attempt. It now writes one providerRetrying
row per retry run, warning-toned, revised by each attempt with Codex's own
progress sentence. A run is the retry frames of one turn with nothing else the
thread journals between them; every attempt still publishes, so the idle sweep
keeps seeing activity. Errors Codex will not retry are unchanged.

* fix(native-chat): a Codex retry row says it is retrying and keeps the frame behind Details

The quoted retry sentence leads with "is retrying", which holds for any
provider's progress text. The Codex retry row also keeps the whole bounded
frame behind the row's Details, as the generic row did, so Codex's
additionalDetails stays available to diagnose a retry.

* test(native-chat): a Codex retry re-handled after backpressure keeps its one row

Pins the run being opened before the write: a first attempt whose publish is
refused and is handed back must revise the row it already wrote, not open a
second run. Also stops the fixture claiming Codex sends an idle thread status
beside each retry, which the app server does not do.

* perf(native-chat): a Codex frame with no retry run open is not classified

Ending a retry run classified every non-retry frame, and classifying walks the
whole payload: every streaming delta, and every large item/completed, paid a
walk about as costly as parsing the frame. Only a thread with a run open needs
the answer, so the classification now runs only there.

* fix(native-chat): each Codex retry attempt is its own warning row, with what failed on its second line

A stream error Codex says it will retry is written as its own warning row
with a providerRetrying fact, under the same per-frame identity every
Codex frame row gets. The host no longer tracks retry runs or rewrites
one row in place, so there is no run state to open, end or clear, and no
frame has to be classified to end a run. Every attempt publishes, which
keeps renewing the idle clock while Codex retries.

Codex's additionalDetails, which its own UI shows under the progress
message, is kept on the fact as the retry's cause and printed on the
row's second line. Errors Codex will not retry are unchanged.

* fix(native-chat): a transcript draws only the latest row of a provider retry run

The shared structured message projection, which both the desktop and the
mobile transcript read, collapses a run of retry rows into its latest
row. A run is retry rows from the same agent with no other drawn row
between them; a row that draws nothing, or a queued send drawn after the
conversation, does not split it. The earlier attempts stay in the
journal.

* test(native-chat): the retry-run render test uses the message list's current props

* fix(native-chat): a Codex frame row is named for its connection, so a later one never revises it

Frame rows were named provider-frame:codex:<n> from a counter that starts
over with every connection, so the first frame row after a reconnect in the
same session revised an earlier connection's row in place, at its old spot.
Each connection's frame rows now carry the acquisition generation, minted
once before the translator is built: provider-frame:codex:<generation>:<n>.
Rows already written keep their identities.

* fix(native-chat): agents retrying at once each keep one row, read from the row's own agent

* test(native-chat): a reconnect's Codex rows are named for the acquisition that received them

* fix(native-chat): a retry run is one agent's, so another agent's row never splits it

Each agent's rows are drawn apart: the session's own rows are the conversation,
and a subagent's rows open in that subagent's section. Splitting a run on any
other row in the flat list left two adjacent retry rows on screen whenever
another agent wrote between two attempts: a subagent finishing a command while
the session reconnected, or the session working while a subagent reconnected.

A run is now per agent: an agent's retry rows with none of its own other rows
between them, drawn as its latest.

* test(native-chat): a subagent's retry run is checked in its own section, and the run rule's words say same agent

* test(mobile): the retry-rows test typechecks, so the test ratchet keeps checking it
2026-09-30 15:05:22 -07:00
Brennan BensonandClaude 24edf0f64b fix(agent-status): a turn a crash cut off reads Interrupted, an unproven end Couldn't confirm (#23467)
* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* fix(native-chat): name a chat write by its target, not the owner generation

A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.

Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.

Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.

* fix(native-chat): every journal append reaches the chats that are open

A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.

A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.

* test(native-chat): an epoch replacement reaches the open chat

* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

* perf(native-chat): a publish behind a delivered commit reads nothing

Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.

* test(native-chat): state why the teardown test's fake journal is safe to cast

* docs(native-chat): say mutation admission checks only the writer lease

* docs(native-chat): drop the send rebase from comments that still described it

* fix(native-chat): a message is accepted, then delivered

A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".

A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.

Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.

A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.

Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.

* fix(native-chat): settle queued messages only for the child that ended

A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.

A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.

The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.

* fix(native-chat): an adoption that fails to import keeps the conversation open

The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.

* perf(native-chat): the recovering open reads the journal once

Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.

* fix(native-chat): an attach that fails after indexing its child leaves no child behind

A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.

* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer

The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.

A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.

* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down

The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.

* fix(native-chat): a message rejected while its chat was closed reads as not sent

A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.

* test(orchestration): name why the readiness settlement fakes are cast

* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent

* docs(native-chat): drop the fence from the admission the send effects run behind

* docs(native-chat): give the fence move on release the reason that still holds

* docs(native-chat): stop citing a write fence check in launch and mailbox comments

Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.

* refactor(native-chat): the provider child is its own record

A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.

- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
  patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
  own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
  the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
  dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
  is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
  the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.

* fix(native-chat): the delivery loop alone settles a message its start or child failed

A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.

- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
  reads how it ended: a Stop continues; anything else writes one failure row and rejects every
  queued message with the same words, then stops. A child still starting whose start the adapter
  says did not land fails the same way. The exit, eviction and the settlement retry only settle
  the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
  nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
  failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
  closed, with or without a child, and a start the loop already has in flight is waited for so the
  child it produces is stopped rather than left behind.

* refactor(native-chat): a stopped child ends on the one reading of its stop

The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.

* feat(native-chat): the host says it accepts a send before any agent has it

The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.

* refactor(native-chat): an attach never opens a journal of its own

The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.

* fix(native-chat): a moved fence resends nothing on a host that accepts first

The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.

The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.

* refactor(native-chat): a child's end says whether the user or the host stopped it

The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.

* fix(native-chat): a chat whose only work is a queued message is not offered for resume

A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.

* test(native-chat): type the queued-message fixtures in the resume-offer tests

* fix(native-chat): a start that dies while a message waits on it is that message's failed start

Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.

* fix(native-chat): a request that failed reads as failed

A structured chat whose only message the agent's start refused read as a
green finish, and a cancelled structured turn did too: the host published a
verdict only for turn records, and structured rows carried no `interrupted`.

The host projection now reads the session's latest request: its turn's
outcome, or `failure` for a send the agent or its start refused. A send
that was withdrawn, or left undelivered by a restart or a close, fails
nobody and makes nothing listable. The ingest publishes `interrupted` as the
hook lanes do, and every reader decodes the verdict through one accessor, so
a failure reads Failed on the dot, the rollups, history and `worktree ps`,
behaves like a cancellation in every clean-finish policy, and notifies as
"failed".

* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now

* test(native-chat): a verdict change republishes the mobile status projection

* refactor(native-chat): the store's retention trigger keeps its flag compare

A verdict change always moves the completion clock the same check already
reads, so a second verdict compare there caught nothing new.

* test(native-chat): a user message the provider journaled keeps its session listed

* test(native-chat): pin what a failed start settles, and what a resume offer names

A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.

* test(native-chat): the failed-start pins fail on what the message became, not on a timeout

* fix(native-chat): a late provider-session update keeps a failed recovery record failed

A provider-session heartbeat that rewrites a completed recovery record kept
its interrupted flag but dropped the outcome it was copied with, so a live
failed checkpoint read as a clean finish until the next status write.

* test(orchestration): the preamble's host stub is typed, not cast

The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.

* test(native-chat): the terminal-bell check asserts the renamed verdict field

The bell notification test still checked for agentInterrupted, which no
longer exists, so it could not catch a verdict leaking into a bell dispatch.

* fix(native-chat): a failed turn ranks like a completion for attention

Attention readers (completion time, Smart Sort, sticky retention, Cmd+J
Recent) now demote only a turn the user stopped. A failure is news the
user has not seen, so it keeps its completion time, ranks in the Done
class, stays retained after its pane goes away, and a retained failure
reads failed in the worktree rollup instead of done. Clean-finish
policy (hibernation, pane ownership, the value moment) still treats a
failure like a stop.

The retention trigger compares verdicts again: success -> failure no
longer moves the completion clock.

* fix(native-chat): a failed main agent reads failed while its subagents still work

The verdict is now read from the main agent's own state, not the folded
row: a main agent that is done and failed has a verdict even while its
subagents keep the row working. Without mainAgent (history, worktree ps,
older hosts) the old combined-done rule stands.

Display marks the verdict through agentVerdictDisplayMark: a failure
outranks every combined state on the agent's dot, label, tab badge,
dashboard and activity rows; a stop marks only a done row, so a
successful or stopped main agent with live subagents still reads
working. Subagent rows keep their own state. The worktree card, terminal
tab and Cmd+J rollups share one pane fold and rank a pending question,
then failed, then working, monitoring, interrupted and done.

worktree ps publishes the main agent's outcome on a working row, and the
mobile mirror reads it. The store's change check, the paired-client
mirror's equality and its epoch now see a verdict change on a working
row, which otherwise moves no state or clock and left the worktree card
reading working. Clean-finish policy is unchanged: a working row is never
hibernated and has no completion time.

* docs(native-chat): the worktree ps outcome comment no longer claims old hosts send it

The field is new: an old host sends no outcome at all, so a reader falls
back to interrupted. The removed clause said old hosts send it on done
rows, which never shipped.

* docs(native-chat): the status-store listing rule names provider-journaled user messages

* fix(native-chat): a refused send notifies failed through the completion feed

The host's completion feed followed only the newest turn, so a send the
agent or its start refused, which creates no turn, read Failed on its row
but sent no notification. The feed now follows the session's latest
request, read from the projection the status feed already makes for the
commit: a turn keeps its id, a refused send is named by its journal item
key. It announces only while the session is idle, as the row reports a
verdict, so queued sends refused one commit at a time notify once, and a
withdrawn send falls back to a request already announced.

* fix(native-chat): every copy of a row carries the main agent's own status

History entries, sleep records and `worktree ps` rows carried a flattened
top-level `outcome`, copied under different gates and without the main agent's
clock. They now carry `mainAgent` (state, outcome, stateStartedAt), the type
the live row already persists and sends, and every copy site takes it with
`interrupted` through one function, `agentVerdictFields`.

- The accessor reads `mainAgent` then the legacy flag; the mobile mirror
  matches it line for line.
- Sleep records admit `mainAgent` with `normalizeMainAgentStatusField`, so a
  malformed value drops the field, never the record.
- Mobile dates a main agent that failed under live subagents by its own clock,
  as desktop does, and its row equality compares `mainAgent`.
- The activity feed reads a history entry's own `mainAgent` instead of
  rebuilding one; the sync key and history equality compare it.

* test(native-chat): pin the worktree ps verdict across host and phone versions

Pairs the real v1.4.212 host and phone row reader with this build: an old phone
reads a new host's rows by `interrupted`, a new phone reads an old host's rows
(no `mainAgent`) the same way, and a new phone reads a failure under live
subagents as Failed, dated by `mainAgent.stateStartedAt`. The release checkout
now carries the phone's self-contained row reader, and the lane runs when the
`worktree ps` row producers change.

* test(mobile): name the parity table's row for its role

* fix(native-chat): a request that settles while the user is asked something notifies once

The completion edge waited for an idle session, and a pending prompt (including a
subagent's approval) is not idle. Structured chat has no other attention producer,
so a main turn that finished while a subagent waited on the user sent nothing
until the prompt was answered.

The edge now waits only on owed work (a running turn or an unanswered send), which
the projection reports even beneath a pending prompt. A request that settles with
a prompt pending announces once; the renderer words it "needs input" from the
host status mirror's `attention`, and answering the prompt keeps the same request
identity, so it does not announce again. The wire shape is unchanged.

* fix(native-chat): the completion says when the user is being asked

A request that settles while a prompt waits on the user was worded "needs input"
from the renderer's status-feed mirror. Remote clients receive the status and
completion streams over separate sockets, so they can arrive in either order and
the wording could be wrong both ways.

The host already knows at emit time, so the completion now carries an optional
`awaitingUser: true` in that case and omits it otherwise. The renderer words the
notification from that field alone and no longer reads the status mirror. Old
clients ignore the field and word by outcome; old hosts never send it.

* fix(worktree-status): a departed agent's failure yields to live work on the worktree card

A retained failed agent has no expiry, so ranking it with a live failure pinned the card to Failed over other panes' live work. It now ranks below working, monitoring and permission, and above every finished outcome.

* docs(agent-status): a departed agent's failure ranks below live work on the worktree card

* fix(native-chat): a view never restarts a chat whose last start failed

A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.

* test(native-chat): start the child the loop waits on with an attach, not a second view

A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.

* fix(native-chat): settle a gone generation's turn wherever a conversation opens

A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.

* test(native-chat): prove the next child's start settles the turn an earlier child left

The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.

* test(native-chat): count a failed start's rows by row, not by text

Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.

* test(cross-version): load the phone row readers without mobile's toolchain

Vite transforms a file against its nearest tsconfig, and mobile/tsconfig.json
extends expo/tsconfig.base.json, which the root-only cross-version lane never
installs. The worktree ps verdict suite imported the current phone row reader
from mobile/ directly, so CI failed with TSConfckParseError before any test ran.

The harness now imports a copy of the working-tree reader placed under the
checkout cache, where the root tsconfig applies, as it already does for the
release checkout's copy. Both readers are still the real files.

* test(cross-version): keep the checkout path-guard message and justify the copy import's cast

* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget

* test(native-chat): pin the open's and the send's start and row counts, however the view binds

Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.

* fix(agent-status): a turn a crash cut off reads Interrupted, an unproven end Couldn't confirm

When the provider gave no verdict, the structured status projection now derives
one from the newest turn's lifecycle: interrupted -> interruption, unverifiable ->
unconfirmed. Nothing new is journaled, the completion feed stays provider-only, and
the legacy interrupted flag stays a user stop only. Every verdict reader handles
both arms explicitly.

* fix(native-chat): settle a gone generation's turn at every open but an acquisition's

The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.

* fix(native-chat): a folded turn a crash cut off reads Interrupted after N

The settled-turn timing now carries the turn's verdict, derived by the same
agentTurnVerdict the status row uses. A turn that ended interrupted with no
provider verdict heads its fold 'Interrupted after N'; a user's stop keeps
'Worked for N'.

* test(native-chat): hold the create's start open until the views bind

The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.

* fix(native-chat): the user's close of a chat records the turn it cuts short as their cancellation

The expected-close settle writes outcome cancellation when the user aimed the
stop at this chat: a Stop while the agent starts, the chat's tab closed (the
agentSession.close RPC, or session.tabs.close with a user reason), or /clear.
A quit, an idle eviction, a worktree teardown or an orchestration stop leaves the
turn with no verdict, so it still reads Interrupted.

* test(native-chat): a Claude turn a newer send superseded reads Interrupted

The supersede fires for any send Orca dispatched, the user's or another agent's,
and nothing at that site records the sender, so the turn keeps no verdict and
folds as Interrupted after N.

* fix(native-chat): the user's close records cancellation on the turn the provider settled on its way out

The Codex and Claude adapters settle their open turn as interrupted, with no
verdict, while the host stops them. The expected-close settle then found no
running turn, so a user's close of a mid-turn chat read Interrupted. The stop
now reads the running turns before it reaches the provider and records the
user's cancellation on each one it cut short, unless the provider gave a
verdict of its own. The close-verdict test's adapter now settles its turn on
close the way the real adapters do.

* test(native-chat): update the close and settled-turn expectations for the host-observed verdict

agentSession.close now passes the user's word to the host, and a settled
interrupted turn with no provider verdict carries `interruption`. Also merge a
duplicate import the code-quality gate rejects.

* refactor(native-chat): drop the composer's second error formatter

After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.

* test(native-chat): pin the reason on a message rejected while its chat was closed

The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.

* fix(native-chat): the user's close cancels a turn whose start landed as the provider stopped

The close read which turns were running before the stop. A turn whose start was
still in flight (a send echo not yet journaled) was absent from that read, so the
provider's verdict-less settle on the way out left it Interrupted. The close now
reads which turns were already over instead, and records the user's cancellation
on every other turn the stop left running or interrupted with no verdict. A turn
cut off earlier, or finished during the stop, keeps its end.

* fix(native-chat): a send the provider never received after a restart has no verdict

Restart reconciliation rejects a crash-stranded send that is absent from a
trustworthy provider history with reason 'not_delivered'. Nobody failed that
send, but the verdict allowlist did not name it, so after a crash the chat
read Failed, was listed, and could notify "failed". Give the reason a shared
constant (persisted value unchanged), add it to the no-verdict set, and treat
it as an internal marker so the Retry row no longer shows the raw string.

* refactor(native-chat): the adapter settles the turn a stop cuts with the stop's typed cause

The host hands its stop's cause to the adapter's close. Each adapter settles its own
open turn on 'ended' through one mapping, turnVerdictForChildEnd, and the host's
dead-generation fallback uses the same mapping for any turn no adapter settled. A
user's close or stop of this chat is their cancellation; a quit, eviction, teardown
or an exit the adapter saw first is news.

Deletes the snapshot-and-diff reconstruction (endedTurnItemIds,
userStoppedTurnRevisions, settleUserStoppedTurns) and the requestedByUser flag.
host.close now takes a required cause.

* test(native-chat): a user's close drops the chat's status row like an eviction

* test(native-chat): expect the eviction cause on the host closes of idle release, worker stop and worker discard

The stop's typed cause now travels into host.close and the adapter's close, so these
three non-user closes assert the 'evict' they pass.

* refactor(native-chat): every stop names its cause, so none defaults to the user's cancellation

stopStructuredAgentSessionAgentUnderSerialize defaulted its ending to 'user-stop', which now
settles the cut turn as the user's cancellation. Every caller already passes a cause; the
parameter is now required, and a type-level test fails to compile if the default returns.

* test(native-chat): pin who a chat's session.tabs.close speaks for, older clients' reasonless close included

The mapping lived inline in a type-unchecked file, and only the explicit user reason had a test: an older client's reasonless close, or a lifecycle echo read as the user's, stayed green. It is now one exhaustive, type-checked function with a case per reason.

* style(mobile): draw the unconfirmed dot in the theme's status amber, not an inline hex

* docs(agent-status): the main agent's outcome also carries the host-observed end, interruption or unconfirmed

* fix(native-chat): a chat the user closed while its agent started is not a failed start

A still-starting child the user's close cut counted as a failed start, since only 'user-stop' was excluded: a start-failure row, and queued messages rejected as a provider failure. Whether an ending fails its start is now one exhaustive switch, shared by the failed-start read and the delivery loop's handover: a user's stop or close never does; an exit, a failed attach and the host's own stops still do.

* fix(native-chat): a chat the user closed closes its queued messages, and starts no agent for them

After 876b6989f1 a user's close of a still-starting chat went on like a Stop, so when the close did not complete the delivery loop started a new agent for the message queued behind it. A child's end now has three dispositions, not a failed-start boolean: a user's Stop lets the queue go on, the user's close closes what was queued before it, and any other end fails it. The close is the one a completed close does (the provider-closed rejection, no verdict), applied at the top of each delivery step and ordered against the close so a later send still goes on.

* fix(activity): a crash-cut turn draws the interrupted glyph; only a user's Stop keeps the done check

The Activity page drew every Interrupted row with the done check, which #2569 chose for a user's Stop. With a crash now reading Interrupted, that put a green check on a turn nobody asked to stop. The row's glyph is now an exhaustive switch over the verdict: a cancellation keeps the done check, and an interruption draws the existing interrupted dot. An unconfirmed end already drew its own glyph.

* fix(activity): the Interrupted group header draws the done check only when every row is a user's Stop

A user's Stop and a crash share the Interrupted group, and its header took its newest row's glyph, so a Stop newer than a crash put a green check over the crash. The header is now folded over the group's rows: the done check only when every row draws it, the interrupted dot otherwise.

* fix(native-chat): a failed close of what the user closed starts no agent for it

Rejecting the messages a user's close left queued swallowed a journal write failure, so the delivery step went on to start an agent for a message in a chat the user closed. The rejection now reports whether it landed, and a step whose rejection failed stops instead; the next wake re-derives and retries it. Also pins that the ordering against the close holds only within its epoch, since a later epoch's sequences restart.

* test(native-chat): the idle sweep's stop is an eviction, so its close carries that cause

Main's idle sweep now stops an idle agent through the conversation lifetime, which this branch gives the 'evict' cause; its expectations name it.

* fix(native-chat): a retried stop keeps the cause of the stop it finishes

A user's Stop or close whose wind-down failed after the child was proven gone was finished by the idle sweep as an eviction, so the turn it cut read Interrupted. The owed wind-down now carries its stop's cause, and a retry with no child settles with it.

* test(native-chat): the idle sweep's close of a retrying Claude chat carries the eviction cause

Main's new test expected the adapter close with the session id alone; every stop now names its cause, and the idle sweep's is 'evict'.

* fix(status): a user's Stop marks done on the tab and sidebar; red Interrupted is only a turn cut short by something else

The tab, the worktree card and the sidebar rows drew a Stop with the same red dot as a crash. The
verdict mark now maps a cancellation to done, still saying "Interrupted by user" in the row text,
and the mobile mirror follows. The Activity page keeps grouping a Stop under Interrupted with the
done check, as before.

* test(cross-version): a new phone reads a user's Stop as done; an old phone still draws it interrupted

* fix(native-chat): a user's Stop inside a live Claude chat reads as their cancellation

Stopping a running Claude turn interrupts it and keeps the session, so the turn's end comes from
the CLI's result frame. Claude CLIs before 2.1.91 send that frame with no terminal_reason, and later
ones may still omit it, so the user's own Stop was recorded as a failure with an error row.

Orca now records the stop on the open turn when it sends the interrupt. An error result for that
turn reads as the user's cancellation whatever reason the CLI gives. The stop belongs to that one
turn, so it cannot reach the next, and it is withdrawn when the CLI refuses the interrupt.

* docs(agent-status): a user's stop marks done; name the tab close cause by its type

The reference still said a stop marks a row interrupted and ranks between live work and an
unconfirmed end. A cancellation now marks done, and only a turn cut short by something else ranks
as interrupted. The runtime's tab close restated the close cause's union; it now uses the type.

* test(native-chat): a proven crash reads as an interruption on the status feed and in the chat

A crash the relaunch proves now settles its turn interrupted, and the status feed works the verdict
out from that record, so the restart test expects interruption for a proven crash and unconfirmed
for one it cannot prove, never a cancellation. A chat read before the proof lands reports
unconfirmed, then interruption and a folded "Interrupted after 27s" once the proof revises it.

* fix(status): a user's Stop reads Interrupted, and a turn anything else cut short reads Failed

The verdict mark now maps a cancellation, the user's own Stop, to interrupted, and an interruption,
a turn cut short by a crash or a killed agent, to failed, the same as a failure, which outranks live
subagent work. An unconfirmed end is unchanged. This applies to every agent, in a terminal or a chat,
on the tab, the sidebar rows and worktree card, the dashboard row, Cmd+J and the phone. A Stop is
not news, so the rollups rank it below an unconfirmed end, and notifications word an interruption
"failed". Recording is unchanged.

* fix(activity): group a user's Stop under Interrupted and a crash with failures

A user's Stop draws the interrupted glyph and sits alone in Interrupted, and a turn anything else cut
short sits in Failed, titled "Agent failed". Every row in a status group now draws the group's own
glyph, so the header is the group's status and the rule that folded a Stop's done check into the
header is gone. Interrupted ranks below an unconfirmed end, as in the sidebar.

* fix(status): draw a user's Stop in the muted tone, not the fault red

The interrupted dot, which now means only a user's Stop, draws in the muted foreground token on the
agent rows, the sidebar card and the phone. Red stays for a failure or a turn cut short by anything
else, and green for a finish.

* fix(native-chat): fold a stopped turn as "Interrupted after N" and a failed or crash-cut one as "Failed after N"

The settled turn header now follows the verdict mark: a user's Stop reads "Interrupted after N", and
a failure or a turn anything else cut short reads "Failed after N", under the new key
components.native-chat.status.failedAfter in all six catalogs and the boot catalog. Desktop and
phone share the one description, so they agree.

* docs(agent-status): describe the Interrupted and Failed marks

The reference and the phone's turn bar still described a user's stop as done and a crash as
interrupted. A fault now reads failed, a user's stop reads interrupted in the muted tone, and the
rollups rank an unconfirmed end above a stop.

* test(status): a crash the relaunch recovers marks failed

The recovery test still expected a recovered interruption to mark interrupted; it now marks failed,
as a failure does. Formatting only elsewhere.

* fix(native-chat): record a turn a newer request replaced as superseded, and show it Interrupted

A Claude turn that a newer send replaced before its result arrived was recorded as interrupted with
no verdict, which reads as a turn cut short by something else, now "Failed". It is now recorded with
its own outcome, `superseded`, where the replacement is detected. That outcome names no sender, so a
dispatch from another agent is never recorded as the user's Stop, and it sets no legacy flag.

Every reader handles it in an exhaustive switch: it draws the muted Interrupted mark with the plain
text "Interrupted", folds as "Interrupted after N", and attention demotes it with a Stop, through
the renamed agentTurnEndedOnRequest. Older builds read an arm they do not know as no verdict, which
is what this turn carried before, so their rows keep reading done; the cross-version suites pin an
older desktop's journal and status readers and an older phone.

* refactor(status): name the attention predicate for a turn ended on purpose

agentTurnEndedOnRequest becomes agentTurnEndedOnPurpose: a user's Stop or a newer request's
replacement, never a fault. The Claude turn-end comment no longer says a replaced turn carries no
verdict.

* test(native-chat): a turn cut off by a restart or by quitting Orca reads Failed after N

On main a restart-cut turn shows the done tick. Pin the chat's turn bar and
the tab's mark for both cuts, through the recovery settlement and the quit's
child-end mapping, and pin the quit's turn bar through the host's own quit.

* style(native-chat): format the superseded turn-bar expectations

* fix(claude): a Stop that names no turn is the user's stop of the open turn

The chat's Stop button names no turn. Claude's conversation Stop recorded the
user's stop only for a named turn, so an older CLI's error result after that
Stop read Failed. It now records it on the open turn through the same intent,
dropped when Claude refuses the interrupt and never carried to the next turn.

* fix(native-chat): keep the attach context's publishStatus required

The lifetime context type makes publishStatus optional, so the attach context
that spreads it no longer satisfied its own type once its duplicate
publishStatus went. The host's lifetime context is now inferred, and checked
with satisfies, so the spread carries the member it always sets.

* fix(activity): rank a user's Stop below live work in the status grouping

The Activity page's status grouping put the Interrupted group (a user's Stop, or a
turn a newer request replaced) above Working and Monitoring, so a Stop still sorted
like news there while the sidebar, worktree card and Cmd+J rank it below live work.
It now follows live work and stays above Done; Failed and Couldn't confirm keep
their places above live work.

* docs(agent-status): say which turn outcomes the journal records and which are derived

The journal now records superseded as well as the provider's verdict and a stop;
interruption and unconfirmed are derived on read. The resume row no longer claims
interrupted renders red.

* test(native-chat): the idle sweep's held-send rest closes with the evict cause, like its siblings

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-30 15:02:44 -07:00
Brennan Benson 5cda0f4508 refactor(native-chat): keep agent-session records in the chat journal database (#24006)
* fix(native-chat): report a failed startup chat reconcile instead of failing app startup

At startup the chat host re-checks every saved chat's lease and writes the
result to agent-sessions.json. If that write failed (the file lock gave up,
the file could not be written, or the file was written by a newer Orca and
is read-only here), reconcileRestartLeases rejected, the startup IPC call
rejected, and the renderer fell into its degraded "Session restore failed.
Changes won't be saved until restart" mode.

The reconcile is bookkeeping: a lease left unreconciled grants no writer,
and every attach, send and read of a chat reconciles its own lease again.
So the startup reconcile now reports its failure through a new optional
host dependency, onStartupReconcileFailure, and resolves. The runtime
routes it to its onError sink under the scope
structured-agent-session-startup-reconcile, or logs it when no sink is
installed (the desktop installs none).

* fix(native-chat): read restored chats without waiting on lease bookkeeping

With native chat on and a chat tab open at quit, the renderer's startup
also awaits the chat tab restore (session.tabs.listAll). That restore
re-ran the lease reconcile before reading each chat and rethrew its store
failure, then recorded each restored tab as visible through a store
transaction that throws on a held lock or a read-only store. Either one
failed the restore, so startup still fell into "Session restore failed".

Reading a chat grants no writer, so the reconcile startup and the restore
run is now a reader's: createReaderReconcile never throws, answers whether
every lease is settled (recovery is resolved only then; the journal opens
either way), and reports each distinct failure once until a reconcile
settles. Attach and agent start keep the strict reconcile. The restore's
tab republish logs a failed visibility write and still publishes the tab,
since a client drops every unpublished chat tab; user-driven publishes
still refuse.

The host dependency is renamed onLeaseReconcileFailure (scope
structured-agent-session-lease-reconcile), since it now also reports for
reads.

* fix(native-chat): keep every record-store write off the startup chat read path

Round-2 review found two more writes on the startup chat restore that
could still fail it and put the app into "Session restore failed":
republishing a /clear replacement recorded its tab visibility strictly,
and resolving a chat's recovery rethrew its store error. The restore
also paid one lock wait per tab and per batch of chats while the lock
stayed held.

The restore now derives tabs from state it already holds:
- publishStructuredAgentSessionTab splits into the strict write and
  projectStructuredAgentSessionTab, which only updates the runtime's
  snapshot. The restore and /clear replacements only project: a saved
  tab index already lists every restored chat, and a /clear moves the
  tab in the same write that commits it. visibilityWriteMayFail is gone.
- Chats a legacy profile restores that the index does not list are
  recorded in one best-effort transaction (store.showSessionTabs), so a
  failure leaves the index absent to seed again rather than partial.
- The read restore's recovery resolution is caught and reported through
  onLeaseReconcileFailure, deduplicated with the reconcile's reports.
- Once lease bookkeeping fails in a restore pass, the rest of that pass
  skips it, so a held lock costs one wait for the startup reconcile and
  one for the restore, however many chats are open.

User actions (create, reveal, attach, send, the /clear commit) keep
their strict writes.

* test: open, seed and read the agent-session record store through one harness

Tests that open the durable agent-session record store, seed it, or read
back what it persisted now go through agent-session-record-store-test-harness.ts
instead of calling AgentSessionRecordStore.open or touching agent-sessions.json
themselves. A later change that moves the store into the chat database then
changes the harness instead of every test. No production code changes.

Tests whose subject is the JSON file itself (its .bak recovery, salvage,
schema versions, permissions, and what older builds read back) keep reading
and writing the file directly; the storage move rewrites or deletes them.

* fix(native-chat): start each restore pass from one lease check and stop its bookkeeping at the first failure

The restore now runs one reader lease check for the pass and lets each chat
re-check and resolve recovery only while the pass is still settled. The
first refusal or failed write clears it for the rest of the pass, and every
chat is still opened for reading. With another process holding the lock,
startup waits on it once in prepare and once in the restore, however many
chats are open; a legacy profile waits once more for its tab-index seed.

* docs(native-chat): correct restore comments and a test name to match the final design

* test: address the record-store harness by the host's state directory

The harness took the store's own folder, so each caller picked one
(join(root, 'store'), or 'agent-sessions' where a test read the store the
runtime owns). A later change that moves the store into the state
directory's journal database could not tell those apart, and would have
had to edit every caller again.

Every harness function now takes the state directory, the one the test's
journal database and recovery capsule already live in, and keeps the
store in the same subfolder the runtime uses. Callers pass that directory;
store-only tests pass their temp directory unchanged. Format tests that
share a directory with harness calls take the file path from
testAgentSessionStoreFilePath.

The folder name moves from a private constant in the runtime to
AGENT_SESSION_STORE_DIR_NAME beside the store's file name, so the harness
shares it without importing the runtime. Its value and every path built
from it are unchanged.

* refactor(native-chat): keep agent-session records in the chat journal database

The record store's records, operation ledger, retired claim keys and chat tab
index become tables in agent-session-journal.db (user_version 4). The version-4
migration copies agent-sessions.json in its own transaction and never writes,
renames or deletes that file or its .bak. Each store write is one journal
transaction over exactly the rows it changed, checked with the load rules; the
file lock, the external-change refresh and its hash, the .bak rotation, salvage
and the hot-path recovery fence are gone from the store.

* wip: importer tests

* test(native-chat): cover the records migration, the import, row writes and read-only records

* docs(native-chat): retire comments that describe the records file as the live store

* test(native-chat): drop the record-store harness's leftover file path and type the import fixture

* test(native-chat): let the host harness cleanup wait out a recovery-offer read's lock

* fix(native-chat): let Stop reach the agent when its ledger row cannot be written

Stop's operation-ledger row now shares the database with the chat history, so
damage, a full disk or a stranded transaction on that write refused the Stop
before the interrupt. A cancel plan now takes its decision from the committed
ledger in memory, runs without settling, and warns that the row was skipped.
Other mutations answer proven damage with the typed "Unable to load this chat."
refusal instead of the raw SQLite error.

* fix(native-chat): answer whether a profile holds chats from the database's rows

Every host install creates agent-session-journal.db, chats or not, and the
version probe created it too, so its mere existence made every profile that
ever installed the host wait on host install and reconcile at startup. The
check now opens the database read-only and looks for a record or tab row,
lets the records file answer while its import is still owed, and counts an
unreadable database as present. The version probe no longer creates the file.

* fix(native-chat): open a chat from history when its tab index cannot be written

Over records a newer Orca wrote, every write is refused, so opening a closed
chat from Agent Session History failed on the tab-visibility write and the chat
read as unreachable. Like closing a tab, opening one now reports a failed
restore-index write and still publishes the tab.

* fix(native-chat): keep the records import owed when the backup read fails transiently

A torn records file whose .bak could not be read (EACCES, EIO) was reported as
unusable, so the migration completed with nothing copied and never retried.
A non-ENOENT read failure of either copy now carries its cause, which the
importer classifies as a read that can clear.

* test(native-chat): pin that an unreadable records file never falls back to its backup

* fix(native-chat): restore imported chats' tabs when the records file had no tab index

A chat created while the import was owed recorded a tab index holding only
itself. When the file it later imported had no index, that index still read
as recorded, so the imported chats' tabs never came back. The import now
clears the recorded marker in that case, and restore falls back to the
profile's tabs.

* refactor(native-chat): drop the unused in-transaction store write

Nothing called it, and it bypassed the write queue and the read-only refusal.

* docs(native-chat): say that an unusable records file is left untouched but never re-imported

* refactor(native-chat): keep the provider handle chain check as main has it

The chain-validation refactor has no measured need in this change.

* docs(native-chat): retire lease-renewer comments that describe the records file as the live store

* fix(native-chat): keep a throwing failure sink from failing the startup chat read

The lease bookkeeping failure reporter called the host's failure sink
directly, so a sink that threw turned a reported, recoverable store failure
back into a rejected startup reconcile or read restore. The reporter now
catches a sink throw and logs both the original failure and the sink error
with console.warn.

* test(native-chat): wait for a replaced host's restart-offer writes before cleanup

A restart test replaces the host without tearing the old one down, so the old
host's fire-and-forget restart-offer withdrawal could still hold the recovery
capsule's lock directory when cleanup removed the test directory (ENOTEMPTY).
The harness now hands hosts a capsule that tracks running operations and waits
for them before removing the directory, replacing the rm retries.

* docs(native-chat): retire the abandon helper's note that the store re-creates its directory

* fix(native-chat): restore a chat opened while the import was owed beside the profile's chats

When the imported records file had no tab index, restore fell back to the
profile's saved tabs, which never list a Claude chat, and the seed then
rewrote the tab table without the chat opened while the import was owed.
The tab rows that chat left are now loaded as unrecorded, restore takes them
together with the profile's chats, and the seed keeps their tab ids.

* test(native-chat): pin that a create whose tab index write fails still opens the chat

* docs(native-chat): say why restore puts chats opened while the import was owed first

* test(native-chat): replace a ledger row rather than change it in place in the Send-now rerun test

The record store freezes published rows in tests, so setting a row's outcome
in place threw; the test now swaps in a changed copy, as its sibling cases do.
2026-09-30 14:45:32 -07:00
Brennan Benson cfa43e7eab fix(codex): opening a terminal no longer strips Codex hooks from the real ~/.codex (#23552)
* fix(codex): a real-home restore leaves a file alone once someone else changed it

Orca writes ~/.codex/hooks.json (and a trust rebase writes config.toml), then
runs a Codex trust session for up to 10 s, then restores the original bytes if
the session fails. The restore wrote unconditionally, so a save that landed
during the session, from the user or another Orca, was silently reverted.

Each restore now compares first: it writes the original back only while the
file still holds the generation Orca's mutation left, and otherwise logs and
leaves it alone. This covers the real-home install and opt-out sweep
(restoreRealHomeHooksJson), the legacy sweep's hooks restore, and config.toml
rollback (restoreCodexTrustConfig).

For hooks.json the generation is the exact bytes Orca wrote. For a config.toml
that a trust rebase changed it is the file as the rebase left it. When Codex
itself wrote config.toml inside the session that just failed, Orca never knew
those bytes, so that rollback compares against the file as the session settled.
The next commit keeps other Orca instances out of that window; a user edit made
during such a session can still be rolled back.

* fix(codex): serialize real-home Codex writes across Orca instances

Every Orca on one HOME (a dev and a packaged app, or an offline CLI) writes the
same ~/.codex/hooks.json, config.toml and ~/.orca/agent-hooks/codex-hook.sh.
The per-file lane that orders capture, mutate and restore was in-process only,
so another instance could write inside this one's restore window, or undo it.

The lane for the user's real config.toml now also holds the existing
crash-safe managed-hook install lock (~/.orca/managed-hook-install.lock, the
one relay installers take for the same home). It is taken only by the
outermost acquire, because the lock file is not reentrant and grants and trust
rebases nest inside an install. Managed-home installs, the real-home install
and opt-out sweep, and the legacy sweep all enter through it. Compare-and-swap
on restore stays as the backstop.

A lock that cannot be taken within its 10 s wait fails that install, which is
already best effort: launch prep logs it, and the real-home lane falls back to
the managed lane until its retry.

* fix(codex): opening a terminal no longer strips the shared Codex entry from ~/.codex

Every Orca instance on one HOME writes the same status-hook entry into the
user's ~/.codex/hooks.json, with its trust in config.toml. Launch prep runs on
every pane spawn, and under a managed Codex account it ran the legacy system
sweep. That sweep matched Orca entries by script file name, so it removed the
current shared entry and the trust blocks the grant ledger recorded. On a live
laptop hooks.json went 4139 -> 18 bytes about 150 ms before a new pane opened.
With hooks off, the real-home lane's launch prep swept the same way.

Now nothing automatic removes the current entry or its trust:
- The legacy sweep removes only an enumerated list of retired command forms
  that no build writes any more (#1019's double-quoted form, #1536's
  exec-guarded form, and Windows' per-userData bare path), plus their trust.
- ensureRealHomeCodexHookState with hooks off writes nothing; that covers
  launch prep, session resume and startup.
- Only the user's explicit opt-out (codexHookService.remove()) strips the entry
  and its ledger-recorded trust from the real home.
- The sweep-suppression gate existed only to stop the sweep from deleting the
  current entry, so it is deleted with its main-process wiring.

Startup with hooks off already skipped the real-home install; with this change
the first pane's launch prep with hooks off also leaves ~/.codex untouched.

* fix(codex): a pane's prepare-codex only repairs a home its own HOME's app installed

On macOS a pane starts through login(1), so it gets the user's real HOME even
when its Orca app runs with another one. The pane's `codex()` preflight
installed hooks in the CLI process with that real HOME: it rewrote
~/.orca/agent-hooks/codex-hook.sh, promoted trust into the real config.toml,
and wrote the real HOME's script path into the app's managed home.

The preflight now acts only when the managed home's hooks already run this
process's own shared script, which proves the app that installed them shares
its HOME. Otherwise it writes nothing; the app installed the home at spawn.

Why not a no-op: the preflight was added (#14326) because trust can go stale
between opening a pane and typing `codex`, for example in a pane that survives
an app update, and Codex then stops in hook review. For a same-HOME pane it
still repairs that. Why keep promotion: the install drops runtime trust the
system config does not back, so skipping promotion would delete approvals the
user gave inside Orca-launched Codex.

* test(agent-hooks): await every installer in the refresher coverage test

The test fired each managed installer without awaiting it and read
~/.orca/agent-hooks straight after. Codex's install now takes the
cross-process real-home lock before it writes its script, so the script
landed after the read. Await the installers, and stub Codex's trust sessions
so the awaited install cannot start a real `codex app-server`.

* fix(codex): retire the two real-home command forms the list missed

The real-home lane wrote two Codex hook forms into ~/.codex that no build
writes any more and that the enumerated retired list did not name:
- POSIX, #9501 until #10885: the file-guarded form draining with a bare `cat`.
- Windows, #9501 until #10221 took Windows off the real-home lane: the encoded
  PowerShell launcher for a non-cmd-safe script path.

The file-name sweep removed both before; the enumerated sweep left them in
place, trusted, still passing the script's exit status to Codex. Both now
match as frozen literals.

Also corrects the startup ordering comment: the real-home install runs first
so its in-slot upgrade lands before the managed install's sweep retires the
prior command; nothing re-arms a legacy sweep any more.

* fix(codex): take the real-home lock only when a write is needed

The previous commit made every entry to the real-home config lane take the
cross-process lock. That lane runs on every pane spawn and every typed
`codex` preflight, so the steady state paid an owner probe (a `ps` spawn on
macOS) and could wait up to 10 s behind another instance's trust session,
even though it wrote nothing.

Each real-home writer now compares the desired state with the files on disk
first, without the lock. Only when a write is needed does it take the lock,
re-read and recheck, then write:
- real-home install: the planned hooks.json, the shared script and the
  ledger-recorded grant are compared; the locked path re-plans from disk.
- legacy sweep: locks only when a retired entry is present; the sweep re-reads.
- approval promotion: locks only when there is something to promote; the
  promotions are recomputed under the lock.
- the shared ~/.orca/agent-hooks script: locks only when its bytes differ.
The explicit opt-out always takes the lock. The lock is reentrant through
async context, since grants and rebases nest inside an install, so the
config-lane option the previous commit added is removed.

* fix(codex): a shared script without its exec bit is not the steady state

The compare-first check matched the shared ~/.orca/agent-hooks script on bytes
alone. writeManagedScript also restores 0755 on every call, and the POSIX hook
guard skips a script that is not executable, so a script whose mode was lost
(a dotfiles restore, a plain copy) now stayed that way: every Codex hook
drained stdin and reported nothing until an app restart refreshed the script.

The check now also requires the mode the writer sets, so that case takes the
lock and the write path repairs it.

* test(codex): the retired encoded launcher never matches today's shared one

The shared encoded Windows launcher is still current for other agents, so the
comment claiming today's launcher is never encoded was wrong. What keeps the
retired matcher off it is the exact payload: since #14825 the shared launcher
prefixes its payload and drops -ExecutionPolicy Bypass. Pin that with a case.

* fix(codex): the pane step recognises its own script under a home path with an apostrophe

The same-HOME check looked for the script path wrapped in bare single quotes,
but both hook writers escape an apostrophe inside the quotes. A home such as
C:\Users\O'Brien never matched, so the pane-step repair never ran there.

* fix(codex): the trust-RPC escape hatch still keeps the real home off its lane

The no-write check reported a recorded grant as current, so with
ORCA_DISABLE_CODEX_TRUST_RPC set the real-home lane stayed in use. The grant
itself refuses before reading its ledger; the check now does the same.

* fix(codex): the shared script write no longer waits on the real-home lock

The write is atomic and skips identical bytes; waiting behind another
instance's trust session could only fail a pane's managed-home install.

* fix(codex): an in-Orca approval survives a launch that cannot get the real-home lock

The install drops runtime trust the system config does not back, so a
promotion skipped for want of the lock lost the approval for good. It
now writes unlocked, as it did before the lock existed.

* refactor(codex): take the cross-process real-home lock back out

The lock fixed no observed failure. The three that were observed each have
their own fix in this series: the legacy sweep matches only frozen retired
command forms, hooks-off launch prep writes nothing, and a pane's
prepare-codex repairs only a home its own HOME's app installed. The lock
instead brought its own defects: a steady-state spawn waiting behind another
instance's trust session, a compare-first split to avoid that, a script
write and an approval promotion that could fail for want of the lock.

Removed, with their tests: the real-home write lock and its async-context
reentrancy, the plan/compare split that kept it off steady-state spawns, the
compare-first legacy sweep, the locked approval promotion and its unlocked
fallback, the compare-first shared script write (writeManagedScript already
skips identical bytes and restores the exec bit), and the CLI tsconfig
entries the lock pulled in.

Kept: the retired-forms matcher, the hooks-off no-op, removal only on an
explicit opt-out, the pane own-script check, and the compare-and-swap
rollbacks. Every instance now writes identical bytes idempotently.

* fix(codex): an opt-out that cannot read hooks.json keeps Orca's trust and ledger

The opt-out swept the real-home entry, then dropped Orca's ledger-proven trust
whenever a ledger existed, even when the sweep could not read hooks.json. The
entry could still be there, now untrusted, and the ledger that proves ownership
was gone for the retry. Drop that trust only after a sweep that read the file.

* refactor(agent-hooks): one predicate for whether an agent's status hooks are on

"Global switch on and this agent not turned off" was spelled out separately
in the startup controls, the settings reconcile, the retained-home
reconcile, the WSL preflight RPC, the CLI preflight and the OpenCode plugin
selection. They now share one function, in a module light enough for the
CLI's per-launch Codex preflight to load. The PTY spawn env derives the
Codex flag from the switch and opt-out list it already carries, the same way
it does for OpenCode and Pi, instead of receiving a second copy.

* fix(codex): launch and resume prep honour Codex's per-agent hook opt-out

Turning Codex off in the per-agent hook settings removes Orca's Codex hook
entry, but launch prep and session resume read only the global hooks switch,
so the next Codex launch or resume wrote the entry straight back into the
real ~/.codex or the account's home. Both now read the per-agent predicate,
which the PTY spawn env and startup already honoured.

* fix(codex): turning Codex off per agent clears the real ~/.codex entry

While the real-home lane owns ~/.codex/hooks.json, the legacy system-home
sweep stands down. That gate read only the global switch, so turning Codex
off per agent ran remove() with the sweep still suppressed and left Orca's
entry in the real ~/.codex. The gate now reads the per-agent predicate, the
same as turning every hook off.

* test(codex): cover the system ~/.codex sweep gate for Codex turned off

The gate that lets the legacy system-home sweep run was an inline closure in
startup, so reverting it to the global switch left CI green. It is now a
pure function beside the gate it feeds, with a table test and a remove()
test on a seeded ~/.codex: turning Codex off strips Orca's entry and keeps
user hooks; with Codex on the entry stays.

* fix(cli): keep the agent-status hooks predicate loadable by the packaged CLI

The CLI's prepare-codex handler imported the predicate from src/main, but
the Electron build rebuilds out/main from its declared entries only, so the
packaged `orca agent hooks` commands could not load it (package jobs and the
CLI bundle-parity test were red). The predicate reads only settings, so it
now lives in src/shared, which the CLI compiles itself.

* feat(codex): every Orca build writes one frozen Codex hook command

The Codex hook command was built from this build's wrapper, so two builds
on one HOME disagreed about the bytes of the shared ~/.codex entry and kept
rewriting it, with a Codex trust session each time.

The command is now fixed per form and carries its form number:
- POSIX: one command with no path in it. It runs the shared script only in
  an Orca pane with hooks on (pane key and hook port set), drains stdin
  everywhere else, and always exits 0. A branch for a per-build script root
  is written now and stays dormant until Orca sets ORCA_AGENT_HOOK_ROOT, so
  that change will not move these bytes.
- Windows: the bare forward-slash path to the shared .cmd, which runs under
  PowerShell 7 and 5.1, Codex's hook hosts. A profile path that is not one
  PowerShell token gets a plain PowerShell form with the same branches.

The literals live in the form module, so a change to the shared hook
constants cannot move them; goldens pin the bytes. Every form keeps
`agent-hooks/codex-hook.*` in plain text, so older builds still recognize it.

* fix(codex): one main-process owner adds the real-home entry; nothing restores files

Each Orca writer of ~/.codex decided what Orca's entry must be from its own
build and instance, then removed or reverted whatever differed: launch prep
rewrote any Orca-shaped entry to this build's command and stripped Orca
entries from events this build does not use, and a failed trust session
restored hooks.json and config.toml from snapshots. With several instances
and builds on one HOME, every disagreement became a deletion or a revert.

The main process is now the one writer, and its writes are add-only:
- A launch or resume adds Orca's frozen entry to an event that has none and
  leaves every Orca entry it finds, so a running older build is never fought.
- App start also converts an older Orca form to the frozen command, once, in
  its own slot: one hooks.json write (one .bak) and one trust grant per home.
- A newer form is never rewritten or appended beside, and Orca entries in
  events this build does not use are kept.
- After a failed trust grant, only an entry this call wrote that is still
  untrusted is withdrawn, putting back the handler it replaced. Both files
  are re-read, so a concurrent edit, or the identical entry another Orca
  trusted meanwhile, survives.

Deleted: the compare-and-swap hooks.json restore, the config.toml snapshot
restore after a grant session and after a user-trust re-key, and the
rollback module. A grant session writes trust only at Orca's own keys, and
every caller settles those keys itself. A failed re-key of moved user hooks
now keeps the write and reports it; Codex lists those hooks for review.

* fix(codex): the pane CLI asks the app to prepare its Codex home

`orca agent hooks prepare-codex` ran Codex's install inside the pane. That
process can have the real HOME (login(1)) and runs outside the app's
in-process queues, so it was a second writer of ~/.codex and ~/.orca beside
the app. A check that the home ran "its own script" guarded it.

The pane step now only asks the app, over the same kind of local RPC the WSL
pane step already uses (agentHooks.prepareCodexForPane). The app checks that
the pane's CODEX_HOME is one its own userData owns, reads its own hooks
setting, and installs on its own queue. An app that is not running, or is
too old to know the method, makes the step a no-op, as it is on WSL. The
own-script check and the CLI's settings read are gone, and the preflight
module leaves the CLI bundle.

* fix(codex): delete the pane step on native hosts

The previous commit had `orca agent hooks prepare-codex` ask the app to
prepare the pane's Codex home. The case it existed for (#14326, a pane that
survives an app update with stale hook trust) did not reproduce, and no other
desktop agent host writes agent config from a terminal or launch wrapper.

- Deleted: the agentHooks.prepareCodexForPane RPC method, its params and
  catalog entry, and prepareManagedCodexHomeBeforeShellLaunch with its module,
  tests and CLI build entry.
- `agent hooks prepare-codex` is a no-op on native hosts. It stays for one
  release so shell wrappers from older builds, which still call it, exit 0.
- WSL panes are unchanged: they still ask the app over
  agentHooks.prepareCodexForWslPane.

The shell wrappers and ORCA_CODEX_LAUNCH_PREFLIGHT stay, because WSL panes
use the same wrappers and variable (forwarded through WSLENV). A native pane
still starts the CLI once per `codex` it runs; skipping that is a follow-up.

* test(codex): a failed trust session keeps concurrent edits to both files

QA case 9 at host level, on a real file system in a temp HOME: Codex's trust
session fails after another writer saved hooks.json and config.toml.

- Both saves survive, and no Orca entry is left that Codex would list for
  review: this call's entry is withdrawn.
- A failed one-time conversion puts the older Orca entry back in its slot and
  keeps both saves.

Both tests fail on the previous head, which restored config.toml from a
snapshot and left the untrusted entries in hooks.json. Removing the
withdrawal turns both red.

* feat(codex): read whether an Orca entry's stored trust is still current

A Codex release that changes how it hashes a hook leaves Orca's stored trust
stale: the entry is present, but Codex lists it as modified. Checking only
whether the entry is missing cannot see that.

readOrcaEntryTrust sorts a present entry into four states:
- trusted: the stored hash is the current one;
- untrusted: there is no stored hash;
- stale: the stored hash is not the current one;
- disabled: the user turned the entry off.

The caller can pass Codex's current hash, for example one a grant recorded.
The failed-grant withdrawal now uses it, and also keeps an entry the user
turned off. Nothing re-grants on 'stale' yet.

* fix(codex): a slow Codex start retries on the next launch, never for minutes

On a loaded Mac a cold `codex app-server` took over 10 s (QA case 4). The
grant timed out, the entry was withdrawn, and a 5-minute cooldown in both the
grant and the real-home install then refused every retry.

- The native session deadline is 30 s, the same as WSL's.
- A timeout starts no cooldown in the grant or in the real-home install. The
  next launch retries. Other failures keep their cooldown.
- Launches that queue behind a slow session share one follow-up run, so a
  launch waits for at most two sessions, not one per earlier launch.

Tests: a 15 s cold start still grants and keeps the entry; after a timeout,
the next launch runs a session at once; four queued launches run two
sessions. Each is red on the previous head, and each mechanism was removed in
turn to confirm its test turns red.

* fix(codex): Orca's automatic writes never move a user hook

Codex keys a hook's trust by its position in hooks.json. App start's collapse
of Orca duplicates removed every Orca entry and appended one at the end. That
moved any user hook that followed a removed entry, so the write waited on a
session to re-key the moved hook's trust.

App start now:
- converts the first Orca entry that sits in a plain slot to the frozen
  command, in place;
- drops any other Orca entry only when that moves no user hook;
- keeps a duplicate that a user hook follows, and trusts every frozen copy,
  so none is listed for review;
- appends only when no frozen entry is left.

Tests check user positions and user trust blocks byte-for-byte for each
automatic write: add-missing (append), the one-time conversion (in place),
a trailing duplicate, a duplicate before a user hook, and older duplicates
normalized to one entry. The three collapse cases fail on the previous head.
Removing the position check, or the in-place conversion, turns its tests red.
Only the explicit opt-out still removes an entry that user hooks follow.

* fix(codex): removing an Orca entry never waits on a Codex session

Removing an Orca entry from ~/.codex/hooks.json moves every user hook behind
it up a slot, and Codex keys trust by slot. The retired-form sweep, the
opt-out and a failed-grant withdrawal all asked a `codex app-server` session
to list the old trust before writing, and to re-key it afterwards. A timeout
there threw before the write and latched a 5-minute cooldown, so a slow cold
start blocked the retired-form sweep at boot (QA case 4).

Each moved hook's [hooks.state] block now moves to its new key, body bytes
unchanged, straight after the hooks.json write. Codex hashes a hook's content,
not its position or its file path, so the moved block stays exactly as valid
as it was: a trusted hook stays trusted, an untrusted one stays untrusted, and
one the user turned off stays off. No removal waits on or depends on a
session. A failed config.toml write keeps the hooks write and logs.

Deleted: the inspect and repair sessions, their client, and their cooldown.
The generation guards on the hooks.json writes stay, for other processes.

Tests: the retired sweep removes the retired entry and carries the trust of
the user hook behind it while every Codex session times out (red on the
previous head); the opt-out carries an appended user hook's trust; the move
carries trusted, disabled and untrusted states byte for byte. Removing the
move turns all of them red.

* fix(codex): a Codex launch never waits on Codex's approval of Orca's entry

A launch on the real-home lane awaited Codex's trust grant for the entry it
had just added. A cold `codex app-server` on a loaded Mac took over 10 s, so
the launch could wait that long, and a failure then latched a 5-minute
cooldown.

- Codex's approval runs in the background, with a 30 s cold-start budget.
- A launch uses the real home only when the ledger shows trust is already
  current. Otherwise it goes to the managed home at once, and the next launch
  picks up the finished grant.
- A launch that arrives while a grant runs does no work and does not queue
  behind it.
- A resume into the real home has no managed home to fall back to. It waits
  for the grant, but no longer than the 10 s a launch always could.
- A background grant that times out starts no cooldown; the next launch
  retries. Any other failure backs off for 10 s instead of 5 minutes.
  Success is what the ledger remembers.
- A failed grant still withdraws only what that install added and is still
  unapproved. The log now says how many entries it took back and when the
  next try comes.

Managed-home grants keep their 10 s deadline and stay on launch prep, as
before; they fall back to Orca-computed trust.

Tests:
- A 15 s start: the launch returns in under a second on the managed home, a
  second launch starts no session, the grant lands in the background, and the
  next launch uses the real home.
- A timeout sets no cooldown, withdraws its adds and logs it.
- Another failure retries after 10 s, not before.
- A resume waits only as long as allowed.
- Case 9 checks the log line and the retry.

Making the launch await the grant, a 10 s budget, either timeout cooldown, and
a 5-minute backoff were each tried, and each turns its test red.

* fix(codex): move a hook's trust only when every stored key has the known shape

Orca now edits Codex's trust store directly when a removal moves a user
hook. Three safeguards keep that honest:

- Fail safe. If any [hooks.state] key in config.toml does not have the
  shape `<path>:<event>:<group>:<handler>`, nothing moves and Codex asks the
  user to review. That shape was checked unchanged from Codex 0.141 to 0.158.
- Targeted. The file is read immediately before the atomic rename, and only
  the moved keys' blocks change. Every other byte stays, and no snapshot is
  restored.
- Verbatim. Each block's body moves as Codex wrote it, including fields
  Orca does not know. No hash is ever computed, and a hook with no block
  gets none.

Tests:
- An unknown key shape stops every move.
- Everything except the moved block survives byte for byte, and the moved
  body keeps an unknown field.
- In case 9, a hook the user approved during the failed session keeps its
  approval when the withdrawal moves it, beside the concurrent project edit.

Removing the shape check, or writing a computed block instead of the stored
body, turns these tests red.

* refactor(codex): keep only the trust read the failed-grant withdrawal uses

A capture across Codex 0.141, 0.150 and 0.158, switching in all six
directions, showed Orca's entry keeps the same hash and stays trusted. A
Codex upgrade does not make its trust stale, so nothing needs to re-grant
on staleness.

readOrcaEntryTrust keeps the four states the withdrawal needs, but loses
the parameter that let a caller pass a different current hash, and the test
for a Codex that hashes differently.

* fix(codex): native panes no longer start the Orca CLI before each codex

The pane step is a no-op on native hosts, but native panes still carried
ORCA_CODEX_LAUNCH_PREFLIGHT, so every `codex` typed in a pane started the
Orca CLI for nothing. Only a packaged Windows build's WSL pane now gets the
variable; the app prepares every native Codex home itself.

The resolver loses the dev-launcher path and its userDataPath option, which
only native panes used.

Tests: a native macOS, Linux and Windows pane gets no preflight, packaged or
not, even with the bundled CLI present; a WSL pane still gets the verified
absolute launcher. Letting native panes through again turns them red.

* chore(cli): say when the native prepare-codex no-op can go

Native pane wrappers from builds up to v1.4.216 still call it. It can be
deleted once no supported build's wrapper does.

* test(codex): check the WSL launcher path instead of asserting it

* fix(codex): a launch no longer waits behind the background real-home approval

The background grant ran its whole codex app-server session inside the shared
~/.codex/config.toml lane, and on a cold host its session was also the shared
capability probe. A launch sent to the managed home then waited on both: the
managed install and the project-trust write queue on that lane, and the
managed install's own grant waited for the probe. On a cold app-server that
was up to 30 s per launch.

The lane was held across the session only to protect the retired
capture-and-restore. Codex writes its own records, so the lane is now taken
only around Orca's own pre-grant write. The background grant runs its session
without publishing it as the shared probe, and the whole grant is bounded by
its deadline, so a hang outside the session cannot leave the lane 'granting'.

* fix(codex): a failed re-grant no longer strips Codex's own approval of Orca's entries

Before each trust session, the grant deleted every Orca record whose hash
matched the one Orca computes. That exists because a managed home's fallback
writes Orca-computed trust under both Windows path-separator spellings, and
Codex rewrites only its own spelling, so the other copy would linger. On
failure the managed and WSL fallbacks write that trust back, and before this
fold a snapshot restore covered it.

The real ~/.codex has neither: Orca never writes computed trust there (the
real-home lane does not run on Windows at all), so a matching record there is
Codex's own approval. After a ledger miss (another Orca profile, a Codex
update, a lost ledger) and a failed session, nothing put it back, and every
Orca entry showed "Hooks need review".

The clear now runs only for homes whose fallback writes that trust.

* fix(codex): a real-home resume spawns only once Orca's entry is approved or withdrawn

A resume that must run in ~/.codex waited at most 10 s for the background
approval, then spawned anyway. On a cold app-server that left Codex beside an
unapproved Orca entry, so the resumed pane showed hook review.

The resume now waits for the grant to settle. Settled means Codex approved the
entry, or the grant failed and withdrew its own unapproved write; the grant's
deadline bounds the wait (30 s, the cold-start budget), and a failed approval
never fails the resume.

Why this over the alternatives:
- Spawning at 10 s keeps the review prompt this fold exists to remove.
- Withdrawing at 10 s from the resume races the still-running session: Codex
  can write the frozen entry's hash after the withdrawal, and for a converted
  entry that marks the older command Orca put back as modified.
- A resume cannot use the managed home: the session lives in ~/.codex.
So the only states that cannot race Codex are the grant's own settle. The cost
is a longer worst case on a cold app-server (up to the 30 s deadline, plus any
managed-home install that holds the config.toml lane); a warm approval takes
seconds, and an approved entry costs no wait.

* fix(codex): keep the 5-minute trust cooldown for launch-path grants

The fold shortened the host's trust-grant cooldown from 5 minutes to 10
seconds for every grant. That was meant for the background ~/.codex approval,
which blocks no launch. The managed-home and WSL grants run inline on the
launch path, so with a hung app-server every launch more than 10 s after the
last failure paid the full inline timeout again (10 s native, 30 s WSL).

Cooldowns are now kept per lane: inline grants keep 5 minutes, the background
grant retries after 10 s, and neither lane's failure cools the other down. A
success, or a proven-missing surface, still clears both. The real-home
install's own retries (an unreadable hooks.json, unknown keys) are back on the
5-minute interval they had before the fold.

The cooldown moves to its own module so the grant stays within the file limit.

* fix(codex): a failed grant withdraws the exact copy it wrote

The withdrawal re-found "this call's" entry by command, taking the first
frozen handler in the event. When app start converted a later slot while an
earlier frozen copy sat in a matcher group (which conversion skips), a failed
grant acted on that earlier copy: it put the older command into it, or skipped
it, and left the converted, unapproved copy in place.

Each write now records where its handler landed, after any duplicate drops,
and the withdrawal acts only on that slot. A copy that has since moved is left
alone; the next launch's grant retries it.

* fix(codex): the failed-grant withdrawal checks hooks.json is unchanged before writing

The install and the retired-form sweep both refuse to replace ~/.codex/hooks.json
if it changed since they read it. The withdrawal did not: a save landing
between its read and its atomic replace was lost. The window is small, since
the withdrawal is synchronous, but it now carries the same guard.

* refactor(codex): drop rationale left over from the snapshot restore; name the trust-move module for what it does

Comments on the config.toml lanes still justified them by a grant's
capture-and-restore window, which the fold deleted, and the trust-write
deadline still counted a grant session holding the lane. They now give the
reason that remains: Orca's own multi-step reads and writes, and managed-home
installs that hold the lane across their inline grant.

codex-user-hook-trust-rebase no longer rebases through Codex; it moves stored
trust records, so it is now codex-user-hook-trust-moves.

The grant test that pinned two sessions on one config.toml to run one at a
time is removed: its reason was an interleaved capture and restore. Callers
that write config.toml around a grant hold their own lane, which the nested
installer test still covers.

* build(cli): list the trust-grant cooldown module in the CLI program

The CLI's agent-hooks handler loads the hook controls, which reach the Codex
trust grant; the CLI project is composite, so every module in that graph must
be listed.

* docs(codex): say which Windows hosts each hook command form runs under

Codex runs a hook under the turn's shell (PowerShell 7 or 5.1 in every
captured session) and, with no single local turn shell, under %COMSPEC% /C.
The bare forward-slash path ran under all three in the Windows host census.
The PowerShell form used for a profile path with a space does not parse under
cmd.exe; no form valid in all three hosts has been run for such a path, so the
form stays and the gap is stated here and in the PR.

* test(codex): type the withdrawal seam without an assertion

* fix(codex): a real-home resume starts at once, trusting Orca's entries for that process

A resume that must run in ~/.codex waited for Codex's background approval of
Orca's newly written hook entry: up to 30-40 s on a cold app-server. That made
the user's resume wait on bookkeeping, and the alternatives (start at 10 s with
Codex's hook review showing, or withdraw the entry and race Codex's own write)
were worse.

Codex reads hook trust from its session-flag config layer as well as the user's
config.toml, merged per key, and has since hook trust shipped. So the resume no
longer waits. When Orca's own frozen entries in ~/.codex are untrusted (or hold
a stale hash), the resume command carries
`-c hooks.state={'<key>'={trusted_hash='<hash>'},...}` for exactly those entries:
the key under both the logical and the real path of ~/.codex (Codex keys an
explicit CODEX_HOME by its real path), and the hash of that entry's content, so
it can trust nothing else at that slot. The user's hooks are never included,
nothing is written, and the background approval still runs for later plain
`codex` launches. An approved entry adds nothing; a Codex known to lack hook
trust gets nothing.

One inline table, because Codex splits a `-c` key on every `.` and the key holds
`.codex/hooks.json`. TOML literal strings keep `"` out of Windows native-argument
quoting. The flag goes before `resume <id>`, quoted for the pane's shell (portable
Unix, PowerShell or cmd), in the launch command and in the setup-sequenced copy of
it; a cmd line whose path cmd would expand, or a key with an apostrophe, is left
unchanged. SSH and WSL resumes get no preparation, so no local path reaches them.

* Revert "fix(codex): a real-home resume starts at once, trusting Orca's entries for that process"

This reverts commit 1bd30651d6.

* fix(codex): a real-home resume starts at once, without waiting for approval

A resume into the real ~/.codex waited until the background approval settled,
up to its 30 s deadline on a cold app-server: bookkeeping for later launches
gating the resume the user asked for. It now starts at once. If the approval is
still running, that first resume can show Codex's hook review once; the
approval then lands and later resumes and plain codex launches are trusted.

Trusting Orca's entries per process was the alternative, but the resume command
is typed into the pane's shell, and hook settings stay out of typed commands.

* test(codex): read real-home hook groups with the installer's own type

* fix(codex): a background approval is bounded only by its session's own deadline

Review loop 2, L3. grantWithinDeadline raced a second 30 s timer against
the background approval. Loop 1 added it so that a hang upstream of the
session could not leave the lane 'granting' forever.

That hang cannot happen. The only caller is the native real-home grant
(its plan is always host 'native'; the real-home lane is off on Windows,
so WSL never reaches it). Everything before the session is synchronous
there: command resolution and binary stamp, the ledger read, the
state-db backfill check, the capability and cooldown checks, and
runUnshared awaits no shared probe. A synchronous hang would freeze the
main thread, which no timer can rescue. The session itself starts a kill
timer right after spawn (runCodexAppServerSession), with the same 30 s,
and it kills the app-server tree when it fires.

So the outer timer was a second copy of that bound. Because it started
first, it won by the spawn time. It then settled the lane and cleared
backgroundGrant while the app-server was still alive, and the next
launch could start a second concurrent session. It abandoned the
session rather than cancelling it. Deleted, not moved: the session's
own timer is the one bound, and it cancels.

Test: codex-real-home-slow-app-server.test.ts "runs one session at a
time, ended by its own deadline". The fake session starts its timer
after a simulated spawn, as the real one does. A launch at 30 s finds
the session still running and starts none; the lane settles when the
session times out. It replaces the "settles a grant that never answers"
test, whose never-answering session could not time out at all.

* fix(codex): a background approval's retry has one schedule, the real-home lane's

Review loop 2, L4. A non-timeout background failure set two 10 s
schedules for one failure: the real-home lane's installRetryAfterMs,
which gates ensure, and a `<host>#background` cooldown in the grant
module. ensure's gate always tripped first, so the second one was
consulted only after something reset the first (turning hooks off).
Then it answered 'retry-cached', which wrote the entry into
~/.codex/hooks.json only to withdraw it again: churn, not protection.

Background plans now neither start nor consult a grant-module cooldown.
The real-home lane (installRetryAfterMs) is the one source of truth for
when a background approval runs again, and its 10 s interval moves into
codex-real-home-background-grant.ts, the module that sets it. The
cooldown module is back to one host-keyed map for launch-path grants,
with the same 5-minute interval as main. A success or a proven-missing
surface from either lane still clears the host's cooldown.

Tests:
- codex-hook-trust-grant.test.ts "neither starts nor waits on a
  cooldown for a background grant": two failing background grants each
  run a session and leave no cooldown; an inline failure still cools
  down inline grants and not the background one.
- codex-real-home-slow-app-server.test.ts "has one retry schedule:
  turning hooks off and on after a failure retries at once": after a
  failed approval, hooks off then on runs a session and installs,
  instead of a retry-cached write-and-withdraw.

* fix(codex): hooks turned off and on during an approval re-add Orca's entry

Review loop 2, L1. ensure returned at once whenever a background
approval was running, whatever the lane. Turning hooks off during an
approval sets the lane to 'removed' (usable), so turning them back on
returned 'removed' without re-adding the entry. Launches in that window
spawned in ~/.codex with no Orca hook and got no status for their
lifetime, for up to 30 s, until the approval settled and a later launch
re-added it.

ensure now returns early only while the lane is 'granting', which is
what the early return exists for: a launch never waits on Codex's
approval and uses the managed home until it lands. Any other lane runs
the normal add-missing install.

That install can start a second approval while the first is still
running. Approvals are now chained, so Codex still runs one session at
a time, and a finished approval clears the handle only if it is still
the latest one (before, an older approval's finally could clear a newer
one's handle). The older approval's result is already dropped by the
lane generation check.

Test: codex-real-home-slow-app-server.test.ts "re-adds the entry when
hooks go off and on during an approval, one session at a time". While
the approval hangs: opt-out removes the entry; re-enable re-adds every
entry, keeps launches on the managed home, and starts no second
session; once Codex answers, the lane is installed and every entry is
approved.

* test(codex): a launch during the real-home approval shows what it waits on

Review loop 2, M2. The launch test's fake Codex failed every
managed-home session at once with ENOENT, so the managed home's own
approval was an instant "unsupported" fallback, and the test could not
show that a launch sent to the managed home still waits on that home's
inline approval when its ledger misses (first use, a Codex update, a
lost ledger), up to 10 s, as on main.

Now the managed-home session behaves like a real one:
- "settles on the managed home with its hooks and the project trust
  written": the managed app-server answers; two launches settle in
  under 2 s while the real-home approval hangs, and the second launch
  finds the managed approval in its ledger (one managed session).
- new "waits up to the managed home's own 10 s approval when that home
  is cold too": the managed session fails at its own deadline, as the
  real one does. The first launch is still pending at 9.999 s and
  settles on the managed home at 10 s; the request asked for 10 s. The
  next launch settles at once, because the failed inline approval cools
  down for 5 minutes.

No product change.

* refactor(codex): the managed and WSL installs own their pre-approval trust clear

Review loop 2, L7. Before a Codex approval session, a managed or WSL home
clears the approvals Orca itself computed, because on Windows its
fallback writes them under both path spellings and Codex's canonical key
may not overwrite the other one. The fallback writes them back if the
session fails. ~/.codex has no such fallback, so there the clear would
only delete Codex's own records (loop-1 H2). The grant module carried
this as a plan flag, fallbackWritesSelfComputedTrust, and took the
config.toml lane around the clear itself.

The reviewer proposed moving the clear into the two callers. A literal
move, clearing before the grant call, is NOT behaviour-neutral, so this
does not do that:
- The grant first checks its ledger, which compares the stored hash
  with the one Codex recorded. Codex's hash equals Orca's computed one
  (the premise of readOrcaEntryTrust), so a clear before that check
  deletes exactly the record the ledger proves. Every managed launch
  would then miss the ledger and run an inline session (up to 10 s).
- Checked, not inferred: with the clear moved before the call in the
  managed install, codex-launch-during-real-home-grant.test.ts "settles
  on the managed home..." fails (2 managed sessions instead of 1).
  Log: ~/orca-qa/codex-real-home-leak/fb6/l7-literal-move.log

What this does instead: each caller passes its clear as the grant's
`beforeSession` step, which the grant runs only when a session will
actually run (after a ledger miss, and not on a cooldown or cached
fallback), exactly where the flag ran it. So:
- the flag and its "never set for the real home" rule are gone; the
  real-home grant passes no step, so the grant module has no path left
  that deletes a trust record in ~/.codex;
- the grant module's own lane acquisition around the clear is gone. It
  was always a pass-through: both callers already hold that file's
  lane (the managed install holds the runtime and system lanes, the
  WSL install holds its config.toml lane) across the whole grant.

No behaviour change. The loop-1 probes still pass as fixed: trust-strip
prints every entry trusted after a failed re-grant, and lane-hold
prints managedInstall=settled projectTrust=settled.

Tests (codex-hook-trust-grant.test.ts):
- "removes equivalent Windows fallback keys before the RPC writes
  canonical trust" now passes the managed caller's step;
- new "runs the caller's pre-session step only when a session runs":
  the step runs once for a session and not on the ledger hit after it.

* chore(codex): comments stop describing a lock held across the session, or a rollback

Review loop 2, L6 comment sweep (comments and one test name only):
- codex-trust-config-concurrent-launch.test.ts: the test named "does not
  let a failing launch roll back a concurrent launch" said the per-file
  lane was the only thing left and that the doomed run's rollback must
  not resurrect the file. There is no lane across a session and no
  rollback now. Retargeted to what it covers: "leaves a concurrent
  grant's records in place when a sibling grant fails" (a restore would
  still turn it red).
- codex-trust-grant-ledger.ts: "a grant session blocks launch prep" is
  true only of inline grants; the background one still costs an
  app-server start. The drift clause no longer says "before the pane
  launches", which is false for the real home.
- agent-trust-write-deadline.ts: a stray hard wrap.
The install.ts:105 comment was fixed with L1. A sweep of src/main/codex,
src/main/startup, src/main/agent-hooks, the trust presets and the CLI
handlers for rollback, restore, rebase, capture/restore, and a lane held
across a grant or session found nothing else stale; the remaining "no
restore" comments state the current rule.

* fix(codex): a real-home resume waits for the one running approval, up to its 30 s limit

Review loop 2, M1; coordinator ruling. A resume into ~/.codex has no
managed home to fall back to. 5a737261d8 let it start at once beside an
Orca entry still awaiting Codex's approval. Codex's TUI then shows a
full-screen hook-review picker before the session and waits for keys:
"Trust all and continue" also trusts the user's own unreviewed hooks,
and "Continue without trusting" leaves that session with no Orca status
for its whole life, because Codex does not reload hooks when Orca's
approval lands later. Panes restored at app start after an update hit
it too, since the start-time conversion leaves every entry awaiting
approval.

The resume now waits, but only while Orca's entry in ~/.codex is
written and a grant is approving it (lane 'granting'). Every resume
waits on that same in-flight grant: ensure never starts a second one
while the lane is 'granting', so panes restored together share one
session. The bound is the grant's own session limit (30 s). The grant
settles only after Codex approved the entry, or after it withdrew its
own unapproved adds, so the resumed session starts either trusted or
with no Orca entry: never beside an unapproved one, and no picker. On
a withdrawal that session has no Orca status, as on main after its
10 s wait. A failed approval never fails the resume.

Tests (codex-launch-during-real-home-grant.test.ts):
- "waits for a warm approval, and spawns with the entries approved";
- "spawns at the approval session limit with Orca entries withdrawn"
  (fake timers: pending at 29.999 s, spawns at 30 s with no Orca entry);
- "makes panes restored together wait on one approval session" (three
  resumes, one session, all settle once it lands).
codex-launch-per-agent-hook-opt-out.test.ts: a resume into ~/.codex
awaits the approval; a resume into a managed account home does not.

* fix(codex): repeated background approval timeouts back off, growing to 5 minutes

Review loop 2, M3; coordinator ruling. A timeout of the ~/.codex
approval starts no cooldown, so the next launch retries at once. On a
host where codex app-server never starts within 30 s, every launch then
wrote Orca's entry into ~/.codex/hooks.json, withdrew it again, and
started another 30 s session, for the rest of the process: an unbounded
retry with no exit.

After 3 timeouts in a row the retry now waits 10 s, then 1 minute, then
5 minutes for every later one. The first two timeouts still retry on the
next launch, so a slow cold start is not punished. Any other outcome
ends the streak (a success, or any other failure, which keeps its own
10 s wait). The streak lives only in memory, so every app start begins
at zero and a slow boot can never latch.

Tests (codex-real-home-slow-app-server.test.ts):
- "backs off after three timeouts in a row, growing to 5 minutes, and a
  success resets it": the first two timeouts retry at once, then 10 s,
  1 min, 5 min, 5 min; after a success, a fresh approval gets two
  immediate retries again and a 10 s backoff after the third;
- "keeps trying after timeouts during a slow first start, once the app
  server answers": three timeouts, then the next attempt at 10 s
  installs.

* refactor(codex): one approval at a time, decided under the config.toml lane

The real-home check kept a lane label, a generation stamp, a promise chain of
ensures and a chain of approvals, and decided from the label at call time.
Concurrent resumes from any state other than 'granting' each started their own
approval (N x 30 s), a chained approval ran a plan an earlier failure had
withdrawn, a hooks-off check during an approval released a waiting resume beside
unapproved entries, and an app-start conversion during an approval was dropped.

Now each check is one step under the real config.toml lane: an approval in
flight answers 'approving' (unusable), hooks off answers 'removed', an open
retry window answers 'unavailable', and otherwise the unchanged install runs and
starts at most one approval. The approval settles under the lane: it withdraws
its own unapproved adds on failure, sets the retry, and derives the verdict from
the settings and the outcome, then runs an owed conversion. A resume waits only
while an approval runs and an unapproved Orca entry is on disk. The opt-out
sweep moves verbatim into its own module.

* fix(codex): only a success or app start resets the approval timeout streak

The ruling is that three timeouts in a row back off, and the count resets on
success and at app start. A non-timeout failure or an unexpected error also
reset it, so a host alternating those with timeouts never backed off.

* fix(codex): a Windows profile path the shells cannot carry bare runs through cmd.exe

The Windows hook command was the bare forward-slash script path, or, for a
profile path that is not one PowerShell word, a PowerShell script. That script
cannot parse under cmd.exe, which Codex uses when a session has no single local
turn shell, so such a profile got no status there.

A path of only letters, digits and _ . : / ~ - stays bare. Any other path,
including one with a space, & ^ $ ` ' ! ( ) or a non-ASCII character, is written
as cmd --% /d /c @"<path>", which ran under PowerShell 7, Windows PowerShell 5.1
and cmd.exe for each of those characters with a real Codex 0.158.0. The choice
depends only on the path, so every build on a machine writes the same bytes. A
machine holding the earlier PowerShell spelling converts it once at app start.

* build(cli): list the real-home hook sweep module in the CLI program

* fix(codex): the Windows cmd spelling names the system cmd.exe and turns off delayed expansion

A profile path the shells cannot carry bare was written as
cmd --% /d /c @"<path>". Under Codex's cmd.exe host the outer cmd.exe resolves
a bare `cmd` from the hook's working directory first, so a repo holding
cmd.bat (or .cmd, .com, .exe) at the session cwd would run on every hook event.
And with delayed expansion turned on in the registry, a `!` in the path was
dropped.

The spelling is now <SystemRoot>/System32/cmd.exe --% /d /v:off /c @"<path>",
unquoted (PowerShell reads a quoted first token as an expression) and with
forward slashes. The Windows directory comes from %SystemRoot% when written,
else from the directory above %ComSpec%'s System32, so both give the same bytes;
if neither is a drive-absolute path it can spell unquoted, it is C:/Windows,
which is still absolute. The bytes stay a pure function of the profile path and
that directory, so every build on a machine writes the same command. Safe
profile paths keep the bare path. Older Orca forms, including the bare-cmd
spelling, convert once; the new spelling is never swept as retired.

* refactor(codex): an approval's settle runs no deferred conversion

An app-start conversion that arrived while an approval ran was remembered and
run by that approval's settle. The settle then rewrote an older entry in place,
unapproved, and started a second approval inside the same wait that releases
every resume, so a resume could start beside an entry Codex would put up for
review.

That path could not happen: the only conversion caller is app start, and it is
the process's first check, so no approval can be running when it arrives. The
deferral and the settle's second check are deleted. A conversion that met an
approval would now be skipped until the next start, and the test for this case
pins that the settle writes nothing new and runs one session.

* fix(codex): an approval's settle keeps a failed opt-out's verdict and ends only its own flight

With hooks read off, an approval's settle always concluded 'removed', which the
routing check treats as usable. If an opt-out during that approval could not
read hooks.json, it had concluded 'unavailable' because the entry may still be
there, and the settle overwrote that. The settle now keeps 'unavailable' when
hooks are off; the next hooks-off check or opt-out re-derives it as before.

The settle's fallback when it cannot run now clears the running approval only
if it is still its own, and the routing check's comment states its rule: never
usable while an approval runs.

* fix(codex): spell the system cmd.exe with backslashes

Under Codex's cmd.exe host the outer cmd.exe hands the typed program text to
the child verbatim, and cmd.exe scans its whole command line for switches, so a
forward-slash C:/Windows/System32/cmd.exe is read as switches: the hook never
runs ("The syntax of the command is incorrect.") and /d is lost. Measured live
on Windows; both PowerShell hosts rewrite argv0 and were unaffected. The script
path after @" keeps forward slashes.

* chore(codex): say why the cmd.exe path is absolute, as measured on Windows

* test(codex): Windows managed-install tests expect the frozen command

They still asserted main's PowerShell text and a backslash bare path; they only
run on Windows, so nothing here caught it. Also correct the /v:off comment: a
lone ! is never dropped, only a !NAME! pair expands.

* ci: run the Codex managed-install tests in the Windows job

Its Windows-only cases skip everywhere else, so nothing ran them; three of them
still asserted a command this branch no longer writes.

* ci: a change to the Codex managed-install tests starts the Windows job

Also say what the missing-script case asserts: a non-zero exit, which
PowerShell reports as 1.

* chore(codex): name the hook trust key pattern for what it matches

* test(codex): the managed-install tests remove folders with the retrying helper

Now that they run in the Windows lane, a raw recursive rm there can throw EPERM
after the assertions pass.

* refactor(codex): one Codex hook-trust key pattern for the trust move and #23958's carry

* test(codex): the trust move carries a block in Codex's quoted spelling and leaves no second table
2026-09-30 14:26:23 -07:00
Brennan Benson afa81dc3ad fix(native-chat): chat failure messages appear in the app's language (#23674)
* fix(native-chat): a read whose history will not open is refused with its reason

History, subscribe, snapshot and options reads reach a chat through one accessor, whose open had no
catch: a journal that would not open reached every client as a runtime error carrying the storage's
own text (a path, "file is not a database"). The accessor, and the options read's own open, now throw
the classified journal refusal: journalCorrupt when SQLite reports damage, journalUnavailable
otherwise. The storage text goes to the log only.

The wire code stays runtime_error and the message becomes the bare code, as for every thrown
refusal; the reason rides in the error's data.

* refactor(native-chat): the idle sweep's stop of a hung start carries no hand-written reason

The sweep passed an English sentence as the stop's reason. It lived only in memory and nothing read
it: the delivery loop words the error row and the rejection from the hostStopped fact. Dropped, with
the display-name lookup that built it.

* fix(native-chat): a refusal the host throws is worded from its data, never its message

A thrown agent-session refusal reaches the client as runtime_error with the bare code as its message
and the typed refusal in error.data. Stop, answers, options and goals, a launch's held option pick,
the option picker's failure toast, and the Retry line of a chat that could not start now word it
from that refusal through the shared notice table. What each caller decides about the outcome is
unchanged: only the words move.

The Retry line of a failed start no longer prints the host's message or a thrown error's text; it
keeps the refusal as a fact and says the cause and step its reason names, or only that the chat
could not be started. The option toast keeps a local option surface's own sentence.

* fix(native-chat): an unreadable history is worded from its refusal, and damage stops the retry

The structured chat's read failure showed the host's text on the status line, and the pane always
said Orca keeps trying. The read transport now takes the refusal from the error's data (a stream
payload or a thrown RPC error), the reducer keeps it beside the failure text, and the pane and the
status line word it through the notice table, once: on the pane when the failure took it, else
beside the transcript that stays.

A damaged journal says "Unable to load this chat." and the read stops reconnecting for that run;
reopening the chat reads again. An open that can clear names its cause without "Try again", since
the pane retries on its own. A failure that names no reason keeps today's generic line. Finality
comes from the refusal's reason, never its message, which is the bare code for both.

* fix(native-chat): a rejected message is worded from its stored fact

A message the host recorded and then rejected keeps the host's typed fact beside its reason, but
the Retry words re-read the reason alone. Now the fact decides: a hand-over failure says Orca
couldn't reach the agent, a kind whose reason may be a legacy marker gets its fact's own sentence
(a full queue now says so instead of only "not sent"), and any other kind shows the sentence the
host wrote for it, which carries the agent's name and any words the provider wrote for a person. A
row with no fact reads as before.

* fix(native-chat): each message that did not go through says why on its own row

The structured chat showed one Retry strip under the transcript for whichever single entry it
picked, so a second failed message had no reason and no Retry of its own. The terminal-backed
chat's existing per-row delivery marker now carries a notice and an optional Retry, and the
structured pane derives one per message from the outbox on each render: every rejected message,
and the one the queue stopped on (read through the drain's own rule, so a Retry never names a
message waiting behind it). Each is worded from that message's stored failure. The single strip is
deleted. Nothing new is stored, and the shared message projection is untouched.

* fix(native-chat): say each chat failure's words where its own control already acts

Three wording rules for the desktop chat:

- A chat that could not start shows Retry beside its reason, so the reason stops at its cause
  where the Retry is the step: a reason whose action is to retry, and a start failure's "send
  your message again". Any other step stays (quit the terminal agent, start a new chat). The
  start-failure sentences take a retryControl context for this; what the host writes is unchanged.
- A history that couldn't open right now still reconnects, so the pane keeps "Orca keeps trying
  to load it" under its cause. Only a damaged history, which no retry reads past, drops it.
- A read failure that names no reason while the transcript is shown is only the pane
  reconnecting: the status line says "Reconnecting to this chat…" in muted text, not an error.
  New key components.native-chat.state.reconnecting, hand-translated for es/fr/ja/ko/zh.

* fix(native-chat): a rejected message offers Retry only once the queue is moving

Each rejected message's row offered its own Retry even while the queue was stopped on another
message. Any Retry clears the stopped queue, so pressing a rejected message's Retry also sent the
message the queue was holding, which the user had not retried; behind a message whose delivery is
unconfirmed, the retried one instead went back into the queue with no notice and waited there.

While the queue is stopped, only the message it stopped on offers Retry, as the single Retry strip
this replaced did. A rejected message keeps its words on its row and gets its Retry back once the
queue moves.

* fix(native-chat): a chat whose history will not open logs once, not on every reconnect

A reader reconnects every 750 ms while a journal open can clear, and each attempt logged the
failure with its full stack. The read door now logs a session's failure once until that session
opens, closes, or fails differently; every attempt is still refused with its reason.

* fix(native-chat): a message's own Retry is its resend step, so its notice stops at the cause

A rejected message offering Retry read "Claude stopped before it finished starting. Send your
message to try again." beside that button. Its row now takes the rule the launch strip already
follows: beside its own Retry the words leave out sending or trying again, worded from the stored
fact with the chat's agent name. The stored fact keeps less than the host wrote from (a refusal,
the provider's words), so a reason it cannot rebuild exactly is kept as written. A rejected
message without a Retry, while the queue is held, keeps the step. What the host writes and the
phone's notice are unchanged.

* fix(native-chat): a message's Retry sends only that message, never the one the queue is held on

Retry released the queue's refusal hold whichever message it was pressed on. While a queued message waited ahead of a held one, every rejected message offered Retry, and pressing it also sent the held message the person had not retried. Retrying an unconfirmed message ahead of a held one did the same. Retry now releases the hold only for its own message.

* fix(native-chat): a not-signed-in failure beside Retry still says to sign in first

Beside a Retry the notice dropped the whole next step, so a chat that could not start because the agent was not signed in read only the cause. Pressing Retry without signing in fails the same way again. The words now keep the sign-in step and leave out only the resend, which the Retry button is.

* test(native-chat): a rejected message's hidden Retry only avoids waiting unseen

* fix(native-chat): every Retry beside a notice leaves out the retry step the same way

A message the queue stopped on worded its refusal with no agent name and with its retry step, beside its own Retry, while a rejected message next to it named the agent and left the step to the button. The launch strip and the history pane each had their own copy of the same rule. One wording context now goes through the one notice table for every surface: a Retry beside the words, or a pane that reconnects on its own, is the step for a reason whose action is to retry, and every other step stays. What the phone and the host write is unchanged.

* test(native-chat): read the sent message id without a type assertion

* fix(native-chat): a rejected message is worded from the journal's own fact, never by comparing sentences

A message the host recorded and then rejected kept only the rejection's kind on the message, so its notice was reworded from that smaller copy only when it rebuilt the host's sentence word for word. A different agent name, an older host's wording, or anything the copy dropped (why a start failed, the provider's own words) left the host's sentence in place, beside a Retry that repeated its resend step. The notice now reads the journal's own rejection for that message, found by id, with the pane's agent name and Retry, and shows the provider's words only when they were written for a person. The message's smaller copy words it only when that journal row is not loaded. Nothing new is stored.

* fix(native-chat): a message rejected before a restart retries under a new id the first time

Whether a Retry needed a new message id was remembered in memory for one message, or read from the journal row when it was loaded. After a restart, or for an older message whose row was not loaded, the first Retry resent under the old id, the host answered with the same settled rejection, and nothing visibly happened. The message now says so itself: one the host recorded and rejected always retries under a new id, including after a restart. A refusal that already gave the message a fresh id, and a message whose delivery is unconfirmed or in flight, keep their id as before.

* fix(native-chat): a rejected message older than the loaded history keeps the provider's words

When the journal row that rejected a message is not loaded, the message's own copy of the
fact has no provider detail or start refusal. For the kinds worded from those, the row now
shows the sentence the host wrote for the person instead of a thinner rebuilt one.

* test(native-chat): the chat pane words a rejected message from its loaded journal row

Nothing covered the pane handing the journal's rows to the per-message notices, so a pane that stopped passing them would quietly fall back to the message's smaller copy of the rejection and show the host's sentence, resend step and all. The new case renders the pane with a rejected message whose journal row is loaded and checks that it reads that row's refusal in the chat's own agent name.

* fix(native-chat): a chat whose history won't load says why in one line

A read the host refused for a named reason put its sentence under the generic
"Could not load conversation" title, so a damaged history read as two lines
saying the same thing. The pane's own sentence now takes the title's place; a
history that can come back keeps its line saying Orca keeps trying. A failure
that names nothing keeps the generic title.

* fix(native-chat): a message a failed start rejected says only that it was not sent

When an agent stopped before it finished starting, the chat showed the start's
red row ("Claude stopped before it finished starting. Send your message to try
again.") and then repeated that cause under every message the start rejected.
Each of those messages now reads "Your message was not sent." beside its Retry.

The match is made on typed facts, not on the words: the host writes the start's
row and the rejection of its queued messages from the same failure fact, and the
row is keyed by the start. The pane finds the loaded start-failure rows by that
key and shortens a message's notice only when its loaded journal submission was
rejected with the same fact. Any other rejection, or one whose submission or row
is not loaded, keeps its full notice. The row key moves to a shared module so the
host that writes it and the pane that reads it use one definition.

* fix(native-chat): a chat whose history keeps failing to open retries less often

A read the host kept refusing (its history store could not be opened right now)
reopened every 750 ms for as long as the chat stayed open, about 40 opens every
30 seconds. Each reconnect now waits twice as long as the last, from 750 ms up
to 30 seconds, and never gives up; the first read that delivers anything starts
the wait over at 750 ms. A damaged history still stops reconnecting at once.

Reset happens on a delivered read, not on connect: a local subscribe resolves
before the host's open refuses, so resetting there would keep the 750 ms loop.

* test(native-chat): the pane harness types its journal rows without a cast

* test(native-chat): the admission test passes no start-failure rows to the notices

* test(native-chat): import the journal types once

* fix(native-chat): a remote chat reads again as soon as its host is back

The read retry doubles its wait up to 30 s during an outage, and nothing
reset it when the remote runtime reconnected, so the transcript could
lag the reconnect by up to 30 s. The read now watches the runtime
status store's contact-regained edges (hostContactEpoch for a
same-runtime return, connectionGeneration for a new runtime session)
and, when one lands, runs a waiting retry immediately with the wait
reset to its base.

* refactor(native-chat): build each failure sentence from whole pieces

Every sentence agentSessionFailureWords writes is now assembled from a
table of whole English pieces, so a reader can supply its own words for
each piece. The host still fills them in English, byte for byte as
before.

* fix(native-chat): translate the failure sentences desktop notices show

A refused start and a rejected message now carry their failure fact to
the notice instead of its English sentence, and desktop words that fact
through translate keys whose English defaults are the host's own
pieces. The host keeps writing English into rows and reasons, and a
host sentence with no fact beside it still shows as written.

* fix(native-chat): the history pane says only that Orca keeps trying

When the pane's title already says Orca couldn't open this chat's
history right now, the line under it no longer repeats that the
transcript could not be read; it says only that Orca keeps trying to
load it. The pane with no named reason keeps its two-part line.

* test(native-chat): type the failure pieces a refusal notice shares

* test(native-chat): the Chinese failure words use no Japanese-only characters

* fix(native-chat): every history pane that says it didn't load says only that Orca keeps trying

A pane whose title is a code's own words ("This chat's history couldn't
be loaded.") now gets the short retrying line too. Only the pane with no
named reason keeps the two-part line.

* test(native-chat): a provider's words with nesting and markup stay as written in a translated notice

* fix(native-chat): French and Spanish say a withdrawn message was withdrawn before the agent began working on it

* fix(native-chat): Japanese and Chinese notices run their sentences on without a space

A notice joined its sentences with a space in every language, so Japanese and
Chinese read "Claude 无法启动。 请重新发送消息。" with a stray gap after the full
stop. The failure-sentence builder now takes the joiner alongside its words, and
desktop joins in the UI language: no space in Japanese and Chinese (including a
plugin pack that declares either), one space elsewhere. The host and the phone
keep English, joined with a space as before.

* fix(native-chat): a failed /clear or /compact says why in the app's language

The line under the composer printed the host's English sentence although the
result carries the typed failure beside it. It now words that failure the way
the host does (the chat's agent and /clear for a failed /clear, nothing for
/compact), in the app's language; an older host that sends no failure keeps
its sentence.

* fix(native-chat): Spanish says a rate limit, and Korean says a withdrawn message was never processed

The Spanish retry notice said the agent hit a usage limit, a different thing
from the rate limit the English names. The Korean withdrawn-message notice
said the agent had not started, which reads as the agent not launching; it now
says the agent had not begun processing the message, as the other languages do.

* refactor(native-chat): one rule says which words already say the history didn't load

* fix(native-chat): an image size limit says its unit the way the reader's language does

* test(native-chat): the option picker's i18n stand-in knows the reader's locale, which a refusal notice now reads

* fix(native-chat): a sentence a language pack left in English keeps its space

Sentences were joined by the UI language: no space in Japanese and Chinese,
one elsewhere. A plugin pack for a Chinese or Japanese variant that predates
the failure words falls back to English for them, so a notice read
"您的訊息未傳送。Claude couldn't start.Send your message to try again."

Each gap now follows the sentence before it: none after a full-width 。!?,
one space after anything else. The built-in Japanese and Chinese catalogs end
every sentence in 。, so they read as before, and English is unchanged. Since
the rule no longer needs the language, one joiner serves every surface and
the failure-sentence builder no longer takes one alongside its words.

* fix(native-chat): a failed /clear this build only partly understands shows the host's own sentence

The line under the composer words a failed /clear or /compact from the fact
the host sends beside its sentence. The reader drops any part this build
cannot place, such as a refusal code a newer host added, and the rest of the
fact can then give different advice: "Run /clear again." where the host said
"Start a new chat to continue."

When any part the host sent did not survive the read, the line now shows the
host's sentence as written, the same as for a host that sends no fact. A fact
this build reads whole is still worded in the app's language.

* fix(native-chat): a failure fact this build reads only in part shows the host's sentence everywhere

The previous check compared only a fact's top-level parts, so a known refusal
code carrying a reason a newer host added still counted as read: the reader
dropped the reason and the notice re-worded what was left, which can advise
differently from the host ("Run /clear again." against "Start a new chat to
continue.").

One shared reader now answers whether this build read the whole fact: it reads
the fact and keeps it only when the read equals what arrived, at every depth.
Every place that chooses between wording a fact and showing the host's text
uses it: the line under the composer after /clear or /compact, a rejected
message's notice from the journal's fact, and the smaller copy a rejected
message keeps for when its journal row is not loaded. Matching a rejected
message to the start row that already says why still uses what this build can
read, since that is identity, not wording. Facts this build reads whole are
worded as before.

* fix(native-chat): a failed /compact names /compact as its next step on desktop too

The host names the command a failed start was waiting on, for /clear and /compact alike.
The line under the composer re-worded only /clear with it, so a /compact whose start failed
read "The agent couldn't restart. Send your message to try again." in the reader's
language. It now words every command the host answers with the agent and the command the
host used, so desktop English matches the host and French or Japanese keep /compact.

* fix(native-chat): a /compact whose start failed no longer says the operation was not confirmed

A /compact on a chat whose agent is not running starts it first, and that start takes a new
lease, so the chat's fence moves before the command's reply arrives. The write settles as one
for a fence this pane no longer shows, and the composer line read that as "Conversation
operation was not confirmed." Such a write now says nothing there, as every other write
already does: the chat's own start-failure row says why, and the command's message is
rejected in the journal. The sentence it printed is gone from the catalogs.

* fix(native-chat): keep the message outbox within its line limit after main's growth

* fix(native-chat): a command's own reply is kept when its start moved the fence

A /clear or /compact on a chat whose agent is at rest starts the agent first, and that start
moves the chat's fence before the command's reply arrives. Every reply from an earlier fence was
discarded, so a /clear whose new chat failed to start said nothing at all, and a /compact that
started left "/compact" in the composer.

A conversation command's reply is now kept while the pane still shows the chat it was sent for;
a closed pane or another chat still drops it, and every other write keeps the fence rule. The
line under the composer says nothing only when the failure is the chat's own start and that
start's loaded row already says why, the rule a message that start rejected already follows.

* fix(native-chat): a returned queued message shows the host's sentence for a fact read in part

The caption under a returned queued message re-worded its failure from whatever this build could
read of the fact. A newer host's fact with a refusal code or reason this build drops read as a
shorter sentence with different advice. It now re-words only a fact read whole, and otherwise
shows the host's own sentence, as every other surface that re-words a fact does.

* test(native-chat): a command reply after any fence move, for /clear and /compact

The fence-move tests now state the rule as the code has it: a conversation command's reply is kept
whenever the pane still shows its chat, whatever moved the fence. They cover a /clear that
completed, and a failed start for each command in French with the fact that command really meets
(a /clear's new chat fails to start; a /compact's chat fails to restart).

* fix(native-chat): only the reply a pane still waits on outlives a fence move

A conversation command's reply was kept across a fence move whenever the pane still showed the
same chat. A reply the pane had stopped waiting on, because it left the chat and came back or a
newer command replaced it, was applied as if it answered the current one.

The pane now remembers the one command request it waits on; only that request's reply is kept
after the fence moves, and closing the pane or showing another chat forgets it. Tests now drive the
fence move during the request itself rather than through /clear starting an agent, and cover a
/clear that stops a running agent.

* fix(native-chat): a newer-Orca history error keeps a whole retry line, and the phone shows the host's words for a fact it reads in part

When a chat's history was saved by a newer Orca, the pane's title reads "Chats were saved by a newer
Orca. Update Orca to keep using them." and the line under it said only "Orca keeps trying to load
it.", with nothing for "it" to mean (in French and Spanish the pronoun also disagreed with "chats").
That title names every chat rather than this one, so the pane keeps the full line: "The transcript
could not be read. Orca keeps trying to load it."

On the phone, a returned queued message whose failure fact this build reads only in part was
re-worded from what it could read, dropping advice the host gave; it now shows the host's own
sentence, as the desktop card already does.

The comment on the reply the pane waits on now says what forgets it: a newer command, or disabling
the pane.

* fix(native-chat): a command reply this build can't place shows the host's own words

A newer host can answer a conversation command with a command name this build doesn't know. This
build re-worded that reply from its failure fact as if it were a command it knew, naming the
command in its own words, or said nothing when a loaded start row matched the fact. It now shows
the host's sentence as written, as it already does for a fact it can read only in part.
2026-09-30 14:15:07 -07:00
FenjuFuandJinjing 2caa79e097 fix(source-control): drop reasoning model think blocks from generated messages (#24005)
* fix(source-control): drop reasoning model think blocks from generated messages

Custom commit-message commands that run a reasoning model print the
reasoning before the answer, either as a <think>...</think> block or,
when the chat template prefills <think>, as text ending in a lone
</think>. The cleaner kept all of it, so the first reasoning line became
the commit subject.

Strip everything through the first </think> when the output starts with
<think> or has no <think> before the close tag. Output that quotes both
tags is left unchanged.

Fixes #24004

Signed-off-by: FenjuFu <fufenjupku@gmail.com>

* fix(source-control): scope lone </think> stripping to custom commands and add Kimi-VL tags

A lone closing tag is only stripped for custom commands, so a built-in
agent's message that mentions </think> is kept. PR fields parse the raw JSON
first and strip reasoning only when that fails, so a body quoting the tag
still parses. Adds Kimi-VL-Thinking's ◁think▷ tags.

---------

Signed-off-by: FenjuFu <fufenjupku@gmail.com>
Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com>
2026-09-30 12:12:21 -07:00
Brennan Benson cb363444f3 feat(native-chat): register Codex default-mode helpers as subagents (#22619)
* feat(native-chat): register Codex default-mode helpers as subagents

Codex's default multi-agent mode announces a helper only as the
collabAgentToolCall that spawned it; it sends no subAgentActivity. The
roster and the background-task tracker registered children only from
subAgentActivity, so such a helper had no record, no strip row and no
roster row, and its commands read as the session's own.

One announcement reader now yields a child from either wire shape, and
both the journal roster and the tracker register through it into the
same executions, so a session sending both keeps one child per thread.
A finished closeAgent ends the helper's running turn as stopped through
the executions, beside the child's own turn and thread frames.

Each collab call renders as a tool row (spawn_agent, wait_agent,
close_agent, ...) naming the helper the way the roster does, with what
the helper said back as its output, instead of the raw provider row.

* fix(native-chat): name a spawn row by its prompt until the roster holds its helper

Also pin that the roster row appears when the helper's first turn arrives
before the spawn call finishes.

* test(native-chat): a spawn call keeps its own row beside the roster row it starts

* fix(native-chat): pair a structured tool row's output with its own call

A run paired results to calls by position alone, so once one call finished
with no output (a Codex spawn_agent row) every later output drew under the
call before its own. A structured row carries its call and output together,
so the projected result now names its call id and pairing honors it,
falling back to position for results that name none.

* fix(native-chat): the Codex roster row follows the executions, so every child ending settles it

The subagent-group row was rewritten only on a child's turn/started and
turn/completed. A child turn ended any other way — a fatal error, its
thread closing, its caller closing it — settled the strip and the host
record through the executions but left the transcript row reading working
until the session ended.

The executions now say when a child's execution changes, and the roster
revises its row from that, whichever frame changed it. handleTurn still
re-derives to hand back its write admission; the revision is idempotent.

* fix(native-chat): a restored Codex call row names its helper, not its thread id

History replay never runs the live item router, so the roster never learned
the helpers a restored thread had spawned, and a restored wait_agent,
close_agent or send_input row labelled its helper with the raw thread id.
Replay now registers each announced helper for its name and membership only:
register starts no execution, so no strip entry, record or roster row claims
the helper runs until a live turn of its own says so.

The replayed-item handling moves into the restore module beside the replay
that calls it. Also name the roster's render contract in handleItem, and say
why the collab tool-name map is not a spelling fix.

* fix(native-chat): register a Codex helper whose spawn call failed but created its thread

Codex reports a spawn as failed when the helper it created errored at birth,
yet the call still names the thread it created, and that thread can run. The
spawn was registered only when the call completed, so such a helper existed
nowhere and its shell read as the session's own bare command: the original
default-mode bug. A spawn now registers whenever it names a receiver; one in
progress, or one that created nothing, names none. A failed closeAgent still
ends nothing, because the helper was not closed.

* refactor(native-chat): each Codex thread's latest token total gets its own home

Every thread's running total, member or not, with the rule that the newest
frame replaces the last and the recency-ordered cap, moves out of the roster
into codex-thread-token-totals.ts. The roster still selects its children's
totals at write time. With the producer linkage the roster now builds, it
was past the size limit.

* refactor(native-chat): the messages a session list quotes get their own module

The readers of a session's newest own prompt and own assistant prose move from
the status projection into structured-agent-session-latest-messages.ts, and
are re-exported so their consumers keep one import site. With the status
clock the projection now carries, a tool result naming its call put it past
the size limit.

* test(native-chat): a Codex helper's command is a live record until its process reports its exit

A helper's in-turn shell is a live command record owned by the helper and is also its open
Bash operation; its exit removes the record. A caller's closeAgent kills the helper's
processes, and each still reports its exit on the helper's thread, so the close leaves the
command to that exit rather than ending it itself.

* fix(native-chat): every reader of a tool run pairs a result with the call it names

The folded desktop run now pairs a result with the call it names, but mobile's tool run, the desktop edit cards and the task lists still paired by position through `pairToolBlocks`. In a Codex default-mode session the new `spawn_agent` row finishes with no output, so on mobile the helper's reply drew under `spawn_agent` and `wait_agent` showed none, and a desktop edit card could take the next command's output as its own.

`pairToolBlocks` now follows the same rule as `pairNativeChatToolResults`: a result that names its call answers that call, and one that names none answers the oldest unanswered call as before.

* fix(native-chat): a Codex helper's turn ends on its own frames, not on its caller's closeAgent

A finished `closeAgent` call ended the helper's running turn as `stopped`. But the call's status is not the close's outcome: Codex sets it from the helper's own agent status, so a close that errors on a running helper still reports `completed`. Orca then marked a helper that was still running as cancelled in the strip, the host record and the roster row, and because the first ending a turn gets stands, its real ending could never correct it.

A close that works does not need the edge either. Codex's shutdown of the helper aborts its running turn, and the app-server reports that as the helper's own `turn/completed` with status `interrupted`, which Orca already maps to `stopped`. So a helper's turn now ends only on its own frames (its `turn/completed`, a fatal `error`, `thread/closed`) or the session ending, and the close is only its call row.

This also drops the child-work evidence re-keying that existed only for the close: every remaining frame names the thread that sent it.

* fix(native-chat): a tool result that names a call is never given to a different call

A result that named a call id with no unanswered match fell back to the oldest unanswered call. On mobile, a run shows at most six calls, so a command past that window whose output named its own call gave that output to an earlier call that finished with none, such as `spawn_agent`.

A result that names its call now answers only that call and otherwise stays unpaired. A result that names no call still answers the oldest unanswered one. Both pairing readers share the one rule through `answeredToolCallIndex`.

* test(native-chat): name the Codex default-mode test file after the subagents it covers

* fix(native-chat): a Codex helper is known from any call that names it, not only its spawn

A helper whose spawn Orca never saw (history replay after compaction dropped the spawn, or a resume or message to a helper from earlier) stayed unregistered, so its shell read as the session's own command: the bug this PR fixes, in another shape.

Every collab call now announces each helper it names, skipping a receiver the finished call reports notFound. The spawn's prompt still labels a helper when it was seen; otherwise the helper has no label and reads as any unnamed subagent does. A send_input, like the spawn, names the parent turn the helper's next run belongs to.

The session's own thread is excluded once, in the reader, instead of at each of its three callers.

* fix(native-chat): every finished Codex collab call row carries what the call reports

A client that predates result call ids (an older mobile app, or an older desktop reading a newer host) projects the journal itself and pairs each tool result with the oldest unanswered call. This PR publishes collab calls as tool rows, and a finished spawn_agent, send_input, close_agent or resume_agent on a running helper had no output, so such a client drew each later output in the run under the call before its own (the parent's shell output under spawn_agent, the helper's reply under the shell).

Every finished call now has an output taken from the item: a wait's reply from each helper that finished (an errored helper's error), and for every other call the helper's reported status in Codex's own words (Pending init, Running, Completed, ...). A close or resume no longer shows the helper's last reply as if the call returned it. A wait whose end names no helper (it timed out, or v2) reads Finished waiting; any other call with no state reads its own status.

A call also keeps naming the helpers its started item named: Codex ends a timed-out wait with no receivers, so its row lost the helper's name when it finished.

* refactor(native-chat): the desktop run's tool pairing is a view of pairToolBlocks

Desktop runs and every other reader (mobile, edit cards, task lists, the ask row) each had their own pairing loop sharing one index rule. The desktop's pairNativeChatToolResults now reads the pairs pairToolBlocks makes, so a run is paired by one loop and a new reader cannot add a third. No behaviour change.

* test(native-chat): type the positional-client pairs without assertions

* fix(native-chat): a finished Codex collab call row says what the call did, in Codex's own words

A close_agent row read `Running` and a spawn_agent row `Pending init`: each showed the helper's status snapshot from before the call took effect, so a close that stopped its helper read as though it had not worked.

Each finished call's output now follows Codex's own client: spawn reads `Spawned` (or `Agent spawn failed` when no helper was created), send_input `Sent input`, close `Closed`, and resume the helper's status summary. A wait keeps each finished helper's reply, with other statuses in Codex's summary wording (`Error - <message>`). The row label already names the helper, so the output does not repeat it.

* docs(native-chat): the collab call reader says which calls show the helper's snapshot as output

Since spawn, send_input and close rows say what the call did, the reader's header was wrong to claim every call row shows the reported snapshot as its output: only a wait or resume does.

* fix(native-chat): a Codex spawn call no longer claims it ran an agent

Codex's spawn_agent call ends as soon as the helper starts, so counting it as running an agent drew "Ran 1 agent" directly above "Kicked off 1 subagent · working" while the helper was still working, and "Ran 1 agent · 1 failed" when Stop cancelled a spawn before any helper existed. Claude's Task call lasts as long as its subagent, so it keeps the agent category; the Codex spawn row is now a plain tool call and the roster row alone stands for the helper.

* test(native-chat): a Codex helper's row settles on the failed completion that follows a fatal error

Codex follows every turn-ending `error` with the turn's failed `turn/completed`,
on a helper's thread as on the primary, and only that completion ends the turn.
The roster test now sends both and checks the row, strip and record stay
working through the error and settle on the completion.

* fix(native-chat): a Codex helper's section opens while the parent spawns, waits on or messages it

A running chat holds a subagent's section open while the parent's newest row delegates to that subagent. For Codex that rule only knew the raw collab status row; this PR writes each collab call as a tool row (spawn_agent, wait_agent, send_input, close_agent, ...) naming its helpers in `input.agents`, so a default-mode helper's section stayed shut for its whole run while Claude's opened.

The delegation reader now reads a Codex collab tool row as a delegation to the first helper it names, the same rule the raw row keeps for journals written before this change; a call naming no helper stays ordinary output. The row names and the `agents` key live in one shared module the host writes from and the reader reads, instead of a second list.

The new test drives the real adapter's rows through the transcript projection to the delegation; the collab frame harness moves to a fixture shared with the positional-clients test.

* fix(native-chat): a Codex collab call with a long prompt still names its helper

A collab call row bounded its whole input as one value, so a spawn or send_input whose prompt passed the 16 KB journal limit was stored as a clipped wrapper: the row lost the helper's name and thread ids, and with them the label and the delegation that opens the helper's section.

The prompt is now clipped on its own with the journal's inline-text bound, marker included; the helper's name, ids, model and effort are bounded as before and always survive.
2026-09-30 11:36:57 -07:00
Neil db73d28e51 test: retire backlog cases that assert a shim, a literal, or an unread branch (#24120)
Deep-reads the 596 files earlier auditors explicitly disclosed as reviewed at
title-and-import level only, never against production. Six ~100-file chunks, chosen so
depth was achievable rather than optional. 27 case declarations removed across 19 files,
1 test file deleted, 522 lines gone. No production code touched.

Two chunks found nothing, and that is reported as the result rather than padded:
`src/main/agent-hooks` (89 files, 50 read case-by-case in full) returned zero deletions;
`src/main/claude` (90 files, 44 read in full) found three near-identical pairs via a
normalized case-body hash and kept all three after reading them.

What went:

- Identity copiers with a type-level title. `ui-new-workspace-draft.test.ts` (deleted, 3
  cases) tested `setNewWorkspaceDraft: (draft) => set({ newWorkspaceDraft: draft })` — a
  bare pass-through — by passing a literal in and asserting a subset of that literal back.
  Its titles named field-shape facts that the typed parameter at
  `ui-slice-contract-core.ts:197` already enforces in production, so the annotation that
  does the work lives in production, not the test.
- A test that asserted its own mock. `destroyRemovedBrowserWebview(id)` is literally
  `destroyPersistentWebview(id)`, and the case mocked `destroyPersistentWebview` — so it
  checked that the mock received the argument the shim passed through.
- Replays across a bare re-export. `pane-tree-ops.ts:18` is
  `export { equalizePaneSplitSizes, findPaneChildren } from './pane-tree-equalization'`;
  two cases used `MockHTMLElement` with no dividers while the owner asserts concrete flex
  values on real happy-dom elements and runs a 400-seed differential against an
  independently reimplemented weight walk.
- Table rows varying a value production never branches on.
  `mergeCurrentOrchestrationContext` tests only `dispatchStatus !== undefined`, so
  `it.each(['failed','circuit_broken'])` ran the path the surviving `'completed'` case
  already proves. Per-value behavior is owned by the one consumer that reads those
  literals.
- Duplicate invocations whose surviving sibling asserts strictly more, including a push
  error case differing from its neighbour only in which verb produced the same error
  string.

Kept deliberately: bound and cap guards (a 50-entry nav-history cap, a skill-cache
eviction pinned to size 2 after 512 retired runtimes, a cleanup-concurrency ceiling);
`buildDefaultTerminalOptions` cases restating declared constants, because each records a
cross-cutting UX decision whose comment preserves the v1.4.51 ZWJ table-corruption history
that once forced scrollbar width 0; three per-grammar `tokenizer.root` invariants for
astro, svelte and vue, which are separate grammars rather than replays; and a test reading
`@xterm/{headless,xterm}/package.json` out of node_modules to prove both were built from
the same upstream commit, which is legitimate because the installed package is the shipped
contract.

Six `src/shared` files were trimmed, and each was checked against the risk that matters
there: `src/shared` is the OWNER four earlier waves deferred to when deleting
main/renderer/mobile tests, so hollowing one out would orphan several callers at once.
Each retains 9 to 46 cases after removing 1 or 2.

Coverage is partial and stated as such. Read case-by-case: pane-manager 81 of 81, slices-a
96 of 121, slices-b 94 of 121, agent-hooks 50 of 89, claude 44 of 90. Every auditor listed
the paths it did not reach.

Verified: 1,294 test files / 14,257 cases pass across the touched areas, plus one
pre-existing `it.fails` marker; `check-reliability-gates.mjs` 140 gates; the deleted file
is absent from the gate manifest, `cloud/package.json` and
`mobile/tests-typecheck-baseline.txt`.
2026-09-30 03:59:58 -07:00
Neil a781a602a8 test: retire duplicate cases that replay an owner across a re-export or provider shim (#24114)
Resolves 208 candidate pairs where the same case title appears verbatim in two or more
files, produced by a repo-wide scan calibrated against a known positive. 46 case
declarations removed across 32 files, 798 lines gone. No file deleted whole, no
production code touched.

The headline result is the measurement, not the deletions: across the three buckets that
reported in detail, the signal ran roughly 86% false-positive (3/42, 9/42, and the rest).
It has good recall and poor precision, and it reorders a reading queue rather than
replacing one. Calibrating a detector against a known positive proves recall, not
precision.

What the deletions were:

- Duplicate invocation through a re-export shim. `native-chat-tool-summary.ts` is a
  ten-line `export {...} from '../../../../shared/native-chat-tool-summary'`, and
  `agent-status.ts:161` is `export { isExplicitAgentStatusFresh } from
  './pane-agent-evidence'`. Cases on the shim side were byte-equivalent to the owner's
  with no rendering or transport hop.
- Provider-local replays of a shared helper: three `repository-ref` providers that are
  each `createRemoteRefProbeCache(parseXRef)` and contribute nothing to transient
  handling; two `local-pty` and `daemon/session` tables replaying
  `shell-startup-output-scanner`, whose owner additionally checks every split point.
- A reader-side replay of store policy. `runtime-worktree-agent-rows-structured.test.ts`
  asserted an attention-to-blocked mapping; the reader contains zero `attention` or
  `blocked` tokens and copies `state` through. The mapping lives in
  `structuredAgentSessionAgentStatus`. Consistent with
  `docs/reference/agent-status-store.md`: readers keep only presentation policy.
- Constructor-only subclass duplication: the shared capability-cache case is covered by
  `codex-app-server-capability-cache.test.ts`, whose ten cases include the identical
  title plus all four risks `docs/reference/git-compatibility.md` names — first fallback,
  later cached call, concurrent probes, per-host isolation.
- A private predicate duplicated at a real boundary, varying only a path passed straight
  into the shared predicate.

Why most pairs were KEPT, because the false positives are principled rather than noise:

- Two independent execution hosts. `src/relay/git-handler-*` and `src/main/git/*` are
  separate Git implementations that cannot import each other and hold separate capability
  caches, exactly as the compatibility doc requires; the repo already ships
  `status-branch-line-total-relay-parity.test.ts` to pin the duality deliberately. Neither
  side's argv, timeout or cache regression is visible to the other.
- Deliberately duplicated production siblings: Codex vs Claude (different account fields,
  different CLIs, different wire protocols), gitea vs bitbucket (`/pulls/42` vs
  `/pullrequests/42`), gl-utils vs gh-utils (separate in-flight maps). Same contract
  shape, different implementations — an identical title is the correct naming.
- Shared-predicate consumers: one side tests the predicate, the other tests a caller's
  wiring to it. A caller that forgot to call the predicate passes the shared test.

In a codebase with intentional provider and host symmetry, identical test titles are
expected, and the signal cannot distinguish "copied" from "parallel by design" because
both produce the same prose. Only reading both bodies separates them.

Verified: 6,968 desktop test files pass; the three modified mobile files pass (39 cases);
`check-reliability-gates.mjs` 140 gates; nothing under
`mobile/src/test-support/rpc-recording/` or `mobile/rpc-foundation/goldens/` touched.

62 local failures across 12 files were each accounted for and none is caused by this
change: `browser-manager-tab-identity`, `browser-manager-viewport-ownership`,
`session-scanner-codex-workers` and `managed-hook-script-refresh` all fail identically in
a pristine `origin/main` worktree; five `mobile-web-app-*-render` tests need Playwright
browsers this machine lacks; `structured-agent-session-restart-ownership` and
`ssh-remote-commands` pass in isolation and fail only under concurrent load.
2026-09-30 03:30:23 -07:00
Neil cef66fbab8 test: retire long-tail cases whose input cannot reach the behavior they name (#24101)
Sweeps the triage-only backlog: 2,269 files that earlier waves saw and skipped for
size, reconstructed from the unread lists five waves of auditors disclosed. 35 case
declarations removed across 15 files, 2 test files deleted, 487 lines gone. No
production file touched.

These are large integration suites, so the junk here is individual cases buried among
real coverage rather than whole bad files. The dominant defect was again a case whose
input cannot reach the behavior its title names:

- `resume-sleeping-agent-session-remote-compat.test.ts` (deleted) — two cases titled
  for "transport-level host authority on a capable host" and "host authority is not
  known". `resume-sleeping-agent-session.ts` has no host-authority or capability
  concept at all, and its only read of `origin` is
  `if (!record.origin && record.state === 'done')`, unreachable for both rows. Both
  executed one identical path. The surviving contract is owned by
  `resume-sleeping-agent-session-execution-host-scope.test.ts`, which drives a real
  host catalog.
- `project-group-header-drag.test.ts` (deleted) — four cases setting
  `data-project-group-header-id`, which the predicate never reads. Its subject,
  `isProjectGroupHeaderActionTarget`, is byte-identical to `isRepoHeaderActionTarget`
  apart from the function name and imports the same `REPO_HEADER_ACTION_SELECTOR`, so
  all four cases were a strict subset of `project-header-drag.test.ts` using identical
  `data-repo-header-*` fixtures.
- `remote-worktree-history-cleanup.test.ts` — "repeats idempotent cleanup through the
  PTY owner" against a six-line best-effort forward with zero dedupe state. The case
  called it twice and asserted the mock recorded two calls, which is arithmetic over
  the test's own loop; nothing about idempotence was established.

Also removed:

- Runtime assertions of type-level facts, where production already makes the check at
  a stronger boundary: `const adapterSatisfiesPort: AdapterIsPort = true` followed by
  `expect(...).toBe(true)` — unconditionally true, while
  `createExpoGenerationFileSystem(): GenerationFileSystem` is explicitly annotated and
  passed into `createGenerationStore` at a typed call site. And a case named "does not
  typecheck" whose runtime assertion is a length check on its own literal, declaring
  its own local annotation so it could never notice the production annotation
  weakening.
- Private predicate tests duplicated at a real boundary: four `repo-slug-cache` cases
  delivered by `repo-slug-index.test.ts`, which drives the same resolution through the
  hook, the real store and the preload bridge, while the cache-level versions hand-seed
  the internal map and break on a cache-key format change.
- Duplicate invocations owned at the shared boundary, including commit and push
  recovery cases owned by `src/shared/source-control-recovery-agent-command.test.ts`.

Kept deliberately, verified rather than assumed: the production duplication behind the
deleted drag test was left alone, because `REPO_HEADER_ACTION_SELECTOR` ends in generic
`button, a, input, textarea, select`, so genuine action targets inside a group header
still match — it is an unspecialised copy-paste, not a live bug, and collapsing two
functions is a refactor. Reported instead.

Auditors' probes produced 20, 11 and 13 candidate hits for the signature-versus-title
shape across their chunks; every one was inspected and every one was genuine coverage.
No deletion in this wave rests on a probe alone.

Coverage is partial and stated as such: of 2,269 files, roughly 100 were read
case-by-case and the remainder reviewed at title-plus-import level. Each auditor listed
its own unread set. The largest remaining surfaces are `src/main/agent-hooks` (95),
`src/main/claude` (100), `src/renderer/src/lib/pane-manager` (62) and the 20 largest
sidebar suites.

Verified: 2,583 desktop test files / 25,636 cases pass, plus one pre-existing
`it.fails` marker; the two modified mobile files pass (57 cases);
`check-reliability-gates.mjs` 140 gates; `check:code-quality:changed` 0 new findings.
Both deleted files confirmed absent from the gate manifest, `cloud/package.json` and
`mobile/tests-typecheck-baseline.txt`.
2026-09-30 02:45:32 -07:00
Neil 5fa290fa50 feat(secrets): warn in Settings when a credential is stored unencrypted (#24048)
* feat(secrets): warn in Settings when a credential is stored unencrypted

When no OS keyring is usable, the MiniMax stores write the credential as a
plaintext envelope and say so with a console.warn nobody reads. The users this
affects are exactly the ones who never see a main-process log, so in practice
they were told nothing (#21827).

Report it where the credential is managed instead. Each store gains a
protection reader, the status IPC carries it, and Settings renders a warning
next to the credential it applies to.

Keyed on the stored bytes, not isEncryptionAvailable(): a credential saved
before a keyring existed stays plaintext until it is saved again, so reporting
current capability would call it protected while the file says otherwise. The
readers parse the envelope kind without decrypting, so opening Settings cannot
provoke a keychain prompt.

The console.warn stays. It carries no secret material, and it is still the only
signal on a headless host with no Settings window.

* feat(secrets): extend the unsealed-credential warning to every affected store

The speech key, Linear tokens, Jira tokens and the Bitbucket credential have
the same plaintext fallback the MiniMax stores do, and the same console-only
warning nobody reads.

Add a shared `readCredentialFileProtection` for the four stores that write bare
ciphertext with no envelope, classifying with the same printable-UTF-8 test
`readStoredCredentialToken` already uses — so the reporter cannot drift into
disagreeing with the reader about the same bytes.

Linear and Jira report across every stored workspace/site rather than the
active one: sealing is a host-wide property, so a second workspace stored while
the keyring was missing is exposed even when the active one is sealed. Both
fields are optional, so an older remote host that omits them reads as unknown
rather than as sealed. Bitbucket reports null for env-supplied auth, where Orca
stores nothing and has no claim to make.

Also fixes the credential-connection test double, whose identity-function
`encryptString` wrote a readable token — faithful enough for a round-trip
assertion, but it made the suite assert that a sealed credential was exposed.

* chore(i18n): extract the unsealed-credential notice strings

CI's localization-extraction gate requires every translate() key to exist in
the primary catalog. Inserted in place rather than re-sorting the file, which
is not fully sorted and would have produced a 17k-line diff.

* test(web): pin the null protection fields on the desktop-only MiniMax bridge

The web bridge reports no protection because it stores nothing; the shape
assertions had to move with it.
2026-09-30 02:13:03 -07:00
NeilandBrian Grablin f4092c06d6 fix(runtime): treat Hermes session start as idle, not a running turn (#24064)
Hermes fires `on_session_start` when a session is opened, switched or reset;
its payload carries no turn. Orca mapped it to `working`, so a freshly launched
Hermes held a fresh first-party `working` row for the whole 30-minute staleness
window: `terminal wait --for tui-idle` never settled against an idle composer,
and the sidebar span a phantom spinner.

Map it to `done` with `sessionBoundary`, matching what Claude and the
compatible-lifecycle providers already do for SessionStart, and let the tui-idle
evidence lane settle on a fresh boundary row. A boundary row claims a new
session owns the pane and awaits its first input, which cannot arrive mid-turn,
so it carries none of the #6011 risk that scoped that lane to DSH; a turn-end
`done` still does not settle.

Fixes #13653.

Co-authored-by: Brian Grablin <5216789+bgrablin@users.noreply.github.com>
2026-09-30 01:25:50 -07:00
Brennan Benson 59c05d32b0 fix(rate-limits): stop driving a hidden Codex TUI to read usage (#23806)
* fix(rate-limits): stop driving a hidden Codex TUI to read usage

When the headless Codex usage call failed, Orca opened a hidden
interactive Codex, typed /status and pressed Enter without reading the
screen, then killed it after 15 s. If Codex showed its "Update available"
prompt, that Enter picked "Update now", Codex started its installer, and
the 15 s kill interrupted it, leaving the global install broken.

Drop the hidden-terminal fallback. When the headless call fails with a
non-sign-in error, read the same usage from the HTTP endpoint Orca
already calls on WSL and for the 5-hour window, so a usage check can no
longer answer any Codex startup screen.

Fixes #17415

* test(rate-limits): pin the RPC error when the Codex HTTP fallback fails

Also drop comments that still described the removed hidden-terminal fallback.

* test(rate-limits): use real Response objects in Codex fetcher tests

The changed-code gate rejects the new type assertions this PR added.
2026-09-30 01:04:43 -07:00
Brennan BensonandHarshul Rathod 707d3dc96b fix(chat): decode Claude pastes and report terminal delivery uncertainty (#23788)
* fix(chat): decode Claude pastes and track terminal delivery uncertainty

Keep queued prompts pending while the existing agent status reports work, and check fresh history after a later idle fact. Preserve draft text and distinguish write rejection from unconfirmed delivery.

Co-authored-by: Harshul Rathod <harshulrathod1640@gmail.com>

* fix(native-chat): break the observed-send import cycle and keep renderer tests out of main

The observed-send path imported the clear helpers from native-chat-runtime-send,
which imports it back. Move the input-clear layer into its own module both use.
The Claude paste decoder test imported the renderer pending module from
src/main, which the node typecheck project cannot see; the echo-retirement
assertions now live in a renderer test.

* fix(native-chat): still submit a Claude chat send whose write acknowledgment was lost

A remote write whose acknowledgment is lost (timeout, dropped link) is not a
refusal, but the observed path stopped there and never sent Enter, leaving a
body that did land sitting unsubmitted in Claude's input line until the next
send's clear wiped it. Continue to the next write without re-sending the bytes,
as the unobserved path always did.

* perf(native-chat): keep terminal Chat pending delivery from re-rendering every row

The delivery notices were merged into a new Map on every render, which
invalidated the transcript row context and re-rendered every memoized row on
each stream update. The pending hook also wrote a fresh array on every status
ping and prune pass even when nothing changed, and the phone mapped its pending
list on every render, rebuilding the chat list data. Memoize the merged
notices, skip no-op pending writes, and memoize the phone's rendered pending
list. The phone also skips a transcript read when no send is due.

* fix(native-chat): never flag a queued Claude send, and flag one an idle Claude never starts

Two gaps in when terminal Chat calls a Claude send "Delivery unconfirmed":

A prompt sent while Claude is mid-turn is queued, and Claude folds it into the
running turn as a queued-command record. The transcript reader drops those
records, so once the turn ended the prompt Claude did run read as unconfirmed,
inviting a duplicate resend. A send made while the agent is busy is now never
checked; it keeps the pending behaviour it had before.

A prompt sent to an idle Claude that never starts a turn (Claude exited to the
shell, or the paste went nowhere) left the status at the same idle fact
forever, so the check never ran and the bubble stayed pending. An idle agent
starts a turn on a delivered prompt at once, so a send whose idle status is
unchanged after the existing 20 s bound is now checked against a fresh
transcript read.

* fix(native-chat): add the delivery notice strings to the English catalog

The Dismiss action's translate key was missing from en.json, which fails the
localization catalog and extraction gates. The desktop "Message not sent" and
"Delivery unconfirmed" notices were hard-coded English; route them through
translate with the same wording.

* fix(mobile): sync the held-send refs after commit instead of during render

Moving the acknowledgment-loss hold into its own hook made its render-time ref
writes new lines, which the React Doctor changed-lines gate blocks. Held sends
report after commit, so syncing those refs in a layout effect keeps them
current where they are read.

* fix(native-chat): report only definite terminal Chat send outcomes

The delivery rule inferred "Delivery unconfirmed" from "the turn ended and the
transcript has no matching row". Claude records a prompt sent mid-turn only as a
queued-command attachment, which the transcript reader drops, so that rule
flagged prompts Claude had answered. It also never fired for an idle Claude that
lost the write, because no newer turn arrives.

Keep only facts the transport reports:
- a refused write reads "Message not sent", keeps its text, and can be dismissed;
- a lost write acknowledgment holds the echo for 20 s, the phone's existing
  rule, then reads "Delivery unconfirmed" unless its row has landed.
An ordinary send, including one Claude queues mid-turn, stays pending as before.

Remove the agent-status subscription, the status-epoch origin, the fresh
500-row transcript read, the confirmed state and the no-status clock. The phone
already implements this rule, so its changes revert to main; only a test for
old-host paste envelopes remains.

* fix(i18n): translate the terminal Chat delivery notices

Add the Dismiss, "Message not sent" and "Delivery unconfirmed" strings to the
es, fr, ja, ko and zh catalogs, reusing each catalog's existing Dismiss wording.

* fix(native-chat): let a resend replace its failed terminal Chat echo

A "Message not sent" or "Delivery unconfirmed" echo kept its transcript
occurrence, so resending the same text numbered the resend as the second
copy: the one landed row retired the failed echo and pinned the resend
below the reply forever. Appending a send now drops a failed echo with the
same content first.

* fix(native-chat): unwrap a Claude paste that quotes pasted_content tags

The envelope parser refused any body containing a pasted_content tag, so a
pasted prompt that itself quotes one (a transcript excerpt, or code that
handles these tags) kept its wrapper and its echo stayed pinned below the
reply. Claude's per-paste id exists to disambiguate exactly that; only a
same-id tag inside the body is now ambiguous. Wrappers without an id keep
the strict rule.

* test(native-chat): pin which terminal Chat sends observe write outcomes

Only a Claude chat send (text or images) reports a refused or unacknowledged
write to its pending echo; other agents and slash commands keep the
unobserved write path exactly as before.

* fix(native-chat): keep failed terminal Chat sends through Stop

Stop cleared every optimistic echo, including a "Message not sent" or
"Delivery unconfirmed" bubble whose send had already settled. Stop cannot
affect that send, and the bubble is the only place its text stays copyable,
so it now survives until the user dismisses or resends it. Also moves the
observed-send import below the file header comment.

---------

Co-authored-by: Harshul Rathod <harshulrathod1640@gmail.com>
2026-09-30 01:04:21 -07:00
5b93c6216a Fix Chat UI paste intake and pane routing (#23784)
* fix(chat): separate text paste from attachments and route by pane

Keep composer text independent of image checks and saving, and route pastes
caught underneath chat to the originating pane's mounted input. Preserve
native event data, selection replacement, undo, and target lifetime checks.

Co-authored-by: Wooseong Kim <innocarpe@gmail.com>
Co-authored-by: lurunzi <lurunzi@gmail.com>

* fix(chat): keep focus and quiet text paste after routing it to chat

- A paste inserted into the composer now moves focus there, as the old
  menu-paste insert did; otherwise a paste routed from the hidden terminal
  left the next keystrokes going to that terminal.
- With a remote-server or not-ready workspace, pasted text no longer shows
  the "Local attachments are not available" refusal because the clipboard
  also held an image rendition (common for Office copies). The menu path
  probes for an image only when the text read is empty, so a paired browser
  does one permission-gated clipboard read for a text paste, not two.
- Latest-value refs update in a layout effect instead of during render.

* perf(clipboard): answer "is there an image?" from the format list

The chat composer asks the main process whether the clipboard holds an
image before explaining an image-only paste on a remote-server workspace.
That probe decoded the whole image (readImage().isEmpty()) on the main
thread just to return a boolean. Read clipboard.availableFormats() instead,
and share the MIME check with the paired-web probe.

* fix(chat): a chat cover owns focus, so input never reaches the hidden terminal

When a Chat UI tab opened over its terminal, the terminal's xterm kept
keyboard focus until the composer claimed it a frame later, and forever if
the composer never became ready (still starting, a question card, a phone
holding input). The previous commits rerouted paste from that hidden
terminal to the chat, but typing and Enter still went to the terminal, an
image-only or refused paste left focus there, and about twenty
terminal.focus() call sites could put it back.

Make "a covered terminal cannot hold focus" structural instead:
- The chat cover takes focus in the commit that mounts it and marks the
  covered xterm inert, so every terminal.focus() path is refused by the
  browser. Split siblings are untouched. When the chat goes away the xterm is
  un-inerted, and gets focus back only if focus was inside that chat.
- Terminal paste listeners skip anything inside a chat cover (previously only
  inside a mounted chat root). The reroute from terminal to chat is gone;
  terminal-only paste is back to main's code.
- A paste that finds no chat input (before the chat mounts, or an approval
  card with no text field) gets a visible refusal from the cover. A disabled
  composer shows the same notice inline instead of dropping the paste. New
  copy: "Can't paste — this chat isn't accepting input right now." (the old
  "Worktree not ready" toast was wrong for a chat that is still starting).
- The terminal context menu, which names its pane, keeps a small request
  event to that pane's chat, now without a clipboard payload and using the
  existing covered-pane check.
- Cmd/Ctrl+V or Shift+Insert on a non-input part of the chat focuses the
  composer (or question answer) first, so the paste lands there.
- The composer-scope check used to decide whether a text field inside the
  chat keeps its own paste matched the whole pane (the file-drop surface
  carries the same attribute). It now asks the composer whether the target
  is inside its input.

* fix(chat): don't paste a copied file's name next to the file

Copying a file in Finder or another file manager puts its name on the
clipboard as text/plain beside the file itself. Since text and images are
now pasted independently, pasting such a copy into a local or SSH chat
inserted the file name into the prompt as well as attaching the image.

On the paste-event path, text/plain that is exactly the names of the pasted
files (one per line) is the file's label, not prompt text, so it is dropped
when an image from that paste is being attached. Rich-text copies (text plus
an image rendition) still insert their text, and a copied non-image file,
which is not attached, still pastes its name as before.

* fix(chat): don't type a Finder file's name on Cmd+V either

On macOS, Cmd+V in the chat goes through the app-menu paste, which reads the
clipboard text and saves the clipboard image separately. A file copied in
Finder also puts its name on the clipboard as text, so the composer typed
the name next to the attachment. On main the menu path never read text once
an image saved.

The main process now reports the paths of the files a file manager copied
(macOS filenames plist or file URL, Explorer's FileNameW, a Linux uri-list).
Text that only labels those files waits for the image outcome: dropped when
an image is attached (or refused on a remote owner), typed when none came.
The same label rule now also accepts a path or file URL per line, which is
how Linux file managers label copied files on the paste-event path.

* fix(chat): pane focus aimed at a chat lands on the chat

Since the covered terminal became inert, focusing a pane that shows a chat
(keyboard pane navigation, focus-follows-mouse, split activation) was refused
and focus stayed on the pane the user left, so typing went to that visible
sibling terminal. The one place a pane's focus is requested now puts it on the
pane's chat cover, which hands it to the composer when the pane is revealed.
Focus already inside the chat is left alone.

* test(terminal): give fake panes the container pane focus now reads

Pane focus checks the pane's container for a chat cover, and these two
fixtures built panes with only a terminal, so four tests threw.

* refactor(native-chat): move composer paste handle and chat-root key routing into their own modules

Brings NativeChatComposer.tsx and NativeChatResolvedView.tsx back under the
400-line limit after merging main. No behavior change.

---------

Co-authored-by: Wooseong Kim <innocarpe@gmail.com>
Co-authored-by: lurunzi <lurunzi@gmail.com>
2026-09-30 01:03:25 -07:00
Brennan Benson ca7c14db08 fix(mobile): start + menu, quick command and diff-note agents through agent.launch (#22954)
* fix(mobile): start + menu, quick command and diff-note agents through agent.launch

The session screen's + menu, agent quick commands and diff notes' New agent
session now ask the host to start the agent with agent.launchReplay, so the
host picks chat or terminal from the desktop's default and delivers any
prompt. Hosts without the launch capabilities keep today's paths.

The phone's pending tab choice is one value (a tab, a terminal by handle, or a
launched surface) instead of two refs, and a launched chat is found by its
session id in the next snapshot rather than a predicted tab id. A launched
surface waits a bounded number of snapshots for its tab.

* test(mobile): add the + menu and diff-note launch scenarios to the recording corpus

* test(mobile): repin bridged-parity tallies for the four launch goldens; drop test casts

The corpus grows from 790 to 794 goldens; all four new ones replay identically.

* fix(mobile): show a refused agent launch as a toast beside open tabs

The inline create error renders only in an empty session, so a host refusal
(for example a disabled agent) from the + menu in a session with tabs showed
nothing. Always toast the failure: the caller's own copy when it gave one,
otherwise the host's reason.

* test(mobile): type the launch reply helper with the shared launch outcome types

* fix(mobile): record a launched agent's tab as this device's pick on the host

A launch carries no navigation, so the phone selected the new tab only
locally while the host kept this device on the tab it had before. Leaving
the session and coming back, or a reconnect that reset the screen, reopened
that old tab. The "+" terminal path this replaced asked the host to select
the tab for the caller.

When a launched surface's tab lands in a snapshot, activate it for the
caller exactly as a tap does. The resolver now names the landed tab in
place of the unused `missed` flag. Route parity re-pinned for the new
activation body, identity payload and strings.

* fix(mobile): land on a launched agent's tab without a 500 ms wait or a blank pane

The host publishes a launched tab before it replies, so the tab list the
phone already holds usually has it by the time the reply arrives. The
launch paths still waited for a refetch 500 ms later, leaving the phone on
the old tab for that long after every launch. Read the tab list at once.

On hosts without agent.launch, the chat path also unsubscribed the open
terminal and cleared its handle before the chat's tab landed, while the old
terminal tab stayed selected: a blank pane until the next tab list. Leave the
open tab live until the chat lands, as the launch path does; applying that
tab list tears the old terminal down.

Route parity re-pinned for the two bodies.

* fix(mobile): keep a tab the user picked while a prompted launch was still replying

A quick command or review-notes launch now waits for the host to deliver the prompt, which can take up to a minute. The launched tab shows up in the tab row well before that, so a user who tapped another tab meanwhile was pulled back onto the launched one when the reply arrived, and that pick was recorded on the host.

The launch now remembers which tab the phone was on when it started and only takes focus if the phone is still there when the reply lands. Any move made in between, by a tap or by the computer navigating this phone, wins. Session route parity re-pinned for the handleCreateTerminal body only.

* fix(mobile): name a launched agent's tab before asking, and land on it when it is listed

A "+" menu, quick-command or review-notes launch now reserves its tab before it asks the host: a fresh pane key (tab and leaf UUIDs) and, for an agent the host may start as a chat, a session id. Both are minted once per launch and sent unchanged on every replay, since the host's replay fingerprint covers them.

The phone arms its pending selection with that reservation before sending, so it lands on the terminal (matched by pane halves) or chat (matched by session id) as soon as the tab is listed. For an agent whose prompt is pasted after start, that is long before the reply, which waits for delivery. Landing also frees the "+" lock; the lock holds the create's id, so an older launch's reply cannot free a newer one's. The reply now only adds its own handle or session id (an older host ignores the reservation), starts the fallback countdown, and reports prompt delivery.

A tab the user picks mid-launch replaces the pending selection, so the launch-start tab check is gone. A reservation the host refuses as already taken reads "Couldn't start the agent. Try again." on the first send, and as unconfirmed after a replay. The mobile UUID fallback now yields a v4 UUID, because a pane key's leaf must be one. The host launch path moved to new-tab-agent-host-launch.ts; session route parity re-pinned for that move and the landing's lock release.

* test(mobile): expect the launch reservation in the four launch scenarios

The four launch scenarios now expect the pane key and session id the phone sends (the scripted ids come first, so the operation id moves from ...001 to ...004).

* fix(mobile): don't say an agent may not have started while the user is looking at it

When a launch's reply was lost after its tab had already landed, the phone said "Couldn't confirm the agent started", although the listed tab proves it did. Now a listed tab narrows the doubt to the prompt or notes ("The agent started, but couldn't confirm the notes were sent."), the notes stay unsent, and a bare launch says nothing. Only the nested-function parity pin moves, for handleCreateTerminal passing the tab list to the launch.

* test(mobile): read the launch's sent reservation through the host's params schema

The anti-slop audit rejects Reflect.get; parsing with AgentLaunchReplay also
asserts the host accepts the params the phone sent.

* test(mobile): check the launch reservation against the host without importing its schema

Mobile code may import the params contract only as types. The phone's tests
now read the sent reservation by narrowing, a chat reservation is checked
through the real host dispatcher, and the older-host drop is pinned host-side.

* fix(mobile): don't send the same review notes to a second new agent

The "+" lock is now freed when the launched tab lands, but review notes are
only cleared when the launch's reply confirms delivery, which for a prompted
launch can take up to a minute. In that window "Send review notes to AI" still
offered the same notes, and choosing a new agent session started a second
agent with them.

The notes a new agent session is being started with are now held from the tap
until that launch settles: the Send button no longer counts them, the sheet no
longer offers them, and a stale tap on the old sheet starts nothing. Notes the
host did not deliver become sendable again once the reply arrives.

* test(mobile): record the + menu and diff-note launch goldens

4 added (+ as a terminal, + as a chat, notes delivered, notes not delivered). 10 existing create-terminal goldens move only because the recorded state now shows one pending selection instead of two refs; their requests are unchanged.
2026-09-30 00:42:52 -07:00
Brennan Benson fceca5cece fix(sidebar): an agent's row stays while it runs, whatever its tab title (#23948)
* fix(sidebar): keep hook-less agent rows while the agent runs, whatever its title

Codex retitles its pane to the project name, so the sidebar's title-derived
row (which required the title to name an agent) vanished while Codex kept
running (#23767). Rows now take identity from the canonical pane resolver
over the pane's foreground-process read and launch record, then the title;
the title only decides idle/working/needs-input. The row still goes away when
the PTY exits, the process tracker proves the shell is back, or the title is a
shell or default title.

* test(dashboard): justify the partial store fixture's type assertion

* fix(sidebar): only a live process read keeps a plain-title agent row

Review of the previous commit found ghost rows: the tab launch record is a
latch nothing clears on WSL, after an SSH exit, or for a launch that never
started, and a parked pane's process read went stale because only the
mounted tracker re-derives it.

- The launch record returns to main's role: a fallback only for titles that
  show activity, ranked below a title naming another agent (pane reuse),
  matching the tab icon's order.
- A parked pane's command boundary retires its unconfirmable process read,
  like the mounted ladder's unavailable path; reveal re-reads it.

* fix(sidebar): confirm before a parked marker retires an agent; read Git Bash prompt titles as the shell

- A parked pane's end-of-command marker can be a nested shell's leak under a
  still-running full-screen agent, so confirm the foreground first (as the
  mounted ladder does) and retire the process read only on a shell or no
  answer. SSH/remote parked panes hold no incarnation to fence a host read
  with, so they still retire.
- Git Bash emits no command marks; its `$MSYSTEM:$PWD` prompt title
  (MINGW64:/c/repo) is now shell evidence, so a stale Codex read there no
  longer keeps a ghost row after Codex exits.

* fix(sidebar): trust only process-read agents for plain-title rows; per-worktree foreground selector groups by tab

A daemon reattach seeds the pane's foreground entry with its launch agent, which
can outlive the process while Orca is closed. The entry now records where its
agent came from (agentEvidence), and the sidebar/dashboard title-derived rows
only keep a plain-title row on an actual process read. Routing and the tab icon
are unchanged.

selectPaneForegroundAgentsForWorktree grouped every pane key per worktree; it now
groups by tab once per map identity and skips worktrees with no tabs.

* fix(sidebar): a parked pane's reattach keeps its own process read of the same agent

The reattach seed marked a returning parked Codex pane as launch-record
evidence, over the process read this session already took, so its row
blinked out on reveal and stayed hidden if the user left the tab before
the visible read landed. Keep the read when it names the same agent; the
seed still drops byte-routing trust.

* test(terminal): foreground confirmation publishes process-read evidence

* fix(sidebar): a cleared pane title retires the agent's process read

Codex clears its title when it exits, and the tab then shows its default
title. A pane without shell command marks never re-reads its foreground
process, so the retained read kept a "Codex · Idle" row after /quit
(permanently for a hand-typed Codex; about 15 s while the marked-pane
confirm ladder ran). Treat a blank title like the default title it shows.

* fix(sidebar): the pane's process monitor retires an exited agent's process read

A hook-less pane keeps its sidebar row from the tracker's foreground-process
read, but nothing re-derived that read in a pane without OSC 133 command
marks. After Codex exited there, a "Codex · Idle" row stayed: permanently
when the shell titles its prompt, or when a killed Codex leaves its last
title.

The pane's agent-completion process monitor already confirms an agent's
exit (no agent and no child processes, held past its settle window). It now
reports that exit to the tracker, which retires its own process read and
runs the confirmed-shell path the visible-pty read uses. A tracker read that
names an agent seeds the monitor, so hidden panes and panes the monitor had
not polled yet are watched too. A command read in flight still decides the
pane, and launch records or other agents' reads are left alone.

* fix(sidebar): a monitor-confirmed exit leaves the next agent in an unmarked pane identifiable

The process-exit retire published shellForeground:true and left the one-shot
visible sample settled; a pane without command marks has no command start to
lift either, so a Codex typed again after quitting was never read and lost its
row on retitle. Publish shellForeground:false and reopen the sample.

* test(terminal): justify the pane binding cast in the process-exit relaunch test
2026-09-30 00:17:53 -07:00
Brennan Benson f4068747ac fix(claude): a restarted provider continues the subagent roster earlier runs journaled (#23758)
* fix(claude): a subagent resumed after a restart keeps its canonical id

A backgrounded Claude subagent resumed by a message after its session's provider
restarted parents its frames to its ORIGINAL spawn call, while the announcement the
new provider sees names only the resuming call. The spawn call's alias lived only
in the old provider's memory, so every row the resumed child wrote fell through to
its raw call id: one child shown as two, a named roster entry with no rows and an
unnamed section holding them.

The alias table now also recalls what an earlier run of the session resolved,
re-derived from the agent rows it journaled (canonical id beside the call its frames
arrived under), read once per bound journal epoch. Nothing new is persisted.

* fix(claude): a restarted provider continues the subagent roster earlier runs journaled

The roster's state lived in one provider process while everything it writes is
the session's. After a restart a resumed child was re-rostered in a second group
row with its attempts restarted, its resumed frames lost the alias they still
carry, and the one group row no turn owns was rewritten from empty, erasing the
children an earlier run had listed there.

A run now reads what earlier runs journaled, once per bound journal epoch: group
rows give each child's entry and group, agent rows give the spawn-call aliases
and the latest attempt. A group is inherited only when this run's own events
reach it, and inheriting writes nothing. An inherited child is reopened by an
announcement exactly as an in-process resume reopens it, and by nothing else.
The alias recall is one facet of that read. Nothing new is persisted.

* fix(claude): an earlier run's subagent takes Claude's restart verdict in one row

At restart Claude reports each agent the previous session left running as
stopped ("didn't finish before the previous session ended"). The roster ignored
it, leaving the entry unverifiable, while the background-task lane, which never
saw that agent announced, opened a second, unnamed row for the same agent.

The roster now records the verdict on the inherited entry, keeping the time the
earlier run lost contact, and still reopens the entry on an announcement. The
background-task lane leaves any agent an earlier run rostered to the roster,
from the same journal-derived reading.

* fix(claude): a restart verdict on a child whose host died invents no stop time

A journal reopened after its host died leaves a child unverifiable with no stop
time. Stamping Claude's later restart verdict with the current time would show
the whole outage as how long the child ran, so the earlier run's stamp, or its
absence, is kept. The journal-liveness comment no longer describes the roster as
unable to continue from the journal.

* fix(claude): a child two journaled rows list stays live in one of them

An older build could re-roster a resumed child in a later turn's row, so two
group rows list it (67 children in 39 local sessions). Inheriting the second
row re-pointed the child to that row's copy, so the copy the resume had
reopened was left at working and swept to unverifiable, while the outcome
landed on the stale copy. The first row reached keeps the child.

* fix(native-chat): no run length for a settled child with no stop time

A subagent whose host died without sweeping it has no stop time. After a
restart, Claude's own verdict on it ("stopped") is now recorded, but the rule
that hides the group's duration only covered `unverifiable`, so a mixed group
showed the sibling's duration as the whole group's. The rule now keys on the
missing stop time, whatever the settled state.

* fix(claude): a twice-listed child resumes in the row the journal reading chose

A child an older build listed in two rows was placed in whichever row this
run's frames reached first, so a sibling's frame reaching the older row made
the resumed child reopen there. Placement now follows the reading's one
tie-break, and evicting a row drops only placements that point at it.

* fix(native-chat): a resumed subagent's clock never counts the idle gap

A reopened Claude child now starts a new run: its startedAt is reset to the
reopening frame's time on every reopen path (resume after a restart and a
same-process reactivation). The group clock reads the union of the children's
latest runs, so an idle gap between runs is never counted; an ordinary
overlapping fan-out reads the same as before.
2026-09-29 23:47:48 -07:00
Brennan Benson 21124db4d5 refactor(native-chat): a subagent's rows live in its own section, not in the parent's conversation (#23752)
* refactor(native-chat): a subagent's rows live with that subagent, not in the conversation

A subagent's rows were drawn in its parent's conversation, each captioned with
the subagent's name. They now belong to the subagent: the transcript projection
keeps the session's own rows as the conversation and each subagent's rows apart,
keyed by the agent id its roster entry already carries, folded on their own.

Desktop: a subagent's rows open in a section under the roster row that names it,
from that agent's roster entry, and are windowed like any other rows. A subagent
no loaded roster names opens where its first row happened, inside the section of
the agent that spawned it or in the conversation. Its edits still count in the
turn they were made, and revealing one opens the sections around it.

Mobile shows the conversation, with each spawn's roster line. Worker reads and
structured terminal reads serve the worker's own rows.

Removes what the move makes redundant: the per-row caption and its copy, the
producer check in the tool fold and the turn answer, the per-agent frontier
interleaved in the conversation, worker-text subagent tags, and the agent id on
worker-read messages.

* refactor(native-chat): a diff target names the sections its row sits in

Revealing a subagent's edit opens the sections around it from the target the
rollup already holds, instead of looking the row up at click time. The section
head keeps to the agent's name and dot; its state in words stays on the roster
entry. The worker page test stubs the host through its module rather than a cast.

* fix(native-chat): a working subagent's section is open; a worker page windows its own rows

A subagent's section is open while its agent works and closes once it settles,
the way the turn's own live run does; a section the reader opened or closed by
hand keeps that choice. A subagent another subagent spawned opens inside that
one's section, so a working grandchild shows inside its working parent. Openness
is derived from the roster's state and the reader's choices; nothing stores an
automatic open.

A worker page is now the newest page of the worker's own rows. The host windows
the read over them before the limit, so a subagent's burst can no longer crowd
the worker's rows off the page, and "older" still means older worker rows. The
scope is an in-process argument of the host's history read; no wire request
carries it.

* fix(native-chat): a subagent section head names the turn it sits in, for the outline rail

* fix(native-chat): a subagent section's rows sit in the turn the section is shown in, for the outline rail

A background subagent's rows written during a later turn carried that later
turn onto their slots, so scrolling through its section lit the later turn's
rail tick and then snapped back. The rollup still counts each edit in the turn
it was made; only the slot, which the rail reads, takes the shown turn.

* fix(mobile): Load earlier reads past pages that hold only a subagent's rows

Mobile draws only the session's own rows, so an older page made entirely of a
subagent's rows landed as nothing: the reader tapped Load earlier, saw the
spinner, and got the same transcript back. One load now reads on (up to 8 pages)
until a page holds a row of the session's own, then applies the pages in order.

* test(mobile): stub the RPC client the way the other structured-session hook tests do

* perf(native-chat): order subagent rows for the changed-files rollup once per change to them

The rollup flattened and re-sorted every subagent row on each update, including
every token the parent streamed. The ordering now keys on the projection's
subagent rows, which keep their identity while only the conversation changes.

* refactor(native-chat): order subagent rows in the sections hook, keeping the list under its line limit

* fix(mobile): a transcript whose newest page is only a subagent's rows reads back on its own

Opened while a subagent is busy, the newest page can hold nothing but that
subagent's rows. Mobile draws none of them, so the reader saw an empty chat with
a Load earlier button, and an empty list cannot be scrolled to page. The hook now
reads back once from each such head, and the read runs on to the session's own rows.

* fix(native-chat): count the live window in the session's own rows, so a subagent's burst keeps its roster

The live window kept the newest 1,024 rows of every agent. A subagent writing
more than that trimmed its own spawn's roster row and the prompt, and its
section fell back to a closed, unnamed header. The window now keeps the newest
1,024 of the session's own rows and everything after, with an 8,192-row cap on
every agent's rows as the memory backstop. A transcript with no subagent rows
trims exactly as before.

* fix(agent-session): window history pages by the session's own rows, with a subagent's rows riding along

A history page held the newest 200 rows of every agent, so a subagent's burst
could fill a page on its own: the phone opened on an empty chat and "Load
earlier" landed nothing. A page now starts at the oldest of the newest `limit`
rows of the session's own and serves every row from there, so the subagent's
rows come with the conversation they happened in. The page stays contiguous,
the cursor still names its first row, and the byte bound still applies. A
transcript with no subagent rows gets the same pages as before.

Clients already take a page larger than its limit: both reducers raise their
retained window to the page's size. The mobile read-on and read-back stay for
older hosts.

* test(agent-session): a page reaches back to the start rather than leaving a subagent-only page

* fix(native-chat): an own-row trim takes a trimmed roster's subagent rows with it

The live window trimmed to just after the own row it dropped, so a subagent
whose roster row went kept its rows at the top as an unnamed section until
the parent wrote again. Trim to the oldest own row kept instead; it still
fires only once an own row passes the limit, so a paged-in run of subagent
rows at the head stays until then. With no subagent rows nothing changes.

* perf(native-chat): cap the live window at 4,096 rows, bounding each delta's re-derivation

Every live batch re-derives the transcript over every retained row. On the
largest real window (7,374 rows) that cost 7-8 ms a delta on desktop against
0.6 ms at the old 1,024-row window, and held about 26 MB of row content.
4,096 halves both. The most rows any local journal puts between a roster and
its subagent's last row, with the parent inside its own-row limit, is 3,005,
so no observed subagent loses its roster to the lower cap.

* fix(native-chat): a subagent section opens only while its roster is the running scope's live frontier

A section used to open whenever its roster said the subagent was working, anywhere
in the transcript and whether or not the session was running, so a background
subagent's section stayed open and grew mid-transcript while the parent moved on.

It now opens by default only while the session runs and the roster row naming the
subagent is the newest thing the parent produced, user rows aside. Newer parent
output closes it even while the subagent still works; the roster row keeps
showing that live state. A subagent still working is a running scope of its own
for the sections it spawned; a settled one closes its scope. Derived every
render, no latch; the reader's own open or close still wins.

* fix(native-chat): name a subagent's section from a client roster the window never trims

A section took its name and state from a roster row in the loaded window. Once a
burst trimmed that row, or the row sat on an older page, the section fell back to
an unnamed, closed "Subagent" header.

The shared reducer now keeps a roster keyed by agent id, folded from every roster
row and revision the client receives: pages, older pages and live batches,
including revisions of roster rows outside the window, which live batches already
carry. The first roster naming an agent wins and its revisions update it; a
removed roster row drops its entries; it is rebuilt on every page that replaces
the window and bounded to 512 agents. Sections take their name, state and
live-frontier place from it; placement stays under the loaded roster row, else
at the section's first loaded row. Only a subagent no roster ever named stays
unnamed.

* feat(agent-session): a history page names the subagents whose roster row is older than it

A page is a contiguous run of the journal whose older-page cursor is its first
item, so it cannot pull an older roster row in without skipping the rows between.
When a page held a subagent's rows but not the roster row naming it (about 11% of
the moments a reader could open a session on local journals), that subagent drew
as an unnamed "Subagent" header.

History and hydration pages now carry an optional `subagentRoster`: the first
roster entry naming each subagent whose rows are on the page and whose roster row
is not, with the row's id, sequence and revision; bounded to 64 entries and
16 KB. Items and cursor are unchanged. The client seeds its roster from it.

Rule 1 in docs/reference/remote-wire-compatibility.md: an optional field on an
existing frame, no capability gate. An older client ignores it (the released
reducer reads a page with it exactly as one without); against an older host the
field is absent and the section falls back to an unnamed header.

* Revert "fix(native-chat): an own-row trim takes a trimmed roster's subagent rows with it"

This reverts commit 22078656b2.

Its only purpose was to stop a subagent whose roster row an own-row trim had
dropped from showing at the head of the window as an unnamed section. The client
roster now names that section whatever the window holds, so the cut is back at
just after the own row the limit passes. The retention test that pinned the
unnamed-section case now asserts the section at the head keeps its name.

* chore(native-chat): state the retention limits' own reasons, now that no name depends on the window

Own-row retention keeps the conversation a reader sees from being crowded out by
rows drawn as a one-row section on desktop and not at all on mobile; the
every-agent cap bounds memory and each live delta's re-derivation. Neither is
about keeping a roster row loaded any more.

* fix(native-chat): hold the roster fold's draft map where type narrowing can see closure writes

* fix(native-chat): a roster row's newer revision replaces it in the client roster too

A revision that stops naming an agent (the host drops an entry it learns is not a
subagent, or re-keys a provisional one) left the client roster holding the old
entry, often still "working", with nothing to re-derive it. The section then read
as working forever and could auto-open, while a fresh read of the same journal
left it unnamed. The fold now drops an entry when a newer revision of the row that
named it no longer does, before any roster takes it over.

* fix(native-chat): a parent's spawn and wait calls keep the subagent they name open

A subagent section auto-opened only while its roster row was the running session's
newest row, so any later row closed it: a Codex wait on the agent, or the parent's
text before its next spawn call. Now a row that is part of delegating to a subagent
keeps that subagent open:

- a Codex collab call (spawn, wait, resume, message, close) opens each agent its
  receiver thread ids name; one naming none is ordinary output;
- a Claude spawn call names no agent, so it counts toward the roster announcing it;
- a roster row at the frontier opens its most recently added agent, not all of them.

A roster or call naming only agents one subagent spawned is that subagent's output,
so a grandchild's roster, which the host journals as the session's row, no longer
closes the spawner's section.

* fix(native-chat): a parent's call right after the roster closes its subagent's section

A parent's tool calls after a roster row fold into the tool run drawn above
the roster, so the roster stayed the newest drawn row and its section stayed
open while the parent was already reading or running commands. The fold now
records the newest journal position among the rows it merged, and the live
frontier orders rows by that newest part. The layout is unchanged. A spawn
call folded there still counts as part of the roster announcing it.

* fix(native-chat): a Codex call naming several subagents delegates to the first

A Codex collab call that names several agents opened every one of their
sections. It now counts as delegating to the first agent it names, so one
section opens, the same as a call naming one agent.

* fix(native-chat): closing a roster's list closes the sections under it

Collapsing a roster row's list of subagents left their open sections drawn,
so the section's own head became the only way to close them. And the list's
open state lived in the row, so a row the window unmounted came back
collapsed.

The transcript now holds each roster list's open state beside the section
choices. A closed list hides every section it anchors; each section keeps
its own open or closed choice for when the list reopens. With no choice from
the reader, a list is open while a section under it is open. Closing a
section from its entry keeps the list open, and revealing a subagent's edit
opens the list it sits under.

* perf(native-chat): a reveal finds the roster lists it opens with one set lookup per entry

* fix(native-chat): a subagent's roster entry heads its own rows

An open section drew the agent's name twice: its entry in the roster's list,
then a separate section head above its rows. The entry is now the head. The
roster row draws its entries through the first open one, that agent's rows
follow, then the entries after it, each run in its own windowed slot. A
section no loaded roster row holds (an older page, a grandchild, an unnamed
agent) keeps its own head.

A roster list is open while the live frontier or a reader's choice is on
one of its agents, unless the reader closed the list, so closing an agent
from its entry no longer needs to pin the list open.

The section emitter moves to its own module, and the trailing-run
predicates it shares with the slot builder to theirs, to keep the slot
builder under its line limit.

* fix(native-chat): the entries after an open subagent's rows set in its roster's type

The roster row's list inherits the system row's small muted type; the entries that
follow an open section sit outside that row, so they now carry the same type.

* docs(native-chat): a current host can also serve a page of only a subagent's rows

A page is bounded by bytes after it is windowed by the session's own rows, so a
burst that fills the bound yields a page, or an opening page, with none of the
session's own rows. Mobile's read-on and read-back therefore serve current hosts
too, not only older ones; the comments said otherwise. The retention comment
still described a closed section as a row of its own; it now sits behind its
roster entry.

* fix(native-chat): a section's prose keeps its copy/timestamp controls inside the section

An assistant row's hover controls (copy, scroll-to-top, timestamp) hang 20px
below the row into the gap before the next one (`-mb-5`). Inside a subagent's
section that put them below the section's left border, and on the section's
last row they touched the parent's next row with no gap.

Inside a section the controls now stay in flow, so the border covers them and
the next row sits the normal gap below. The row-height estimate reserves the
same 20px for a section's prose so windowing does not jump on measure.

* test(agent-session): state each appended row's turn scope, as the journal now requires

* refactor(native-chat): the client's journal retention policy lives in its own module
2026-09-29 23:47:02 -07:00
Neil dcaef9dee5 test: retire relay, preload and shared cases that re-prove an owned contract (#24007)
Audit sweep over `src/relay`, `src/preload` and `src/shared` (1,087 test files
reviewed). 101 case declarations removed across 40 files, 6 test files deleted
outright, 1,143 lines gone. Executed-case count falls further, since several
removals were `it.each` tables.

Dominant patterns, by frequency:

- Self-comparisons that cannot fail: `expect(f(x)).toBe(f(x))`,
  `JSON.parse(JSON.stringify(literal))` deep-equalling the literal for a type
  with no codec, and `normalizeKeyToken(t) === normalizeKeyToken(t)` presented as
  proof of memoization.
- Object literals asserting their own fields back, where the guarantee comes from
  the type annotation and the runtime assertion cannot fail.
- Copied inventories: constants compared to their own initializers, and a
  function returning a copy of an exported constant checked against that
  constant's literal contents.
- Duplicate invocations of a contract owned at a stronger boundary, including
  provider-local replays of a shared helper.
- Table rows varying a field production never reads, so every row runs one path.
- Names promising more than the input exercises: a "Windows launch" case in a
  module with no platform input, and a case whose named branch is never entered.

Two production symbols go with them, each a test-only export whose sole caller
was a deleted case:

- `getGitHubProjectRefInputByteLength` — a one-line forward to
  `getClipboardTextByteLength`. The real bound
  (`GITHUB_PROJECT_REF_INPUT_MAX_BYTES`) and its guard stay.
- `GRAB_STYLE_PROPERTIES` — an intended shared source of truth that nothing ever
  consulted; the property set is hand-enumerated at three independent sites.

One case was deliberately restored and strengthened rather than dropped. The
relay integration suite is the only place the real `SshChannelMultiplexer` is
wired to `RelayDispatcher`, so it reaches transport behavior the handler suites
cannot (they use `createMockDispatcher`). Its `fs.writeFile` roundtrip is the one
case producing a void result, and `JSON.stringify` drops an absent `result`
member — a shape no other surviving case exercises. Restored with an assertion
pinning what the client actually observes: `null`, not `undefined`. That
assertion failed on first run, so the fact was previously unasserted anywhere.

One deletion was reverted mid-audit. A case asserting that optional fields stay
invisible to "old attach and ready decoders" builds those decoders from `z.object`
schemas declared in the test file, so it demonstrates zod's unknown-key stripping
rather than anything shipped. It is nonetheless the only forward-compatibility
coverage these envelopes have, and `reliability-gates.jsonc:6232` names it as
evidence verbatim, so it stays. Note that `check-reliability-gates.mjs` passed
both with and without it: the script resolves manifest paths and commands, and
does not check that a named assertion still corresponds to a live case.

Kept deliberately: everything a reliability gate cites as evidence; the three
`registers all expected handlers` RPC manifests (a dropped registration is a
silent wire break no type checker catches, and one carries the STA-4571
`pty.ackData` ratchet); the `child-process` direct-import ratchet; and
prototype-spy cases paired with a `.repeat(10_000)` input, which assert a real
memory bound rather than merely forbidding a technique.

Verified: `pnpm test src/shared src/relay src/preload` (1073 files, 11996
passed, 1 pre-existing `it.fails`, 131 skipped), `pnpm tc` after clearing
`.tsbuildinfo`, `check-reliability-gates.mjs` (140 gates),
`check:code-quality:changed` (0 new findings).
2026-09-29 22:17:13 -07:00
Neil b7209b5ae9 perf(git): relist only the repo whose worktrees changed, and stop blocking main on sync git (#23998)
* perf(git): stop blocking main on the open-on-remote git cascade

`getRemoteFileUrl` ran up to 6 sequential `gitExecFileSync` calls on the Electron
main thread — `remote get-url`, then `getDefaultBaseRef`'s `symbolic-ref` plus up
to four `rev-parse --verify` probes — each with its own 15s timeout and no yield
between them.

A complete async twin already existed (`getDefaultBaseRefAsync` ->
`resolveDefaultBaseRefViaExec`, sharing DEFAULT_BASE_REF_PROBES), so the sync
cascade is deleted rather than converted. `getRemoteUrl`, `getRemoteFileUrl` and
`getRemoteCommitUrl` become async; all four downstream callers were already async
(`filesystem-git-url-handlers` inside `ipcMain.handle`, `runtime-git-diff-commands`
async methods) and the provider contract already typed both wrappers
`Promise<string | null>`, so no new async plumbing was needed.

Removes 3 of the 10 `gitExecFileSync` sites and the confusing name collision with
the unrelated async `getDefaultBaseRef` in hosted-review-creation-git-state.

The base-ref regression tests keep their coverage, repointed at the public async
`getBaseRefDefault`.

* perf(git): resolve the repo root in one sync spawn instead of two

getGitRepoRoot ran `rev-parse --is-inside-work-tree` and then `rev-parse
--show-toplevel` as separate blocking spawns. Each sync git call holds the main
thread for up to its whole 15s timeout, so the spawn count is the cost — and this
function is called twice per "Add Project" on a linked worktree, once directly and
once through getLinkedWorktreeMainRepoRoot's self-recursion.

Combined into one invocation. Safe only here: in a bare repo the combined form
exits non-zero, and both that throw and the plain `false` already land on the same
marker-scan fallback. probeGitRepo deliberately does NOT combine — it has to read
`false` cleanly to go on and detect a bare repo, which the combined form's exit 128
would misread as indeterminate.

* perf(git): rebuild only the repos whose authorized roots actually changed

One worktree create called `invalidateAuthorizedRootsCache()`, which dirties every
registered owner. The next authorization-requiring IPC then rebuilt by listing EVERY
repo — and the rebuild never consulted `dirty` when choosing what to list, so `dirty`
gated only whether a rebuild ran, not its scope. At 58 repos that is 58
`git worktree list` spawns, roughly ten seconds of git wall-clock through an
admission budget of four, to rediscover roots one repo changed.

Both halves were needed; scoping the invalidation alone changed nothing.

- `markAuthorizedRootsOwnerDirty` dirties a single owner, reusing the per-owner
  primitives `registerWorktreeRootsForRepo` already used. It leaves `baseRevision`
  and the per-repo revision map alone — that pair is the global side-effect-token
  fence, and bumping it would retire in-flight tokens for untouched repos.
- `rebuildAuthorizedRootsCache(store, onlyDirty)` re-lists only owners that are
  dirty, have no listing yet, or still hold recovered roots (those are retired by
  comparison against a fresh listing, so skipping them would strand them as
  authorized). Only `ensureAuthorizedRootsCache` passes `onlyDirty`; an explicit
  rebuild keeps re-listing everything because callers use it to force a refresh —
  `filesystem-auth.test.ts` pins that contract.

`invalidateAuthorizedRootsCacheForRepo` wraps the primitive and falls back to the
global form for an unknown owner or a missing store, rather than silently skipping an
invalidation and leaving a stale allowlist. Applied to the worktree-create path.
Changes that can alter the owner SET (store swap, host/WSL re-routing, nested-repo
import, folder->git upgrade) stay global. Removal paths are not converted yet.

The allowlist contents are unchanged and the failure direction is a false denial
rather than a false allow. The relist predicate is split into its own module so it is
testable alone and the cache file stays inside its line budget without a suppression.

* test(perf): measure what git orchestration actually costs the main thread

The existing churn probe (ORCA_MAIN_THREAD_DIAGNOSTICS=1) reported spawn-initiation
cost for git/gh/glab only — its 7 call sites all sit inside git/command-runner — so
it was blind to `spawnProcess`/`runProcess`, the repo's own mandated wrapper, and to
the blocking `execFileSync('ps')` per PTY resize. That understated total churn across
115 main call sites.

- `spawn-observer.ts`: a settable seam, since shared code cannot import src/main.
  Unregistered in the daemon/relay/CLI, where it costs one boolean check.
- `spawnProcess` brackets `nodeSpawn` and reports; exec-file-capture's own report is
  removed because it routes through runProcess and would double-count.
- `posix-pty-foreground-group` now reports its full blocking duration. Note this
  lands on the daemon, not main, whenever the daemon hosts the PTY.
- `ORCA_UNMINIFIED_MAIN=1` build flag, because a minified main bundle cannot
  attribute CPU-profile self time to real function names. Defaults unchanged.
- `main-thread-git-cost.spec.ts` + `analyze-main-cpuprofile.mjs`: sweeps concurrency
  against real registered repos, captures the churn lines and a V8 CPU profile of
  main per phase.

What it found, which is why this is worth keeping: at the width-4 admission ceiling
(~90 git:status/s) main sees ZERO event-loop gaps over 50ms and a worst gap of 23ms,
and is 85% idle. Git orchestration does not stall the main thread. Of the cost it
does incur, spawn-init is 58%, parse 5%, stdout drain 4%.

* test(perf): name the inspector params type the anti-slop gate requires

The broad `object` parameter trips anti-slop(no-object-parameters); the only
Profiler call that passes params sends `{ interval }`.
2026-09-29 22:01:08 -07:00
Jinwoo Hong 80786ddccb fix(onboarding): build the agent step around skills, not CLI registration (#22720)
* fix(onboarding): build the agent step around skills, not CLI registration

The checklist step "Enable Orca CLI" was marked done once the agent skills
were installed, while Settings -> Browser still showed CLI registration as a
pending step. Orca terminals already put the bundled CLI on PATH, so
registration only matters for shells Orca did not launch, plus WSL, where
`orca-ide` exists only once registered.

- Rename the step to "Give agents Orca skills"; setup registers the CLI only
  for WSL (isOrcaCliRegistrationRequired), via onboarding-cli-registration.ts.
- Settings -> Browser drops the CLI step outside WSL (2 steps instead of 3).
- Skills panel: "All skills installed" + "Update skills" replaces the disabled
  button; status pills sit top-right; no "Installed" beside "Unavailable".
- Full Disk Access moves to the "Start work in multiple repos" step.

Fixes STA-8306 / #22524.

* refactor(onboarding): simplify agent-skill step state after review

- One done rule: isAgentCapabilitiesDone in feature-wall-setup-progress.ts,
  reused by the skills panel instead of a mirrored copy.
- 'unavailable' is an install-status tone instead of a second boolean;
  pill/note rendering moves to AgentCapabilityStatusBadges.tsx.
- "Update skills" skips Computer Use when it can't run (no warning toast);
  setup takes an explicit selection.
- The WSL gate lives in registerOnboardingCliIfRequired; onboarding deps drop
  the now-unreachable host CLI branches.
- BrowserUsePane: one cliRequired/cliReady pair, no host CLI status fetch,
  WSL-only enable path, single "Finish the steps below." string.
- Full Disk Access placement goes through a SelectedStepFooter switch.
- Prune orphaned locale keys (and their boot-bundle entries); fix a stale
  comment.

* fix(skills): require CLI registration only for WSL setup

* fix(skills): retain CLI install labels for WSL

* docs(skills): clarify remaining WSL registration fallback

* fix(skills): skip WSL registration when the host confirms managed CLI access

* fix(wsl): prepare managed shell wrappers before onboarding probes

* test: update daemon capability and terminal hook expectations

* refactor(onboarding): remove CLI registration checks from skill setup

* refactor(setup): remove redundant state and obsolete registration scaffolding

* fix(settings): stop registering the CLI before installing the CLI skill

The General > Orca CLI skill panel still registered `orca` on PATH before
opening the install terminal, contradicting the rest of skill setup. Orca
terminals already provide the CLI, so the shell command toggle now says it
is only for terminals outside Orca.

* style(onboarding): polish the setup checklist and first-run steps

Make onboarding monochrome: completion is a neutral check, selection a
neutral outline, and color only flags real problems. Tighten the checklist
rail and header, single-line agent cards with a grid that scrolls only when
it runs out of room, a labeled permission switch under the grid, calmer
notification and skill cards, sentence-case copy, and a labeled "Hide
checklist from sidebar" action. Workspace setup leads with "Add project"
when no git project exists.

* fix(emulator): drop the Enable Orca CLI step from agent control setup

Agents that drive the emulator run in Orca terminals, which already provide
the `orca` command. Agent control setup in the emulator card and Settings is
now a single step: install the Orca CLI skill.

* fix(onboarding): hide the Full Disk Access card once access is granted

A granted card has no remaining action and only takes space on the add
projects step. It also no longer flashes a "Checking" state before the
first status arrives.

* fix(onboarding): address review on permission warning, hide button, and translations

- Name the permission switch "Yolo mode" (matching Settings > Agents) and
  state the risk: agents act without asking and some bypass their sandbox.
- Hide the modal's "Hide checklist from sidebar" button below sm, where the
  header centers its title under it; the sidebar entry keeps its own control.
- Translate every string this PR adds into es, fr, ja, ko, and zh.
2026-09-29 23:55:17 -04:00
Neil fb52c0602a fix(terminal): release xterm's DEC 2026 render hold instead of waiting out its 1s timeout (#23920)
* fix(terminal): release xterm's DEC 2026 render hold instead of waiting out its 1s timeout

xterm paints nothing while DEC mode 2026 (synchronized output) is open and only
force-flushes after 1000ms. Codex wraps every draw in mode 2026, so any byte gap
or chunk split that loses the closing \x1b[?2026l freezes the pane for a full
second and then repaints in one burst.

Orca never emitted \x1b[?2026l anywhere, and three paths could destroy a TUI's:
the per-PTY pending cap drops buffered output wholesale (mode 2031 was already
salvaged there, 2026 was not), main sliced pending data at a blind 16KB offset
that can land inside an open frame or sever the 8-byte marker, and the renderer's
backlog warnings replace a queued tail that may hold the close.

- salvage the 2026 latch across dropped output, mirroring the existing 2031
  salvage, and append the release on both delivery sites
- ground 2026 in RESET_AFTER_BYTE_GAP and the replay baseline, and in both
  backlog warnings, so every drop path is self-healing
- make main's 16KB flush split frame-aware instead of a blind byte offset
- lift the synchronized-output scanner into shared/ so main and the renderer
  use one implementation

Closing a frame early costs one premature repaint; leaving it open costs a
second of blank screen, so the asymmetry favours always closing.

Also adds the reproduction this needed: the pre-existing typing bench observes
the xterm BUFFER, which the parser fills while rendering is held, so it scored
these freezes as fast echoes.

* fix(terminal): stop the renderer's queue drain cutting inside an open DEC 2026 frame

takeQueuedChunk sliced a queued chunk at a blind byte offset to fit the 16KB
coalescing budget, which can strand a frame's closing \x1b[?2026l in the residual
until a later drain. Same defect as main's flush split, same fix: reuse the
frame-aware split helper.

Usually masked because the drain coalesces adjacent chunks and reassembles what
main split, but not when the budget boundary falls inside a frame.

* fix(relay): keep the SSH path's bounded slice outside an open DEC 2026 frame

pty-handler split pending output at a byte offset with a surrogate-pair guard but
no synchronized-output awareness, so a frame straddling the 16KB wire slice had
its closing \x1b[?2026l stranded in the remainder — the same defect just fixed on
the local path, on the path AGENTS.md requires us to consider.

Placed before the surrogate guard so that guard keeps the final say, and floored
at 2 so frame alignment can never walk a healthy slice into the guard's
decrement and then into the chunkChars <= 0 pause-and-retry path.

Also drops a dead `splitAt === 0` branch in takeQueuedChunk: both callers pass a
positive limit and the helper never returns 0 for one.

The two new split tests were each confirmed to fail without their fix.

* test(terminal): sweep the DEC 2026 split helper over escape-sequence shapes and every limit

Covers OSC 52, DCS, repeated open/close markers and limits 1..len+3, asserting the
result never exceeds the limit, never reaches 0, and stays byte-exact. Also pins
that a buffer beginning inside an open frame degrades to the blind offset rather
than doing something worse, and documents that callers do not thread latch state.

* fix(terminal): ground DEC 2026 on the daemon slice, the recovery replays, and the process boundary

Four more sites could strand the latch, found by sweeping every path that drops,
splits, or replays terminal bytes.

- daemon-stream-data-batcher: the 64KB bulk-write slice used a surrogate-only
  clamp, and its remainder is HELD until 'drain' — "seconds for multi-MB
  backlogs" per the file's own note. A frame straddling that boundary parked its
  \x1b[?2026l behind the hold, blanking the pane past xterm's 1s timeout once per
  frame for as long as the backlog lasted. This is the default daemon-backed pane
  path, so it is the one users actually hit. The new
  clampToSafeBulkWriteSplitIndex frame-aligns first and surrogate-clamps last,
  and lives in daemon-stream-data-split alongside the policy it belongs to.
- replay-data-drain and remote-runtime-terminal-binary-snapshots wrote a bare
  \x1b[2J\x1b[3J\x1b[H, which does not clear mode 2026 — so on the SSH/remote
  reconnect path, the very event most likely to sever a frame, the whole replay
  could paint nothing.
- ipc-pty-attach: trimIncompleteTerminalControlTail can cut a half-written
  \x1b[?2026l while its opening marker survives in the replayed prefix.
- PROCESS_BOUNDARY_GROUND: the "process that armed these modes is gone" ground
  omitted 2026, the last unexplained gap in that file. A disable, so it still
  satisfies the recovery barrier's ownership scan (only ?25h may be an enable).

Recovery-path expectations updated where they pin the emitted bytes. Deliberately
NOT touched: apply-reattach-payload and ssh-snapshot-prepaint already ground via
buildSnapshotReplayPrologue.

Still unfixed, deferred with reason: terminal-output-frame-chunks.ts splits the
remote wire on accumulated UTF-8 byte width and needs a different shape than the
char-index helper; desktop clients reassemble in main's pending buffer, so the
exposure is mobile/web only.

* fix(terminal): emit the DEC 2026 release before the mode-2031 tail, and stop claiming the drop path writes it

Two corrections from adversarial review of the earlier commits.

1. Ordering bug I introduced. getDroppedMode2031RendererData ends with
   `state.tail`, which extractPrivateModeScanTail deliberately retains as an
   INCOMPLETE private-mode sequence so the next chunk can resolve it. Appending the
   2026 release after it put an ESC behind a dangling CSI, aborting it and silently
   losing whatever mode spanned the drop boundary. The release now goes first.

2. The drop-path release does not reach xterm in the dominant case, and the comment
   now says so instead of implying otherwise. live-data-callback's droppedOutput
   branch discards `data` and salvages only queries
   (salvageRendererQueriesFromDiscardedRestoreData handles CPR/DA1/OSC colour;
   \x1b[?2026l is not a query), so for hidden panes and visible panes outside
   foreground-restore backpressure the synthesized release was dropped. The grounded
   snapshot replay releases the latch instead.

   I tried writing it through writePtyOutputToXterm there and reverted: it consumes
   the pending hidden-output snapshot and broke
   pty-connection-hidden-snapshot-resize-signals ("re-restores a skipped alt frame"),
   so the release rides the restore rather than perturbing that state machine.
   Residual gap, documented: a cap-dropped pane whose restore never arrives.

The salvage is still load-bearing on the fall-through path, so it stays.

* fix(terminal): release DEC 2026 on the reattach clears, floor the split, and correct the freeze framing

Remaining findings from adversarial review.

- apply-reattach-payload's three bare-clear branches (:63 daemon snapshot, :229
  relay replay, :269 cold restore) had no release anywhere in their sequence: I
  checked all seven POST_REPLAY_* profiles reachable via chooseReattachReplayReset
  and none contains \x1b[?2026l. Only the buildMainModelSnapshotReplayWrites branch
  was grounded, so covering the streamed replay path and not the main reattach path
  was inconsistent. Verified no production code matches these clear strings — the
  three test updates are mock equality, and each was confirmed to fail without the
  source change.
- clampToSafeBulkWriteSplitIndex could return 0 (('\u{1F600}aaaa', 1) — alignment
  returns 1, the surrogate clamp decrements to 0), which would leave a zero-length
  slice that never shifts the batcher's queue entry and spin its drain loop.
  Unreachable from today's only caller, but it is exported with an unstated
  precondition. Floored at 1.
- Frame alignment could halve per-PTY flush throughput: main re-queues the
  remainder with eligibleRound = round + 1, so the shortfall cannot be refilled in
  the same round, and aligned size is floor(W/F)*F — 50% worst case in the 8-16KB
  band, which is exactly the full-screen redraw burst that reaches the pending cap.
  Alignment is now rejected below half the window, preferring throughput and
  letting the reset profiles release the latch.

Framing corrected throughout: bufferRows records a row range and clears nothing, so
the pane freezes on its last painted frame — it does not go blank. The real trade is
"stale but coherent for <=1s" versus "immediate partial frame", and
RESET_AFTER_BYTE_GAP (written alone, with no repaint behind it in the same write) is
the one site that can newly flash a partial frame. Said so at the constant instead
of implying the release is free.

* fix(terminal): rename the shape-flagged symbols the anti-slop audit rejects

CI's anti-slop gate rejects "shape" in symbol names as structural rather than
domain language: `shapes` -> `outputSamples`, and
`writeCodexShapedEchoProbeScript`/`codexShapedEchoProbeScript` ->
`writeCodexEchoProbeScript`/`codexEchoProbeScript`.
2026-09-29 20:27:30 -07:00
Brennan Benson ad2e1b5efa fix(terminal): restore the mouse format with mouse tracking, so phone swipes don't type into Codex (#23946)
* fix(terminal): restore the mouse encoding with mouse tracking in every snapshot

Swiping to scroll Codex from the phone on a Windows host typed legacy
`ESC [ M` mouse reports into the Codex composer (#23818). SerializeAddon
re-arms mouse tracking (?1000h/?1002h/?1003h) but never the SGR encoding
(?1006h/?1016h). Any snapshot taken from a desktop pane's xterm (the
runtime seeds its headless model from it after a reattach, and serves it
to remote viewers when no model exists) therefore restored "tracking on,
legacy encoding", and the phone encoded wheel events as X10 bytes, which
ConPTY hands to Codex as keystrokes.

serializeWithAbsoluteCursor, the one wrapper every Orca snapshot producer
uses, now appends the encoding xterm itself parsed, read from xterm's
mouse state service. The daemon/runtime headless model reads tracking and
encoding from xterm too, so its regex mirror of the DECSET stream is
deleted (one source of truth; one less regex pass per PTY chunk).

Mixed versions: no wire field changes. A new host's snapshot carries an
extra DECSET that old desktop and phone clients already parse; an old
host's snapshot restores exactly as before. With tracking off the encoding
alone sends no reports, so the wheel still scrolls scrollback.

* test(terminal): pin the mouse-encoding read against the renderer xterm build

* fix(terminal): type the xterm mouse-state read behind named shapes
2026-09-29 19:23:25 -07:00
Brennan BensonandClaude 6b36a2c3fb feat(native-chat): mid-turn messages wait as editable cards above the composer (#23731)
* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* fix(native-chat): name a chat write by its target, not the owner generation

A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.

Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.

Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.

* fix(native-chat): every journal append reaches the chats that are open

A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.

A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.

* test(native-chat): an epoch replacement reaches the open chat

* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

* perf(native-chat): a publish behind a delivered commit reads nothing

Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.

* test(native-chat): state why the teardown test's fake journal is safe to cast

* docs(native-chat): say mutation admission checks only the writer lease

* docs(native-chat): drop the send rebase from comments that still described it

* fix(native-chat): a message is accepted, then delivered

A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".

A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.

Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.

A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.

Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.

* fix(native-chat): settle queued messages only for the child that ended

A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.

A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.

The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.

* fix(native-chat): an adoption that fails to import keeps the conversation open

The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.

* perf(native-chat): the recovering open reads the journal once

Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.

* fix(native-chat): an attach that fails after indexing its child leaves no child behind

A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.

* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer

The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.

A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.

* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down

The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.

* fix(native-chat): a message rejected while its chat was closed reads as not sent

A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.

* test(orchestration): name why the readiness settlement fakes are cast

* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent

* docs(native-chat): drop the fence from the admission the send effects run behind

* docs(native-chat): give the fence move on release the reason that still holds

* docs(native-chat): stop citing a write fence check in launch and mailbox comments

Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.

* refactor(native-chat): the provider child is its own record

A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.

- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
  patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
  own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
  the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
  dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
  is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
  the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.

* fix(native-chat): the delivery loop alone settles a message its start or child failed

A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.

- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
  reads how it ended: a Stop continues; anything else writes one failure row and rejects every
  queued message with the same words, then stops. A child still starting whose start the adapter
  says did not land fails the same way. The exit, eviction and the settlement retry only settle
  the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
  nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
  failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
  closed, with or without a child, and a start the loop already has in flight is waited for so the
  child it produces is stopped rather than left behind.

* refactor(native-chat): a stopped child ends on the one reading of its stop

The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.

* feat(native-chat): the host says it accepts a send before any agent has it

The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.

* refactor(native-chat): an attach never opens a journal of its own

The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.

* fix(native-chat): a moved fence resends nothing on a host that accepts first

The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.

The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.

* refactor(native-chat): a child's end says whether the user or the host stopped it

The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.

* fix(native-chat): a chat whose only work is a queued message is not offered for resume

A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.

* test(native-chat): type the queued-message fixtures in the resume-offer tests

* fix(native-chat): a start that dies while a message waits on it is that message's failed start

Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.

* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now

* test(native-chat): pin what a failed start settles, and what a resume offer names

A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.

* test(native-chat): the failed-start pins fail on what the message became, not on a timeout

* test(orchestration): the preamble's host stub is typed, not cast

The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.

* fix(native-chat): a Stop that names no turn stops what the conversation has in flight

Between handing a message to the agent and the agent opening its turn, there is no turn id a
client could name, so a Stop in that gap was refused as "already finished" while the agent went
on to answer. A cancel's turn id is now an optional precondition instead of its target: with
none, the host withdraws what is queued and, when the journal still reads working, asks the
adapter to stop whatever the child has in flight. Claude's interrupt is session-scoped, so it
is guarded by fence and acquisition generation rather than a turn identity. Codex interrupts
the turn its latest turn/start answered with until the journal shows one.

A cancel that names its turn behaves exactly as before.

* fix(native-chat): Stop is there from the moment a message is sent

The composer showed Stop only once the agent had opened a turn, so for the second or two after a
send the chat read "thinking" with no way to stop it. Against a host that takes a Stop naming no
turn, Stop now shows whenever the chat reads working (a turn, a queued message, or a handed-over
one still unanswered) or this client still has a message on its way. Pressing it, or Escape,
first drops every outbox entry the journal does not hold yet, so nothing goes out after the
Stop, then sends the conversation-wide cancel. A send already on its way reaches the host ahead
of the cancel, which withdraws it there. Against an older host Stop still needs a running turn.

The unconfirmed-send probe moves into its own hook so the outbox hook stays in budget.

* fix(native-chat): Stop before a turn is gated on its own host capability

A host that accepts sends first (agent-session.accepted-send.v1) can still predate the cancel
that names no turn and would refuse it as invalid, since clients and hosts ship independently.
Hosts that take that cancel now advertise agent-session.conversation-stop.v1, and the renderer
shows Stop before a turn opens, and sends the no-turn cancel, only to a host advertising it.
Every other host keeps a Stop that needs, and names, a running turn.

The host capability probe the accepted-send hook used is generalized so both read one path.

* test(native-chat): a build advertises conversation stop exactly where its cancel may name no turn

* fix(native-chat): a view never restarts a chat whose last start failed

A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.

* test(native-chat): start the child the loop waits on with an attach, not a second view

A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.

* fix(native-chat): settle a gone generation's turn wherever a conversation opens

A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.

* test(native-chat): prove the next child's start settles the turn an earlier child left

The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.

* test(native-chat): count a failed start's rows by row, not by text

Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.

* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget

* test(native-chat): pin the open's and the send's start and row counts, however the view binds

Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.

* fix(native-chat): settle a gone generation's turn at every open but an acquisition's

The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.

* test(native-chat): hold the create's start open until the views bind

The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.

* fix(native-chat): Stop reads the one working rule every session list reads

While Claude retries a rate-limited request it never echoes the message, so no
turn opens: the sidebar read Working from the unanswered send while the composer
showed Send. The chat's working state, the host's session-list status and the
host's no-turn Stop check now call one shared rule instead of three copies.

* test(native-chat): a rate-limit retry pins only that no turn opens, not how its rows are kept

* fix(native-chat): Stop leaves a message waiting on its Retry, and does not show for one

A send that failed holds the queue until the user retries it, and one the host restarted under is
parked the same way. Stop counted both as still on their way, so it showed in an idle chat and
could never go away, and pressing it dropped the failed message along with its Retry.

* test(native-chat): the chat's Stop and a session list read the main agent alike over their own copies

The chat reduces its stream and a list reads the status feed. Driven through the real host for a
rate-limit retry with no turn, a subagent still running after the main turn, and the handed-over
child exiting.

* refactor(mobile): the chat reads the main agent's working state through the shared rule

Behaviour is unchanged: the same two terms, now from the one function the host projection and the
desktop chat read.

* fix(codex): a Stop naming no turn never interrupts an earlier turn

It fell back to the id an earlier turn/start answered with when the latest start went unanswered,
or when the journal showed a compaction Codex had not started, and reported that as stopped.

* fix(native-chat): a Stop naming no turn never says a turn had already finished

When the provider found nothing left to stop, for instance a turn that ended between the host's
check and the interrupt, the chat got "The provider had already finished this turn." for a turn
the Stop never named. It now ends quietly, as a Stop with nothing in flight does.

* fix(native-chat): one Stop the host could not settle no longer refuses every later one

A Stop naming no turn has one operation key per session. When the host could not settle one, it
answered every later Stop under the same id as unknown until the id expired. Once the host says
so, the next press is a new Stop; transport doubt still replays the same id.

* refactor(native-chat): drop the composer's second error formatter

After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.

* test(native-chat): pin the reason on a message rejected while its chat was closed

The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.

* test(native-chat): read Stop operation ids without a cast

* fix(native-chat): a Stop whose answer was lost no longer swallows the next one

A Stop that names no turn has one operation key per chat. When its answer was lost in transit, the
chat kept the id, so every later Stop replayed it; the host answers a replay as already handled, so
for up to a day Stop stopped nothing. The id is now dropped once the call settles, however it
settles. A second press while the first is still on its way still shares its id.

* refactor(native-chat): a Stop naming no target keeps its operation id only for its own call

The chat kept each write's operation id per payload across calls, and dropped it only on some
settle paths. That is right for a write naming what it acts on, but a Stop naming no turn, and a
stop of every background task, share one payload with every later one, so any path that kept the id
made the next Stop replay as already handled and stop nothing. One path was still open: an answer
that arrived after the chat moved to a new fence.

Whether a write names its target is now decided once, before its id is picked. One that names none
keeps its id only while its call is in flight, so a press made meanwhile joins it, and releases it
when the call settles, however it settles. The release runs only while the key still holds that
call's id, so a joined call settling late cannot drop a newer one's. This replaces the per-path
exceptions for a thrown call.

* test(native-chat): read the Stop fences without a cast

* test(native-chat): pin the new id for a named cancel the host could not settle

After the Stop naming no turn moved to a per-call id, the only test of the unknown-refusal release
was gone, and the half that stays, for a cancel naming its turn, could be removed with every test
green.

* fix(native-chat): a Stop pressed after a new message stops it, even while the last Stop is unanswered

A Stop naming no turn shared its operation id with any press made while it was still in flight. The
host runs a chat's writes in order, so a message sent between two presses was accepted after the
first Stop ran, and the second press replayed that Stop as already handled and left the message
running, although the chat had already withdrawn it from the outbox.

A write naming no target now gets a new id on every press and is never kept, so each Stop acts on
whatever is running when the host reaches it. A write naming its target keeps its id exactly as
before. A double press can ask the provider to stop the same turn twice, which it tolerates.

* fix(native-chat): Stop no longer blinks off as Claude opens the turn for a message

Claude's echo of a sent message both answers the send and opens its turn. The echo settled the send
first, so the host published the message as answered one frame before the turn it opened, and for
that frame the chat read nothing running: Stop turned back into Send, and Working blinked off in
every session list, for tens of milliseconds on each turn.

The echo now settles the send after the turn it opens has been emitted, so the running turn is
published first.

* fix(native-chat): a message a Stop withdrew comes back to its sender's composer

A Stop withdraws every message the host holds but has not run, and S also
drops the ones this client had not handed over yet. Either way the message
left the chat and its text survived only in a hidden journal row and the
in-memory ArrowUp history.

The sending client now puts the withdrawn text and images back in that
pane's composer, after whatever is typed there. Withdrawn is read from the
rejection reason through one shared check, which the outbox reconcile now
uses too. The composer is written before the entry leaves storage, so a
failure between the two repeats the text instead of losing it, and an entry
storage no longer holds is never given back again, so a replay, a second
view or a remount restores it once. Only this client's outbox holds the
entry, so other viewers still see the message disappear. A failed Stop
withdraws nothing on the host and gives nothing back.

* fix(native-chat): withdrawn text put back during an IME composition is not lost

While the IME owns the field, the composer ignores a programmatic draft, and
the next composed keystroke wrote the draft without the restored text, after
its outbox entry had already been dropped. The composer now holds text
appended mid-composition, keeps it in the cache after each composed write,
and shows it once the composition settles, the way attachments that land
mid-composition already wait for it.

* test(native-chat): pin that only a withdrawn message comes back to the composer

* test(native-chat): set up the composer's window API for every describe in the composition-race file

* docs(native-chat): note that the withdrawn check reads the legacy reason until a typed category lands

* test(native-chat): pin that text put back mid-composition shows once, even beside a mid-composition clear

* feat(native-chat): host-owned queued-message draft store in the session journal

A queued mid-turn message is a draft row in the session's journal.db,
created idempotently at every writable open with no user_version bump so a
downgrade stays writable. Consume converts one draft into an ordinary
submission inside the journal writer's own transaction (exactly-once), and
a standing writer hook returns a consumed draft only when a committed row
newly settles its current consumed submission to a non-withdrawn rejection
— the same decision the reducer folds rows through. Open-time repair
re-derives returned state behind the stored fact; retention never prunes a
row whose refusal could still return it.

* feat(native-chat): queued-messages wire contract, dark capability, and send classifiers

The send result becomes a union: today's submission arm unchanged, plus a
capability-gated queued arm only clients that sent delivery:'queue-if-active'
ever receive. Whole-list queuedMessages fields ride the subscribe events and
history pages; Stop gains withdrawQueued with the withdrawn bodies in its
result; clear's result carries withdrawn drafts too. Both classifiers treat
queued as accepted/spent. agent-session.queued-messages.v1 is defined but
deliberately NOT advertised: the rollout prerequisites (Claude fold receipt,
integrated Codex steer matrix) are not in this host.

* feat(native-chat): queue a capable mid-turn send as a draft, drain it at turn end, and let Stop and clear return its text

A send carrying delivery:'queue-if-active' while the session owes work — or
behind an actionable backlog — becomes a host-held draft instead of a
submission. A serialized drain woken by journal commits, draft mutations and
conversation opens re-derives its gates from live facts (streamed-event
barrier first, backlog never a gate) and converts the oldest actionable
draft through the exactly-once consume; from that instant today's delivery
pipeline runs unchanged. Stop pauses the withdrawable frontier at the stop
step (a process-level pause set that survives handle eviction and, via the
per-process host instance, restarts), then withdraws it with the text in the
result for capable clients; /clear does the same for the superseded source.
The draft list publishes whole per emit with identity dedup, rides only the
final catch-up page, and attaches to history pages. queuedMessageSend
overrides queue policy only; queuedMessageDelete hands the body back.
Replays for all of it answer from op-stamped tombstone receipts.

* test(native-chat): pin mid-turn queueing against the real host

Accept (working/backlog/text-only/budget/replay), the one-per-settle drain,
returned cards with N1 overtake and the N4 re-send loop, Stop withdraw with
tombstone replays, the process-level pause across evict/reopen, Delete
receipts, /clear returning the withdrawn text, and publication (hydration,
unchanged-cursor insert, same-frame consume, identity dedup).

* test(native-chat): read the queued receipt ids before the wait closures

* chore(native-chat): SAFETY rationales on the sqlite row casts and a cast-free mobile narrowing

* fix(native-chat): queued-draft bookkeeping never costs a publish, an open, a clear or a history read

- Cache the draft list per draft-table revision. The drain re-checks on every
  journal publish, so each streamed delta was running a SELECT and parsing
  every draft body the handle had ever written (tombstones included).
- Open-time repair/prune failures are reported and skipped; they no longer
  fail opening the chat.
- /clear on a source with no drafts answers exactly as before: no empty
  `withdrawnQueued`, no empty write transaction, no extra publish. A draft read
  failure after the committed clear no longer turns it into a refusal.
- History pages read drafts through the same guarded reader as subscribers.
- Publication moves to its own module; the held-draft rule lives with the
  pause state; one pending-prompt check; drop an export nothing calls.
- Tests: restart-held drafts, pre-consume failure pause + Send retry, failed
  open repair, clear with no drafts.

* fix(native-chat): a Stop that withdraws a consumed draft's send gives its text back

A queued draft converted into a submission leaves the sender's outbox, so when
a Stop withdrew that submission before the agent received it, the text had no
holder: the draft stayed `dispatched` forever and nothing restored it.

- The returned-card rule now follows every effective `rejected` settlement of
  a consumed draft's submission, a Stop's withdrawal included, with the
  withdrawal reason stored as the fact (`dispatchWasWithdrawn`). The writer
  hook and the open-time repair share the rule, so no rejected submission can
  leave its draft `dispatched`.
- A capable Stop withdraws the cards it returned itself along with its
  frontier, stamped with its caller-scoped key: the text comes back once in
  `withdrawnQueued` and replays from the tombstone. An old client's Stop
  leaves a returned card.
- Stop's draft steps move to structured-agent-session-queued-stop.ts.
- Tests: Stop between consume and the agent's receipt for both client kinds,
  its replay, a crash after the withdrawal, restart in the window, and the
  repair of a hookless withdrawal.

* perf(native-chat): the queued-draft drain takes no serialized step while the agent works

The drain was woken by every journal publish and, with a draft waiting, queued
a serialized step (streamed-event flush included) per publish, only to find the
session still working. During a streamed turn that is one step per delta,
contending with Stop and every other mutation for the session's queue.

The pre-check now also skips while the session is working. Whatever ends the
work is itself a commit that schedules again, and the step still re-reads every
gate after its flush, so no wake is lost.

- Test: queued sends during a turn take no drain step; settling the turn drains.

* fix(native-chat): a clear withdraws queued text only for a caller that can take it back; paused reasons are markers

An older client running /clear had its source's waiting and returned drafts
withdrawn and their text returned in a `withdrawnQueued` field it does not
read, so the text was lost. Clear now mirrors Stop: `withdrawQueued: true` on
`agentSession.conversationCommand` (strict params, sent only when the
queued-messages capability is advertised) withdraws the drafts and returns
their text once, replaying from the tombstones. Without it the source keeps
its cards: the supersession fence already blocks the drain, and Delete still
hands the text back.

A paused card's reason was host-authored English on the wire. It is now a
typed marker (`send_failed`) the client localizes, like `returnedReason`; a
client treats an unknown marker as a plain pause.

- Tests: an old client's clear leaves the cards and its replay stays
  field-free, then Delete returns the text; a capable clear returns the text
  once and replays it; the paused marker.

* fix(native-chat): a draft pause that commits no journal row still reaches live subscribers

A pause writes no journal row, so it reaches subscribers only on the next
publish. Two pauses had none behind them: the drain's pre-consume failure
(the session is idle by then, so nothing else commits) and an old client's
Stop that interrupted nothing. A live card kept reading as waiting, with no
failure marker, until some unrelated commit arrived.

The drain now publishes after pausing a draft it failed to convert, and an
old client's Stop publishes when it paused a frontier.

- Tests: a failed conversion and an idle old-client Stop each reach a live
  subscriber as a paused card; both fail without the fix.

* fix(native-chat): a failed clear wakes the queued drain, a failed Stop withdrawal still publishes its pause

A conversation command can settle on the record alone (a retried clear that
fails), so drafts held behind its prepared phase waited for an unrelated
journal commit; the command controller now re-derives the drain when any
command finishes. A capable Stop whose withdrawal write failed never
published the pause it set, and a publish failure after a committed
withdrawal (Stop or clear) dropped the bodies from the answer; publishing now
happens outside the withdrawal and can no longer discard its result. Tests
reset the process-level pause set between cases: operation ids repeat per
test, so a shuffled order held later tests' drafts.

* refactor(native-chat): the draft store notifies through the journal's commit listener, the hold is a stored row fact, and one typed gate decides every queue hold

R1: every standalone draft-table transaction that changed rows (insert,
withdraw, hold, open-time repair) fires the journal's own commit listener
after COMMIT, so a draft or hold change publishes and wakes the drain through
the same path a journal row does — no call site can forget. All hand-written
publish/wake plumbing for draft changes is deleted; wakeQueuedDrain survives
only as the record-input wake (a conversation command can settle on the
record alone).

R2: the process-level pause set becomes a hold_reason column on the draft row
(pre-ship, so no migration): holds survive eviction and restart, keep their
send-failed marker across restarts, die with the session's journal, and are
cleared by consume and withdraw in their own UPDATE. The host-instance
derivation stays the one restart mechanism.

R3: one typed structuredQueueHold (blocked | command | prompt | working)
consumed by admission, the drain step and Send-now, with each caller's
override set written beside it. A capable send during a late-result /compact
now queues instead of being refused (PLAN §3.1); the dead prepared-command
branches and the drain's duplicated gate list are gone. prompt outranks
working so Send-now's one override cannot swallow it.

R4: one isUnsettledQueuedMessage predicate for the withdrawable/budget
filters.

Loop 4: a replayed send whose draft was refused answers with the returned
card, never the rejected submission, so the text cannot render twice. Rewind
completion was verified to publish after the record clears (the rewind path's
own publish; the open path's recovery precedes the open snapshot).

* fix(native-chat): a Stop with no drafts writes nothing, and a failed hold still lets a capable Stop withdraw

The stored hold turned Stop's in-memory pause into a draft-table write, so
every Stop (drafts or not, capability advertised or not) opened a BEGIN
IMMEDIATE/COMMIT. An empty hold now returns before the serialized write.

A hold that threw also emptied the frontier, so a capable Stop withdrew only
returned cards and left the waiting drafts unheld to auto-send after the
interrupt. The frontier is read once and survives a failed hold.

* fix(native-chat): a capable Stop with no drafts writes nothing

The empty-hold guard from the previous fix did not reach withdraw, so every
capable Stop still opened a write transaction after the interrupt, and a
closed handle turned its empty answer into a missing field. The draft store
now answers an empty withdraw without a transaction, for every caller.

* refactor(native-chat): Stop and /clear never withdraw queued drafts; no text rides the wire back

Adopt the host-owned-queue model end to end: a Stop holds the waiting
frontier ('stopped') for EVERY client and interrupts — the cards stay
published as paused, Send-now overrides per card, and the pause dies when
the user next starts a turn (an ordinary dispatched send lifts 'stopped'
holds in the same serialized step; 'send_failed' holds still need their
explicit Send). /clear carries the source's unsettled drafts to the
replacement session as born-held rows — identical for every client
version — then tombstones the source. Delete answers with no body: the
card leaving the published list is the outcome.

Removed (never shipped; the capability was dark and unadvertised, so no
wire compatibility is affected): CancelParams.withdrawQueued and its
refine, ConversationCommandParams.withdrawQueued,
CancelResult.withdrawnQueued, ConversationCommandResult.withdrawnQueued,
AgentSessionWithdrawnQueuedMessage, the Delete result body,
settleStopQueuedWithdrawal and the cancel finisher,
withdrawClearedSourceQueuedMessages, replayWithdrawnQueuedMessages, and
cancelPlan's tombstone replay. This also removes the defect where a
withdrawal took every row regardless of which client sent it (a phone
Stop pulled desktop-typed text): nothing moves text anymore, so a Stop
from one client can never relocate another client's drafts.

Hold and carry writes are bookkeeping: a failure is logged and never
gates the interrupt or the clear.

* feat(native-chat): a restart hold lifts like a Stop's, and paused cards say why

The user's next dispatched send lifts every stop-shaped hold in one
UPDATE: stored 'stopped' rows, and restart-held rows (host_instance
mismatch), which are adopted into the running instance — the same fact
the derivation reads, so no second copy of the hold exists. 'send_failed'
still requires its explicit Send. Publication now marks stop/restart
holds with pausedReason 'stopped' (an additive optional value on a dark
capability), so clients can caption them "sends after your next
message" and keep "couldn't send" for 'send_failed'.

* fix(native-chat): only a client's own send lifts a Stop's queue pause

The lift ran for every accepted host send, so orchestration mail, a
restart continuation and a launch prompt released drafts the user had
stopped (and adopted restart-held rows into the running instance). The
client-facing agentSession.send RPC now marks its sends as the user's
own; host-internal senders leave the pause alone. Also drops comments
still describing the withdrawn return-text rule.

* fix(native-chat): a Stop's queue pause lifts when the user's send starts its turn

The pause lifted as soon as the host accepted a user send, so a send the
provider then refused (a failed child start, a refused turn/start) had
already released the stopped drafts into the same failure. The host now
remembers a client's own send, in memory, until the provider answers it:
acceptance lifts the stop-shaped holds, a refusal forgets it with the
holds intact, and a later Stop supersedes it. Nothing is persisted, so a
restart between the send and its turn start leaves the cards held for the
user's next send rather than sending them unasked.

* fix(native-chat): a consumed draft's turn starting lifts a Stop's queue pause

Drafts are only ever a client's own sends, so a drained draft or a
Send-now is a user send for the pause: its submission joins the same
in-memory set a direct send uses, and the provider accepting it lifts the
stop-shaped holds. Before, a message typed while a stopped turn wound
down drained as a draft and left the older stopped cards held, so their
"sends after your next message" caption was false. A refused consumption
lifts nothing, a later Stop still clears the set, and orchestration mail
and restart continuations still never lift.

* fix(native-chat): queue a capable send behind a /compact and re-scope /clear's carried drafts

- A text send with queue-if-active during a /compact in flight is admitted on the
  compact's side lane as a held draft instead of being refused; it may only become
  a draft, so one the gate no longer holds is refused rather than dispatched.
- Drafts /clear carries to the replacement are fingerprinted for the replacement
  session, so the provider's echo folds into the sent bubble.
- The in-memory set of user sends awaiting their turn is capped; sends settling
  unknown no longer grow it without bound.
- Correct the userSend comment: the renderer's launch prompt goes through the
  client RPC and does set it.

* feat(native-chat): queued mid-turn drafts become editable cards above the composer

Against a host advertising agent-session.queued-messages.v1, Enter stamps the
send 'delivery: queue-if-active' (chat-wide 'Queue follow-ups' setting, on by
default) and the host's published drafts render as compact cards between the
transcript and the composer — never as transcript bubbles — with Steer
(send-now, Cmd/Ctrl+Enter for the newest), Delete, and a menu with Edit message
and Turn off queueing. Returned cards show the stored effective rejection with
the same words a rejected submission gets (a Stop-withdrawn one says so);
paused cards localize the host's typed marker, and an unknown marker reads as
a plain pause. Hold captions are derived client-side; the wire carries none.

Restore is write-ahead: Stop, Edit and a capable /clear (withdrawQueued on
conversationCommand, fingerprint-matched to the host's digest) persist their
operation identity before the RPC and append the withdrawn bodies to the
composer draft exactly once — replays answer from the durable restored record,
and a /clear's text lands in the replacement session's pane. A marker left by
a crash is RELEASED, never replayed: an unadmitted operation-id replay would
execute the command, so a reopened chat can never be cleared, nor new work
stopped, by a press from before a crash; unwithdrawn drafts stay visible as
cards. Text never duplicates: an outbox entry the host visibly holds as a
draft (same id) or answers for in withdrawnQueued retires without a local
restore, and Stop's host-side restore skips ids the outbox withdrawal already
put back.

The renderer carries the list everywhere frames flow: reducer (live over stale
history, omitted means unchanged) and the frame coalescer (latest wins, like
commands). Everything is capability-gated: an older host sees byte-for-byte
today's requests — no delivery key, no withdrawQueued, no queuedMessage RPCs.
The capability stays dark; nothing here advertises it.

* fix(native-chat): queued-draft restore survives a lost answer and an unconfirmed /clear

- A capable /clear reuses the operation id write keeps for an unconfirmed
  clear, so the next press replays it; a fresh id each press was refused by
  the host for as long as the first stayed unconfirmed.
- Edit, Stop and a capable /clear replay a lost answer (the call threw) under
  the same operation id, bounded and in-session, so withdrawn text still comes
  back after the card has gone. A refusal or fence move stays final; a crash
  marker is still only released on remount.
- Restored-id bookkeeping lives in memory beside the draft cache it guards;
  storage holds only in-flight markers, validated per element, removed when
  empty. The /clear marker is written only when the clear actually sends.
- A mid-turn queue send awaiting its answer, or already held as a draft, no
  longer paints as a transcript bubble next to its card.
- One action per card at a time; Edit/Delete hand focus to the composer.
- Revert unrelated en.json reflow.

* fix(native-chat): a lost Stop never lands on newer work; the steer chord never skips typed text

- A Stop whose answer was lost is replayed only while the turn and sends it was
  aimed at are still what is in flight; once another turn opens or a newer
  send lands (e.g. a queued message drained), the Stop is reported unconfirmed
  instead of interrupting work begun after the press.
- Cmd/Ctrl+Enter steers the newest queued card only from an empty composer;
  with text or an image in the composer it stays a plain send.
- A mid-turn queue send hides from the transcript only while it is on its way:
  from the entry the drain is stopped on (read through the drain's own rule),
  sends stay visible as bubbles beside the Retry row. A rejected entry holds
  nothing up, so what follows it still becomes a card.

* fix(native-chat): a send the host visibly holds as a draft frees the outbox's single flight

The published draft list is the host answering the send, exactly as a journal
row is: retiring the in-flight entry now also releases single-flight and voids
the unsettled reply. Before, a slow or lost reply kept the next mid-turn
message waiting, hidden (neither card nor bubble), until the RPC timed out.

* fix(native-chat): Steer hands focus to the composer like Edit and Delete

A steered card leaves the list once the host sends it; focus on its Steer
button fell to the document body, so the next keystroke went nowhere.

* fix(native-chat): a lost /clear stops replaying within seconds, so sends never wait on bookkeeping

Sends are refused while a clear settles. Each clear call can run for its full
195 s timeout, so three lost-answer replays could hold the composer for about
13 minutes. Replays now start only within 10 s of the press: a slow first call
is never followed by more, and at most one replay can outlast the window.

* refactor(native-chat): queued drafts stay paused cards; no draft text ever rides a wire answer

Stop and /clear go back to main's plain writes: the host pauses its drafts and
carries them across a clear, so nothing needs restoring and cards stay visible
on every device. Edit copies the text the card already shows into the composer
before a plain Delete, so no RPC outcome can lose it. The write-ahead restore
journal, replay loops, the Stop wrong-turn guard, and the clear replay window
are deleted with the contract that needed them. Stop's local outbox step keeps
an issued queue send whose answer is still out — the host may already hold it
as a card, and its answer settles it — so the same text can never appear twice.

* test(native-chat): drop the removed tabId option from the queued gating test

* fix(native-chat): paused cards caption per published reason; first card reaches the live region

A Stop's hold ('stopped') says it sends after your next message, a failed
consume ('send_failed') asks for Send, and an absent or unknown marker reads
as a plain "Paused" instead of promising a resume the host may not do. The
live region now stays mounted while empty so the first queued card is
announced.

* fix(native-chat): a paused or returned card's Send tooltip no longer promises to skip a turn

* fix(native-chat): show the queue follow-ups switch only when the host queues messages

The switch rendered whenever structured chat was on, even though a host that
does not advertise agent-session.queued-messages.v1 ignores the preference.
It now reads the local host's capability through the existing structured
host-capability hook and stays hidden until the host says it queues.

The copy now also says that messages with images send right away, since
image messages never queue. Updated in all six catalogs.

* fix(settings): find the Queue follow-ups switch when searching "queue"

The switch renders inside the Chat UI settings entry, whose search keywords
never included "queue", so settings search hid it. Add a localized "queue"
keyword to that entry in every locale catalog.

* fix(native-chat): a returned queued card carries the typed rejection fact, like a rejected submission

A consumed draft the agent never ran comes back as a returned card. The card
kept only the rejection's sentence, while its submission now also records the
typed fact a client classifies from. A host-restart rejection's sentence
carries no legacy marker, so such a card could not be told apart from a
provider's refusal.

The draft table stores the submission's fact next to its reason
(`returned_rejection`, written by the same settlement that sets the reason,
and read back with the reducer's own fact reader), and the card publishes it
as `returnedRejection`. Both are overwritten on every return, so a re-sent
card never keeps an earlier refusal's fact, and a /clear carry inserts a plain
held draft with neither.

Retention moves to queued-message-retention.ts to keep the table module
within max-lines.

* fix(native-chat): say why Stop keeps a dispatching queue send that is not the in-flight one

A pending answer frees single-flight but leaves the entry dispatching until its journal row lands.

* fix(native-chat): word a returned queued card from its typed rejection fact

A returned card is classified and worded exactly as a rejected submission: returnedRejection decides, returnedReason is the fallback. A host-restart card now says Orca restarted instead of the generic not-sent line.

* fix(native-chat): fit the queue to main's typed rejections and compaction result

Main (#23026) dropped the disposition's fresh-id retry field, gives a
rejected dispatch a typed sentence plus fact, and types /compact's result.
The queued-draft disposition and the queue tests now use those shapes.

* fix(native-chat): a returned queued card's words leave out sending again

The card offers its own Send, so its caption is worded with the retry control present, as the delivery notices are.

* fix(native-chat): a queued send in doubt that survives a Stop waits for the user's Retry

The unconfirmed probe resent it onto the session the user had just stopped, starting a new turn when the host never got the first attempt. A Stop now parks it the way a recovered unknown is parked.

* fix(native-chat): a withdrawn send the host returns as a card is not also put back in the composer

When the withdrawn submission and the returned card arrived in one frame, the journal reconcile restored the text before the card retired the entry, so it showed twice.

* fix(native-chat): a send stops asking the host to queue it once the host no longer can

delivery was fixed at enqueue, so after a host rollback every Retry of a queued send was refused on the same strict field. It is now decided per attempt: an id already sent keeps it while the host can read it, an id never sent takes the current choice, and a host without the capability never sees it.

* fix(native-chat): draft bookkeeping can never roll back the journal row it rides

The queued-draft returned transition runs inside every journal append's
transaction. A throw there (a draft table an earlier build created without the
returned_rejection column) rolled back the journal's own rejection row, so a
Stop, a failed start or a provider refusal could not be recorded. The standing
hook now runs in its own savepoint: its failure is logged and rolls back alone,
and the open-time repair re-derives the missed transition from the committed
row. The draft table also gains any missing nullable column at open.

* fix(native-chat): a draft a Stop or restart took back waits again instead of blocking the queue

Cards A, B and C wait; the turn ends and the drain consumes A, but the agent
has not taken it yet. A Stop then pauses B and C and withdraws A's submission,
which made A a returned card. The user's next send lifted B and C, yet a
returned card blocks everything behind it, so B and C never sent although they
read "sends after your next message". A restart or close before hand-over did
the same.

Nobody failed the user there, so the draft now goes back to waiting at its own
position, under the hold that same event put on the drafts behind it: a Stop's
'stopped', or no stored hold after a restart, whose hold derives from the host
instance. It carries no refusal, and records its spent submission id in
consumed_as, so its next consume (the drain, or Send on the card) mints a fresh
id through the same path a returned card's re-send uses. Provider refusals and
other failures still return the card. The live settlement hook and the
open-time repair share one decision. After a Stop and the user's next turn,
A drains first, then B, then C, one per turn.

* fix(native-chat): Delete and Send on a queued card answer at once during a /compact

A /compact holds the chat's serialized lane for its whole provider call, and
the queued-card Delete and Send ran on that lane, so both hung until the
compaction finished. They now run on the side lane a draft-only send already
uses while a compaction is in flight: Delete completes at once, and Send
reaches its readable "wait for the conversation operation" refusal at once.
The drain stays on the main lane and keeps its command hold, so nothing sends
until the compaction settles.

* fix(native-chat): a re-sent returned card drops the refusal it came back with

Re-consuming a returned card left returned_reason and returned_rejection on the
now-dispatched row, so the row described a refusal that no longer applied. The
consume clears both in the same update that moves the card to dispatched.

* perf(native-chat): the queue gate reads pending prompts without rendering the journal

The prompt check ran on every send admission and drain step, and read
journal.snapshot(), which copies and sorts every item in the chat. It now walks
the reduced items in place with journal.visitItems; the answer is the same,
since the snapshot only sorts those items.

* fix(native-chat): a Stop that fails leaves the queued cards as it found them

Stop holds the waiting cards before it withdraws queued sends and interrupts
the agent. When a later step threw or the Stop was refused, the cards stayed
paused ("sends after your next message") although a failed Stop is meant to
change nothing. A failed Stop now undoes exactly what it added: each card it
held gets back the hold it replaced, a consumed card its withdrawal sent back
to waiting is released, and the user sends it had set aside can again lift the
pause. Holds an earlier Stop or a restart put on the cards stay.

The hold SQL moves to its own module, and the draft store's standalone
transactions share one helper.

* docs(native-chat): confirmed cancellation is no longer a queue rollout prerequisite

Stop withdrawing queued sends with a typed cancellation landed on main with
#23026. The comment gating the queued-messages capability now lists only what
remains: the Codex steer matrix (#21062), the Claude fold receipt, turn-owner
bars, and the desktop and phone clients.

* docs(native-chat): the Claude fold receipt and turn-owner bars have landed; Codex steer and the clients remain

* test(native-chat): type the returned-card restore test's hook props

* fix(native-chat): a draft a Stop put back stays visible as a card

The card list hid a waiting draft whose id already had a submission. A Stop that withdraws a consumed draft requeues it under the same id while the first submission stays rejected, so the draft vanished from both the cards and the transcript. Only a submission that was not rejected now hides its card.

* fix(native-chat): Send on a queued card during a /compact is refused before it takes a lane

Send-now chose its lane once, at entry. During a /compact it took the side
lane, where it could wait behind a Stop, then run after the compaction had
settled and append a real submission unserialized against the main lane.
While a compaction is in flight, Send-now is now answered with the "wait for
the conversation operation" refusal before entering any lane, and otherwise it
runs on the main lane. Only Delete keeps the side lane, whose compare-and-set
withdrawal is safe on either.

* fix(native-chat): a Stop that fails after reaching the agent keeps the queue paused

A failed Stop undid its queue holds whenever it threw, including after the
interrupt had already gone to the provider (a status-note write failing after
cancelTurn, or after stopping a starting agent). The turn could be stopped
while the cards drained as if no Stop was pressed. The Stop now marks the step
that reaches the provider, and undoes its holds only when it failed before
that. A Stop the agent refused answers ok and keeps its holds; the comment no
longer claims otherwise.

* fix(native-chat): an unanswered capability probe no longer rewrites a queued send

The per-attempt delivery decision was stored on the entry, so a replay during the window before the host's queued-messages probe answered, or after it failed, was saved without delivery; the host's ledger then refused every later replay of that id. The entry now keeps the user's intent, the wire field is decided per request, a queue send waits while the capability is unknown, and a failed probe is asked again when contact with a remote host is regained.

* fix(native-chat): a skipped draft settlement heals on the next drain step, not only at reopen

The draft settlement rides each journal append as bookkeeping, and a failure
there is logged and skipped. Only the open-time repair re-derived it, so a
consumed draft whose submission was rejected stayed dispatched (invisible, and
blocking nothing it should) until the chat reopened. The re-derivation is now
its own function, shared by the open-time repair and the drain: whenever a
dispatched draft's submission is already rejected, the drain step applies the
owed settlement first.

* fix(native-chat): a queued send in flight when Stop lands is never resent by the probe

Stop parked only sends already unconfirmed; one still dispatching whose answer later came back unknown was left to the unconfirmed probe, which resent it onto the stopped session. The entry now records that a Stop outlived it, the probe skips it, and only the user's Retry, which clears the mark, sends it again. This replaces the retryAfterUnknownSubmittedAt parking for the unconfirmed case.

* fix(native-chat): one id is never recorded as a submission twice

A second submission row under an id the journal already holds replaces the
submission with a fresh pending one, so a rejected message could be handed
over again under its own id. Send on a queued card could do exactly that: if
the host died after it consumed the card under the operation's id but before
its answer settled, the rerun consumed again under the same id.

The journal now refuses a submission under an id it already records, so no id
is delivered twice whatever the caller does. And a Send-now rerun that finds
the card consumed under its own operation id answers with that submission
instead of consuming again.

* fix(native-chat): a waiting draft whose first send the agent echoed is withdrawn, never resent

A consumed draft goes back to waiting when its submission is rejected as never
delivered (a Stop's withdrawal, a restart, a close), and then sends again
automatically. That rests on the "never delivered" claim. If the provider then
echoes that message, the first delivery happened, and the automatic resend
would give the agent the same message twice.

The reducer already keeps such an echo apart, since a rejected submission may
not claim it, so the draft store reads it from the appended row itself: a
provider echo of a user message that no live submission claims, matching a
waiting draft whose spent submission is rejected, withdraws that draft the way
a Delete would. The echo-claiming rule is split out of the reducer's aliasing
so both read the same decision, and the per-row draft hook moves beside the
settlement re-derivation.

* feat(native-chat): a submission names the queued draft it hands off

Clients told a queued card's hand-off apart from other sends by comparing the
draft's id with the submission's id. That holds only for a draft's first
hand-off: a re-send, or a draft that goes back to waiting and drains again,
goes out under a fresh id, and the clients showed the card and the sent
message together, or restored text the host still held.

Every submission the host creates by handing off a draft now carries
queuedMessageId, the draft's id. It is written on the submission's journal row
as an optional key (older readers keep it and ignore it), carried by the
reducer, listed in the published submission schema (which otherwise strips
it), and stamped where the row is built from the consume itself, so no
hand-off path can leave it off; a caller naming a different draft is refused.
A direct send names none. The queued-messages capability comment makes the
link part of v1.

* refactor(native-chat): every queued draft goes out under a fresh submission id

A draft's first hand-off reused the draft's own id as the submission id, so
comparing a draft id with a submission id looked right in every first-send test
and failed only on a re-send or a requeued draft. Every hand-off now uses a
fresh id (the drain mints one; Send on a card uses its operation's id), so id
equality is never true and a reader must use the submission's queuedMessageId.

The host gets simpler: queuedMessageNeedsFreshSubmissionId is gone, consumed_as
is set on every dispatched row and cleared when a withdrawal sends the draft
back to waiting (its spent submissions stay findable by their link), the
consume refuses the draft's own id, and the consumedAs ?? messageId fallbacks
collapse. The delivered-echo check finds spent hand-offs by link.

A send this host queued, asked again (a lost answer's replay, or a rerun the
operation ledger no longer covers), answers from its draft and then from the
hand-off that names it, through one function. The rerun path used to be kept
from sending twice only because a submission sat under the send's own id;
with fresh ids that guard is now explicit. A Send-now rerun recognises its own
consume by the link instead of consumed_as.

* refactor(native-chat): a queued send is matched to its hand-off by queuedMessageId, never by id

The host now hands every queued draft off under a fresh submission id and names the draft on the submission. The card list hides a waiting card only for a live hand-off linked to it; an outbox entry a submission links to belongs to the host in any state (one rule, in a queue-aware reconcile both readers use); a withdrawn hand-off is never restored to the composer; a send answered with the hand-off settles as held; and the Stop and in-flight checks read the same link. This replaces the rejected-submission filter and the published-id restore skip.

* test(native-chat): read the outbox only after the replayed send's answer lands

* fix(native-chat): an echo withdraws a draft only if its rejected hand-off reached the agent

The delivered-echo rule withdrew a waiting draft when a provider echo matched
any rejected hand-off of it, including one a Stop rejected before it was ever
handed over. That hand-off is provably unwritten, so a matching unclaimed echo
is some other message, and the rule silently deleted the card. Only a hand-off
that was handed over and then rejected as never delivered can be disproved by
an echo now.

* fix(native-chat): a skipped echo withdrawal is re-derived before the draft can send again

The delivered-echo withdrawal rides each journal append as bookkeeping, and a
skipped hook left the draft waiting, so it later sent the same message a
second time. Nothing re-derived it. The draft store now also withdraws, in its
owed-settlement pass, each waiting draft that an echo already in the journal
proves delivered: an unclaimed provider user message (still stored under its
own id), carrying the draft's payload, appended after a hand-off that was
handed over and rejected. The live hook and the re-derivation share one
predicate. The pass runs at open and in the drain step, right before a draft
would send; it reads every item, so it never runs per streamed row.

* fix(native-chat): a rolled-back journal append leaves no draft state cached

The draft store caches its row list by revision. The per-row hook read that
list eagerly inside the append's transaction, after the consume in the same
transaction had already written and bumped the revision, so a failed COMMIT
left the cache showing a hand-off that never happened. The hook now reads the
drafts only once a row holds an unclaimed echo, and any rollback of a journal
append or of its bookkeeping savepoint invalidates the cache, so no other read
inside the transaction can leave it stale either.

* fix(native-chat): a replay of a deleted queued card answers withdrawn, not refused

Once a deleted card's tombstone is pruned, a replay of the send that queued it
found the draft through its last hand-off. When that hand-off had been
rejected (the card came back, and the user then deleted it), the replay
answered with the rejected submission, which clients show as a failed send
with a Retry. Only a withdrawn row is pruned while its last hand-off stands
rejected, so the replay now answers queued, withdrawn.

* fix(native-chat): a queued send records what it sent instead of waiting on the capability

The round-two hold kept a queue send back while the host's capability was unknown, which hid its text, wedged every later send, made Retry a no-op and let Stop restore text the host held. There is no hold now: the entry records what its first attempt sent and every replay sends exactly that (dropping it only for a host known not to read it); a first attempt asks to be queued only of a host known to queue with the setting on, and otherwise goes out plain as before; the transcript hides only a send whose request asks to be queued; Stop keeps any queue send that has gone out and parks it, unconfirmed, for the user's Retry, which the drain and an owner change now respect too.

* test(native-chat): a replayed send of a deleted card is spent, with no restore and no Retry

* refactor(native-chat): name the queue's pause-lift for what it releases

* chore(native-chat): one import of the mutation helpers

* fix(native-chat): a Stop marks a queue send without rewriting its state; the entry stores only what it sent

Stop set an in-flight queue send to unconfirmed, so a settled refusal answering its first attempt kept the old id, and every Retry replayed into the same recorded refusal. Stop now only marks the entry, and the drain never admits a marked queued entry, so its state and refusal notice stay what its answer made them. The stored delivery intent is gone: an entry keeps only sentDelivery, recorded by its first attempt; a never-attempted send decides at attempt time.

* docs(native-chat): no send waits on an unknown queued-messages capability

* test(native-chat): one import of the outbox module in the owner-change test

* test(native-chat): read a stale outbox entry without a JSON round-trip

* fix(native-chat): keep a first attempt's recorded delivery a literal

* fix(native-chat): an interrupted send attempted before a Stop reads as unconfirmed, never as not sent

* feat(native-chat): a Stop pauses the whole queue, derived from the journal, with an explicit Resume

After a Stop, each waiting card was held on its own row ('stopped'), lifted
when the host saw, in memory, that a user send made after the Stop had its
turn accepted. The cards read "sends after your next message" one by one,
there was no way to resume the queue without sending something, and the
in-memory record of user sends was lost on a restart or eviction.

The pause is now the queue's, and derived rather than stored as a flag:
- 'stopped': the user's last Stop took effect at a recorded journal position
  and no turn a person asked for has started since. "A person asked for it"
  is the new `origin: 'client'` on the submission row (a send over the client
  send RPC, or a card they sent now); orchestration mail, a restart
  continuation, a host-sent launch prompt and the queue's own drain record
  `host` and never lift it.
- 'restarted': a waiting card was written by another host process and no
  person's turn has started since this conversation opened.
Resume (`agentSession.queuedMessagesResume`) lifts either. Send-now sends one
card; the rest stay paused until that card's turn starts, which is a person's
turn like any other.

The journal's row kinds are closed (an older build truncates a journal at a
row kind it does not know), so the one event the journal cannot carry, where
the Stop took effect, is recorded beside the drafts in `queued_message_pauses`;
everything after it is read from the journal. A Stop records it only once it
takes effect (after withdrawing queued sends, as it reaches the agent), so a
Stop that fails first leaves nothing to undo, and the per-row hold, its undo
and `userSendsAwaitingTurn` are gone. A card keeps a hold of its own only when
its conversion failed ('send_failed').

The pause is published once, as `queuePause` beside `queuedMessages`, on live
frames, catch-up and history. A /clear starts its replacement paused, as after
a Stop, since the carried cards were written for the context it discarded.

* feat(native-chat): a /clear's replacement queue reads paused because of the clear, not an interrupt

The replacement's pause was recorded as 'stopped', which clients show as
"Queue paused because you interrupted" although the user cleared the chat.
It is now its own reason, 'cleared', on queuePause.reason
('stopped' | 'restarted' | 'cleared'). It lifts and resumes exactly like a
Stop's: through Resume, or the user's next turn starting on the replacement.

* feat(native-chat): a paused queue shows one header row with Resume; cards keep Steer

The host now publishes the queue's pause once (queuePause: stopped, restarted or cleared) beside the list, and per-card holds mean only a failed send. The card list shows a header row above the cards naming why the queue is paused, with a Resume button that calls agentSession.queuedMessagesResume (a failure is the usual toast). Cards keep Steer, Delete and More actions while the queue is paused; the old per-card paused caption is gone. Steer's tooltip now reads Submit without interrupting the model.

* test(native-chat): the coalescer keeps the pause of the latest list

* test(native-chat): fit the queue tests to the queue-level pause types

* fix(native-chat): a queue pause covers only the cards it paused

A Stop recorded its pause fact even when the queue had no cards, and the fact
outlived the cards it did pause. The published list hid a pause over no cards,
but the drain still treated the queue as paused, so a card typed much later —
during an orchestration-mail turn, or a correction typed before the stopped
turn ended — sat under "paused because you interrupted" with no Stop of its
own.

A Stop now records its pause only if the queue holds a card when the Stop takes
effect (the hand-offs its withdrawal sent back included). The fact is retired
in the same transaction as the Delete, consume or withdrawal that empties the
queue, never from an async publish. A /clear's carry now lands each card with
its 'cleared' pause in one transaction, so a failed insert leaves no pause over
an empty replacement.

* perf(native-chat): the queue's pause reads the latest person's turn in O(1)

The pause is derived on every publish, per subscriber, and each derivation
copied and scanned every submission to find a person's accepted turn after the
Stop. The reducer now keeps that fact as it folds rows: the submission row of
the latest accepted turn whose origin is `client`. The Stop's and the
restart's lift both read it directly.

* fix(native-chat): each Resume press is its own operation, and a pause-only frame updates state

Resume names no target, so reusing its operation id after a failed press replayed a stale answer or the same refusal. The reducer's no-change check also ignored the queue pause, dropping a frame that changed only the pause.

* fix(native-chat): a card handed off after a restart belongs to the process that sent it

A draft's host_instance was only ever the process that first wrote it (or
adopted it while waiting). A returned card from before a restart, sent again
in this process and then withdrawn back to waiting, still carried the old
process, so it raised a 'restarted' pause although no restart happened since
it was sent. Every hand-off (the drain, Send on a card) now stamps the
handing-off process on the draft in the consume's own update.

* fix(native-chat): the paused-queue row shows only over cards Resume can send, and matches the queue's icons

The header row appears only when a card waits on nothing but the queue's pause; Resume shows it is pending, hands focus back to the composer like the card actions, and the truncated line keeps its full text as a title. A paused queue outranks a pending prompt in the card's hold, as on the phone. Steer carries the corner-down-right arrow, a queued card leads with the list-end glyph, and a card whose send failed leads with the alert, as a returned one does.

* test(native-chat): Resume reports itself in flight until it settles

* fix(native-chat): a queue pause shows only while Resume would send something

After a Stop whose only remaining card was a returned one, or after a restart
with only a card held by its own failed send, the queue published a pause with
a Resume that could send nothing: a returned card waits for the user anyway,
and a held one for its own Send. The pause is now published, recorded by a
Stop, and kept only over a card it can hold back — waiting, with no hold of
its own. The fact is retired in the same transaction as the write that removes
the last such card, a hold or a refusal included.

The publication's dedup also compared only the pause's reason, so a pause
appearing or clearing with no readable reason could read as unchanged; it now
compares presence first.

* fix(native-chat): a queue pause counts only cards Resume would actually send

A waiting card behind a returned one is blocked until the user acts on the
returned card — the drain never sends past it — so a pause over only such
cards still offered a Resume that sent nothing. The rule for "a card Resume
would send" is now one function: waiting, no hold of its own, and not behind a
returned card. The publication, a Stop's record and the fact's retirement all
read it; retirement reads the rows in position order inside the same
transaction as the write that took the last such card.

* fix(native-chat): a returned card that blocks the paused cards hides the pause but keeps it

The last change retired a Stop's pause as soon as a returned card blocked every
paused card. Deleting that returned card then sent the cards behind it at once,
with no Resume — not what the user asked for.

The two rules are now separate. The pause is KEPT (recorded by a Stop, retired
in the same transaction as the write that takes the last one) while any waiting
card with no hold of its own exists, wherever it sits. It is PUBLISHED only
while such a card is not behind a returned one, so the header never offers a
Resume that sends nothing. Deleting the blocking card shows the pause again,
and the cards behind it wait for Resume or the user's next turn.

* test(native-chat): build the pause-only batch through the typed helper

* fix(native-chat): a Stop pauses a card its withdrawal sent back even when that settlement was skipped

The Stop checked the draft table for a card to pause. When the per-row hook
that settles a withdrawn hand-off was skipped, that card was still
'dispatched', so the Stop recorded no pause; the drain later healed it back to
waiting and sent it, although the user had pressed Stop. Retirement had the
same blind spot and could drop a pause while such a card was owed.

What a pause holds back is now one predicate, judged inside the transaction
that records or retires it: a waiting card with no hold of its own (one SQL
EXISTS), or a dispatched card whose consumed submission was rejected with a
settlement back to waiting (read against the journal's submissions). The Stop
first runs the owed settlement, as the drain does; if that fails, the owed
card still counts, so the pause is recorded rather than skipped. recordPause
now checks inside its own transaction and returns whether it recorded, and any
draft-table write (and the per-row hook, the consume and the open-time
repair) retires a pause that no longer holds anything back.

* test(native-chat): pin the per-row hook's pause retirement; skip the judgement when no pause exists

The retirement test recorded its second pause over a queue with nothing to hold
back, so the recording returned false and the "retired" assertion proved
nothing; ablating the per-row hook's retirement passed every test. The hold
case now asserts the pause was recorded, and a new test has a delivered echo,
through the per-row hook, withdraw the last card a recorded pause holds back.

Retirement runs on every appended journal row, so it now checks the pause row
by key first and judges nothing when no pause is recorded. Two comments were
brought in line with the owed-hand-off rule and rewrapped.

* fix(native-chat): the queued area is one bordered box, and the Steer tooltip spaces its shortcut

The pause row, when shown, is the box's first row and each card a row below it, divided rather than individually bordered. The Steer tooltip groups its hint and the shortcut chips with the house gap, so the chips no longer touch the text.

* test(native-chat): match main's append and dispatch shapes in the queue tests

* fix(native-chat): read a compaction's settled submission through the send-result union

* test(native-chat): a queued card Steered into a turn joins it under the opener's bar

A Steer hands the draft over under a fresh submission id, linked by queuedMessageId, and the host scopes its row to the running turn; the transcript keeps it inside that turn with no bar of its own, settled or running.

* fix(native-chat): queued messages sit above running shells and agents, which stay next to the composer

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-29 19:12:26 -07:00
Brennan Benson 59ef74876f fix(terminal): the terminal's owner answers colour queries for the terminal's whole life (#23925)
* fix(terminal): the PTY owner answers OSC 10/11 for the terminal's whole life

Codex and Claude's `theme: auto` ask the terminal for its foreground and
background colours (OSC 10/11) and pick their colours from the reply. Orca
answered in the process that owns the PTY only for agent launches and only
for 5 s; after that the query was handed to whichever viewer was attached.
On Windows ConPTY the owner kept swallowing the query but stopped answering
it, so a Codex started from an older shell tab lost its message shading
(#22332). On a headless `orca serve` host no viewer existed yet, so a Codex
started before anyone attached got no reply at all (#22500).

The owner (in-process provider, terminal daemon, SSH relay) now answers
every OSC 10/11 query for the PTY's whole life and strips it, so no
downstream view ever sees one to answer twice. It answers from, in order:
the host-wide viewer theme pushed to that process, the creating viewer's
colours sent at spawn (now for every PTY, not only agents), and Orca's
default dark theme. The desktop pushes its renderer theme to every owner
on change and on (re)connect: a daemon request gated on protocol v38, and
an SSH relay notification that older relays ignore. The answer-once rule,
the 5 s colour window and the colour-authority handoff are removed; Kitty
keyboard queries keep their startup window.

Viewer-side answerers (renderer xterm, main's hidden-pane model responder,
the mobile webview) stay as the fallback for older owners, which still hand
queries off; they are never reached for a new owner.

* fix(terminal): answer OSC 10/11 with the colours the pane is really painted with

Review follow-ups to the lifetime PTY-owner colour answerer.

- The theme catalog moves to src/shared so the renderer and the PTY owners
  read one source; the owner's last-resort default is derived from it
  rather than copied.
- Main seeds every owner from the host's saved theme settings (light or
  dark, custom themes, colour overrides) at startup, so a headless host
  and a desktop pane that queries before the renderer's first push are
  not told dark to a light-theme user. The renderer's push replaces it.
- Colours an app sets with OSC 10/11, and clears with OSC 110/111, are
  tracked per terminal and reported back, as a viewer paints them; a theme
  change drops them, as a viewer's theme apply does.
- A terminal a paired client created with its own colours answers with
  those, not the host's theme (`colorSource: 'remote-viewer'` on the
  spawn intent), so a light client on a dark host is told light.
- After the 5 s startup window a reply's echo is watched for 512 bytes
  instead of 256 KB, and a torn query candidate is released after 500 ms
  rather than held indefinitely.

* fix(terminal): keep the long echo watch for relayed replies; one theme lookup

The 512-byte post-startup echo watch now applies only to replies the PTY
owner produced itself. A viewer's reply relayed through
answerLiveQueryReply keeps the 256 KB watch, because a cooked-mode app can
keep printing after it queries and the echo then trails that output.

The renderer's getTerminalTheme now calls the shared lookupTerminalTheme,
so the custom-vs-built-in theme lookup exists once.

* perf(terminal): scan colour overrides in one pass over each PTY chunk

Two indexOf searches per OSC went quadratic on long runs of ST-terminated
hyperlinks, and the tracker now sees every chunk of every terminal.

* fix(terminal): one host viewer colour value, set by whichever viewer acted last

A paired client's colours reached the host only as frozen spawn colours on
terminal.create, tagged remote-viewer. UI-started agent sessions on a headless
host answered OSC 10/11 with the host's saved theme, and a client's theme flip
never reached panes it had created.

The host now holds one viewer colour value that every PTY owner answers with.
The desktop renderer's push, a window focus on the host, the new
terminal.setViewerColors RPC, and terminal.create colours from older clients
all set it; equal values do not re-notify daemons or relays. The remote-viewer
tag (colorSource / terminalColorQuerySource / spawnFromRemoteViewer) is gone;
it never shipped in a release.

* fix(terminal): paired clients push their terminal colours on connect, change and focus

The renderer publisher now hands each published fg/bg to subscribers. A new
remote-runtime-terminal-color-push module calls terminal.setViewerColors on
every host this client is connected to when it connects (or the host restarts),
when the colours change, and when the window gains focus. A host that answers
method_not_found or forbidden is not asked again until it reconnects.

The app shell also republishes terminal view attributes on settings and system
theme changes, so a theme change reaches main and paired hosts with no
terminal pane open.

* fix(terminal): a host with its own window answers OSC 10/11 with its own theme

Round 1 kept one host-wide viewer colour value set by whichever viewer acted
last, so a paired client's push (reconnect after sleep, a dusk theme flip)
took over the host desktop's own panes until its window regained focus.

The value is now derived: this host's renderer colours when a local window
has pushed, otherwise the last paired client's push (terminal.setViewerColors
or terminal.create colours), otherwise the saved theme. Only a headless host
takes a client's theme. The window-focus reassert and the identical-re-push
takeover are gone; owners are notified only when the derived value changes.

* perf(terminal): scan only OSC starts for colour queries once the Kitty window closes

The PTY owner answers OSC 10/11 for the terminal's whole life, and it tried
every ESC as a query start: a 240 KB SGR-heavy read cost about 2 ms and 256 KB
of bare ESC about 15 ms, long after startup.

Once the Kitty query window closes only an OSC colour query can match, so the
scan jumps between ESC ] starts, plus a trailing lone ESC so a query torn right
after its ESC still resolves on the next read. Output and replies are
unchanged.
2026-09-29 19:06:57 -07:00