Commit Graph
1367 Commits
Author SHA1 Message Date
Brennan Benson 444e0b1cf9 fix(codex): recognise Codex's quoted spellings in config.toml, and repair Orca's duplicates (#22592) (#23958)
* fix(codex): recognise Codex's quoted project-trust spellings in config.toml (#22592)

Codex's settings screen writes project trust as ["projects"."/p"] and
"trust_level" = "trusted". Orca's matchers only knew the bare spelling, so a
trust write appended a second [projects."/p"] table (or a second trust_level
line) and every codex command then failed with "duplicate key". The config
mirror kept both spellings in Orca-managed homes for the same reason.

- Project table headers are now read through the existing TOML key-path
  parser, so bare, quoted, literal-quoted, mixed and spaced spellings are the
  same table for trust writes and the managed-home mirror/dedupe.
- trust_level is found by decoded key, in both the trust writer and the
  mirror's trust reader, and an existing key is rewritten, never duplicated.
- On the next trust write, a table older Orca appended (exactly
  [projects."<p>"] holding only trust_level = "trusted") that duplicates the
  user's table, or the bare line it inserted under a quoted "trust_level", is
  removed; the user's table wins and the atomic writer keeps config.toml.bak.
  Any other duplicate, or a repair that would still leave one, leaves the file
  untouched and logs once.

* build(cli): list the new Codex trust modules in the CLI project

* fix(codex): recognise Codex's quoted hooks.state spellings and repair Orca's copies (#22592)

Codex writes hook trust as ["hooks"."state"."<key>"] (and the parent as
["hooks"."state"]). Orca's hook-trust writer, parent-table check and mirror
only knew the bare spelling, so a hook-trust write appended a bare copy and
the file failed to parse with "Cannot declare ... twice".

- The hooks.state header, parent-table and mirror checks now use the TOML
  key-path parser, like project tables.
- The duplicate repair now also removes Orca's own hooks.state tables (an
  exact [hooks.state."<k>"] with only enabled + trusted_hash, or an empty
  [hooks.state]) that repeat a table in another spelling, and runs on hook
  trust writes too, so a file with both project and hook duplicates is fully
  repaired. The Orca-shaped copy is removed whichever order the two tables are
  in, only when exactly one other table (the user's) remains; anything else is
  left untouched and logged once.

* fix(codex): carry plain-Codex plugin and project hook trust into Orca's Codex homes (#22592)

Codex keeps hook trust in $CODEX_HOME/config.toml under hooks.state, keyed
by the hook's source. Plugin keys (`id@mkt:path`) and project keys
(`<repo>/.codex/...`) are the same in every home, but the mirror dropped
every hooks.state table from ~/.codex, so Codex inside Orca asked users to
re-trust plugin and project hooks they had already trusted in plain Codex.

- classifyHookTrustKey splits keys into home-scoped (the home's own
  hooks.json/config.toml, re-keyed by install as before) and shared.
- The mirror now carries shared hook trust from ~/.codex in every spelling.
  A key the managed home already holds keeps the managed copy, a key
  repeated in ~/.codex is carried once, and the parent [hooks.state] table
  is never copied, so the result never declares a table twice.
- mergeSystemCodexConfigIntoRuntime moves to codex-config-mirror-merge.ts
  to keep codex-config-mirror.ts under the line limit.
- Tests cover plugin/project carry in each spelling, user-hook keys staying
  out, repeated launches, managed-copy precedence, Windows key spellings,
  parent tables, and user-hook trust re-keying (trusted_hash and enabled)
  from every ~/.codex spelling.

* fix(codex): carry session_end and interrupt hook trust into Orca's Codex homes (#22592)

The shared-trust classifier parsed hook keys with Orca's own trust-key
parser, which only knows the ten events Orca installs hooks for. Keys for
Codex's session_end and interrupt events did not parse, so their plugin and
project trust was treated as home-scoped and left out of the managed home.

The classifier now reads the source path from Codex's key shape
`{source}:{event}:{group}:{handler}` for any event label. A key without
that shape is still never carried. Tests cover both events for plugin and
project keys in both spellings, user-layer keys for both events, and five
unattributable key shapes.
2026-09-30 00:09:26 -07:00
Brennan Benson 5b3366f78e test: unit tests can no longer write a developer's real agent or Orca settings (#23979)
* fix(agent-trust): write per-user trust under the home the launched agent reads

The Cursor, Copilot, Qoder and Antigravity writers and the local Codex config list
resolved ~ with os.homedir() at write time, so any test that reached them wrote into
the developer's real ~/.codex, ~/.cursor, ~/.copilot or ~/.gemini. Each writer now
takes the home, derived once from the launch env (HOME, or USERPROFILE on Windows,
else this host's home) by launchedAgentHome, which the SSH relay already used.

* test: give tests that wrote the real agent or Orca home a temp one

The structured Codex adoption replay pre-trusted /repos/workspace-1 in the real
~/.codex/config.toml; it now runs with a temp HOME and userData. The Codex
session-resume and WSL hook tests created Orca's managed Codex home under the live
userData, and the Claude Agent Teams tests wrote their tmux shim into ~/.orca; each
now runs against a temp userData or HOME.

* test: fail any unit test that writes the real agent or Orca home

A vitest setup file wraps the node:fs mutating calls and refuses a target under the
account's real ~/.codex, ~/.claude(.json), ~/.orca, ~/.cursor, ~/.copilot, ~/.gemini,
~/.qoder or Orca userData, found through os.userInfo() so a test that swaps HOME
cannot hide it. The refusal is recorded and rethrown after the test, since trust
writers swallow errors. Reads are untouched. It stands down only while an opted-in
real-agent suite's own switch is set. It also unsets what an Orca terminal exports
toward the live app (userData, Codex and Claude homes, and the Codex launch preflight
CLI, which a shell test would otherwise run), so a local run matches CI.

* test: type the guarded fs call from its narrowed original
2026-09-30 00:02:35 -07:00
Neil d414033400 fix(packaging): stop shipping the relay bundles inside app.asar (#24027)
resources/relay is the only relay copy a packaged build resolves, but out/relay
was also packed into app.asar — 14.2MB of unreachable duplicate. Kaspersky
flagged app.asar as a compound object precisely because relay.js was inside it,
so one script-heuristic verdict on relay.js gutted the whole install. Excluding
it decouples app.asar from that verdict and drops the duplicate bytes.
2026-09-29 23:30:46 -07:00
Brennan Benson e594cb06af test(mobile): record RPC goldens without a pinned commit, and check recorded requests against the desktop's params rules (#23732)
* test(mobile): add rpc:diff to decode what a recording change moved

The RPC recording goldens are content-addressed JSON, so their raw git diff is
pool hashes. `pnpm --dir mobile rpc:diff [<base>]` decodes both sides and prints,
per golden, the checkpoint, field and JSON path that moved with both values,
grouped across checkpoints, plus added and removed goldens. `--summary <file>`
appends a Markdown report capped for GitHub's step-summary limit.

It reads any pooled format, so it can prove the next commit's format change
moves no recorded value. Checkpoints are matched by occurrence because an id can
repeat within one golden.

This commit adds files under the recorder directory, which moves the header
digest every golden pins; the next commit removes that header.

* test(mobile): record RPC goldens without a pinned commit or input digests

Every golden carried a pinned `baseline` commit plus digests of the recorder,
its mount adapter and its scenario, and the record script refused to run unless
the product tree matched the pin. So every behaviour change repinned to its own
branch commit and rewrote all ~790 files, the squash made that commit
unreachable, and main's pin job stayed red until a hand-made repin pull request
landed (22 of them in 12 days). The digests could only fail when an input moved
and the recording did not, which is exactly the change that carries no
information; every run already re-derives each golden from the current tree and
compares it.

Format 6 keeps the format version, operation, family, named deltas, the value
pool and the recording. Removed: the pin and fence, the three digest modules and
their test, the pin guard and its CI job, and the dead scenario `version` field
(the manifest reader now refuses `baseline` and `version` with a message).

- `pnpm --dir mobile rpc:record [<golden-id>...] [--prune]` records all or some
  goldens; orphans are listed, and deleted only with `--prune`. Every derived
  test title now starts with its golden id so an id selects it.
- `compareGolden` reports every difference in one failure (identity fields by
  name, the checkpoint list, each checkpoint/field/path grouped), keeps the
  final byte compare, and ends with the command to re-record that golden.
- `unhandled-recording.test.ts` now drives a detached rejection through
  `runRecording` into a checkpoint and the cleanup checkpoint; no golden carries
  one, and disconnecting the capture passed every suite before.
- Seam rules that existed only to keep a digest honest are gone; the
  mutant-reachability, register-completeness and one-exposure rules stay.
- CI: `Mobile tests on main` runs the whole mobile suite on every merge that
  touches mobile/, src/shared/, the root lockfile or the host RPC paths, since
  `verify` never runs on main. A new `Mobile RPC Recording Replay` workflow
  replays the recordings on pull requests that touch src/shared/ or the root
  lockfile without touching mobile/. `verify` writes the `rpc:diff` report to
  the job summary.

Proof: `rpc:diff` against the parent reports no recorded behaviour moved; each
golden only loses its ten header lines.

* test(mobile): check every recorded request against the host's params contract

The goldens script the host's replies, so a scenario could record a success
for a request the real host would refuse, and a desktop change that tightens a
params schema moved no golden at all.

`recorded-request-params.test.ts` parses every distinct request the corpus puts
on the wire with the host dispatcher's own `parseRpcRequestParams` and the
schema `rpc-params-catalog.generated.ts` binds to that method. It fails on a
method the host lacks, params it refuses, params sent to a method that takes
none (the dispatcher never reads them), and keys the schema silently strips
unless an inventory entry gives the reason; a stale entry fails too. Each rule
is also shown firing on a made-up request, since the corpus has no instance of
three of them. It imports the desktop dispatcher, so it sits beside the other
Node-side tests outside the RN test program, and the params-contract boundary
now exempts test files, which are never bundled.

It found twelve requests the host would refuse, all from invented fixture
values, not product code, fixed at their source:
- git.branchDiff sent `base-oid`/`head-oid`/`merge-base` where the host needs
  full object ids (diff-review and source-control adapters, and the branch
  compare replies in the manifest that feed them);
- an iOS push registration without `apnsEnvironment`, which a real iOS token
  always carries (`push-token.ts`); the adapter now defaults to `production`;
- `settings.update` given Linear's `assigned` filter as a GitHub preset, which
  the product type forbids; the scenario now picks `my-issues`;
- GitLab `projectRef` as a string where the host and the product type take
  `{ host, path }` (7 methods, 5 adapters and the manifest).

46 goldens move, and a decoded comparison of every one of them shows no change
other than those substitutions; `rpc:diff` lists them.

* ci(mobile): detect a mobile change without a SIGPIPE-prone grep pipe

Under the runner's pipefail, grep -q exiting on its first match SIGPIPEs git
diff on a long file list, so a large pull request touching mobile/ read as
uncovered and replayed the recordings a second time.

* test(mobile): drop comments that still describe the golden header and digests

Eleven adapters justified an import rule by the header a golden no longer
carries, and that rule's test is gone. The census failure now names the
rpc:record and --prune commands.

* test(mobile): refuse a golden that keeps a key no recording writes

Decoding dropped unknown top-level keys, so an old header left behind by a
hand-resolved merge conflict passed every compare unseen.

* ci(mobile): summarize RPC recording changes after a failed test step too

* test(mobile): stream rpc:record output instead of capturing it

A captured run stayed silent for its whole duration and clipped its tail,
where the failure summary sits, past 8 MB.

* test(ci): let the Ruby-gate contract skip the always-run RPC summary step

fef088d8f4 gave the summary step an `if: ${{ !cancelled() }}`, and this test lists every gated step in `verify` and expects each to be gated on the Ruby scope.

* test(mobile): replay only a golden file that is exactly what rpc:record writes

Replay compared two re-encodings of decoded values, so anything decoding drops (a leftover header
key, a hand edit) sat in the committed file uncompared; a key allow-list covered one case of that.
Replay now passes only if the file text equals the formatted golden for the run, sharing one
formatter with writeGolden, and keeps the field-level report as the failure message. The allow-list
goes; the value-based compareGolden stays for the bridged run, which has no file.

* test(mobile): end a corrupt or hand-edited golden's failure with the re-record command

A hand edit to a pooled value failed in decode with only "Golden value <hash> does not hash to its
pool key": no golden id and no command to fix it. readGolden now prefixes parse and decode failures
with the golden id and ends them with the rpc:record command. The rpc:diff header also said it
always exits 0; it exits non-zero when git or a golden cannot be read, and now says so.

* ci(mobile): run Mobile Checks on every src/shared and root lockfile change

Replaces the replay-only workflow: mobile imports hundreds of shared modules, so a shared edit can
move a golden or break mobile's typecheck, and the full job catches both before merge. A root
lockfile-only change skips the Ruby release checks, which read no root Node dependency.

* test(mobile): list or prune orphaned goldens even when the recording run fails

Orphans come from the manifest, not the run, so a failed or timed-out rpc:record still reports
them; the exit code stays non-zero. README: say what a failed replay reports (first differing
path per field, capped groups) and what rpc:diff compares with and without a base.

* test(mobile): pin the RPC recording goldens to LF so a CRLF checkout still replays

Replay now requires the committed golden text to equal exactly what rpc:record writes, which is
LF. A Windows checkout with core.autocrlf=true converted every golden to CRLF and failed all 790
with "holds the same recording but is not the file rpc:record writes for it".
2026-09-29 23:26:55 -07:00
Neil fb52c0602a fix(terminal): release xterm's DEC 2026 render hold instead of waiting out its 1s timeout (#23920)
* fix(terminal): release xterm's DEC 2026 render hold instead of waiting out its 1s timeout

xterm paints nothing while DEC mode 2026 (synchronized output) is open and only
force-flushes after 1000ms. Codex wraps every draw in mode 2026, so any byte gap
or chunk split that loses the closing \x1b[?2026l freezes the pane for a full
second and then repaints in one burst.

Orca never emitted \x1b[?2026l anywhere, and three paths could destroy a TUI's:
the per-PTY pending cap drops buffered output wholesale (mode 2031 was already
salvaged there, 2026 was not), main sliced pending data at a blind 16KB offset
that can land inside an open frame or sever the 8-byte marker, and the renderer's
backlog warnings replace a queued tail that may hold the close.

- salvage the 2026 latch across dropped output, mirroring the existing 2031
  salvage, and append the release on both delivery sites
- ground 2026 in RESET_AFTER_BYTE_GAP and the replay baseline, and in both
  backlog warnings, so every drop path is self-healing
- make main's 16KB flush split frame-aware instead of a blind byte offset
- lift the synchronized-output scanner into shared/ so main and the renderer
  use one implementation

Closing a frame early costs one premature repaint; leaving it open costs a
second of blank screen, so the asymmetry favours always closing.

Also adds the reproduction this needed: the pre-existing typing bench observes
the xterm BUFFER, which the parser fills while rendering is held, so it scored
these freezes as fast echoes.

* fix(terminal): stop the renderer's queue drain cutting inside an open DEC 2026 frame

takeQueuedChunk sliced a queued chunk at a blind byte offset to fit the 16KB
coalescing budget, which can strand a frame's closing \x1b[?2026l in the residual
until a later drain. Same defect as main's flush split, same fix: reuse the
frame-aware split helper.

Usually masked because the drain coalesces adjacent chunks and reassembles what
main split, but not when the budget boundary falls inside a frame.

* fix(relay): keep the SSH path's bounded slice outside an open DEC 2026 frame

pty-handler split pending output at a byte offset with a surrogate-pair guard but
no synchronized-output awareness, so a frame straddling the 16KB wire slice had
its closing \x1b[?2026l stranded in the remainder — the same defect just fixed on
the local path, on the path AGENTS.md requires us to consider.

Placed before the surrogate guard so that guard keeps the final say, and floored
at 2 so frame alignment can never walk a healthy slice into the guard's
decrement and then into the chunkChars <= 0 pause-and-retry path.

Also drops a dead `splitAt === 0` branch in takeQueuedChunk: both callers pass a
positive limit and the helper never returns 0 for one.

The two new split tests were each confirmed to fail without their fix.

* test(terminal): sweep the DEC 2026 split helper over escape-sequence shapes and every limit

Covers OSC 52, DCS, repeated open/close markers and limits 1..len+3, asserting the
result never exceeds the limit, never reaches 0, and stays byte-exact. Also pins
that a buffer beginning inside an open frame degrades to the blind offset rather
than doing something worse, and documents that callers do not thread latch state.

* fix(terminal): ground DEC 2026 on the daemon slice, the recovery replays, and the process boundary

Four more sites could strand the latch, found by sweeping every path that drops,
splits, or replays terminal bytes.

- daemon-stream-data-batcher: the 64KB bulk-write slice used a surrogate-only
  clamp, and its remainder is HELD until 'drain' — "seconds for multi-MB
  backlogs" per the file's own note. A frame straddling that boundary parked its
  \x1b[?2026l behind the hold, blanking the pane past xterm's 1s timeout once per
  frame for as long as the backlog lasted. This is the default daemon-backed pane
  path, so it is the one users actually hit. The new
  clampToSafeBulkWriteSplitIndex frame-aligns first and surrogate-clamps last,
  and lives in daemon-stream-data-split alongside the policy it belongs to.
- replay-data-drain and remote-runtime-terminal-binary-snapshots wrote a bare
  \x1b[2J\x1b[3J\x1b[H, which does not clear mode 2026 — so on the SSH/remote
  reconnect path, the very event most likely to sever a frame, the whole replay
  could paint nothing.
- ipc-pty-attach: trimIncompleteTerminalControlTail can cut a half-written
  \x1b[?2026l while its opening marker survives in the replayed prefix.
- PROCESS_BOUNDARY_GROUND: the "process that armed these modes is gone" ground
  omitted 2026, the last unexplained gap in that file. A disable, so it still
  satisfies the recovery barrier's ownership scan (only ?25h may be an enable).

Recovery-path expectations updated where they pin the emitted bytes. Deliberately
NOT touched: apply-reattach-payload and ssh-snapshot-prepaint already ground via
buildSnapshotReplayPrologue.

Still unfixed, deferred with reason: terminal-output-frame-chunks.ts splits the
remote wire on accumulated UTF-8 byte width and needs a different shape than the
char-index helper; desktop clients reassemble in main's pending buffer, so the
exposure is mobile/web only.

* fix(terminal): emit the DEC 2026 release before the mode-2031 tail, and stop claiming the drop path writes it

Two corrections from adversarial review of the earlier commits.

1. Ordering bug I introduced. getDroppedMode2031RendererData ends with
   `state.tail`, which extractPrivateModeScanTail deliberately retains as an
   INCOMPLETE private-mode sequence so the next chunk can resolve it. Appending the
   2026 release after it put an ESC behind a dangling CSI, aborting it and silently
   losing whatever mode spanned the drop boundary. The release now goes first.

2. The drop-path release does not reach xterm in the dominant case, and the comment
   now says so instead of implying otherwise. live-data-callback's droppedOutput
   branch discards `data` and salvages only queries
   (salvageRendererQueriesFromDiscardedRestoreData handles CPR/DA1/OSC colour;
   \x1b[?2026l is not a query), so for hidden panes and visible panes outside
   foreground-restore backpressure the synthesized release was dropped. The grounded
   snapshot replay releases the latch instead.

   I tried writing it through writePtyOutputToXterm there and reverted: it consumes
   the pending hidden-output snapshot and broke
   pty-connection-hidden-snapshot-resize-signals ("re-restores a skipped alt frame"),
   so the release rides the restore rather than perturbing that state machine.
   Residual gap, documented: a cap-dropped pane whose restore never arrives.

The salvage is still load-bearing on the fall-through path, so it stays.

* fix(terminal): release DEC 2026 on the reattach clears, floor the split, and correct the freeze framing

Remaining findings from adversarial review.

- apply-reattach-payload's three bare-clear branches (:63 daemon snapshot, :229
  relay replay, :269 cold restore) had no release anywhere in their sequence: I
  checked all seven POST_REPLAY_* profiles reachable via chooseReattachReplayReset
  and none contains \x1b[?2026l. Only the buildMainModelSnapshotReplayWrites branch
  was grounded, so covering the streamed replay path and not the main reattach path
  was inconsistent. Verified no production code matches these clear strings — the
  three test updates are mock equality, and each was confirmed to fail without the
  source change.
- clampToSafeBulkWriteSplitIndex could return 0 (('\u{1F600}aaaa', 1) — alignment
  returns 1, the surrogate clamp decrements to 0), which would leave a zero-length
  slice that never shifts the batcher's queue entry and spin its drain loop.
  Unreachable from today's only caller, but it is exported with an unstated
  precondition. Floored at 1.
- Frame alignment could halve per-PTY flush throughput: main re-queues the
  remainder with eligibleRound = round + 1, so the shortfall cannot be refilled in
  the same round, and aligned size is floor(W/F)*F — 50% worst case in the 8-16KB
  band, which is exactly the full-screen redraw burst that reaches the pending cap.
  Alignment is now rejected below half the window, preferring throughput and
  letting the reset profiles release the latch.

Framing corrected throughout: bufferRows records a row range and clears nothing, so
the pane freezes on its last painted frame — it does not go blank. The real trade is
"stale but coherent for <=1s" versus "immediate partial frame", and
RESET_AFTER_BYTE_GAP (written alone, with no repaint behind it in the same write) is
the one site that can newly flash a partial frame. Said so at the constant instead
of implying the release is free.

* fix(terminal): rename the shape-flagged symbols the anti-slop audit rejects

CI's anti-slop gate rejects "shape" in symbol names as structural rather than
domain language: `shapes` -> `outputSamples`, and
`writeCodexShapedEchoProbeScript`/`codexShapedEchoProbeScript` ->
`writeCodexEchoProbeScript`/`codexEchoProbeScript`.
2026-09-29 20:27:30 -07:00
Neil 45c63a66e9 test: delete the source-grep tests an earlier detector's regex missed (#23976)
A rebuilt detector found 195 source-grep candidates where the original found 111.
The gap was one over-specific regex: the first scanner required a literal `.ts`
path inside `readFileSync(...)`, so every test that built its path from variables
(`join(dirname, '..', 'foo.tsx')`) was invisible to it. Roughly 84 files of a
pattern an earlier wave reported as cleared had in fact survived.

Deleted whole, every case asserting on production source text:
- `app-startup-routing.test.ts` (27 cases) — exact import statements
  (`"import('../components/UpdateCard').then"`), relative-path spelling, and
  `indexOf` source ordering. A file move or a `lazy()` refactor breaks it.
- `pull-request-page-host-boundary.test.ts` (13) — `toContain` on whole argument
  expressions concatenated across 20+ component files.
- `SmartWorkspaceNameField-source-boundaries.test.ts` (7) — placeholder copy, a
  Tailwind class string, and `not.toContain` on an already-deleted symbol.
- `github-project-repo-list-load.test.ts` (9) — `indexOf` statement ordering
  inside `loadTasks`.
- `github-enterprise-slug-routing-boundary.test.ts` (4) —
  `toContain('host: githubProjectHost(parsed?.slug.host)')`.
- `web-viewport-shell.test.ts` (3) — a regex demanding exact CSS selector-list
  ordering and whitespace.
- `agent-catalog-links.test.ts` (1) — restates two `homepageUrl` literals straight
  out of `agent-catalog.ts` with nothing in between.

Trimmed, keeping only what nothing else can reach:
- `desktop-startup-ordering.test.ts` 549 -> 66 lines, retaining the three cases
  named as `assertionRefs` by the `ssh-filesystem.stream-inactivity-lifecycle` and
  `agent-browser.owner-boundary-cleanup` gates; 15 source-order greps went.
- `ResourceUsageStatusSegment.session-polling.test.ts` keeps its census that no
  `setInterval` exists and `listSessions()` is called exactly once — an added poll
  multiplies a global daemon scan and no behavioral test sees it. The
  `indexOf('if (!open)')` ordering pair and four `not.toContain` lines went.
- `agent-skill-installed-command-callers.test.ts` 231 -> 86, keeping the
  `readdirSync` census that discovers every `<AgentSkillSetupPanel` caller and
  asserts set equality against the allowlist, so a new panel host cannot silently
  show a default Update action.

Also in this wave, from the renderer lib/runtime sweep: 22 cases whose routing
signal the production path never reads — verified by mutation, stripping
`connectionId`, the WSL preference and the UNC path from four of them left all 29
tests passing — plus braille-spinner rows collapsed onto one regex range, copied
`WELL_KNOWN_LABELS` rows, and a whole `resolveAiVaultResumeStartupShell` describe
whose four darwin/linux fixtures all return before the login shell is read.

`config/reliability-gates.jsonc` drops the two `app-startup-routing.test.ts`
references; the manifest still validates for 140 gates.
2026-09-29 19:55:50 -07:00
Brennan Benson ad2e1b5efa fix(terminal): restore the mouse format with mouse tracking, so phone swipes don't type into Codex (#23946)
* fix(terminal): restore the mouse encoding with mouse tracking in every snapshot

Swiping to scroll Codex from the phone on a Windows host typed legacy
`ESC [ M` mouse reports into the Codex composer (#23818). SerializeAddon
re-arms mouse tracking (?1000h/?1002h/?1003h) but never the SGR encoding
(?1006h/?1016h). Any snapshot taken from a desktop pane's xterm (the
runtime seeds its headless model from it after a reattach, and serves it
to remote viewers when no model exists) therefore restored "tracking on,
legacy encoding", and the phone encoded wheel events as X10 bytes, which
ConPTY hands to Codex as keystrokes.

serializeWithAbsoluteCursor, the one wrapper every Orca snapshot producer
uses, now appends the encoding xterm itself parsed, read from xterm's
mouse state service. The daemon/runtime headless model reads tracking and
encoding from xterm too, so its regex mirror of the DECSET stream is
deleted (one source of truth; one less regex pass per PTY chunk).

Mixed versions: no wire field changes. A new host's snapshot carries an
extra DECSET that old desktop and phone clients already parse; an old
host's snapshot restores exactly as before. With tracking off the encoding
alone sends no reports, so the wheel still scrolls scrollback.

* test(terminal): pin the mouse-encoding read against the renderer xterm build

* fix(terminal): type the xterm mouse-state read behind named shapes
2026-09-29 19:23:25 -07:00
Jinwoo Hong 26bb7c23f1 fix(terminal): run Orca's cmd.exe, path-named and setup-gated Codex launches without the shared server (#23933)
* fix(terminal): give plain shells and cmd.exe Codex launches --no-daemon

Plain bash, zsh and fish tabs were never wrapped, so a typed codex skipped the
shell function that adds --no-daemon. Wrap them (bash keeps its prompt and
DEBUG trap untouched unless Orca asked for command markers), add --no-daemon
host-side where no function can run (cmd.exe, path-named binaries), and move
new tabs to a v38 terminal daemon so they get the new wrappers.

* fix(terminal): keep plain bash a login shell and give the setup gate the codex function

Plain bash and Git Bash tabs launch exactly as before again: the rcfile
wrapper would have made every one a non-login shell. Plain tabs on the
user's configured shell args stay unwrapped on both transports. The
wait-for-setup gate's bash -lc now defines the codex function, so a
sequenced Codex launch gets --no-daemon from the binary it actually runs.

* fix(terminal): define the setup gate's codex function after setup finishes

Setup can be what puts codex on PATH, so defining the function before the
marker wait found no binary and skipped --no-daemon.

* refactor(terminal): fold the SSH/WSL guard into the Codex launch planner and bound the gate test

* fix(terminal): honour the pane's env deletions in the Codex opt-out check

Also pin the setup-gate test's fake codex ahead of path_helper's PATH.

* revert(terminal): launch plain zsh and fish tabs exactly as on main

Drops the always-wrap for plain zsh and fish, the configured-args guard
that only served it, and the v38 daemon bump: the daemon's launch configs
and generated wrappers are byte-identical to main again. Keeps the
host-side --no-daemon for cmd.exe and path-named launches and the setup
gate's codex function.
2026-09-29 21:40:53 -04:00
Brennan Benson 1aa943798a fix(codex): a Codex Stop is the interrupt alone, so it no longer says "Cancellation was not confirmed" (#23850)
* fix(native-chat): a Codex Stop that ended its turn no longer reads "not confirmed"

Codex answers a turn's interrupt only as that turn ends, then sends the turn's
own end. Orca also required its sweep of the turn's processes to prove them
gone, and when that sweep could not read the process table it wrote
"Cancellation was not confirmed." a few milliseconds before the turn's
interrupted end arrived. A Stop that named its turn said "The provider had
already finished this turn." in the same case.

Codex's answer is now what confirms the Stop. The sweep still runs and still
holds the turn's end until it finishes, but its result is logged, not shown.
A Stop Codex never answers, which is a turn that never ends, still reads
"Cancellation was not confirmed."

* docs(codex): say why an interrupted turn's un-echoed send is withdrawn, for a steer and for a turn's own input

* fix(codex): a Codex Stop is the interrupt alone

A Codex Stop sent turn/interrupt and, alongside it, swept the turn's
processes: it compared a snapshot of the app-server's descendants taken at
each send and compaction against a fresh one and killed what was new, then
waited for those processes to exit. The turn's end was held back until the
sweep finished.

Codex already owns this. It kills a turn's one-shot commands on interrupt
and deliberately keeps unified-exec background terminals running across one,
so the sweep killed work Codex means to keep, read the process table on
every send, and delayed the interrupted end behind a kill-and-wait.

A Stop is now turn/interrupt only. Codex's answer confirms it and the turn's
end publishes the moment it arrives, through the same path as any other
notification. A long-running command started in that turn keeps running
until the thread ends, as it does in Codex itself.

Removed: the turn-process module and its integration test, the snapshot
taken at dispatch and compaction, the held turn/completed, the adapter
dependencies that injected both, and the prompt-claim re-check that only
covered the wait for the snapshot.

* docs(crash-reporting): justify the self-kill ring size by the documented window-close burst
2026-09-29 18:13:18 -07:00
Jinwoo Hong 9f4311598f fix(codex): trust the worktree Codex starts in, not a guessed repo root (#23937)
* fix(codex): trust the path Codex checks for bare-repo worktrees

Codex keys a linked worktree's trust on the main checkout only when that
checkout's .git leads back to the common git dir; otherwise (bare repo,
--separate-git-dir) it keys on the worktree itself. Orca always wrote the
main-checkout key, so Codex showed its trust prompt and worker-start
failed at agent_readiness.

Mirror trust.rs exactly, and pin it with a real-binary contract that
runs in the existing Codex contract CI job.

Fixes #23847

* fix(codex): trust the worktree path itself instead of mirroring trust.rs

Codex looks up the cwd's own [projects] entry before any repo root
(config_toml.rs get_active_project, loader decision_for_dir), so trusting
the workspace realpath satisfies every git layout. Drops the
resolve_root_git_project_for_trust mirror: simpler, cannot drift from
Codex, and never widens trust past the folder Orca launched in. Cost is
one config entry per worktree; entries older Orca wrote on main
checkouts stay valid.

The six real-git layout tests now assert the workspace key and, under
the contract, that real Codex starts workspaceWrite for each. The
contract probes the binary version once and fails at load when required
but missing; its CI path filter now includes config-toml-trust.
2026-09-29 20:27:02 -04:00
Neil fed1eca486 test: stop restating internal tuning constants, keep the ones that are contracts (#23950)
Removes ~74 assertions of the form `expect(SOME_CONSTANT).toBe(<literal>)` where
the literal is an internal tuning value — a timeout, retry count, debounce
interval, cache TTL, circuit-breaker window, Tailwind class string. Those cannot
fail for any reason a user would notice: they fail only when someone deliberately
changes the number, and then the test is simply updated. They are copies of the
declaration.

The same pattern is NOT junk when the exact value is observable outside this
process, so those were deliberately kept:
- terminal byte contracts: `\r`, `\x03` ETX, Kitty escapes, `\x1b[?1;2c`;
- wire and capability values: `agent.launch.v2`, protocol 3 / min-compatible 2,
  daemon per-feature boundary versions (a daemon survives app updates, so those
  pin what an old field daemon may be trusted with), relay header tokens;
- security invariants: the `127.0.0.1` bind default, an empty iframe `sandbox`;
- values external processes read: exit code 78 (EX_CONFIG) and exit code 3
  (systemd `RestartPreventExitStatus`), `ORCA_AGENT_SESSION_SPAWN_TOKEN`,
  `npx skills …` commands users paste, on-disk journal schema versions,
  the `orca_<hash>` filename prefix the fish sweeper matches;
- third-party names: expo-router's `unstable_settings` / `ErrorBoundary`,
  iOS Safari's 16px zoom threshold.

Where a case asserted a relation rather than a literal — `A < B`, a sum of parts,
a cap compared against a sibling budget — the relation stays and only the literal
went.

Test-only changes: no production file is touched and no test file is deleted.
2026-09-29 17:06:09 -07:00
Brennan Benson 59b746ff3c feat(native-chat): one structured-chat journal database per host, owned by one process (#23613)
* feat(native-chat): one structured-chat journal database per host, owned by one process

Every structured chat on a state directory now lives in one SQLite file,
agent-session-journal.db, opened once by the process holding
agent-session-journal.owner: an empty SQLite file whose held BEGIN EXCLUSIVE is a
kernel byte-range lock, refused while another process holds it and released when
the holder dies.

- Stores own no connection: the per-chat handle, its close contract and the
  close-retry registry are gone; closing a conversation drains its writes, and the
  one connection closes last at teardown.
- The owner lock is taken at runtime start, before orca-runtime.json is written;
  a process that does not own the chats is not published and refuses every
  structured request with journalUnavailable and words that say what to do. It
  retries the lock with backoff and runs the full install once it holds it.
- A journal that will not open fails the host install: every chat says "Unable
  to load this chat." (journalCorrupt), and nothing is renamed, deleted or
  rebuilt. A newer build's database is refused and left byte-identical.
- An append is one INSERT. The listing status is a column, written after the
  rows it describes and keyed by (epoch, sequence).
- A per-chat journal from an earlier build is copied in verbatim (epoch UUID and
  every sequence) on that chat's first open, and its directory is retired only
  after the copy commits.
- auto_vacuum = INCREMENTAL, with freed pages handed back in bounded steps after
  every delete.

* perf(native-chat): key journal rows by block so one chat's rows sit together

Each chat's live epoch owns a block of row ids, block * 2^32 + seq, so a chat's
rows share leaf pages with nobody else's, a replay is one range scan, and
replacing or rewinding a chat deletes one contiguous range. Measured on the
largest real chat (61 MB) beside 19 interleaved peers: 39 ms and 7.5 MB of WAL,
against 214 ms and 102 MB for a (session_id, epoch, seq) key.

- Ids are computed in Number arithmetic, never bitwise. A sequence is refused
  outside [1, 2^32) and a block at 2^21, which keeps every id below 2^53.
- A replace, roll or import allocates a fresh block, moves the chat's pointer,
  and deletes the old block in the same transaction, so no orphan block exists.
- The listing status write moves into its own writer beside the column.

* feat(native-chat): copy a chat's per-chat journal again when an older Orca wrote it after a downgrade

A per-chat journal.db that reappears after its chat was copied in is the newer
history: an older build, run after a downgrade, attached the chat and wrote it.

- journal_imports records the (epoch, tip) each chat was copied from, in the
  copy's own transaction. A file already copied is never copied again, across
  any number of restarts after a failed rename; a file that differs always is.
- Newest writer wins, per chat, with a row saying the chat was continued in an
  older version of Orca. When both builds wrote past the recorded tip under one
  epoch, the copy takes a fresh epoch, so readers reset instead of skipping rows.
- Each copied directory retires to its own .imported-<epoch8>-<ms> name, so a
  second downgrade and re-upgrade never collides with the first.

* test(native-chat): fixture deps match the host journal database shape

Attach-flow and reconcile-attach fixtures stop passing a journal database those inputs do not take, and host and restore fixtures pass the one they now require instead of the removed journal root.

* test(native-chat): state why the runtime-state fixtures' existing casts are safe

* fix(native-chat): start up normally when this process cannot open the chat journal

A process refused the chat journal, because another Orca owns it or because its own journal will not open, failed startup restoration: the window booted in degraded no-save mode and a paired phone could not list any tabs. Startup restoration now treats the refusal structured requests are getting as having no structured host; terminals, tabs and saving go on, structured requests are still refused by the gate, and the install is retried on the next one. Any other install error fails startup as before.

* test(native-chat): name the owner-lock sweep test after the two sweeps it runs

* fix(native-chat): session history and terminal resume work while chats are refused

Session history (listing and preparing a resume) and a terminal typing a resume command only check whether a structured chat owns a provider session. In a process refused the chat journal they failed outright. They now take the refusal chats are getting as having no structured host, the same treatment startup restoration gets, through one shared helper; any other install failure still fails them. Chat requests keep the gate's refusal.

* test(native-chat): the first-work rename's fake journal saves the listing status

* fix(native-chat): open a chat whose per-chat journal file never got its schema

A crash between creating a chat's journal.db and creating its tables left an empty or schema-less file. Each chat used to open that file as an empty chat; the importer instead refused the open as "try again" forever. A file with no journal_sessions table is now read as never written, the same as one with no rows. A file that is not a database, or whose read fails, is still refused.

* fix(native-chat): let the event loop run between chats during startup restore

Opening a chat's journal is synchronous SQLite now that no per-chat directory
is created first, so the restore of every visible chat ran as one main-thread
task. Each chat now waits for a macrotask before it opens.

* fix(native-chat): import a per-chat journal in bounded batches

The one-time copy of an earlier build's per-chat journal ran as one
transaction, which blocked the main thread for 650 ms on the largest chat.
Rows now copy 512 at a time, each batch its own transaction, yielding to the
event loop between batches. The rows go into a block journal_import_blocks
reserves, which no reader follows and no other chat is allocated; the last
batch publishes the chat's pointer, repair marker and import marker together
and releases the reservation. A copy that stops midway leaves only that
block, which the next open clears and copies again. Two opens of one chat
import one after the other.

* fix(native-chat): refuse chats when the owner lock file cannot be opened

A lock file that is not a database, or cannot be opened, made the claim throw
before any refusal was recorded, so startup restoration failed on every
launch. The claim now sits in the same try as the database open and records
the same typed refusal.

* fix(native-chat): finish reclaiming pages a delete frees during a running pass

A reclaim pass ended as soon as the freelist stopped shrinking between steps,
so a second delete that freed more than one step's worth mid-pass ended it
early and left those pages on the freelist. A pass now ends only when a step
itself frees nothing, or the freelist is empty.

* test(native-chat): desktop session history is served while chats are refused

* fix(native-chat): a send to a chat holding a newer Orca's rows says to update

A chat opened read-only because a newer Orca wrote rows to it answered a send
with the generic write failure. It now refuses the way a database a newer Orca
wrote does, with the same reason and words.

* fix(native-chat): keep chat tabs while this process cannot list its chats

A process whose chats another Orca owns, or whose chat journal will not
open, has no structured host. Its session-tabs inventory still answered,
with no chat rows, and the renderer read that as "every chat was closed":
it removed the restored chat tabs and the next session save persisted
their placement away.

The inventory now says `agentSessionsUnverifiable` when the last tab
restore ran with chats on disk but no host to list them. The flag is set
and cleared at the per-client projection point beside the client-hosted
page hold, and the restore is memoised only once a host answered, so a
later lock takeover or journal open republishes the chats and clears it.
The renderer keeps agent-session tabs, and keeps cancellation tombstones,
against an inventory that does not affirm its chat set.

* fix(native-chat): say chats are open in another Orca, with this process's way past it

A process refused because another Orca owns the profile's chats sent the
generic `journalUnavailable` reason, so current desktop and phone surfaces
said "couldn't open this chat's history right now. Try again." — a step
that never helps while the other Orca runs.

The refusal now names its own reason, `journalOwnedElsewhere`, with the
refused process's kind (dev desktop, packaged, orcad) as a fact. Each kind
gets its own step: quit the other Orca, or give this one its own profile
(ORCA_DEV_USER_DATA_PATH) or data folder (ORCA_USER_DATA). The sentences
are added to the shared notice copy, the desktop catalogs in all six
locales, and the boot catalog.

A client that predates the reason reads it as none and keeps the code's
words; an unknown kind reads as the packaged app's step. The `message`
released clients print is unchanged. Which requests refuse does not change.

* fix(native-chat): restore chats on taking ownership, without a list to ask

A refused startup kept its hostless result, so after the owner quit this
process never installed a host, never reconciled restart leases, and kept
telling clients it could not list its chats until a desktop chat request.
Taking the lock now reruns startup restoration once and pushes the chats.

* test(native-chat): a navigation reply says chats are unverifiable while refused

* fix(native-chat): retry a refused owner lock at most every 5 seconds

The lock frees as its holder exits, but a refused process only learns that
on its next retry, and the 30 s cap left a second Orca refusing chats for up
to half a minute after the owner quit. One retry is an open and BEGIN
EXCLUSIVE on an empty file.

* fix(native-chat): install before deciding whether a takeover must republish chats

A list that landed on the refused startup after the lock was taken finished
after the takeover had already checked, so nothing republished. The takeover
now installs first, waits for any restore in flight, and restores only then;
the restore that clears "cannot tell" pushes the frames itself, so a list
that heals the inventory first reaches subscribers too.

* fix(native-chat): no takeover lands a host after the runtime stop

Quitting cancelled a refused claim's retry only at its end, so a retry firing
during the stop's awaits took the lock and installed a host the stop never
tore down, and the lock was then released under an open journal. The stop
now cancels the retry first, keeping the refusal, and repeats its teardown
while an install that began during it (a takeover already under way) is
pending, so no journal connection outlives the lock.

* fix(native-chat): show a thrown refusal in its own words, not its code

A refusal the host throws reaches the client as an RPC error whose message is
the bare code; its reason and facts ride only in the error's data, which no
client read. The chat pane's status line therefore printed
agent_session_journal_unreadable, a send took the bare "not sent" path, and
other writes said the outcome was unconfirmed.

One shared reader, agentSessionThrownRefusal, now reads the refusal from the
error data. A failed history read shows the refusal's read-history words, a
send keeps the refusal behind its Retry exactly as a returned refusal does, and
the other writes (desktop and phone) name the refusal instead of doubting the
outcome. The phone's read failure goes through the same reader.

* fix(native-chat): log a failed journal open once per distinct failure

Every chat request retries a journal open that failed, which is intended, but
each retry also logged the failure with its full stack: a junk database file
logged the same "file is not a database" error 189 times in a minute. The open
now logs a failure only when its code and message differ from the last one
logged, and forgets it once an open succeeds. The retry is unchanged.

The open moves to its own module beside the runtime, which had no room left.

* fix(native-chat): restore lists a chat from its per-chat file and copies it on first use

Startup restore opened every restored chat, and that open copied the chat's
per-chat file into the host database, so the first boot after an upgrade paid
the whole one-time copy before the chat list appeared.

A restore open now reads a chat that is still in its per-chat file straight
from that file, read-only, with the importer's own reader, and closes the file
before moving on. That read drives the listing, the status row and the
restart offer, as it did when every chat had its own file. The copy becomes
owed work on the chat's write queue: it runs before the chat's first write,
and a reader that reaches the chat awaits it. A chat the host already holds,
or that was copied before, still opens through the import and its reimport
rules, and so does a file whose read needs a repair written.

* fix(native-chat): no host stays registered after a stop an install spanned

Each teardown pass clears the registered host before it awaits an install in
flight, and that install registers its host when it finishes. The pass then
tore the host down but left it registered, so a request after the stop was
served by a host whose journal was closed. The stop now clears the slot once
its passes are done.

* fix(native-chat): checkpoint the journal with a full flush on macOS

synchronous = FULL fsyncs each commit, but macOS fsync leaves the drive cache
unflushed, so FULL alone does not survive a power loss there. With
checkpoint_fullfsync, each checkpoint uses F_FULLFSYNC; elsewhere it is a no-op.
The comment that said FULL alone was enough is corrected.

* fix(native-chat): delete a per-chat journal once its copy verifies

An imported chat's per-chat file was kept under an `.imported-*` name, which
doubled the disk its history takes. The copy now reads back from the host
database before it is published: its items, submissions, epoch and tip must
match the file's. Only then does one transaction publish the chat with its
import marker, and the file and its WAL files are deleted, the directory too
when nothing else is in it (a pre-SQLite transcript there is kept).

A copy that does not match is never published: the file stays, the chat is
refused as unreadable ("Unable to load this chat."), and the mismatch is logged
once. A file left behind by a failed delete or a crash matches the marker, so
the next open deletes it rather than copying it again; a file an older build
wrote after a downgrade still differs, and is still copied again.

* fix(native-chat): verify an imported chat a batch at a time

The check that a copied chat reads back as its per-chat file folded both whole,
each in one synchronous task: over half a second on the largest chat. Both
reads now go a batch at a time between turns of the event loop, like the copy
itself, and count rows as well, so a copy that lost a row with no item in it
is caught too.

* fix(native-chat): restore reads a chat's per-chat file a batch at a time

Restore folded a chat still in its per-chat file in one task, so the largest
chat's file held the main thread for about half a second at startup. The fold
now takes the file a batch of rows per turn of the event loop, into the same
fold a replay uses, and nothing reads it before it is done. The file is still
closed before restore moves on.

* fix(native-chat): end a per-chat copy on a turn of its own

A chat's first open ran the copy's last steps (the verified publish and the
per-chat file delete) and the replay of what was copied in one task. The copy
now yields before it returns, so the replay, which every open runs, is a task
of its own.

* fix(native-chat): commit a per-chat copy's batches without an fsync each

Each 512-row batch of a chat's first-use copy committed under synchronous =
FULL, so a large chat paid one fsync per batch, about a quarter of its first
open. The batches now commit under NORMAL, set and restored in the batch's own
task so no other chat's commit runs under it. The publish that makes the copy
visible still commits under FULL, and under WAL that sync makes every earlier
batch durable with it. A crash before it leaves only the unpublished block,
which the next open clears and copies again.

* fix(native-chat): roll back a chat journal transaction whose COMMIT fails

The shared connection's transaction rolled back only when its body threw. A
COMMIT that failed left the transaction open, so every later write, for any
chat, failed with "cannot start a transaction within a transaction", and reads
saw rows that never committed. Under the unsynced copy the failure also tried
to restore the sync level inside the open transaction, which SQLite refuses,
so the caller got that error instead of the COMMIT's.

One transaction helper now covers the body and the COMMIT, rolls back whatever
transaction survives, and rethrows the original error. Schema creation uses it
too. If that ROLLBACK fails as well, the connection is marked stranded: each
later use retries the ROLLBACK, and until one goes through every chat gets the
same "history unavailable, try again" refusal a journal that will not open
gives. The rollback that frees it also restores the FULL sync level.

* fix(native-chat): keep the chat journal connection until its close succeeds

Closing the journal dropped its connection handle before closing it. A close
that failed left the database reporting itself closed with the connection still
open, so the stop that retried the teardown found nothing to close and released
the owner lock over a live connection.

The handle is now dropped only once the close succeeds. A failed close keeps
the runtime pending and the lock held, and the next stop closes that same
connection before it releases the lock.

* fix(native-chat): publish the runtime only once it holds the chat journal lock

When this process could not open the owner lock file at all (a permission
error, or a file that is not a database), the runtime counted that as owning
the chats and wrote orca-runtime.json. That overwrote the real owner's entry,
so the CLI was sent to a process that cannot serve its chats.

A claim that throws is now refused like one another process holds: the runtime
starts but does not publish, the claim's existing retry keeps asking for the
lock, and discovery publishes once the retry takes it. Chats still get the
refusal for the failure itself, and startup restoration reruns on the takeover
the same way it does after another owner quits. A sole process whose lock file
never opens is not found by the CLI until it does.

* fix(native-chat): keep a chat's history when an older build started it over

The first copy deletes a chat's per-chat file, so an older build run after a
downgrade finds no file and starts the chat from nothing. On the re-upgrade that
fresh file was copied in as the newer history, replacing everything the shared
database held for the chat, and then deleted.

A file whose epoch is not the one last copied and that opens with
`session_created` is now kept: neither copied nor deleted, and the chat keeps
the history it has. A file that carried the copied epoch on is still copied
again, as before.

* test(native-chat): pin which chats startup restore copies

Restore copies a chat still in its per-chat file only when restore itself has
to write to it: settling what the last run left open, here a running tool call
or a send handed over and never answered. Every other restored chat stays in
its file until its first use.

* test(native-chat): pin the copy wait on a read that opens a chat restore opened

A read queued behind restore's open of the same chat reaches the conversation
through its own open rather than the listing. It must still wait for the
owed copy, or it reads the chat before its history is in the one database.

* fix(native-chat): record a set-aside per-chat file so no later open reads it

Setting aside a file an older build started over is decided once and kept in
the new `journal_set_aside` table (schema 2, additive), with the file's epoch
and tip as they were. Every later open of the chat skips the file without
opening it, across restarts and after the older build writes more to it:
anything written there grows from that build's own start, never from this
build's history.

The best-effort delete moves beside the per-chat file reader.

* fix(native-chat): set aside any per-chat file at an epoch this build never copied

A chat's per-chat file is deleted once its copy verifies, so a file that
reappears at another epoch was never this build's history, whatever its first
row says: an older build started the chat over, possibly rewinding it after
(`handle_forked`), or rolled the epoch of a file whose delete had failed.
Copying any of them would replace everything the chat holds, so each is set
aside. Only a file still at the copied epoch is copied again (it grew) or
deleted (it did not). The first-row check is gone.

* fix(native-chat): copy a reappearing per-chat file again only while this build has not written past the copy

A per-chat file that an older build carried on under the copied epoch was
copied again even when this build had also written to the chat since the
copy, or had rolled its epoch. The second copy replaced the chat's block,
so what was sent in this build after the copy was gone for good.

Now the file is copied again only when the chat still stands exactly as it
was copied: the same epoch and tip the import marker recorded. Otherwise it
is set aside like any other file that is not this build's history, left on
disk untouched and recorded so no later open reads it. A second copy
therefore never replaces rows this build wrote, keeps the file's own epoch,
and the fresh-epoch rewrite goes away. The row it adds now says the history
includes what the older version recorded, not that anything was replaced.

* test(native-chat): pin that a chat founded here keeps its history, and the v1 schema upgrade

A chat this build founded has a pointer and no import marker, so a per-chat
file an older build later starts for it is set aside. Nothing pinned that
half of the rule: letting such a chat be copied again replaced its history
and every test still passed. A second test pins that a database written at
schema version 1 upgrades in place, gaining the set-aside table and keeping
its import markers.

* test(native-chat): drop a lost copied row by patching the source, not wrapping it

* chore(mobile): restore the mobile lockfile to main's

* fix(native-chat): pass a classified journal refusal through a send or Stop unchanged

* fix(native-chat): refuse a read whose owed copy fails as a failed open does

* test(native-chat): measure only the replace's WAL in the block-key case

Opening the chats starts a free-page pass that waits one event-loop turn,
and the seed never yields one, so that pass was still pending when the
replace committed. It woke during the async stat and reclaimed the pages
the replace freed, adding ~500 KB of WAL whenever the stat lost the race
(Linux CI). Drain that pass before measuring and stub the replace's own.

* test(native-chat): the RPC fixture's status journal can save its listing status

The status feed now hands every projection to the journal, which decides whether it is worth saving.

* test(native-chat): state why the RPC fixture's status journal cast is safe

* fix(native-chat): refuse a per-chat copy whose rows differ from the file, not only its counts

* fix(bench): build the replay benchmark's baseline arm from the base tree and release its handles on failure

* fix(native-chat): retry a failed listing status save on the next read of a cached status

* refactor(native-chat): drop the chat journal owner lock; the process instance lock already guards the profile

The journal carried its own exclusive lock, with a retry loop, an in-process
takeover, lock-gated runtime discovery and a "chats are open in another Orca"
refusal. Every shipped process kind (packaged desktop, serve mode, orcad)
already refuses a second instance on one profile before the journal opens, so
the lock only ever mattered for dev desktops, which the next commit covers at
the process level instead.

The host now opens its one journal connection at install with no lock. What a
sole process whose journal will not open needs stays: the install refusal
recorded for the gate, the no-host startup path, and the unverifiable chat
inventory, now in structured-agent-session-host-refusal.ts. The unreleased
journalOwnedElsewhere reason, its processKind fact and their copy are removed.

* fix(startup): dev desktops take the single-instance lock, and a second one says why it quit

Dev skipped Electron's single-instance lock so parallel `pnpm dev` runs from
several worktrees would not quit silently, but two dev processes on the
default orca-dev profile then write the same stores at once. Dev now takes
the lock like packaged builds: a second launch on the same profile focuses
the first window and exits with code 3, printing one stderr line that names
the taken profile and how to run another copy (ORCA_DEV_USER_DATA_PATH).

Serve mode, the macOS diagnostic bypass and the E2E harness are unchanged:
an E2E launch still skips the lock unless it sets
ORCA_E2E_ENFORCE_SINGLE_INSTANCE_LOCK=1.

* refactor(native-chat): key journal rows by chat, epoch and sequence

Rows in the host's journal database are now addressed by the chat's own
identity, with `(session_id, epoch, seq)` as the primary key, the same
shape each per-chat file already used. The block-keyed layout goes with
everything built on it: the block column and its allocator, the 2^21
block ceiling, and the import's reserved block table.

A first-use copy writes its rows under the file's epoch, which the chat's
pointer does not name until the verified copy publishes it, so no reader
sees a half-copied chat. A try that stopped midway leaves only rows no
pointer names, and the next try deletes them before it copies again.
Replace, rollover and repair delete by (chat, epoch).

This build's history always wins: once a chat was copied or founded here,
any per-chat file that reappears is set aside, and the same-epoch copy
again after a downgrade is removed.

The bounded free-page reclaim after every delete is dropped;
`auto_vacuum = INCREMENTAL` stays at file creation, so a later periodic
reclaim can still be added. Session search keeps its own step.

The schema moves to version 3. Versions 1 and 2 were written only by
unreleased builds of this change and are refused as found, not migrated.

* fix(native-chat): open a chat journal a newer Orca wrote read-only instead of refusing it

After a downgrade, the host's journal database carries a newer user_version. It was refused
outright, so every chat's history disappeared. It now opens on a read-only connection, as the
per-chat journals did: each chat shows what this build can read, from the database or a per-chat
file never copied in, and every write is refused with "Chats were saved by a newer Orca. Update
Orca to keep using them." Nothing is written, copied, repaired or founded, and the file stays
byte-identical. A table the newer schema changed reads as the same read-only refusal, not damage.

* refactor(native-chat): leave the saved listing status to the change that reads it

Nothing in this change reads the per-chat listing status column: it was a stored copy of a fact
the status feed derives, written after every turn end and cleared on every epoch change. The
status_json / status_seq columns, their writer, the saved-status type, the status feed's save and
its retry on a cached projection all go, with their tests. The change that lists chats from a
saved status adds the column back beside its reader.

* fix(native-chat): a chat saved by a newer Orca says to update Orca, not to try again

When a newer Orca wrote the chat journal, this build opens it read-only. A send or a Stop was
refused with the reason `journalUnavailable`, so today's desktop and phone clients chose the
words for an open that can clear: "Orca couldn't open this chat's history right now. Try again."
Retrying never cleared it; only updating Orca does.

The refusal now names its own reason, `journalWrittenByNewerOrca`, whose words are "Chats were
saved by a newer Orca. Update Orca to keep using them." A read refused the same way names it
too. An older client does not know the reason, drops it, and falls back to the code's words
("Orca couldn't read this chat's saved history."), and released clients still print the message.

* fix(native-chat): a chat journal from an unreleased build reads as unusable, not as retryable

A chat journal database stamped with schema 1 or 2 was written only by unreleased development
builds of this change. Opening it threw a plain error, which every chat reported as "Orca couldn't
open this chat's history right now. Try again." Retrying never cleared it.

It now throws a named error that is classified as unusable, so every chat says "Unable to load
this chat." The one log line names the file, says an unreleased development build wrote it, and
says to move it aside. Nothing migrates or renames it.

* docs(native-chat): drop the second-Orca-owns-the-chats case from three comments

The chat-only owner lock is gone, so only a chat journal that will not open leaves a runtime
unable to list its chats.

* docs(native-chat): correct three chat-journal comments the redesign left behind

A per-chat file left without its WAL is set aside, not copied again; nothing runs an incremental
vacuum yet, so the auto_vacuum mode is kept for a later pass; and the idle sweep drops a chat's
in-memory fold, since a chat holds no journal connection.

* refactor(native-chat): stop exporting chat-journal names nothing imports

Each is used only inside its own module now; the teardown's export served a deleted test.

* test(native-chat): name the version-0 test for what it covers, and check every journal table

The test named 'migrates an older user_version forward' covers only a version-0 file that already
has its tables; versions 1 and 2 are refused. The table test now also checks journal_imports and
journal_set_aside.

* fix(startup): a second dev launch's exit line no longer claims it focused a window

The running dev instance may be a background launch or a server, which show no window. The line
now says only that this launch passed its request to that instance.

* fix(native-chat): a failed structured-chat install closes the journal connection it opened

The install opened the chat journal database and closed it only if the record store then failed
to open. A later failure, such as the model catalog wiring or the host constructor, left the
connection open, and the next install opened a second one in the same process. Every failure
after the open now closes it.
2026-09-29 16:42:29 -07:00
Neil 6194a7a1b6 test: drop private-internal and boundary-census tests with behavioral owners (#23941)
Fourth audit wave, cut short by a session restart, so this lands the verified
subset rather than the full batch.

Removes private-predicate cases whose behavior is already covered through the
module's real entry point, and de-exports the seams they reached for. Also drops
three whole files whose every case was a duplicate or a call-shape grep.

The source-grep vein is close to exhausted. One auditor reviewed 15 remaining
flagged files and deleted nothing: what is left is mostly legitimate
architectural ratchets that no type checker and no behavioral test can reach —
AST fences banning `as`/`any` in an RPC operation region, discovered-vs-listed
set equality over subscription sites, count ceilings on unchecked reply readers,
and assertions on generated WebView bundles (no CDN URL, no `</script`
tokenizer escape, parses at the Chrome 74 floor). Those stay.
2026-09-29 16:35:49 -07:00
Brennan Benson 21d4ae9448 feat(agents): pre-trust the folder wherever Orca starts an agent (#23744)
* feat(claude): pre-trust worktrees Orca creates

Claude Code asks "Do you trust this folder?" on first launch in any folder it
has not seen, which blocks unattended launches in worktrees Orca itself made.
Orca now records where a worktree's content came from when it creates it, and
before each Claude launch writes Claude's own folder-trust entry for that
worktree's root (never the main checkout) when the new setting is on and the
content is the user's repository. Forks, bare commits, folder workspaces and
external checkouts keep Claude's prompt. The write takes Claude's lock, never
creates or breaks the file, runs on the SSH host itself, and is revoked when
the worktree is removed or the setting is turned off.

Launches that already pass --dangerously-skip-permissions also skip the trust
prompt for that one process only, via CLAUDE_CODE_SANDBOXED=1 on the command.

* fix(claude): parse the relay trust request with a schema and ship its search keys

* fix(claude): never write a WSL guest's trust into the Windows config

A WSL worktree's Claude reads the guest's own config. Two paths still wrote
its trust into the Windows host's ~/.claude.json instead: the Claude auth prep's
fallback (runtime 'wsl' but the host config dir, when the WSL home cannot be
resolved), which wrote a Linux-path key the removal revoke can never delete;
and the Agent Teams leader, which passes no auth or distro and wrote a UNC key.
Require the guest's own config dir, and treat any WSL worktree path as guest-only.

* revert(claude): drop the skip-permissions trust shortcut

Pre-trust stays limited to worktrees Orca creates from the user's own
repository. The per-launch CLAUDE_CODE_SANDBOXED prefix skipped Claude's
trust question in every folder for launches carrying the skip-permissions
flag (Orca's default Claude args), including the user's own folders and fork
PR worktrees, and the setting could not turn it off. Remove the prefix, the
inherited-variable strip that existed only for it, and the Agent Teams
leader-to-teammate propagation; restore the tests that pinned the prefixed
launch string.

* fix(claude): revoke SSH trust in the config file the grant used

At spawn the relay resolves Claude's config from the launch env, which carries
a CLAUDE_CONFIG_DIR set in Orca's Claude default env. claudeTrust.converge had
only the relay's own process env, so removing the worktree or turning the
setting off revoked in the default file and the grant outlived the worktree.
Send the config-file keys with the request, as the local revoke already uses.

* i18n(settings): translate the Claude worktree trust setting

* fix(settings): say Claude trust applies when Orca starts Claude

The description said Claude skips its trust prompt in any worktree Orca
created. Trust is written only when Orca itself starts Claude there, so a
`claude` typed by hand in a fresh worktree still asks. Say that, and bring
the es/fr/ja/ko/zh translations in line with the new text.

* fix(worktrees): treat a base on an Orca-added fork remote as fork content

A worktree based on a named ref was always stamped as the repository's own
content, so picking the fork remote Orca adds for a pull request (or a local
branch tracking it) as the base made a fork's code eligible for Claude trust.
At create time, read the repo's `remote.<name>.orca-created` markers and each
branch's tracked remote in one `git config` call; a base on such a remote is
stamped as a fork's content, and a read failure is not vouched for. Remotes
the user added, such as `upstream`, stay first-party.

* perf(claude): revoke worktree trust once per config file, not per worktree

Turning "Trust worktrees Orca creates for Claude" off read and parsed the
whole Claude config once per Orca worktree on the main process. Group the
revocations by config file locally and by SSH connection, and make
claudeTrust.converge take a batch of requests.

* feat(settings): one agent-wide "trust the folder" setting in Settings > Agents

Replace the Claude-only worktree trust toggle with a single setting,
agentWorkspaceTrustEnabled (on by default; unreleased, so no migration).
The row says what it does for every agent: agents Orca starts skip their
"trust this folder?" prompt in that worktree or folder, turning it off
stops new trust while existing trust stays, and while it is off unattended
launches (orchestration workers, automations, the phone) stop at the
agent's trust question until someone answers.

Translations for es/fr/ja/ko/zh. Also restores the `awaitingUnnamed` chat
catalog keys an earlier merge of main dropped from this branch.

* feat(agent-trust): pre-trust the workspace for every preset agent at PTY spawn

Every Orca-started agent PTY passes through one of the two spawn builders
with its declared launchAgent, which survives setup-script wrapping. The
builders now call one hook that, for a fresh launch (never a reattach or
restored pane) with the setting on, applies the agent's trust preset to
the worktree, folder workspace or main checkout it starts in.

- One dispatcher, applyAgentWorkspaceTrust(preset, workspacePath, launch
  context), carries what a writer needs: the final spawn env, the Claude
  managed-account auth prep, the WSL distro and the SSH connection.
- Claude joins the presets on both the claude and claude-agent-teams
  entries. Its writer stays grant-only in claude-folder-trust-file.ts:
  the file Claude reads (CLAUDE_CONFIG_DIR / custom-OAuth suffix / legacy
  .config.json / a WSL guest's own file), Claude's <file>.lock never
  broken and taken only when a write is due, atomic temp+rename keeping
  mode and symlinks, never creating the file, NFC + realpath keys.
- SSH Claude launches forward the optional claudeFolderTrust spawn field
  so the relay grants with its own spawn env; old relays ignore it and
  Claude asks. Other presets keep the SFTP writer. A WSL launch never
  writes the Windows home: non-Claude presets skip it, Claude writes the
  guest's file or nothing.
- Codex keeps the 20 s deadline its shared config lane needs; every other
  preset gets 1.5 s. A miss means the agent asks; trust bookkeeping never
  fails or blocks a launch.

Removes the Claude-only machinery this replaces: the eligibility/host/
lifecycle/spawn modules, the persisted creation content-origin field and
its classification, revoke-on-removal, the setting-off sweep, the
claudeTrust.converge relay method and the Agent Teams leader special case
(the leader pane now spawns through the hook with the claude preset).
The agent config types move to tui-agent-config-types.ts so the config
table stays under the line budget.

* refactor(agent-trust): delete the pre-spawn trust writes the spawn hook replaces

The spawn hook is now the only owner of agent folder trust, so remove
every other writer:

- the agentTrust:markTrusted IPC channel, its preload bridge and types,
  and all renderer callers (agent-trust-preflight and its callers in the
  background session, work-item direct launch, session continuation,
  worktree creation, folder workspace composer and session fork);
- the main pre-spawn sites: the createdWithAgent preflight in
  worktree-remote.ts, markLocalWorktreeTrusted/markRemoteWorktreeTrusted
  and the runtime's markWorkspaceTrustedForAgent family with the
  markTrusted ports of the runtime create flows;
- Codex's own launch-prep and resume-prep trust writes.

Each of those launches reaches a spawn builder with launchAgent set, so
the hook covers it. This also fixes a live gap: the worktree-remote.ts
copy of the preset switch omitted Antigravity, so an agy agent started
from a desktop worktree create still asked; the single dispatcher covers
it. Trust is also written on the host the PTY actually spawns on, which
removes the #11163 class of writing the wrong host's config.

* test(agent-trust): type the spawn-builder trust fixtures and prove the spawn waits for trust

The builder test passed untyped args (a string launchAgent) and cast its deps,
which failed tc:node. It now builds both spawn states from a fully typed deps
fixture and a typed restored pane, with no casts.

Adds a case that holds the trust write pending and checks the builder does not
finish until it settles, the ordering the deleted renderer and launch-prep
tests used to cover.

* fix(agent-trust): give SSH trust writes the 20 s deadline again

The dispatcher gave every non-Codex preset a 1.5 s budget, including the SSH
writers for Cursor, Copilot and Qoder, which make several round trips over the
link. Before this PR those writes had 20 s (desktop) or no limit (runtime), so
on a slow link an unattended SSH worker would now stop at the agent's trust
question. SSH writes get the 20 s deadline back; local non-Codex writers keep
the short budget, and Codex keeps 20 s.

The relay's Claude grant keeps the short budget: it writes the relay host's own
disk and does not cross the link.

* fix(agent-trust): never pre-trust a home folder or a filesystem root

A folder workspace can be the user's home folder or a disk root. Claude and
Copilot let a trusted folder cover every folder under it, so pre-trusting one
of those would silently trust everything on the machine for those agents.

One check, isHomeOrFilesystemRoot, now refuses them for every preset: the
dispatcher checks roots and this machine's homes (including the spawn env's
HOME and a cached WSL guest home), the SSH writer checks the remote home it
already resolves, and the relay checks its own home. The agent then asks, as
it would without Orca.

* refactor(agent-trust): drop the Codex launch plumbing that only carried trust

The spawn hook replaced the trust writes in Codex launch prep and resume prep,
which left the fields that fed them unread: CodexHomeLaunchContext.workspacePath
and .launchAgent, the resume prep's workspacePath, and the structured Codex
launch input's workspacePath (plus the extra target lookup that produced it).
Remove them and their plumbing; unavailableManagedHomePath stays.

Also removes test stubs of runtime trust methods this PR deleted, whose
not-called assertions could no longer fail, and two comments that still
described the old trust preflight.

* chore(reliability-gates): point the trust gate at the spawn-time trust tests

The agent-session trust gate still listed three test files this PR deleted
(the renderer preflight, the Codex launch-prep deadline and the e2e trust
completion suites), so check:reliability-gates, which runs in PR CI and in
pnpm lint, failed on missing files. Its invariant also described the deleted
IPC handler and pre-spawn writers.

The gate now covers what replaced them: the spawn builders holding the spawn
until trust settles, the fresh-launch and setting gates, the per-preset
deadlines, and the home and root refusal.

* fix(agent-trust): skip the relay Claude grant for a WSL shell

On a Windows SSH host whose pane shell is wsl.exe, Claude runs inside the WSL
guest and reads the guest's config. The relay still granted trust in the
Windows host's own .claude.json, writing the Windows home for a WSL launch,
which the local path never does. The relay now skips the grant there, so that
Claude asks, as a local WSL launch does when Orca cannot reach the guest file.

* perf(agent-trust): only agent launches wait on the trust hook

Both spawn builders awaited the trust hook on every spawn, including plain
shells, reattaches and agents without a preset. Awaiting even a resolved
promise adds microtask ticks ahead of the pane-spawn reservation check, and
this handler already keeps non-Codex spawns off an await because an extra tick
reorders those reservation races.

The hook now returns null when there is nothing to write, and the builders
await only a real trust write.

* test(agent-trust): keep the home and root cases off any real Claude config

The home and root cases ran the real Claude writer with the test process's
env, so a regression in the guard would have written trust for the home
folder and / into whatever Claude config that env named. They now point
CLAUDE_CONFIG_DIR at a folder that does not exist, and the writer never
creates a config.

* test(runtime): drop needless casts from the launch-host test

The renamed launch-host test kept three `as never` casts on launch options
that already match launchAgentTerminal's parameter type. The changed-lines
casting gate reads the renamed file as new and failed on them.

* test(agent-trust): type the Claude grant mock with the real writer's signature

The mock took an unknown target, so installing the real writer as its
implementation would not typecheck under strict function types.

* fix(codex): drop the launch context the trust move left unread in local spawn env

* fix(agent-trust): queue Claude grants per config file so a launch burst keeps them all

Concurrent grants in one process retried Claude's file lock in lockstep, so each
retry round admitted about one winner. Starting 12 Claude agents at once left 6
of them at the trust question with nothing logged. Grants for one config file now
queue in-process; only Claude's own writes contend for the lock. The relay shares
the writer, so bursts of SSH launches are covered too.

* fix(agent-trust): never pre-trust a folder above a home either

The guard refused only an exact home or a filesystem root. A folder workspace at
/Users, /home or C:\Users was still pre-trusted, and Claude walks up parent folders
for a non-git folder, so every non-git folder in the user's home became trusted.
The guard now also refuses any folder that contains a home, on every host, and is
renamed to say what it decides.

* perf(agent-trust): skip the SSH round trips for Antigravity, which has no remote writer

Every Antigravity launch over SSH now reaches the remote trust writer, which
resolved the remote home and realpath'd the workspace over the link before
writing nothing (the known remote gap). That delayed each launch by two SSH round
trips, and up to the 20 s deadline on a stalled link. It now returns first.

* fix(settings): keep the hidden folder trust row out of web-client settings search

The paired web client hides the host-only "Trust the folder" row, but settings
search still listed it, so searching "trust" opened the Agents pane with no
matching row. Its search entry is now filtered the same way as Agent Awake.

* fix(agent-trust): a failing breadth guard skips trust instead of failing the spawn

* docs(qoder): New Tab now pre-trusts through the agent-wide spawn hook

* fix(agent-trust): never pre-trust a home reached through a symlink

Every trust writer stores the workspace's resolved path, but the breadth
guard compared only the path as given. A folder workspace that is a
symlink to the home folder (or a real home picked while HOME names a
symlinked one, as on distros that link /home to /var/home) passed the
guard, and Claude, Copilot and Cursor then trusted the home itself.

The local dispatcher and the relay now compare given and resolved forms
of both the workspace and each home. The SSH writer resolves the remote
home alongside the workspace, in parallel, so it adds no round trip.
Local non-Claude WSL launches still skip before any filesystem call.

* fix(relay): a failing breadth guard skips Claude trust instead of failing the SSH spawn

The relay ran its home/root guard and homedir() before its catch, so a
throw there rejected the relay's terminal spawn. Same fix as the main
dispatcher's: the whole grant, guard included, is best-effort.

* fix(agent-trust): guard the path each writer stores, not the path Orca was asked to trust

The breadth guard checked the launch's workspace while each writer stored a
transformed path, so every new transformation opened a hole. Codex stores a
linked worktree's main checkout: with a git repo rooted at the home, a Codex
launch in one of its worktrees wrote trust for the whole home.

One relay-safe host module now computes the stored path (Codex's main-checkout
hop, then given and resolved forms of it and of each home), refuses a root, a
home or a folder above one, and only then writes. Main uses it for local and
WSL launches and the relay for Claude. An unknown home writes nothing, and the
WSL home cache is keyed case-insensitively by distro.

* fix(ssh): the relay writes every preset's trust on the SSH host itself

Codex, Cursor, Copilot and Qoder trust over SSH was written from the desktop
over SFTP: four or five round trips per launch, so it needed a 20 s deadline
that outlasted the 8 s draft paste, the 10 s phone wait and the 15 s web-client
create. It also skipped Claude's atomic rename, ignored CODEX_HOME, and stored
the worktree where local Codex stores the main checkout.

The unreleased `claudeFolderTrust` spawn field becomes `agentWorkspaceTrust`,
sent for every preset. The relay derives the preset from the `launchAgent` it
already receives and runs the same host writer main uses, on its own disk,
within 1.5 s and with no extra round trip. Antigravity still returns early on
the relay (its writer is unverified on SSH hosts), a WSL shell still skips,
and any throw means the agent asks.

Deleted: the SFTP preset writer, the remote Qoder writer, the SSH deadline
clause and the desktop-side SSH root pre-check.

* test(e2e): keep CLAUDE_CONFIG_DIR out of isolated Electron launches

The spawn hook now writes Claude folder trust into the config
CLAUDE_CONFIG_DIR names, so an e2e run started from a shell that sets it
could add trust entries to the developer's real Claude config. Also drops
a stale comment that still named Codex launch prep as the trust owner.

* fix(agent-trust): guard Claude's resolve() form of the workspace too

Claude's writer stores both resolve(path) and the realpath. The breadth
guard compared only the given path and its realpath, so a workspace
path that does not exist and climbs back with `..` (for example
<home>/missing/..) passed the guard while Claude stored a key for the
home itself. The guard now also compares resolve(path), so it sees
every form a writer stores.

* test(relay): pty.spawn writes agent trust before the agent's process starts

Nothing exercised the relay handler's call into the trust writer, so
removing that call, or no longer awaiting it, left every suite green
while SSH launches silently stopped pre-trusting. The new case holds the
trust call pending and checks the spawn waits for it, and that the call
gets the request, the declared agent and the final spawn env.

The reliability gate lists the new suite and records the resolve() form
the breadth guard now compares.

* fix(agent-trust): refuse a home only for agents that inherit trust from it

The home and root refusal applied to every preset, so Codex, Cursor and
Antigravity started asking in a home folder workspace, where they did not
before. Only Claude, Copilot and Qoder let trust on a folder cover the
folders below it; Codex matches its start folder or that folder's repo
root, Antigravity the exact folder, and Cursor itself never inherits from
a home, a folder above one or a shallow path. The refusal now reads a
per-preset table in the host module, so the local and relay writers share
the rule.

* fix(agent-trust): trust Codex at the folder it starts in, as before

Before this PR, Codex launch prep trusted the spawn's start folder. The
spawn hook trusted only the workspace root and skipped terminals with no
workspace, so Codex began asking in a floating terminal and in a subfolder
of a non-git folder workspace: its lookup checks the start folder, then
that folder's repo root, and a plain folder above it is neither. The hook
now passes the resolved start folder for presets marked as keyed by it
(Codex only), falling back to the workspace root.

* fix(agent-trust): pre-trust a structured Codex chat's folder, as before

Before this PR, creating a structured (native) Codex chat pre-wrote Codex
trust for its folder through launch preparation. The PR removed that write
and routed trust through the PTY spawn builders, which a structured chat
never passes. Codex's app-server trusts the folder itself only when the
chat's permissions can write it, so a read-only chat started running
untrusted and ignored the project's .codex config. Creating the chat now
calls the same dispatcher, behind the same setting, before launch prep.

* fix(settings): plainer folder trust setting text
2026-09-29 16:03:28 -07:00
Brennan Benson 2807332735 refactor(shared): bring constants.ts back under the max-lines limit (#23923)
* refactor(shared): move onboarding, notification and terminal platform defaults out of constants.ts

* chore(lint): keep the shapedSidebar naming exemption on the file that now holds it

* chore(i18n): regenerate the runtime catalog so it covers main's shipped keys
2026-09-29 14:38:47 -07:00
Jinwoo Hong 9420d49bcb fix(terminal): run Codex in Orca terminals without the shared background server (#23900) 2026-09-29 10:22:27 -07:00
Neil 31012aeb09 test: remove assertion-free probes, copied inventories and export-shape checks (#23816)
Second audit wave, targeting three more junk patterns:

- assertion-free cases that run code and assert nothing, so they pass no
  matter what the code does;
- inventory literals re-typed from a production declaration, where the only
  way the assertion can fail is someone editing one of the two copies;
- export key-set and export-shape loops (`typeof x === 'function'` over every
  export) that restate what TypeScript already enforces.

Yield is much smaller than wave 1 on purpose: the assertion-free scanner has
a high false-positive rate, because many flagged blocks assert through a
shared helper or their oracle is "this must not throw". Those were kept.

`mobileWebCheckArgs` in `config/scripts/run-mobile-web-app-checks.mjs` is
de-exported — after the inventory comparison went away, nothing outside the
module read it.
2026-09-29 02:21:47 -07:00
Neil 6e1b7e7fa3 test: remove junk tests that assert source text instead of behavior (#23815)
Deletes 101 test files and trims 112 more, all matching documented junk
patterns: exact source/import/string greps, copied inventories and export
lists, duplicate invocations of a contract another test already owns,
typeof-shape checks TypeScript already enforces, and self-comparisons.

The largest group read a production `.ts` file and asserted on its text —
for example a TaskPage test that required the source to contain
`selectedRepos.find((r) => r.id === newIssueRepoId) ?? selectedRepos[0] ?? null`.
Any behavior-preserving rename broke it; no behavior change ever did.

Production-side follow-through: exports that only these tests imported are
de-exported or deleted, stale comments pointing at removed censuses are
dropped, and the reliability-gate registry, `cloud/package.json` test lists,
and orphaned source-reading helpers are updated so nothing references a
deleted file.

Two files kept their real coverage and lost only the census scaffolding:
`agent-status-producer-census.test.ts` now drives all five producers end to
end instead of grepping the source tree, and `config-toml-trust-stale-writes`
replaces an export-list parity check.
2026-09-29 01:21:53 -07:00
OrcaWinandm4air e1362ada4c fix(terminal): stop inline-image decoders exhausting the renderer's wasm memory budget (#23499)
V8 reserves an 8 GiB guard region per wasm memory inside its 1 TiB sandbox,
so an Electron renderer can hold only ~124 live wasm memories regardless of
free RAM. @xterm/addon-image instantiated a SIXEL decoder per terminal at
activation (and kept IIP decoders after the first image), so ~120+ terminals
exhausted the budget: new panes raised 'WebAssembly.instantiate(): Out of
memory' rejections, and the next Kitty/IIP image threw 'WebAssembly.Memory():
could not allocate memory' out of the parser, permanently wedging that
terminal's write queue.

The addon-image source patch now borrows SIXEL decoders from a shared pool
only while a sequence is open (color registers stay on the terminal), drops
IIP decoders after each image, and turns a failed decoder allocation into a
dropped image instead of a parser throw. Bundles regenerated with
regenerate-xterm-patches.mjs --write.

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-29 01:20:37 -07:00
Neil f8f656ca19 perf(ci): spend fewer concurrency slots per pull request (#23810)
A concurrency slot is charged per job, not per core, and the account's cap is
the scarce resource: standard runner minutes are free and unlimited on a public
repository. Two paths spent slots that bought nothing.

The unit matrix ran eight fixed shards averaging 6.5 minutes each, 3384
job-slots a day and 68% of all slot demand, while the arm pool queued 10.5
minutes at p95 — the queue was the oversharding. Five shards run the same work
in ~10.5 minutes each for three fewer slots per run.

Bun profile persistence escalated to all six platforms on `config/`,
`resources/` and `.github/` wholesale, which took 36.5% of the last 1100
commits through the full matrix where a platform-flavoured predicate takes 19%.
A pull request now qualifies one platform unless the change is platform-
flavoured, and the push to main re-qualifies all six, so an unescalated miss
surfaces minutes after merge rather than at the next cron. Missing changed-file
evidence and an unavailable dependency graph still fail closed to all six.
2026-09-29 00:13:33 -07:00
Neil 25d9c57e3a fix(native-chat): keep a child turn settling on an error Codex will not retry (#23808)
#23801 removed the child-path reading of Codex's turn-ending `error` along with
the import of the module #23682 deleted. That was the wrong half to remove: the
primary journal path can rely on Codex's failed `turn/completed` arriving within
~32 ms, which is what #23682 established, but a child turn has no such
guarantee, and without the error as its end the child's lifecycle row latches on
`working` for the life of the session. Three tests assert exactly that and could
not run, because the unresolved import had been skipping the unit matrix since
#23682 merged.

The reading is restored inline against `readCodexErrorWillRetry`, itself restored
to `codex-structured-thread-facts.ts`, rather than by reviving the deleted
module: its `thread-stopped-running` arm lost its only consumer when #23682
rewrote the primary path, so restoring the file would re-add dead code.

Also drops `pr-workflow-parallelism.test.mjs`'s read of
`.github/workflows/track-community-prs.yaml`, which #23796 deleted while leaving
the assertion behind. Same failure class, and it fails the same shard.
2026-09-28 23:42:01 -07:00
Neil 2ea3fb1d46 perf(ci): take advisory unit-selection evidence off the gate (#23776)
selection_evidence is continue-on-error on both the job and its comparison step,
so it can never fail a PR -- it downloads the shard reports, compares selection
against the full results and uploads a review artifact. But a caller's
`needs: test` waits for every job in the called workflow, so living inside
unit-tests.yml it held verify for ~36s after the last shard finished.

It moves to its own reusable workflow called as a sibling, so it still runs on
every PR and still uploads its artifact, but verify no longer waits for it. It is
deliberately absent from verify's needs, and a contract test pins both that and
its advisory status so it cannot drift back onto the critical path.

Measured on a recent run: the shards finished, then selection_evidence ran 36s,
then verify 3s. Only the last of those gates anything.
2026-09-28 22:13:35 -07:00
Brennan Benson c5fc0c6f26 fix(ci): keep a squash-merged RPC recording pin reachable through its pull request (#23720)
* fix(ci): keep a squash-merged RPC recording pin reachable through its pull request

Main's "RPC recording pin" check has been red since #22762: that branch pinned
the recording corpus to its own commit 03995ae, and the squash-merge left that
commit out of main's history. Every behaviour-change squash did the same, and
each needed a hand-made repin PR to clear it (#23565, #23535, #23046 and more).

The guard now accepts a pin that is either in this history or in the head of the
pull request whose squash wrote it into the manifest. It finds that pull request
from the `(#n)` subject of the commit that added the pin and fetches
`refs/pull/<n>/head`, which GitHub keeps after the branch is deleted. The
reproduce step uses the same lookup, so it can still check the pinned tree out.

* fix(ci): give the recording pin lookup room to walk a blobless clone

In CI's blobless clone, `git log -S` fetches the manifest's blobs one commit at a
time, a few seconds each. Under the 30 s process default the walk was killed after
a handful of manifest commits, which main's history already exceeds (up to 7
manifest commits between a pin landing and the next pin change), and the guard
then failed with an empty "Could not find the commit that pinned" error. The
lookup and the pull request fetch now carry explicit budgets and say when they
timed out.

The not-an-ancestor instruction now names the pull request whose head was
checked, or says the commit that pinned it names none.

Adds the two merge-preview shapes the guard runs on: a branch opened after a
squash resolves main's pin through the squash's pull request, and a branch whose
rebase dropped its own pinned commit fails on its pull request instead of on main.

* fix(mobile): tell a missing recording pin apart from product drift

After a squash the pinned commit can live only in its pull request's head, so a
clone that never fetched it makes `git diff --quiet <baseline>` exit 128. The
recorder reported that as "Product sources or lockfile differ from the pinned
main baseline", which sends the developer to repin a tree that may match. It now
prints git's error and the command that fetches the pin.

* fix(ci): ask GitHub which pull request holds a squash-dropped recording pin

The recording pin guard found the pull request that keeps a squash-dropped
pin by walking main's first-parent history for the commit that wrote the pin
into the manifest and reading "(#n)" off its subject. A merger who edits the
squash title loses the number, and the push to main turns red anyway. That
already happened on main: of the 22 squashes that left a pin outside main's
history, #21674's title had no "(#n)".

The guard now asks GitHub for the pull requests associated with the pinned
commit (GET /repos/{owner}/{repo}/commits/{sha}/pulls) and, for each in turn,
fetches refs/pull/<n>/head and accepts only when git proves the pin is an
ancestor of that head. GitHub only nominates candidates, so a wrong answer can
fail the guard but never pass it. The endpoint named the right pull request
for all 22 historical cases, #21674 included, and names none for commits a
force-push orphaned.

This removes the pickaxe walk, its 600 s budget and its lazy blob fetches in
a blobless clone, the first-parent subtlety, and the subject regex. A revert
that restores an older pull-request-only pin now resolves too, because the
lookup is by the pin itself rather than by the commit that last wrote it.

CI passes the job token to both guard steps and grants the job
pull-requests: read. Local runs work without a token on this public repo and
send GITHUB_TOKEN or GH_TOKEN when set. A failed lookup throws with the HTTP
status, and names the rate limit when an unauthenticated call is refused.
2026-09-28 21:17:07 -07:00
Neil 3976ad4c59 perf(test): remove obsolete structural snapshots (#23777) 2026-09-28 20:55:05 -07:00
Neil ec9f35e2ee perf(ci): plan the unit shards before the static-analysis gate instead of behind it (#23743)
A caller's `needs` gate the whole called workflow, so while the plan job lived in
unit-tests.yml it could not start until static analysis and typecheck had both
finished and passed -- and the shard matrix then waited on it. The two hops were
serial when they did not need to be: planning reads the checkout, a git diff
against HEAD^1, the import graph and the checked-in timing baseline in
config/scripts/ci-shard-timings.json, and consumes nothing that static analysis,
typecheck or the native-cache primer produce.

Planning moves to its own reusable workflow so pr.yml can run it against
code_paths alone, overlapping it with the gate. Measured across 99 runs, the
shard matrix is created a median 93s earlier (p25 47s, p90 241s, never later).
Planning stays a required predecessor of the shards, so an empty assignment
cannot expand the matrix.

The gate itself is deliberately left in place. It fires on 22% of runs, and the
shard queue wait knees hard above ~9 concurrent ARM jobs -- 4s median below that
against 218s at 15-19 -- so admitting 8 doomed shards per failed run would cost
more in queue pressure than it returns in latency.

Cost is one 37s ubuntu-latest job, which does not touch the ARM pool the shards
contend for.

A planning failure still fails the PR: the shards are skipped, and verify's
check_job requires success whenever the classifier says tests should run, so it
reports `test: expected success, got skipped`.
2026-09-28 19:22:10 -07:00
Neil ccdb324b63 Add CodeBuddy as a built-in coding agent (#23740)
* feat(agents): integrate CodeBuddy launch, status and session history

* docs: record CodeBuddy lifecycle verification

* fix(codebuddy): backfill scoped history and negotiate remote resume

* test(cli): include CodeBuddy in known search agents
2026-09-28 18:11:25 -07:00
Brennan Benson d68eee3757 fix(runtime): retire an exited terminal before its stream end (#23492)
* fix(runtime): retire an exited terminal before its stream end

An exit's durable retirement became asynchronous, so onPtyExit released
the terminal stream before the retirement landed. A paired client answers
a stream end by re-activating its pane; that activation still found the
exited leaf, materialized it under the same session id, and registerPty
dropped the pending retirement. The exited split pane came back as a
fresh shell.

The exit now stages the retirement into the in-memory session and
publishes it synchronously, then notifies exit listeners, and only then
makes it durable. A failed durable write is logged and left in memory for
the next profile write instead of being rolled back, since the process is
gone either way. This removes the pending-retirement latch and its
post-await incarnation fence: there is no longer a window for them to
guard.

* test(runtime): a failed exit retirement still reaches disk

Pins the no-rollback contract through a real Store and SQLite authority:
when the retirement's own durable write fails, the in-memory retirement
is carried by the next unrelated profile write, and by the app-quit
flush when no other write happens. The delayed authority fixture can now
fail its next write, and the acknowledged-retirement fixture reads the
database a relaunch would load and models the quit flush.

* test(runtime): a stream end observes the exit retirement already published

The re-activation check alone passes with the listener ordering reverted,
because activation awaits before its lookup. Record the session binding
and publication count at the moment the exit listener fires so the
ordering itself is pinned.

* fix(runtime): an exit cleanup fault still ends the terminal stream

* perf(runtime): exits retired together share one durable write

* test(runtime): a refused staging write still retires the pane and ends the stream

* refactor(runtime): describe exit retirement as staged, not durably accepted

The retirement result is staged in memory before any write, and the removable-surface comment and the replacement-admission test name still described the old publish-after-durable rule.
2026-09-28 16:12:26 -07:00
Brennan Benson c5330d0d52 fix(native-chat): stop killing processes that only inherited a chat's spawn tag (#23460)
* fix(native-chat): stop signalling processes that only inherited a spawn token

A spawn token is an environment variable, so every descendant of a provider child
carries it. The Linux-only startup scan treated any carrier no lease claimed as a lost
provider child and sent it SIGTERM, which also hit editors, tmux servers and nested
Orca processes the agent had started. Remove that scan's killing consumer; the token
scan stays for the reservation probe, and recorded owners are still stopped by
identity during recovery.

* fix(codex): remove the token-scan kill path from app-server teardown

Every descendant inherits the spawn token, so killing each pid that carries it can
reach processes the agent started that are not the provider. Production never
injected this path; teardown always uses the process-group and descendant-snapshot
proof. Drop it, its deps, and the now-unused spawn-token argument.
2026-09-28 15:25:24 -07:00
Brennan Benson 2ca4ecbc61 feat(orchestration): let a structured chat run orchestration as itself (#22568)
* feat(orchestration): inject the Orca session id into structured children and let the CLI act as it

Every structured session's child (native Claude, native Codex, and the terminal
view) carries ORCA_AGENT_SESSION_ID and reaches the Orca CLI. The CLI sends the id
in the orchestration envelope; when present it is the caller, and a caller flag
naming anyone else is refused before any request. The id is stripped from
inherited PTY env and from the SSH host-CLI passthrough, and crosses into WSL so
the host can refuse the cross-host claim.

* test(orchestration): pin session id injection for native Claude, native Codex, the terminal view, WSL, PTY inheritance and SSH

* test(orchestration): pin one caller precedence rule across every CLI verb that names its caller

Adds the per-verb table (flagless acts as the session; a conflicting --from or
--terminal is refused before any request; the session's own spellings are
accepted), the enumerated guess population with its positive control, the
structured worker's own handle, the identity-less refusal for an older child,
the unchanged terminal agent, and the envelope. dispatch-show's --from only fills
preview text, so it passes through unfenced and a session's flagless preview
names the address the real dispatch writes.

* refactor(orchestration): keep the identity-less marker reader to the marker; the id is checked first

* test(orchestration): pin that a host refusal of the session surfaces verbatim from the CLI

* fix(orchestration): keep the identity-less marker beside the id for CLIs that predate it

A CLI older than the id, reached through a global install when a shell rc resets
PATH, would otherwise guess a sibling's terminal in a chat that no longer carries
the marker. It refuses on the marker instead; a current CLI checks the id first,
so the marker never makes a session with an id identity-less.

* fix(orchestration): refuse a conflicting --from on gate-list and task-list scoped by --run

A --run listing needs no caller, so both handlers skipped the resolver and a
--from naming another actor was dropped silently under a session. The conflict
check now runs on that branch too; terminal callers are unchanged.

* fix(orchestration): name this app's CLI by absolute path for a structured session's login shells

A provider can run each command in a login shell: Codex runs zsh -lc, and the
profile rebuilds PATH, putting a global install (possibly an older Orca) ahead of
the directory Orca prepended. ORCA_CLI_COMMAND, which an agent resolves the CLI
from first, is now the absolute launcher in that directory (the native launcher
on Windows), so no shell's startup files can swap it. The PATH prepend stays for
shells that read no profile. Found by the live coordinator run of the next PR.

* test(orchestration): pin a structured worker's CLI command as this app's absolute launcher

* test(orchestration): run the zsh login-shell arm in the real-shell lane that installs zsh

The ordinary Linux unit lane has no /bin/zsh, so the zsh arm failed there with
ENOENT. It moves to a live-shell file registered in the shell-contracts lane; the
bash arm keeps running in every lane. The lane guard's detector now also sees a
zsh spawned through the ProcessSpec program field, which is how this test
escaped it.

* fix(orchestration): omit a structured child's CLI command when no launcher resolves, and pin its instance

A bare `orca` fallback named GNOME's screen reader on packaged Linux, and an inherited value named
another app's CLI. The builder now deletes any inherited value, sets the absolute launcher only when
one resolved, and pins ORCA_USER_DATA_PATH so a current CLI dials the instance that minted the id.
Renames the marker reader to hasStructuredSessionMarker and records why the terminal view carries
the id without the marker.

* fix(terminal): name this app's CLI launcher by absolute path in every local terminal

ORCA_CLI_COMMAND meant three things by lane: an absolute launcher for a structured session, a bare
name for WSL, and nothing for any other terminal, so a structured session's terminal view lost it.
Local terminals now get the same absolute launcher the structured lane gets; WSL keeps its guest
command name, and a terminal whose launcher does not resolve still gets none.

* feat(cli): hand a command to the session's own CLI when another Orca CLI was invoked

A login shell can reorder PATH behind a global install, and an agent or its helper script can run
bare `orca`, so the binary that answered depended on the agent following instructions. Orca's
packaged launchers and bare-orca shims now export ORCA_CLI_SELF (outermost wins). At the CLI entry,
when it names a different launcher than ORCA_CLI_COMMAND, the command re-runs once through the named
launcher with ORCA_CLI_REEXEC=1 and exits with its status; both variables are consumed so no child
inherits them. Dev launchers export no self on purpose, WSL and SSH names never qualify, and a
launcher that cannot start leaves the command to run here. The Windows launcher no longer rewrites
ORCA_CLI_COMMAND; the legacy ask protocol normalizes its resume command itself.

* refactor(orchestration): declare which flag names the caller on each spec and refuse at the CLI entry

Each handler hand-classified its --from/--terminal as the caller or a target, and the refusal of a
conflicting caller flag ran inside the caller resolver plus two standalone calls for --run listings,
so a new verb that read its flag raw would pass a sibling's handle to a pre-session host. Specs now
declare identityFlagRoles, the CLI entry refuses a conflicting caller flag once from the spec, the
resolver only applies the id-wins rule, and a test fails any orchestration verb that accepts --from
or --terminal without classifying it.

* perf(cli): keep the session caller check off the actor codec's module graph

The check runs at the CLI entry for every command, and the actor codec pulls zod through the session
record. Compare the session's own spellings as plain strings instead.

* refactor(cli): spell a session's address from the one prefix constant, off the codec's module graph

The Orca session address prefix moves to a leaf module with no imports, re-exported by
the address codec, so the CLI entry check derives `session:<id>` from that constant
instead of re-typing it and still stays off the codec's zod graph. Prose and test names
say caller or Orca session id, not actor.

* refactor(orchestration): drop the session id's terminal-view spawn now that the handoff is gone

The terminal handoff was removed, so no terminal is ever a structured session:
- delete the terminal-view identity env and its WSL passthrough, and their tests;
- strip the session caller keys from every terminal's env unconditionally;
- the CLI's own-address spelling moves beside the injected id in src/shared, with
  a test pinning it to the address the host's party resolver gives that session.

* fix(terminal): run the Codex launch preflight through the CLI the terminal names

Packaged Linux names the userData shim in ORCA_CLI_COMMAND, while the preflight
ran the bundled launcher behind it. The CLI saw a different launcher and handed
the preflight off to the shim, booting Electron twice before every codex launch.

* revert(terminal): keep terminals on main's ORCA_CLI_COMMAND and Codex preflight

Only a structured session needs an absolute ORCA_CLI_COMMAND; local terminals go back to
naming none (WSL keeps its guest command), and the Codex launch preflight goes back to the
bundled launcher. The CLI handoff is scoped to sessions, so a terminal's preflight can no
longer be handed off and start Electron twice.

This reverts commit d2cefb6c03 and commit dd2853a5a9.

* fix(cli): hand off to the session's CLI only inside a structured session

The handoff ran whenever an Orca launcher's ORCA_CLI_SELF differed from an absolute
ORCA_CLI_COMMAND, so any process with both - a terminal, a script - ran another install's CLI
instead of the one invoked: a beta's --version lied, and an AppImage command from a terminal
that outlived its Orca failed. It now requires the injected session id, the identity it exists
to deliver. The launcher variables are still consumed in every process.

* fix(cli): name the packaged Windows command after the handoff decision

The launcher stopped writing orca/orca-ide over ORCA_CLI_COMMAND so the handoff could see a
session's absolute launcher, which also changed what every Windows terminal's CLI read. The CLI
entry now applies the launcher's rule itself once the handoff is decided, so terminals and the
legacy ask resume command see exactly what they saw before, and the resume-command reader
goes back to its original form.

* refactor(cli): decide the session handoff from the CLI's own entry, not a launcher export

Every packaged launcher, shim and dispatcher exported ORCA_CLI_SELF so the CLI could tell which
launcher ran it, and compared that with the session's ORCA_CLI_COMMAND. Two launchers of the same
app are different files, so a session that reached its own app through a global orca-ide on Linux
still handed off and started Electron twice, and the export rode artifacts every terminal uses.

A structured session now also names the JS entry its launcher runs (ORCA_SESSION_CLI_ENTRY), and
the CLI compares its own argv entry with it: any launcher of the same app stays, another install
hands off. The launcher scripts, Linux shim and dispatcher go back to main; the Windows launcher
keeps only leaving ORCA_CLI_COMMAND for the CLI to name after the handoff decision.

* refactor(cli): drop the session CLI handoff; the pinned instance and injected id already bind any current CLI

Every current Orca CLI dials the instance ORCA_USER_DATA_PATH names and sends the injected
session id in the orchestration envelope, so a bare `orca` that reaches another install's
current CLI already acts as the session. An older CLI has no handoff code and refuses on the
marker. The handoff only lined up versions between two current CLIs, and comparing two
separately derived paths kept misfiring (an AppImage's mount against its registered
extraction started the CLI twice on every call).

Removes the re-exec, ORCA_SESSION_CLI_ENTRY and ORCA_CLI_REEXEC, and the CLI-side Windows
command naming; the packaged Windows launcher rewrites ORCA_CLI_COMMAND again, as on main,
inside its own process only. resolveHostCliEntryPath goes back to the SSH passthrough.

* test(orchestration): say why the registered worker case pins the handle, now that every session's env is populated
2026-09-28 15:19:44 -07:00
Neil 8c61a5df1f fix(windows): require signed release binaries and identify CLI launcher (#23680) 2026-09-28 14:34:08 -07:00
Neil 0f52bb8be5 perf(ci): use four ARM test workers and remove repeated compilation (#23685)
* ci: benchmark per-job Node compile caching on full unit shards

* ci: measure unit shards with three and four workers

* ci: benchmark localization extraction CLI patch

* perf(build): reuse identical relay bundles across platforms

* ci: compare Vitest 4 and 5 on complete ARM shards

* perf(ci): upgrade localization extraction to skip irrelevant syntax walks

* perf(ci): use all four ARM cores and remove benchmark workflows

* ci: preserve failures while capturing unit source revision

* fix(ci): preserve commented and escaped localization calls

* ci: remove corrected localization benchmark harness
2026-09-28 14:05:30 -07:00
Brennan Benson a84bd1c4fd fix(claude): write only the hook events and statusLine the user's Claude accepts (#23614)
* refactor(claude): name the Claude version module after the hook events it gates

Pure move of claude-session-end-hook-capability.ts and its tests; the next
commit turns its one-event SessionEnd floor into a per-event version table.

* fix(claude): write only the hook events the resolved Claude knows

Claude 1.0.81 through 2.1.100 validate settings.json `hooks` against a
closed event enum and discard the whole file on one unknown name, so
Orca's install made Claude <= 2.1.77 silently ignore the user's env,
permissions and hooks. Each managed event now carries the first Claude
release that knows it (pinned to per-release enums read from the npm
packages), and install, status and the SSH/WSL relay installer write
only the events the resolved Claude accepts. An unresolved version gets
the set every tabled Claude knows; a downgrade removes only Orca's own
entry for an event the older Claude would reject.

* refactor(claude): move the managed Claude hook events into their own module

hook-settings.ts is at its line limit; the event list and its version gate
move out whole so the next change has room.

* fix(claude): an unresolved Claude version never removes Orca's hook entries

A failed or timed-out version probe is no evidence of an old Claude, so it
must not strip StopFailure, PermissionRequest and the other newer events a
version-aware install wrote. With the version unknown, install adds only
the set every tabled Claude knows and leaves every other entry exactly as
it is; only a known version that lacks an event retires Orca's entry.

* fix(claude): gate the core hook events on the Claude release that added them

Claude validates hooks against a closed event list from 1.0.23, not 1.0.81.
The table treated SessionStart, UserPromptSubmit, Stop, SubagentStop,
PreToolUse and PostToolUse as known by every resolved version, so a Claude
from 1.0.23 to 1.0.61 was still sent names it rejects, and it dropped the
whole settings file. Pin each to its first release from the packed enums and
keep the unresolved-version set as its own policy.

* fix(claude): write Orca's statusLine only for a Claude that knows it

Claude 1.0.49 through 1.0.66 also reject any unknown top-level settings
key, and statusLine joined that schema only in 1.0.64. Orca wrote its
statusLine for every Claude, so 1.0.49 to 1.0.63 still dropped the whole
settings file even with the event gate. Gate statusLine on 1.0.64, pinned
by the packed schemas; a known older Claude has Orca's own statusLine
removed along with the opt-out marker, so an upgrade re-adds it. An
unresolved version is now assumed to be 1.0.64, which knows the same core
events and keeps the statusLine install it had before.

* test(claude): a user statusLine opt-out survives a downgrade and upgrade

Retiring Orca's statusLine for a Claude older than 1.0.64 forgets the
install marker only when Orca's own statusLine was removed. Pin that, so
a user who deleted Orca's statusLine is not opted back in by an upgrade.

* test(claude): check the whole written settings file against each strict schema

Claude 1.0.49 through 1.0.66 discard the whole settings file over any
top-level key their schema lacks. The fixture recorded only whether each
release knew statusLine, so a new top-level key Orca wrote would pass every
test. Record each release's top-level keys instead (statusLine is derived
from them), add the hook enums for every packed release in that window,
and check that a real install and a downgrade write only keys and events
each strict release accepts.
2026-09-28 12:04:59 -07:00
Jinjing 95a16e3f67 fix(release): stop the release policy from deleting pipeline-cut releases (#23669)
* fix(release): stop the release policy from deleting pipeline-cut releases

The policy judged a release by who created the release object. Cut Release
reuses an existing draft, so a CI-built v1.4.216 whose draft a person had
created was deleted (tag included) when its notes were edited, and Latest
fell back to v1.4.214 because v1.4.215 was also published by a person.

- Authorize a desktop release when its annotated tag was created by the
  release pipeline and points at its `release: vX` commit, not only by author.
- Only delete on `published`; an edit never deletes a release or tag.
- Pick Latest from the highest authorized stable using the same check.
- Move the policy into config/scripts/release-policy.mjs with tests.

* fix(release): load the policy module from the tagged commit

Release events run the workflow file from the tag's commit, so checking out
the default branch could pair an old workflow with a newer module.
2026-09-28 12:03:58 -07:00
Jinwoo Hong e8d8f2f0b3 fix(deps): take Electron 43.7.5 so detached webviews stop blanking browser tabs (#23586)
Electron 43.7.0 threw 'Invalid guestInstanceId' from <webview>'s
disconnectedCallback for a loaded guest (electron/electron#53989), so a
webview React removed and re-inserted kept a dead guest id: the tab went
blank, reload did nothing, and the destroyed listener never fired.
43.7.4 (electron/electron#54097) returns early when the guest is gone.

Raises the runtime floor test to 43.7.4 so a downgrade cannot re-ship it.

Fixes STA-8757
2026-09-28 13:58:48 -04:00
Neil 21f8e0f9db ci: use faster gzip for temporary Linux test packages (#23609)
* ci: benchmark faster Linux package compression

* ci: pass compression options through typed builder configuration

* ci: retain original configuration for benchmark baseline

* test: preserve release settings in CI compression configuration

* ci: normalize generated changelog dates in package comparison

* ci: remove completed Linux compression benchmark
2026-09-28 03:48:07 -07:00
Neil cf20423ff3 ci: skip unrelated installs and share xterm build dependencies (#23607)
* ci: pilot shared xterm installed dependencies

* ci: bound xterm cache production to verified main entries

* ci: benchmark xterm reuse on the production ARM runner

* ci: avoid installing Orca dependencies for standalone xterm checks

* ci: use Node-only setup in the production xterm job

* ci: remove completed xterm benchmark workflow
2026-09-28 03:46:22 -07:00
8b410b4893 feat: add first-class Qoder CLI support (#23581)
feat: add first-class Qoder CLI support

Integrate Qoder launch, identity, canonical hook status, trust and resume.
Verify with captured Qoder 1.1.64 transcripts and hidden Electron sidebar checks.

Builds on and cross-reviews #7502, #8611, #9655, #12910, #13311 and #15291.

Co-authored-by: dalveytech-vincent <vincent@dalveytech.com>
Co-authored-by: Eridanus117 <45489268+Eridanus117@users.noreply.github.com>
Co-authored-by: xingqingzzp-gif <xingqingzzp-gif@users.noreply.github.com>
Co-authored-by: jyang2004 <jyang2004@users.noreply.github.com>
Co-authored-by: yunqian <yunqian@alibaba-inc.com>
Co-authored-by: huzhening.hzn <huzhening.hzn@alibaba-inc.com>
2026-09-28 02:59:50 -07:00
400e4e7957 feat(agents): add Freebuff launch and sidebar status support (#23567)
Add Freebuff launch support and execution-host status reporting for the sidebar, including running, question, blocked, and settled states. Validate against captured CLI transcripts and real rendered sidebar evidence.

Cross-referenced community implementations #17065, #20839, and the Freebuff portion of #18790. Preserve their agent/catalog/mobile/documentation coverage and add canonical status publication and regression tests.

Co-authored-by: Harkaran Brar <18134082+harkaranbrar7@users.noreply.github.com>
Co-authored-by: Prarambha369 <98906077+Prarambha369@users.noreply.github.com>
Co-authored-by: Lesley Murfin <260182349+LesleyMurfin@users.noreply.github.com>
2026-09-28 02:32:41 -07:00
Neil d05080175b perf(ci): stop duplicating shared Linux download caches (#23578)
* perf(ci): pilot pnpm verification record caching on Linux

* test(ci): review pnpm verification record in mobile cache audit

* perf(ci): share Electron downloads and clean closed PR caches

* test(ci): retain cache ordering checks for restore-only consumers

* perf(ci): limit archive sharing rollout to primed Linux hosts
2026-09-28 01:58:46 -07:00
Neil e6fbbdf684 perf(ci): cache pnpm verification records on Linux (#23568)
* perf(ci): pilot pnpm verification record caching on Linux

* test(ci): review pnpm verification record in mobile cache audit
2026-09-28 01:50:30 -07:00
Neil 3dd7d29455 Revert "perf(ci): shard the anti-slop audit across processes instead of one JS runtime (#23543)" (#23575)
This reverts commit fae0ae7a46.
2026-09-28 01:47:07 -07:00
Jinwoo Hong ba858ee446 feat(mobile): the keyboard covers the page like a native screen, and the shell says its height (#23110)
The shell no longer shortens the WebView for the keyboard; it publishes the keyboard height like the safe-area insets, so native's keyboard lift, refit hold and dismiss key run on the page unchanged. Keyboard and inset arithmetic read the shell's OS through a host-os seam. One page-version floor (manifest pageVersion, shell floor 1) replaces per-feature accept negotiation; a page below the floor gets the existing update wall, a desktop with no bundle keeps native screens. iOS shell drops the form accessory bar and its own keyboard observers. Native session screens untouched.
2026-09-28 04:07:35 -04:00
Neil 080c562898 perf(ci): diff against the merge commit's first parent so PR checkouts can be shallow (#23562)
Every changed-path gate asked git for `--merge-base "$BASE_SHA" "$HEAD_SHA"`,
which needs the event payload's base SHA to be in the local graph. That is the
only reason two jobs cloned all 8127 refs' history. On a pull_request checkout
HEAD is already the merge commit, so its first parent is the base side and no
merge base has to be computed. config/scripts/git-pull-request-diff-base.mjs
resolved that for the two Node gates; the workflow's inline gates now use the
same helper through a small CLI rather than open-coding it.

code_paths gates all 22 jobs, so its checkout is charged to the start of every
one of them: measured 20.7s to 1.6s, keeping blob:none because its sparse tree
is ~7 files and leaves no blobs to refetch. Static analysis drops the filter
instead, since populating all 30,226 files makes blob:none force a second
promisor fetch: 23s to ~11s.

Verified on a real merge ref. At depth 50 the old and new forms produce
identical changed-file sets. At depth 2 the new form still works and the old one
fails with `fatal: bad object`, which is the failure a stale base would have
caused once the checkout stopped being complete.

Also drops the dead resolveBase + merge-base prelude in the changed-code gate,
whose result resolvePullRequestDiffBase already discarded on every PR.
2026-09-28 00:37:26 -07:00
Jinwoo Hong 2077956254 fix(mobile): size a terminal's first subscribe from the document's reported cell box (#23080)
* fix(mobile): size a terminal's first subscribe from the document's reported cell box

#22960 sent phone dims on a terminal's first subscribe by opening a throwaway
empty terminal (init 80x24 ""), awaiting its ready and measuring, behind a
per-document first-subscribe mark whose lifetime was tied to web-ready. That
cost a second xterm/WebGL instance and ~150 ms per open, plus lifecycle state.

The document now measures the cell box without a terminal (xterm 6's
CharSizeService strategy, rounded as the renderer rounds it) for every
text-size preset and reports it with its viewport in web-ready; a table,
because the text scale only reaches the document after that notify. Each
init's ready reports the box xterm actually laid out, which replaces the
probe's entry. The controller answers fitDimensions/measureFitDimensions from
that table and the view's layout with no message; without a table it asks the
document as before.

The session seeds an unmeasured viewport synchronously in subscribeToTerminal,
so the first subscribe carries dims by construction. Deleted: the empty init,
its awaitReady gate, deferFirstSubscribeUntilViewportMeasured and the
subscribedDocuments mark. The fit pass is unchanged and still covers a
document that reports no cell box.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): read the reported cell box through in-narrowing, not Reflect.get

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): correct the probe's cell-box guess from the box xterm lays out

The web-ready probe is a guess: building the WebGL addon creates no context,
so a context that fails on load lands on the DOM renderer, whose width is not
snapped and depends on the column count. Before, a ready box that differed was
only logged; the first subscribe had carried the wrong column count, the host
echoed it, the fit pass saw the viewport equal to the host's dims, and the grid
stayed slightly shrunk. The store also kept the WebGL width after a context loss.

The document now reports the box xterm laid out whenever it changes (from
onRender, which covers a renderer swap and a DPR change that
onDimensionsChange does not fire for, and at ready). The store replaces the
guess; when that changes the current text size's entry, the view calls
onCellBoxChange with xterm's grid and the session re-fits, running the bounded
fit pass if the dims moved (one resubscribe). Equal boxes do nothing.

The RN layout box now survives a document reload; the document's own viewport
only stands in until the view reports a layout (on the page, web-ready arrives
first). The mismatch console.log is gone, and the probe's rounding names the
xterm version it copies.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): correct each cell-box guess at most once, so a DOM renderer cannot loop

On the DOM renderer the cell width is the rounded canvas width divided by the
column count, so every re-init at new cols reported a new box. Each one counted
as a correction, a floor over floats could flip the fit between two sizes, and
each flip landed converged, which reset the resubscribe budget: an unbounded
series of full-snapshot resubscribes.

Only the first laid-out box for a guessed text size may be a correction; later
reports still update the store, so fits stay truthful, but never resubscribe on
their own. The fit's floor gains a 1e-6 epsilon so floating-point error at an
exact boundary cannot flip a column or row. New document tests pin the render
report after a renderer swap and the report at ready for a paused renderer.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): make xterm the only terminal cell measurer

The document builds its real terminal before web-ready, at the app's
text scale, and reports the box xterm laid out; the first init reuses
that terminal. The page-side prediction, the per-scale guess table and
the once-per-document correction are gone. The app remembers the box
per text scale for its lifetime, so a later open at a known scale
subscribes with phone dims at once. A box that changes at the same grid
(renderer swap, pixel ratio) refits the open terminal in place.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): keep commands queued before the terminal WebView first loads

A subscribe sized from the stored cell box can queue init before the
native WebView reports its first load start, which cleared the queue
and left the terminal blank. Only a reload now drops queued commands.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): re-init a document that lacks the subscription's init, and fit one frame width

- Web-ready now says whether the document holds the terminal's latest init
  (a reload before the first ready drops a queued one); the session
  resubscribes any initialized terminal whose document lacks it.
- One grid fit, shared by the app and the document, fed the unrounded frame
  width React Native laid out; it keeps exact fits whole at fractional pixel
  ratios. The document's viewport-width fits are gone.
- The page builds every document at the scale the view mounted with, as the
  native WebView does.
- A new document's first cell box is compared against the grid the
  subscribe fitted from the stored box.
- The terminal built before ready stays hidden until its first init.
- The cell-box census matches glyph-measurement techniques, not names;
  the store's unused clear() is gone.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): build the terminal before ready only for the view shown at mount

A session mounts one terminal view per tab, and each built xterm and a
WebGL context before ready: 20 tabs made 20 contexts at load, past the
~16 a page (or Android's shared WebView renderer) holds, and native logged
32 context losses. Only the view shown when it mounts builds early now;
the rest build at their first init as before. Deferring the WebGL addon
instead would change the reported box: the DOM renderer lays out 7.8x15
where WebGL lays out 7.667x15 at the same font.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): write a WebView document's start values into its page, not an injected script

Android ran the pre-content injected script after the document's own in
1 of 22 documents on the emulator; that document started with no text
scale or shown flag and built a terminal it should not have. The values
now sit in the page ahead of the document script, one source object per
start pair so a render never reloads the WebView.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): the pre-ready terminal measures and reports while hidden

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): measure only the laid-out frame, and refit on a new grid, not a new width

- A measure needs both of the frame's dimensions from React Native; the
  document's viewport-height fallback is gone, and before the first layout
  the handle answers no fit without asking the document.
- A frame width change that still fits the PTY's grid from the stored box is
  a no-op, so sub-pixel layout jitter no longer re-measures. The width ref is
  written in that effect rather than during render (react-doctor).
- One "last grid" ref: the last reported grid, or the one a subscribe fitted
  from the stored box.
- The page render rig measures through the frame it laid out, as the session
  does, and lets the replay's fit settle before its resize-refit witness.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): let only the current terminal document's ready flush

A reload kept the WebView and its onMessage, so the old document's late
web-ready flushed the queue into the reloading view and the new document
got a second init. Each document now gets its own view (keyed on a
generation the controller owns), every notify carries the generation of
the view that received it, and a web-ready from a replaced document
flushes nothing and stamps nothing.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): drop every notify from a replaced terminal document

One rule at the receive boundary: a notify from any generation but the
current one is dropped, whatever its type, not only web-ready.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): make fitDimensions a pure question; name each generation counter

- fitDimensions no longer records the grid. A width change to a new grid
  asked it first, so the DOM renderer's report of that grid's box read as
  "same grid, new box" and refit again. Only the first-subscribe seed
  (seedFitDimensions) records the grid the document's first report is
  checked against.
- viewGeneration counts the views, readyGeneration counts web-readies.
- replaceDocument no longer resets the load flag; the load-start reset
  stays as the guard for a view that reloads itself.
- The name-based lifecycle census is replaced by a behavioural test.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): typecheck the handle mocks, drop the unused cell-box get

- The two handle mocks carry both fitDimensions and seedFitDimensions,
  and the fake-timer acts return nothing, so the three test files check
  under tsconfig.test.json again.
- terminalCellBoxes.get had no product caller; the store's tests assert
  through fit.
- The load-start comment says what the controller does now.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): hold the grid the document has, ignore a replaced view's load start, dispose a failed pre-ready terminal

- The document reports a new grid even with an unchanged box, so an
  in-place reflow on WebGL is held before a later renderer swap at that
  grid; the swap then refits. The app's apply paths do not hold the grid
  themselves: the DOM renderer's box follows cols, and a grid held on
  apply would read its own box as a renderer change and loop. One
  writer (holdGrid) holds the seeded or reported grid.
- A load start from a view a replacement unmounted is ignored, as its
  notifies already are.
- A terminal whose open throws before ready is disposed, not only
  unreferenced.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): ignore every native event from a replaced terminal view

One wrapper binds each WebView lifecycle event (load start, error, HTTP
error, render process gone, content process terminated) to the view's
generation, so a replaced view's late event cannot reset, replace or
put an error over the current document.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): a DOM seed refits once on its first report, not on the refit's own

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): subscribe a terminal only after its document is ready

The document still builds its terminal before ready and reports the cell
box xterm laid out in web-ready; the app now subscribes after that ready
and fits from that box, so nothing is sent to a document before it is
ready. Everything that made a pre-ready subscribe safe goes: the
app-lifetime box store, the seed fit, the per-document view generations
and their event filtering, the init tracker and the hasInit resubscribe.
The native view reloads in place again and web-ready keeps main's reload
rule. Boxes are kept per view; the grid a document last reported still
guards the in-place refit against the DOM renderer's cols-dependent box.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): hold one reported cell box and the grid the subscribe fitted

The controller keeps only the box the current document last reported,
not a per-text-size store: the document re-reports on a scale change.
The subscribe after ready fits from that box and holds the grid it
fitted, so the DOM renderer's first report at that grid (a new box)
refits once in place and converges; refit and apply paths hold nothing.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): fit only a ready box at the app's scale; forget a reloaded document's box and grid

A reload keeps the document's mount scale, so a ready after a text-size
change reports a box at the old scale; that box no longer sizes the first
subscribe, which then takes the no-box path. A readiness reset drops the
old document's box and held grid, so the new document's first DOM report
at the same grid does not refit.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): give the terminal document its frame at init, and fit text scale over it only

A subscribe sized from the ready box sends no measure, so the document
had no frame when the text size changed and reported the pre-refit row
pitch. The app's init now carries the frame it laid out, in the fields a
measure uses; the router takes it from either. The text-scale fit reads
only that frame, with no viewport fallback.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): say why a frameless text-scale change skips the resize

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): one cell box per terminal notify, not an array

web-ready and cell-metrics carry `cellBox: {fontScale, cellWidth,
cellHeight} | null`; the document's `laidOutCellBox` returns one or
null and the parser validates one object. The text-scale match moves
from web-ready into `handle.fitDimensions`, the one place a box is fitted.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): fit terminals in the app from the reported box; drop the measure round trip

The app already holds the box the document reported, so the refit and
the fit pass await the init's ready and call `handle.fitDimensions`
instead of posting `measure` and waiting on `measure-result`. The
document's measure, its retries, and the measure promise and timeout go.
The document still resizes locally on a text-size change, so every grid
the app sends (init, resize, reflow) carries the laid-out frame it was
fitted to. `holdSubscribedGrid` replaces `subscribeFitDimensions`, so
the only fits are `fitDimensionsFromCell` and `handle.fitDimensions`.
The render rig reads its fit from the ready box. The recorder adapter
mounts the new handle with the same recorded effects; the goldens it
mounts move on their adapterSha256 header only.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): keep the terminal frame in one ref, and notify a new width imperatively

The session held the frame in a height ref, a width ref, a width state
and the refit's own width ref. It now holds one `terminalFrameRef`
({width, height} | null until the first layout; a hidden 0x0 layout
keeps the last box). onLayout notifies a new width imperatively, as it
does height, and the refit's notify skips a width whose fit is the grid
the PTY has. `terminal-frame-width-refit.ts`, the width state and its
effect go. The subscribe's layout gate reads "no frame yet" directly.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): subscribe a held-back terminal on the frame's first layout only

`handleTerminalFrameLayout` ran on every onLayout; it now runs once, when
the frame first has a size. Later layouts only notify a new width.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): size the first subscribe inline in subscribeToTerminal

`sizeTerminalViewportFromCellBox` wrapped five lines in a 37-line
module; the subscribe now fits the ready box against the frame, holds
that grid and records the diagnostic itself. The helper's tests fold
into the subscription tests, which move to the subscription's name.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): drop the unreachable font-size guard on the reported cell box

xterm 6.1.0-beta.303 updates the render service's cell box in the same
task that sets `options.fontSize`: CharSizeService.measure fires
onCharSizeChange, and RenderService.handleCharSizeChanged runs the
renderer's `_updateDimensions` (DomRenderer.ts:359, WebglRenderer.ts:229).
`term.onRender` fires from RenderService._renderRows after the rows
are drawn (RenderService.ts:213, CoreBrowserTerminal.ts:538), and the
document writes its text scale and the font size in one task
(text-scaling.ts applyTextScale, terminal-init.ts init). So no report
can read a box between the font and the scale; the guard and its test
go. A new test pins the real order: no report when the font is set,
the new box at the new scale on the next render.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): one start seam, no source cache, the reported box as an object

- `useState` already pins each view's WebView source at mount (a new
  test re-renders at another text scale and gets the same object), so
  the module-level `webViewSources` Map goes.
- `initialTextScale` and `buildsTerminalBeforeReady` become one
  `start(): { textScale, shown }` seam.
- `reportedCellBox` holds the last reported box and grid, not a string key.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the RPC recordings to this branch and re-record

The terminal refit now fits in the app from the reported box and reads
one frame ref, so the recorder's terminal adapter mounts the new handle
(`awaitReady` + `fitDimensions`) and options (`terminalFrameRef`),
keeping its recorded effects. `baseline` is repinned to 21954dbd2f, the
last commit to touch a fenced path, and every golden is re-recorded.
Proof by class against HEAD: 787 header-only, 0 body moved, 0 added,
0 deleted. Header keys moved: `baseline` on all 787, and `adapterSha256`
on the 14 goldens `terminal-mount-adapters.ts` mounts (query-reply 3,
accessory-raw-send 4, takeover-report 4, viewport-refit 3). No recorded
traffic moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): hold the reported cell box and its grid in one ref

The controller kept the box in `cellBoxRef`, the grid in a string
`lastGridRef` and wrote it through `terminal-held-grid.ts`. One
`heldRef` now holds `{ cellBox, grid }`, as the document's own
`reportedCellBox` does: web-ready writes the box, every cell-metrics
report writes both, `holdSubscribedGrid` writes the grid, and a
readiness reset clears it. Same write points, so the one-refit bound
holds; the DOM-loop and refit-once tests pass unchanged.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the RPC recordings to the hold-rule commit and re-record

H (f00bebba48) touched a fenced path after the last repin, so
`baseline` moves to it and every golden is re-recorded. Against the
corpus before this branch's refreshes (21954dbd2f): 787 header-only,
0 body moved, 0 added, 0 deleted; `baseline` on all 787 and
`adapterSha256` on the 14 goldens `terminal-mount-adapters.ts` mounts.
Against the previous refresh: `baseline` only. No recorded traffic moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the RPC recordings to the main merge and re-record

The merge (b3b1b0def2) is the last commit to touch a fenced path, so
`baseline` moves to it and every golden is re-recorded. Against
97b5bb2b9a: 787 header-only, 0 body moved, 0 added, 0 deleted;
`baseline` on all 787, and `adapterSha256` on the 14
session.diff-review-actions goldens whose adapter #22951 edited. Against
origin/main: 787 header-only, 0 body moved/added/deleted; `baseline` on
all 787 and `adapterSha256` on this branch's 14 terminal goldens.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): the terminal document holds the grid and decides each refit

The document already kept the last reported box and grid; the app kept
a mirror of both to decide the refit. Now the document decides: its
`cell-box` notify carries `{ cellBox, refit }`, sent only when the box
changes, with `refit` a box that changed at a kept grid. web-ready
records the pre-ready terminal's box at its 80x24 grid, and the first
init that reuses that terminal holds the init's grid, so the DOM
renderer's first report refits once, as the subscribe's hold did. A
re-init no longer clears the record, so a new renderer at the same grid
still refits. The app keeps one `cellBoxRef` and `holdSubscribedGrid`,
`heldRef` and the grid on the notify go. The one-refit, DOM-loop and
renderer-swap tests move to the document with the same scenarios.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): one init options object, and a frame on every grid

`init` takes `{ cols, rows, data, preserveScroll, oscLinks, frame }`
instead of six positionals, and `init`, `resize` and `reflow` (handle
and messages) require `frame: TerminalFrame | null`. The refit's reflow
check reads `!dims` alone, and the controller's test file is named for
the `cell-box` notify it now covers.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): one notifyTerminalFrame for the frame's layout

The frame's onLayout made four calls and held the classification
itself. It now calls `notifyTerminalFrame({ width, height })`, and the
session's terminal-webview hook keeps the one frame ref, notifies the
height, subscribes the document held back for the first layout, and
notifies a later width change. `handleTerminalFrameLayout` is named for
what it does: `subscribeIntendedActiveTerminal`. The layout tests move
to that hook.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the RPC recordings to the round-8 head and re-record

85d421963c is the last commit to touch a fenced path. Against
a676c1b65a: 787 header-only, 0 body moved/added/deleted, `baseline`
only. Against origin/main: 787 header-only, 0 body moved/added/deleted;
`baseline` on all 787 and `adapterSha256` on this branch's 14 terminal
goldens. No recorded traffic moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* refactor(mobile): name the init option initialData, as the message does

The init option `data` becomes `initialData`, the message field's name,
so the controller passes it through unrenamed. The `preserveScroll` why
stays on the message type only, and the document test's title names the
three grids that carry the frame.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the RPC recordings to the round-9 head and re-record

486566c82b is the last commit to touch a fenced path. Against
3371c39715: 787 header-only, 0 body moved/added/deleted, `baseline`
only. Against origin/main: 787 header-only, 0 body moved/added/deleted;
`baseline` on all 787 and `adapterSha256` on this branch's 14 terminal
goldens. No recorded traffic moved.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* ci: rerun checks against main with #23560 landed

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-28 03:26:49 -04:00
Brennan BensonandClaude 7a24d3d335 fix(native-chat): the conversation outlives its agent (#22835)
* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* fix(native-chat): name a chat write by its target, not the owner generation

A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.

Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.

Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.

* fix(native-chat): every journal append reaches the chats that are open

A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.

A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.

* test(native-chat): an epoch replacement reaches the open chat

* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

* perf(native-chat): a publish behind a delivered commit reads nothing

Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.

* test(native-chat): state why the teardown test's fake journal is safe to cast

* docs(native-chat): say mutation admission checks only the writer lease

* docs(native-chat): drop the send rebase from comments that still described it

* fix(native-chat): a message is accepted, then delivered

A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".

A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.

Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.

A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.

Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.

* fix(native-chat): settle queued messages only for the child that ended

A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.

A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.

The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.

* fix(native-chat): an adoption that fails to import keeps the conversation open

The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.

* perf(native-chat): the recovering open reads the journal once

Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.

* fix(native-chat): an attach that fails after indexing its child leaves no child behind

A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.

* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer

The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.

A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.

* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down

The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.

* fix(native-chat): a message rejected while its chat was closed reads as not sent

A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.

* test(orchestration): name why the readiness settlement fakes are cast

* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent

* docs(native-chat): drop the fence from the admission the send effects run behind

* docs(native-chat): give the fence move on release the reason that still holds

* docs(native-chat): stop citing a write fence check in launch and mailbox comments

Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.

* refactor(native-chat): the provider child is its own record

A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.

- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
  patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
  own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
  the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
  dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
  is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
  the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.

* fix(native-chat): the delivery loop alone settles a message its start or child failed

A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.

- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
  reads how it ended: a Stop continues; anything else writes one failure row and rejects every
  queued message with the same words, then stops. A child still starting whose start the adapter
  says did not land fails the same way. The exit, eviction and the settlement retry only settle
  the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
  nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
  failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
  closed, with or without a child, and a start the loop already has in flight is waited for so the
  child it produces is stopped rather than left behind.

* refactor(native-chat): a stopped child ends on the one reading of its stop

The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.

* feat(native-chat): the host says it accepts a send before any agent has it

The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.

* refactor(native-chat): an attach never opens a journal of its own

The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.

* fix(native-chat): a moved fence resends nothing on a host that accepts first

The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.

The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.

* refactor(native-chat): a child's end says whether the user or the host stopped it

The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.

* fix(native-chat): a chat whose only work is a queued message is not offered for resume

A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.

* fix(native-chat): the conversation outlives its agent

Opening a chat no longer starts its agent. A conversation is reached through one host
accessor that opens its journal at rest, and a send is what starts the agent, through
the delivery loop. One idle sweep, every five minutes, stops an agent that has been
quiet for thirty minutes and owes no work, then drops an open journal handle that is
only a cache. Its record, tab, status row and readers stay.

- hold and release are no-ops; hold still builds the host for shipped mobile builds.
- The holders, the holds, the release clock and the exit respawn are deleted.
- Options, the model list, the goal and the context meter answer at rest; a model pick
  at rest is recorded as intent for the next start.
- Compact, rewind, clear and goal changes start the agent first. A send does too when
  a rewind is still in doubt after the conversation opens.
- Orchestration routes mail and group addresses on ownership (the record plus the chat
  tab), not on whether the process runs. An open dispatch keeps its worker running.
- The restart continuation is a send; Resume all holds each slot until the message is
  handed over or rejected.
- A read error never replaces a loaded transcript, and shows the host's own words.

* test(native-chat): type the queued-message fixtures in the resume-offer tests

* fix(native-chat): a start that dies while a message waits on it is that message's failed start

Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.

* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now

* test(native-chat): pin what a failed start settles, and what a resume offer names

A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.

* test(native-chat): the failed-start pins fail on what the message became, not on a timeout

* fix(native-chat): a restart offer ends when the chat's agent starts again

The offer used to end only when the chat's newest user message changed,
because opening a chat started its agent and that start could not be told
apart from real activity. Opening a chat starts nothing now, so the host
reads the fact it already publishes: a chat's status row goes from not
host-owned to host-owned exactly when its agent is started. At that edge the
offer and any failure record for the chat are withdrawn, unless the start is
a resume action's own (its continuation is the oldest undelivered message).

A continuation and a message racing to be first are decided at acceptance:
the continuation is refused, quietly and with nothing filed, when any other
message was accepted since the restart. A failed continuation start leaves
the offer retryable, and each resume action sends its own message id.

Deleted: the newest-user-message comparison, its journal reader, the
continuation filter, and the failure ledger's own "answered by the chat"
check. The marker still carries its message id for one release, so the
previous build can read it.

* fix(runtime): end a transcript stream when its client unsubscribes

Desktop: the IPC subscription controller was dropped as soon as the streaming
handler returned, which for most streams is right after it binds. A later
runtime:unsubscribe then found nothing to abort, so the host kept the subscriber
and derived and sent every publish to a channel no one listened to. The controller
now lives until the renderer unsubscribes, resubscribes the same id, or goes away.

Mobile: disposing an agentSession.subscribe stream now sends agentSession.unsubscribe
with the stream's frame id, so the host ends that subscriber and leaves a sibling
stream on the same socket running. The direct path now passes the frame id the relay
path already passed.

* test(orchestration): the preamble's host stub is typed, not cast

The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.

* fix(native-chat): one fact ends a restart offer: the chat moved on since the restart

The offer is live while no other message has been accepted in the chat since the
restart and its agent has not proved a start since. The offer list, the resume's
reservation check and the continuation's acceptance check all read that one fact,
so a message whose start then failed withdraws the offer too, and a stale click
finds nothing to act on.

The fact is read off the conversation's open handle, which the restart closed, so
it is retired durably whenever it may have changed: a message accepted, a start
proven. A close and reopen within the same run therefore cannot bring the offer
back. A continuation rejected before it reached the agent does not count, so a
retry after a failed start still runs.

Deleted: the quit-time gate on withdrawal, which changed nothing because the
withdrawal and the quit's own offer write share one queue; the per-action
"withdrawn" flag and the separate acceptance check it paired with.

* test(native-chat): an older build reads the restart offer this build records

The offer lives in a file the previous release reads after a downgrade. Pin that
against the pinned release's own capsule, and run the lane when the marker or the
capsule changes.

* fix(native-chat): read a restart offer against where the journal stood when it was taken

"Since the restart" was read off the conversation's open handle, which the idle
sweep closes: after a reopen, a message the user had already sent looked older
than the handle and the withdrawn offer came back.

The offer now records the journal position (epoch and sequence) at the moment
it is taken, and a message accepted after that position, or a journal on another
epoch, means the chat moved on. That is derived from the journal, so it holds
across any number of closes and reopens. An older build's offer has no position;
only a start withdraws it. Because the message half is now durable, the offer is
no longer rewritten in the recovery file on every accepted message; a proven
start still writes it, since only the host that saw the start knows of it.

* test(native-chat): wait for the listing's retire write before reading the recovery file

* fix(native-chat): keep the terminal-backed chat's read error over its local echoes

Messages winning over a read error is right for the structured chat, whose read retries and whose
messages came from the transcript. The terminal-backed view assembles its list from local echoes
too (a launch prompt, a pending send), so a failed read there showed only those bubbles and no
error. Only the structured pane now keeps messages over an error.

* fix(native-chat): a start retries the exit settlement a failed journal write left owed

An agent exit whose journal settlement write failed releases the lease latched until a retry lands.
Reopening the chat used to be that retry; with reveal now only opening the journal, nothing retried
it before the next app launch, and every send was refused. The start the send needs now runs the
retry first, where the attach would.

* perf(native-chat): answer the owner check without opening the chat

Worktree activation calls agentSession.handoffStatus for every chat tab in the worktree, and the
answer comes from the session record alone. Reaching it through the accessor opened each resting
chat's journal (a full read, the crash-boundary write and a restored status publish), then kept it
open for the idle window. It now checks the record and the adapter's support, as before this series,
and opens nothing.

* fix(native-chat): a read waiting on the session lock opens nothing once quit began

The accessor checked for quit before queueing the open, so a read queued behind a session task ran
its open after teardown had begun and indexed a journal no teardown step would close. The check now
runs at the open itself.

* fix(native-chat): read a failed resume's chat before calling it retryable

Whether a failed resume is retryable is the offer's own rule: the chat has not moved on since the
restart, read from its journal. The failure list read it only for a chat already open, so once the
idle sweep closed a chat the user had moved on in, its failure showed Retry again, and the click did
nothing. The list now opens the failed chats first, as the offer list does.

* test(native-chat): type the provider event sink the settlement test reaches for

* fix(native-chat): say the structured read keeps trying only where it does

The structured pane's "Orca keeps trying to load it" line never showed: the view state filled in an
untranslated fallback whenever the read error had no text, and the empty state prefers any message.
The view state now leaves the message out, so the structured pane shows that line and the
terminal-backed pane its own translated one. Mobile's structured lane does not resubscribe after an
error frame, so it no longer makes the claim.

* test(native-chat): await the send's settlement instead of polling for the start

The at-rest send tests polled for the provider start with vi.waitFor's one-second default, which a
loaded machine outran. They now await the host's own settlement of the message.

* fix(native-chat): a restart offer resumes any time after the quit, and knows its own continuations

The continuation's message id was dated by the quit, and the ledger refuses a new id dated more than
a day back, so Resume or Retry a day after quitting was always refused (on main too). It is now
dated by the resume action.

Telling a rejected continuation from the user's own message read the operation ledger, whose rows
expire after about a day; after that a failed resume stopped being retryable. The offer now
records the continuation each action sends on its own capsule entry, bounded to the newest 16, so
the ids end with the offer. The ledger read is deleted.

* fix(orchestration): route no mail to a structured worker its orchestration released

A structured worker is routed on ownership, and a resting worker's lease is released, so ownership
held while its chat tab stayed listed. A worker the coordinator abandoned and then released, found
at rest by the release, therefore still took peer mail and @worktree: broadcasts, and each one
restarted its agent. Routing now also reads the orchestration's own resource row: once it is
released, direct mail, group addressing and worker-show's addressable answer drop the worker, as
they would a terminal worker whose terminal closed. The chat tab stays, and nothing new is stored.

* fix(native-chat): a failed retry names the user's prompt, not Orca's continuation

A resume's continuation is written to the chat before its start, so after a failed attempt the chat's
newest user message is that rejected continuation. A second failure then showed Orca's own restart
text as the chat's prompt. A retry now keeps the prompt its first failure named.

* fix(orchestration): read the released row optionally, as the authority does

worker-show's observation called the row lookup directly, which a runtime double without it threw on
and failed the structured tab-retirement release.

* fix(native-chat): the status bar drops a restart offer the chat moved on from

The renderer re-read the host's restart offer only when a failed chat showed activity, so after a
message withdrew a pending offer the host answered no chats while the status bar kept counting one,
and clicking it opened nothing. The same watch now covers pending offers: a status change in an
offered chat asks the host again, once.

* test(native-chat): a roster of idle or finished children does not keep an agent awake

The sweep reads owed background work through the shared child-work liveness that upstream's
release clock adopted; a child that went idle or finished is not work the agent still owes.

* fix(orchestration): a task dispatched into a resting structured worker keeps it running

The sweep's open-dispatch check read only the worker-start dispatch that owns the worker's terminal
resource, so a task later dispatched to the same worker (orchestration dispatch --to, which writes a
dispatch with no worker row) did not count: after thirty quiet minutes the worker was stopped while
that task was open, and its coordinator read exited. Any unsettled dispatch addressed to the worker's
process incarnation now counts, derived from the existing rows.

* docs(native-chat): comments stop describing the hold this PR removed

Eight comments still justified orderings and teardown choices by a viewer or dispatch hold that
pinned the provider child. Nothing holds any more; the orderings stand for the binding's redrive
subscription and parked mail, and a chat's agent runs from a send until the idle sweep rests it.
Comment-only.

* fix(native-chat): a restart offer keeps the start its own continuation made

Whose start ended an offer was decided at read time, from whether the offer's continuation was
still the queued message. Once the provider refused that continuation, the child it had started
read as someone else's start, so the offer ended and its failure showed no Retry. The delivery
loop now records which queued message a start is for on the in-memory child, and the child's end
carries it; the offer counts a start as its own when that message is one of its continuations.

* fix(native-chat): an agent gets a full idle window after its owed work ends

The sweep measured quiet only from the last journal row, so once a subagent, command, monitor or
dispatch that had outlived the window ended, the agent was stopped at the next tick. A child can
read done before the lead's wake-up turn writes anything, and stopping in that gap loses the
wake-up. The sweep now counts owed work it observes as activity, which gives the agent the full
window afterwards, as the release clock it replaced did.

* test(claude): the options-read fixture runs a live child

The fixture marked its conversation running with a hasProviderChild field the
session type does not have, so the read took the at-rest path and refused a
session with no record. It now carries a child, which is what the read checks.

* test(native-chat): host tests reach its collaborators through a typed seam

The rest-test rig and three test files read the host's private members with
Reflect.get and cast the result. The host now exposes one test-only accessor,
collaboratorsForTests(), and the subscribers class a subscriberCountForTests()
beside its existing retainedActivityCountForTests(), so the tests are checked
against the real types and the casts are gone.

* refactor(orchestration): one owner answers a structured worker's custody

Routing, group addressing, worker-show and the idle sweep each composed their own reading of
whether orchestration still holds a structured worker, so each new obligation or retirement state
had to be added to every reader. structured-worker-custody now derives both answers from the
worker-terminal list state coordinators see in worker-list: addressable is owned and not released,
and owed work is an active custody or an unsettled task dispatched to the same incarnation. The
owner's state is read through the remote dispatch attachment too, as the terminal transfer lookup
already does. Behaviour is unchanged; a settled worker awaiting its coordinator still rests.

* refactor(orchestration): owed work is an open dispatch on the worker's incarnation

A supervised worker's own dispatch context stays open exactly while the worker is active, so the
separate active-custody branch only repeated it. Owed work is now one fact, which also states the
policy that a worker awaiting its coordinator's decision may rest, and both custody decisions are
written once at the top of the module.

* fix(native-chat): a restart offer knows its continuations by a tag in their id

The offer recorded each continuation id in a list on its capsule entry, capped at 16, and a running
action's id in memory. Both could disagree with the journal: past the cap an old rejected
continuation read as the chat moving on, and a crash during a retry restored the failure's older
entry, which lacked the retry's id. Each continuation id now carries a tag derived from the offer
(its teardown and chat), then the action's own part, so any continuation of this offer, queued or
rejected, is recognised from the journal row and the marker alone. The persisted list, its cap and
the in-memory action map are deleted; the agent-start withdrawal keeps an offer whose own
continuation the start was for, read against the stored marker.

* test(runtime): the legacy-worker reveal test judges its stale snapshot inside the wait

The tui-idle probe reads through readTerminal, which now awaits the structured
worker check before the PTY read, so the probe's snapshot request starts a
microtask later. vi.waitFor missed it on its first check and polled again at
50 ms, the same moment the wait's own 50 ms timeout fired. The stale snapshot
then resolved after the wait had already timed out, so the test passed without
judging it, and the rejection landed before any handler was attached. Vitest
reported that as an unhandled error and failed the shard.

Polling every 1 ms sees the request within a few ms, so the snapshot is judged
while the wait is still pending.

* fix(native-chat): a view never restarts a chat whose last start failed

A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.

* fix(native-chat): the idle sweep reads owed work every tick

Owed work counted as activity, but the sweep read it only once the idle window had elapsed, so it
refreshed the clock at most once a window. Work that ended just before the next read left the
agent to be stopped at that read, moments after the work ended, which is the gap the refresh was
meant to cover. The sweep now reads owed work on every tick for a started agent, so the window
always runs from the last tick that saw work owed.

* fix(native-chat): a continuation handed to the agent stays sent

The offer read its own continuation as not reaching the agent while its dispatch was pending, which
also covered one already handed over and still unanswered. When the wait for that answer ended first,
the failure it filed read as retryable, and a retry sent a second continuation to an agent that may
have acted on the first. Only a continuation still queued, or rejected, is now read as unsent.

* test(native-chat): start the child the loop waits on with an attach, not a second view

A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.

* fix(native-chat): settle a gone generation's turn wherever a conversation opens

A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.

* test(native-chat): prove the next child's start settles the turn an earlier child left

The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.

* test(native-chat): count a failed start's rows by row, not by text

Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.

* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget

* test(native-chat): the interrupted create's own retry continues again

The merge of main's lease-latch fix replaced that test's retry of the interrupted create, under its
own operation id, with a fresh start whose result nothing read. That fresh start passes with the
released-reservation continuation deleted, so the case the fix exists for went untested. The retry
and its assertion are main's again.

* docs(native-chat): three comments that still had views starting agents

A start with nothing queued now comes from a command, goal change or rewind; an interrupted compaction
left alone would refuse every send, so no agent would ever start to finish it; and a current host
raises the unattached read refusal only once quit began, with the attach window belonging to an older
host.

* test(native-chat): pin the open's and the send's start and row counts, however the view binds

Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.

* test(native-chat): a reader's open settles the turn a failed exit settlement left running

An exit whose settlement write failed leaves its turn running in the open journal. PR 1's open now
settles it, and this pins the two reads that reach it here: a reader reopening a chat the idle
sweep closed, and a read that opens the chat before the restart restore reaches it.

* test(native-chat): the view-start test's starting window outlasts two subscriptions on a loaded runner

A subscription reads the conversation before it returns, so under load the two views took longer
than the create child's 300 ms start, which then exited before the test checked that it had not.
The child now takes a second to fail.

* fix(native-chat): settle a gone generation's turn at every open but an acquisition's

The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.

* test(native-chat): hold the create's start open until the views bind

The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.

* test(native-chat): a read that reaches a crashed chat before the startup reconcile settles its turn

On desktop the chat on screen at relaunch reads before startup reconciles the leases, while the
dead process's lease still reads live. The open settles the turn it left running anyway, and the
restore that follows finds it settled.

* refactor(native-chat): drop the composer's second error formatter

After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.

* test(native-chat): pin the reason on a message rejected while its chat was closed

The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.

* docs(native-chat): drop the removed dispatch hold from six comments

A worker's session no longer takes a dispatch hold, and no release clock
rests a chat by visibility; the agent-launch comments, the abandon test,
the teardown test and the refusal census still said so.

* test(native-chat): rest the owner-status chat through the idle sweep, not a hold

The activation-gate test from #22808 put its chat at rest by holding and
releasing it, and passed the release-clock grace. This branch deleted both,
so the case threw before it reached its assertions. It now moves the host's
clock past the idle window and lets the sweep stop the agent and close the
conversation, then asserts the same owner answer and activation gate.

* fix(native-chat): show the structured pane's retrying line when a read fails

The read transport always hands the pane the host's words, so the error
state's "Orca keeps trying to load it" line, which showed only when there
were none, was never seen: the pane showed the host's text twice, as its
subtitle and on the status line under it. The structured pane now always
says its read keeps retrying, and the host's text stays on the status line.
The terminal-backed chat is unchanged.

* test(native-chat): wait for a send's background start before the refusal oracle removes its store

An accepted send wakes the delivery loop, which starts the agent in the background. The oracle's teardown disposed the loop but did not wait for that start, so its lease write could create a temp file in the store directory while the directory was being removed, failing the test with ENOTEMPTY about one run in four. The teardown now drains tracked starts before it closes the journals.

* fix(native-chat): a start a message waited on gets one failure row, the delivery loop's

When a queued message's start failed, two writers could report it under the same row: the delivery loop, when the adapter settled the start without proving it, and the exit settlement, when the child's exit landed. The last one won, so the chat's row could name a different cause than the one the message was rejected with, or be written twice.

The exit settlement now writes the start's row only when no message is queued and the loop has not already recorded that start. A start for a command, goal change or rewind, with nothing queued, still gets its row from the exit.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-27 23:46:14 -07:00
Brennan Benson 5e5f4f6603 fix(native-chat): keep a Codex ask's questions in the order it asked them (#23502)
* fix(native-chat): keep a Codex ask's questions in the order it asked them

Codex journals every question of one ask in a single write, so the
questions share a timestamp. The transcript list sorted rows by
timestamp and broke ties by row id, and a question's id ends in the
question id the model chose, so answered, cancelled and still-pending
rows of one ask came out in the alphabetical order of those ids.

The transcript projection now breaks timestamp ties by the order its
source gave: the journal's order for structured sessions. The session
assembler keeps its id tie-break, since the sources it merges share no
order of their own.

* refactor(native-chat): name the id tie-break comparator for what it does

Two shared comparators differed only in whether they break timestamp ties by id,
under near-identical names; spell the id tie-break in the name.

* fix(native-chat): order structured chat rows by their journal position

The desktop list and the host's conversation outline sorted structured rows
by timestamp. The journal's contract is that the sequence orders the
timeline and the timestamp is the provider's clock: a Codex ask writes all
of its questions in one write (one sequence, one timestamp), and a row
recovered after a crash carries an earlier clock at a later sequence.

The host now records each item's place within the write that created it,
keeps it across revisions like the sequence, and sends it as an optional
field. Rows projected from the journal carry that position, and both the
desktop list and the outline order journal rows by it. Rows outside the
journal keep their rules: rank for the streaming and pending tail, outbox
sends after every journal row, and terminal-backed chats keep time then id.

* fix(native-chat): keep a refused send at its journal place, and keep list positions off worker reads

A send the host journalled before the provider refused it is shown from the outbox, and it
sorted after every journal row, so it dropped below whatever the agent wrote after it. It now
takes the journal position of the submission the host recorded.

Structured worker reads and their archives projected journal rows through the same
projection, so they returned the list-only journal position on every message. The worker
payload bound now drops it.

* test(native-chat): a failed restart's row draws below the message it failed

Since a message is accepted before its delivery starts the agent, a restart that fails is
journalled after the message, and the refused message keeps that journal place in the chat.
Both host paths now pin it through the chat's own projection: a start the host could not make,
and a restarted child that exits before proving its start.

Folds the journal reducer's batch item write onto fewer lines, which the merge of main pushed
past the file's line limit.
2026-09-27 23:44:36 -07:00
Neil 9179b93ebf ci: reduce repeated runner work and validate affected-test selection (#23540)
* ci: stage heavy checks and measure affected-test selection

* fix(ci): exercise the real Git boundary in unit selection planning

* Harden review cancellation and CI demand reporting
2026-09-27 23:25:20 -07:00
Neil fae0ae7a46 perf(ci): shard the anti-slop audit across processes instead of one JS runtime (#23543)
config/oxlint-anti-slop.json turns every native oxlint category off and runs its
rules through jsPlugins, so oxlint's threaded Rust engine does no work and the
pass is one JS runtime per process. Measured, it does not scale with --threads:
11.68s at 4 threads against 12.38s at 16. Parallelism has to come from more
processes, so the audit now splits its file set across them.

Locally, 11.72s single pass against 3.83s at 4 shards (3.1x) and 2.91s at 8.

Sharding is sound because every anti-slop rule is a single-file analysis; the
only mutable module state is a WeakMap keyed on each file's own Program node.
Verified by running both shapes with all 18 rules enabled: 392,398 findings from
one pass and from the shard union, identical as sorted multisets.
2026-09-27 23:15:19 -07:00
Neil 67581dd090 perf(ci): stop the orcad smoke idling 15s and the static job fetching mobile packages it skips (#23541)
The shutdown race in the orcad terminal smoke never cleared its losing timer, so
the process sat on a live 15s timer after PASS had already printed. Measured
locally: 21.4s -> 6.52s, with the round trip and the shutdown assertion intact.

Static analysis also asked for the mixed root+mobile pnpm store (537 MB, 8.6s to
restore) on every run, while installing mobile dependencies only when the diff
needs them. Most runs paid 216 MB for packages they never linked.
2026-09-27 22:51:59 -07:00
Neil 45f3512a33 feat(agents): add first-class DeepSeek Harness (dsh) support (#22468)
* feat(agents): add first-class DeepSeek Harness (dsh) support

Register DSH as a supervised Orca agent: catalog entry and detection for its
dsh-tui profile, status/question hooks through DeepSeek's own Claude-Code hook
bridge, composer-ready prompt delivery, session resume, headless Source Control
AI, and title identity that no longer collides with Gemini's.

* fix(dsh): reach Orca through DSH's credential scrub and stop reading its title as Gemini

DSH runs command hooks through its own shell executor, which drops every env var whose
name contains KEY, TOKEN, SECRET or PASSWORD — taking ORCA_PANE_KEY and
ORCA_AGENT_LAUNCH_TOKEN with it, so every hook exited without posting. Mirror both onto
scrub-safe aliases at spawn and restore them at the top of the DSH hook script.

Its title collided too: DSH rests on the same glyph Gemini works on, so a resting DSH
pane was relabelled Gemini CLI and reported working forever. Defer both the Gemini
classifier and the title status detector on DSH's whale, in the base module both copies
of that classifier read.

* test(mobile): repin the session-route closure for the DSH agent icon

* fix(dsh): address review — never splice user rows, cover remote panes, keep the diff off argv

- findManagedDshPatchRegion paired an orphan start marker with a later block's end, so a
  truncated write made install/remove delete the user's own rows. Pair each end with the
  nearest preceding start; regression test fails without the fix.
- The relay PTY env builder never applied the scrub-safe aliases, so remote DSH status
  silently never appeared even with the remote hook installed.
- Source Control AI sent the whole diff on argv; send it over stdin with DSH's '-' marker.
- dsh-tui/dst already chose the interactive profile, so a workspace folder named 'web' or
  'plugin' no longer marks a live agent pane non-interactive.
- Isolate USERPROFILE as well as HOME so a Windows run cannot edit the real home.
- Drop the duplicate README badge and revert an incidental doc reformat.

* refactor(dsh): share the managed-hooks reader and tighten the new modules

Reuse before reimplementing: readManagedDshHookEvents was a near-verbatim copy of Muse's,
with byte-identical private helpers. Both now call one readManagedHookEventsFromJson.

Also: one readTextOrAbsent instead of two spellings of the same read (dropping an
existsSync TOCTOU), one status() builder instead of four inline literals, rmSync(force)
instead of exists-then-unlink, and a redundant empty-string guard before JSON.parse.
The patch-file transforms lose their index juggling for a predicate plus a filter.

* fix(dsh): refuse a flow-style patch file, keep its mode, and stop the relay inheriting a pane

- applyManagedDshPatch matched only an exact `[]`, so `[] # keep empty` or a non-empty
  flow sequence got a block entry appended after it — invalid YAML that would leave DSH
  unable to load the user's own patch layer either. It now strips the token from an empty
  sequence (keeping a trailing comment) and returns null for a non-empty one; install
  reports that and changes nothing.
- The patch rewrite dropped an owner-only file to the umask default (CWE-732); pass
  preserveMode.
- The relay PTY env never dropped inherited pane identity the way the local and daemon
  builders do, so a spawn that specified none could inherit the relay's own and every
  agent's hook would report against that pane.

* fix(dsh): keep the flow-style refusal in every status read, and scope the mode test to POSIX

A refused patch file carries no managed region, so getStatus() fell through to a bare
not_installed with detail null — the actionable 'rewrite it as a block sequence' message
only ever reached the one-shot install() return. Export the predicate and check it first,
behind one shared message constant.

The owner-only mode assertion cannot hold on Windows, where chmod only toggles the
read-only attribute and mode & 0o777 reads 0o666 for any writable file.

* docs(readme): restore the DeepSeek Harness badge lost in the rebase

* test(mobile): repin the session-route closure to the measured 4221

Measured, not derived: 4220 without the DSH icon entry, 4221 with it. Two of the three
modules above main's 4218 pin are not this change's — they arrived with the mobile work
after #22570 and were never repinned; the changelog records that split explicitly.

* fix(dsh): settle tui-idle on the agent's own hook, so supervised workers see it ready

Reported by a tester on the adhoc build: `terminal wait --for tui-idle` ran to its 90s
timeout against an already-ready DSH composer, so a supervised worker never sees the agent
as ready.

Every existing tier reads the title, and DSH deliberately carries no title status: its rest
prefix is Gemini's working glyph, so the detector reports none. A fresh first-party `done`
is better evidence than any title anyway — it is the agent's own account of its own turn,
and normalizeDshEvent drops subagent events, so it is the lead's. Scoped to DSH: for agents
whose hooks report child turns, a mid-turn `done` is the #6011 class this file prevents.

* test(daemon): record the DSH transcript's true-colour I2 divergences

Adding the dsh-tui capture to __fixtures__ enrolled it in the serialize replay sweep, where
it reports 10 I2 divergences and failed the unlisted-transcript default of 0.

Every one is the same shape — visible-grid row=0, a 24-bit background the round trip does
not restore to default — which is DSH's whale intro painting whole rows of true colour.
Verified as an upstream limitation rather than a regression by replaying against the
previous build (build-serialize-addon-at-ref.mjs --ref origin/main): I1 and I3 both hold.

* fix(dsh): return the new tui-idle verdict from the first-party done lane

Main refactored isTuiIdleSatisfied into evaluateTuiIdle, which returns a verdict rather
than a boolean. The DSH lane still returned `true`; it is tier-1 positive evidence, so it
returns READY_STRONG like the title/body lane above it. Re-verified the regression test
still fails without the lane.

* test(relay): pin the scrub-safe pane-identity aliases on the relay spawn path

The relay builds a remote pane's env itself, so the alias mirroring there had no
test: removing the call left every suite green while remote DSH status silently
vanished. Both cases fail without it.

* docs(dsh): point the hook service at the integration reference

The reference doc had no inbound link from anywhere in the repo.
2026-09-27 22:44:18 -07:00