Commit Graph
355 Commits
Author SHA1 Message Date
Brennan Benson 21d4ae9448 feat(agents): pre-trust the folder wherever Orca starts an agent (#23744)
* feat(claude): pre-trust worktrees Orca creates

Claude Code asks "Do you trust this folder?" on first launch in any folder it
has not seen, which blocks unattended launches in worktrees Orca itself made.
Orca now records where a worktree's content came from when it creates it, and
before each Claude launch writes Claude's own folder-trust entry for that
worktree's root (never the main checkout) when the new setting is on and the
content is the user's repository. Forks, bare commits, folder workspaces and
external checkouts keep Claude's prompt. The write takes Claude's lock, never
creates or breaks the file, runs on the SSH host itself, and is revoked when
the worktree is removed or the setting is turned off.

Launches that already pass --dangerously-skip-permissions also skip the trust
prompt for that one process only, via CLAUDE_CODE_SANDBOXED=1 on the command.

* fix(claude): parse the relay trust request with a schema and ship its search keys

* fix(claude): never write a WSL guest's trust into the Windows config

A WSL worktree's Claude reads the guest's own config. Two paths still wrote
its trust into the Windows host's ~/.claude.json instead: the Claude auth prep's
fallback (runtime 'wsl' but the host config dir, when the WSL home cannot be
resolved), which wrote a Linux-path key the removal revoke can never delete;
and the Agent Teams leader, which passes no auth or distro and wrote a UNC key.
Require the guest's own config dir, and treat any WSL worktree path as guest-only.

* revert(claude): drop the skip-permissions trust shortcut

Pre-trust stays limited to worktrees Orca creates from the user's own
repository. The per-launch CLAUDE_CODE_SANDBOXED prefix skipped Claude's
trust question in every folder for launches carrying the skip-permissions
flag (Orca's default Claude args), including the user's own folders and fork
PR worktrees, and the setting could not turn it off. Remove the prefix, the
inherited-variable strip that existed only for it, and the Agent Teams
leader-to-teammate propagation; restore the tests that pinned the prefixed
launch string.

* fix(claude): revoke SSH trust in the config file the grant used

At spawn the relay resolves Claude's config from the launch env, which carries
a CLAUDE_CONFIG_DIR set in Orca's Claude default env. claudeTrust.converge had
only the relay's own process env, so removing the worktree or turning the
setting off revoked in the default file and the grant outlived the worktree.
Send the config-file keys with the request, as the local revoke already uses.

* i18n(settings): translate the Claude worktree trust setting

* fix(settings): say Claude trust applies when Orca starts Claude

The description said Claude skips its trust prompt in any worktree Orca
created. Trust is written only when Orca itself starts Claude there, so a
`claude` typed by hand in a fresh worktree still asks. Say that, and bring
the es/fr/ja/ko/zh translations in line with the new text.

* fix(worktrees): treat a base on an Orca-added fork remote as fork content

A worktree based on a named ref was always stamped as the repository's own
content, so picking the fork remote Orca adds for a pull request (or a local
branch tracking it) as the base made a fork's code eligible for Claude trust.
At create time, read the repo's `remote.<name>.orca-created` markers and each
branch's tracked remote in one `git config` call; a base on such a remote is
stamped as a fork's content, and a read failure is not vouched for. Remotes
the user added, such as `upstream`, stay first-party.

* perf(claude): revoke worktree trust once per config file, not per worktree

Turning "Trust worktrees Orca creates for Claude" off read and parsed the
whole Claude config once per Orca worktree on the main process. Group the
revocations by config file locally and by SSH connection, and make
claudeTrust.converge take a batch of requests.

* feat(settings): one agent-wide "trust the folder" setting in Settings > Agents

Replace the Claude-only worktree trust toggle with a single setting,
agentWorkspaceTrustEnabled (on by default; unreleased, so no migration).
The row says what it does for every agent: agents Orca starts skip their
"trust this folder?" prompt in that worktree or folder, turning it off
stops new trust while existing trust stays, and while it is off unattended
launches (orchestration workers, automations, the phone) stop at the
agent's trust question until someone answers.

Translations for es/fr/ja/ko/zh. Also restores the `awaitingUnnamed` chat
catalog keys an earlier merge of main dropped from this branch.

* feat(agent-trust): pre-trust the workspace for every preset agent at PTY spawn

Every Orca-started agent PTY passes through one of the two spawn builders
with its declared launchAgent, which survives setup-script wrapping. The
builders now call one hook that, for a fresh launch (never a reattach or
restored pane) with the setting on, applies the agent's trust preset to
the worktree, folder workspace or main checkout it starts in.

- One dispatcher, applyAgentWorkspaceTrust(preset, workspacePath, launch
  context), carries what a writer needs: the final spawn env, the Claude
  managed-account auth prep, the WSL distro and the SSH connection.
- Claude joins the presets on both the claude and claude-agent-teams
  entries. Its writer stays grant-only in claude-folder-trust-file.ts:
  the file Claude reads (CLAUDE_CONFIG_DIR / custom-OAuth suffix / legacy
  .config.json / a WSL guest's own file), Claude's <file>.lock never
  broken and taken only when a write is due, atomic temp+rename keeping
  mode and symlinks, never creating the file, NFC + realpath keys.
- SSH Claude launches forward the optional claudeFolderTrust spawn field
  so the relay grants with its own spawn env; old relays ignore it and
  Claude asks. Other presets keep the SFTP writer. A WSL launch never
  writes the Windows home: non-Claude presets skip it, Claude writes the
  guest's file or nothing.
- Codex keeps the 20 s deadline its shared config lane needs; every other
  preset gets 1.5 s. A miss means the agent asks; trust bookkeeping never
  fails or blocks a launch.

Removes the Claude-only machinery this replaces: the eligibility/host/
lifecycle/spawn modules, the persisted creation content-origin field and
its classification, revoke-on-removal, the setting-off sweep, the
claudeTrust.converge relay method and the Agent Teams leader special case
(the leader pane now spawns through the hook with the claude preset).
The agent config types move to tui-agent-config-types.ts so the config
table stays under the line budget.

* refactor(agent-trust): delete the pre-spawn trust writes the spawn hook replaces

The spawn hook is now the only owner of agent folder trust, so remove
every other writer:

- the agentTrust:markTrusted IPC channel, its preload bridge and types,
  and all renderer callers (agent-trust-preflight and its callers in the
  background session, work-item direct launch, session continuation,
  worktree creation, folder workspace composer and session fork);
- the main pre-spawn sites: the createdWithAgent preflight in
  worktree-remote.ts, markLocalWorktreeTrusted/markRemoteWorktreeTrusted
  and the runtime's markWorkspaceTrustedForAgent family with the
  markTrusted ports of the runtime create flows;
- Codex's own launch-prep and resume-prep trust writes.

Each of those launches reaches a spawn builder with launchAgent set, so
the hook covers it. This also fixes a live gap: the worktree-remote.ts
copy of the preset switch omitted Antigravity, so an agy agent started
from a desktop worktree create still asked; the single dispatcher covers
it. Trust is also written on the host the PTY actually spawns on, which
removes the #11163 class of writing the wrong host's config.

* test(agent-trust): type the spawn-builder trust fixtures and prove the spawn waits for trust

The builder test passed untyped args (a string launchAgent) and cast its deps,
which failed tc:node. It now builds both spawn states from a fully typed deps
fixture and a typed restored pane, with no casts.

Adds a case that holds the trust write pending and checks the builder does not
finish until it settles, the ordering the deleted renderer and launch-prep
tests used to cover.

* fix(agent-trust): give SSH trust writes the 20 s deadline again

The dispatcher gave every non-Codex preset a 1.5 s budget, including the SSH
writers for Cursor, Copilot and Qoder, which make several round trips over the
link. Before this PR those writes had 20 s (desktop) or no limit (runtime), so
on a slow link an unattended SSH worker would now stop at the agent's trust
question. SSH writes get the 20 s deadline back; local non-Codex writers keep
the short budget, and Codex keeps 20 s.

The relay's Claude grant keeps the short budget: it writes the relay host's own
disk and does not cross the link.

* fix(agent-trust): never pre-trust a home folder or a filesystem root

A folder workspace can be the user's home folder or a disk root. Claude and
Copilot let a trusted folder cover every folder under it, so pre-trusting one
of those would silently trust everything on the machine for those agents.

One check, isHomeOrFilesystemRoot, now refuses them for every preset: the
dispatcher checks roots and this machine's homes (including the spawn env's
HOME and a cached WSL guest home), the SSH writer checks the remote home it
already resolves, and the relay checks its own home. The agent then asks, as
it would without Orca.

* refactor(agent-trust): drop the Codex launch plumbing that only carried trust

The spawn hook replaced the trust writes in Codex launch prep and resume prep,
which left the fields that fed them unread: CodexHomeLaunchContext.workspacePath
and .launchAgent, the resume prep's workspacePath, and the structured Codex
launch input's workspacePath (plus the extra target lookup that produced it).
Remove them and their plumbing; unavailableManagedHomePath stays.

Also removes test stubs of runtime trust methods this PR deleted, whose
not-called assertions could no longer fail, and two comments that still
described the old trust preflight.

* chore(reliability-gates): point the trust gate at the spawn-time trust tests

The agent-session trust gate still listed three test files this PR deleted
(the renderer preflight, the Codex launch-prep deadline and the e2e trust
completion suites), so check:reliability-gates, which runs in PR CI and in
pnpm lint, failed on missing files. Its invariant also described the deleted
IPC handler and pre-spawn writers.

The gate now covers what replaced them: the spawn builders holding the spawn
until trust settles, the fresh-launch and setting gates, the per-preset
deadlines, and the home and root refusal.

* fix(agent-trust): skip the relay Claude grant for a WSL shell

On a Windows SSH host whose pane shell is wsl.exe, Claude runs inside the WSL
guest and reads the guest's config. The relay still granted trust in the
Windows host's own .claude.json, writing the Windows home for a WSL launch,
which the local path never does. The relay now skips the grant there, so that
Claude asks, as a local WSL launch does when Orca cannot reach the guest file.

* perf(agent-trust): only agent launches wait on the trust hook

Both spawn builders awaited the trust hook on every spawn, including plain
shells, reattaches and agents without a preset. Awaiting even a resolved
promise adds microtask ticks ahead of the pane-spawn reservation check, and
this handler already keeps non-Codex spawns off an await because an extra tick
reorders those reservation races.

The hook now returns null when there is nothing to write, and the builders
await only a real trust write.

* test(agent-trust): keep the home and root cases off any real Claude config

The home and root cases ran the real Claude writer with the test process's
env, so a regression in the guard would have written trust for the home
folder and / into whatever Claude config that env named. They now point
CLAUDE_CONFIG_DIR at a folder that does not exist, and the writer never
creates a config.

* test(runtime): drop needless casts from the launch-host test

The renamed launch-host test kept three `as never` casts on launch options
that already match launchAgentTerminal's parameter type. The changed-lines
casting gate reads the renamed file as new and failed on them.

* test(agent-trust): type the Claude grant mock with the real writer's signature

The mock took an unknown target, so installing the real writer as its
implementation would not typecheck under strict function types.

* fix(codex): drop the launch context the trust move left unread in local spawn env

* fix(agent-trust): queue Claude grants per config file so a launch burst keeps them all

Concurrent grants in one process retried Claude's file lock in lockstep, so each
retry round admitted about one winner. Starting 12 Claude agents at once left 6
of them at the trust question with nothing logged. Grants for one config file now
queue in-process; only Claude's own writes contend for the lock. The relay shares
the writer, so bursts of SSH launches are covered too.

* fix(agent-trust): never pre-trust a folder above a home either

The guard refused only an exact home or a filesystem root. A folder workspace at
/Users, /home or C:\Users was still pre-trusted, and Claude walks up parent folders
for a non-git folder, so every non-git folder in the user's home became trusted.
The guard now also refuses any folder that contains a home, on every host, and is
renamed to say what it decides.

* perf(agent-trust): skip the SSH round trips for Antigravity, which has no remote writer

Every Antigravity launch over SSH now reaches the remote trust writer, which
resolved the remote home and realpath'd the workspace over the link before
writing nothing (the known remote gap). That delayed each launch by two SSH round
trips, and up to the 20 s deadline on a stalled link. It now returns first.

* fix(settings): keep the hidden folder trust row out of web-client settings search

The paired web client hides the host-only "Trust the folder" row, but settings
search still listed it, so searching "trust" opened the Agents pane with no
matching row. Its search entry is now filtered the same way as Agent Awake.

* fix(agent-trust): a failing breadth guard skips trust instead of failing the spawn

* docs(qoder): New Tab now pre-trusts through the agent-wide spawn hook

* fix(agent-trust): never pre-trust a home reached through a symlink

Every trust writer stores the workspace's resolved path, but the breadth
guard compared only the path as given. A folder workspace that is a
symlink to the home folder (or a real home picked while HOME names a
symlinked one, as on distros that link /home to /var/home) passed the
guard, and Claude, Copilot and Cursor then trusted the home itself.

The local dispatcher and the relay now compare given and resolved forms
of both the workspace and each home. The SSH writer resolves the remote
home alongside the workspace, in parallel, so it adds no round trip.
Local non-Claude WSL launches still skip before any filesystem call.

* fix(relay): a failing breadth guard skips Claude trust instead of failing the SSH spawn

The relay ran its home/root guard and homedir() before its catch, so a
throw there rejected the relay's terminal spawn. Same fix as the main
dispatcher's: the whole grant, guard included, is best-effort.

* fix(agent-trust): guard the path each writer stores, not the path Orca was asked to trust

The breadth guard checked the launch's workspace while each writer stored a
transformed path, so every new transformation opened a hole. Codex stores a
linked worktree's main checkout: with a git repo rooted at the home, a Codex
launch in one of its worktrees wrote trust for the whole home.

One relay-safe host module now computes the stored path (Codex's main-checkout
hop, then given and resolved forms of it and of each home), refuses a root, a
home or a folder above one, and only then writes. Main uses it for local and
WSL launches and the relay for Claude. An unknown home writes nothing, and the
WSL home cache is keyed case-insensitively by distro.

* fix(ssh): the relay writes every preset's trust on the SSH host itself

Codex, Cursor, Copilot and Qoder trust over SSH was written from the desktop
over SFTP: four or five round trips per launch, so it needed a 20 s deadline
that outlasted the 8 s draft paste, the 10 s phone wait and the 15 s web-client
create. It also skipped Claude's atomic rename, ignored CODEX_HOME, and stored
the worktree where local Codex stores the main checkout.

The unreleased `claudeFolderTrust` spawn field becomes `agentWorkspaceTrust`,
sent for every preset. The relay derives the preset from the `launchAgent` it
already receives and runs the same host writer main uses, on its own disk,
within 1.5 s and with no extra round trip. Antigravity still returns early on
the relay (its writer is unverified on SSH hosts), a WSL shell still skips,
and any throw means the agent asks.

Deleted: the SFTP preset writer, the remote Qoder writer, the SSH deadline
clause and the desktop-side SSH root pre-check.

* test(e2e): keep CLAUDE_CONFIG_DIR out of isolated Electron launches

The spawn hook now writes Claude folder trust into the config
CLAUDE_CONFIG_DIR names, so an e2e run started from a shell that sets it
could add trust entries to the developer's real Claude config. Also drops
a stale comment that still named Codex launch prep as the trust owner.

* fix(agent-trust): guard Claude's resolve() form of the workspace too

Claude's writer stores both resolve(path) and the realpath. The breadth
guard compared only the given path and its realpath, so a workspace
path that does not exist and climbs back with `..` (for example
<home>/missing/..) passed the guard while Claude stored a key for the
home itself. The guard now also compares resolve(path), so it sees
every form a writer stores.

* test(relay): pty.spawn writes agent trust before the agent's process starts

Nothing exercised the relay handler's call into the trust writer, so
removing that call, or no longer awaiting it, left every suite green
while SSH launches silently stopped pre-trusting. The new case holds the
trust call pending and checks the spawn waits for it, and that the call
gets the request, the declared agent and the final spawn env.

The reliability gate lists the new suite and records the resolve() form
the breadth guard now compares.

* fix(agent-trust): refuse a home only for agents that inherit trust from it

The home and root refusal applied to every preset, so Codex, Cursor and
Antigravity started asking in a home folder workspace, where they did not
before. Only Claude, Copilot and Qoder let trust on a folder cover the
folders below it; Codex matches its start folder or that folder's repo
root, Antigravity the exact folder, and Cursor itself never inherits from
a home, a folder above one or a shallow path. The refusal now reads a
per-preset table in the host module, so the local and relay writers share
the rule.

* fix(agent-trust): trust Codex at the folder it starts in, as before

Before this PR, Codex launch prep trusted the spawn's start folder. The
spawn hook trusted only the workspace root and skipped terminals with no
workspace, so Codex began asking in a floating terminal and in a subfolder
of a non-git folder workspace: its lookup checks the start folder, then
that folder's repo root, and a plain folder above it is neither. The hook
now passes the resolved start folder for presets marked as keyed by it
(Codex only), falling back to the workspace root.

* fix(agent-trust): pre-trust a structured Codex chat's folder, as before

Before this PR, creating a structured (native) Codex chat pre-wrote Codex
trust for its folder through launch preparation. The PR removed that write
and routed trust through the PTY spawn builders, which a structured chat
never passes. Codex's app-server trusts the folder itself only when the
chat's permissions can write it, so a read-only chat started running
untrusted and ignored the project's .codex config. Creating the chat now
calls the same dispatcher, behind the same setting, before launch prep.

* fix(settings): plainer folder trust setting text
2026-09-29 16:03:28 -07:00
OrcaWinandm4air e1362ada4c fix(terminal): stop inline-image decoders exhausting the renderer's wasm memory budget (#23499)
V8 reserves an 8 GiB guard region per wasm memory inside its 1 TiB sandbox,
so an Electron renderer can hold only ~124 live wasm memories regardless of
free RAM. @xterm/addon-image instantiated a SIXEL decoder per terminal at
activation (and kept IIP decoders after the first image), so ~120+ terminals
exhausted the budget: new panes raised 'WebAssembly.instantiate(): Out of
memory' rejections, and the next Kitty/IIP image threw 'WebAssembly.Memory():
could not allocate memory' out of the parser, permanently wedging that
terminal's write queue.

The addon-image source patch now borrows SIXEL decoders from a shared pool
only while a sequence is open (color registers stay on the terminal), drops
IIP decoders after each image, and turns a failed decoder allocation into a
dropped image instead of a parser throw. Bundles regenerated with
regenerate-xterm-patches.mjs --write.

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-29 01:20:37 -07:00
Neil f8f656ca19 perf(ci): spend fewer concurrency slots per pull request (#23810)
A concurrency slot is charged per job, not per core, and the account's cap is
the scarce resource: standard runner minutes are free and unlimited on a public
repository. Two paths spent slots that bought nothing.

The unit matrix ran eight fixed shards averaging 6.5 minutes each, 3384
job-slots a day and 68% of all slot demand, while the arm pool queued 10.5
minutes at p95 — the queue was the oversharding. Five shards run the same work
in ~10.5 minutes each for three fewer slots per run.

Bun profile persistence escalated to all six platforms on `config/`,
`resources/` and `.github/` wholesale, which took 36.5% of the last 1100
commits through the full matrix where a platform-flavoured predicate takes 19%.
A pull request now qualifies one platform unless the change is platform-
flavoured, and the push to main re-qualifies all six, so an unescalated miss
surfaces minutes after merge rather than at the next cron. Missing changed-file
evidence and an unavailable dependency graph still fail closed to all six.
2026-09-29 00:13:33 -07:00
Neil ccdb324b63 Add CodeBuddy as a built-in coding agent (#23740)
* feat(agents): integrate CodeBuddy launch, status and session history

* docs: record CodeBuddy lifecycle verification

* fix(codebuddy): backfill scoped history and negotiate remote resume

* test(cli): include CodeBuddy in known search agents
2026-09-28 18:11:25 -07:00
8b410b4893 feat: add first-class Qoder CLI support (#23581)
feat: add first-class Qoder CLI support

Integrate Qoder launch, identity, canonical hook status, trust and resume.
Verify with captured Qoder 1.1.64 transcripts and hidden Electron sidebar checks.

Builds on and cross-reviews #7502, #8611, #9655, #12910, #13311 and #15291.

Co-authored-by: dalveytech-vincent <vincent@dalveytech.com>
Co-authored-by: Eridanus117 <45489268+Eridanus117@users.noreply.github.com>
Co-authored-by: xingqingzzp-gif <xingqingzzp-gif@users.noreply.github.com>
Co-authored-by: jyang2004 <jyang2004@users.noreply.github.com>
Co-authored-by: yunqian <yunqian@alibaba-inc.com>
Co-authored-by: huzhening.hzn <huzhening.hzn@alibaba-inc.com>
2026-09-28 02:59:50 -07:00
Neil 9179b93ebf ci: reduce repeated runner work and validate affected-test selection (#23540)
* ci: stage heavy checks and measure affected-test selection

* fix(ci): exercise the real Git boundary in unit selection planning

* Harden review cancellation and CI demand reporting
2026-09-27 23:25:20 -07:00
Brennan BensonandClaude 85067494a1 fix(native-chat): a request that failed reads as failed (#22944)
* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* fix(native-chat): name a chat write by its target, not the owner generation

A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.

Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.

Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.

* fix(native-chat): every journal append reaches the chats that are open

A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.

A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.

* test(native-chat): an epoch replacement reaches the open chat

* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

* perf(native-chat): a publish behind a delivered commit reads nothing

Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.

* test(native-chat): state why the teardown test's fake journal is safe to cast

* docs(native-chat): say mutation admission checks only the writer lease

* docs(native-chat): drop the send rebase from comments that still described it

* fix(native-chat): a message is accepted, then delivered

A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".

A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.

Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.

A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.

Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.

* fix(native-chat): settle queued messages only for the child that ended

A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.

A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.

The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.

* fix(native-chat): an adoption that fails to import keeps the conversation open

The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.

* perf(native-chat): the recovering open reads the journal once

Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.

* fix(native-chat): an attach that fails after indexing its child leaves no child behind

A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.

* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer

The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.

A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.

* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down

The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.

* fix(native-chat): a message rejected while its chat was closed reads as not sent

A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.

* test(orchestration): name why the readiness settlement fakes are cast

* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent

* docs(native-chat): drop the fence from the admission the send effects run behind

* docs(native-chat): give the fence move on release the reason that still holds

* docs(native-chat): stop citing a write fence check in launch and mailbox comments

Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.

* refactor(native-chat): the provider child is its own record

A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.

- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
  patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
  own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
  the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
  dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
  is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
  the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.

* fix(native-chat): the delivery loop alone settles a message its start or child failed

A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.

- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
  reads how it ended: a Stop continues; anything else writes one failure row and rejects every
  queued message with the same words, then stops. A child still starting whose start the adapter
  says did not land fails the same way. The exit, eviction and the settlement retry only settle
  the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
  nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
  failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
  closed, with or without a child, and a start the loop already has in flight is waited for so the
  child it produces is stopped rather than left behind.

* refactor(native-chat): a stopped child ends on the one reading of its stop

The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.

* feat(native-chat): the host says it accepts a send before any agent has it

The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.

* refactor(native-chat): an attach never opens a journal of its own

The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.

* fix(native-chat): a moved fence resends nothing on a host that accepts first

The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.

The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.

* refactor(native-chat): a child's end says whether the user or the host stopped it

The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.

* fix(native-chat): a chat whose only work is a queued message is not offered for resume

A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.

* test(native-chat): type the queued-message fixtures in the resume-offer tests

* fix(native-chat): a start that dies while a message waits on it is that message's failed start

Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.

* fix(native-chat): a request that failed reads as failed

A structured chat whose only message the agent's start refused read as a
green finish, and a cancelled structured turn did too: the host published a
verdict only for turn records, and structured rows carried no `interrupted`.

The host projection now reads the session's latest request: its turn's
outcome, or `failure` for a send the agent or its start refused. A send
that was withdrawn, or left undelivered by a restart or a close, fails
nobody and makes nothing listable. The ingest publishes `interrupted` as the
hook lanes do, and every reader decodes the verdict through one accessor, so
a failure reads Failed on the dot, the rollups, history and `worktree ps`,
behaves like a cancellation in every clean-finish policy, and notifies as
"failed".

* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now

* test(native-chat): a verdict change republishes the mobile status projection

* refactor(native-chat): the store's retention trigger keeps its flag compare

A verdict change always moves the completion clock the same check already
reads, so a second verdict compare there caught nothing new.

* test(native-chat): a user message the provider journaled keeps its session listed

* test(native-chat): pin what a failed start settles, and what a resume offer names

A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.

* test(native-chat): the failed-start pins fail on what the message became, not on a timeout

* fix(native-chat): a late provider-session update keeps a failed recovery record failed

A provider-session heartbeat that rewrites a completed recovery record kept
its interrupted flag but dropped the outcome it was copied with, so a live
failed checkpoint read as a clean finish until the next status write.

* test(orchestration): the preamble's host stub is typed, not cast

The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.

* test(native-chat): the terminal-bell check asserts the renamed verdict field

The bell notification test still checked for agentInterrupted, which no
longer exists, so it could not catch a verdict leaking into a bell dispatch.

* fix(native-chat): a failed turn ranks like a completion for attention

Attention readers (completion time, Smart Sort, sticky retention, Cmd+J
Recent) now demote only a turn the user stopped. A failure is news the
user has not seen, so it keeps its completion time, ranks in the Done
class, stays retained after its pane goes away, and a retained failure
reads failed in the worktree rollup instead of done. Clean-finish
policy (hibernation, pane ownership, the value moment) still treats a
failure like a stop.

The retention trigger compares verdicts again: success -> failure no
longer moves the completion clock.

* fix(native-chat): a failed main agent reads failed while its subagents still work

The verdict is now read from the main agent's own state, not the folded
row: a main agent that is done and failed has a verdict even while its
subagents keep the row working. Without mainAgent (history, worktree ps,
older hosts) the old combined-done rule stands.

Display marks the verdict through agentVerdictDisplayMark: a failure
outranks every combined state on the agent's dot, label, tab badge,
dashboard and activity rows; a stop marks only a done row, so a
successful or stopped main agent with live subagents still reads
working. Subagent rows keep their own state. The worktree card, terminal
tab and Cmd+J rollups share one pane fold and rank a pending question,
then failed, then working, monitoring, interrupted and done.

worktree ps publishes the main agent's outcome on a working row, and the
mobile mirror reads it. The store's change check, the paired-client
mirror's equality and its epoch now see a verdict change on a working
row, which otherwise moves no state or clock and left the worktree card
reading working. Clean-finish policy is unchanged: a working row is never
hibernated and has no completion time.

* docs(native-chat): the worktree ps outcome comment no longer claims old hosts send it

The field is new: an old host sends no outcome at all, so a reader falls
back to interrupted. The removed clause said old hosts send it on done
rows, which never shipped.

* docs(native-chat): the status-store listing rule names provider-journaled user messages

* fix(native-chat): a refused send notifies failed through the completion feed

The host's completion feed followed only the newest turn, so a send the
agent or its start refused, which creates no turn, read Failed on its row
but sent no notification. The feed now follows the session's latest
request, read from the projection the status feed already makes for the
commit: a turn keeps its id, a refused send is named by its journal item
key. It announces only while the session is idle, as the row reports a
verdict, so queued sends refused one commit at a time notify once, and a
withdrawn send falls back to a request already announced.

* fix(native-chat): every copy of a row carries the main agent's own status

History entries, sleep records and `worktree ps` rows carried a flattened
top-level `outcome`, copied under different gates and without the main agent's
clock. They now carry `mainAgent` (state, outcome, stateStartedAt), the type
the live row already persists and sends, and every copy site takes it with
`interrupted` through one function, `agentVerdictFields`.

- The accessor reads `mainAgent` then the legacy flag; the mobile mirror
  matches it line for line.
- Sleep records admit `mainAgent` with `normalizeMainAgentStatusField`, so a
  malformed value drops the field, never the record.
- Mobile dates a main agent that failed under live subagents by its own clock,
  as desktop does, and its row equality compares `mainAgent`.
- The activity feed reads a history entry's own `mainAgent` instead of
  rebuilding one; the sync key and history equality compare it.

* test(native-chat): pin the worktree ps verdict across host and phone versions

Pairs the real v1.4.212 host and phone row reader with this build: an old phone
reads a new host's rows by `interrupted`, a new phone reads an old host's rows
(no `mainAgent`) the same way, and a new phone reads a failure under live
subagents as Failed, dated by `mainAgent.stateStartedAt`. The release checkout
now carries the phone's self-contained row reader, and the lane runs when the
`worktree ps` row producers change.

* test(mobile): name the parity table's row for its role

* fix(native-chat): a request that settles while the user is asked something notifies once

The completion edge waited for an idle session, and a pending prompt (including a
subagent's approval) is not idle. Structured chat has no other attention producer,
so a main turn that finished while a subagent waited on the user sent nothing
until the prompt was answered.

The edge now waits only on owed work (a running turn or an unanswered send), which
the projection reports even beneath a pending prompt. A request that settles with
a prompt pending announces once; the renderer words it "needs input" from the
host status mirror's `attention`, and answering the prompt keeps the same request
identity, so it does not announce again. The wire shape is unchanged.

* fix(native-chat): the completion says when the user is being asked

A request that settles while a prompt waits on the user was worded "needs input"
from the renderer's status-feed mirror. Remote clients receive the status and
completion streams over separate sockets, so they can arrive in either order and
the wording could be wrong both ways.

The host already knows at emit time, so the completion now carries an optional
`awaitingUser: true` in that case and omits it otherwise. The renderer words the
notification from that field alone and no longer reads the status mirror. Old
clients ignore the field and word by outcome; old hosts never send it.

* fix(worktree-status): a departed agent's failure yields to live work on the worktree card

A retained failed agent has no expiry, so ranking it with a live failure pinned the card to Failed over other panes' live work. It now ranks below working, monitoring and permission, and above every finished outcome.

* docs(agent-status): a departed agent's failure ranks below live work on the worktree card

* fix(native-chat): a view never restarts a chat whose last start failed

A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.

* test(native-chat): start the child the loop waits on with an attach, not a second view

A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.

* fix(native-chat): settle a gone generation's turn wherever a conversation opens

A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.

* test(native-chat): prove the next child's start settles the turn an earlier child left

The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.

* test(native-chat): count a failed start's rows by row, not by text

Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.

* test(cross-version): load the phone row readers without mobile's toolchain

Vite transforms a file against its nearest tsconfig, and mobile/tsconfig.json
extends expo/tsconfig.base.json, which the root-only cross-version lane never
installs. The worktree ps verdict suite imported the current phone row reader
from mobile/ directly, so CI failed with TSConfckParseError before any test ran.

The harness now imports a copy of the working-tree reader placed under the
checkout cache, where the root tsconfig applies, as it already does for the
release checkout's copy. Both readers are still the real files.

* test(cross-version): keep the checkout path-guard message and justify the copy import's cast

* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget

* test(native-chat): pin the open's and the send's start and row counts, however the view binds

Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.

* fix(native-chat): settle a gone generation's turn at every open but an acquisition's

The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.

* test(native-chat): hold the create's start open until the views bind

The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.

* refactor(native-chat): drop the composer's second error formatter

After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.

* test(native-chat): pin the reason on a message rejected while its chat was closed

The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.

* fix(native-chat): a send the provider never received after a restart has no verdict

Restart reconciliation rejects a crash-stranded send that is absent from a
trustworthy provider history with reason 'not_delivered'. Nobody failed that
send, but the verdict allowlist did not name it, so after a crash the chat
read Failed, was listed, and could notify "failed". Give the reason a shared
constant (persisted value unchanged), add it to the no-verdict set, and treat
it as an internal marker so the Retry row no longer shows the raw string.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-27 22:23:49 -07:00
OrcaWinandm4air 27b823f934 ci: compile the E2E CLI once for all consumers (#23384)
* ci: share compiled CLI output across E2E consumers

* ci: preserve CLI setup and old-ref fallback for shared artifacts

* docs: record shared E2E CLI benchmark evidence

* docs: include final CLI reuse timing range

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-27 02:02:29 -07:00
OrcaWinandm4air c15f082031 ci: build independent Electron targets together for E2E (#23378)
* ci: reuse parallel Electron targets for E2E builds and guard cache action setup

* test: recognize the top-level cache repository preload

* docs: record E2E build timings and exact output parity

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-27 01:17:27 -07:00
Neil 17690e6b9a style: settle oxfmt 0.70 drift and stop formatting vendored licences (#23377)
The oxfmt 0.65 -> 0.70 bump landed without a repo-wide reformat, so 36 files
already in the tree no longer matched what the new version emits. Anyone running
`pnpm format` picked all of them up alongside their own change.

Also excludes `resources/licenses/**`: `oxfmt --write .` was rewriting the
vendored PCRE2 licence, turning its `*` redistribution bullets into `-`. Third
party licence text has to be reproduced verbatim, so formatting must not touch it.
2026-09-27 01:14:53 -07:00
OrcaWinandm4air 47cebbf5d2 ci: use ARM unit runners, overlap web builds, and reuse verifier fixtures (#23376)
* test: reuse isolated mobile bundle fixtures for verifier checks

* ci: run PR unit shards on ARM and overlap independent web builds

* docs: record controlled CI overlap and runner measurements

* test: observe WebRTC packets with the host clock

* ci: isolate Windows installer CIM probe from native test load

* docs: record native probe scheduling validation

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-27 01:06:54 -07:00
Neil bc78acc43e fix(editor): detect all bundled Monaco language associations (#23371) 2026-09-27 00:32:55 -07:00
OrcaWinandm4air 9f5a8a5b8a Reuse mobile recording compilation and refresh desktop CI timings (#23343)
* ci: reuse recording compilation, split families, and refresh shard timings

* Keep recording suite intact after hosted performance comparison

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 22:44:15 -07:00
OrcaWinandm4air 4b6fe95943 fix(windows): preserve relocated terminals and native process scans (#22872)
* fix(windows): ship the process-table addon to the relocated daemon host

The Windows terminal daemon runs from a copy of the app under
%LOCALAPPDATA%\Orca\daemon-host\<version>. That copy took node-pty but not
@vscode/windows-process-tree, so the daemon's bare require of the addon found
nothing and every process-table read (foreground tracking, descendant sweeps)
fell back to a powershell.exe Get-CimInstance scan (#16905).

- Copy the addon's runtime files (package.json, lib/, the .node binary) into
  the host; the ~25MB of gyp intermediates beside them are filtered out.
- Treat a host missing those files as unmaterialized, so hosts built before
  this are rebuilt, and skip relocation if the install itself lacks them.
- Log the daemon's native/CIM capability at startup and warn once when the
  process table falls back to CIM.

Revives #19525 on current main.

* test(windows): locate update-survival loss before relaunch

* test(windows): preserve daemon tree before update-survival proof

* test(windows): distinguish Electron exit from launcher close timeout

* test(windows): verify process exit when inherited pipes delay close

* test(windows): trace installer process checks in isolated survival runs

* fix(windows): probe process-query capability before installer sweep

* fix(windows): match installer probe and process-check profile behavior

* fix(windows): use NSIS separators for the process-check include

* test(windows): dismiss session-search overlay in survival harness

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 14:31:07 -07:00
OrcaWinandm4air d17a17684b Reduce redundant CI runs, pnpm uploads, and fixture startups (#23145)
* Reduce redundant CI runs, store uploads, and fixture processes

* Avoid repeating draft-independent mobile checks on readiness

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 01:17:51 -07:00
OrcaWinandm4air 9b30c7f60a ci: verify mobile disposal and balance unit-test costs (#23114)
* ci: verify mobile disposal and reduce unit scheduling costs

* docs(ci): clarify timing assignment validation

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 00:12:06 -07:00
OrcaWinandm4air d9bc75752c fix(terminal): a remounted new SSH tab keeps the shell its old pane was still spawning (#22578)
* fix(terminal): a pane disposed mid-spawn no longer kills its successor's shell

A new terminal tab whose pane remounts while its first pty:spawn is in flight
is handed the SAME PTY by main's pane-spawn reservation. The disposed first
transport then killed that PTY as an orphan, so the tab closed on pty-exit
(focus fell back to tab 1) or stayed bound to a dead shell. Reported on SSH
worktrees (scan 22).

A transport destroyed mid-spawn now asks the pane surface first and keeps the
PTY while the tab exists, the worktree is not being deleted, and any layout
still names the leaf (leak-over-kill). A live transport refusing the id via
admitPtyId still kills unconditionally (#11003).

Adds rate-limited, id-hashed crash breadcrumbs for the next report:
terminal_fresh_spawn_retired (killed vs retained), terminal_tab_pty_exit
(host kind, ms since spawn, synthetic), terminal_active_tab_auto_move
(active-terminal repair, createTab orphan sweep). The two duplicated tab
pty-exit handlers now share handleTerminalTabPtyExit, and the FNV id hash
used by two crumbs moves to crash-breadcrumb-id-hash.ts.

Ports and supersedes #19386 (disposed-spawn-retention and its unit tests,
credit to its author); its Docker SSH e2e specs are not included.

* fix(terminal): scope spawn retention to its host and keep the PR focused

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-25 22:59:27 -07:00
OrcaWin 6fc3cdcad6 Bundle Bun for headless Orca and profile persistence (#22635)
Bundle a pinned, verified Bun runtime for headless Orca so existing Node launch commands can hand off before opening a profile. Keep desktop execution on Electron.

Add the Bun SQLite adapter and terminal backend, bounded shutdown, process inspection and cross-platform artifact qualification. Keep future managed SSH deployment separate from current production launch paths.
2026-09-25 22:49:06 -07:00
OrcaWinandm4air 38bcdf76ac perf(ci): reduce queue pressure without paid runners (#23053)
* ci: measure complete unit file costs for shard balancing

* perf(ci): reduce repeated PR setup and capture complete shard timings

* perf(ci): seed reusable main-branch native and typecheck caches

* fix(ci): stop superseded unit workflows from resisting cancellation

* perf(ci): reuse bundle fixtures and share the baseline Git build

* ci: record hosted gains and refresh main hook parity

* ci: retain default workers after performance-budget regression

* docs: record hosted mobile timing flake

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-25 22:48:28 -07:00
OrcaWinandOrca Worker 58d1ff3b6a Provide the Orca CLI automatically in managed WSL terminals (#22761)
* Provide the Orca CLI automatically in managed WSL terminals

* Simplify managed WSL CLI provisioning

Never block a shell on CLI availability, keep the shared WSL login-shell
builder unchanged, provision from PTY env assembly only, drop the error
variable and command probing, and reuse the existing WSLENV helper.

* Scope the managed WSL CLI to WSL terminals

Provision only for WSL panes and publish the directory through
addOrcaWslInteropEnv, so daemon terminals keep inherited WSLENV entries and
non-WSL builds never see the variable. Write the bridge with a UTF-8 BOM so
Windows PowerShell 5.1 keeps non-ASCII user-data paths, give the dev bridge the
dev launcher's app-launch env, and drop the unused skill-setup wiring and
runtime capability.

* Tighten the managed WSL CLI bridge and setup

Launch the bridge child exactly like the registered bridge (no hidden
window or output relay; verified through WSL with Node and Electron), give
the dev bridge the dev launcher's NODE_OPTIONS stash, clear the guest-only
directory before starting Windows processes, collapse setup into one
function, warn once, and guard WSL env routing with tests.

* Harden managed WSL CLI quoting and inheritance

PowerShell also ends single-quoted strings at typographic quotes, so a
user-data path such as O'Brien with a curly apostrophe broke the managed
bridge. Fix the shared quotePowerShellLiteral and reuse it. Drop an inherited
ORCA_WSL_CLI_DIR on the daemon path, remove the unreachable PATH dedupe, and
cover failed setup with a stale caller value.

* Cover the managed WSL CLI in zsh and on POSIX CI

Add a live zsh case that reaches a real prompt, a POSIX test that runs the
PATH restore snippet in bash and zsh under set -u, and a null result for
unwritable user data. Say what a failed write actually costs, and document
per-spawn write logging and older-daemon behaviour.

* Keep system bashrc out of the PATH restore test

CI runners make bash -c read /etc/bash.bashrc, which fails under set -u.

---------

Co-authored-by: Orca Worker <orca-worker@localhost>
2026-09-25 21:55:21 -07:00
Brennan BensonandClaude 16784c1a67 fix(native-chat): name a chat write by its target, not the owner generation (#22812)
* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* fix(native-chat): name a chat write by its target, not the owner generation

A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.

Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.

Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.

* docs(native-chat): say mutation admission checks only the writer lease

* docs(native-chat): drop the send rebase from comments that still described it

* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent

* docs(native-chat): drop the fence from the admission the send effects run behind

* docs(native-chat): give the fence move on release the reason that still holds

* docs(native-chat): stop citing a write fence check in launch and mailbox comments

Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-25 18:31:29 -07:00
Brennan Benson a0e24905f6 fix(agent-status): a cancel never hides live work (#22476)
* fix(agent-status): a cancel never hides live work

After the user cancels a turn, a background shell, scheduled check or
subagent that is still running keeps reading as it truly is in both
lanes. The fold no longer takes a verdict input; the cancellation
survives only as lead.outcome, restated as the row's interrupted flag on
a settled row for readers that predate lead.

* fix(agent-status): keep a cancel's verdict and clock on every settle path

A Grok turn cancelled while a task ran now reads monitoring, and the
idle_prompt backstop that later settles it restated done without the
row's `interrupted` flag, so notification readers announced the
cancelled turn as a clean finish. Derive `interrupted` from the main
agent's outcome, as the Claude builder already does.

The inferred Claude cancel now folds through the host's local main
agent record, which a relayed pane never refreshes, so a second cancel
on an SSH pane inherited the first cancel's clock. The caller admits
only a working main agent, so the cancel always starts a new done clock.

* fix(agent-status): keep the shell fact on an inferred cancel so restart can seed it

An inferred Ctrl+C cancel beside a working subagent publishes a row held
open by child work, but the synthesized event dropped the row's paired
claudeRunningNonAgentTask fact because mainAgent changed. Hydration seeds
a settled main agent only when that fact says no shell ran, so after a
restart the child's drain left the row working with no mainAgent. Carry
the fact forward: a cancel does not change what the shell inventory said.

* fix(agent-status): a Ctrl+C at an idle main agent's prompt cancels nothing

Every row that publishes the main agent fact now admits an inferred cancel
only while that main agent is working. Grok's Ctrl+C at the idle prompt
leaves its background task running, so settling the monitoring row to
done hid live work. Rows without the fact keep the evidence guard, and
Codex keeps it too because its synthesized row is a plain done.

* fix(agent-status): fold a relayed pane's cancel from its row, not the desktop's records

The inferred Claude cancel read and wrote the desktop's own listener
records for every pane. For an SSH pane those records are not the relay's:
hydration seeds them from the saved row and nothing reaps them, so a
subagent that finished on the remote after a desktop restart kept a
cancelled row spinning with nothing running. A local pane still records
the verdict on its listener and folds its own roster; a relayed pane
folds only the child work its row carries. The relayed-pane parameter and
forced clock the shared record path grew for this are gone.

* fix(agent-status): hold a cancel verdict in the store until a new turn or the provider's own

A relay never learns of the cancel the desktop infers from Ctrl+C, so its
next child hook or reconnect replay restated the main agent as working and
flipped the row back. The late-hook suppression that guarded this keyed on
a done row flagged interrupted, which a cancel held open by a shell or
subagent no longer is; it also dropped Grok's own stop_cancelled when the
inference won the settle race, hiding the task that hook reported.

The suppression is replaced by a latch derived from the row: its main
agent reads cancelled (or, from an older host, a done row flagged
interrupted). A settled incoming main agent, another prompt, an explicit
prompt or a session start releases it. Child and replayed events keep the
latched main agent and are re-folded with their own child evidence; late
main agent work is held as before, and Codex keeps its record re-mark.

* test(agent-status): pin Codex's evidence guard beside the main agent fact

* fix(agent-status): a prompt submission ends the cancel verdict latch

The task notification Claude starts when background work ends is a real
turn, but it keeps the cached prompt and carries no explicit prompt, so
within 15 s of a cancel the latch held its prompt submission and every
tool event after it: the turn read as monitoring under a cancelled main
agent until its Stop. The captured shell cancel has exactly this: the
notification lands 0.17 s after the cancel key.

* fix(agent-status): derive a Codex row's interrupted flag from its main agent

The cancel verdict latch lets any settled mainAgent through, so a late
root Stop after an inferred Codex cancel now applies where the old
same-prompt window held it. It restates the cancellation on mainAgent
but, unlike Claude and Grok rows, carried no interrupted flag, so mobile,
the dashboard and notification text read the cancelled turn as finished.
Codex rows (local and relayed) now derive the flag from the main agent
record, like the other providers that publish one.

* docs(agent-status): describe cancel admission for every provider and the store's cancel-verdict hold

* docs(agent-status): correct the idle-prompt Ctrl+C claim to the measured CLI behavior

* fix(agent-status): preserve waiting relay children on cancel

* fix(agent-status): resolve the cancel hold before a child's permission card adopts a relayed main agent

The permission-card hold took the incoming event's mainAgent before the cancel hold ran,
so on an SSH pane a child's next tool under a sticky card restated the relay's stale
working main agent and dropped the cancellation the desktop had inferred.

* test(agent-status): pin that a cancelled turn's drained subagent settles as stopped, not completed

* fix(agent-status): keep a cancel through a restarted relay's child hook and a teammate's idle

A relay that restarts after a desktop-inferred cancel has lost its prompt cache,
so the child's next hook arrived with an empty prompt, read as a new turn, and
replaced the cancelled main agent with none; the row then stayed working after
every child stopped. A child's empty prompt is now unknown, not another turn; a
non-empty different one still releases, since it is the listener's newer prompt.

TeammateIdle names its child by teammate_name and carries no agent id, so the
latch treated it as the main agent's and let the late-hook window apply it after
15 s, reviving the cancelled turn. It is now re-folded as child work.
2026-09-24 21:24:59 -07:00
Brennan Benson b4d732685c feat(agent-status): combine Codex child work through the shared main-agent status fold (#22475)
* feat(agent-status): combine Codex child work through the shared main-agent status fold

* docs(agent-status): correct two comments the waiting child-work arm made stale

A child failure reported in place as `blocked` now pins the row `waiting`, not
`working`; and no relay ever sent an unfolded `working` beside a waiting child.

* fix(agent-status): only a waiting child asks for a human

A child's `blocked` state means its task failed (the only producer maps a
failed background task to it, and the background-task view labels it
"failed"), not that a human must act. Folding it into the waiting arm would
surface a failed child as needs-you. It stays live work, as before this
series.

* docs(agent-status): say a waiting child, not a blocked one, makes the row wait

A child's blocked state means it failed; only its waiting state feeds the
waiting arm. Two fold comments, a test describe and two parity story names
still called the waiting child blocked.

* docs(agent-status): name where a child's wait is still lost, and pin the structured lane's real input

The doc said the Claude hook lane's rows match Codex and that every lane feeds a
child's wait into the fold. Neither holds: Claude keeps the wait in one slot the
next main agent event overwrites, the structured lane turns a child's prompt
into the main agent's own attention, and Codex drops its roster on a root Stop
when it tracks no child transcripts. The parity story now drives the structured
lane with the input it actually receives.
2026-09-23 22:09:07 -07:00
Brennan Benson 80f5aae0f9 feat(agent-status): publish the main agent's own state beside the combined row state (#22452)
* feat(agent-status): publish the lead agent's own state beside the combined row state

Every status producer folded the main agent's state together with live child
work into one `state`, so a lead that had finished while a subagent still ran
read `working` and its own state was lost. The row now also carries
`lead: { state, outcome?, stateStartedAt }`, admitted by the one payload
normalizer on the relay wire, IPC and disk, and published from the Claude hook
lane, the structured host ingest and renderer bridge, Grok (now on the shared
fold) and Codex (own combine kept). The persisted child-only boundary flag is
derived from `lead` plus child evidence and no longer written; old rows map
onto `lead` at hydrate. Combined `state` and `workingMode` are unchanged for
every reader; a cross-lane parity table pins that, with the cancelled-turn
watch-loop story recorded as a known divergence.

* fix(agent-status): make Orca's inferred interrupt the primary source of a Claude lead cancellation

Current Claude Code sends no hook at all on a cancel and no is_interrupt on
Stop, so the cancellation enters the lead record from the server's inferred
interrupt and rides into the next real Stop; is_interrupt on a turn boundary
stays as the secondary source for builds that send it. Comments, the store
reference and the parity table say so; no suppression changes.

* docs(agent-status): the child-only boundary comment now describes the persisted shell fact

The old sentence said a hydrated row no longer carries the shell fact, which is
the opposite of the mechanism: claudeRunningNonAgentTask is persisted precisely
so hydration can read it, and only a pre-lead row lacks it — reading as
shell-free, the same assertion its legacy flag made at write time.

* rename the lead fact to mainAgent: the main agent's own state

* docs(agent-status): the inferred cancel comes from Ctrl+C, not Esc

* fix(agent-status): an inferred interrupt keeps an already settled main agent, and the row verdict docs name its inferred source

* fix(agent-status): a child-induced wait publishes the main agent state it displaced

* fix(agent-status): decide child-held Claude rows from the saved main agent fact

Restart seeds the Claude main agent from the row's saved mainAgent whenever it
settled and no shell held the row, instead of re-deriving a child-only shape.
OSC cannot settle or repaint a row child agents hold open, including a row
waiting on a child's permission prompt. A sticky child permission prompt still
records the main agent's own progress, and OSC repaints and inferred answers
keep the shell fact beside the main agent they preserve.

* fix(agent-status): keep a finished turn's main agent verdict and clock with that turn

A Claude SessionStart restarts the main agent's clock instead of inheriting the
previous session's last Stop. A Grok idle prompt or session end, and a late
Codex root Stop after an inferred cancel, restate the same finished turn, so
they keep its recorded verdict; only a new turn clears it.

* test(agent-status): publish the Grok verdict restatement past the late-event window

* docs(agent-status): describe hydrate seeding and the OSC refusal from the saved main agent fact

* fix(agent-status): push a held child permission row when its main agent changes

* fix(agent-status): keep the shell fact on a held child permission row so restart does not settle it

* docs(agent-status): note the held child permission row carries the shell fact and is pushed

* fix(agent-status): pair the Claude shell fact with the main agent at the one row-build point

Every non-hook rewrite (terminal-title repaint, inferred answer, held child
permission) had to re-carry the shell fact beside `mainAgent`, and each one that
forgot let a restart settle a row while a shell still ran. The row builder now
pairs the fact once: a listener event restates it, any other write keeps it only
while `mainAgent` is unchanged. Restart seeds a settled main agent only when the
row says no shell ran, and legacy child-only rows map to that explicitly.

A held child permission now also accepts the main agent event's background
evidence, as it already accepts its `mainAgent`, so the child's drain no longer
settles a row a shell still holds. The renderer keeps a previous `mainAgent`
only for writers that never carry one, so a hook row without it matches the
host snapshot.

* test(agent-status): pin that restart never seeds a main agent from a row silent about its shell

* docs(agent-status): the row builder pairs the shell fact with the main agent, and restart seeds only on an explicit no-shell

* test(agent-status): name the legacy-row case parameter for what it holds

* docs(agent-status): name which rows carry the main agent fact
2026-09-23 17:45:50 -07:00
Brennan Benson 4c696a1e2a fix(agent-status): a structured session with live child work reads as working (#22295)
* fix(agent-status): a structured session with live child work reads as working

An idle native-chat session whose subagent was still running showed a green
check in the sidebar, the collapsed worktree pill, and worktree ps, while a
terminal Claude session in the same situation showed working. The two lanes
folded child work into the parent's status with different code: the hook
listener did, the structured lane did not.

Both lanes now share one child-work liveness vocabulary and one lead-status
fold. Live agent work makes a settled lead working; shells and monitors alone
make it monitoring. The structured lane derives liveness from the background
task list already on the wire, in both its readers, so the sidebar, the CLI,
the dashboard and mobile agree. The Claude task-kind table is one shared file
covering both the hook inventory and SDK stream names, and the renderer bridge
reuses the shared child-work projection instead of carrying its own copy.

* fix(agent-status): a blocked or out-of-contact subagent still holds its session working

Child-work liveness retired an agent-kind child on any state but working/monitoring,
while the shell beside it stayed live on everything except done/idle. A subagent
waiting on a permission prompt, or one whose host lost contact, therefore counted
for less than a backgrounded sleep and let the session read done. Both kinds now
share the settlement rule `resolveAgentChildWorkFreshness` already reads rows by:
only an explicit done/idle retires child work.

Also keep empty task labels out of the shared background-task projection candidate,
so a host that publishes `name: ''` cannot beat the child-row fallbacks.

* test(agent-status): pin the widened hook-inventory agent names, and correct two stale claims

The hook inventory now classifies through the shared kind table, which also maps the
SDK stream's `local_agent` / `local_subagent`. Nothing pinned that widening, so add
cases for all four agent names — including `teammate`, whose pane state stays `done`
under the #8825 idle-squat rule.

Two comments the fold made false:
- the teardown marker rule's comment claimed it could not disagree with what the UI
  calls working; it is deliberately lead-only, so now it says that and why;
- the agent-status store reference still described the structured row's `state` as the
  deleted `structuredAgentSessionStatusState`, and omitted the `workingMode` the ingest
  now writes.

* fix(agent-status): the state clock restarts when monitoring becomes a real turn

`stateStartedAt` carried forward whenever the prior `state` matched, which was sound
while `state` meant "a turn is running". Now that it folds in child work, an idle lead
watching a `sleep 3600` publishes `working`/`monitoring`; the user's prompt 45 minutes
later keeps `state: 'working'`, so the row inherited the watch loop's clock and read
"Working for 45m" the instant the turn began. Monitoring is its own displayed label
(`worktree-card-compact-agent-row.tsx:40`), so the continuity key is now the whole
published work identity — state AND workingMode — in both writers.

Also record two facts the code stated wrongly: the structured lane's `interrupted: false`
is inert (a projected session status has no interrupted member) rather than a decision,
and the child-work liveness rule's escape hatch is the roster's session lifetime, not a
settled state.

* fix(agent-status): a workflow is watch work, and child work dates itself

Two defects the fold introduced.

`isAgentChildWorkKind` counted `workflow` as agent work, so a structured session
whose only live task was a backgrounded `local_workflow` published a full working
spinner while the children projection — which admits `kind === 'agent'` only —
rendered nothing to expand, and the same workflow in a terminal pane showed the
monitoring badge instead. The repo already decides this: `isClaudeSubagentTask`
excludes workflows by name, and MATERIALIZED_TASK_KINDS leaves "the backgrounded
shell command and the workflow" to the non-agent owner. The predicate is now
`kind === 'agent'`, and the three sites that restated the same test route through
it, so a new kind is decided in one place instead of three that merely agree.

`evidenceObservedAt` dated every row by `summary.updatedAt`, the journal's last
activity. The journal cannot date child work: its clock stopped when the lead's
turn did, so a genuinely live roster aged past the 30-minute staleness window and
mobile's dot decayed a running session to idle. The fold now reports whether child
work alone holds the row open, and only then does the host's observation clock
stand in — keeping "a restart's republish is not new evidence" for lead turns.

* fix(agent-status): the sidebar dates child work the same way the host does

`fromChildWork` reached the host ingest but not the renderer bridge, so after ~30
minutes of live child work with no journal activity the sidebar's row aged into
staleness while `worktree ps` and mobile stayed fresh — two writers for one session
answering differently, which is the defect this PR exists to remove. For a remote
host the client's own receipt time is also the more honest clock, since the journal
stamp is the host's and is never comparable against this machine's now.
2026-09-22 13:48:07 -07:00
OrcaWinandm4air ba742a86bb fix(linux): release orphaned processes when their owner exits (#22247)
* fix(linux): release orphaned processes when their owner exits

* fix(linux): handle inhibitor errors until streams close

---------

Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
2026-09-22 05:10:03 -07:00
Jinjing 0232c03c43 Refactor editor header file rename to breadcrumb morph UI (#21265)
* Refactor editor header file rename to a breadcrumb morph UI

- Display full breadcrumb (repo name + parent dirs) during rename for context
- Separate basename field from extension suffix to clarify what users edit
- Replace blur-to-commit with explicit confirm/cancel buttons
- Auto-attach extension to basename; respect explicitly typed extensions
- Add comprehensive tests for rename scenarios and edge cases

* Fix markdown rename: drop blur-commit, handle IME, track active file

- Blur no longer commits, preventing accidental renames when focus moves
- IME composition keys properly handled for CJK input support
- File switches during rename now cancel the operation
- Extension display improved to show final filename
- Comprehensive test coverage added for edge cases

* Simplify markdown rename to inline field without breadcrumb or buttons

Replaces breadcrumb-morph rename with confirm/cancel buttons with a
simple inline field accepting full filenames. Commits on Enter, blur, or
Escape to cancel, matching tab bar and file explorer behavior. Adds
renameCancelledRef to prevent blur-commit after Escape. Removes unused
i18n strings for buttons and simplifies state by dropping extension
pinning and breadcrumb display.

* Handle blur race when file changes during rename

When switching files mid-rename, React may deliver the old input's
blur event after the new file renders, causing a stale rename commit.

Mark the rename as cancelled when the active file changes, and add
test coverage verifying stale blur events are ignored.

* Test blur-race condition in hook unit test

Move blur-commit-after-file-change test from EditorPanelHeaderPath
integration tests to useEditorHeaderFileRename unit test. Tests the
blur-handling logic at the hook level where it belongs.
2026-09-21 22:55:43 -07:00
88f2f01061 fix(daemon): escape the terminal daemon into its own systemd scope so a service restart no longer kills every live PTY (#19430)
* fix(daemon): escape the terminal daemon into its own systemd scope so a service restart no longer kills every live PTY

Root cause: daemon-launched-child.ts forks the detached terminal daemon with
detached: true, which escapes the POSIX process group (setsid) but never the
systemd cgroup. Every PTY the daemon owns is itself an undetached direct
child of the daemon (native-pty-spawn.ts). Under a combined systemd unit
(Type=simple, KillMode=mixed, per docs/reference/headless-linux-server.md),
a systemctl restart/stop SIGKILLs every process still in the cgroup at the
stop timeout -- the daemon and every live terminal -- even though the
codebase already has a fully-built adoption/reattachment path for a
surviving daemon (orcad-entry.ts's refreshRestoredOrchestrationAuthority +
reconcileLegacyWorkerTerminals, gated on daemonOwnsFreshPersistentPtys()).
That path never fires today because the daemon never survives long enough.

Fix: when systemd is actually supervising the process and the OS user has a
reachable systemd --user manager (isDurableDaemonScopeSupported(), Linux
only), launch the daemon via systemd-run --user --scope so it lands in a
cgroup that is a sibling of the service unit's cgroup, not a descendant of
it. A systemctl restart of the combined unit then never reaches it. Any
failure of the scoped launch (no reachable bus, D-Bus policy rejection,
etc.) falls back transparently to the existing plain fork() launch, so
every platform/environment without this capability is unaffected.

The daemon self-detects its own resulting cgroup scope via /proc/self/cgroup
(detectOwnCgroupScopeUnit()) rather than trusting the launcher's intent, and
publishes it as cgroupUnit in its pid record and orcad's health/readiness
payload (health.terminalDaemon.cgroupUnit), so a running deployment can be
observed to confirm the fix actually engaged.

No new session registry is added: the existing daemon pid-record + adoption
protocol (publishDaemonPidFile, daemon-pid-record-quarantine.ts's
dead-record reclaim, refreshRestoredOrchestrationAuthority) already
implements durable, crash-safe reattachment for a surviving daemon -- it
was simply never exercised against a full unit restart before now.

Proven via a systemd-in-Docker recovery test: a live PTY session's shell
process, its daemon, and the daemon's cgroup scope were all confirmed
unchanged across a real systemctl restart of a Type=simple/KillMode=mixed
unit, while the main process pid changed (confirming the unit actually
restarted) and the new process's health payload recognized the surviving
daemon as adopted and live. A fresh write into the same PTY post-restart
reached the same running shell. Ordinary terminal create/work/release and
the #18789/#18790 worker-release reap-fix regression tests are unaffected.

Fixes stablyai/orca#19408

* fix(daemon): probe the real per-UID XDG_RUNTIME_DIR before trusting the process's own env

isDurableDaemonScopeSupported()/buildDurableDaemonScopeCommand() trusted the current
process's own XDG_RUNTIME_DIR env var first, falling back to /run/user/<uid> only when
that var was unset entirely. On mtl-02, orca-serve@factory.service's RuntimeDirectory=
hardening directive makes systemd export XDG_RUNTIME_DIR=/run/orca_serve/factory into the
unit's process -- a private scratch dir that shares the env var's name but has nothing to
do with the user session bus. /proc/<pid>/environ on that host confirmed exactly that path
plus DBUS_SESSION_BUS_ADDRESS=disabled:, while the real bus was reachable the whole time at
/run/user/985 (confirmed via systemctl --user is-system-running with that dir exported by
hand). The probe treated the hardened override as authoritative, found no bus socket there,
and reported unsupported on every launch -- so the cgroup-escape fix from #19408/#19430
never actually engaged on real hardware, even though tonight's factory deployment picked it
up.

Fix: resolveUserRuntimeDir() now always tries the conventional /run/user/<uid> path first
(computed independently via getuid(), never trusted from env), checking for a genuinely
connectable bus socket via statSync(...).isSocket() rather than a bare existsSync. It falls
back to the process's own XDG_RUNTIME_DIR only when that canonical path has no reachable
bus -- covering hosts that legitimately have no /run/user/<uid> at all but do have a
working bus wherever their own environment points. buildDurableDaemonScopeCommand() now
explicitly sets XDG_RUNTIME_DIR to whichever path this resolution picked, rather than
inheriting the spread env's (possibly hardened-wrong) value.

Both isDurableDaemonScopeSupported() and buildDurableDaemonScopeCommand() gained an
injectable canonicalRuntimeDir parameter (defaulting to the real computed path) so tests
can exercise the hardened-override scenario deterministically with a real, connectable
AF_UNIX socket fixture instead of the live host's actual runtime directory.

Docker's stock jrei/systemd-ubuntu test container never had this hardening directive, so
this gap was structurally invisible to the container-based verification in #19430 -- only
caught against real mtl-02 hardware.

* fix(daemon): report the daemon's own pid over the ready handshake, not systemd-run's

The launcher used to infer the daemon's identity pid from the immediate
spawned child (`child.pid`). On the durable-scope path that child is
`systemd-run --user --scope`, not the daemon, so the launcher was asserting
an identity it had no authority over.

`DaemonReadyIdentity` now carries a required `pid` populated from
`process.pid` inside the daemon itself, and `daemon-launched-child.ts` takes
`launchedIdentity.pid` from that self-report. Both sides of the
`holdDaemonAdoptionLease` pid comparison therefore originate inside the
daemon process, which is the idiom this branch already uses for cgroup
membership (`detectOwnCgroupScopeUnit` reads `/proc/self/cgroup` rather than
trusting what the launcher intended).

Note on the reported consequence: `systemd-run --scope` registers its *own*
pid on the transient scope unit and then `execvpe()`s the target command --
same pid, no intermediate process -- so adoption did not in fact fail on
systemd >= 206 (verified against systemd 255.4-1ubuntu8.17 and current main,
`src/run/run.c` `start_transient_scope()`). The fix stands on its own merits:
it removes a silent dependency on that exec-vs-fork implementation detail,
which a `systemd-run` shim earlier in PATH or any future systemd change would
have broken with no diagnostic.

`terminateLaunchedDaemonChild` was audited and deliberately left on
`child.pid`: for the same execve-preserves-pid reason that pid is either
still systemd-run mid-scope-setup (killing it correctly aborts the launch) or
already the daemon, so it targets the right process either way.

Regression coverage: `daemon-launched-child-identity.test.ts` pins the
identity source, and `daemon-ready-identity.test.ts` gains pid-validation
cases. Ready-message fixtures across the `daemon-init-*` suites were updated
for the now-mandatory field.

Addresses:
https://github.com/stablyai/orca/pull/19430#discussion_r3953722704
https://github.com/stablyai/orca/pull/19430#discussion_r3954346518

* test(daemon): assert cgroupUnit in the pid-file parse contract

`parseDaemonPidFile` returns `cgroupUnit` on every branch as of the
durable-scope commit on this branch, but five exhaustive `toEqual`
assertions in daemon-health.test.ts still described the pre-scope shape, so
they failed on the branch independently of any later change.

Adds the field to those expectations. Deliberately not relaxed to
`toMatchObject`: asserting the full parsed shape is what makes these tests
catch a field silently dropped from the pid-file contract.

* refactor(daemon): resolve the canonical user runtime dir at one point

The per-UID path cannot change for a live process, so compute it once into a module
const instead of threading the same default call through three signatures, and drop
the try/catch around a getuid() that cannot throw once it exists. Trims the module
prose to the non-obvious facts and corrects the pid-file record comment: an unscoped
daemon writes null; only records no daemon wrote are absent.

* test(daemon): clean up the cgroup-scope fixtures and assert a verdict

The cgroup fixture tracked only the file it wrote, leaking one temp dir per case.
Drains both fixture lists with splice so the pop-may-be-undefined guards go away,
and replaces a not-throw/typeof-boolean pair with the verdict it was circling:
no resolvable runtime dir means unsupported.

* refactor(daemon): share the detached child options across both launch paths

cwd, detached and stdio were repeated in the fork and systemd-run branches, which
left the two comments explaining them hovering over the env block instead. Names
them once so each branch carries only its own delta.

* refactor(daemon): validate the ready pid like every other field

typeof-first narrows the value, so the two 'as number' casts the isSafeInteger check
needed disappear and the pid guard reads like the startedAtMs guard below it.

* fix(daemon): don't retry the launch unscoped after losing the endpoint race

A scoped attempt that lost the endpoint to another daemon was retried unscoped: a
second doomed fork, a misleading 'cgroup-scope launch failed' warning, and the same
DaemonEndpointUnavailableError the caller was already going to adopt on. Rethrows it
instead, since no launch mode can win a race that is already lost.

Also drops a private alias for DaemonChildSpawnOptions and the two 'as number' casts
on child.pid in the startup-failure cleanup.

* fix(daemon): unlink the pid record by the pid the daemon published

The record holds the daemon's self-reported pid, so match on that rather than on the
immediate child's, which is the systemd-run wrapper's until it execs.

* fix(daemon): route the scope launch through the child-process chokepoint

The two files this PR added imported `node:child_process` directly, which
`child-process-import-boundary.test.ts` fails on deterministically: the
offender count went 155 -> 157 against a pin of exactly 155. Raising the pin
or listing the files is what that test explicitly forbids, and the allowlist's
own note says a split "moved the import, it did not add one" -- so the fix is
to get both new files off the module and put the count back at 155.

- `daemon-cgroup-scope.ts`: the `systemd-run --version` probe now uses
  `runProcessSync` instead of `execFileSync`, so it gets the shared spawn
  decisions. Kept synchronous deliberately: `launchDaemonChild` attaches the
  readiness listener in the same tick it is called, and an await before the
  spawn moves the child past that tick. A non-zero exit is data rather than a
  throw here, so the verdict now checks `code === 0 && !timedOut`.
- `daemon-launched-child-spawn.ts`: the scoped launch uses `spawnProcess`, and
  the long-standing unscoped launch keeps `fork` semantics through a new
  `forkProcess`.
- `src/shared/child-process/fork-process.ts`: the fork arm of the chokepoint.
  `spawnProcess` cannot express a Node child with an IPC channel started from
  a module path under an overridden `execPath`, and the existing launch tests
  are written against `fork`'s contract, so a spawn rewrite would have changed
  module resolution, `execPath` and `execArgv` at once. It passes
  `windowsHide: true` -- the flag every other call site in that directory
  sets, reachable via an assertion because `ForkOptions` omits it -- which
  keeps `windows-console-visibility.test.ts` at its pin of 65 too.

Both ratchets pass with both pins and both allowlists untouched.

Docs: `orcad-operations.md` and `headless-linux-server.md` still described the
limitation this PR removes as permanent. Both now describe the durable-scope
survival path and its preconditions (systemd as PID 1, a reachable user bus /
`loginctl enable-linger`, `systemd-run` on PATH), and scope the old text to
the unscoped-fallback case, pointing at `health.terminalDaemon.cgroupUnit` as
the way to tell the two apart on a running host.

* fix(daemon): seal the cgroup capability probe from the host and correct KillMode=mixed docs

The capability probe consulted the host's own /run/systemd/system marker and
spawned the real systemd-run binary, so the hermetic unit tests could only pass
on a systemd host (and fail closed otherwise, even with faked bus sockets).

- Thread systemdBootPath and runVersionProbe as test seams through
  isDurableDaemonScopeSupported, defaulting to the real boot marker and
  systemd-run --version probe in production.
- Narrow the injected probe to the ProcessResult slice it consumes.
- Cover: no-systemd-boot, non-zero probe exit, and probe-timeout cases.
- Correct KillMode=mixed semantics in the docs: the cgroup-wide SIGKILL fires
  the instant the main process exits, not after TimeoutStopSec; document the
  Docker-container caveat and add KillMode=mixed to the multi-service template.

* fix(daemon): satisfy assertion checks in scoped launch

* fix(daemon): satisfy anti-slop and console guards

* test(serve): update shutdown docs assertions for daemon scope

* fix(daemon): migrate adopted legacy scopes

* docs: qualify restart safety by daemon scope

* docs(daemon): qualify Upgrade restart prose with durable scope caveat

Align the Upgrade section in docs/reference/headless-linux-server.md with
the earlier preservation section and docs/reference/orcad-operations.md:
a service restart terminates live processes only when running under the
unscoped fallback, and stops should be treated as destructive unless
health.terminalDaemon.cgroupUnit names an orca-daemon-*.scope.

Update the shutdown workflow test assertion in
config/scripts/headless-serve-shutdown-workflow.test.mjs to match.

* fix(daemon): harden legacy scope migration

---------

Co-authored-by: Lesley Murfin <260182349+LesleyMurfin@users.noreply.github.com>
Co-authored-by: m4air <m4air@Mac.localdomain>
2026-09-21 17:23:30 -07:00
Neil 35fe67b610 fix(perf): measure terminal latency with presented CI frames (#22096)
* fix(perf): present benchmark frames only on isolated CI display

* fix(perf): wait for the benchmark page before presenting its window

* docs(perf): record full scale pass with unchanged latency budgets

* test(perf): document and verify the isolated display exception
2026-09-21 16:03:40 -07:00
Neil 8cf0e81ced fix(perf): calibrate report budgets without masking latency stalls (#22075)
* fix(perf): calibrate report budgets without masking latency stalls

* docs(perf): record historical evidence for report limits
2026-09-21 14:13:09 -07:00
Neil 5fed670a50 fix(rebuild): refuse to compile node-pty source that lacks the MSYS breakaway denial (#21968)
The addon gate already rejects a conpty.node without the L"msys-2.0.dll"
marker, in the Electron probe and after the rebuild. But the rebuild compiles
whatever node_modules/node-pty holds, and pnpm only materializes that from the
patch at install time. On a Windows dev checkout whose node_modules predated
the denial, --force compiled for minutes, rewrote conpty.node byte-identical
and unpatched, and the gate then advised "rebuild from source" -- the step
that had just run.

Read src/win/conpty.cc before compiling. If it lacks the literal, stop before
the rebuild and say to run pnpm install, which re-applies the current patch.
An absent source file is not judged; the addon gate still reads the binary.
2026-09-21 03:00:44 -07:00
Neil 37dbb6cf6e fix(windows): prune the unpatched conpty prebuild by header, not host arch (#20048)
* fix(windows): prune the unpatched conpty prebuild by header, not host arch

`prunePackagedNodePty` deleted the published `prebuilds/win32-<arch>/conpty.node`
only when `electronArch === process.arch`. That proxy stood in for "build/Release
holds an addon this slice can load", and it is false for the arm64 slice
cross-built on an x64 Windows host — a rebuild that DOES emit a correct arm64
addon. That slice shipped the unpatched prebuild.

Nothing loads it today: `verifyPackagedConptyBreakawayMarker` resolves the addon
the way node-pty's loader does, so the patched `build/Release` wins and the
release passes correctly. But the loader swallows every require failure and falls
through, so an AV quarantine or a missing dependency on `build/Release` hands the
pane to that unpatched prebuild — the silent downgrade the gate exists to close,
with the binary still sitting in the package.

Read the PE `Machine` field instead of guessing, reusing `readPeMachine` from the
verifier's `windows-pe-machine.cjs` so prune and verifier ask one question. A
missing, truncated or non-PE `build/Release` reads as unloadable and keeps the
prebuild, which is what the true cross-host case needs: packaging Windows from
macOS leaves no Windows binary in `build/Release`, and removing the prebuild there
would leave the package with no ConPTY at all.

Mutation-proven: restoring the `electronArch === process.arch` guard fails exactly
the two new rows in packaged-node-pty-prebuild-prune.

* docs(windows): note the cross-arch conpty slice is real but not yet built
2026-09-20 23:57:17 -07:00
Wooseong KimandNeil 30f2bc60f9 fix(antigravity): recognize non-Gemini tui-idle prompts (#21231)
* fix(antigravity): recognize non-Gemini tui-idle prompts

* fix(antigravity): reject stale composer caret in model picker

* fix(antigravity): do not treat a wrap continuation caret as ready

An unsent composer can show `> draft` then an indented `>`. That continuation is not an empty input box, so tui-idle must stay false.

* test(antigravity): align later bare-caret status expectation

---------

Co-authored-by: Neil <neil@stably.ai>
2026-09-20 18:01:45 -07:00
Neil 85a3ba6d42 fix(terminal): align CJK IME preedit spacing (#19367)
* fix(terminal): align IME preedit to terminal cell grid

* fix(terminal): preserve native shaping and reuse IME preedit on repaint

* fix(terminal): preserve native shaping with bounded IME spacing runs

* test(terminal): account for inline preedit subpixel rounding

* test(terminal): keep the IME grid fixture wide at every DPI

* chore: regenerate xterm patch after rebase

* test(terminal): remove IME assertion lint findings

* test(terminal): avoid reflective IME fixture access

* test(e2e): run IME renderer matrix with WebGL available

* fix(ci): restore editor line budget
2026-09-19 16:44:38 -07:00
Neil e8a7be4ce2 fix(omp): recover retired pane status with validated restart authority
Merged after fresh run 35448889017 passed all required checks, including static analysis, typecheck, package jobs, all test shards, changed E2E, Docker SSH E2E, and verify.
2026-09-19 08:05:33 -07:00
Neil 3ba7cb4de9 feat(diagnostics): trace terminal startup delivery phases
Merge fully verified: all required CI checks pass. This lands bounded startup timing instrumentation for the open Windows OMP first-paint investigation in #19333; it does not claim the latency fix itself.
2026-09-19 06:58:31 -07:00
Neil b0ec11f5b0 fix(omp): redact credential references before status transport (#21673) 2026-09-19 05:58:09 -07:00
Neil ff934256ae fix(omp): retain recorded transcript paths when resuming (#20634)
Based on the resume-locator proposal in stablyai/orca#16276 by @CodeHourra. Retains UUID-based ownership and existing reattach behavior.
2026-09-19 02:02:50 -07:00
Neil 605a4ef868 fix(omp): start new tasks without auto-resuming old sessions (#20622)
* wip(omp): prove fresh settings overlay without redirecting storage

* fix(omp): guard fresh launches with execution-host settings

* fix(omp): preserve unmodelled shell launch commands

* test(omp): consolidate shell fixture path import

* preserve fresh OMP launch status

* test: cover preserved OMP launch status

* fix: recognize wrapped fresh OMP launches

* chore(ci): refresh validation against fixed main baseline

* fix(omp): recognize generated fresh launch guards across shells

* fix(omp): preserve draft status and clear prefill across Unix shells

* fix(omp): run cmd draft cleanup after either guard branch

* fix(omp): launch drafts safely with nounset enabled

* fix(omp): select draft shell without parser diagnostics

* test(omp): await relay environment augmentation
2026-09-19 01:42:14 -07:00
Neilandstevelliu ea02d90704 fix(omp): answer startup Kitty queries before renderer handoff (#20620)
* fix(omp): answer startup Kitty queries before renderer handoff

Forward actual renderer capability through local and remote spawn. Preserve source ranges and following keyboard mode pushes, and retain independent ConPTY color authority.

Refs #17081. Secondary review: #17082.

Co-authored-by: stevelliu <stevelliu@tencent.com>

* test(omp): cover fragmented keyboard modes and ConPTY handoff

* fix: preserve keyboard startup intent without terminal colors

* fix: negotiate keyboard support for host-authoritative agent launches

* fix: keep terminal creation within line budget

* fix(omp): negotiate keyboard support for paired web launches

* test: remove obsolete message type import after main integration

* fix: validate paired launch results and retry incomplete SSH test snapshots

* fix(omp): negotiate keyboard support for background paired launches

---------

Co-authored-by: stevelliu <stevelliu@tencent.com>
2026-09-19 01:29:16 -07:00
Neil 9841f4f790 fix(omp): keep child sessions from reporting into the parent pane (#20611)
* fix(omp): fence pane status to the root session manager

* test(omp): preserve root preview and recovery through child hooks

* test(omp): exercise status ownership through actual runtime runner

* fix(omp): honor runtime subagent provenance when available

* test(omp): avoid writes to the read-only hook status view

* fix(omp): preserve child status ownership provenance

* fix(omp): normalize child transcript paths across platforms

* fix(omp): clean status handler rebase

* fix(omp): keep prefill inside session ownership fence
2026-09-19 00:29:55 -07:00
Neil a913f1b457 fix(omp): resolve transcripts from active XDG and profile roots (#20633)
* fix(omp): resolve XDG session root fallback

* fix(omp): share profile-aware transcript root resolution

* test: verify actual OMP XDG transcript discovery and reading

* test(omp): satisfy reader smoke loop lint

* test(omp): resolve vault fixture paths without redundant cache

* fix(omp): refuse normalized Windows drive-relative roots
2026-09-19 00:17:49 -07:00
Neil 2bf538a4e1 fix(runtime): detect a same-size terminal artifact swap the granted stat cannot see (#21436)
* fix(runtime): detect a same-size terminal artifact swap the granted stat cannot see

A local terminal-artifact grant pinned the file as `dev:ino:nlink:size:mtimeMs`.
On Linux every one of those can survive an unlink+recreate: ext4 reuses the
just-freed inode (measured: 100% of the time), nlink and size are unchanged for a
same-size replacement, and the mtime clock is tick-quantized to 1ms, so a swap
inside one tick produces a byte-identical identity string. The grant then served
the attacker's bytes as if nothing had changed.

Local grants now also pin a sha256 of the artifact's content, taken from the same
handle as the stat so nothing can swap the file between them, and every local
read, preview and write re-checks it before returning or committing content.

The stat identity string itself is unchanged: the relay recomputes it verbatim to
honour `expectedStatIdentity`, so its format is a wire contract. Remote grants
keep the stat-only check and are untouched.

This is also the mechanism behind the intermittent
`orca-runtime-files-terminal-artifact-io.test.ts` failure on
`rejects stale absolute terminal artifact previews before returning changed
content`: it replaces an 8-byte artifact with 8 different bytes, so whenever the
two writes shared a 1ms tick the product genuinely could not tell them apart.

* docs(runtime): record what the terminal artifact grant checks do not close

The digest makes the same-size swap detectable; it does not make the sequence
atomic. A reader arriving at the access module would reasonably assume otherwise,
so write down the measured limits of the stat identity, why the identity string
cannot change, and the four windows that stay open — the write path's surviving
rename() gap above all.
2026-09-18 23:33:21 -07:00
09073086a8 feat(terminal): inline images via @xterm/addon-image (perf-first) (#19512)
* feat(terminal): inline images via @xterm/addon-image, perf-first

Add opt-in inline terminal images (SIXEL, iTerm2 IIP, Kitty graphics)
through @xterm/addon-image, designed to keep idle terminals unaffected.

Performance:
- The addon (base64-inlined wasm decoders + protocol handlers) loads off
  the boot critical path via a deferred loader that mirrors the WebGL
  addon: primed after first paint only when the setting is on, read back
  synchronously at attach, with a 3-attempt cap so a transient failure
  never disables images for the session and a missing chunk never
  refetches per pane. renderer-boot-graph guards against eager import.
- enableSizeReports:false so the addon never sets windowOptions and
  double-answers Orca's own CSI 14t/16t responder.
- Perf-tuned decode/storage limits (storageLimit, sixel/iip/kitty size
  caps) in one place.

Correctness:
- Orca's DA1 handler wins over the addon's (last-registered-first), and
  the default DA1 response never advertised Sixel (;4), so DA1-detecting
  tools (chafa, img2sixel, viu, timg) never emitted it. The winning
  handler now appends ;4 while the setting is on, resolved per query so a
  live toggle changes the next DA1; idempotent against the ConPTY
  response that already lists it.
- ORCA_IMAGE_PROTOCOL=kitty is exported to spawned shells (local, daemon,
  relay/SSH) and forwarded across the WSL boundary, so image-capable
  agents can pick an encoder. Unknown image sequences are swallowed by
  xterm when the addon is detached, so this never garbles output.
- Settings toggle (default on) gates rendering and DA1 advertisement.

Cross-checked against community PRs #7775, #11706, and #19201 at the end;
credited below.

Co-authored-by: s546126 <s546126@users.noreply.github.com>
Co-authored-by: XRX193 <XRX193@users.noreply.github.com>
Co-authored-by: lmsh7 <lmsh7@users.noreply.github.com>

* fix(terminal): bound inline image memory and classify Kitty replies

* fix(terminal): bound image decode and release image resources on cleanup

* fix(terminal): address image addon review feedback

* test(terminal): stub setPaneInlineImagesEnabled in appearance manager fakes

* fix(terminal): evict unplaced kitty payloads before displayed images

Byte-budget eviction dropped the oldest transmitted blob regardless of
placement, so a new upload could erase a visible image while abandoned
blobs still held budget. Unplaced payloads now go first and displayed
ones only when that is not enough. The incoming image is always stored,
so an oversized one overshoots the cap by one payload instead of being
dropped after the protocol already acked OK.

* fix(terminal): gate DA1 Sixel on real addon attachment; claim SSH image spec in CI

- DA1 advertised Sixel from the setting alone, so a pane whose lazy addon
  chunk was still loading (or had failed all three attempts) told
  feature-detecting tools to emit DCS that nothing could render. Track the
  attached decoder per terminal and require it before setting the ;4 bit.
- tests/e2e/terminal-inline-images-ssh.spec.ts was Docker-gated but claimed
  by no lane runner, so pr-e2e-gate-contract failed and the spec would have
  self-skipped green forever.
- Reject non-positive PNG IHDR dimensions before decode: they are parsed with
  signed shifts, so a dimension >= 0x80000000 came back negative and slipped
  past the pixel-limit comparison.
- One resolveTerminalInlineImagesEnabled() for the default-on setting; the
  four call sites mixed '?? true' with '!== false', which disagree on null.
- One readInlineImageResources() walk of the addon internals instead of two
  copies that could drift against the patched dependency.
- Isolate the deferred-attach drain per pane; make the zoom-invariance and
  backing-storage e2e assertions fail when the feature is dead.

* refactor(terminal): one lazy xterm addon loader for webgl and image

terminal-image-addon-loader was a structural clone of the webgl one — same
memo, attempt cap, and .then(ok,err)-clears-memo recovery. Both now wrap
createLazyXtermAddonLoader; each keeps its literal import() specifier so the
bundler still splits the chunk (verified against a fresh build: addon-image
stays out of the boot graph).

* refactor(terminal): name openTerminal's addon flags; pin image addon limits

Two adjacent optional booleans could be swapped without a type error once
inline images added the second one.

* docs(terminal): state the real per-pane image ceiling; drop test ordering dependency

storageLimit:32 reads like the pane's budget but keys three pools — decoded
pixels, retained encoded Kitty blobs, and pending WASM decoders — so the worst
case is ~98 MB per pane with no cross-pane governor. Say so at the constant.

pane-inline-images.test.ts's deferred case needed to run first; it now takes a
fresh module instead, and the rest prime in beforeAll. Verified by running the
file with that test moved last.

* fix(terminal): satisfy rebased static analysis gate

* fix(terminal): complete casting gate cleanup

* fix(terminal): recover failed image addon loads

* fix(terminal): bound image decoder allocations

---------

Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
Co-authored-by: s546126 <s546126@users.noreply.github.com>
Co-authored-by: XRX193 <XRX193@users.noreply.github.com>
Co-authored-by: lmsh7 <lmsh7@users.noreply.github.com>
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
Co-authored-by: Neil <neil@stably.ai>
2026-09-18 16:32:49 -07:00
Neil a61119ceb0 refactor(runtime): name the four answers a host probe can give (#21207)
The renderer expressed every non-answer as one nullable `status`, so a probe in
flight, a probe that failed, a host that refused us and a retired pairing all
reached readers as the same `null` -- and readers spent that `null` on decisions
of very different weight, including destructive ones.

`RuntimeHostContact` names the four. Nothing changes yet: the connection-state
derivation is rewritten on top of it and a 384-case parity table asserts the
result is identical to a frozen copy of the old one on every combination of
verification, transport, retired, answered and remote-control state.
2026-09-17 21:22:39 -07:00
Neil ea01cd0ccd fix(windows): reject a node-pty addon that predates the MSYS breakaway denial (#20047)
* docs(windows): record the measured MSYS job-breakaway mechanism

The per-PTY job already denies JOB_OBJECT_LIMIT_BREAKAWAY_OK for Cygwin/MSYS
shells (#19068), but nothing records why, and a conpty.node built before that
commit fails windows-msys-job.win32.test.ts in a way that reads as a source
defect. Measured on a real Windows 11 host: both the plain and the exec-
replacement Git Bash shapes leak, the escape is the MSYS runtime's own
spawn/exec (fork keeps membership), and a single-variable A/B on
usesCygwinRuntime flips the result 0/2 -> 4/4.

Also names the gap the failure hid behind: node-pty-job-ownership.cjs asserts
symbol presence, which cannot distinguish patch revisions.

* fix(windows): reject a node-pty addon that predates the MSYS breakaway denial

The native-runtime gate asserted only that terminateJob, listJobProcessIds and
assignCurrentProcessToJob were exported. All three predate the Cygwin/MSYS
breakaway denial, so an addon built before it passes every gate,
isPtyJobOwnershipAvailable() returns true, and windows-pty-job.win32.test.ts
passes 6/6 -- while every Git Bash child is created outside its pane's job and
survives terminatePtyJob.

Read the resolved .node and require the wide msys-2.0.dll literal that
usesCygwinRuntime holds, the way stagedRelayAddonIsUnpatched() already tells a
patched windows-process-tree addon from a published one. An addon the caller
cannot name is refused rather than skipped: a gate that cannot see its subject
is not a gate.

Verified against real binaries on a Windows 11 host: the shared checkout's
pre-#19068 build errors, a build from current patched source passes, a missing
path errors.

Also closes the cross-host packaging skip. The export half has to load the
addon so it cannot run when the packaging host is not the target, which is how
a Windows release built elsewhere could ship this. The marker is a file read
and needs neither; an unrecognised layout warns rather than fails a release
that was packaging fine.

* fix(windows): check the MSYS breakaway denial on the rebuild path too

The Electron probe carried the marker check, but it lives inside
probeElectronNativeModules, which returns early whenever the Electron package
binary is unusable. Covered by another path is not this path checks -- and the
defect this whole change closes was a gate that looked like it checked.

Reading the binary needs neither a loadable Electron nor an executable target
arch, so assert it after the rebuild, beside the windows-process-tree
assertion that exists for the same reason: this is the addon copied into the
packaged app. Absent warns (a cross-platform rebuild need not leave a win32
addon on this disk); present and unmarked is fatal.

The fixtures now write a real addon file, because the gate reads the binary it
was told about rather than trusting the exports. Verified against the two real
binaries measured on the Windows host: the pre-#19068 build fails this path,
the build from current patched source passes.

* fix(windows): check the marker on every ConPTY path the packaged app can load

The packaged marker check read one hard-coded path, `build/Release/conpty.node`,
and warned when it was absent. `loadNativeModule` tries `build/Release`, then
`build/Debug`, then `prebuilds/win32-<arch>`, swallowing each failure, and
`prunePackagedNodePty` drops the published prebuild only when a same-arch
`build/Release` exists to replace it. So the two packages the check was added for
were the two it could not see:

- cross-host: no host but Windows can build conpty.node, so there is no
  `build/Release` and the prebuild is what ships. The check warned and returned.
- cross-arch: `build/Release` is the packaging host's own arch, patched and
  marked, so the check printed OK -- while the target app cannot load it and
  falls through to the unmarked prebuild underneath.

Measured, not assumed: both published Windows prebuilds in the node-pty tarball
contain neither `msys-2.0.dll` nor `cygwin1.dll` in any encoding. They are the
binary that leaks every MSYS pane child out of its job.

It now sweeps every candidate present for the *target* arch and refuses a package
with no candidate at all, which is a package with no ConPTY backend rather than a
layout to shrug at. It runs for every Windows slice instead of only the branch
the export check skips, so deleting the export check cannot silently take it too.
A stale source build keeps the rebuild advice; the prebuild gets the advice that
actually works, which is to package the slice on a Windows host of that arch.

Also: the marker constant was re-typed in four places and was tied to the C++
literal that produces it by nothing at all, so editing the patch would have left
a gate that fails every correctly rebuilt addon and tells the developer to do the
one thing that cannot help. The fixtures now take the constant from the gate, and
a test asserts the patch still adds `L"msys-2.0.dll"` to conpty.cc.

And the rebuild path treated a missing addon as a warning even on the host that
will run the install, where node-pty would fall through to that same prebuild.
The verdict is now a value, so it is tested without a platform gate.

* fix(windows): resolve the packaged ConPTY the way its loader does

Sweeping every candidate and demanding the marker on all of them was wrong in
the one case it was meant to make safe. `beforeBuild` runs
`rebuild-native-deps.mjs --platform=win32 --arch=<target>`, so a cross-arch slice
normally does get a patched `build/Release` for the target; `prunePackagedNodePty`
keeps the prebuild anyway because its guard is `electronArch === process.arch`
rather than the arch of the binary. That package is correct and its leftover
prebuild is never reached, and the sweep failed it -- telling whoever ran it to
package on a Windows arm64 host, which is both the wrong remedy and one no runner
here can offer.

Presence cannot separate that package from the one whose cross-arch rebuild
quietly emitted the host's architecture, because the only difference is the arch
of `build/Release`. So the gate now resolves the addon the way `loadNativeModule`
does -- first candidate whose PE `IMAGE_FILE_HEADER.Machine` matches the target,
walking root-then-lib for each layout in node-pty's own order -- and checks the
marker on the one that will actually run. A package with no candidate, or none of
the target's architecture, is refused: it has no ConPTY backend either way, and
the second is exactly what a silently host-arch cross-build looks like.

The PE machine reader already existed, privately, in the relay addon builder that
needed the same "a cross-build cannot silently emit host arch" guarantee. It is
now shared rather than copied.

Two seams were unreachable from anything but Windows, so nothing tested them:

- the afterPack hook's win32 block was an inline if/else that only a source-text
  assertion could inspect, and that assertion could not tell the difference
  between the check running and the check being wrapped in `try {} catch {}`. It
  is now `verifyPackagedWindowsNodePty`, and "the marker check runs even where
  the export check cannot" is four spied assertions instead of a string match.
- the rebuild path's verdict read `process` directly, so the branch that fires
  only on the host being rebuilt for was dead on every other host. It now takes
  the host as arguments, and the fs checks, the warning and the failure are all
  exercised from macOS.

Fixtures write a real PE header rather than `MZ fake addon`, since the gate now
reads one. The machine table is pinned to the documented IMAGE_FILE_MACHINE
values, because every fixture builds its header from that table and a table wrong
in both entries would otherwise agree with itself.

* fix(windows): say why the packaged ConPTY fell back, not just that it did

The previous commit resolved the addon by architecture but still had one message
for every way the resolution could land on the published prebuild. Those ways
want opposite remedies, and the one it printed was the remedy the commit before
it had just called wrong:

- no source build in the package at all — the slice has to be built somewhere
  that can build node-pty for the target arch.
- a source build that is there but is the packaging host's architecture, because
  the cross-arch rebuild did not honour `--arch` — re-running that rebuild is the
  fix, and "package on a Windows arm64 host" is neither necessary nor possible.

The second is the common one, since node-pty publishes a prebuild for both
Windows arches and prune keeps the target's on every cross-arch package. So the
old text fired mostly on the case it described least. It now reports which source
builds were skipped and the machine field each carried, and names the rebuild
command.

"Nothing the target can load" had the same problem in reverse: a zero-length or
truncated `conpty.node` got a cross-architecture diagnosis. Every candidate is
now named with what was actually read, including "not a PE image".

The rebuild path asserts the architecture too. A rebuild that ignored `--arch`
was otherwise only visible at packaging, two steps from the command that fixes
it. Arches with no known machine value are left unjudged rather than guessed at.

Two things the extraction broke or nearly broke, both found by mutation:

- the shared PE reader answers `null` where the relay builder's private copy
  returned a number, which would have turned its "node-gyp ignored --arch" error
  into a `TypeError`. Both callers now go through `describePeMachine`.
- the rebuild fixtures stage a script's co-located modules by walking its
  imports, and the walker only understood `from '...'` — so the gate's new
  `require('./windows-pe-machine.cjs')` was left behind and every subprocess test
  failed with a resolution error, which is the exact failure its own comment
  warns about. It now follows `require` and bare side-effect `import` as well,
  and has tests; the fixture stages the gate by walking it rather than by naming
  one file.

Fixtures write real PE headers through one shared builder instead of three
hand-rolled ones.

* fix(windows): run the node-pty addon gates on the Windows job that can

`rebuild-native-deps-node-pty.test.mjs` carries four `skipIf(platform !== 'win32')`
tests. The full suite runs on ubuntu, and the Windows PR job runs an explicit
file list that never named this file -- so those tests were skipped on Linux and
never reached anywhere else. Three of them predate this branch. The Windows job
is added the four node-pty addon suites plus the module-walker one; the comment
above that list already says why it is the right place, which is that the addon
assertions only hold once natives have been rebuilt. Running the path-joining
suites there also covers the separator this gate's candidate list is built from.

The rest is round-three review:

- the rebuild-time arch assertion told a reader "node-gyp did not honour --arch"
  about a file that was not a PE image at all, which is a truncated or
  quarantined artifact and a different command to run. The two now read
  differently, and neither claims the other's cause. Same fix the packaged gate
  had one commit ago, in the place that had not had it yet.
- the missing-addon error said node-pty "would load" a prebuild without checking
  it is there. It says "fall through to" now, which is true either way.
- `isLoadableByArch` had no caller left once the packaged gate started needing
  the raw machine field for its message. Removed rather than kept warm.
- each candidate's header is read once instead of up to three times.
- the module walker's comment claimed every shape that reaches a co-located
  module; it does not follow `projectRequire`/`requireLocal`, and it must not --
  those specifiers resolve against the project root, so following one stages the
  wrong path and the copy fails. Proven by trying: widening the pattern to
  require-shaped names broke nine tests on
  `projectRequire('./config/scripts/...')`. The comment now says what it follows
  and why it stops there.
- a new test resolved a file URL with `.pathname`, which keeps the drive-letter
  slash on Windows -- the very job this commit adds it to.

* docs(windows): put the superseded export-only gate in the past tense

It describes what used to pass a broken addon, so present tense reads as a
description of the gate the same document then explains replacing it.

* fix(windows): repair what running the node-pty suites on Windows exposed

Putting these files on the Windows job turned four assertions red on the first
run. Three of them were in tests that carried `skipIf(platform !== 'win32')` and
had therefore never executed anywhere, on any branch.

- `writeFakeElectronRebuild` emitted the `windows-process-tree` addon a real
  rebuild leaves but never node-pty's, so every Windows test of the rebuild path
  ran against a tree no real rebuild can produce: node-pty "rebuilt" with nothing
  in `build/Release`. The new same-host check reads that state correctly and said
  so. The fake rebuild now writes `build/Release/conpty.node` when it was asked
  to rebuild node-pty for win32, with the marker and the target machine.
- `mkTempProject` never staged `windows-process-tree-creation-time.cjs`. The
  rebuild script reaches it through `projectRequire`, which resolves against the
  project root, so the module walker cannot follow it and must not try. Staged by
  name, with a comment saying which of the two it is. Without it the
  windows-process-tree probe failed to load its own checker and the module joined
  `modulesToRebuild`, which is the second and third red assertion.
- the two `nodePtyAddonPath` cases compared against a literal POSIX string.
  `resolve` returns a drive letter and backslashes on Windows, so they could only
  ever pass off it. Built from segments now, which still pins the `..` traversal
  that is the point of the test.

Verified on macOS: ensure-native-runtime-job-ownership,
verify-packaged-node-pty-job-ownership, windows-pe-machine,
script-module-dependencies, rebuild-native-deps-node-pty, rebuild-native-deps,
rebuild-native-deps-windows-process-tree, ensure-native-runtime -- 109 passed, 6
skipped. The 6 are the Windows-gated rebuild tests, which is the job this change
is aimed at; Windows CI is the arbiter.

* fix(windows): give the packaged fallback a third verdict, for a file that is no image

The packaged gate had two remedies for landing on the published prebuild and
picked between them on `!prebuilt`, which puts a truncated, empty or quarantined
`build/Release/conpty.node` in the cross-arch bucket: "the source build beside it
is the wrong architecture ... re-run with --arch". It is not the wrong
architecture, it is not an architecture, and `--arch` is not the command. The
rebuild-path gate was split for exactly this a commit ago; this is the same split
in the place that had not had it.

Also from review of the settled state:

- the stale-source-build branch ended in a call that happened to throw, so a
  reader could not see it was terminal and the file was read twice to get there.
  The verdict is now an Error the caller throws, built once from the read it
  already did, and shared with `assertCygwinBreakawayDenied` rather than copied.
- four injection seams had no consumer in production or in tests
  (`deniesBreakaway`, `peMachine`, and `exists`/`peMachine` on the rebuild
  verdict). An unused seam is a way for the tested path and the real one to drift
  apart; the tests drive both with real files. Removed.
- the loader table existed in a docblock and in the reference doc, already
  disagreeing about row four. The docblock cites the doc now.
- `peImage` stamped machine `0x0000` for an arch it had no value for, because
  `writeUInt16LE(undefined)` coerces to zero. A fixture that quietly invents the
  field the gates read is the same species of silent lie the gates exist to
  catch; it throws, and a test holds it to that.
- a test named for refusing an unreadable candidate asserted only that something
  threw. Renamed to what it proves.

* fix(windows): make the rebuild fixtures represent a tree that can exist

Second round of what running these suites on Windows exposed. The module the
walker could not stage is now staged, so the probe reached its own checker and
the real reasons surfaced:

- `writeFakeWindowsProcessTree` exported `{}`. The creation-time gate reads
  `supportedProcessDataFlags` off the addon and calls its absence "the tarball
  prebuilt, not a build of the patched source" — correctly. The fixture predates
  that gate and, being Windows-only, never met it. The healthy fake now reports
  the flag, taken from the gate's own constant. Two tests were failing on this,
  the second only because the module then joined `modulesToRebuild`.
- `rebuilds a loadable ConPTY native that lacks Orca job ownership` asked for a
  node-pty rebuild in a tree where node-pty had none of the payload its package
  ships. It gets `writeFakeNodePtyConptyPayload` like its two siblings.

I also tried making the fake rebuild emit `build/Release/conpty.node` the way a
real one does, and backed it out: `restoreNodePtyWindowsConptyRuntime` keys off
that file and then reads `third_party/conpty`, so emitting it in a tree without
the package payload turns one honest gap into an ENOENT two steps away. The
payload fixture is where "node-pty has its addon" belongs.

macOS: ensure-native-runtime-job-ownership, verify-packaged-node-pty-job-ownership,
windows-pe-machine, script-module-dependencies, rebuild-native-deps-node-pty,
rebuild-native-deps, rebuild-native-deps-windows-process-tree,
ensure-native-runtime — 112 passed, 6 skipped. The 6 are the Windows-gated
rebuild tests; Windows CI is the arbiter and is why they are on that job now.

* fix(windows): register the node-pty addon suites in the scope list too

Putting the five suites in the Windows lane's vitest argv gets them run once the
job starts; `WINDOWS_PACKAGE_TESTS` in `pr-code-change-scope.mjs` is what decides
whether the job starts at all. Only the argv was updated, so a PR touching just
`rebuild-native-deps-node-pty.test.mjs` would not have started the Windows job,
and its four Windows-only cases — including the same-host-absent one added here —
would have run on no machine for that PR. Exactly the shape of gap this branch is
about. Both lists now name all five, and `windows-pe-machine`,
`windows-pe-image-fixture` and `script-module-dependencies` join
`NATIVE_RUNTIME_PREFIXES` so a change to the modules themselves starts it too.

`win32-test-lane-registration.test.mjs` exists to catch precisely this and did
not, because its matcher only recognises suite-level gates (`describe.runIf` /
`describe.skipIf`) and a `.win32.` filename. These tests gate per `it`. Widening
it is not this branch's change to make: about thirty files across the repo carry
per-`it` Windows gates and are unregistered, so the ratchet would move far beyond
node-pty. Flagged rather than done.

Message repairs from the same review:

- the non-PE arm of the rebuild-time arch error read "... is not a PE image, so
  nothing can load it, so node-pty would fall back ...". The shared consequence
  clause already opens with ", so".
- the no-source-build packaging error ended "Package this Windows slice on such a
  host", which is wrong advice for the case where the host IS such a host and the
  rebuild simply left nothing — reachable when the artifact is removed before
  prune runs. It now names both readings and points at the beforeBuild output.
- the relay-addon builder blamed `--arch` for a build output that is not a PE at
  all, the same guess the node-pty gate was taught to stop making.
- the patch-drift assertion was a bare `toBe(true)`, so a real drift read as
  "expected false to be true". It now names the two things that can have drifted
  and what happens until they agree.
2026-09-16 22:23:30 -07:00
Neil 6c3b97b950 fix(mobile): a scope refusal is not a missing method on the Relay pairing probes (#19952)
* fix(mobile): a scope refusal is not a missing method on the Relay pairing probes

The desktop's mobile allowlist gate runs before its RPC dispatcher, so a method an
older desktop predates is absent from both and the phone is answered `forbidden`,
never `method_not_found`. Keying the "too old for Relay, stay on LAN" fallback on
`method_not_found` alone therefore never fired against the exact desktop it exists
for: first-time pairing threw instead of committing a LAN host.

`isPairingRelayRpcUnavailable` accepts both codes at the three pairing probe sites.
It is pairing-scoped on purpose - `isMethodNotFoundRefusal` has four other consumers
that must keep reading `forbidden` as a refusal, not as absence.

The main-side test pins the claim the fallback rests on: the dispatcher really does
answer `forbidden` to a mobile-scoped device and `method_not_found` to a runtime one,
and this build allowlists both probes, so `forbidden` on either can only mean an
older desktop.

* fix(mobile): leave a breadcrumb when a desktop refuses relay pairing

The LAN fallback now commits a host instead of throwing, so the refusal code
was the only record of why a phone ended up without a relay endpoint and
nothing wrote it down. Log it on the path that swallows it.

Narrow `isPairingRelayRpcUnavailable` to the two codes it matches rather than
to `RpcFailure`: a plain failure guard would collapse the *false* branch to
`RpcSuccess`, which a refusal carrying any other code still reaches.

Rename the `'method-not-found'` sentinel in the direct-upgrade reader, which
stopped describing what it covers, and correct two comments that named a
`method_not_found` mechanism the desktop cannot produce for these methods:
both probes have been allowlisted and registered by the same commit since
Relay landed, and an unwired pairing provider answers `runtime_error`.

* docs(wire): record that the mobile surface refuses by scope, not by absence

Two comments cited this page for "a scope refusal is not a missing method" and
the page did not say it — the only nearby statement says the opposite, because
it describes the runtime-scoped surface, where the dispatcher does answer
`method_not_found`. The allowlist gate makes the mobile surface the exception,
and the harness does not run that surface, so this note is the only record.

* docs(mobile): name the pairing site the scope refusal actually reached

The comments and the wire-compat note said this fixed first-time QR pairing.
It cannot: the `relay` block on the pairing offer, both RPC handlers and both
allowlist entries all landed in 77b154d5dd, so a desktop old enough to refuse
the probe also omits the offer block, and that flow commits a LAN host without
ever probing. The site that reached is `upgradeDirectMobileRelay`, which
re-probes every LAN-only host on reconnect: the refusal threw into the
controller's swallowing catch, so the write-once journal it had just written —
and the pending resume secret in it — was never retired.

Also drop two overclaims: the phone's Files and Git fallbacks have read both
codes since they shipped, so this is settled practice rather than a new rule,
and the reason the allowlisted-but-unregistered case cannot ship is
mobile-rpc-allowlist.test.ts, not a convention about what lands together.
2026-09-16 22:22:05 -07:00
Jinwoo Hong a28085adbf refactor(mobile): checked reply readers for the source-control domain (step 7 pilot) (#20950)
* test(mobile): ratchet the 201 unchecked RPC reply readers

Step 4 moved every call-site cast into an RpcOperation's `read`, but 201 of those
readers still answer `compatible: true` for any payload: `rpcUncheckedPayloadReader`
(163), `rpcReadUnchecked` (26 outside its own module) and `rpcUncheckedMemberReader`
(12), across 42 files. The cast moved; it did not become true.

Held as data with an AST boundary test, shaped on the raw-request-port ratchet: a file
that is not listed fails, a listed file that no longer has one fails, and a count that
rises fails. Only a call counts, so an import is not a reader and prose never is.

No behaviour change: this commit adds a list and a test.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): validate the source-control domain's RPC replies at arrival

Replaces all 17 unchecked readers in mobile/src/source-control/ with
`rpcResultVariant(variant, schema)`, so a malformed reply is an
`RpcIncompatibleReplyError` naming the operation instead of a TypeError three
frames downstream. The inventory drops 201 -> 184 and the five source-control
operations files leave it entirely.

This is a behaviour change, scoped to malformed replies. Six reply-matrix
goldens move; every named-scenario golden and every `normal` partition is
byte-identical, which is the parity claim.

Schemas live one module per reply domain, beside the operations that read them:
git-status, git-compare, git-history, hosted-review and worktree-metadata. A
member is required only where a consumer reads it unguarded, and each schema
records the consumer line that justifies it. Nothing is `.strict()`; every
reply a consumer publishes verbatim keeps `z.looseObject` so an undeclared host
member still passes through. Six replies have no reader anywhere in mobile and
get `z.unknown()`, which is the honest schema for them, not a holdout.

Three readers stay total by construction, because their contract is that an
unreadable reply is a value rather than an error: the `git.status` projection
(a null status three screens route on), the `session.tabs.list` reveal (a null
list means poll again) and the generated commit message (a screen's copy, never
a decode error in a text field). They gain the salvage report, not a verdict.

Consumers take the schema's output type, so `MobileGitStatusResult` and the
branch-compare aliases now name what mobile reads rather than the desktop
aggregate, and seven call-site casts are gone.

Three requirements came from the goldens, not from the host types:
`git.history` sends `timestamp: null`, `hostedReview.getCreationEligibility`
sends a `reviewLookupOutcome` the shared union does not list, and the
`git.status` projection writes an absent member as a present `undefined`.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): re-record the six source-control reply-matrix goldens step 7 moves

Six goldens, all on malformed partitions. Every named-scenario golden and every
`normal` partition is unchanged, which is the parity claim for this step.

  git.history-read / git.history#1
    result-absent, result-null, inner-ok-missing, inner-false-string-error,
    inner-false-object-error: the load rejected with a TypeError reading 'items'
    or 'map' off undefined/null; it now rejects with
    `incompatible_reply: git.history-page (git.history)`.

  hostedReview.eligibility + create-intent / hostedReview.getCreationEligibility
    result-absent, result-null, inner-ok-*: the fetch fulfilled with the error
    envelope itself, re-typed as an eligibility and published into the compose
    prefill; it now rejects, and both callers already route that to the same
    "eligibility unavailable" state a null answer produced.

  hostedReview.create-chain + create-intent / hostedReview.create
    result-absent, result-null, inner-ok-missing, inner-false-object-error: the
    create form showed the raw TypeError text "Cannot read properties of
    undefined (reading 'ok')"; it now shows the incompatible-reply message.

Every header digest is unchanged -- baseline, recorder, adapter, scenario and
lockfile all match -- so the diff is the behaviour and nothing else.

Recorded from this branch into a scratch directory and copied in, because there
is no scoped honest alternative: scripts/rpc-recording.mts refuses to run unless
the product tree equals the pinned baseline, and the README's remedy for an
intended behaviour change is to repin, which rewrites the `baseline` header of
all 667 goldens. So these six now carry a pin whose tree no longer produces
them. That is a real gap in the oracle's design for behaviour changes, not a
detail of this step, and it needs a decision before this lands.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): pin the four reply-schema properties the goldens found

Each of these cost a reply-matrix golden while writing the source-control
schemas, and none of them follows from reading the consumers or the host types:
a newer host's undeclared members must still decode, `git.history` sends
`timestamp: null`, `hostedReview.getCreationEligibility` sends a
`reviewLookupOutcome` the shared union does not list, and the `git.status`
projection writes an absent member as a present `undefined`.

The `.strict()` case is the one worth stating twice: at the top level it rejects
the reply, and on the entry it drops the row, which shows a dirty worktree an
empty Changes list. The fifth test pins the salvage report that makes such a
drop visible instead of silent.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): give an unreadable reply a message a user can read

`RpcIncompatibleReplyError` put `incompatible_reply: <op> (<method>)` in
`message`, and `message` is what the screens hand to a toast. Step 7 is the
first change that can reach this error at all, so the token would have shipped
to users as its own error copy.

Fixed at the boundary rather than per site: `message` is now plain copy, and the
machine token moved to `code` (`incompatible_reply`) and `name`
(`RpcIncompatibleReplyError`), both readable by callers. The cross-bundle
fallback in `isRpcIncompatibleReplyError` matched on the old message prefix, so
it now matches on `name`, which a foreign copy of the module still carries.

No existing test pinned the old text. Two new ones pin the copy, the token and
the foreign-copy match.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the recording baseline to this branch and re-record

Commit adeb5f9531 recorded the six moved goldens into a scratch directory and
copied them back, which left them pinned to `e7206f62`, a tree that no longer
produces them. That is the one claim the `baseline` header exists to make, so
this replaces it with the README's remedy done in full.

`baseline` is now f741b2ea82, the last commit on
this branch that touches a fenced path, so the recording fence passes in place
and every golden is pinned to the tree that produced it. All 667 were
re-recorded through `scripts/rpc-recording.mts --record`; none were hand-edited.

Decoding every value pool against the branch point b8d4cde09f sorts the corpus
into 661 header-only moves where `baseline` is the only key that moved, 6 whose
body moved as well, 0 added and 0 deleted. The 6 are the disclosed step-7 delta,
unchanged at 69 moved observation fields across malformed reply partitions, plus
the readable incompatible-reply copy from f741b2ea82. No `normal` partition and
no named-scenario golden moved.

`scenarioSha256` hashes the derived scenarios, not the manifest, so the repin
moves no other header key; the README section this adds records that, the
scratch-copy failure mode, and the follow-up repin main needs after a squash
merge.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): narrow the incompatible-reply error by instanceof, not by cast

The two new tests in f741b2ea82 read the error through `as` casts, which the
changed-code casting gate rejects. An `instanceof` guard narrows the same value
and checks the class at the same time.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the recording baseline to the branch tip and re-record

71d8c6a1e2 touched a fenced path (`mobile/src`), so the pin from 5f3f184fdf no
longer named the tree that produces these goldens. The fence compares the whole
of `mobile/src`, and a test file is inside it, so the pin follows the last commit
that touches a fenced path rather than the commit whose behaviour moved.

Re-recorded all 667 in place through `scripts/rpc-recording.mts --record`.
Decoding every value pool against the branch point b8d4cde09f still gives 661
header-only moves with `baseline` the only moved key, 6 body moves, 0 added and
0 deleted; the six and their 69 moved observation fields are unchanged.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): record the four source-control reads that had no oracle

git.status (host payload), git.branchCompare, git.commitCompare and
git.branchDiff were migrated to checked readers with no recording observing
them, so a required member a host omits would have surfaced only in production.

Three families mount the owners rather than the senders, because each reply is
only visible in what the owner then publishes: the Changes screen's loader hook
(git.status, and the base-ref chain and git.branchCompare it triggers), the
history list screen (git.history and the per-commit git.commitCompare), and the
committed-diff opener hook (git.branchDiff). Ten goldens: three pilot recordings
and seven reply matrices.

Two adapter capabilities this needed. An inert FlatList never calls `renderItem`,
so the history adapter renders one row through the screen's own callback, both to
reach the handler that expands a commit and to read the file list back; without
that the commit-compare reply changes nothing observable. And `lowlight` joins
`react` and `zod` as a real library rather than a refusing proxy, because the
branch diff highlights on its success arm before the preview reaches state, so
the shipped text arm was otherwise unrecordable. No golden recorded its absence,
so only `recorderSha256` moves.

Recording the same scenarios against 4b0009d414, the pre-refactor tree, is the
before column. Decoding every value pool across the two gives 11 body moves and
666 header-only, 0 added, 0 deleted: the 6 already disclosed, plus the 5 new
matrices at 63 moved observation fields. What moved is the point. A malformed
git.status used to leave Changes `ready` over the malformed payload and go on to
fetch a branch compare; it now says the host sent a reply it could not read. An
absent git.branchDiff result used to put "Cannot read properties of undefined
(reading 'kind')" on the screen. An unreadable git.commitCompare used to spin the
expanded commit forever; it now says "No file changes".

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the recording baseline to the merge commit and re-record

The merge is the last commit touching a fenced path, so it is the only tree
the recorder's fence can match. Every golden moves `baseline` and picks up
main's `recorderSha256` from #20920; the six the checked readers changed are
the only bodies that move against main.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): note the merge-commit pin and unwrap the recipe's record command

`format:check` from `mobile/` caught the wrapped inline command the recipe
had been carrying since it landed; pointing at the command above removes the
duplicate and the wrap together.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): open the source-control reply enums so a newer host's arm degrades

A closed `z.enum` in a reply schema is a version claim, and it refused replies
every declared reader could have rendered: a `git.branchCompare` summary status
of 'shallow-base' failed the whole Changes compare, a 'codeberg' provider failed
the whole eligibility, and a 'typechange' entry status dropped the row. Main
passed all three through.

`openEnum` in zod-salvage declares the arm set open: an unrecognised arm reads as
a member the consumers already handle, while absence and a non-string stay fatal.
Not `.catch()`, which would swallow those two as well.

`area` stays closed and says why: every arm grants stage, unstage or commit, so
there is no member to degrade to that would not offer an action against a row
this build cannot place. Main rendered such a row in no section either.

Also drops two claims the code does not back. Nothing reads the salvage report,
so the two comments promising a dropped entry "arrives as salvage.droppedPaths"
are gone, and `hostKind` on the non-text diff arm had no reader.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs: write down the open-enum rule and the header keys a branch moves

Rule 4 in the wire-compatibility page, beside the three rules it belongs with:
an enum arm set is a wire surface, unknown arms degrade rather than reject, and
leaving one closed is a decision to state where the schema is declared.

The recorder recipe's step 4 said `baseline` would be the only moved header key,
which is only true of a branch that never touched the recorder. It now names the
three digests a branch's own edits move, so a reader recognises a clean result.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): stop the recorder's own timeout killing a full re-record

The corpus records in ~110s warm and 160s under load, against a 120s budget, so
a full re-record was killed roughly half the time. A killed run wrote a partial
reporter banner and exited 1, which reads as a failing scenario rather than as a
run that never finished — it cost two investigations here. The budget is now ten
minutes, and a killed run says so.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the recording baseline to the open-enum commit and re-record

`baseline` is the only header key that moves and no golden body moves: no matrix
partition scripts an unknown enum arm, so the corpus cannot see this change. The
eight schema unit tests are its only oracle.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): stop an unresolvable eligibility claiming the branch is not ready

Both fallback prefills set `canCreate: false`, which is a determination nobody
made. It short-circuits getMobilePrCreateBlockMessage before reviewLookupOutcome
is read, so a malformed, refused or rejected eligibility told the user "This
branch is not ready for a pull request yet." instead of asking them to retry.
Dropping it leaves `canCreate` undefined, which is what "unproven" means here.
Only a host that determined `canCreate: false` still gets the blocked copy.

`area` now degrades to absent rather than staying closed. Dropping the row also
dropped it from the unresolved-conflict gate, which grants create on a conflicted
worktree; absent withholds stage, unstage and commit while keeping the row, since
every area reader is an equality check. Its four consumers narrow explicitly: the
diff-review queue filters unplaceable rows, the opener withholds the route, and
the commit-failure prompt pins 'staged' where its own filter already did.

`git.branchCompare` entries are nullish, matching the `?? []` its consumers use.

Deletions: `MobileGitStatusProjection` and `uncheckedReaderCount` lose `export`,
the boundary test drops its dead inventory self-file (the AST counter finds zero
calls there, only prose), and `isRpcIncompatibleReplyError` is gone — it had no
caller in mobile, desktop or e2e.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* style(mobile): formatting and a thrown rejection in the round-2 tests

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the recording baseline to the round-2 tip and re-record

The round-2 eligibility fix is a behaviour change, so the corpus has to be
re-recorded at a pin that includes it. Four goldens move body: the two
create-intent eligibility matrices on every non-normal partition, and the two
prefill scenarios that lose the fallback's `canCreate: false`.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the recording baseline to the main merge and re-record

The merge is now the last commit touching a fenced path, so the corpus has to
carry its sha. No body moves against the pre-merge corpus: main's engine change
shifts `recorderSha256` on every golden and nothing else, and main's fifteen
step-6 goldens re-record byte-identical.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): admit the three unchecked readers #20954 landed

The ratchet is a ceiling against this branch adding readers, not a claim about
what main may land. #20954 brought `notification-stream-closed`,
`native-chat-session-page` and `terminal-buffer-cleared`, so the merge has to
raise those lines and say where they came from.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): repin the recording baseline to the inventory commit and re-record

The ratchet inventory is a fenced path, so admitting #20954's three readers
moved the fence head again. Baseline only; no body moves against the merge
re-record.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): send the host's own provider token back instead of a fallback

`provider` is not a member mobile only reads. The eligibility reply names it and
the create call returns it, so `openEnum(..., 'unsupported')` did not soften a
reading — it rewrote the bytes, and a host that had just named `codeberg` refused
its own provider as unsupported. The action-sheet Create path has no provider
gate, so nothing caught it.

Passes the token through as a string from the reply to the create params. The
allow-list that decides whether mobile may create stays supportsHostedReviewCreation(),
which already answers no for a token this build does not know; its parameter
widens to `string`, since answering for an unknown token is the whole job. The
worktree-link switch gains a default, which also fixes an older hole: an
unrecognised provider used to fall out of the switch as `undefined` params.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): pin the provider pass-through in the corpus

Repins to the provider fix and records `sc-create-intent-unlisted-provider`,
whose eligibility reply names `codeberg` and whose recorded `hostedReview.create`
params carry it back unchanged. Restoring the old enum fallback fails that
golden on `Request params mismatch: hostedReview.create#1` and nothing else.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-16 13:32:09 -04:00
Jinwoo Hong b997fcc77a fix(session-search): try phrase and AND routes for prose queries before OR (#20754)
* fix(session-search): try phrase and AND routes for prose queries before OR

An exact sentence pasted out of a transcript was not returned. The route
ladder only ran the phrase and AND rungs for a literal-looking query, so
prose fell straight to OR, where the sentence's common words filled the
candidate limit with recent sessions and the old session holding the
sentence never reached ranking.

The planner now carries a `phrase` candidate: the query's tokens in order
with stop words kept, which is what the sentence is actually indexed as.
The ladder runs phrase then AND over those tokens for every query of two
or more tokens. A one-token query still takes the rung only when it
looked literal. `incomplete` is reported by the rung that answered rather
than accumulated across every rung tried.

* fix(session-search): mark a snippet with the route that retrieved it

A phrase hit was highlighted with the OR expression over the stop-word
stripped terms, so an exact sentence rendered as scattered bold words with
its stop words plain. The snippet now uses the expression the route
matched by: one run for a phrase, every typed word for AND, the terms for
OR.

* fix(session-search): repair a prose phrase without dropping its stop words

Typo repair re-planned the query from `plan.body`, which prose has already
had its stop words removed from. `relay is droppng frames` therefore came
back as the plan for `relay dropping frames`, and the phrase rung searched
for a sentence nobody wrote: the transcript holds `relay is dropping
frames`, so the exact match fell through to AND.

The repair now maps over `plan.phrase`, the tokens as typed, and re-plans
from those. Only terms the body holds are offered to the corrector, so a
stop word is still never repaired, and the re-plan recomputes the body
from the corrected sentence exactly as before.
2026-09-16 12:27:05 -04:00
Jinwoo Hong 3e5eb0329a feat(cli): orca search over the agent session index (#20514)
* feat(cli): orca search over the agent session index

`orca search <query>` calls PR 5's `aiVault.searchSessions` over the CLI's
existing runtime RPC, against the host `--environment` / `--pairing-code`
selects and no other. `orca search --index-status` calls `aiVault.searchStatus`.
It is the proof the contract works with no panel.

Every flag maps onto a contract field and nothing else: `--scope`, `--fresh`,
`--limit`, `--cursor`, repeatable `--agent` and `--path`, `--since`, `--sort`,
`--debug`, `--json`. No fan-out, no merged output, no `--host`.

One command rather than a `search status` subcommand: the query is a bare
positional, so `orca search status` could not be told apart from searching for
the word "status". `--status` is unavailable because `orchestration task-list
--status <state>` already owns the name as a valued flag.

No new runtime capability. PR 5 decided an explicit `method_not_found` refusal
maps to `unavailable/no-service`, so reusing `createSessionSearchClient` gives
an old host a plain "this host runs no session search service" answer at exit 0
instead of a raw JSON-RPC error.

`CommandSpec.repeatableFlags` scopes repeatability per command, because
`--agent` must repeat for search and stay single-valued for `worktree create`.
`help.ts` sat exactly at max-lines, so `skills-command-flag-help.ts` becomes
`command-scoped-flag-help.ts` carrying both tables at the same call-site size.

* refactor(cli): drop the search type assertions main's casting gate now rejects

Main gained a `consistent-type-assertions: never` scan in the changed-code gate
after this branch was cut, and it reported twelve assertions in the new files.

The four in the argument parser were avoidable. `readEnum` now keeps the value
`find` returns, which already carries the narrow type, and the agent filter goes
through an `isAiVaultAgent` predicate over a `Set<string>` instead of widening
the agent tuple.

The test now narrows the printed envelope by shape and re-reads the printed
result through `AiVaultSearchResponseSchema`, so the JSON assertions are checked
rather than claimed, and the flag table is typed so its callback needs no cast.
One assertion is left, for the structural fake client, with the SAFETY rationale
AGENTS.md requires.

* fix(cli): sanitize host strings and scope pre-command repeatable flags

Route every host-supplied string the search formatter prints through the
escape stripper, and resolve the repeatable-flag set from the command
tokens ahead when a flag sits before the command.

* refactor(cli): resolve repeatable flag rules once per command

* fix(cli): clarify session search availability and SSH scope

* feat(cli): hide orca search until the settings toggle ships

`orca search` stays dispatchable but leaves every discovery surface: root
help, group help, unknown-command suggestions, and `agent-context --json`.
`buildAgentContext` did not filter hidden specs, so it also stops leaking
the hidden `terminal stop`.
2026-09-16 12:03:38 -04:00