Commit Graph
12978 Commits
Author SHA1 Message Date
5cafefe726 Phase 3: every SSH host runs a managed Orca server (orcad), replacing the relay (#24863)
* Revert "revert: take the 26 Phase 3 (#16741 port) PRs back out of main (#24559)"

This reverts commit 5f308bfa9c.

* feat(orcad): Windows remote primitives for managed orcad hosts (W1) (#24525)

* feat(orcad): Windows remote primitives for managed orcad hosts (W1)

* refactor(orcad): run Windows host ops as node.exe with plain argv, no PowerShell hop

* fix(orcad): refuse secret-shaped names on the breakaway launcher's --env

* fix(orcad): name the secret env guard for its role

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(orcad): stage and commit a dormant migration catalog on the managed server (#16741 T6-9) (#24521)

* feat(orcad): stage and commit a dormant migration catalog on the managed server (#16741 T6-9)

The destination half of a catalog migration: an orcad stages a T6-7 manifest
(repositories, project groups, folder workspaces, dormant session, client,
automation and worktree metadata, retired names, scrollback snapshots) with
exclusive claims, then commits it with a receipt so a retried commit returns
the same receipt and never imports twice. Served as orcad.migration.* runtime
RPC behind the orcad.migration-catalog.v1 capability; the client refuses a
host without it or with method-not-found, and any other failure is left for
the caller to recheck. Dormant only: no live PTY projection. Inert on the
desktop until T6-10.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(rpc): catalog the orcad.migration params in the shared contract; name the catalog-import install target

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(ssh): journal and fence an SSH host for dormant migration, gated on proven terminal exit (#16741 T8-c1+c2) (#24522)

A migration from a relay-hosted SSH target into a managed orcad now starts with
a journal in its own sidecar directory, then the target's managed-owner fence,
then a profile flush, before any remote call. A fence with no journal is
unverifiable and never released; a journal whose fence is gone is stale and
grants nothing; an unreadable journal fails closed. The fence requires every
terminal the target ever leased to be proven exited, checked before the fence
(with the relay's process list) and again under it. Same-owner claims now need
the durable record that explains them. Inert until T8-c4/T6-10.

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(orcad): run orcad itself on Windows hosts (W2) (#24529)

* feat(orcad): run orcad itself on Windows hosts (W2)

* test(orcad): load the ConPTY smoke's addon from out/orcad so the temp slot can be removed

* test(orcad): skip the foreign-uid lock case when running as root

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): a connected but unused SSH host previews as movable (#24609)

The untransferred-dependency census counted activeConnectionIdsAtShutdown naming the
target as workspace-session state. The renderer rewrites that list on every connection
change, so merely connecting to an empty host blocked the move. It is a reconnect hint;
the remote work it can stand for is counted on its own. The empty-target claim check
likewise ignores global-field copies inside the host's session partition.

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(ssh): stage, commit and abort a dormant migration against its destination (#16741 T8-c3) (#24523)

* feat(ssh): stage, commit and abort a dormant migration against its destination (#16741 T8-c3)

The coordinator re-checks before every stage and commit that the fenced source
still exports the journaled manifest, carries no untransferable state and
started no terminal. A lost answer is re-read from the destination's catalog
state; only a committed read whose receipt matches the journal advances it,
and the journal is on disk before anything returns. Abort releases the fence
only on proof the destination holds nothing, or on an unsupported destination
before anything was staged, and never once the destination committed. Codes
against T6-9's catalog client through an injected interface. Inert until
T8-c4/T6-10.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor(ssh): import the T6-9 client's unsupported refusal instead of mirroring it

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(orcad): deploy, activate and roll back orcad on Windows SSH hosts (W3) (#24563)

* feat(orcad): deploy, activate and roll back orcad on Windows SSH hosts (W3)

* test(orcad): exhaustive op switch in the Windows lifecycle fake

* test(ssh): narrow the Windows host-cell descriptor by lane before building a relay cell

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(serve): run orca serve on the local orcad slot behind ORCA_SERVE_RUNTIME=orcad (T6-11) (#24608)

* feat(serve): run orca serve on the local orcad slot behind ORCA_SERVE_RUNTIME=orcad (T6-11)

* fix(serve): keep orcad selection app-side and wait out Windows temp cleanup

* refactor(orcad): move the data-root privacy check out of the instance lock

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* test(serve): prove D7 and the profile lock across a real Electron/orcad serve switch (#24619)

* test(serve): prove D7 and the profile lock across a real Electron/orcad serve switch

* ci(e2e): install ripgrep for the serve mode-switch job's window-manager wait

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(ssh): convert an SSH host with Orca state into a managed server through a journaled migration (#16741 T8-c4) (#24562)

* feat(ssh): convert an SSH host with Orca state into a managed server through a journaled migration (#16741 T8-c4)

The conversion entry resumes or takes the fence, deploys and pairs the
managed server into it, marks the server as migrated, then stages and commits
the dormant catalog. Every step is keyed by the journal, so a repeat after a
crash, deferral or lost reply resumes the same migration. Status reports an
unfinished migration, and rollback is refused while one runs or when the
rollback snapshot predates the migrated catalog. Inert until T6-10.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* style: oxfmt the c4 conversion and maintenance files

* fix(ssh): name the fake migration destination's type so declarations stay portable

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(orcad): decommission, managed stop and GC on Windows SSH hosts (W4) (#24570)

* feat(orcad): decommission, managed stop and GC on Windows SSH hosts (W4)

* fix(orcad): accept a managed stop request whose lock path is spelled with Windows client separators

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(ssh): retire a migrated SSH host's source state only after a proven commit (#16741 T8-c5) (#24565)

* feat(ssh): retire a migrated SSH host's source state only after a proven commit (#16741 T8-c5)

Once the journal records destination-committed, the source profile drops the
manifest's repositories, folder workspaces and unreferenced project groups,
its dormant session, automation, client and worktree state, and the leases
the fence proved exited. The profile flushes, the retirement is verified,
the journal moves to source-retired and compacts once the server matches.
A retry after any crash repeats idempotent work. The fenced target stays: it
carries the managed server's tunnel. Conversion now ends retired. Inert until
T6-10.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(orcad): retirement drops the migrated host from the reconnect hint

The census no longer treats activeConnectionIdsAtShutdown as untransferable (#24609), so
retirement must remove the target from it; otherwise a restart dials a host that is now a
managed server.

* style: oxfmt the c5 conversion file

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(orcad): convert Windows relay-hosted SSH targets to managed orcad (W5 part 1) (#24579)

* feat(orcad): convert Windows relay-hosted SSH targets to managed orcad (W5 part 1)

* test(orcad): start the Windows lane's exec spy after the relay gate prelude restores its own

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(settings): managed servers and "Move to managed server", behind an experimental setting (#16741 T6-6 + T6-10 UI) (#24590)

* feat(settings): managed servers and "Move to managed server", behind an experimental setting (#16741 T6-6 + T6-10 UI)

Adds a Managed servers section under Remote servers (deploy an empty server,
status with deferred-update and migration states, update, rollback, recover,
stop and cancel-stop, and SSH access for paired servers), and a Move to managed
server action on connected macOS and Linux SSH hosts with a preflight summary,
a terminals-closed confirmation and a resumable progress view. Main wires the
conversion to the relay's process list, the direct session and the T6-9 catalog
client. Everything is hidden until the new experimental setting is turned on,
and Windows SSH hosts are never offered. Merges the T6-9 branch (#24521) until
it lands on the integration branch.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(settings): align managed-server form controls and name the section the setting reveals

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(orcad): ask the relay with an absolute deadline via the W5 terminal-gate lister

The conversion wiring passed a relative 10 s as listProcesses' deadlineMs, which the
provider reads as an absolute time, so every relay inventory timed out after 1 ms and
the terminal gate could never prove exit. Adopt #24579's lister verbatim so the stacks
merge cleanly.

* feat(settings): name blocking saved state in plain, localized words

The move preview listed internal dependency ids such as workspace-session; each kind
now has its own catalog entry.

* feat(settings): offer managed servers and the move on Windows SSH hosts

W1-W5 are on the integration branch, so a Windows relay-hosted host can deploy, convert
and retire like a POSIX one. The move still waits for a connected relay that reported
its platform.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(ci): list the serve mode-switch e2e job in the release-cut permissions matrix (#24865)

* test(ci): list the serve mode-switch e2e job in the release-cut permissions matrix

* test(ci): expect the orcad Windows host cells in the SSH Windows hosts workflow

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* test(orcad): wait for killed terminal daemons to exit before removing their temp profiles (#24871)

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(updater): read rollout kill switches from the update-campaign payload, all inactive (#24867)

The nudge request Orca already polls may now carry an optional versioned rollout block
naming the Node runtime flips. A typed reader resolves each flip with kill-switch, version
range and install-id-bucketed percent semantics, falling back to the baked value (every flip
inactive) when the block is absent, invalid or never read. No consumer reads it yet.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(telemetry): report Windows security-software refusals and unverifiable or failed runtime checks (#24866)

ssh_remote_runtime_resolved dropped the self-test's security_software refusal to 'none' and
sent nothing when a self-test was unverifiable or failed, because those attempts throw before
a rung settles. Add the refusal value, self_test 'unverifiable', and an outcome field
(resolved | unverifiable | failed) deduplicated per host and outcome per session.

Co-authored-by: m4air <m4air@Mac.localdomain>

* refactor(settings): move WorktreeVisibilityDefaults into its own module (#24884)

global-settings-types.ts sits at the 300-line max-lines ceiling; merging main's two new
agent-state-rules settings with experimentalManagedServers put it at 301. The worktree
visibility defaults type moves next to the other visibility types and is re-exported so its
38 importers are unchanged.

Co-authored-by: m4air <m4air@Mac.localdomain>

* refactor(watcher): move the supervisor's child, terminating child and canary into a child slot (#24887)

* refactor(watcher): move the supervisor's child, terminating child and canary into a child slot

* fix(watcher,runtime): take the child slot's child type from the shared wrapper, and stub main's title-display clear in the projection test

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>

* ci(adhoc): build and ship the orcad template in adhoc macOS and Windows builds (#24969)

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>

* feat(orcad): bound the desktop slot cache and prune proven-stopped orcad versions after each managed deploy (#24973)

* feat(orcad): bound the desktop slot cache and prune proven-stopped orcad versions after each managed deploy

* fix(orcad): keep the in-use slot plus the two most recent others, and prove same-version reuse survives eviction

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>

* Phase 3: managed orcad on every SSH connect, with downgrade-safe fence (#24975, #24979)

* feat(orcad): fence managed hosts outside owner and keep converted projects for downgrades

Shipped builds hide any SSH target with an owner, so a downgrade made a converted host and its
projects vanish. The managed fence moves to a new orcadFence field (older builds keep but ignore
it), phase-3 owner fences migrate on load, and managed hosts stay visible but refuse a direct
relay. Conversion now stops at destination-committed with sourceRetainedAt; source retirement
waits on the baked-off orcad-source-retirement rollout flag. This build hides retained rows, and
a start that finds an older build changed them marks the host sourceChangedAt (relay, needs a
new move) instead of merging a second manifest.

* feat(ssh): every SSH host runs managed orcad, decided on each connect (#24979)

* feat(settings): managed servers are no longer experimental

Remove experimentalManagedServers and its gates; the Managed servers section always shows.
Loading drops a stored value, which an older build reads back as its own default (off).

* feat(ssh): every SSH host runs managed orcad, decided on each connect

Before a relay is started, the connect decides the host's server: a converted host connects
through its tunnel (retiring a retained source once orcad-source-retirement is on), an empty
host deploys orcad, and a host with Orca state converts through the journaled migration. A host
with live or unproven relay terminals keeps the relay this session and converts later. A
refusal for any other reason keeps the relay and names the blocker. A host orcad can't run on
(unsupported target, no template, native preflight, runtime self-test) releases any claim or
fence it took, records why with this app version, and keeps the pinned-relay ladder.

Progress and the decision ride an optional SshConnectionState.managedServer field (dropped by
older clients' admission). SSH Hosts shows each host's server status; the manual move dialog
is removed.

* fix(ssh): always await the connect's server decision, after the provider authority rotates

The decision now runs after the old session and transport are torn down and the authority has
rotated synchronously, so concurrent connects still join one attempt; a shutdown that began
during the decision wins over the rotation. The IPC tests use an async double, no sync branch.

* test(renderer): the IPC events store double carries its SSH connection states

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>

* test(ssh): prove connect-time conversion to managed orcad on real Linux and Windows hosts (#24981)

* test(ssh): prove connect-time conversion to managed orcad on real Linux and Windows hosts, and the downgrade view

* fix(e2e): read the SSH host's session partition, provision the convert cell's account, and keep rollout file overrides out of packaged builds

* test(e2e): report the full relay state when the convert poll times out

* test(e2e): require the relay to list no terminals before the converting connect

* test(e2e): require no running-terminal lease before the converting connect

* test(e2e): report the host's terminal leases when the converting connect keeps the relay

* test(e2e): end the relay era with no relay shell left to respawn

* test(e2e): settle before the converting connect and name the tabs a blocking shell belongs to

* test(e2e): use the exited relay tab as the session tab, since any mounted tab starts a shell

* test(e2e): carry an editor tab through the conversion instead of a terminal tab

* test(e2e): log the conversion census inputs before the converting connect

* fix(orcad): log which state blocked a refused conversion

* test(e2e): log both session partitions before the converting connect

* fix(orcad): a source partition's copy of focus on another host no longer blocks conversion

* fix(ssh): an ssh2 forward on port 0 reports the port it bound, so managed tunnels pair

* test(e2e): give the conversion its full budget again

* test(e2e): report the migration journal phase when the conversion stalls

* test(e2e): report the connect's own result and main's state when the conversion stalls

* fix(ssh): a converted host's managed state reaches the renderer instead of staying on 'converting'

* test(e2e): print a failed server call's response

* test(e2e): give server calls the budget a fresh server's first inventory needs

* test(e2e): read the converted worktree's tabs with a scoped session.tabs.list

* test(e2e): log the converted worktree's tabs instead of asserting them, pending the server-side fix

* test(e2e): prove retirement by the dropped source rows; the journal compacts away after it

* test(e2e): drop the conversion diagnostics now the cell passes

* test(e2e): fail a hung disconnect or connect with main's state instead of the whole budget

* test(e2e): convert an upgraded relay-era profile's host on its first connect, on Docker and Windows

* test(e2e): seed the relay-era target the way addTarget registers it

* fix(ssh): a stale ssh2 forward drops a late connection instead of crashing main on 'Not connected'

* fix(orcad): the active-slot readiness probe reads Windows hosts through the host script

* ci(ssh-windows): let only the convert cell's account open the SSH local forward its managed server needs

* test(e2e): match server paths in their JSON-escaped form, for Windows backslashes

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(orcad): move what an older build added to a converted host, or keep the server's version (#24980)

* feat(orcad): move what an older build added to a converted host, or keep the server's version

A host an older build changed stays on the relay with two actions. 'Move the new projects' runs a
fresh journaled conversion of only what the host's earlier migrations didn't move: the source is
viewed with those migrations retired from it, so nothing that overlaps the server is merged, and
a row the server already holds fails the whole move at stage, before any commit. Its journal
supersedes the chain head; retained-source checks compare against its baseline, and retirement
retires every manifest in the chain only after the newest committed. 'Keep the server's version'
records the current source as the baseline and returns the host to its managed server.

* test(ipc): the runtime environment handler contract lists the delta-move channels

* test(native-chat): snapshot the journal directory after the attach's restart-offer lock is released

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): converted hosts keep their editor tabs, forwarding-refusing hosts stay on the relay, listAll settles (#25099)

* fix(orcad): publish the headless graph so session.tabs.listAll settles instead of hanging

* fix(ssh): a system SSH forward on port 0 picks a free port first and reports it

* fix(runtime): a headless host lists and closes the editor tabs its session holds, so migrated editors reach clients

* fix(ssh): keep a host that refuses TCP forwarding on the relay, and release a conversion it stranded

* test(e2e): assert the migrated editor tab and a settled listAll, and keep a forwarding-refusing host on the relay

* fix: restore the journal import after rebase, type the probe's failure code, and update headless-graph test seams

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): client focus and 'local'-stamped tabs no longer refuse a host's move to its managed server (#25110)

* fix(orcad): client focus and 'local'-stamped tabs no longer refuse a host's move to its managed server

* fix(orcad): a v1.4.218 profile focused on the SSH worktree converts, and a refusal names what blocks it

The debounced session writer never patched activeWorkspaceKey or activeWorkspaceExecutionHostId,
so the first focus stayed on disk. The all-dependency census also counted global focus copies in
every non-source partition.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(ssh): offer to move a host whose open terminals keep it on the relay (#25097)

A host with any open SSH terminal never converted: the connect gate keeps the relay while relay
terminals are live, and an open tab respawns them on every connect. The first such connect per
host per app version now marks the relay status with offerMove (recorded as
managedServerMoveOffered), which toasts "Move <host> ... Its N open terminals will restart." The
SSH Hosts status line keeps a "Move to managed server" action while terminals are live.

Confirming runs ssh:moveToManagedServer: stop the host's relay terminals through the extracted
ssh:terminateSessions path, re-run the connect gate's terminal census, refuse on anything but
exited, then reconnect so the connect-time decision runs the journaled conversion.

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(orcad): reach managed orcad over an SSH stdio bridge where sshd refuses port forwarding (#25120)

* feat(orcad): reach managed orcad over an SSH stdio bridge where sshd refuses port forwarding

Hosts with AllowTcpForwarding no were kept on the relay. The managed tunnel now
probes forwarding each time it starts and, on refusal, serves the same local port
through a second provider: each accepted socket opens one SSH exec channel running
a small bridge on the host's pinned Node, which dials orcad's loopback port.
POSIX hosts run it with node -e; Windows hosts run it as the content-addressed
host script's stdio-bridge op with base64 line framing. Bridges are capped at 8
per connection under sshd's MaxSessions default, and a lost channel only drops its
socket. Only a host where even the bridge cannot run keeps the relay, recorded as
ssh_tunnel_unavailable; the older tcp_forwarding_refused record is retried.

* test(e2e): prove the stdio bridge by refused forwarding plus a working call

* test(e2e): connect the refused-forwarding host without the relay-only repo step

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(ssh): report each connect's server decision, and show orcad.log's last lines when setup fails (#25118)

* telemetry: one enum-only event per connect decision (outcome, reason, tunnel transport, host platform, duration), plus conversion start/commit/fail, deploy failures, and the per-host move offer and its result
* errors: deploy, launch and rollback failures carry the last 40 redacted lines of the host's orcad.log, read over SSH (Windows through the node.exe host script)
* SSH Hosts: a deferred or failed setup offers its reason, log tail included, under Details

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(startup): Windows never crashes resolving userData when roaming AppData is unavailable (#25113)

* fix(startup): pin Windows appData and userData before anything resolves them

A Windows session without a loaded profile (e.g. orca serve over SSH) can fail the
roaming AppData known-folder lookup. Electron 43 then falls through to Chromium's
userData provider and crashes natively. Resolve appData first (falling back to
APPDATA, then USERPROFILE\AppData\Roaming), and set userData explicitly so
Electron's provider never runs.

* test(startup): remove the AppData fixture through the retrying helper

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): keep a terminal the previous Orca version's relay still runs instead of replacing it (#25124)

After an app update the new relay answers "not found" for a PTY the previous build's relay still
runs, because the old relay refuses this build's handshake. The client read that as absence: it
expired the lease and the pane cold-restored into an empty shell while the user's shell kept
running, unreachable. Each deploy now takes a census of this target's older relay endpoints; while
one is live or unverifiable, a not-found reattach keeps the lease and the pane binding, and the
pane says the terminal is still running under the previous Orca version. A detached lease also
keeps blocking managed-server conversion until that terminal exits.

The cross-version harness now extracts src/relay, and a new test drives v1.4.218's relay socket and
grace lifecycle with this build's endpoint probe: the probe leaves no grace deadline and reads the
old relay as live work.

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(orcad): update a managed server on connect when it runs an older build and is idle (#25122)

* feat(orcad): update a managed server on connect when it runs an older build and is idle

A connect to a managed SSH host now runs the Managed servers update when the host's
orcad differs from this app's bundled build, the template carries the host's target,
and the update planner finds no live or uncounted terminals. A rejected candidate is
restored through the activation journal; the reason is recorded per app version so
later connects don't retry it. A host a newer Orca activated is never downgraded:
the activation record now names the app version behind each build, and an explicit
rollback holds the build it left.

* test(e2e): connect without a racing disconnect after relaunch, and report each attempt

On launch the app already reaches the managed server through its tunnel; a disconnect racing that
restore cancelled the connect that runs the update.

* feat(orcad): run the update check when the launch restores a managed server's tunnel

An auto-restored host may never see an SSH connect, so its server would never update. The tunnel
restore now runs the same check, once per server per session and off the caller's path, through
the shared update-check module; a server mid-migration is left alone.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): retry the pinned Node download through transient network failures (#25154)

A Chromium network change (net::ERR_NETWORK_CHANGED) during the pinned Node
download failed the deploy and sent the host back to the relay. The archive
download now retries up to three more times, after 1s, 3s and 9s, on dropped
connections, timeouts, stalls and retryable HTTP statuses, removing the partial
file first; checksum mismatches, other HTTP errors and cancels stay final. The
transient-error classifier moves from the speech download to
src/main/network/transient-download-error.ts so both share it.

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(serve): run orca serve on orcad by default, with ORCA_SERVE_RUNTIME=electron as the opt-out (#24972)

* feat(serve): run orca serve on orcad by default, with ORCA_SERVE_RUNTIME=electron as the opt-out

* test(serve): read Electron serve's pretty-printed readiness in the CLI mode-switch e2e

* feat(serve): gate D7 on Windows and serve on orcad there by default

* test(serve): take the profile lock in the CLI mode-switch e2e, and keep Windows profile logs on failure

* test(serve): tell Electron and orcad serve apart by readiness health, and trace Windows startup

* ci(e2e): dump Electron's native log and stack on the Windows serve mode-switch job

* fix(serve): keep Windows on Electron serve until it can adopt orcad's daemon

The Windows D7 job shows Electron serve exiting before its window when it relaunches
onto a terminal daemon orcad forked. Restore the win32 fallback and skip that case
there as a known gap; the follow-up PR fixes it and re-flips Windows.

* fix(serve): let ORCA_SERVE_RUNTIME=orcad opt in on Windows while Electron stays the default

* test(serve): skip the Windows D7 cases where orcad forks the daemon, and stop cleanup hiding a failed relaunch

Test 3 hits the same Windows gap as test 2: orcad forks its own daemon there, and Electron
crashes at startup beside it. A failed relaunch also made dispose close the old, already
closed app, whose throw replaced the launch error.

* test(serve): run every Windows D7 case, with the isolated home's AppData in place

Electron 43 crashes natively (0xFFFF7003) when it resolves userData and Windows cannot find
roaming AppData. The e2e home isolation points USERPROFILE at a fresh folder with no AppData,
so later launches hit that. The harness now creates it, every D7 case runs on Windows again,
and a new case proves Electron serve starts beside another profile's live orcad daemon.

* test(serve): retry removing a Windows e2e profile while a killed daemon releases it

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(ssh): resume a terminal the previous Orca version's relay still runs, through that relay's own bridge (#25170)

* feat(ssh): resume a terminal the previous Orca version's relay still runs, through that relay's own bridge

After an app update the previous relay keeps the user's shells alive but refuses this build's
handshake. Its own relay.js --connect, run from its own version directory, presents its own bundle
hash, so on POSIX hosts the client now reaches it that way: a pane whose reattach the current relay
held for an older relay opens a route through the old bridge, takes the PTY owner role without
output flow control, reattaches the PTY with its replay, and routes every later operation on that
id to the old relay. When the last pane a route serves exits, the route hangs up and the old relay's
own idle grace retires it. Windows hosts, relocated short sockets and unreachable bridges keep the
held-pane behaviour.

The cross-version harness now builds v1.4.218's relay from its tagged sources and runs it as the
real detached daemon: a shipped client leaves a shell in it, this build resumes the pane through the
old bridge, types into it, sees its output, and watches the old relay exit on its own after the
shell does.

* fix(ssh): read the legacy relay router through its instance

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): a host re-upgraded after a downgrade can move what the older build added (#25179)

The delta view left the moved projects' session state in place whenever it held anything
unmovable, so it then counted against the delta. Relay PTY bindings, shutdown markers and the
relay consumer's recovery record blocked every move although the terminal gate already proves
those terminals exited before any move commits; the move now drops them instead.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(worktree-ps): report a lost-contact host's terminals as unverifiable, not live:0 pty:no (#25167)

During a network drop to a relay-served SSH host, every PTY on it reads as an
unconfirmed exit, so worktree ps skipped them and printed live:0 pty:no for a
terminal that was still running. The listing now counts terminals whose liveness
verdict is unverifiable into a new optional unverifiableTerminalCount, and the CLI
prints live:unverifiable pty:unverifiable (JSON carries the same word) instead of
zero. A host-confirmed exit still reads as zero.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(terminal): a held pane shows no client OS or shell, and the boundary doc says POSIX hosts resume it (#25194)

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): tunnel to the port managed orcad bound and verify it is ours (#25182)

When another Orca already listens on 6768, orcad binds a different port. The
tunnel kept forwarding to 6768, reached the other runtime, was rejected with
4001, and the connect still reported a managed server.

The tunnel now reads orcad's bound port from its active slot's readiness
(falling back to the persisted port for slots without one) and, after the
forward is up, proves the server answering is the paired runtime. On a
mismatch it re-reads the port once and fails with orcad_identity_mismatch
instead of reporting managed.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): a converted host shows only its managed server's rows, and its old editor tabs load there (#25193)

Retained relay-era project groups and the per-host SSH catalog now hide like repos and folders.
On the managed transition the renderer reloads server names, groups, folders and worktrees, then
drops the host's relay-era rows a local refresh would keep. A restored tab with no host stamp in
a workspace now owned by a managed server takes that server as owner, in place.

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude <noreply@anthropic.com>

* fix(orcad): ship the port-scan worker so managed servers detect workspace ports (#25197)

orcad resolves port-scan-command-worker-entry.js beside orcad.js, but the orcad
build never emitted it and ORCAD_ARTIFACTS never listed it, so every managed
server logged 'probe worker unavailable' and had no port detection. The build now
emits it with the other children, and the artifact list carries it, so it is
uploaded, hashed and covered by the template contract.

A new test bundles orcad and its children and fails when the bundle names a
worker or child entry the slot does not ship. Two existing gaps it found,
session-scanner-service-entry.js and wsl-transcript-fs-process-entry.js, are
listed as known and may only shrink.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): keep retrying when a reconnect loses the socket before the SSH banner (#25195)

After a network outage, a port forwarder or NAT can accept the TCP
connection and then close it before the server sends its banner. ssh2
reports that as 'Connection lost before handshake' with no errno, so the
reconnect ladder classified it as permanent and published 'error' with
no retry scheduled. Treat it as recoverable on the bounded ladder only;
the initial connect keeps its narrow classifier.

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(serve): serve on orcad by default on Windows too (#25162)

D7 now runs every case on Windows (orcad-serve-mode-switch-windows), including Electron
serve adopting a daemon orcad forked, so Windows no longer needs the Electron default.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): the terminal gate asks the relays before treating a detached terminal as running (#25200)

* fix(orcad): the terminal gate asks the relays before treating a detached terminal as running, and retiring a host drops its relay recovery record

* fix(ssh): earlier-relay census gaps and an unanswered relay stay unverifiable; asking the terminal gate changes nothing

- the gate is read-only; the conversion and delta move retire proven detached leases themselves
- an expired lease also needs every earlier-build relay to answer before it reads as exited
- a failed, input-less or truncated census marks its older relays unverifiable, so reattach holds
- a legacy relay route closes only when no attach or listing still awaits it
- a disposed session forgets the census it started

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(migration): recover interrupted conversions and delta moves, and trim dead migration code (#25226)

* refactor(migration): drop the unused full dependency census

* refactor(migration): one table for the routed UI fields a migration carries

* refactor(migration): share the session focus field list and fix stale claim comments

* fix(session): keep a fenced SSH host's source session partition intact through renderer saves

The renderer cannot see a fenced host's repos or folders, so hydration drops their
worktree rows; a save that still routes any row to ssh:<target> (a folder workspace
key stays valid) rewrote that partition without them. That changes the migration's
source manifest and loses the session a downgraded build reads back. Main now
ignores renderer writes to a fenced, unchanged host's partition.

* fix(migration): resume or back out unfinished conversions and delta moves, serialize delta moves, clean stale journals

- On connect, a registered server whose conversion never committed resumes the commit; a failure aborts it through the destination, unregisters the server and releases the fence.
- An interrupted delta move keeps its journal and gets its changed mark back; the next move resumes it from the journaled manifest or aborts it when the source has moved on.
- Delta moves run under the target lifecycle queue and re-check the head, so two concurrent moves cannot write two journal heads.
- Delta checks re-read the live source instead of a frozen copy.
- Stale journals from a stopped or never-registered server are cleared.
- Retiring a chain goes newest first and compacts only at the end.
- The journal schema tolerates fields a newer build adds.

* fix(migration): hide source rows only after commit, and retire what an older build added to a moved project

- Source rows and the renderer session guard share one rule: hidden once the server committed, shown while a migration is still in flight.
- A delta move a crash interrupted gets its changed mark back on the next start.
- Retirement removes worktree metadata and automations by moved-project scope, so an older build's additions no longer fail the leftover check forever; that check ignores rows of other migrations in the chain.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): survive a slow first daemon start and close review gaps (#25213)

- Give the terminal daemon 30 s to start on Windows and retry the spawn once before
  falling back, so a cold first deploy is not refused as daemonless.
- Say in the activation refusal that orcad.log holds the daemon's error and that the
  next connect retries.
- Managed stop completion now waits out daemon retirement plus the shutdown deadline.
- Publish the instance lock atomically and reclaim an abandoned torn lock.
- Fix the stop listener closing before it was defined on an already-present request.
- `orca serve`'s cache prune keeps the slot a running local orcad uses.
- A timed-out daemon retirement reopens admission once it is refused; a retirement that
  may have reached the daemon stays fenced.
- A headless host keeps an editor tab's unsaved draft unless the close is forced.
- Unverifiable terminals are attributed by the host's PTY record, like live ones.
- An explicit --user-data-dir is no longer overridden on Windows.
- Read Windows daemon process identity from the process table, not PowerShell.
- Remove the unread data-event incarnationId and daemon health runtime fields, share
  one process-alive and error-code check, and fold the websocket limits file back.

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(cli): stop, update, roll back and recover managed Orca servers from the CLI (#25201)

* feat(cli): stop, update, roll back and recover managed Orca servers from the CLI

orca environment status|update|rollback|recover|stop|cancel-stop call the same managed-server
actions as Settings > Managed servers, over new managedServer.* runtime RPC methods. The desktop
main process registers those actions; the runtime advertises managedServer.v1 only then, and the
CLI refuses on a runtime without it or one that answers method_not_found. stop requires --yes.

* fix(ipc): keep a missing managed-server selector a rejection, not a synchronous throw

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): Move to managed server converts the host it is offered for (#25196)

The move stops the relay terminals and passes the terminal census; with #25179 the conversion no
longer refuses on the stopped terminal's saved tab, layout leaf and pane incarnation. A new test
drives one live relay terminal through Move to a conversion whose manifest carries the tab without
its relay PTY, so it spawns a fresh shell on orcad.

ssh:terminateSessions also no longer records a not-found shutdown as terminated while an older
Orca build's relay may still run the PTY (#25124): that terminal is reported unverifiable and its
lease kept, so the move refuses instead of converting over a running shell.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): deploy the glibc 2.17 runtime on old-glibc hosts, and fall back to the pinned relay (#25180)

* fix(orcad): deploy the glibc 2.17 runtime on old-glibc hosts, and fall back to the pinned relay

Managed orcad picked its runtime from the host's libc flavour alone, so a CentOS 7 host
(glibc 2.17) got the default linux-x64-glibc Node, whose self-test fails there. The deploy
now picks the runtime by glibc the same way the relay ladder picks rung A or B, and the
host-side slot checks accept the compat target it ships.

When orcad still can't run, the host is recorded as such and the relay connect that
follows now runs the pinned-runtime ladder instead of defaulting to the host-Node relay,
so a host with no Node lands on rung B rather than failing.

The hostile-host lane deploys managed orcad on a fresh CentOS 7 host and asserts it runs on
the compat runtime with no default runtime uploaded.

* test(ssh): CentOS 7 cell asserts the compat runtime pick, then the refusal and relay rung B fallback

The compat template still ships the base @parcel/watcher binary, which needs GLIBCXX_3.4.20; CentOS 7's
libstdc++ stops at 3.4.19, so the candidate's preflight refuses. The cell now pins that end-to-end
behaviour: compat runtime picked and uploaded alone, refusal classified native_preflight, and the relay
that follows settles on rung B with no runtime setting.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* test(serve): prove D7 on an installed Windows app whose daemon runs from the relocated daemon-host (#24976)

* test(serve): prove D7 on an installed Windows app whose daemon runs from the relocated daemon-host

* test(serve): read the pre-switch scrollback best-effort and wait for it after reattach

* test(serve): opt the packaged Windows serve switch into orcad explicitly

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(migration): a retirement that fails after deleting rows resumes instead of reading as changed (#25234)

The chain head records sourceRetiringAt before any row is retired; the startup change check and the chain's own comparison skip a head that carries it.

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(orcad): managed orcad idles out after 15 minutes and is started again whenever it is down (#25121)

* feat(orcad): a managed orcad stops after 15 idle minutes and starts again on the next connect

A client-launched orcad now exits, like the relay, once no client, terminal,
working agent, staged migration or activation fence has been seen for 15
minutes. The exit is the normal graceful shutdown, which leaves the terminal
daemon running; the daemon retires only if it proves itself empty. A record in
the data root tells the next start (and its readiness health) that the stop
was an idle one rather than a crash.

On connect and after host resume, a fresh tunnel whose server does not answer
starts the activated slot under the activation fence, but only on a proven
exit, so a stopped server reads as not running rather than a failure.

* fix(orcad): keep orcad-entry under max-lines; idle e2e connects without a relay repo

* fix(orcad): deploy and rollback launches carry the managed idle-exit fence

The candidate launch in activation and the rollback launch built their
own launch spec without the activation root, so a freshly deployed orcad
never enabled idle exit; only the wake path did. The field is now required
on every launch spec, so the type system covers each launch site.

* feat(ssh): start a stopped managed orcad on connect, on restore and after resume

Once orcad stopped (idle, kill or host reboot), a connect still resolved
managed over a forward to a dead port and every call failed. Every connect
now checks the server behind its tunnel, as does a call through a restored
environment; a server proven stopped is started from its activated slot
under the activation fence, adopting a surviving daemon and its terminals.
The status line shows the start, and a start that fails keeps the host
managed with the reason and orcad.log's tail, never as a terminal verdict.

* fix(ssh): reuse a serving verdict only on the same SSH transport, for 5s

A reconnect right after a reboot was answered from the previous
transport's cached verdict, so the stopped server was never started.

* fix(ssh): key the serving verdict on the tunnel's remote port too

* test(ssh): a stopped server starts before the update counts its terminals

* feat(ssh): check serving at the bound port, and follow a restarted orcad to a new one

The serving check uses the port the tunnel forwards to (the one orcad bound).
A restart that binds a different port drops the forward and rebuilds it at
the new port, within the same ensure or on the explicit connect check. The
tunnel manager class moves to its own file to stay under max-lines.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a connect with no leases asks the host's relay endpoints before converting (#25223)

* fix(ssh): a connect with no leases and no relay session asks the host's relay endpoints before converting

* refactor(ssh): one isLiveSshPtyLease for every lease-liveness check

* fix(ssh): the host relay census asks a relay its PTYs before reading it as live work

An accepting relay whose holders or children the probe could not read (no lsof, an unrecognised
service child) read as live work, so a relay-era host that had exited its last shell never
converted. The relay's own bridge answers pty.listProcesses without the owner role.

* fix(ssh): the host relay census runs each relay's probe and bridge on the runtime it runs on

Pinned-ladder hosts often have no Node on PATH, so a PATH Node read every relay as unverifiable.
Each daemon's own argv names its pinned runtime, or the host Node a legacy relay started with.

* fix(ssh): an older relay's bridge runs on that relay's own runtime

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(orcad): build @parcel/watcher into the glibc 2.17 compat slot, so CentOS 7 runs managed orcad (#25199)

The compat target swapped in only node-pty, so it shipped the base target's upstream
watcher.node, which needs GLIBCXX_3.4.20; CentOS 7 stops at 3.4.19. orcad's preflight
refused it, and relay rung B lost file watching without saying so.

The compat slot now compiles @parcel/watcher from the package's own sources against the
pinned headers with the C++ runtime static, like node-pty. The slot gates (glibc 2.17
symbol floor, no shared libstdc++, N-API 8) and the smoke load cover it, the template
stages it into the compat target, and both orcad and relay rung B pick it up from there.

COMPAT_SLOT_ADDONS names a compat slot's own addons: a compat slot missing one fails
--require-slots, and a compat template target that would ship any native file without a
compat build fails the template build. The CentOS 7 cell expects an activated managed
server again, and every launched relay cell loads its watcher directly.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(phase3): close e2e routing holes, localize managed-server outcomes, trim harness and CI (#25206)

* fix(phase3): close e2e routing holes, localize managed-server outcomes, trim harness and CI

- Route the auto-convert spec into needs_build and route the new orcad/serve
  e2e helpers to the specs that use them, with routing tests.
- Run the auto-convert Docker lane only when routed; fold the missing-AppData
  check into the Windows mode-switch job and drop crash-hunt diagnostic env.
- Delta-move dialog keys its preview on the target id and cannot close or
  resubmit mid-move; Resume has an in-flight guard.
- Managed-server toasts and status lines show localized messages instead of
  raw codes or main-process English; add singular and zero-count wording.
- Settings style fixes (quiet Cancel, section header, labelled fields,
  progress labels); delete dead i18n keys and the unused previewConversion.
- Harness: kill the serve child on readiness timeout, share the isolated
  profile and spawn-until-ready helpers, hooks and a condition wait in the
  auto-convert spec, shared cross-version exec helper.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(phase3): localize the refusal toast, share serve readiness and liveness helpers, correct D7 docs

- The refusal toast no longer shows main's English blocker detail; the SSH
  Hosts status line keeps it under Details. A missing terminal count is no
  longer defaulted to 0.
- startOrcadServe uses spawnUntilReady, so a readiness timeout kills orcad;
  one isPidAlive replaces the spec's copy and the lock-holder loop.
- The port-6768 auto-convert test uses the shared hooks.
- The docs say the Windows D7 job checks rather than gates.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ci): give the idle-exit spec its app build and route the convert harness to it

Co-authored-by: m4air <m4air@Mac.localdomain>

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* ci(adhoc): pin every job of an adhoc build to the commit resolved at dispatch (#25311)

Each job read the requested branch name and checked out whatever it pointed at when that
job started. A push mid-run mixed commits: in run 37234548638 the glibc217 slot lane
built 2ddea8736e, then #25199 merged, and desktop_template checked out 70948d597f, whose
merge step requires the compat watcher that lane never built.

A first resolve-ref job (no secrets) resolves the branch, tag or full SHA once, and every
job checks out that commit. The mac job still vets it for reachability before signing;
the requested name stays the concurrency key and the release label.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(terminal): report a confirmed relay PTY exit as exited and give a cold Windows PTY probe more time (#25304)

A worktree terminal close returned as soon as the relay confirmed the stop, but the
exit frame reaches the runtime record only after the SSH output intake drains, so the
verdict read straight after saw a still-connected PTY and answered unverifiable. The
stop now waits, bounded by its deadline or 10 s, for that record.

The bundled runtime's PTY probe gives the first Windows spawn 15 s and retries once
after a timeout only; spawn errors and non-zero exits still fail at once.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(relay,orcad): keep new wire messages forward compatible, drop unshipped relay.reset, fix shutdown and serve gaps (#25298)

* fix(orcad-migration): keep migration wire forward compatible with newer peers

* docs(pty): record why pty.resumeClient negotiates by method-not-found

* fix(relay): remove the unused relay.reset method and keep a deferred shutdown serving

* refactor: drop dead hold API, release gate and migration pass-throughs; fix serve and delegation gaps

* fix(lint): keep reopen hooks within file size limits

* revert type dedupe in orcad-incumbent-recovery to avoid a parallel conflict

* fix(lint): name stop-reply request fields for their role

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* refactor(ssh): one managed-tunnel ownership check and forward bookkeeping (#25316)

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a stuck managed server can recover, cancels finish their change, unknown stays unknown; delete unused reset/drain code (#25227)

* refactor(ssh): delete the unused connection reset and drain machinery

* fix(ssh): surface a stuck managed server, finish activations past their first change, and fall back from Windows launch refusals

- A rejected build that changed profile state now refuses as orcad_recovery_changed_state
  (unverifiable) with a Recover path that restores the snapshot once the operator accepts;
  an interrupted activation is no longer reported as a quiet update deferral.
- Activation and rollback drop the abort signal after their first journaled change.
- Windows launch refusals and unsafe command lines send the host back to the relay.
- Fence refusals keep unverifiable blockers unverifiable; the POSIX liveness probe reads
  kill errors in the C locale and treats permission errors as unknown.
- A host with no relay fallback surfaces its real connect error; disconnect and removal
  clear the setting-up status, and a cancelled decision's progress is dropped.
- Remove tcp_forwarding_refused leftovers, dedupe incumbent stop/restore, exec-or-empty,
  census client, errorMessage and the blocking-blocker predicate; split activation and
  snapshot files under 300 lines.

* fix(ssh): one import per module in the rollback transition

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(migration): committed reads trust the receipt, conversions survive UI churn, journals are cached, and round-2 low items (#25327)

* fix(ssh): keep a newer build's fence and note fields, and validate a fence beside a legacy owner

Normalization now validates known fields but passes unknown ones through on orcadFence and the managed-server notes. A legacy managed owner next to a malformed fence falls back to the owner's environment id instead of keeping the malformed fence.

* fix(migration): delete a retired migration's scrollback files once retirement is durable

Retirement dropped the moved terminals' scrollback refs from session state but kept the files forever. After the journal records source-retired, the files the manifest names are deleted, except refs any session partition or pending export still names. A crash before the delete repeats it on the next retirement pass.

* fix(orcad): a committed migration reads as committed from its receipt, survives receipt eviction, and abandoned stages expire

- Committed-state reads trusted only a byte-identical copy of every moved row and snapshot, so a live server that had been used could never confirm its own commit. The receipt alone now proves it; full equality stays inside the commit.
- A receipt that ages out of the 64-entry list keeps a compact record, so the commit never reads as absent.
- A stage no client returns to within a week stops holding the server awake and is dropped at the next stage.

* fix(migration): freeze a host's session from the fence on, ignore UI churn in the frozen-source check, and cache parsed journals

- Renderer writes to a fenced host's ssh: partition are skipped from fence time, not only after commit, so tab work mid-conversion cannot change the source.
- The frozen-source check leaves out workspace session and client routing state, which the UI rewrites as the user works (including through the local partition and UI state); the server gets them as journaled.
- Parsed journals are cached per directory, keyed by each file's inode, size and mtime and dropped on every write and remove, so list and session calls stop re-parsing every manifest.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): an older relay a census cannot rule out stays held, and a reconnect resumes held PTYs through it (#25331)

- A census that does not know the host platform is unverifiable, never "no older relay". On Windows
  it now probes each older version directory's pipe for the target, the way relay GC does, so a
  live older Windows relay keeps its PTYs held instead of respawning their panes; those endpoints
  are held, never bridged.
- On reconnect, a PTY the current relay disowned while an older relay holds it is reattached through
  that relay's own bridge, with its runtime restored and its replay forwarded. When no route serves
  it, it is left for recovery like an exhausted reattach.
- A superseded or disposed deploy no longer starts the census, so it cannot replace the current
  attempt's.
- A route stops holding an unserved PTY that exits.
- Drop the unused censusPreviousRelays export.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(phase3): round-2 ui-infra review fixes (#25325)

- An unverifiable move refusal with no count says so instead of "0 terminals".
- One move per host: the dialog stays mounted while open, joins a run already
  in flight on remount, and the toast shares the same guard.
- A managed server's start, wake or update no longer reloads every host's
  catalog; only a new environment for the host loads, scoped to that host
  and the local catalog.
- A failed host-partition session write is no longer acknowledged as written,
  so the writer re-queues those fields.
- The deploy picker lists only hosts with no managed server or pending move.
- Harness: orcad and the released relay daemon are stopped when startup fails.
- Windows host CI: the convert cell always runs last and fails on a failed
  native switch or build; ssh-host-server unit tests no longer trigger it.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(relay,cli): never report an unreachable terminal as exited; keep PTYs through a deferred shutdown (#25328)

* fix(relay,cli): never read an unreachable relay or terminal as exited; keep PTYs on a deferred shutdown

* fix: stop an unrecorded breakaway child, route orcad serve through the spawn chokepoint, and drop review-flagged leftovers

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): reclaim a crashed desktop lock, close the stop-cancel race, drop dead daemon and RPC code (#25338)

* fix(orcad): reclaim a crashed desktop lock, close the stop-cancel race, drop dead daemon and RPC code

- A desktop reclaims a stale desktop lock record (Electron's lock proves it), and a real
  orcad hold shows a dialog instead of exiting silently.
- A failing quit handler no longer skips closing observability.
- Completion withdraws its request when a cancel lands mid-write; the listener removes a
  cancelled leftover.
- An idle stop's clean record is retracted when the shutdown fails or overruns.
- The managed-stop request tolerates unknown fields from a newer client.
- Shared error-code checks, one win32 coverage rule, no redundant isAlive filters.
- Remove the recovery-only daemon provider, requirePinnedWsPort/strictPort and the
  superseded orcad.migration.importCatalog RPC.

* fix(startup): keep main-process-preflight under the line cap

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): fail fast on a held fence, restore untouched rollbacks, atomic Windows host script, honest tunnel ensure (#25332)

* fix(ssh): fail fast on a held activation fence, restore a rollback the target never touched, stage the Windows host script atomically

- withOrcadActivationLock no longer waits up to 15 minutes inside the target lifecycle: a held
  fence answers at once (orcad_activation_recovery_required); deploy probes it before uploading.
- A rejected rollback target that left the restored snapshot untouched puts the newer build back
  unasked; only a real change waits for the operator. Crash recovery applies the same rule.
- The Windows host script is written only when missing, through a partial file and a rename; a
  bridge exit without a sentinel is unverifiable unless the shell could not find the command.
- The managed tunnel's ensure() throws when superseded, and a caller arriving after close()
  builds a fresh run instead of joining the doomed one.
- An unparseable stop answer after a stop was sent keeps the fence (new 'unconfirmed' outcome).
- stopRemote starts over instead of reporting live when another run settled the journal.
- Releasing an unreachable setup stops the orcad it activated and clears active on proven exit.
- A failed startup is torn down quietly so its error stays published; port-forward listeners
  keep an error handler; a linked SSH access connect is cancelled when its window closes.
- Shared errorMessage, one SFTP transfer helper, getConnectGeneration only, test-cell names.

* fix(ssh): stage the Windows host script through the pinned node.exe, which cmd.exe and PowerShell both run

* fix(ssh): host-script staging runs as host-script ops; a joiner of an overtaken tunnel run builds its own

- The presence check and the install are ops in the uploaded host script (script-present, and
  script-install run from the partial upload itself), invoked as node.exe <script> <op> <args>
  like every other Windows host op: no inline code.
- ensure() callers that joined an in-flight run no longer inherit its 'superseded' end; they
  start a fresh run, which also covers a close() in between.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a connect re-checks unverifiable relay terminals once its relay session can answer (#25385)

The server decision runs before any relay session exists, so a terminal this desktop left
detached could never be asked about and read unverifiable for as long as it ran. On Windows no
endpoint census can fill that gap. Once the session is up, the relay that holds the PTY answers
and a terminal it still runs is reported live.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(settings): a managed server whose status never loaded reads unknown, not "Not running" (#25389)

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(relay,runtime-env,cli): keep a deferred relay's AI Vault and skill uploads, make managed re-pair crash-safe, show unverifiable in worktree ps text (#25393)

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): keep a surviving daemon's slot through the cache prune, and drop an idle record a signal stop took over (#25387)

- The slot prune also protects any slot a live terminal daemon's PID record points into, and
  evicts nothing while a daemon record is unreadable.
- The shutdown trigger reports whether it took ownership; an idle stop that another source
  took over discards its clean-idle record.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): edit a managed host's connection, show rows no migration owns, drop the unused quit-drain predicate (#25403)

- SSH settings may edit a managed host's connection fields; its fence and generation are never written from the renderer, and its tunnel is closed so the next use redials. Removing it is refused with a pointer to Stop under Managed servers, which stops and removes the server; the renderer no longer ends the host's terminals before that refusal.
- A fenced host with no journal (an empty host's deploy, or one whose journal compacted after retirement) no longer hides its rows: no committed move owns them, so projects an older build added stay visible. An unreadable journal still hides them.
- The quit drain's mayDetach predicate and disconnectAll's filter had no production caller and are removed.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): input to a terminal an older relay holds but no route serves is refused, not dropped (#25404)

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): refresh a changed host's status once its delta move or keep-server choice lands (#25416)

The status line kept saying the host was changed on an older Orca until a manual reconnect. Main now publishes the host as managed by its server after a successful move or keep, with a disconnected state when the move released the relay session.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a stopped managed server stops showing on its SSH host without a reconnect (#25409)

* fix(ssh): a stopped managed server stops showing on its SSH host without a reconnect

On unlink, forget the host's managed-server decision and republish its connection state. The
republish also drops the host's cached worktree scans, so worktree.ps stops naming the removed
server.

* fix(runtime-environments): a removed server's workspace session partition goes with it

Host-scoped listings enumerate session partitions as known hosts, so worktree ps kept naming a
stopped or removed server as an omitted, unselectable host. Unlinking or removing a server now
drops its runtime:<id> partition.

* test: give the removal-storage fake store casts their SAFETY rationale

* fix(runtime-environments): drop orphaned runtime workspace sessions at startup

A crash between unlinking a server and dropping its session, or a build that unlinked before the
drop existed, left a runtime:<id> session listings name as an unselectable host. Startup now drops
runtime sessions whose server is not in the environment store; it never touches local or ssh
sessions, and skips entirely when that store is missing or unreadable.

* test(migration): retirement re-aims focus without creating the destination's session

Locks in the ordering the startup reconcile relies on: a runtime:<id> session is never written
before that server is registered.

* docs(runtime-environments): name the ordering invariant the startup reconcile relies on

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a connect whose relays prove its leases ended retires them, so the next connect converts (#25405)

Every connect decides before a relay session exists, so a detached or expired lease reads
unverifiable there. The post-session re-check proved such leases ended but left them in place, so
the host stayed unverifiable on every connect and each respawned pane added another lease.

Co-authored-by: m4air <m4air@Mac.localdomain>

* ci(ssh-windows): check out preload and renderer for the orcad-convert e2e build

Main's #25359 narrowed this workflow's checkout to the server and test trees. Phase 3's
orcad-convert cell builds the full e2e app with electron-vite, which also needs src/preload
and src/renderer, so the x64 inbox cell failed with UNRESOLVED_ENTRY.

* test(orcad): model a really converted host in the v1.4.218 downgrade wire test

#25403 shows a fenced host's rows when no migration journal explains the fence. The downgrade
test fenced the host without a journal, so this build showed the retained project it is meant
to hide. Write the destination-committed cutover journal a real conversion leaves.

* fix(ssh): a briefly held fence waits and is retried, not recorded as a failed update; terminate keeps held leases (#25420)

- The activation fence is retried for a few seconds. A fence still held answers
  orcad_activation_fence_busy (a waiting deferral, never an update failure) unless a journal or
  a lock past its stale age shows an interrupted run, which stays recovery-required.
- acquireInstallLock reports Busy only when a holder answered; a lock command that only ever
  failed surfaces its own error.
- Terminate detaches instead of disposing when any PTY was unverifiable, so the final teardown
  never bulk-marks a lease an older relay may hold as terminated.
- A wake whose connection dropped while holding the fence releases that fence on this client's
  next wake (no journal, slot proven exited), so a relaunch-then-connect is not left fenced.
- The orcad e2e reconnect helper surfaces a connect's error text instead of a JSON parse error.

Co-authored-by: m4air <m4air@Mac.localdomain>

* ci(ssh-windows): check out all of src for the orcad-convert cell

Its e2e global setup also compiles the bundled CLI from src/cli, which the narrowed checkout
left out.

* fix(orcad): idle stop — drop the record when a signal stop wins, and read the activation lock, not its root (#25464)

* fix(orcad): drop the idle-stop record when a signal stop finishes first

Every stop's cleanup now discards the record unless the idle trigger owns the stop, so a
takeover that exits before the idle request runs no longer leaves a false idle stop.

* fix(orcad): idle check reads the activation lock, not the transaction root

An interrupted acquire can leave the root without a lock; the client already treats that as
unfenced, and managed orcad now does too instead of never idling out.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): main refuses a managed server host's removal before ending its terminals (#25465)

The remove flow skipped ending terminals for a managed host only when the renderer's cached target list already showed the fence, so a fence that landed after the list loaded still lost the host's terminals before main refused the removal. The removal's terminate call now carries forRemoval, and main refuses it for a managed host with the Stop… message before touching anything; the renderer no longer decides.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(terminal): closing an SSH workspace counts a host-confirmed terminal exit as stopped (#25479)

* fix(terminal): a workspace close counts an SSH terminal's confirmed exit as stopped

The close's verdict treated any SSH terminal whose record was still present as
unconfirmed, even when that record held a host-confirmed exit. SSH records outlive
their exit, so every successful close of a relay terminal answered unverifiable.
Also, a relay reattach that finishes registering after the stop no longer revives an
incarnation whose exit is already recorded.

* test(terminal): cover a reattached relay exit confirmed without an incarnation

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(cli,ssh): managed-server actions outwait their own deadlines; cap the unverifiable serving detail (#25473)

orca environment update/rollback/recover/stop/cancel-stop waited the 60 s RPC default while the
desktop runs the whole action inline, so a slow host printed a timeout failure for an action that
kept going. They now wait 20 minutes, and a timeout says the action may still be running and points
at `orca environment status` instead of reporting failure. status keeps the default.

The retained managed-server state now clamps serving.detail to the same byte limit as its sibling
detail fields.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): keep the previous orcad.log on Windows across a restart (#25481)

The Windows breakaway launcher's addon recreates orcad.log on every start, while POSIX
appends, so a crash's log was gone once orcad restarted. The orcad launch now asks the
launcher to move the last run's log to orcad.log.1 first, capped to its last 1 MiB.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(sidebar): a stopped managed server leaves no empty project group behind (#25488)

The removed-runtime purge dropped the server's repos, setups and worktree rows but kept the project
groups and folder workspaces the renderer fetched from it, so an empty heading stayed in the
sidebar. It now drops those runtime-stamped rows too.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): run automations on orcad and keep a managed host up while they can fire (#25475)

* fix(orcad): run automations on orcad and keep a managed host up while they can fire

orcad never built an AutomationService, so with orca serve defaulting to orcad scheduled
runs never dispatched and Run now threw runtime_unavailable. The headless service setup
moves out of Electron startup into automations/runtime-automation-service.ts; orcad
installs, binds, starts and stops it, and managed idle exit counts an enabled schedule or
an unsettled run as busy.

* docs(orcad): list automations among the idle-exit conditions

* fix(automations): precheck reads the SSH manager from its registry, keeping electron out of orcad

The precheck imported getSshConnectionManager through ipc/ssh, which pulled 32 electron
modules and node:sqlite into the orcad bundle and failed build:orcad.

* test(orcad): stub the automation wiring in the push-startup runtime harness

That harness stubs OrcaRuntimeService without an automation surface, so starting the real
service threw setAutomationService is not a function and timed out the next test.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): Move to managed server survives a relay that hangs up on its last terminal exit (#25487)

BUG-15: with two live terminals, one served through an older build's relay, Move failed with
"Failed to terminate SSH host sessions: …: Multiplexer disposed" although both shells died. The
old relay reports the exit before the shutdown reply; that exit closes the route (its last served
PTY), and the disposed mux rejected the shutdown still awaiting its reply.

- The legacy relay route settles a shutdown whose PTY exit it already observed.
- Move no longer aborts on a failed stop: the terminal census (what the conversion trusts) decides.
  Exited closes the relay session and converts; live or unverifiable refuses and republishes the
  relay status, so the stale terminal count is replaced.
- The move dialog offers Try again after a refusal or failure.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): closing a terminal a previous Orca relay runs stops it there and confirms the exit (#25471)

A stop on a PTY an older relay runs reached the current relay whenever no route served it at that
moment (after a reconnect with the pane unmounted, or a shell no pane ever resumed), and the current
relay answers a stop for an id it never minted as done, so the PTY read stopped while its shell
kept running. A stop on a served PTY failed instead: its exit arrived before the stop's reply,
closing the route under the pending request.

- A served PTY stops on its route; a PTY no route serves is stopped through a short-lived route to
  the older relay that lists it, which hangs up once that PTY exits.
- When an older relay may hold the PTY but cannot be asked (incomplete census, Windows pipe, a
  bridge that will not open), or its bridge drops mid-stop, the stop is unverifiable, never reported
  done; terminate keeps the lease.
- A route stays open until its in-flight requests settle.
- The provider's exit stream includes the exits older relays report, so a stop observes the PTY it
  stopped exit on the relay that ran it.

The cross-version harness runs what terminal close --all runs per PTY against a real v1.4.218 relay,
for a pane resumed this connection and for a shell no pane resumed, and sees the old shell exit there
and the old relay retire on its own grace.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a live run's journal reads busy, a wake releases only its own fence, a cancelled re-check never reports connected (#25470)

* fix(ssh): a live run's journal reads busy, a wake releases only its own fence, a cancelled re-check never reports connected

- A held fence asks for Recover only when its lock is stale or a journal has no fence over it;
  a journal under a fresh fence is a run still working and reads orcad_activation_fence_busy.
- A wake writes an owner token into the fence it takes and later releases only a fence carrying
  that token; observing the fence gone forgets it.
- The connect re-checks ownership after the relay-terminal re-check, and the re-check itself
  neither records a decision nor retires leases for a cancelled attempt.

* fix(ssh): a tunnel caller that joined a run a disconnect cancelled builds its own

The launch-time restore's tunnel run connects over SSH; a disconnect then cancels that connect.
An explicit connect that had joined the run inherited its SshConnectAttemptCancelledError and
failed (seen as the idle-exit e2e's 'connect threw: ... cancelled'). A joiner now builds once
anew after any end of the joined run except an auth failure, which it shares rather than prompt again.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(runtime-env): a late reply from a replaced pairing never overwrites the re-paired device identity (#25490)

* fix(runtime-env): a reply from a replaced pairing never overwrites the re-paired device identity

* fix(types): narrow identity fields in markEnvironmentUsed

* fix(ssh): managed tunnel proves its server by runtime id; SSH access linking keeps the strict device check

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(migration): retained scrollback survives a closed tab, local acknowledgements stop blocking, close-intent retirement retries, mobile selections merge per workspace (#25508)

- A transfer reads a retained snapshot straight from storage once its tab closes, and releasing the retention deletes a ref no session names; the frozen-source check leaves the snapshot list to the journaled manifest.
- Only acknowledgements on panes the source host owns count toward the ui-routing blocker.
- Retiring close intents treats an absent source with an identical destination entry as already done.
- Importing a device's mobile selections keeps its selections for other workspaces.

Co-authored-by: m4air <m4air@Mac.localdomain>

* test(ssh): an expired lease an older relay still lists is never retired (#25449)

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(migration): keep each host's state its own, count only saved commits, and never wedge connect on a partial session (#25516)

* fix(migration): a host-qualified owner key belongs only to its own host

Owner matching stripped a key's host qualifier before matching its repo id, so converting host A claimed, moved and on retirement removed host B's session state when the two hosts share a repo id, and destination-qualified focus written by retarget read as leftover source state, failing retirement with orcad_migration_source_ui_routing_reappeared. A qualified key now matches only its own host, and an unqualified key in another host's session partition belongs to that host.

* fix(migration): a partially written session partition never wedges connect, and a marker two partitions agree on stops blocking the move

Real-host BUG-14: the renderer's per-host snapshot leaves out maps a host has no rows in, and main stored host partitions exactly as sent, so a runtime partition lacked tabsByWorktree and the dormant-state collector threw on every connect. Main now fills the required maps on every host-partition write, the migration collectors tolerate a partition persisted without them, and an unreadable session blocks the move instead of failing connect.

The same profile's workspace-session blocker was a false positive: the local and host partitions both carried defaultTerminalTabsApplied for the moved worktree with the same value, and the fragment merge refused any shared worktree key. It now refuses only when the partitions disagree.

* fix(migration): only a flushed commit acknowledgement moves a migration to committed

A committed state read may come from a receipt the server holds in memory but failed to flush. A retry took that read as proof, journaled destination-committed and went on to retire the source, so a later server restart could lose the catalog on both sides. A committed read in stage and in abort is now confirmed through the idempotent commit(), which flushes before it answers; a failure leaves the journal and the fence where they were.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): relay shells no lease here knows count as live, Move stops them, and Windows asks every relay pipe (#25518)

* fix(ssh): a relay shell no lease here knows counts as live, and terminate stops it

A CLI-created terminal has no lease, so with an attached lease the gate answered from leases alone,
a live decision was never re-counted once the relay could answer, and terminate stopped only the
shells it held leases or panes for. The relay's own listing is the authority on what runs.

* fix(ssh): a Windows connect asks every relay version's pipe for its PTYs before converting

Windows pipes cannot be listed, so the connect-time census answered 'unenumerable' and a shell no
lease here knew let the host convert under it. Each version directory's pipe for this target is
derived from its path, so the census probes them all, current included, and asks a live one
through its own bridge; a live pipe it cannot ask is unverifiable.

* fix(ssh): the terminal gate counts what earlier relays still run, leased or not

After an app update a shell a respawn superseded on its tab keeps running on the previous relay
with no live lease here, so a decision counted only the leased shells and Move could not see it.

* fix(ssh): a Windows relay folder with a live pipe it cannot ask stays unverifiable

Each pipe is probed and asked on its own, so one that answered with no PTYs can no longer stand in
for a live sibling the census could not reach.

* fix(ssh): terminate also stops shells only an earlier relay lists

provider.shutdown routes a held id to the older relay that runs it, so the terminate set now takes
listPreviousRelayPtyIds too; one it cannot reach is reported unverifiable as before.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a refused Move to managed server reconnects the host on its relay (#25543)

Stopping the terminals closes the relay session, so a move the census refused (another desktop's
terminals, or an older relay it can't rule out) left the host and its workspaces disconnected
until a manual Connect. The refusal now reconnects the host; the connect-time decision reads the
same census and keeps the relay. A reconnect that converts after all reports the move; a
reconnect that fails still reports the refusal.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(relay): an agent exec ends on its child's exit, not on pipes a background process still holds (#25544)

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): stop the automation scheduler first when a managed stop is dispatched (#25548)

A dispatch could otherwise race the daemon retirement census or write a run record that a
rollback restore then silently discards.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): count a degraded daemon's in-process terminals in the terminal census (#25545)

In degraded mode fresh terminals run on the local fallback inside orcad, but the census read
only daemon adapters, so an update or stop saw 0 live sessions and killed running agents.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): a retained source whose drafts or settings an older build changed is never hidden as unchanged or retired (#25549)

* fix(orcad): a retained source whose drafts or settings an older build changed is never hidden as unchanged or retired

The retained-source fingerprint covered only repo, folder and group identity, so an unsaved draft
edited on an older build read unchanged: the host kept serving the server's older draft and
retirement deleted the newer one. Retention now also records a versioned fingerprint of the source's
drafts and user-authored names and settings; a mismatch marks the host changed, and a journal
without one is never retired automatically.

* fix(orcad): a retained source's automations are part of its state fingerprint

An older build can edit an automation the source keeps; retirement would delete that edit as if
the server held it.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a rollback restores the older snapshot only behind orcad's terminal barrier (#25551)

The rollback's census is taken while orcad still admits work, so a terminal or automation that
starts before the stop had its state wiped while its PTY survived. The incumbent is now stopped
through its managed stop with idle-daemon retirement: orcad closes terminal admission on every
daemon generation, counts live sessions under that fence, and retires the daemon only when none
exist. Only 'retired' lets the older snapshot replace state; live, unverifiable or a missing
answer refuses and relaunches the incumbent on its untouched state. A build without managed stop
is refused before anything changes.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): failed managed setup leaves 'connecting', edited managed host redials, idled-out server starts before its census (#25474)

* fix(ssh): a failed managed setup leaves 'connecting', an edited managed host redials, an idled-out server starts before its census

- doConnect publishes the error and clears the 'setting up' status when the managed-server
  decision throws for a still-current attempt; a cancelled one still reports cancellation.
- Editing a managed host's connection fields closes its tunnel and disconnects its transport,
  serialized with the target's lifecycle, so the next use dials the edited target.
- The terminal census starts a server that idled out behind a forward still up, so Stop, Update,
  Rollback and status no longer refuse with 'census unavailable' on every retry.

* fix(ssh): a fenced failed setup publishes its cause, never-launched slots are collectable, a reused PID is not orcad

- doConnect publishes the relay decision's setup failure (and clears 'setting up') when a failed
  managed setup kept the host fenced, instead of throwing a bare 'serves a managed server'.
- The liveness probe answers NEVER_LAUNCHED for a slot with no process record and no readiness
  file; GC removes such a slot, and every other reader still reads it as UNKNOWN.
- On POSIX a PID whose command line does not run the slot's orcad.js reads DEAD, so a stopped
  orcad behind a reused PID is woken instead of reported serving or unverifiable.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a legacy relay route a new pane started serving stays open when a pending stop settles (#25542)

A served PTY's exit that lands before its stop's reply defers the route's hang-up until the stop
settles. A second pane the same older relay holds could start serving through the route in that
window, and the deferred hang-up then closed it anyway, sending that pane's input and stops to the
current relay. Serving a pane now cancels the deferred hang-up, and a settling request hangs up only
a route that serves nothing.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(migration): the journal holds a migration's scrollback until commit or abort, across failed uploads and restarts (#25550)

Retention was scoped to one transfer call, and its release in finally deleted a closed tab's
snapshot after an interrupted upload, so every retry failed with source_snapshot_changed.
Inline buffers had no file for the retained read at all. Retention now follows the cutover
journal: held from the journaled export through staging, rebuilt at startup, released on
commit or a removed journal. Inline bytes are written to their ref while held.

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude <noreply@anthropic.com>

* fix(migration): a repo id two hosts share never lets legacy keys cross hosts, and a dangling identity alias stops blocking the move (#25558)

- Repo ids are not unique across hosts: the same id may be registered on two SSH hosts. Stores that are not session partitions (worktree metadata, automations, lineage, client state, sparse presets, retired names) can hold legacy keys with no host qualifier, so an id both hosts register said nothing about whose a key was. The scope now records such shared ids; an unqualified key or bare repo id for one only matches with its row's own host evidence (worktree metadata's hostId, an automation's ssh target), so another host's rows are never moved, counted or retired. An automation's target generation now matches only alongside its target id.
- An identity alias whose identities hold no metadata (worktreeMetaByIdentity lost them, as on the B4 profile) is nothing to move rather than a worktree-metadata blocker; retirement drops it.

Co-authored-by: m4air <m4air@Mac.localdomain>

* feat(cli): orca environment recover --accept-changed-state --yes restores over changed state (#25597)

Recover refused when a rejected build changed profile state, and its refusal told the user to run
Recover, which the CLI could not do. --accept-changed-state (confirmed with --yes) maps to the same
acceptChangedState the Managed servers settings pass, and the refusal now names the flags.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): refuse an update that would end a degraded host's in-process terminals (#25598)

The census now reports inProcessSessions separately. Those terminals run inside orcad and
end with any restart, so planOrcadUpdate defers with a non-forceable
orcad_update_ends_in_process_terminals instead of claiming they survive on the daemon.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(migration): a delta move protects what it imported from rolling back across it (#25599)

A delta move never advanced the server's migration mark, so rolling back an update taken before the delta was admitted and dropped the delta's projects. The delta now records the mark before any commit can land, resumed commits included; the mark keeps the latest migration and never moves back; and the rollback gate also counts every journal into the server, so deltas finished before this change stay protected.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a missing relay inventory never proves terminals exited, and the Windows census covers every desktop's relays (#25611)

- The migration terminal gate returned `exited` whenever this desktop held no unresolved lease, even
  when the current relay or the earlier relays could not be asked. A failed or incomplete inventory
  is now `unverifiable` regardless of local leases. With no relay session at all the gate asks for a
  host census (`needsHostCensus`) instead of reading the silence as exit; the connect, conversion and
  delta move pass that census in, and the connect hands its own census result to the conversion it
  starts. A census that cannot list endpoints (`unenumerable`) is unverifiable too.
- The Windows connect-time census derived pipe names from this desktop's target id only, so another
  desktop's relay on the same account was never probed. It now lists every `orca-relay-*` pipe on
  the machine and maps each to the relay instance that owns it through that instance's credential
  file or active-pipe marker, asking each with its own credential. A pipe no version directory
  accounts for is unverifiable unless the host proves it another account's (or gone), and an
  inventory that could not be read is unverifiable.

Co-authored-by: m4air <m4air@Mac.localdomain>

* refactor: drop the unshipped pty.resumeClient relay method and unused SSH provider unregister guards (#25595)

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a connect still deciding its server holds the raw 'connected', closes a transport its cancelled decision opened, and an edit keeps a relay host's session (#25641)

- handleSshConnectionStateChange holds a raw 'connected' while a connect is in flight even before
  any relay session exists (published as 'connecting'), so the census, deploy or conversion that
  dials the pool no longer reports the host up with no providers.
- priorConnection is captured before the server decision; a connect cancelled after the decision
  closes a transport the decision opened, unless a newer connect is using it.
- Editing a fenced host an older build changed (it runs on the relay directly) no longer
  disconnects its transport; only a host reached through its managed server redials.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): a retained source is unchanged only against a pre-commit baseline of everything a user wrote (#25602)

The state fingerprint read drafts, automations and workspace metadata from what a move could
carry, so a session a move refuses hid an older build's draft edit, and retention hashed the source
after the commit, so a crash before retention blessed whatever an older build changed in between.
The baseline is now written with the fence, before any commit is possible, from the source read
directly; a session that cannot be read, or a journal without that baseline, is unverified.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): a delta move refuses a source that changed while it checked terminals (#25691)

The plan and manifest are taken before the terminal check and session release are awaited, but the
journal took its baseline after them, so a draft typed in between became the baseline while the
server received the older one, and retirement deleted the newer draft. The baseline now comes from
the plan's own snapshot, and a source that changed since it refuses the move.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(migration): a save landing mid-retirement no longer defers retirement (#25692)

Retirement removed the source rows, then awaited the profile flush, then checked nothing came back. A session save that landed during that flush re-added a source-owned row, the check failed, and the journal stayed committed until a later connect. Retirement is idempotent, so it now runs one more pass before deferring; a row back after that is reported with the partition and owner key it reappeared under.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(automations): a headless run never reads completed when its agent never ran (#25700)

* fix(automations): a headless run never reads completed when its agent never ran

orcad (and Electron serve) finished a dispatched run on a satisfied
tui-idle wait, and a ready shell prompt satisfies it: a run whose agent is
not installed read 'completed' within seconds. Like the desktop runner,
completion now needs the agent's own status for the run's pane after
dispatch; without it the run fails after the agent-start window with the
reason, instead of claiming completion.

* fix(automations): keep idle-means-done for agents without status; fail only a refused command

Not every automation agent reports status on orcad (no hooks on the host, no
recognised title), so requiring it would fail their runs. A run completes on
the agent's own status, fails when the shell refused the agent's command
(bash, zsh, dash, fish, PowerShell, cmd), and otherwise keeps the old
idle-means-done rule after the agent-start window.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(migration): the retained-source fingerprint sees all worktree metadata retirement owns (#25694)

* fix(migration): the retained-source fingerprint sees all worktree metadata retirement owns

The state view read only legacy worktreeMeta keys and attributed them without meta.hostId, so
identity-backed metadata and unqualified rows a shared repository id leaves to hostId were missing
from the fingerprint: an older build's edit read as unchanged and retirement deleted it. The view
now uses the same attribution as export and retirement, and metadata that claims the source host
but cannot be attributed leaves the source unverified. The fingerprint version moves to v2.

* fix(migration): the retained-source fingerprint skips automations on a repo id another host shares

The state view matched automations by scope.repoIds, which ignores sharedRepoIds, so host A's
fingerprint included host B's automation on a shared repository id; editing it marked A changed and
routed it back to the relay. The view now uses the move's automationTouchesScope.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): a failed readiness read no longer kills a healthy candidate; log unsettled activations (BUG-17a) (#25701)

* fix(orcad): a failed readiness read no longer fails a candidate's launch, and log every unsettled activation

BUG-17: the candidate went ready and the client SIGTERMed it ~1 s later through its reject path,
yet that readiness passes the gate, so the launch itself failed: one readiness-wait exec that
errored failed the launch outright. Retry such reads until the readiness deadline; an
unconfirmed termination still fails at once. The update's outcome never reached the app log,
so every update or rollback that does not go through now logs its code and reason.

* fix(orcad): the host-side readiness wait ends on a wall-clock deadline

On a loaded host each poll's reads outlasted its sleep, so the step-counted loop ran past
the client's 30 s exec timeout. That timeout failed the launch, the reject path SIGTERMed a
candidate still starting, and orcad, which defers a stop until startup completes, published
readiness and then exited (BUG-17).

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(migration): stop a converted host's stale tabs landing in local, and retry retirement only on an exact replay (#25712)

* fix(migration): retirement re-removes only an exact replay of moved session state

#25692's second pass re-ran retirement on any row that reappeared during the flush, which could delete a tab or draft written after the move. Retirement now records the session rows it removes before it runs; a row that reappears is removed again only when it is byte-identical to one of those (in any partition). A new or changed row defers retirement and stays.

* fix(ssh): a converted host's leftover session rows stay in its own partition, never local

After conversion the renderer drops the SSH host's projects and worktrees but keeps their session rows. With no catalog owner left, the next save routed those rows to the local partition, where retirement read them as moved source state reappearing and deferred. Converting now pins each dropped worktree's session key to the host's partition; main's fence guard keeps that partition frozen, so the stale rows are never written.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(test): restore the codexProviderHandle import main's #25078 dropped again

#25722 restored it, then #25078 removed it, so pnpm tc fails on main's tip.

* chore(sync): keep main's own cloud and mobile files byte-identical to main

Earlier syncs added lint-only brace and template fixes to these main-owned files; reverting
them keeps #24863's diff against main free of files Phase 3 does not own.

* fix(terminal): a remote pane's reattach error no longer shows this client's OS and shell (#25693)

* fix(terminal): a remote pane's reattach error no longer shows this client's OS and shell

The error toast appended the client's environment for every non-SSH error. It now shows it only
when the pane's known execution host is this client; an SSH, managed or not-yet-known host omits it.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(terminal): type the pane-host fixture as runtime owner state

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude <noreply@anthropic.com>

* fix(orcad): a retained activation fence is ownerless, so recovery can always take it over (BUG-17) (#25698)

* fix(orcad): a retained activation fence is ownerless, so recovery can always take it over

BUG-17: a recovery that took a stale fence over, failed and retained it left a fresh lock, so
every later recovery read it as still fresh and the host could never be recovered. A run that
keeps the fence once it is done now backdates the lock; one whose remote command may still be
running keeps it fresh.

Also run the in-process terminal deferral before the forced protocol check, so a degraded host
whose daemon is empty names its in-process terminals instead of an unreported protocol.

* test(orcad): the CLI's accepting recover takes over a fence a refused recover just retained

The fake host now answers a stale-only takeover busy while a recovery's own takeover is
fresh, which reproduces BUG-17's 'still fresh' loop without the ownerless mark.

* test(orcad): keep the fake host's fence acquisition void where callers expect it

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a cancelled connect closes only the transport its own decision opened and nothing newer adopted (#25696)

* fix(ssh): a cancelled connect closes only the transport its own decision opened and nothing newer adopted

#25641's cleanup took any transport that differed from the pre-decision one as the cancelled
decision's, guarded only by connectInFlight. A replacement connect that completed (and left
connectInFlight) then had its live transport disconnected by the stale attempt.

The pool now attributes a transport it opens inside a connect's server decision to that attempt
(AsyncLocalStorage), and the latest user adopts it: a connect that connects or publishes a managed
route, or a managed tunnel that records a forward. A cancelled attempt closes the transport only
when it opened it, nothing newer adopted it, and no current replacement is in flight.

* fix(ssh): a still-current connect whose server decision fails closes the transport that decision dialed

The decision's own failure (a throw, or a fenced relay refusal) published 'error' but left the
transport it dialed open, so getPublicSshState read 'connected' and a later auto-reconnect
broadcast a plain 'connected' with no relay. Both branches now close exactly the decision-owned
transport through abandonDecisionTransport, which treats the attempt's own in-flight entry as
no newer owner while that attempt is still current.

* test(ssh): name the stand-in transport type

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): Windows readiness identifies itself by one PID when the process-table snapshot times out (#25723)

* fix(orcad): Windows readiness identifies itself by one PID when the process-table snapshot times out

A loaded Windows runner timed the whole-table snapshot out during the bundled runtime's
readiness preflight, so the candidate failed to start and activation rejected it.

* fix(orcad): fall back to the one-PID query only for a slow process table, never an unreadable one

An unreadable table (EDR-hooked snapshot, restricted token) must still fail qualification.
The table now rejects slowness with a typed WindowsProcessTableTimeoutError, and only that
falls back. Review by win-serve.

* build(cli): list the process-table timeout error in the CLI project's file list

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a connected conversion proves the host idle with the account-wide census, not this target's lists (#25697)

* fix(ssh): a connected conversion proves the host idle with the account-wide census, not this target's lists

The migration terminal gate asked the host-wide census only when no relay session existed. With a
session, it trusted this target's relay listing and its earlier-relay census, both of which name
only this target's instances, and returned `exited` when they were empty, so another desktop's live
shell on the same account, under a different target id, let the host convert under it.

`exited` now always needs a complete host-wide census: this target's lists can prove `live`, but
empty lists only pass the question to the census, and a gate given none answers `unverifiable`
with `needsHostCensus`. The connect-time refinement passes the census too, so a connect retires its
leases only when no relay on the account holds work. The Windows host lane now runs a second
desktop's relay with a live shell and expects the connected gate to read the host live.

* test(ssh): a connected relay's empty lists still ask the account-wide census

* refactor(ssh): drop the gate's unread needsHostCensus flag; its unverifiable reason says why

* test(ssh): the delta snapshot fixtures give their account-wide census

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a reconnected transport starts managed orcad itself instead of inheriting a dropped start (#25689)

* fix(ssh): a reconnected transport runs its own orcad start instead of inheriting a dropped one

The serving check deduplicated in-flight starts by environment only. When a
connect dropped mid-start (as when the launch-time auto-connect is replaced
by a reconnect), the caller on the new transport joined the start bound to
the dead one, got its failure, and reported the host managed with no server
running. In-flight checks now join only on the same connection, connect
generation and port, and a wake's own fence token is cleared only by that
wake.

* fix(ssh): a reconnected wake releases the fence its dropped wake held at any point

The flake's real verdict was 'fenced': the dropped launch-time wake held the
activation fence, and the reconnected wake could not prove it its own. A
wake now claims its token before its first remote step, an absent owner
record under a held token is still its own, and a reconnected wake waits for
this client's dropped wake to settle before reading the fence.

* fix(ssh): release only a fence carrying this process's own wake token

A fence with no owner file could be another client's fresh one. Releasing it
now requires the owner token this process wrote; the token is claimed before
the write so a drop after it still proves ownership.

* fix(ssh): type the wake's fenced fallback

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* test(ssh): stop the exec-stdin test double from failing on EPIPE (#25739)

The truncation test's fake exec channel forwarded the local shell's
EPIPE (or 'Cannot call end after a stream was destroyed') as a channel
error. Whether that error or the shell's exit code won depended on
scheduling, so the test failed under full-suite load. ssh2 silently
drops writes once the remote stops reading; the double now does the
same, and a 1 MB payload makes the early-stop path deterministic.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): Move lets the reconnect's decision run its census, and a census outside a connect holds 'connected' and closes its own transport (#25735)

- Move to managed server no longer runs a separate host census after tearing the relay down,
  which dialed the pool with no connect in flight and broadcast a raw 'connected' with no
  session or providers. It reconnects, and reads the decision the reconnect's census recorded.
  A relay a failed stop left up is detached (leases kept), not disposed, before the reconnect.
- The CLI and delta-move census (censusHostRelayTerminalsFor) runs outside a connect under its
  own owner: the raw 'connected' it causes is held, and a transport it opened that nothing
  adopted is closed afterwards. Reusing a pooled transport inside a scope now adopts it.
- Drop the now-unused publishRelayTerminalsStatus.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(test): drop the restored codexProviderHandle import now that main restored it

* refactor(migration): keep a converted host's source rows instead of retiring them automatically (#25768)

Automatic source retirement leaves Phase 3: nothing deletes a converted host's retained rows on connect, delta move, keep-server's-version or restart. They stay hidden and are removed only by stopping the server, removing the host or uninstalling. Change detection goes back to the catalog-identity fingerprint, so an older build's edits inside an already-moved project stay preserved in the retained rows without marking the host changed. The copy-only helpers the delta view uses are renamed to subtract, and the converted-host session pin now also overrides a boot primary of local.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(automations): a headless run is watched past tui-idle timeouts by the one run observer (#25733)

* fix(automations): a headless run is watched past tui-idle timeouts by the one run observer

The headless dispatcher awaited a single tui-idle wait, which rejects after
its 5-minute default, so a healthy agent working longer was published as
dispatch_failed and never observed again. The dispatcher now hands the run
to its completion watcher, whose runtime observer already re-arms wait
timeouts, honours cancellation and bounds total observation; the agent
status and missing-command checks fold into that observer, and the separate
completion loop is gone.

* fix(automations): an already-idle pane completes when the start window passes

Real-host: a stub that exited before the window left an idle shell with no
agent status, and the observer re-armed a tui-idle wait that never resolves
for an already-idle shell, so it timed out instead of completing. The
observer now keeps judging the pane while its output is unchanged, and only
waits again once the pane changes.

* fix(automations): resolve a watched headless run by its launch handle first

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(runtime-env): a re-paired managed server's subscribers recover without a reload (#25752)

* fix(runtime-env): a re-paired managed server's subscribers recover without a reload

The renderer kept the pairing revision it last read, so after an on-connect update re-paired a
managed server every subscribe and request was refused as 'pairing changed' until a reload. The
first refusal now re-reads the environment catalog, so revision-keyed subscriptions resubscribe
and requests carry the new pairing. A managed server re-pairing for the same host registration is
the same peer, so its workspaces and tabs are no longer purged as a replaced environment.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(runtime-env): a re-paired managed server is the same machine only when its host proves the same identity

Same SSH target registration is not proof: a reinstalled host or a target now pointing elsewhere
keeps it. The runtime id the pairing handshake verifies must be known and unchanged; otherwise the
re-pair retires the environment as before. A proven runtime id change also counts as replaced.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(runtime-env): decide a managed re-pair's same machine by the host's proven key, not its runtime id

The runtime id is minted per process start, so every orcad restart would read as a new host.
The host's E2EE public key persists in its own profile across updates and its pairing handshake
proves it; main now lists a digest of it, and the renderer keeps a re-paired managed server only
when that digest is known and unchanged under the same SSH target registration.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(runtime-env): defer a managed re-pair's same-machine decision until the host key is known

A re-read that lands before the new pairing's host key is listed no longer purges: the decision
waits for a catalog that carries the key and retires only if it differs. Adds the update-flow
store test: same registration and key keeps workspaces and tabs, including a re-read that runs
before the key is known.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(runtime-env): watch for a pairing refusal on a side branch so requests settle on the same tick

Chaining .catch onto every subscribe and request delayed each success by a microtask, which let a
StrictMode cleanup run before a client-event subscription resolved, so its unsubscribe landed late.

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude <noreply@anthropic.com>

* refactor: consolidate pane-ownership, migration-catalog, activation-launch and parser helpers (#25738)

* refactor(orcad): one launch-and-judge helper for activation and rollback

* refactor(orcad-migration): one copy each of the destination projections, selectNewRows, assertSameValue, compareKeys and slotLiveness

* refactor(orcad-migration): one string-list validator and one uniqueness check, error codes passed in

* refactor: one shared pane-ownership and terminal-layout module for migration, profile transfer and split layout

* fix(orcad-migration): row-identity helpers in a leaf module (no import cycle); key order in its own module

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): reopen a managed tunnel to a host another desktop restarted (#25800)

The tunnel's identity check pinned the saved runtime id, which orcad mints per process. A host
updated or woken by another desktop, or restarted while this one was away, failed every reconnect
with orcad_identity_mismatch, and nothing could refresh the id because that needs the tunnel. The
E2EE handshake with the pinned host key and our accepted token now prove the server; the first
authenticated status reply records the new id. A different host is still refused.

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude <noreply@anthropic.com>

* fix(orcad): create the readiness file owner-only so its pairing token is not world-readable (#25809)

orcadLaunchCommand truncated .orcad-readiness before setting umask 077, so under a
login umask of 022 the file that receives the pairing offer (with a runtime-scope
device token) came out 0644. umask 077 now runs first, the readiness file is
chmod 600 after the truncate (a redirect keeps an earlier build's 0644), the pid
and log files are tightened too, and the slot dir and ~/.orca-remote are chmod
700 so files earlier builds left readable are no longer reachable. The state
snapshot capture also sets its umask before creating the snapshot directory.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(terminal): explain a pane whose saved session another host connection owns (#25814)

terminal_pane_owner_host_mismatch reached the user raw, with an issue link. It now reads as a
plain explanation with the open-a-new-terminal action, like the reattach failure.

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude <noreply@anthropic.com>

* test(topology): allow phase3's headless editor-tab retirement in main's boundary ratchet (#25823)

Main's #25329 added the ratchet; phase3's mobile-session-editor-projection.ts writes the host's
own session through setWorkspaceSessionForWorktree, the same way the listed headless
mobile-session tab writers do.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(automations): a headless host closes finished run terminals, keeping the newest few (#25831)

* fix(automations): a headless host closes finished run terminals, keeping the newest few

The desktop closes a run's terminal when the run completes; orcad had no
renderer to do it, so hourly automations left a shell and PTY per run open
forever (28 after ~6h on a real host). The headless service now closes a
finished run's terminal after a 10-minute grace, keeps the newest three per
automation viewable, and never touches a run that has not finished.

* fix(automations): never close a run terminal a client typed into or is viewing

Mirrors the desktop's take-over rule on headless hosts: a finished run's
terminal stays open when any client drove input to it since spawn, is
attached to or viewing it, or when this process cannot tell (it adopted
the PTY rather than spawned it).

* test(runtime): register a viewer through the public subscribe API

* fix(automations): close only completed runs' own panes

A failed run can still hold a live agent (blocked on a prompt, past the
watch window, or after an observer error), so like the desktop only a
completed run's terminal is closed. And only the run's own pane closes, so
a pane a user split into the same tab survives.

* fix(automations): close a run pane only while it still holds the run's PTY

The use check read run.terminalPtyId, but the close hit whatever PTY now
occupies the run's pane. Restart-exited-pane and the Codex account-switch
restart put a new PTY there, so a terminal a user was using could be killed.
The close now resolves the pane's current PTY and closes only when it is the
run's own; otherwise it closes nothing and only clears the run's terminal.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): a snapshot restore past the exec timeout keeps its fence instead of racing a second restore (#25811)

* fix(orcad): a snapshot restore past the exec timeout keeps its fence instead of racing a second restore

A capture, restore or clear ran under the generic 30s exec timeout. On ssh2 the timeout closes the
channel and reads as a confirmed failure, but sshd leaves a pty-less command running, so rollback
ran its rescue restore and recover orphaned the fence for a second one, both in the same stage.

- State mutations run through execOrcadStateMutation: no abort, a client wait past the host's
  deadline, and any closed channel or busy/deadline answer is unconfirmed, so the fence stays fresh.
- POSIX hosts wrap each one in `timeout -s KILL` (where present) and a pid-checked lock dir under
  ~/.orca-remote; the Windows host script takes the same lock.

* fix(orcad): a running state mutation keeps the activation fence fresh

The fence goes stale by its lock dir's mtime after 20 minutes, and a capture, restore or clear can
now run up to 15 under it, so a rollback's rescue capture plus restore could outlast the window and
let a recovery steal the fence from a live run. While a mutation runs, the host now touches the
fence every 60s (POSIX: a background beat that stops with its shell; Windows: an interval in the
host script, whose mutations are now async so the timer runs). A dead process stops refreshing, so
stale takeover still recovers it.

* fix(orcad): the state-mutation fence heartbeat never refreshes a wake's fence

A wake writes .orca-wake-owner into the fence dir and lets its fence age toward takeover; the
heartbeat now skips a fence that holds that token, and only ever changes the dir's mtime.

* fix(orcad): a state mutation releases its host lock before answering, and names its holder by pid and start time

On Windows answer() exits in the stdout write callback, so an op that answered before its first
await (MISSING, EMPTY, FAILED) exited before the wrapper's finally and leaked the lock; a reused
pid then read as alive and every later capture, restore and clear answered busy. Ops now return
their token and the wrapper answers after releasing the lock. The holder is pid plus creation time
(the slot's process-tree addon); one that cannot be identified is stale once its lock misses five
heartbeats. POSIX gets the same heartbeat-age check for a reused pid.

* fix(orcad): a state mutation's host lock is owned by its whole process group

The lock named only the shell's pid, so a shell killed while its rm or tar ran let the next
mutation take the lock and race that child. Each mutation now runs in its own process group
(setsid, or perl setpgrp on macOS), with timeout inside it so a deadline KILL reaches the children
too. The lock records the group, and is taken over only once no member is alive; a host that can
start no group records none, and its lock is never taken over. Windows ops run in-process, with no
children to outlive the holder.

* fix(orcad): record a state mutation's process group without ps -p, and never hold a groupless lock forever

BusyBox ps has no -p, so Alpine hosts recorded no group and their lock read busy forever after a
timeout kill, reboot or OOM. The group now comes from /proc/<pid>/stat (read after the comm field's
last paren), with ps -o pgid= -p as the fallback. A lock that still names no group is taken over
once its pid is dead and its heartbeat has missed three beats.

* fix(orcad): only proof of exit frees a state-mutation lock

A Windows holder whose creation time could not be read was taken over after five quiet minutes
though its pid was alive, so a suspended clear could resume and delete freshly restored profiles.
Both platforms now free the lock only on proof of exit: a dead pid, a different creation time, or
(POSIX) a group with no live member. A live holder of unknown identity stays busy until it exits.
The owner record is written exclusively, so a run that resumes after a takeover backs off.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): never offer Move for terminals another Orca desktop or session runs (#25815)

On a host where another desktop held a live relay shell, this desktop read relay_terminals_live
with offerMove, and its copy ("Its N open terminals will restart") implied they were its own. The
census already attributes them: terminals counted only by the host-wide census, with no lease or
listing of this target naming one, run under another target or session. That verdict now carries
elsewhere / terminalsElsewhere; no move is offered (no toast, no status-line action) and the status
line says the terminals belong to another Orca desktop or session.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(automations): an update releases finished, unused automation shells before counting terminals (#25844)

Hosts with schedules kept completed run shells (the newest three, and any not
yet past their grace), which counted as running terminals and deferred every
on-connect update with orcad_update_terminals_running. The update and
rollback census now ask the server to close completed automation run
terminals no client used, with no grace or keep rule, and count after the
daemon drops them. Used, unknown, failed and running ones still count.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): every activation fence holder carries a generation token its steps and release must match (#25834)

* fix(orcad): clear a bare stale activation fence instead of asking for Recover, and report a restarting update from status

BUG-21: a wake cut short leaves a stale fence with no journal. Every update then answered
'Recover it first' while Recover answered 'none'. The fence-hold check now takes such a fence
over and drops it, and the update retries once; Recover is asked for only over a journal.
The CLI also treats a connection closed by the server's own restart during update or rollback
as expected and reports what status shows once the runtime answers.

* fix(cli): type the reconnect status response explicitly

* fix(ssh): a wake's fence carries its owner token from the moment the lock exists

The idle-exit e2e still read 'fenced' on reconnect: the launch-time wake's
lock landed on the host but its connection dropped before the client saw OK,
so the wake body never ran and never wrote its owner token, leaving a fence
nothing could prove. The token is now claimed before the lock and written by
the same command that creates it, and a wake registers itself before any
remote step so a reconnected wake waits for it instead of racing it.

* fix(orcad): every activation fence holder carries a generation token its steps and release must still match

Astra pass 8: a holder suspended past the stale window resumed, kept acting, and its
unconditional release deleted the successor's fence and recovery journal mid-update. Every
holder (activation, rollback, stop, recover, wake) now writes a token into the lock it creates or
takes over. Each remote step it issues checks that token on the host, in the same command on
POSIX and inside the host script for Windows host ops; release is conditional on the token and
moves the lock aside instead of removing the root. A superseded holder aborts with
OrcadFenceLostError and its release is a no-op.

* fix(orcad): state mutations check the fence token before their lock, and refresh only a fence they still own

On POSIX the fence guard runs outermost in serializedStateMutationCommand, before the mutation
lock and the work, and the heartbeat touches the fence only while the token is still this run's.
The Windows host script records the --fence token and refreshFence compares it. A fence-lost
answer from a state mutation is a refusal, never a FAILED fallback. One owner-file constant
replaces the wake-owner copies.

* refactor(orcad): a state mutation's heartbeat touches the fence directory it checked ownership of (review)

* fix(orcad): a release moves the journal and lock aside and keeps only its own generation's

Astra pass 9: a release that passed its token check and stalled before deleting could, once a
takeover and a successor came and went, delete the successor's journal and lock. The journal is
now stamped with the writing run's fence token (a recovery takeover re-stamps the journal it
adopts), and the release renames the journal and the lock aside, deletes each only if it carries
this run's token, and otherwise moves it straight back.

* test(orcad): a successor restore stays busy beside a paused clear on the Windows host script

* fix(orcad): classify a lost fence from the step's exit and stdout, never the error message

The real exec error quotes the command, and every fenced command carries the
guard's marker text, so any failed or timed-out fenced step read as a lost
fence and dropped its unconfirmed flag. execCommand now attaches exitCode and
stdout to its exit error; execOrcadRemote rethrows unconfirmed terminations
before any reclassification.

* fix(orcad): a wake keeps its fence token until the fence is released

A disconnect fails a wake's next step without the unconfirmed flag, and the
fence release then fails over the dead connection. Forgetting the token on
that error left a fence the reconnected wake could not prove its own, so it
reported the host as held by an update.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* test(recovery): keep Phase 3's recovery-lifetime test on main's legacy-worker ports

Main added a required hasRequestedReleases port and now skips persist when a pass resolves
nothing, so the test mocks the new port and holds the pass at workspace resolution instead.

* fix(ssh): say "1 terminal" when another Orca desktop runs one on the host (#25853)

The terminalsElsewhere status line had no plural forms, so B9 read "while 1 terminals another
Orca desktop … are running". It gains _one/_other entries like the other terminal-count strings
on that line and in the move offer, which were already pluralized.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(automations): close failed and exited runs' terminals once the shell is proven alone (#25859)

* fix(automations): close failed and exited runs' terminals once the shell is proven alone

Failed (command not found, timeout) and forever-dispatched runs kept one
shell per run, unbounded and counted by the update gate. Their terminals now
close like completed ones (unused, past the grace, outside the newest few)
but only on fresh execution-host proof that the spawned shell is alone at
its prompt; a live or unprovable agent keeps its terminal. A still-dispatched
run closed this way is marked failed. Dead terminals no longer take one of
the newest-three keep slots. The update drain follows the same rules.

* fix(automations): prove a run shell alone from the process table, not the daemon's ownership flag

On a real daemon session the daemon's confirmShellForeground stays false
after a plain 'command not found' and after an agent that exited, because its
ownership flag only turns 'shell' after a full-screen command; failed runs
would never have closed. The proof now also reads the host's process table:
on POSIX the PTY's root shell must own the terminal foreground group with
nothing stopped under it, on Windows the host's job-based child census must
be empty. Anything unobservable still keeps the terminal.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* test(ssh): extensive orca CLI matrix on Windows hosts (#25114)

* test(ssh): extensive orca CLI matrix on Windows hosts

Adds three dispatch-only app cells to the ssh-windows-hosts lane that drive the
e2e build and the bundled orca CLI against the provisioned Win32-OpenSSH host:
empty-host deploy/terminal/reconnect/orcad-restart/decommission, seeded
relay-era conversion, and an open relay terminal keeping the host on the relay.

* test(ssh): pin the relay-kept cell's runtime; keep cleanup from masking failures

* test(ssh): run decommission before the orcad restart in the managed cell

* test(ssh): decommission through orca environment stop; accept an unverifiable relay close

* test(ssh): require a confirmed relay close; app cells must run last

* test(ssh): log and accept either relay-kept census reason; keep app-cell test results

* test(ssh): relay-kept requires a live census and its status line again

* test(ssh): match the pluralized relay-kept status line

* test(ssh): orcad restart proves a new process, terminal adoption, and a kill-then-connect relaunch

* test(ssh): restart kills only the orcad server, not its terminal daemon; wait for a released profile

* test(ssh): collect orcad.log.1 so a restarted orcad's previous run is kept

* test(ssh): the managed cell proves a workspace listener is detected and attributed

* test(ssh): start the port listener without $, so a PowerShell terminal doesn't expand it

* test(ssh): the port check proves Windows command-line attribution; retry a dropped version read

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* refactor(runtime): move run-terminal client-use and shell-alone checks out of the branch-cleanup runtime (#25876)

* refactor(runtime): move run-terminal client-use and shell-alone checks out of the branch-cleanup runtime

orca-runtime-preserved-branch-cleanup.ts had grown past max-lines (303) with
the headless run-terminal helpers. Their logic now lives in
run-terminal-client-use.ts and the runtime keeps one-line delegators, with
behavior unchanged.

* fix(ci): the runtime Electron ratchet bundles its entry points once, not 2.5k times

check-runtime-electron-ratchet bundled ~2,532 entry points each in full (format cjs, no
splitting), so esbuild held thousands of copies of the runtime graph: about 2.2GB RSS and 11s per
run, twice per test file. It was in flight in every unit shard that died with "The runner has
received a shutdown signal" (#25815 5/5 twice, #25876 2/5 twice). With esm + splitting the shared
modules land in one chunk: same metafile, about 200MB and 2s.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* test(ssh): wait for the busy relay's child before probing it (#25916)

The fake relay's spawn is not visible to pgrep at READY on Linux under Bun, so
the probe could count zero children. The sibling cases already wait.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): a quit that aborts an upload whose read already ended no longer crashes main (BUG-23) (#25922)

Quitting while an on-connect orcad update was uploading its bundle aborted the connection's
teardown signal. sftp-upload's abort handler destroyed the local read stream with the signal's
reason, but once that read had ended, 'finished' had already removed its listeners, so the
stream emitted an unhandled 'error': [main_uncaught_exception] AbortError: This operation was
aborted. Electron's error dialog then blocked the main thread and the app never exited.

The read stream now always has a no-op error listener; the transfer's outcome still comes from
'finished'. Both the bare upload and the connection-level teardown abort are covered.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): a launch reads readiness at least once; fake hosts match the capture's tar flag, not any -cf (#25918)

A random fence token contains `-cf` about 1 time in 125, and the fake hosts
read any command containing it as a snapshot capture, so a rollback's restore
answered CAPTURED and the rollback never launched. Separately, a client
descheduled between computing the readiness deadline and checking it skipped
every read and failed a ready launch.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): a fence this desktop's exited process left is cleared without the 20-minute wait (#25941)

* fix(orcad): a fence this desktop's exited process left is cleared without the 20-minute wait

The client records the fence tokens its processes hold beside the profile. On
a later launch, a POSIX fence carrying a token from a process that has exited,
quiet for three heartbeats and with no live state mutation, is backdated so
the existing stale rules clear it or hand it to Recover at once. Another
desktop's fence, a live holder's, or one with a mutation still running keeps
the normal stale window.

* fix(orcad): held fence tokens are best effort, pinned to this machine and boot, and pruned after a day

A token-file write that fails no longer breaks a fence operation; an entry
recorded on another machine sharing the profile, or before a reboot, never
proves its holder exited; entries older than 24 hours are dropped. Tests cover
a journal kept for Recover, a successor freshened back after a racing backdate,
a Windows host, and the record across release, supersession, busy and a lost
connection.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): an install lock this desktop's exited process left mid-upload is taken over without the 20-minute wait (#25991)

A quit during the bundle upload leaves the version dir's install lock, not the
activation fence. The lock now carries this desktop's token, recorded in the
held-token store, and is forgotten only once its removal is confirmed. On a
later attempt, before each stale check, a POSIX lock whose token belongs to an
exited process of this machine and boot, quiet for three minutes, is backdated
so the existing stale takeover claims it at once. The fence path now shares
the same helper.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(orcad): a desktop that met another desktop's update fence clears its note once the host answers (#25995)

The serving note "holds this host" and a fence-busy update deferral stayed until a reconnect,
minutes after the other desktop's update finished. The connect now rechecks serving and the
update every 45s while the fence holds, and publishes the first answer without it. A recorded
deferral is dropped once the host runs its candidate or a newer release.

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude <noreply@anthropic.com>

* test(ci): run Phase 3's SQLite-backed tests in the Node runtime project

Main's #25967/#25998 boundary requires every test that opens real SQLite to be listed. This adds
Phase 3's eleven orcad and SSH migration tests, plus main's own agent-launch-instant-tab test
(#25430), which main's tip also leaves unlisted.

* fix(ci): keep Electron probes out of the node-server suites again (#26046)

The runner excluded *.electron.test.ts with a CLI --exclude, but main's switch to Vitest inline
projects (#25967) gave each project its own exclude list, which overrides the CLI one. The
directory selectors then pulled profile-state-writer-stall.electron.test.ts into the glibc-floor
and musl orcad-template jobs, which have no xvfb. Resolve the exact files with vitest list and
drop Electron and cross-runtime ones before running.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): keep the SSH host card quiet while its managed server is healthy (#26072)

The card showed "Runs a managed Orca server" under every healthy host. A managed server is the default, so the status line now appears only for setup progress, updates, the relay, or failures.

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(ssh): Move to managed server keeps the host's terminal tabs (#26077)

* fix(ssh): Move to managed server keeps the host's terminal tabs

Move stops the relay shells; their exits read as a user exit and closed
the tabs before the conversion copied them to the server. Suppress those
exits for the move, restart stopped shells on the relay when the host
stays, re-home the open workspace onto the server, and report stopped
shells to the runtime so terminal list stops calling them connected.

* fix(ssh): mark Move's relay stops in main's intentional-stop register

The renderer-only exit suppression left main retiring the stopped tab from
the saved SSH session before the conversion copied it, left other viewers
unprotected, and swallowed real exits for the whole request. Register
exactly the shells the move stops, from just before each shutdown, as a
'replaced' stop with their incarnation; main keeps the surface and labels
the exit for every viewer, while a confirmed death stays 'exited'. Move
now returns the shells it stopped, and a host that stays on the relay
restarts only those tabs, discarding any buffered exit first.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(hosts): show an SSH host and its managed Orca server as one host (#26076)

* fix(hosts): show an SSH host and its managed Orca server as one host

Phase 3 registers the Orca server it deploys over SSH as its own runtime environment, so every
host list built from the execution-host registry listed the machine twice under the same name.
The registry now folds the pair into one row named after the SSH host. The row routes to the
server, since a managed host has no relay, unless main reports the host back on its relay; the
other id stays as an alias so selections and renames saved under it still resolve. A retired id
that workspaces still point at keeps its own row, and servers no configured SSH host deployed
(manual pairings, orphans) are untouched.

* fix(hosts): keep both ids of a merged SSH host and dedupe only in pickers

Deleting the merged-away id from the registry broke every consumer that matches hosts by exact
id: Add Project fell back to local after a connect, the composer lost ready projects and drafts
(and could swap in an unrelated local project), and a host scope hid folder-only workspaces.

The registry now keeps both entries and marks the pair (aliasHostIds on the row pickers show,
mergedIntoHostId on the other). Pickers show one row per machine, and a choice of that row
expands to both ids: sidebar host scope, jump palette filter, notification toggles, run-target
and repository host offers. Add Project resolves a saved SSH id to its server row and blocks the
actions while that server comes up instead of choosing local. The composer's resolver now fails
closed when a named draft repo isn't actionable rather than picking another project.

* fix(hosts): widen saved host scopes, both-way palette aliases, guard Add Project host

- A sidebar or agents host scope saved by an older build (or before a route flip) can hold one id
  of a merged SSH host; a background gate widens it to both ids so exact-id filters match either
  owner.
- The palette host filter now resolves a saved id to both owners whichever id it names.
- Add Project's create and clone refuse to run while the chosen host is unresolved, and their
  submit buttons stay disabled, instead of falling through to this computer.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

* fix(sync): reconcile main's ratchet bundling and cold-serve hydrate with phase3

The Electron-import ratchet keeps main's single-stdin bundle (cjs); the
auto-merge had also kept phase3's esm splitting, which broke main's
import-graph test. Editor tabs now follow the windowless full-seed rule
from #26022, so a cold serve restart lists persisted editors too.

* fix(ssh): reclaim this desktop's own exited lock on Windows hosts too (#26087)

* fix(ssh): reclaim this desktop's own exited lock on Windows hosts too

The relaunch after a quit mid-update now frees the activation fence and the
version-dir install lock on a Windows SSH host the same way it does on POSIX,
instead of waiting out the 20-minute stale window. The host script ages the
lock only when its token belongs to a desktop process proven exited, it has
been quiet for three heartbeats, and (for the fence) no state mutation is
live, where a mutation holder counts as gone only by pid plus creation time.

* fix(ssh): take an exited holder's lock only through the steal arbitration

Review found the reclaim backdated the lock by path after checking it, so a
live successor that replaced the lock in between could be aged and then
stolen, and an interrupted or failed restore left it aged for good.

The exited-holder check is now read-only. The steal command itself accepts
the proven token and, inside its steal claim and identity recheck, also takes
a lock whose owner file still names that token and that has been quiet for
three heartbeats. Nothing is written to a lock before the steal owns it.
POSIX uses the same path.

* fix(ssh): never take an exited holder's fence while a state mutation can start

Review round 2 found the fence's live-mutation guard ran only in the read-only
proof, so a mutation admitted after the proof, or one whose first heartbeat
landed after the steal sampled the fence's age, kept running under a fence
the steal had replaced.

For the fence, the steal now takes the state-mutation lock inside its claim
(mkdir on POSIX, the exclusive owner.json on Windows) and holds it until the
takeover is done; it refuses when any mutation lock exists. Holding it, it
rereads the owner and only then re-samples the fence identity. A mutation now
rechecks its fence token right after it takes the mutation lock and stops with
the fence-lost marker if it changed. The Windows proof also falls back to the
stale window when its command line would not fit cmd.exe.

* fix(ssh): record the exited-owner steal as a real mutation-lock holder

Review round 3 found the POSIX steal held the state-mutation lock as an empty
directory, which a mutation reclaims after a minute without any liveness
check; a steal stalled that long lost its exclusion and could replace the
fence under a running mutation.

The steal now writes its pid (and group, under the same rule) with the
mutation's own noclobber owner writer, so only proof of its exit frees the
lock, and it removes the lock only while the lock still names it. On
Windows the owner record is moved into place whole, so it never exists
empty, and is removed only while it still names the steal's pid.

* refactor(ssh): keep the relay lock commands off the orcad host-script graph

The mutation-lock owner writers moved into a leaf module, so the relay's
install-lock commands no longer import orcad-state-snapshot and, through it,
the Windows host script, orcad-instance-lock and the daemon process query.
Those modules evaluate imports at load time that existing suites mock
partially. No behavior change.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-07 03:15:06 -07:00
Neil 5b0d38749d test: avoid Store imports in terminal session fixtures (#26166) 2026-10-07 03:03:14 -07:00
Neil 81ba5c3125 test: reuse scratch cells in terminal replay oracle scans (#26162) 2026-10-07 02:32:26 -07:00
Neil e6bcc5a6e0 Keep draft reload tests on the existing cache module (#26145) 2026-10-07 01:54:03 -07:00
Neil d58ea0c362 Advance mocked scheduler clocks without real sleeps (#26135) 2026-10-07 01:53:07 -07:00
Neil 6fd09185d1 Stop pane migration fixture hooks before closing stores (#26131) 2026-10-07 01:49:35 -07:00
25e0a29701 feat: stream desktop audio and video previews with native controls (#26120)
Open local and SSH video and music files with native playback controls. Stream bounded byte ranges through a scoped Electron URL instead of loading entire files into the editor.

Fixes #24859

Co-authored-by: ChangJun Park <40492343+ckdwns9121@users.noreply.github.com>
Co-authored-by: Lirone Levy <lirone88@outlook.fr>
Co-authored-by: dupi <david.li.du@gmail.com>
Co-authored-by: John Cusack <5961784+John-Cusack@users.noreply.github.com>
2026-10-07 01:29:55 -07:00
Brennan Benson 535c83a687 refactor(composer): delete the full-create path no caller reaches (#26052)
* refactor(composer): delete the full-create path no caller reaches

NewWorkspaceComposerModal is the only renderer of the composer card, and it overrides the card
props' onCreate (the full-create submit) with its quick create and never reads useComposerState's
submit. The full path (source and submit preparation, creation execution and its finalization,
issue-command, startup and structured-launch helpers, the orchestration) was reachable only from
its own tests. Delete it with them, and drop onCreate and submit from the composer contracts.

* refactor(composer): delete code the full-create path left orphaned

Removing the full-create path left code whose only readers were gone:
- applyWorktreeMeta and the updateWorktreeMeta plumbing that fed it
- createWorktree and setSidebarOpen on the composer target store
- currentIssueCommand on the composer model
- buildAgentPromptWithContext and getLinkedWorkItemPromptContext (only
  their own tests still called them) and those test cases
- the "Selected agent is disabled" locale string in every catalog
- the export on confirmRuntimeIssueCommandRead

The 'full' create-gate mode goes too. Its only caller passed 'quick', so
the default named a mode with no create action. That removes the option,
the full gate, the issue-automation wait flag and the renderer issue
command preload that ran only in 'full' mode. Quick create is unchanged.

The quick path's name-retirement comment pointed at the deleted full
submit path; it now carries that reasoning itself, including the mobile
counterpart that must change with it. An e2e comment no longer names
applyWorktreeMeta.
2026-10-07 01:26:38 -07:00
Brennan Benson e6c298e8a9 fix(native-chat): a Stop Codex took settles the message whose turn never opened (follow-up to #25217) (#26105)
* fix(native-chat): a Stop the agent took settles the message whose turn never opened

When Codex takes a person's Stop on a turn it never opened, no turn-ended
event follows, so the message stayed pending: the chat read Working with
Stop shown, and the next turn's clock counted from that message.

The Stop's settle now handles that case from its own answer: when the
provider names the turn it took and that turn has no record, the sends it
was for are withdrawn by the same rule a Codex child's end after a Stop
already uses. The send's own row then says it was stopped before the agent
started, and Working, the clock and the opening-send hold follow from it
being settled.

Deletes the opening-send hold's special case for a taken Stop's note, which
this makes dead, and moves the Stop note key back next to its only users.

* test(native-chat): a Claude Stop before the echo settles the send through the CLI's end, and the next message goes out

* fix(native-chat): an interrupted Codex turn that never started records nothing, and the withdrawal reads from the handover

Codex can abort a turn before it starts: it answers the interrupt, then
sends turn/completed for a turn that never sent turn/started. The
translator wrote an empty interrupted turn for it, so the person saw that
turn beside the message's own "Stopped before the agent started" row, and
when the end was read before the interrupt's answer the Stop's settle
found a record and withdrew nothing. Codex records a turn's prompt only
once the turn starts, so an interrupted turn this child never started,
with no item or prompt read, now gets no record. Failed ends keep theirs.

The withdrawal now asks whether a turn opened since the send's handover
row, the same point the opening-send hold reads, rather than since its
acceptance: a turn record written in between is not the send's turn.

* test(native-chat): say why the item-read case pins only that the send is not taken back
2026-10-07 01:25:50 -07:00
Brennan Benson 47f4d3f527 feat(agent-launch): the desktop AI buttons start their agent through agent.launch (#25624)
* feat(agent-launch): host-assigned caller identity and a launch record written when the surface exists

Step 1 of the agent-launch unification, on main.

- The dispatcher stamps every request's caller from what its connection proved (runtime socket:
  the local CLI; the desktop's IPC: the desktop; a paired socket: its device). Params never set it.
- The launch record is written twice: once when the tab exists (what creation settled: on the
  launch command, a draft, or a submit still `unconfirmed`), and again once the prompt's fate is
  known. A restart in between finds the running agent instead of answering "unknown".
- A replay re-derives its terminal handle from the pane key in the running host, and shows
  `unconfirmed` only to callers that advertise agent.launch.prompt-unconfirmed.v1.
- The record store opens in its own slot, without building the chat host; the chat host is built
  on that same store.

Rebuilt from this PR's own commits (b9adf88da0, 0d7d9b5b30, c83f44dd72, 39d15d52c0) onto
main, without #24080/#24081. Conflicts: the delivery doc table (main's "line fits" row plus the
`unconfirmed` row), and main's journal-database open in install(), which now goes through the
record-store slot.

* feat(agent-launch): show an agent's tab at once, where the caller asked for it

An agent.launch now shows its terminal tab before admission and spawn, in
the requested placement (group and/or anchor tab), and the pane attaches to
the agent as soon as it runs. A pane whose agent can't start, or whose start
can't be confirmed, says so instead of becoming a plain shell, and keeps
saying so across restarts. A user's close of the tab or its pane during the
launch stops it and answers agent_launch_tab_closed. Whose view moves is
unchanged from main for every caller.

Rebuilt on main (with #24934) from the previous branch head 8f62858e60.

* feat(agent-launch): the desktop AI buttons start their agent through agent.launch

Part 3 of 3 of the agent-launch unification, on main. The AI buttons (Fix
checks, the source-control actions, commit/push recovery, Explain commit,
notes and annotation sends, session continuation) started their agent from
the window. They now ask the host's agent.launch to start it, with no
prompt, into the tab the window makes at the click, under one operation id
per click; the window's pane waits for the host's agent instead of spawning
a shell. The window then pastes the prompt with main's own paste, moved out
verbatim and started only once the host's agent holds the tab, so delivery
and follow-ups are exactly main's. A typed new-tab prompt and launches while
chat is the default keep main's path. The source-control dialog drops its
launch-command preview, which the host builds now.

* style(runtime): one-line the tab-order map so the headless browser-tabs runtime stays under its line limit

The merge of main (#25724's emitMobileSessionTabsSnapshot metadata) plus this branch's close mark put
the file one line over max-lines. Formatting only.

* fix(agent-launch): watch an AI button's agent for readiness from its first output

The window started watching for the agent's ready signal only after the host
answered agent.launch, so an agent that had already turned on bracketed paste
by then was never seen ready and its prompt waited out the readiness budget.

Readiness is now watched from the moment the tab's terminal exists, as main
watches it; the paste is written, and the chat copy seeded, only once the host
says its agent started in that tab.
2026-10-07 01:18:57 -07:00
Neil 385fe8ab4e perf(tests): avoid repeated persistence imports in automation fixtures (#26111)
* perf(tests): cache Vitest module transforms between runs

* test(native-chat): align resume fixture with mention menu props

* perf(tests): reuse the SQLite store fixture in automation suites
2026-10-07 01:04:36 -07:00
Brennan Benson e5ade4b868 fix(worktrees): when the host can't fetch a remote base, create from the local branch and say so (#26009)
* fix(worktrees): create from the usable local base when there is no tracking ref, and say so, on the host's create

The host's create (agent.launch, worktree.create, the phone and the CLI) asked for a remote base like
origin/main with no tracking ref yet fetched it, and threw offline, even when the local branch existed;
it also never reported a base it fell back to. Like the desktop create, it now uses the local branch
without fetching and returns baseFallback, which the desktop window already turns into a notice.

* fix(worktrees): fetch the remote base first and fall back to the local branch only when the fetch fails

The host's create skipped the fetch whenever a local branch matched a
remote base with no tracking ref, so online phone/CLI/agent.launch creates
silently built on a possibly stale local branch, every time. Fetch first,
as main does; only when that fetch fails use the local branch the remote
names and report it with baseFallback. With no local branch the existing
error is unchanged.

Decide the base before naming so branch reuse and conflict checks, the add
and the persisted baseRef all see the base actually used, and return
baseFallback as its own field of the create result instead of riding on the
git add result.

* fix(worktrees): report the fallback for a local ref of the requested name, as the desktop create does

The tracking ref is still missing there; only the fetch is skipped, as before.

* test(worktrees): pin that the host create names only after the base fetch settles

Also pins the existing not-found-after-fetching error, now owned by the base decision.

* fix(worktrees): fall back to the local base only for creates a person asked for; keep the exact-name ref silent

Scheduled automations, orchestration workers and federation keep main's
"Check your network" error offline: nobody is watching to be told the
workspace was built on a possibly stale local branch. The fallback is a
host-side create option (allowLocalBaseFallback, default off, not on the
wire) that the worktree.create RPC and agent.launch set unless the request
carries automation provenance.

A local ref of the exact requested name goes back to silent, as on main's
host: reporting it showed "may not include the latest remote changes"
online for plain local branches.

* test(worktrees): pin the local-base opt-in at agent.launch and its absence for an automation's worktree.create
2026-10-07 00:55:21 -07:00
Brennan Benson 3427450186 fix(agent-launch): a phone's launch leaves the desktop window where it is (#26025)
* fix(agent-launch): a paired device's launch leaves the desktop window where it is

A launch from the phone (or any paired client) revealed the workspace on the
host desktop: the early tab's viewer rule said `reveal-owner` for callers whose
navigation is their own, and with no early tab the spawn surfaced the owner by
default. One predicate, `agentLaunchMovesHostWindow`, now decides both: paired
callers get `none` for the shown tab and `surfaceOwner: false` at spawn. The
device still selects the tab for itself; the CLI and the desktop are unchanged.

* fix(agent-launch): a paired device's new-workspace launch leaves the desktop window where it is

An agent.launch that creates its workspace never asked agentLaunchMovesHostWindow:
the create payload's activate/runHooks passed straight through, so a paired
caller that set either would switch the desktop to the new workspace and reveal
the startup terminal there. No shipped client sends them through agent.launch
yet, so nothing changes today; it closes the gap before a phone's work-item
creates move here.

For a caller whose launch does not move the host window, the create now runs
unactivated, which provisions setup and default tabs in the background. runHooks
keeps its setup half as setupDecision 'run'. CLI and desktop creates are
unchanged. Also makes the CLI "still reveals" test assert that a terminal was
actually created, so it cannot pass with none.

* docs(agent-launch): a paired client's create no longer activates the host to run setup

The caller-selection header said a launch that creates its workspace relies on
worktree.create's host activation to run setup. Since the previous commit that
holds only for in-process callers: a paired client's create never activates the
host window, and its setup and default tabs are provisioned in the background.
Comment only.
2026-10-07 00:50:41 -07:00
Kelvin Amoaba b062278340 fix(native-chat): let code blocks grow to their full height (#26021)
The 320px cap made long code a scroll-within-a-scroll.
2026-10-07 00:47:05 -07:00
Neil b6f3428714 perf(tests): reuse Vitest transforms between test runs (#26095)
* perf(tests): cache Vitest module transforms between runs

* test(native-chat): align resume fixture with mention menu props
2026-10-07 00:40:53 -07:00
Jinwoo Hong 4077fb4a8d fix(ci): actually exclude Electron probes from the headless node-server lanes (#26093)
vitest 5 does not apply a CLI --exclude to inline projects, so the glibc floor
container ran profile-state-writer-stall.electron.test.ts and failed with
'spawn xvfb-run ENOENT'. Resolve the selectors to files and filter them before
vitest sees them.
2026-10-07 03:10:17 -04:00
NeilandShuhei Konno 67dc092f18 fix(windows): keep generated skills and snapshots LF (#25972)
Pin generated skill JSON and Vitest snapshots to LF so Windows autocrlf checkouts pass byte-for-byte verification. Cover the checkout behavior with a real Git regression test.

Co-authored-by: Shuhei Konno <shuhei.konno@gmail.com>
2026-10-07 00:09:40 -07:00
Jinwoo Hong c8dcffce46 fix(native-chat): pass the @-mention file props in the composer resume test (#26097)
#25835 kept the onAcceptMention prop that #26017 replaced with
onChooseMentionFile and mentionFiles, so the web typecheck fails on main.
2026-10-07 02:51:07 -04:00
github-actions[bot] 98e5f5ca71 Update README downloads badge 2026-10-07 06:44:58 +00:00
Brennan Benson a3ecee0dce feat(native-chat): one notice card above the composer for the chat's own errors (#25835)
Orca's own errors in a native chat were loose lines between the transcript
and the message box: a red "Chat could not be started." line with Retry, a
red paragraph for the chat's line and for a send or slash-command error
(where the second was hidden whenever the first showed), and a muted
attachment notice inside the composer that always used a broken-image icon.
A send that threw showed the raw IPC error as the whole message.

They are now rows of one card attached above the message box, errors first:
each row has an icon, Orca's words, and its action (Retry for a failed start,
dismiss for errors the user caused). Error text Orca did not write, such as a
host or system error, is shown apart in a monospace block with a copy button,
under the words saying what didn't happen. The card also shows where a
question or approval card takes the composer's place.

Also translates "Paste failed." and the attachments-with-a-command refusal.
2026-10-06 23:30:55 -07:00
Jinwoo Hong f31d7a4ed3 fix(acp): fail the journal open through AgentSessionJournal.open in the restore-failed test (#26094)
#26038 removed JournalHostDatabase.legacyDirectoryFor while #25225's test still
spied on it, so typecheck fails on main. Use the replacement spy #26038 gave
its own tests.
2026-10-07 02:30:43 -04:00
Kelvin Amoaba 8f424512b3 feat(native-chat): suggest files on @ and name the triggers (#26017)
Shift+Enter and right-click no longer pick a row in the / menu either.
2026-10-06 23:29:20 -07:00
Kelvin Amoaba eaceac0170 feat(native-chat): searchable model picker (#26011)
The model list was a flat menu that had to be scrolled to find an entry.
2026-10-06 23:28:22 -07:00
Neil 0acf039b5d Keep native contracts on Node and wait for Git upgrade completion (#26034)
* Keep native watcher and supervision contracts on Node

* Route real permission and new ledger contracts through Node

* Synchronize handshake cleanup with forced termination
2026-10-06 23:09:33 -07:00
Brennan Benson acea59c6d8 refactor(native-chat): remove the records-file and per-chat journal imports (#26038)
* refactor(native-chat): remove the records-file and per-chat journal imports

Every native chat user is on a build that already moved chat records and
history into the app-wide database, so the one-time copies are dead code:
the agent-sessions.json import and its owed-copy flag, the per-chat
journal.db import and the write-queue hold behind it, and the pre-SQLite
log.jsonl notice.

* refactor(native-chat): remove what only the deleted imports used

- JournalHostDatabase.unsyncedTransaction and the synchronous-pragma reset
  that undid it; JOURNAL_SYNCHRONOUS is now module-local. The stranded
  rollback test drives transaction() instead.
- deleteUnpublishedJournalRows and its SQL, plus its test.
- boundJournalStatusText.
- Stale comments naming the removed first-use copy (queued messages,
  send refusal example, JOURNAL_SYNCHRONOUS doc).
- Write-queue and readInOrder comments: a write also lags when issued
  behind one still waiting in line.
- Resolved-append ordering test binds the sink from inside a running read,
  so a resolver that read at handover now fails it.

* test(native-chat): run resolved-append in the SQLite runtime project
2026-10-06 23:03:32 -07:00
Brennan Benson 6bc529176b fix(orchestration): word the worker brief for a chat worker (#26036)
* fix(orchestration): word the worker brief for a chat worker

* test(orchestration): prove the chat brief wording through the real send paths

* test(orchestration): read the redispatch paragraphs from the preamble module

* test(orchestration): normalise the chat/terminal wording in the mode-opacity check
2026-10-06 23:01:13 -07:00
Brennan Benson d0b0f13b74 fix(native-chat): hold a queued message until the turn ahead opens (follow-up to the Stop-event plan, fixes STA-9348) (#25217)
* test(native-chat): a Stop over a card sent now into the running turn keeps it paused

Red on main: Codex's turn end withdraws the steered hand-off and the queue
sends the card again as a new host turn, with no pause recorded.

* fix(native-chat): a Stop's queue pause holds a card whose hand-off is still unanswered

A card sent now into the running turn was still pending when Stop judged the
pause, so nothing was recorded; the interrupt then withdrew the hand-off, the
card went back to waiting unpaused, and the queue sent it again as a new turn.
The pause now counts a hand-off that may still return to waiting, judged with
the appended row applied, so a withdrawal lands under the pause and an
acceptance retires it in that same write. Codex and Claude both hit it.

* test(native-chat): the Claude re-send case fails on its diff, inside the test's budget

* test(native-chat): a pause held by an unanswered hand-off ends on every path that ends it

The provider's answer, the provider dying, the chat closing, a restart, a
withdrawal still owed at open, and a /clear (refused until the hand-off ends,
then carrying every waiting card paused 'cleared'); each ends with the queue
sending again.

* fix(native-chat): narrow the pause's settled hand-off, and assert the queued receipt's card

* refactor(native-chat): derive the queue's pause from Stop and Resume journal rows

Stop now appends one journal row where it takes effect, before the interrupt,
whatever the queue holds; Resume appends its own. The pause is a pure function
of the fold: the latest Stop with no later Resume and no later accepted turn a
person asked for. A /clear's carried cards name their source, which is the
replacement's 'cleared' pause. Host-origin turns never lift either.

One predicate decides which cards a pause holds; by default every waiting card
without a hold of its own, including one queued after the Stop. The drain's
consume re-judges it inside its own transaction.

The rows are tombstones of an id no item takes, carrying the mark: a released
host reads an unknown row kind as corruption and truncates the journal there.

Deletes the stored pause (recordPause, the retire hook on every appended row,
the settle-before-record step, mayReturnToWaiting and its row overlay) and the
tests that only proved it retires. The queued_message_pauses table stays in the
schema, unread and unwritten, for downgrade safety.

* fix(native-chat): a card queued after a Stop sends normally, never ahead of held ones

A Stop's pause now holds only the cards queued before its row, plus a steer it
withdrew, which returns to its own place. Each card records the journal
position it was queued at, and the one hold rule compares that with the Stop
row. A card queued after the Stop is a new instruction: it sends as usual, but
the drain still stops at the first held card, so it never overtakes them.
/clear's pause holds the cards it carried. Holding every card again is a
one-line switch in that rule.

* fix(native-chat): the queue's own send re-checks the no-overtake rule in its transaction

The drain's pick and its consume now read one function, nextSendableQueuedCard,
so a Stop row that lands between them holds a newer card behind an older held
one exactly as the pick would. Notes why Stop and Resume ride a tombstone row.

* fix(native-chat): stop creating the unused queue pause table

The queue's pause is derived from journal rows, so nothing reads or writes
queued_message_pauses. It was still created on every open "for downgrade
safety", but an older build creates it itself when it opens the database, so
the table only sat empty in every new database. The tests now pin that no
pause table exists.

* fix(native-chat): a Stop's pause never hides the restart pause

A Stop holds only the cards queued before it. The pause derivation still
returned the Stop alone whenever it was in force, so the restart pause was
never considered: a card queued after the Stop, written by a host process
that has since exited, sent by itself after Orca restarted, with no pause
header and no Resume. A /clear pause that held nothing could hide it the
same way.

Every pause in force is now derived. A card is held if any of them holds
it, and it names the first that does. The drain's pick, the consume
transaction's re-check and the published header all read that one rule;
the header names the pause holding the first card Resume would send.

* test(native-chat): pin the Stop's no-resend, lift and held-card rules

- The Claude and Codex Stop-withdraws-a-steer tests checked "not sent
  again" at one instant, before a queue ignoring the pause re-sends. They
  now wait for the stopped turn to end and re-check after a quiet window.
- The deleted-card test read a card queued after the Stop, which sends
  whether or not a person's turn lifts it; it now reads the Stop's pause
  before and after that turn.
- Unit cases pin that a Stop holds a card with no recorded position and one
  queued before a rewind.

* refactor(native-chat): a Stop writes one Stop event with its reason, turn and caller

The Stop row that paused the queue becomes the general Stop event
{ reason, turnId?, at, caller? }, whose reason is the host's existing stop
cause. It still rides a tombstone of a host-only id (a released host deletes
the journal from the first unknown row kind), and Resume keeps its own marker
on its own id. Only a person's Stop (reason user-stop) pauses the queue.

* test(native-chat): a rewind keeps a lifted /clear pause lifted and restates the same Stop event

* test(native-chat): pin that Stop and Resume rows never reach apps or count as history

* test(native-chat): only a person's Stop event pauses the queue

* test(native-chat): pin that a Stop's event precedes the interrupt and the at-start stop

Through the real host: the event names the turn and who asked and is in the
journal when the interrupt reaches the agent; at an agent still starting it is
there before the start is ended and holds a card queued before it; an idle Stop
writes one only when it withdrew a send; and the queue's claim re-judges a
pause that landed after its pick.

* test(native-chat): a card held at a starting agent is checked before the Stop's timing

Also says precisely what the claim's in-transaction pause check defends
against: the Stop and the drain share one serialized lane.

* test(native-chat): a released build keeps and folds a journal holding Stop events

Replays this build's rows from the released build's own journal database: every
row is kept, the history after the Stop still folds, and an older client is sent
only removed ids no item uses.

* style(native-chat): format the Stop event changes

* test(native-chat): type the released build's exports through one checked helper

* fix(native-chat): the Stop/Resume row guard narrows to those tombstones only

* test(native-chat): run the Stop-event downgrade test in CI, and cover a writable downgrade

The Stop-event downgrade test ran in no CI lane: unit shards exclude the
cross-version folder, and the cross-version lane runs a fixed file list that
did not name it. It is now on that list.

Its only case replayed the rows into a release's own fresh database, because
that release cannot open the current host database. A second case opens the
journal this build wrote with a main build that shares the database: it opens
writable, keeps every row, appends, and this build then reopens it with the
person's Stop still pausing the queue.

* fix(native-chat): a Stop that stops nothing new writes no Stop event

A Stop reaching a running agent wrote a Stop event on every press. Two
presses before the first interrupt landed wrote two events, so a card
queued between them counted as before the latest Stop and was held,
though a card queued after a Stop should send normally. A Stop naming a
turn that had already ended, as a phone sends late, also wrote an event
for a turn it never stopped.

It now writes one only when it withdrew a queued send, or stops something
no event records yet: not a turn the journal no longer runs, and not the
live turn a Stop still in force already names, unless a card was handed
over into it since, which this Stop's interrupt sends back and must hold.
The interrupt and the "already finished" note are unchanged. A Stop at a
starting agent still always writes.

* test(native-chat): pin that a later host, eviction or close Stop never lifts a person's Stop

* chore(native-chat): put each Stop-row doc on its own declaration, and say only user-stop is journaled

* fix(native-chat): any later Stop event ends a person's Stop pause

A person's Stop paused the queue until their next accepted turn or Resume,
and a later Stop of another reason (the host stopping the agent, an
eviction, a close) was ignored. Now the pause is the latest Stop event's:
a later Stop of any reason ends a person's pause, and only a person's Stop
pauses. The fold keeps the latest Stop event whatever its reason.

An eviction of a resting chat writes no Stop event (a Stop that stops
nothing writes nothing), so it cannot release held cards; a test pins that
no event means no lift.

* fix(native-chat): a second Stop press is a repeat even when the first came before the turn showed

A Stop pressed before the agent's turn shows in the journal (before
Claude's echo, or before Codex opens the turn) records no turn. A second
press once the turn showed compared that missing turn with the live one,
wrote a second Stop event, and held a card queued between the presses.

A repeat is now judged by what was sent since the Stop in force: with
nothing sent after it (a refused send aside), a Stop that named no turn,
or named the live one, is repeated and writes nothing. Anything sent since
and not refused, including a send whose fate is unknown, makes the new
press write, since its interrupt may send that card back to waiting.

Tests: the two-press case across the turn showing; a steer between the
presses settled unknown; and a Stop naming a turn that ended while the next
card is sent but shows no turn yet, which writes and holds that card. The
fold test that claimed an eviction path is renamed.

* fix(native-chat): the queue's pause ignores a Stop or Resume row holding a value no build writes

A Stop or Resume row's value is read from disk with no shape check, and
the pause fold stored whatever it found. A stored `stopEvent: null` would
then throw on every pause check for that chat: the queue's pick, its
send, and every queue update to clients. No build writes such a row, so
this is hardening.

The fold now reads a Stop only when it is an object with a string reason
and a finite time, and a Resume only when it is `true`. Anything else is
ignored: it pauses nothing and ends nothing. The row is still not treated
as malformed, which could cut the history short.

* fix(native-chat): a Stop still reads as yours after Orca restarts before the turn ends

Every stop that ends work now writes the Stop's event before it ends the child: a
person's close of the chat, an eviction (worktree teardown, orchestration stop, tab
cleanup) and the idle sweep's stop of a start that never landed. A stop that ends
nothing writes nothing, and quit writes none: its resume marker records why.

The turn-end write reads the latest Stop event where every turn row is built, so the
adapter's settle, the host's fallback and the relaunch's settle all agree: a turn a
person's Stop or close named, ending with no verdict of its own after that Stop, ends
as their cancellation. A relaunch's probe-bounded end is no earlier than a Stop that
found the turn running. When the provider refuses the interrupt and the turn runs on,
a refusal row answers the Stop, so a later crash still reads Failed; pressing Stop
again after a refusal is a new Stop.

* refactor(native-chat): a stop no longer carries its cause; the turn's end reads the Stop event

The cause of a stop was threaded in memory from each entry through the host's stop
step, the adapter router and each adapter's close onto the `ended` it settled with,
and Claude kept a per-turn copy of a Stop it sent. All of that is gone: adapters
settle a turn they cut as interrupted with no verdict, the host's fallback does the
same, and the one rule where a turn row is built (`turnEndAfterStop`) reads the
journal's latest Stop event to say whether it was a person's.

- `closeSession` / `disposeSession` take no cause; `ended` has no `stopCause`.
- Claude reads an error result after a person's Stop as their cancellation from the
  journal's Stop event (through the event sink), not from a per-turn slot, and a
  refused interrupt is the host's refusal row, not `withdrawTurnStop`.
- An owed wind-down keeps no cause: its retry's fallback reads the Stop event.
- The mutation context's Stop passes no cause: its step already wrote the event, and
  the delivery loop's child-end reason is read back from it.
- A Stop pressed before its turn showed applies to the turn that opens under it,
  unless a send a person made since was accepted.

* test(native-chat): a turn a later send opened is no Stop's that named no turn

* test(native-chat): the restart test's death proof carries its detail

* refactor(native-chat): a refused Stop leaves no record; a Stop only ever ends the turn it names

The stop-refused mark is gone: its tombstone kind, its fold, the clock-keyed match that tied it to
a Stop, and the exception that let a second press after a refusal write a new Stop. A Stop that
stops nothing writes nothing. A Codex refusal names a turn that is no longer its active one, and
the Stop names that turn, so the turn running instead never reads as the person's by its id alone.

* fix(native-chat): a Stop pressed before any turn showed stops only the turn opened next

A Stop that named no turn read as the person's cancellation for every later turn that opened
after it, until a send a person made was accepted. The queue's drain, orchestration mail and a
restart continuation send as the host, so a turn they opened long after, cut by a crash, read
"Interrupted" as if the person had stopped it. The Stop now applies only to the first turn
opened after it.

* fix(native-chat): an older Claude's error end after a Stop pressed before its echo reads Interrupted

Claude CLIs before 2.1.91 end an interrupted turn with an error result that names no reason. The
translator judged whether a person's Stop explained it by its own copy of the Stop rule, which
ignored a Stop that named no turn, so a Stop pressed before Claude echoed the send read "Failed".
The translator now writes such an end as interrupted with no verdict and no error row whenever a
person's Stop may name the turn, and the journal's one rule decides as it writes the end.

* fix(native-chat): a person's Stop and /clear each name why they end the agent

The host's mutation path ended the agent with one "recorded" ending for every caller, which read
back the reason of whatever Stop event the journal held last, however old. /clear writes no Stop
event, so its end took an unrelated earlier reason. Each caller now names its own: the chat's Stop
`user-stop`, whose event its own step wrote, and /clear `user-close`, the user replacing this chat.

* fix(native-chat): a host stop judges whether it ends work after the provider's rows land

A close, eviction or host stop decided whether it ended a running turn from the journal as it
stood, while the provider's own rows (the turn its echo opened) could still be in the session's
event sink. A close landing in that gap wrote no Stop event, so the turn it cut read as news. It
now reads after the sink drains, as a person's Stop does, through the same check; a drain that
fails or takes over a second reads working.

* fix(native-chat): a Claude Stop naming a turn that just ended still marks the follow-up it cuts

A phone names the turn it last saw. When that turn had ended and a follow-up was still unechoed,
Claude's Stop interrupted the follow-up and ended the child, but the Stop's event named the ended
turn, so the follow-up's turn the child's end cut read "Failed" under "Cancellation requested.".
A Stop that ends the provider's session ends whatever is in flight, so its event now names the
live turn or none, and a Stop that names none binds the turn opened next. Codex keeps naming only
the turn the Stop names.

The Claude Stop turn-end tests move to their own file, since the session-ending Stop suite is at
its line budget.

* fix(native-chat): the idle sweep reads working by the same rule as a stop's event

The sweep judged a chat resting while a send whose reply was lost was still unanswered, but the
stop's event writer counts that send as work. So the sweep evicted it and wrote an evict event,
which ends a person's Stop pause and let the cards behind it drain on their own. The sweep's owed
work now reads the main agent working the way every session list and the event writer do.

* test(native-chat): an aborted eviction's injected drain failure lands on the eviction's own drain

A host stop now drains the session's sink once to judge whether it ends work, so the tests that
fail the eviction's drain-published step skip that first drain.

* fix(native-chat): the idle sweep's rest writes no Stop event; it evicts a send that never echoes

The previous commit made the sweep count an unanswered send as owed work, which pins a chat whose
admitted send Codex never echoes forever, and the sweep exists to retire exactly that. That rule
returns. The sweep stops only an agent it judged resting, so its eviction now writes no Stop
event, whatever send it retires: a person's Stop pause holds through it.

* fix(native-chat): stopping a start that carries no send writes no Stop event

A host stop, eviction or close of a starting child wrote a Stop event whatever the start carried.
A start with a send already reads working, so the clause only mattered for a start with none,
which ends no turn and no send: its event only lifted a person's Stop pause and bumped the idle
clock, which is why the idle sweep had been changed to close the conversation in the same pass.
The clause goes and the sweep is #24072's again. The child's end still reads host-stop, as before.

* test(native-chat): a Stop's pause across a restart is tested with a restart that writes no event

The rig's restart closes the chat with an eviction, which now writes a Stop event when work runs
and so ends a person's Stop pause. "A Stop never hides a restart's pause" then passed with no Stop
pause left to hide anything. Those tests, and the pause-lift test whose dropped assertion returns,
restart as a process that dies with no close, which like a quit writes no Stop event, and assert
that both the Stop's and the restart's pauses are in force first.

* fix(native-chat): a host stop of a turn a person's Stop is still ending keeps that Stop's reason

An eviction or host stop that landed while a person's Stop or close was already ending the same
turn wrote a newer Stop event, and the turn's end reads only the latest, so the person's Stop of
that turn read as news. A host reason now writes nothing while a person's Stop still decides what
runs: the live turn it names or bound, or, with none, the turn a send opens next. The person's
own close still writes. The E2 tests now open and end the stopped send's own turn, as Codex does,
so the mail turn after it is not the turnless Stop's.

* fix(native-chat): an older Claude's error on a later turn keeps its error text after a Stop

The translator left an error result that names no reason to the journal's Stop rule whenever a
person's Stop named the turn or none, but the rule binds a Stop naming no turn only to the turn
opened next. So a real error on a later turn read "Failed" with its error text dropped. The
translator now asks the journal's rule itself (`personStopDecidesTurn`, the one core
`turnEndAfterStop` and a host stop's in-force check share), so the two cannot disagree.

* fix(native-chat): a Stop of a start that never landed binds no later turn, whatever sent it

A person's Stop pressed while the agent starts names no turn, and the send it stopped is
cancelled before it opens one. The Stop then bound the next turn anything opened (orchestration
mail, a restart continuation, the queue's drain, all of which send as the host), so a host
eviction of that turn wrote nothing and its crash or close read as the person's cancellation. A
Stop that named no turn now binds only a turn no send journaled after it opened: any send since,
of any origin and not refused, opens its own. The E2 test's mail send is accepted as Codex
accepts it, instead of opening the stopped send's own turn first.

* test(native-chat): a rewind's restated turnless Stop binds no turn opened after the rewind

A Codex rewind restates a person's Stop still in force after the turns it keeps, at a new
sequence, so by sequence alone it would bind the next turn opened after the rewind. A send
journaled after the restated row voids that binding (the previous commit), which this pins.

* fix(native-chat): a relaunch settles a person's stopped turn with no "stopped while in progress" row

After a restart, a turn a person's Stop ended reads "Interrupted after N" with the muted mark, but
the relaunch still added the error row saying the provider stopped mid-response, which a live Stop
never writes. The settle now skips that row when every turn it interrupts is the person's Stop's
by the journal's one rule; a crash nobody stopped keeps it.

* test(native-chat): the unexpected-exit settle's journal fake answers whether a person's Stop decides a turn

* fix(native-chat): a host stop whose sink drain fails reads the journal as it stands

A host stop drains the session's sink before judging whether it ends work, and a failed or slow
drain read as working. So an eviction of an agent at rest wrote a Stop event that ended nothing,
which lifts a person's Stop pause, and a close wrote a person's event naming no turn. The drain is
now best effort: the stop goes ahead either way and only its record is at stake, so a failed or
slow drain leaves the journal's read as it stands. A person's Stop keeps its own rule.

* fix(native-chat): a Stop that named no turn applies only to a turn a send it stopped opened

A person's Stop pressed before any turn showed names no turn. It bound the first turn opened
after it, then (5ead1f6bcc) any turn opened by no later send, so a turn the host started for
a card the Stop held, or for orchestration mail, read as the person's cancellation, and a host
eviction of it wrote no Stop event when its send had been abandoned by the close first.

The rule is now the concept itself: a Stop naming no turn applies to a turn opened by a send it
stopped, one already handed to the agent at the Stop's position. Nothing new is stored. The turn's
row names the send that opened it (Codex: the submission's key; Claude: the echo, which the journal
aliases to the submission), and a handed-over send's item sits at its handover, so the target set
is derived from the journal. A card the Stop held is handed over after it, so it is no target; a
Stop of a start whose send never opens a turn binds nothing; a rewind keeps no submissions, so a
restated Stop binds no turn opened after it. With no turn running, a host stop defers to the
person's Stop only while every unanswered send is one it stopped. Claude's translator, which asks
before its echo row lands, passes the send its echo acknowledged.

* fix(native-chat): a host stop whose sink drain runs long reads the agent working; a failed one reads the journal

A drain past its bound may still hold the turn's row, while the echo's acceptance has already
landed, so the journal as it stands read nothing running: a person's close of that turn wrote no
Stop event and the turn read as news. The two drain outcomes now differ: one that failed has
nothing more to deliver, so the journal's read holds (as before); one still running reads working.

* test(native-chat): a host stop with no turn running defers only while every unanswered send is the Stop's

The branch had no test. An eviction with only the stopped send unanswered writes nothing; one
with a send made after the Stop still unanswered writes its event.

* fix(native-chat): a slow sink drain reads working only while an accepted send's turn row is due

The previous commit read every drain past its bound as working, so a host eviction or stop of an
agent at rest during a sink backlog wrote a Stop event that ended nothing and lifted a person's
Stop pause. A slow drain now reads working only when the latest send the agent accepted has opened
no turn the journal holds, the race it was for; otherwise the journal's read holds.

* test(native-chat): the host-stop control keeps the stopped send unanswered beside the later one

With both unanswered, the host writes only because not every unanswered send is the Stop's; a rule
that deferred when any one was would pass the old control.

* fix(native-chat): a steer is no send owed a turn when a slow drain judges a host stop

A slow drain reads working when the latest accepted send has opened no turn yet. A Codex steer or
a Claude fold is accepted into the running turn and never opens one, so a chat at rest whose last
send was a steer still read working, and an eviction lifted a person's Stop pause. Sends delivered
into a running turn, whose item carries that turn's scope, are skipped.

* feat(native-chat): a chat reads Stopping from the person's Stop until the work it stopped ends

The host derives it on each journal publish from the Stop's event, the live turn and the Stop's
own answer, and publishes it as an optional field on the session status and the main agent's row.
Clients present it: the chat's tail line and Stop control, the sidebar row, worktree ps and the
phone's row. The chat and the phone also read their own Stop press until its request answers.

* test(native-chat): pin Stopping on the phone and across mixed versions

* test: give touched fake journals and mocks their SAFETY notes

* test(native-chat): a turn waiting on the person reads attention, never Stopping

* test(native-chat): the sidebar row follows Stopping when it is the only field that moved

* test(mobile): the phone reads Stopping from its own Stop until the request answers

* test(mobile): type the held Stop request instead of casting it

* chore: keep the base lockfile (a local pnpm run rewrote it)

* fix(native-chat): narrow the Stop note's optional failure; type the phone test's reply

* test(native-chat): type the Stop test envelope's fields narrowly

* fix(native-chat): a Stop the agent declined, or whose child end failed, says so while the turn runs on

A Stop naming the turn that still runs, refused by the agent, now writes the
Stop's refused fact instead of 'already finished'. A session-ending Stop whose
child end fails while the work runs on revises its note to unconfirmed.

* fix(native-chat): Stopping holds while any press of the Stop took

A repeat press refused after an earlier press took no longer clears Stopping.
Exit early when the Stop named a turn that is not the live one.

* fix(native-chat): keep Stop enabled while the host says Stopping

Only this client's own Stop request in flight disables Stop and Esc. A repeat
Stop is how a stop the provider took but never answered escalates.

* fix(sidebar): every agent row says Stopping in place of its tool line

The dashboard row, which the sidebar's non-compact mode also draws, read the
last tool line while a person's Stop ended the turn. It now shares the compact
row's rule.

* fix(mobile): the worktree list sees Stopping change on its own

A snapshot whose only change was the host dropping Stopping compared equal and
was thrown away, leaving the row on Stopping.

* refactor(native-chat): fold the status feed in src/shared for both clients

The snapshot merge and the contact-loss strip move out of the renderer feed so
the phone folds the same stream the same way.

* feat(mobile): let phones read the structured session status stream

agentSession.subscribeStatus joins the mobile allowlist. The agent-session
methods move to their own file, which the at-cap allowlist spreads in, and the
allowlist test reads the Set instead of parsing the source.

* feat(mobile): the phone chat reads Stopping from the host, like the desktop

One status stream per client, opened on a host that advertises the status feed;
a refusal to phones reads as no feed. The chat reads Stopping from the host or
its own press, and holds Stop only while its own request is in flight.

* test: give the new fakes checked types or a SAFETY reason

* revert(native-chat): drop the refused-named-turn rewrite of a Stop's note

Codex can send its refusal before the turn's end frames, so reading the turn as
still live after a flush races; the Codex Stop that ends nothing is handled by
ending the process instead. The base's 'already finished' note and its test
expectation return. The Claude wind-down failure revise stays.

* fix(mobile): release the status stream when the host ends it; refusals last one connection

The feed now drops the handle of a stream the host ended or refused, so the
logical client never replays it on a later session. A refusal to phones holds
for one connection, so a host updated while the phone stays paired is asked
again.

* perf(native-chat): read the live turn's opener from its record when deriving Stopping

After a Stop that named no turn, every later commit walked and copied the whole
journal to find the live turn's record. The derivation now reads that record
off the rendered snapshot's tail and decides with the same rule.

* refactor(mobile): move the method-unavailable check into transport

The status feed imported it from the Files tab's fallback. No behaviour change.

* test(native-chat): a send after a Stop reads Working before its turn opens

Pins the derivation's running-only read of the newest turn: the stopped turn,
already ended, must not keep the next send on Stopping.

* fix(sidebar): the compact row leads with Stopping so a narrow sidebar keeps it whole

At the default width the row read 'Codex Chat - Stoppin…': the model and time
keep their room and the line truncates from the end. Stopping now leads the line
the way monitoring already does, so the chat name is what gets cut. Also pins
that the turn bar keeps its running clock while the tail line says Stopping.

* test(native-chat): the retry of a close whose exit was unproven writes no second Stop event

The idle sweep finishes a stop left owed with that stop's own cause. It is the same stop, so its
event stands alone and the child's end keeps the cause, for a person's close and an eviction.

* fix(native-chat): read and write a Stop's answer by the turn its event records

Stop notes are now one row per turn, keyed by the turn the Stop's event
records. Performing a Stop and deriving Stopping share that key. A refused or
unconfirmed answer never overwrites one that took at the same key; only a
session-ending Stop's failed wind-down downgrades it, and that step now revises
the note the Stop actually wrote, carried on the wind-down. Stopping reads the
turn's note whenever it was first written, plus newer notes no other turn owns.
Test fixtures gain the host logger and the phone's quietRepeatedStop.

* refactor(native-chat): read a Stop's target once for its event and its note

The chat's Stop now reads what it is aimed at (the named turn and whether the
Stop ends the provider session) once, and both its event and its note's key
derive their turn from that one value through the same rule. Adds the host test
for a Stop naming an ended turn on a provider whose Stop ends the session.

* fix(native-chat): the host never steers a message into a turn a Stop is ending

A queued card's Send-now, or a send made while a person's Stop ends the turn,
went to the agent as a steer into that turn. The delivery loop now holds any
waiting message while the host's own Stopping reading holds, and sends it as
its own turn once the turn ends. The Stopping reader also stops at the Stop's
position and looks the turn's note up by key, instead of walking the whole
journal.

* feat(native-chat): while Stopping, the composer says a message runs after the stop

Desktop and phone: the composer placeholder reads "Queue a message to run
after the stop" while the chat reads Stopping, and a queued card's Steer (and
the desktop's steer shortcut) is held. New key translated in all 6 catalogs.

* refactor(native-chat): the chat pane's Stop controls live in their own module

The pane went over its line limit once merged with main. Its Stopping reading,
the press that holds Stop, and the steer and placeholder it hands the composer
move to native-chat-structured-stop-controls.ts.

* fix(native-chat): hold a send at its handover, reading the feed's own Stopping

The hold was checked when the delivery step started, but the handover runs in a
later step after waiting on the agent's start, so a Stop landing in between let
a new send steer into the stopping turn. The check now runs at the handover.
It reads the status feed's projection for the commit instead of rendering the
journal again, so holding a send adds no journal read of its own.

* refactor(native-chat): one display status decides Stopping on every surface

agentStopDisplayStatus combines whether the agent works, the host's flag and
this client's own press. The chat pane, sidebar and dashboard rows, and the
phone's chat all read it, instead of each combining the flags.

* fix(native-chat): Stopping holds until the stopped turn ends, whatever the Stop's answer

A Stop the agent declined, or whose end went unconfirmed, used to drop the chat
back to Working. It now stays on Stopping until the turn ends, and Stop stays
enabled so a repeat press escalates. The Stop's answer is no longer read for
Stopping, so its note goes back to the key the base gives it (the restore of
queued-stop.ts and the removed key test landed in the previous commit). The
note still keeps a press that took over a later refusal, and a failed process
end still says the Stop went unconfirmed.

* fix(native-chat): a Stop binds only the turn it actually stopped

A Stop pressed before any turn showed used to claim, at end-write time,
whatever turn the stopped send later opened, even when the Stop stopped
nothing. A turn that then died on its own read as "Interrupted" (your
cancellation) instead of "Failed".

Now a person's Stop that named no turn binds, in memory only, every turn
that ends while the Stop settles, and afterwards only the turn its
interrupt took. The settle ends a still-running stopped turn once. A
relaunch finds nothing in memory, so an unsettled turnless Stop binds no
turn. A Codex Stop whose answered turn does not open within its wait, or
whose send's answer was lost, now answers refused, so the host ends the
child and the turn can never run.

* fix(native-chat): keep the person's queue pause and close binding after a Stop settles

A host stop or eviction with no turn running now defers to a person's
Stop while its queue pause still holds with nothing sent since, read from
rows, so a held card is not handed off on reopen after a Stop that did
nothing or whose kill failed.

A person's close that named no turn opens a settle around its child's
end, so a turn that end cuts reads as theirs.

A press opens its settle only when the latest Stop event is its own or
the one in force it repeats: a late Stop, a card's interrupt or a lost
event row reopens no earlier Stop.

A Codex Stop that cannot reach a turn still able to open says the Stop is
unconfirmed rather than that no turn ran, and a second Stop still reaches
a turn an earlier wait left unopened.

Also drops the unused openedBy plumbing and the unreachable "a written
cancellation stays one" rule, and pins a relaunch after a named Stop.

* fix(native-chat): keep a failed Stop's turn display-only, and settle edges off the commit path

A Stop that failed marks the turn it could not stop for "Stopping…" only
(JournalStopSettle.failedOn): no turn-end rule reads it, so that turn's
own end with no verdict reads as a failure, not the person's.

A settle edge writes no row, so it no longer goes through the journal's
commit listener, which also delivers history, counts as activity for the
idle sweep and schedules the queue drain. A narrow settle-edge hook
republishes the status row and wakes the steer hold's handover, and
nothing else.

* fix(native-chat): a Codex Stop agrees on both presses when a turn is still owed, and pin the close's settle

A Codex Stop that waited for a turn Codex answered a send into now answers
"may still open" whenever that turn neither opened nor ended and its send
is still owed, however the wait ended (it ran out, or the thread went
idle). Before, a first press after an idle thread said no turn was
running and kept Codex, while an identical second press ended it.

Adds a test that a person's close the conversation outlives (as /clear
does) closes its settle, so a later turn that ends on its own reads as a
failure.

* fix(native-chat): a Stop that failed before its turn showed still reads Stopping through that turn

A Stop that failed with no turn open marked nothing, so the chat dropped
to Working and the turn that then opened never read "Stopping…". The
display-only mark now also covers that case: the first turn that opens
after the Stop failed, provided no message was handed to the agent in
between. No turn-end rule reads the mark, so that turn's own end with no
verdict still reads as a failure.

* test(native-chat): a Stop whose event row failed binds no turn to an earlier Stop

With one ordered journal writer the Stop's event is in the fold when its
write returns, so the press reads whether it owns the latest Stop from the
fold instead of awaiting the write. Pins the case the read must refuse.

* test(native-chat): name the settle, not a stream drain, in the Stop's own-end test

* test(native-chat): a Codex Stop answered before Codex ends the turn reads interrupted throughout

Codex answers an interrupt it took before it sends turn/completed (interrupted):
on TurnAborted the app-server answers pending interrupts, then ends the turn, on
one channel. The test fake did the reverse. It now answers first and ends the
turn on a later read, and the tests that read the turn's end right after a Stop
wait for it.

New end-to-end test through the shipped host, journal and Codex adapter: with
the real order, every end row of the stopped turn reads interrupted by the Stop
(named, unnamed, and a Stop pressed while turn/start was in flight). Breaking the
settle window turns the in-flight case red: the Stop's own end row then has no
verdict, which reads as failed until Codex's end lands.

* test(native-chat): Stopping ends with a Codex turn whose interrupt is answered before its end

With Codex's real order (the interrupt's answer, then turn/completed interrupted),
the status shows Stopping while the Stop settles, drops it once the turn ends, and
never carries a verdict other than the person's cancellation.

* fix(native-chat): a Codex Stop interrupts a turn Codex answered but has not opened at once

A Stop that named no turn, made after Codex answered a send but before the turn
opened, used to wait up to 5 s for the turn to open before interrupting, and
ended the Codex process when it didn't. The stated reason, that Codex refuses an
interrupt until it opens the turn, holds only part of the time: with no turn
active, Codex takes an interrupt once its thread runs (turn_interrupt_inner),
and refuses it with -32600 "no active turn to interrupt" before that or once
the turn has ended.

The Stop now sends the interrupt at once. Only on that refusal, while the turn
has neither opened nor ended, does it wait for the turn to open (bounded at
5 s) and send it once more. A turn that ended meanwhile was nothing to stop. One
that never opens, or that an earlier wait already gave up on, fails the Stop,
and the host ends the child as before. Sends still wait for the turn to open
before steering into it.

The test fake models Codex taking an interrupt once the thread runs (run()).

* fix(mobile): name how the phone's status stream is released in the subscription inventory

Main made each inventory entry state its release; the status feed's stream is
released from its subscribe params, as the session event stream is.

* fix(native-chat): every Codex Stop waits for an answered turn to start, as the first did

A Stop whose interrupt Codex refused as finding no active turn skipped the wait
when an earlier wait, a Stop's or a send's, had already given up on that turn.
Every press now waits its own bound and retries once if the turn starts, so a
turn that opens during a later press is still stopped. Both presses still reach
the same verdict when it never starts.

* fix(codex): never steer a turn whose interrupt Codex answered

Codex answers an interrupt as the turn aborts, before it sends that turn's
turn/completed. In that gap the adapter still counted the turn as running, so a
message handed over right after a Stop settled (the Stop's own end row already
reads the turn ended) went out as turn/steer, which Codex refused with -32600
"no active turn to steer", and only then as turn/start. The adapter now marks a
turn whose interrupt Codex answered as aborted until its turn/completed, never
steers into it, and starts the message's own turn directly. Steering a turn
that is genuinely running is unchanged.

* fix(native-chat): a message sent while Stopping is queued as a card, whatever the setting

While the chat reads Stopping (the host's flag or this client's own Stop in
flight) there is no turn left to steer into, so the desktop asks the host to
queue the send even with the queueing setting off, and it is never drawn as a
bubble inside the turn being stopped. The phone already queued every send on a
capable host; a test now pins that it does so while Stopping.

* fix(native-chat): a message queued while Stopping is a card at once, not after the stop

A person's Stop holds the session's lane until Codex answers its interrupt, and
a send was admitted only behind it. By then the turn read ended, so a send that
asked to be queued went out plain: no card for the whole of Stopping, then a
bubble and a new turn.

While the host reads that a person's Stop is ending the work, a text send that
asks to be queued is admitted without waiting for the lane: the same ledger and
lease admission, and a plan that only writes the card through the journal's
ordered writer. The card runs when the stop lands; the Stop's pause holds only
cards queued before it. Anything else, including a Stop that settled by the
time the send runs, takes the lane as before.

The Codex test fake now drops the active turn when it takes an interrupt, as
Codex does before it answers, so a turn/start after the answer opens a new turn.

* fix(native-chat): a Codex Stop that ends the child before any turn opened withdraws its send

A Stop on a Codex turn that was answered but never opened ends the Codex
process. That end settled the send as in doubt (unknown, recovered), and the
client's outbox holds every later send behind a send in doubt until the person
presses Retry, which re-sends the very message they stopped. The chat looked
stuck.

Codex records a prompt only once its turn has started, so a send whose turn
never opened never ran. When the Stop's refusal says so (turnMayOpen), the child
end now settles the unanswered sends as withdrawn, the verdict Codex's own
interrupted-turn end already gives an unechoed send. The flag rides on the owed
wind-down, so a retry after a failed child end withdraws them too. Claude's
child end still leaves its unanswered send in doubt.

* fix(native-chat): derive the withdrawal of a Codex send whose turn never opened

Replaces the flag the Stop carried to the child's end, and its copy on the owed
wind-down, with a reading of the journal at the settlement that lands. A Codex
child's unanswered send is withdrawn when a person's Stop is in force since it
was sent and no turn row ran, or was written, after it; any other end (a turn
that opened, a host's close, a crash, Claude) still leaves it in doubt. A
retried wind-down reads the same rows, so it withdraws the same sends.

* fix(native-chat): a card queued while Stopping runs past the cards the Stop holds

A card queued before a person's Stop waits under its pause until Resume. One
queued after it, as a message sent while Stopping now is, was stuck behind them
too, since the queue never reorders. Such a card was asked for after the Stop,
so it runs when the stop lands, past the cards held only by that Stop's pause;
a returned card and the restart and /clear pauses still hold everything behind
them.

Also: the host's own Stopping reading is gated on working, as the published
flag is, so a failed Stop's mark never reads Stopping on an idle session; a send
that falls back to the lane re-reads the conversation's journal there; and the
end-to-end test asserts the queue's pause rather than a per-card field.

* fix(native-chat): a paused queue labels only the cards it holds

Since a card queued after a person's Stop runs past the cards the Stop holds,
labelling every card "paused" while the queue's pause is published misreads that
card. The host now marks each card its pause holds (heldByPause, a new optional
field), and the desktop and phone label only those. An older host marks none,
so a client keeps today's labels; an older client ignores the field.

Adds a test of the desktop's own send through the real outbox: while the chat
reads Stopping, the request asks the host to queue it and no bubble is drawn,
on a host that advertises the queue.

* revert(native-chat): defer the per-card queue pause label to the queue's rollout

The heldByPause field and its labels are visible only where the host
advertises the queued-messages capability, which shipped hosts do not yet do.
Deferred to that rollout; the real-outbox send test stays.

* fix(native-chat): while Stopping, say and show what a send does where the queue is dark

Shipped hosts do not advertise the queued-messages capability, so a message
sent while Stopping goes out plain: the host holds it until the stopped turn
ends and then runs it as its own turn. The composer still said "Queue a message
to run after the stop", and the message was drawn inside the turn being
stopped.

Now the placeholder reads "Send a message to run after the stop" where the host
does not queue sends, and "Queue a message…" only where it does (desktop and
phone, all six catalogs). A send this client made that the host has not
recorded yet is drawn after the Stopping line while the chat reads Stopping, as
a message held behind a running command already is; once the host hands it
over it opens its own turn. Client presentation only.

* fix(native-chat): keep a send in doubt when a turn was open for it

The derived withdrawal read a turn as open for a send only if it still ran or
was written after the send. A send steered into a running Codex turn whose
interrupt failed met neither once the adapter's end settled that turn ahead of
the host's settle, so it read withdrawn, though Codex drains a steer into the
running turn and may hold it. A turn that ended after the send was handed over
was open for it too: such a send stays in doubt, as before.

Pins that case, and that a send made after the Stop, to a child that then dies
before its turn opens, stays in doubt.

* fix(native-chat): restore the per-card queue pause label

Kept after all: a paused queue labels only the cards it holds (heldByPause),
which is visible only where the host advertises the queued-messages capability.

* fix(native-chat): draw only a send made while Stopping after the Stopping line

Every send the host had not recorded yet was drawn after the Stopping line,
including one made just before the Stop, which the host steers into the turn;
it then jumped up into that turn once recorded. The outbox now marks a send
made while the chat reads Stopping, and only those wait after the line.

* fix(native-chat): withdraw a Codex send by whether it started its own turn, not by timing

Whether a turn was open for a send was read from end times: a turn that ended
after the send's handover counted. A send made while a Stop ended the turn is
handed over once that turn reads ended, yet Codex's own end for it can arrive
later, so such a send whose own turn never opened read in doubt again, and the
chat's queue held behind it.

The handover already records where the send went: its message joins the turn
running then (a steer) or belongs to no turn (it starts its own). Only a send
that started its own turn, with none opened since, is withdrawn; one that
joined a running turn, or has no recorded place, stays in doubt.

The Codex test fake takes an answered interrupt as Codex does, dropping the
turn before its end arrives.

* refactor(native-chat): move queue-while-stopping to its own follow-up

The queued-messages capability is off on every shipped host (#21062), so the
parts of this PR that act only when it is on move to a follow-up stacked on
this one: admitting a queued card while a Stop holds the session's lane, a card
queued after a Stop running past the cards it holds, the per-card pause mark,
and asking the host to queue a send made while Stopping. This PR keeps the
Stopping state, the host's steer hold, the rule that never steers a turn whose
interrupt Codex answered, and what a send while Stopping looks like where the
queue is off.

* fix(native-chat): leave no Stop row when the Stop took back a send that never ran

A Stop on a Codex send whose turn never opened ends the child, and the child's end
takes the send back into the composer. The Stop still wrote "Cancellation
requested." at the conversation level, so with the send gone it sat under the
previous finished turn and read as if that turn had been stopped. A Stop that found
no turn running and whose child end took back every send it found now writes no
row; a Stop of a running turn, or one that leaves a send in doubt, still does.

* fix(native-chat): count a send whose answer was lost when a Stop takes it back

The no-row rule counted only pending sends, but the child's end also takes back a send this process left in doubt when Codex's turn/start answer was lost. That case still wrote "Cancellation requested." under the previous turn. Both now read one predicate, so they cannot drift apart.

* test(native-chat): pin which sends a Stop's child end can take back

A send an earlier process left in doubt is never withdrawn and never holds the row back, and a queued card's send is never counted.

* fix(native-chat): hold the next handover while the send ahead opens its turn

The delivery loop handed queued messages to the provider back to back. Codex steers a second send into the first one's turn once it opens, and Claude folds it into the running cycle, yet its handover row was written before that turn existed, so it read as belonging to no turn and was drawn ahead of the reply. The loop now hands the next message over only once the send ahead has opened its turn, settled, or been stopped, all read from the journal; each is a commit, which wakes the loop again. The message then goes in scoped to the opened turn.

* fix(native-chat): doubt a Codex send whose turn never opened once Codex goes idle

A Codex before 0.148 fails a turn before opening it with only an error. The send stayed pending, and the hold behind it waited on it. When Codex reports its thread not running with no turn open, a send answered into a turn it never opened or ended now settles as doubt with the existing idle reason, which releases the hold. No clock: a slow but healthy turn start is never doubted.

* fix(native-chat): end an unopened Codex turn on its final error, not at idle, and drop the Stop release

Codex publishes the thread idle ahead of turn/completed, and between an aborted turn and the next picked one, so releasing at idle doubted sends in turns Codex did open or was about to. A final error naming a turn Codex never opened (its only end before 0.148) now ends that turn as failed instead, settling its send as a turn/completed failure would. The test fake publishes idle before turn/completed, as Codex does. A Stop no longer releases the hold: the seven rig tests that needed it now settle the stopped send the way the provider does.

* docs(native-chat): put each turn-end settlement comment on its own function

* test(native-chat): echo the first send before the next in the restarted-child test

The test's child admitted the first message and never answered it, which no live provider does; the next send was then held behind a turn still opening. The child now echoes the first message, so the test still checks that the next send restarts nothing.

* fix(native-chat): draw a message held behind an opening turn after that turn's live status

While the send ahead was still opening its turn, the live turn was taken to be the newest user message, the held one, so its live status drew under the held message and the held message drew above it. While a send is opening, its turn is now the live one, and messages sent after it, queued or not yet recorded, wait behind the live turn as a message held behind /compact or a Stop does. The host's hold and the client's drawing read one shared predicate.

* fix(native-chat): keep held messages waiting while Stopping, and leave Retry-only sends in place

While Stopping, a message typed then returned early from the waiting rule, so a message held behind the opening turn moved back above the Stopping line. The rule now waits the union of both. A send only the user's Retry sends again is not held by the host, so it no longer waits behind an opening turn; the projection marks it.

* fix(native-chat): keep a turn on the send that opened it when Codex echoes a steer first

A message held behind an opening turn is steered in moments after that turn opens, and
Codex can echo the steer before the send that opened the turn. Both carry the turn's one
provider key, so the steer was read as the turn's opener: for that moment the live status
moved under it and restarted its clock. While the opener is still in flight ahead of the
turn record, a steer into that turn no longer takes it.

* fix(native-chat): never anchor an opening turn on a message still queued above its send

Two messages queued behind /compact, or sent while Stopping, are accepted above the first
send's handover, and the hold makes the first send's turn record land before the second is
handed over. The turn fell back to the first unechoed send ahead of its record, which was
the queued one, so its live status moved under it. A send still queued is not in flight.

The phone keeps no outbox, so its frame tests now draw only recorded rows, through the
phone's own fold rather than the desktop's transcript order.

* fix(native-chat): read the host's Stopping beside main's startup phase

Main now reads only the startup phase from the status feed and no longer publishes which
child is starting. The chat reads the host's Stopping from its own hook beside it, and the
Stopping bridge test mocks the execution-host lookup main's owner resolution now calls.

* test(orchestration): open the working send's turn before a `now` send joins it

The rig's working send was handed over with no turn record, so the running turn the test
names never existed; the hold rightly kept the `now` send until that turn opened. The setup
now opens it as a provider does. The assertions are unchanged.

* test(native-chat): pin the phone's own echo of an accepted send behind an opening turn

The phone keeps no outbox, but it does show its echo of a send the host accepted until that
send's row arrives, and the shared waiting rule moves that echo behind a turn still opening.
The phone test now builds its list as the phone's view does, echoes included, and covers it.

* test(native-chat): read the outbox reconcile from where main moved it

* test(native-chat): pass the projection's options after main's rejected-in-place rows

* fix(native-chat): a retried message no longer waits behind a later Stop

Retry dropped the Stop it had outlived but kept the mark that it was sent while
a Stop was ending a turn, so a retried message waited behind whatever later,
unrelated turn a Stop was ending. Retry is a new send: drop that mark too.

* fix(native-chat): word a send after a Stop by whether this send will queue

The 'queue a message to run after the stop' placeholder read the host's
queue capability alone. A send queues only when the host queues and this
send asks it to: the queue setting is on and no pending prompt blocks the
queue. Desktop and phone now word the placeholder from that same decision
their send uses.

* test(native-chat): one test per case for the words of a send after a Stop

* refactor(native-chat): the dictation hook owns the composer's dictation state

Keeps NativeChatComposer within its line limit after the Stop props and
main's /context answer both landed in it.

* refactor(native-chat): name the dictation hook for what it owns now

* fix(native-chat): a Stop's note says it took once a joined close proves the exit

A session-ending Stop whose child's end failed revises its note to
'unconfirmed'. Since main's #24862, the next Stop joins that close rather
than stopping again, and wrote no note, so a close that then proved the
exit left 'unconfirmed' under a turn that ended. The close now carries
the note it settles, and its proven exit revises it to 'Cancellation
requested.'. A join that fails again leaves it unconfirmed.

* fix(native-chat): a proven turn end says a Stop's unconfirmed note took

Replaces the note carried on the child's close. Every settlement that
ends turns interrupted on a proven exit, live or after a crash, also
revises an unconfirmed Stop note on those turns to 'Cancellation
requested.', found by the note's turn scope, in the same batch. A note a
Stop wrote before its turn showed is re-keyed onto the running turn when
it becomes unconfirmed, in one batch, so that end finds it. Known limit:
with no turn open yet, the note keeps its key and no turn's end revises
it.

* test(native-chat): a Stop's unconfirmed note says it took when the agent exits on its own

* test(native-chat): a message that joins a Stop's unproven Claude close goes to the resumed child, the next waits for its echo

Covers the case main's #24862 retired with its unproven-stop test: the
first message after an unproven close joins it and reaches the resumed
child, and a second waits for that message's turn to open.

* test(native-chat): name the Codex handle as main's opaque handle does

* test: restore the provider handle import the main merge dropped

* fix: derive Stop note wording from interrupted turns

* test: name the raw replay case for what it covers

* refactor(native-chat): move the waiting-slot split into its own hook

* test: give the opening-send hold host the agent registry main now requires

* test: follow main's chat font-size rename in the stopping tests

* test: follow main's chat font-size change in the opening-send test

* test: follow main's single live-line value in the Stopping tests

* test: give the android live-line fixtures the stopping field

* chore: keep the session host under its line limit after the main merge

* chore: keep the composer test and the phone chat view under their line limits after the main merge

The Stop control now disables itself while Stopping, so the composer passes
the flag through and its test file stays as main has it. The phone chat
header's Stop moves to its own component.

* perf(native-chat): read a Stop note's fields before parsing its key on every snapshot

Every snapshot projects each item through the Stop-note read, and each new
snapshot rebuilds the index of Stop notes by turn. Both parsed every item's
key first; they now check the row's kind and turn scope (and, for the
projection, its unconfirmed-stop failure) before the parse. Every Stop note
is a status row, so what each finds is unchanged.

* fix(native-chat): a Stop ends the hold on the next message, even when the stopped send's turn never opened

* test: a Stop's note ends the hold on the next message

* fix(native-chat): only a Stop that took ends the hold; a refused or unconfirmed one leaves the turn opening

* fix: import the moved Stop-note helpers where the host module uses them; type the test's failure note

* test(native-chat): pin that a taken Stop's note lands after the turn the provider opened before answering it

* test(native-chat): give the opening-send hold test's runtime the launch arguments main now requires

* fix(native-chat): a message a Stop takes back while the turn ahead opens stays after that turn, with its stop row; move the rows that wait behind the live turn into their own module

* test(native-chat): run the two opening-send hold tests, which open a real journal database, in the Node runtime project

* test(native-chat): with a send still opening its turn, a starting child gets only the first message; a failed start still rejects both in its words
2026-10-06 22:57:51 -07:00
Brennan Benson 988f87e743 fix(native-chat): move the attachment failure words to their own module so lint passes on main (#26079)
The failure-words table grew past the 300-line limit when two changes landed back to back,
which fails Lint on main and makes every PR's static-analysis job skip typecheck.
2026-10-06 22:49:50 -07:00
Jinwoo HongandClaude 358e4f92e0 fix(claude): share account history on Windows with junctions and a same-drive hardlink (#26067)
* fix(claude): share account history on Windows with junctions and a same-volume hardlink (STA-3698)

Windows skipped history sharing for Orca-managed Claude account folders on
the mistaken premise that it needs symlink privilege. Session folders now
use directory junctions and history.jsonl a same-volume hardlink, neither
of which needs elevation. A different volume or a failed link keeps that
account's prompt history private and reports it. A small link record tells
a replaced shared file apart from the account's own copy so prompts the
user removed are never merged back. Renames retry on Windows file locks.

* fix(claude-accounts): refuse a cross-volume Windows share before draining set-aside prompt history

A pending set-aside copy was appended to the shared history.jsonl before the
same-volume check refused the share, leaking the account's prompts into it.

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-10-07 01:43:00 -04:00
64bb9373da Claude account profiles: dormant WSL guest setup (Step 3 of 4) (#24384)
* feat(claude): add dormant profile setup and history sharing

* fix(claude): make profile setup one gated, typed, fail-safe entry

Review round 1 of the dormant profile setup found that the pieces could
be called without their safety checks, that one failed write or an
unreadable bookkeeping file could silently stop sharing for good, and
that Windows prompt history could bring back history the user cleared.

- One entry, provisionClaudeAccountProfile: the profile gate (namespace,
  no linked components, outside ~/.claude and ~/.config/claude, and an
  ownership marker beside the home naming the account and target) runs
  first and refuses before creating anything; then history sharing,
  config provisioning, and the hook install after the settings merge.
  Results come back per surface with closed warning codes instead of
  message text.
- The sharing ledger is keyed by surface name, records a value only
  after its write succeeded, and an unreadable ledger starts empty and
  is rewritten instead of blocking every surface.
- The profile state file goes through the same locked writer as folder
  trust (Claude's <file>.lock plus the in-process queue), generalized as
  updateClaudeGlobalConfig. Onboarding and trust are still applied when
  the personal state file is unreadable.
- WSL descriptors build guest POSIX paths; the state-file path style
  follows the injected platform.
- Orca's managed statusLine has one owner in a profile: the settings
  merge never shares it, a user's own statusLine is shared over it, and
  the profile installer follows the default home's slot so a default
  opt-out reaches every profile. remove() takes the same destination;
  the remote installer cannot accept one.
- Prompt history compares file identity (bigint dev+ino) on every
  platform, never drains the shared file into itself, drains retained
  copies in generation order, never reuses a stale cursor, and on Windows
  keeps a replaced default's old copy aside instead of replaying it.
  Directory merges keep going past a failed entry.

* fix(claude): share the user's own hooks and keep merged history whole

A user's own Claude hooks in ~/.claude (notifications, formatters) did
not run under a managed account, because the whole hooks key stayed
private. They are now shared like any other settings key: Orca's own
hook entries and its managed statusLine are stripped from both the
personal value and the profile's current value before the per-key
ledger comparison, so they never travel through the merge and never make
the key look user-owned. Orca entries already in the profile are kept on
write, and the profile hook installer adds them on top as before.

Prompt history: merged bytes that lack a final newline are terminated,
so Claude's next record no longer fuses onto the last merged line. When
a CLI rewrote the profile's history file (old records plus new), only
the lines past the part it shares with the default history are added,
instead of the whole file again.

* fix(claude): close review round 2 gaps in profile setup

Hooks and statusLine sharing:
- When ~/.claude holds only Orca's hook entries, the user's shared hooks
  now read as an empty value instead of a missing key. Removing the
  user's last own hook in ~/.claude therefore reaches profiles that
  never edited it, and deleting the only shared hook inside a profile
  stays deleted.
- A custom statusLine Orca shared, and the profile never edited, goes
  away when the default home drops it. When a shared custom line
  replaced Orca's line in a profile, the profile's statusline marker is
  dropped so Orca's line comes back once the default returns to it; a
  profile that opted out stays opted out. No other key gains deletion.
- install/remove/getStatus with a profile directory refuse when it is
  the default home, or its settings.json resolves to the default one,
  instead of editing System Default's hooks and opt-out state.
- The profile statusline rule reads the default settings under the
  userHome passed to the setup entry, not os.homedir().

Profile state and ownership:
- A malformed `projects` value skips only folder trust (new warning
  code trust-refused); onboarding and shared keys still apply.
- The ownership marker stores only host-local facts (account, runtime,
  distro). The execution host id is the caller's view of the host, so
  it stays in the in-memory descriptor and is not compared.

Prompt history interruption paths:
- With no cursor yet, a retained copy starts past the bytes it shares
  with the default history, so an interrupted share no longer replays
  the whole history.
- A retained name for the shared file itself is removed with its cursor
  instead of lingering until a later scrub makes it look new.
- The Windows link record is read three-state: unreadable stops the
  share instead of reading as "no link". If the record cannot be
  written after linking, the fresh link is undone.
- An unreadable retained copy is reported and no longer blocks linking.

* build(cli): list the new Claude hook modules in the CLI project

hook-service.ts and hook-settings.ts are compiled into the packaged CLI
project, which lists every file explicitly. The statusline policy and
profile destination modules they now import were missing, so the CLI
typecheck failed with TS6307. The CLI still loads hook-service through
the existing managed-agent-hook-controls build entry, which bundles
both modules; neither imports electron.

* fix(claude): close review round 3 regressions in profile setup

- A profile whose hooks hold only Orca's entries and that sharing never
  recorded is no longer treated as a user edit, so the user's first own
  hook in ~/.claude reaches it (for example when the profile was set up
  before ~/.claude had any hooks).
- A retained prompt-history file is removed as a second name for the
  shared file only when the default history does not itself link to it;
  otherwise it holds the only copy and is kept.
- Default-home checks compare file identity: the profile hook
  destination check uses device and inode, and the profile/default
  separation check resolves on-disk case, so a case-only alias of
  ~/.claude is refused on case-insensitive filesystems.
- A test pins that an unreadable leftover session tree no longer blocks
  linking.

* fix(claude): let shared keys leave a profile when ~/.claude drops them

QA found that removing a setting from ~/.claude never reached a managed
account: deleting the whole `hooks` block left the user's hook running
there. Only statusLine followed the default away.

Every shared key now follows the same rule through the existing per-key
ledger: when a key disappears from ~/.claude/settings.json (or
mcpServers/theme from the personal state file), it is removed from the
profile if the profile still holds exactly what Orca last shared. A
value changed inside the account is kept. Keys Orca never shared,
including denylisted ones, are never touched. Deleting the whole hooks
block removes the user's shared hooks and keeps Orca's own entries. A
missing source counts as empty; an unreadable source removes nothing.

* feat(claude): add dormant profile routing and account consumers

* fix(claude): drop the dormant profile selection RPC; clients negotiate by capability

Restores the inline mobile allowlist so its source-scan guard sees every
accounts.* method again, and the generated params catalog to generator order.

* fix(claude): guard the claude shell function and honour a hand-exported config dir

The function is defined only in a routed pane where claude is a real
executable (the codex function's guard), re-reads the pointer only while
CLAUDE_CONFIG_DIR is unset or still Orca's injected twin, accepts Git Bash
drive paths, and starts on its own line after the fish/PowerShell codex text.

* fix(claude): spawn-time profile env, total account listing, setup at lifecycle triggers

Round-1 review fixes for the dormant profile routing:
- Panes get the selected profile's CLAUDE_CONFIG_DIR plus an Orca twin at
  spawn, so nested shells and scripts inherit the account; System Default
  injects nothing and its home is the inherited CLAUDE_CONFIG_DIR.
- An absent routing owner is System Default, never a throw; AI Vault and
  session-search scans receive profile roots from their parent, and the
  capability is advertised only where an owner is installed.
- Account listing never throws: per-account readiness, a stale pointer is
  republished in the background and reported on the snapshot.
- Profiles are set up at select and startup; a launch only sets up one that
  never was, and a worker fault on a prepared profile is a warning. The
  Claude version probe is cached per binary identity.
- Pre-trust goes through the existing deadline- and realpath-guarded writer
  against the launch env's profile config.
- Skill discovery keeps a caller's Claude root and a broken Claude selection
  no longer fails other providers.
- The durable record carries a provider-neutral launchAccountHome, read
  through one helper by the launch fallback and the model catalog.

* test(claude): pin the version-probe cache, launch-account record and temp-home readers

* test(claude): pin dormant bash rc text alongside fish and PowerShell

* test(claude): read the fish launch init without a nullable index

* fix(claude): withdraw the profile pointer when a selection cannot be published

A pointer left naming the previous account would launch it silently; a
missing pointer makes the claude function refuse visibly. A newer selection
that raced the failed one keeps its pointer.

* fix(claude): read the fish profile pointer with read -z for fish older than 3.4

Shell tests skip system config and abort unless claude resolves to the fake.

* fix(claude): only the newest publish withdraws the pointer; total config dir lookup

- An overtaken publish that fails leaves the newer selection's pointer.
- The runtime config dir falls back to the legacy home for an unresolvable
  account or a WSL target, so skill roots never fail for other providers.
- WSL guest reader roots merge verbatim, never realpathed on this thread.
- History readers include ~/.claude, where step-1 setup pools profile history.
- System Default ignores a config dir an outer Orca injected (twin-marked).

* fix(claude): System Default launches and probes use the structured create resolver

A Claude agent-env CLAUDE_CONFIG_DIR the create path stored is now the home
the launch pins and the model probe accepts.

* test(claude): type the System Default launch record as an agent-session record

* fix(claude): install profile hook scripts under the setup job's home

A worker thread's os.homedir() ignores its own env, so the hook and
statusline scripts now go under the home the job names. The worker test pins
the process HOME to a sentinel, refuses to run unless the worker sees it, and
asserts nothing lands there.

* test(claude): skip shell cases whose shell the runner lacks

* fix(claude): remove env vars in the PowerShell claude function instead of setting null

On .NET 9+ (pwsh 7.5+) SetEnvironmentVariable with $null creates an empty
variable, so stripped auth vars reached claude as empty strings and the
restore left CLAUDE_CONFIG_DIR empty in the user's session.

* feat(claude): add dormant WSL guest profile setup

* fix(claude): open WSL panes without guest calls and coalesce same-profile publishes

A WSL pane now gets the same non-throwing, guest-free spawn env as a host
pane; only select, startup and Claude launches publish into the guest.
Overlapping publishes of one target share the newest publish while the
selection still names the same profile, instead of failing as superseded.
Publish issues name their WSL distro and drop out when the target is no
longer routed. A late inspect from an older selection no longer replaces
the newer one's verification, a failed guest request evicts the cached
guest, and readiness is derived per account from the guest's owned homes.

* fix(claude): roll back only the target whose selection failed

With profiles, a failed select or remove republishes just its own target
instead of running startup over every WSL distro, and a rollback failure is
logged instead of replacing the error that caused the rollback.

* fix(claude): scan WSL profile history only in running distros

Vault and usage scans pass Claude profile roots through the same
running-distro filter as every other WSL root, so a stopped distro's UNC
paths are never walked.

* fix(wsl): ship the Claude profile helper only in the WSL bundle dir

The helper only ever runs inside WSL from the desktop, so it moves out of
the SSH relay artifacts (no upload, no relay version change) into
out/relay/wsl beside the other WSL-only guest bundles. The three WSL bundle
resolvers share one candidate list.

* fix(wsl): refuse old glibc before downloading, and keep the shared download per caller

The pinned Node runtime needs glibc 2.28, so a distro below the floor is
refused before any download with a message naming both versions, as SSH
hosts are. The shared download again owns its own deadline and each caller
waits on its own signal, and the OpenCode reader keeps its architecture
error text.

* fix(claude): bound each WSL guest operation and run the helper through the WSL runner

A cached guest no longer carries its 180 s preparation deadline into later
requests. The helper runs through runWslProcess (stdin payload, WSL_UTF8),
the distro is confirmed running once per preparation and once per request,
a failed `claude --version` probe continues with an unknown version like
native setup, the helper resolves from the WSL bundle dir, and the guest
entry decodes stdin once so split UTF-8 survives.

* test(claude): cover WSL profile pre-trust routing and its deadline

* refactor(claude): drop WSL refresh cleanup that the failed publish's withdraw already does

* fix(claude): catch rollback failures only when profiles route the selection

With the gate off, select and remove surface the rollback error exactly as
before; only profile routing logs it and keeps the original error.

* fix(claude): give every WSL pane a guest-relative Claude profile pointer

WSL panes now always carry `~/.local/share/orca/claude-profiles/selected-wsl`,
which the bash/zsh and fish claude functions expand against the guest $HOME
at each invocation, so a pane opened before Orca has met the distro still
follows the selected account instead of falling back to ~/.claude. Absolute
pointers are untouched, PowerShell is unchanged, and a missing pointer file or
profile still refuses visibly. CLAUDE_CONFIG_DIR is set at spawn only when the
selection resolves without a guest call.

* test(claude): assert a missing guest-relative pointer refuses with a visible message

* test(claude): type the WSL runner mock in the transport test

* fix(claude): route only WSL distros that hold an Orca account, and re-derive their publish

A WSL distro is routed only while host settings hold an Orca Claude account
for it, decided from settings with no guest call. An unrouted distro behaves
as before profiles: its panes get no pointer or profile env, and a Claude
launch is System Default with no guest prepare. A distro that loses its last
account has its pointer withdrawn best-effort so older panes stop launching
the removed account.

A routed distro without a current publish (for example stopped at startup)
gets one non-blocking background publish from its next pane spawn, coalesced
per target; its failure stays that distro's issue and a later success clears
it. A late setup result from an older publish no longer replaces the newer
selection's verification. The owner contract moves to its own module so the
routing service stays under the size limit.

* fix(claude): read WSL profile history in native chat and adoption only in running distros

Native chat resolves Claude transcripts from host roots first and reads WSL
profile roots only after a miss, filtered to running distros like Codex's WSL
homes. Structured adoption candidates go through the same filter.

* fix(claude): target registration rollbacks and keep their errors in profile mode

A failed add or re-authentication rolls back only the account's own target.
With profiles, a failed re-authentication rollback is logged instead of
replacing the original error; with the gate off both behave as before.

* fix(claude): spell the guest pointer location once and keep set -u safe

The guest helper, the withdraw script and the pane pointer all derive from
one home-relative constant, and the posix claude function reads ${HOME:-}
so `set -u` with HOME unset refuses cleanly instead of aborting.

* test(claude): cover the IPC preflight and daemon WSLENV paths for WSL profile env

The renderer preflight is tested for wsl.exe and Windows shells with a \\wsl$
cwd (which always launch wsl.exe) and with the gate off, the daemon launch
plan imports the pointer and profile home without a WSLENV flag, and the
Windows launch test uses the guest-relative pointer production sends.

* fix(wsl): report why the guest runtime failed, with download context and trimmed stderr

The install's promote output is classified with the SSH classifier, so a
self-test failure shows the exit code and the loader's words (for example a
missing libstdc++ on Alpine) and a security-software change is named. A failed
runtime download says it was Orca's Node runtime for WSL, while a checksum
mismatch keeps its own text. Guest stderr is trimmed before it reaches a
refusal message.

* test(claude): pin that pointer retirement never runs for host targets or with the gate off

* test(claude): give the routed WSL preflight fixture its required authMethod

* fix(claude): let the pane-triggered WSL publish repair a distro stopped at startup

"Distro not running" is now a typed refusal: it never withdraws the pointer
(the distro's last pointer cannot be stale, and a withdraw racing the boot
could delete a valid one) and never records a distro issue. The background
publish a pane fires now waits a few seconds for the pane's own spawn to boot
the distro, probing three times, and is dropped silently and re-armed if the
distro stays down. It joins any publish already in flight for that target
instead of preparing the guest a second time. Per-target generations and
pointer-write ordering move to ClaudeProfilePointerQueue so the routing
service stays under the size limit.

* fix(claude): remove the last selected WSL account without a guest publish

With profiles, removal writes the account list and the selection in one
update, so a distro losing its last account is already unrouted when it syncs
and its pointer is retired best-effort. Removal no longer needs the distro to
be running or able to run Orca's runtime. The gate-off order is unchanged.

* fix(claude): keep native chat's legacy Claude roots first and unfiltered

Only roots added by WSL profiles are read after a miss and filtered to
running distros; a host CLAUDE_CONFIG_DIR on a \\wsl$ share is searched first
and unfiltered, as before profiles.

* test(claude): cover stopped-at-startup repair, launch join and last-account removal end to end

* test(claude): assert no running probe before the pane has had a turn to boot the distro

* fix(claude): let user-initiated profile work boot an idle-stopped WSL distro

WSL distros idle-stop on their own, and the legacy path boots them with its
spawn or \\wsl$ write. With profiles on, a Claude launch, a select, a remove,
a failed-change rollback and the retire after removing a distro's last
account now skip the running pre-check and let their first bounded guest
command (`wsl -d <distro> --exec ...` through runWslProcess) boot the
distro. They refuse only if that command fails, with wsl.exe's own reason,
for example a distro that does not exist. Startup, the pane-triggered repair
and the history readers keep the running pre-check and its typed refusal, so
background work never boots a distro. With the gate off nothing changes.

* fix(claude): let startup join a launch or select already publishing a WSL distro

Startup no longer overtakes a user's in-flight publish for the same target,
so a launch that is booting an idle-stopped distro is not handed startup's
"not running" refusal.

* fix(claude): remove accounts of a WSL distro that no longer exists, and name the helper once

wsl.exe's own failures (exit 0xFFFFFFFF, empty stderr, the diagnostic and its
WSL_E_* code on stdout) are now read by one shared reader used by the git
runner and the WSL profile transport, so profile refusals show wsl.exe's
message. WSL_E_DISTRO_NOT_FOUND becomes ClaudeProfileHostMissingError: with
profiles, removing an account from a distro that no longer exists keeps the
removal and logs a warning, while select and launch still refuse visibly.
The helper's file name is defined once in shared/relay-artifacts.ts and used
by the relay build and the transport.

* fix(claude): give plain fish tabs the claude function through the codex hand-off

Main now gives a plain fish tab Orca's codex function through a vendor_conf.d
snippet instead of a -C init. The claude function only rode the -C path, so a
plain fish tab would not re-read the account selection per invocation once
profiles are on. Define it at the first prompt beside codex; it stays empty
while the profile gate is off.

* fix(claude): share personal rules, themes, workflows and keybindings into account profiles

A managed account launches Claude with its own config folder, so user-level
rules/, custom themes/ (which a shared `custom:<slug>` theme points at),
personal workflows/ and keybindings.json silently stopped applying. Link the
three directories like skills and commands, and copy keybindings.json with the
same edit-preserving ledger as CLAUDE.md. routines/ stays unshared: routines
belong to the claude.ai account and the folder holds per-run state.

* test(claude): wait for the running child to read its account before switching

The test switched the selection after a fixed 20 ms, so under load the backgrounded claude
had not yet read the pointer and picked up the new account. The stand-in now marks when it has
started, and the test waits for that mark (bounded) before switching.

* fix(claude): accept WSL setup warnings for every shared Claude file

The guest reply schema listed CLAUDE.md by name, so a warning about the newly shared
keybindings.json would have rejected the whole reply. It now takes the shared-file list
from provisioning, like the shared folders.

* fix(claude): import the personal CLAUDE.md into account profiles instead of copying it

Claude also loads ~/.claude/CLAUDE.md as a parent folder's memory for any project under home,
so a copied account CLAUDE.md made every such session read the user's instructions twice
(checked live with Claude 2.1.288). An @~/.claude/CLAUDE.md import resolves to the same real
file, which Claude loads once from home, from projects under home and from folders outside it.

* refactor(claude): simplify account profile setup toward the prior art

- Windows keeps each account's history private; drop the hardlink, link
  record and conflict-copy machinery that only Windows reached.
- Share hooks and statusLine as ordinary settings keys: Orca writes the
  same entries into every folder, so the installer finds them present.
  Drops the Orca-entry carve-out, the per-profile statusline follow
  logic and its marker.
- Unreadable ledger is just an empty ledger.
- Share from the user's own CLAUDE_CONFIG_DIR when they set one (marked
  so Orca's injected value is never mistaken for it), and refuse a
  profile at or around it.
- Pin the one canonical profile path spelling in a test.

* refactor(claude): route launches through one account router, superset-shaped

Replace the routing service, owner interface, setup worker thread, reader-root
merging, persisted launch account and capability string with one
ClaudeProfileRouter: the pointer is written first and setup runs best-effort
after it (superset's order); a missing pointer means System default.

The claude shell function re-reads the pointer on every launch, refuses only
a selected account whose folder is missing, and prints a note when the user's
own CLAUDE_CONFIG_DIR overrides the selected account in that terminal.

Still dormant: claudeProfileRoutingEnabled() is false.

* test(claude): type router test settings instead of casting

* fix(claude): run account setup on a worker thread, never Electron main

publish() writes the pointer and starts setup in the background, so neither
startup nor an account switch blocks on a history merge. Each setup runs in a
one-shot worker (the profile-state backup worker's pattern); one setup per
account at a time, reused by later requests. A launch waits only for a folder
that was never set up, and refuses with a clear message if that setup fails.

* fix(claude): do not await the synchronous pointer publish

* refactor(claude): route WSL distros through a small guest router on the Step 2 shape

Replaces the WSL owner/transport/guest-inspect stack with ClaudeWslProfileRouter:
publish writes the guest pointer with one sh command and kicks Step 1's setup
best-effort; prepareLaunch checks the folder over the distro share and waits only
for a never-set-up folder; preparation returns main's WSL shape, so trust, rate
limits and readers need no new code. Setup runs as Linux in the guest on Orca's
pinned Node via a bundled helper (argv in, exit code out), without hooks.

Restores OpenCode's WSL runtime prep, git's wsl-host-failure, wsl-runner,
workspace trust, readers and account selection/registration to Step 2.
Names the guest pointer per Orca build so dev and packaged never share it.

* test(claude): give the routing launch test the merged resolver deps and handle shape

* test(claude): type the WSL routing mock's original() without an inline import()

* fix(claude-accounts): dedupe merged prompt history, drop drained copies, link setup folders by path

- Prompt-history drain appends only lines the shared file lacks, so a purge never re-adds lines.
- A set-aside history copy whose saved offset reaches its end is deleted on the next run.
- Setup folders link to the default home's own entry, not its resolved target.
- The profile gate and folder creation run once, in provisionClaudeAccountProfile.
- installHooks receives only configDir; drop a duplicate test key that fails CI.

* fix(claude-accounts): refuse a routed resume whose transcript is in another account; zsh claude function; setup timeout

- With account routing, a chat resume checks its transcript is in the launch folder; a missing one
  with a stored leaf refuses with historyInOtherAccount instead of starting fresh.
- The launch folder of a selected account comes from prepareLaunch(); the resolver stays for System default.
- zsh panes get the claude function like bash, fish and PowerShell (empty while routing is off).
- The setup worker is terminated after 60 s so a later launch can retry.
- Document that the setup marker means setup started, not finished.

* fix(claude-accounts): write the WSL account pointer before a launch returns; one relay bundle candidate list

- prepareLaunch awaits writePointer, so a missing or stale guest pointer cannot run another account.
- Startup's WSL republish runs inside serializeMutation, like rollback.
- relayBundleCandidates takes 'wsl'; the hook relay, browser relay and Claude helper use it, and
  wsl-relay-bundle-dirs.ts is gone.
- One setup-marker path helper for host and WSL; the guest pointer path is home-relative and only
  the pane value carries '~/'; drop a no-op esbuild external.

* fix(claude-accounts): refuse a routed resume only when the transcript is found in another folder

A transcript found in no known folder keeps the old stored-leaf resume.

* fix(claude-accounts): a WSL launch writes the pointer for the selection current at write time; bound the pointer read

A selection made while a launch waited on setup was overwritten by the launch's stale account.
A hung \\wsl.localhost read no longer stalls startup's serialized publish.

* fix(claude-accounts): record installed hooks as Orca-shared; skip symlink tests on Windows

After Orca installs its hooks into an account, record the account's hooks in
the settings ledger so a later run can still bring the user's own hooks in.
Tests that create real symlinks now skip on Windows.

* fix(claude-accounts): trim the which-account file in the PowerShell claude function

Co-Authored-By: Claude <noreply@anthropic.com>

* test(claude-accounts): spell the user's own config folder as an absolute path on every platform

Co-Authored-By: Claude <noreply@anthropic.com>

* test(claude): skip the POSIX-only WSL profile test on Windows

A WSL profile's data root is a POSIX path, so building one from a Windows
temp dir fails the absolute-path check there.

---------

Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
2026-10-07 01:25:41 -04:00
Brennan Benson 2f49377425 feat(native-chat): Grok as a structured chat over the Agent Client Protocol (#25225)
* Leave a stopped turn's running tools to the agent's own end

When another writer settles the open turn (a person's Stop), the assembler
now only stops that turn's text and cancels its pending prompts. Running tool
calls stay the agent's: a progress update or completion it reports after the
Stop lands as reported, and whatever is still running settles at the agent's
turn end for that turn, the next turn's open, or the session's end.

An agent's end for an earlier turn while a newer one is open no longer clears
the open turn's activity line or ends its anonymous reply. An unnamed end right
after a Stop ends the stopped turn instead of being dropped. The test rig's
restart no longer writes the dead assembler's window text, matching dispose.

* Pin that a stopped turn's running tools hold budget until the agent's end

* Type the stopped turn's tool progress update as a tool body

* List every event the assembler hands to the decision step

The type-aware lint requires an exhaustive switch with no default case.
Also retitle a Stop test to say what it asserts.

* End a running call as its turn's journal row ends

A call still running when its turn ends takes the state of that turn's
row: a row another writer settled first (a person's Stop) stands, so its
calls read interrupted whatever the provider's later end reports. The
no-ending path that settled calls from the Stop row is gone, since a Stop
now leaves running calls to the provider. Adds the two Spanish strings.

* Say why a Grok turn failed, and keep task rows in Grok's own words

A failed Grok turn ended with no reason on screen: the translator dropped
every copy of Grok's message. The failed turn now gets one status row in
Orca's existing "provider did not accept this message" words with Grok's
reason, read from whichever copy arrives first (the given-up retry, the
turn's end, the prompt's completion notice, or the prompt's error answer);
later copies only fill a reason the row still lacks.

A running background command no longer reads "Background task <id>
started": a task's summary is mapped only once it has settled. A monitor
stays a monitor when the agent reads its output: a frame that names no
kind keeps the known one, and a "[monitor" command is a monitor.

A prompt's turn is marked started, so a late frame for an ended prompt
neither reopens it nor becomes the active turn. A tool's turn is held in
one place at a time.

* Read a monitor from Grok's exact output prefix

* Word a failed Grok turn in Grok's own text, not as a refused message

A turn that started and then failed was told "The provider did not accept
this message", Orca's sentence for a message refused before its turn. The
row now reads as a Codex turn-ending error does: an error status row with the
provider's own words. With no words, the dialect names the failure ("Grok
ended this turn with an error." / "Grok usage limit reached."), else the
agent's display name does.

* Settle a stopped turn's running call as its turn row ended after a restart too

The restart sweep ended every running call by the death evidence alone, so after
a person's Stop with no proof the child died the call read failed under a turn
that read interrupted. The sweep and the live dead-generation settlement now ask
the same rule the assembler does: a call in a turn already settled ends as that
row ended; only a turn still running leaves its calls to the evidence.

* Keep the dead-generation settlement under the line cap

* Register the ACP schema verify step in the PR preflight phase test

* refactor(agent-session): one required agent registry; declarations admit what they claim

A StructuredAgentRegistry, built once from the {definition, adapter}
registrations, is now a required host dependency and the router routes with
it. The adapter interface loses its optional router-only capabilities?() and
definition?(); every reader (options at rest, thread goal, rewind, the
/compact handover) asks the registry. A live session still narrows rewind
through the adapter, and the host combines declared and narrowed in one
helper. The registry refuses a registration that declares compact, a thread
goal or rewind without the adapter method behind it.

/compact is admitted by the declared capability, so an agent that declares
compact:false gets the commandRefused fact instead of a thrown error.

Also: the cut-turn notice names the agent from the catalog, create-support
builds the account home through agentSessionAccountHome, and the turn
status text older clients read names the session's agent instead of
defaulting to Codex (Claude/Codex text unchanged).

* Read ACP permissions, session events and prompt errors through the protocol client's own types

The translator now reads a permission request with the client's lenient reader, a session update
with its session-event reader, and takes only the agent's own error answer as a failed prompt's
reason, so an Orca-side error never reads as the provider's words. Tests cover protocol values
newer than this build.

* refactor(agent-session): the router applies the declared rewind itself

The router holds the registry, so its rewindSupport answers the owner's
declared rewind narrowed by the adapter, in one place no reader can bypass.
The host-side combining helper is gone; readers ask the adapter they hold.

* test(agent-session): register the agents the merged-in tests now need

The adoption replay test builds its host over this build's registered agents
instead of an adapter definition method, and the stored-form test passes the
stored agents the record guard now requires.

* chore(agent-session): the registry is the one lookup; drop the test-only empty capability record

* fix(agent-session): keep "dismiss all" restart offers dismissed for the desktop

The desktop renderer calls the restart methods as a paired runtime client
without the registered-agents capability, so it was handed a Claude/Codex
audience and its "dismiss all" took the scoped path: no dismissal fence and
no clearing of unwritten teardown witnesses, so a late teardown write could
bring a dismissed offer back.

The audience is now derived against this host's registry: a client that can
show every agent the host registered gets none, and dismisses exactly as
before (one fence for every offer). A client that truly cannot show some
agent still gets a final dismissal for what it sees: each cleared session is
fenced on its own (bounded, superseded by a dismiss-all fence) and only the
witnesses of agents it sees are dropped.

One predicate now answers which agents a client renders, for both tab
projection and restart offers; the Claude capability rule is folded in.

* fix(agent-session): a changed agent definition never hides that agent's chats

A record was readable only if every handle used the transport its agent's
definition declares today, and its account variable was declared by some
registered agent. A later build that drives the same agent over another
protocol, or under a renamed variable, would have set every existing chat of
that agent aside: no tab, no history, no error.

Readability now asks only that the record agree with itself: its agent is
registered, every handle shares one namespace owned by that agent, and the
account variable is a well-formed name. Whether this build can drive it (the
chain's transport is the one the agent speaks, and the account variable is
that agent's own, since it becomes the child's environment) is decided in
performAttach, the one admission every start of every agent passes through,
and refused there as hostUnsupported. A record pinning another agent's
variable is refused the same way rather than passing because some agent
declares it.

* refactor(agent-session): each agent's registration says where it runs and which account it pins

createSupport, create and the model catalog's account read still chose a
location rule and an account resolver by name (Claude, Codex, else none),
so a new agent would have needed a third branch beside the registration
list that already decides storage, routing and publication.

Each runtime registration now carries supportsLocation(location) and
resolveAccountHomePath({ launchEnv, location, purpose, workspacePath }),
and Claude's and Codex's rules move into their entries unchanged: a read
still syncs no home and starts no bridge, a Codex launch still trusts the
workspace first, and WSL locations resolve as before. The list exists
before the host is built, so none of these reads installs the host, and an
agent the list does not hold is answered no without opening the journal.

* test(agent-session): drop a duplicate registry key; keep fences out of the capsule state

A spread input already carries the registry, so the explicit key was
overwritten (TS2783). The session-fence helper now returns only the fence,
not the capsule state it was given.

* fix(agent-session): a scoped dismiss-all persists no per-session fence

The per-session dismissal fence added a forever-persisted capsule field
with an arbitrary cap that served no caller: only the local desktop calls
the restart methods, and on a host whose agents it all shows it gets no
audience and takes the unscoped, fenced dismissal. A scoped dismiss-all
(a caller that cannot show some registered agent) keeps that agent's
offers, clears only its audience's unwritten witnesses, and writes no
fence; a late write from this process is serialized behind it. The field
never left this unreleased branch.

* fix(agent-session): refuse an attach whose agent is not the session's own

The start check judges the record's agent while the router starts the adapter of params.agent, and the attach wire let the two differ. Admission now refuses agent !== provider as requestMalformed, so the agent checked is the agent started. Every host-built attach already sets them equal.

* fix(agent-session): offer to start a chat only when the start would accept it

The restart offer, a failure's retryable flag and the pre-send check asked only whether the adapter runs the chat's location, while the start also refuses a record this build cannot drive; such a chat was offered Resume and its Retry failed forever. One host predicate, hostCanStartRecord (location support and agentDrivesSession), now answers all four; adapterSupportsRecord stays the reader gate, and an undrivable chat's offer is kept, not retired.

* fix(agent-session): the model-catalog probe uses a record's account only if this build drives it

A record-scoped catalog read started the agent's lister under the record's account path whatever variable the record pinned. It now uses that path only when agentDrivesSession holds, and otherwise resolves the account as for a read with no record.

* refactor(agent-session): the record store admits agent ids; comments say where transport is checked

The store only ever asked whether an agent is registered, so its admission list is now the registered ids; the definition is the one home of an agent's transport and account variable. Comments that said the record store checks transport, and one that cited a nonexistent function, now point at agentDrivesSession and the attach admission. The isPersistedAgentSessionRecord(value, agents) call shape and the test fixture the cross-version probe imports are kept.

* docs(agent-session): the record store admits the registered agents' ids

* refactor(native-chat): Grok's registration declares where it runs; ACP no longer borrows Codex's location rule

The rule a self-supervised agent child runs under (this machine, no WSL, Windows only with process
start-time proof) is its own module that Codex and the ACP adapter both use. Grok's registration
takes the full account-home resolver signature, and D3's tests build hosts with the agent registry.

* fix(native-chat): Grok follows the ACP runtime's request contract and the managed process's close

A request the agent or a Stop cancels is answered with the agent's own cancelled reply by the code that
owns it (the runtime no longer answers a silent handler), so a Stop needs no separate decline pass. A
permission answer still being saved when the agent stopped waiting is reported unconfirmed, since the
protocol already answered it cancelled. Cancelling the agent's own turn is the plain cancel. Request
rows are matched under their generation-scoped ids. A refusal's reason comes from the dialect's wording
path. The child drops its own stderr tail and close policy for the managed process's, and a close
whose process tree was not proven gone is reported as the adapter contract asks.

* fix(native-chat): a Grok chat Orca already holds resumes without writing what Grok replays

A chat with a saved Grok session reattaches with session/resume where the agent offers it, else
session/load. Either way the call runs inside the translator's load window, so what Grok sends while
it reattaches (its saved exchange, a task the dead process left running, ended by the restart) opens
no turn and writes no row; only context usage reads on. A reply an Orca or Grok crash cut short is no
longer completed from Grok's saved history: it reads like a Claude or Codex chat's, with the existing
notice. The attach window also closes after a failed attach, and a created session that session/resume
reports missing is replaced like one session/load reports missing.

The replay reconciliation is removed: the lane no longer reads the journal, and D3's replayed-input
grammar test and completed-turn check in the assembler go with it.

* refactor(native-chat): a failed Grok reattach needs no window close of its own; its lane is replaced

* test(native-chat): D3's merged tests use the shipped declarations and the launch options main requires

* fix(native-chat): typecheck fallout of the base merges; any agent's empty chat is reusable

Main's idle-empty-chat lookup and launch join now take any registered agent, as the rest of the
launch path does. The refusal check moved into the prompt turns and the prompt-block conversion beside
the turns that send it, keeping both files in their line limit.

* fix(native-chat): a Grok Stop ends the process once Grok settles its turn; the next send resumes

Grok's session/cancel ends only the running turn: work it already moved to the background keeps
running and can begin a turn of its own after the person pressed Stop. Stop is now a session
boundary, as it is for Claude: the cancel answers open requests and lets Grok end the turn, the host
waits a bounded grace for that, then ends the process; the next send relaunches and resumes.
The adapter's own bounded close of a turn Grok began is gone. Its named-turn check stays: the host
ends the session unless the provider declines a Stop naming a turn that has since ended.

* test(native-chat): a Grok Stop ends the process only after Grok answered the cancel

* fix(native-chat): Steer on a Grok card cancels the running prompt, then sends it

A send that reached Grok while a prompt ran was held in the adapter until that turn ended: Steer
on a queued card took the card out of the host's editable queue and meant 'send after this turn'.
It now cancels the running prompt (session/cancel; the session stays) and sends as the next prompt
once Grok answers the cancel, as the common pattern does; a steer behind another cancels it in
turn, so the last one runs. The adapter holds a send only while that cancel lands, so its general
held-send queue and its holdsDispatch report are gone (every send it holds has its turn open in
the journal). An older client's mid-turn send takes the same path. capabilities.steering is
unchanged and still unread.

* refactor(native-chat): a close or Stop cancels a start through the acquire's own abort signal

The host owns the acquire it runs, so it now owns its cancellation: each attach's acquire gets an
AbortSignal, aborted from outside the session's queue by a close and by a Stop admitted now (the
same admission rule as before). The optional abandonStart adapter hook, the router's fan-out to
every adapter and the ACP adapter's session-keyed start map are gone; the ACP adapter keeps an
unkeyed set of starts only so quit can prove their children gone, and keeps a failed start's
unproven child until its exit is proven.
The hook also let a later close ask that child again. The host now does that from state it holds:
a close of a chat with no live child whose record still names an owner process with no death
evidence asks the adapter to release it. The answer is not recorded as proof (the lease probe
does that), so an owner pid an earlier Orca left is never killed or marked gone. Claude and Codex
ignore the signal and hold no such child; their release is a no-op (tested).

* fix(native-chat): a Grok crash that closes stdout before its exit still ends with Grok's last words

On macOS and Linux the agent's stdout ends before its exit is observed, with or without the
supervisor's EOF forwarding, so the connection's loss closed the journal first and its error text
became the session's ended reason, dropping Grok's stderr. The reason is now read at the proven
exit: the agent's last words when it left any, else why the connection closed. The failure already
carried them. Comments that assumed the exit comes first, that early frames past the cap refuse the
start, and that dispatch re-checks image support are corrected.

* fix(native-chat): nothing Grok sends while a held chat reattaches is written, marked as replay or not

The reattach window relied on the dialect's replay verdict, and Grok's frames read as live unless
they carry isReplay, so an unmarked chat frame during session/resume opened a turn that never
ended. D3 now marks every frame inside the window as replay before the translator reads it, so the
translator keeps only context usage whatever the agent marked; options and commands are still
adopted. The translator's load semantics are unchanged.

* test(native-chat): a Stop after a resume finds no turn an unmarked old reply opened

* test(native-chat): a resumed Grok chat keeps its last context reading; the resume refreshes only the window

* test(native-chat): a Grok background task a Stop ended reads as stopped reporting

* refactor(native-chat): quit's stop of each start answers through one promise kind

* fix(native-chat): quit aborts every start the host has in flight before draining attaches

A Grok that never answered its handshake held quit until the start's own 60 s bound, past the
20 s quit deadline. The host's teardown now aborts each in-flight acquire (and any the drain
still begins), so the adapter's own quit controller and its map of starts are gone: a start
has one canceller, the host's signal.

* fix(native-chat): a Grok start's abort stops reaching its child once the start has returned

The listener stayed on the host's signal until the attach finished committing, so a Close in that
window killed the now-live child behind the host's back and it read as Grok crashing. The start
now detaches it when it ends; a later Close goes through the session's own stop.

* fix(native-chat): a close or Stop during any attach phase stops the start before it launches

The attach began its abort controller only after reconciling leases, resolving recovery and
probing the previous owner, so a close or admitted Stop in those phases reached nothing and Grok
launched anyway. The controller now begins first, and the acquisition checks it before asking the
adapter to start.

* test(native-chat): a close during the attach's owner probe asks no adapter to start

Also renames the close test after the hook it no longer exercises.

* test(native-chat): a close's re-ask closes a Claude or Codex child a failed cleanup left

The re-ask is not a no-op for them: when the adapter still holds the child its cleanup could not
prove gone, the close stops it again as a requested close, and Claude persists the handle of the
conversation it ran so the next send resumes it. Corrects the tests' and comment's wording; the
close awaits the re-ask, bounded by each adapter's kill ladder.

* fix(native-chat): Steer during a turn Grok began itself cancels it and sends once it ends

A send while Grok ran a turn of its own (a background task waking it) went straight to Grok, which
queued it behind that turn where Orca could no longer withdraw it, while Stop treated the same turn
as the running reply. The send now waits as a steer, the turn is cancelled once, and the message
goes when the turn ends; a Stop withdraws it and an exit rejects it as never sent.

* test(native-chat): a steer whose cancel Grok never answers ends Grok and is rejected as never sent

Pins the bounded steer cancel kept from the runtime: past the bound the connection closes, the
running reply reads unverifiable, Grok's end reads as its exit, and the waiting steer is rejected
as never sent.

* fix(native-chat): a Grok crash stays a crash when a stop lands before its exit is proven

After the connection broke and the close could not prove Grok's exit, any later stop Orca asked
for (the next start, a Stop, a Close) marked the child as closed by Orca, so the crash read as a
requested close and Grok's last words were dropped; a send meanwhile was recorded unconfirmed.
The connection loss now decides the cause, and a send on that session is rejected as never sent.

* test(native-chat): fixtures this PR's registered Grok and desktop capability made stale

CI's unit shards failed on tests outside the PR's own lists. Each encodes something this PR changes
on purpose: Grok is now a registered agent (the seam test's unregistered agent is now Cursor); the
desktop now advertises registered agents (the restart-offer tests' older client drops that
capability explicitly); the attach context carries the start's abort controllers (the forget-status
double gains them); and the ACP real-host test rig sends to the host directly (listed beside the
other real-host rig in the send ratchet).

* fix(native-chat): a start quit stops is not the queued message's start failure

With quit now aborting a start it would have waited for, the delivery step recorded the aborted
start as the message's failure ("couldn't restart"). After quit has stopped delivery, the step
leaves the message to quit, which settles it as a close does ("The chat closed before this message
was sent."). The test that pinned quit waiting for that start and stopping its child now pins that
nothing is launched behind quit.

* fix(native-chat): a message sent after a Stop or close aborted a start gets its own start

A start the host aborts (an admitted Stop, a close, or quit) returned its refusal to the delivery
loop, which then rejected whatever was queued at that moment with "couldn't restart", including a
message the user sent after the Stop. The attach now reports that the host aborted it, and the loop
re-derives from the journal instead: what the Stop or close withdrew is already settled, a message
accepted since gets a start of its own, and quit's next step stops the loop. This replaces the
quit-only carve-out with the same rule for every abort and every agent.

* test(native-chat): the message sent after an aborted start is answered, so no settlement outlives the test

* fix(native-chat): a Grok model pick Grok never answers no longer holds Stop or Close

The pick runs on the session's queue. It now registers in the host's out-of-queue
abort registry beside a start, so a close, an admitted Stop or quit abandons it, and
the ACP adapter bounds it at 30 s like Claude and Codex. A late answer is still adopted.

* fix(agent-launch): a phone's launch opens a terminal for an agent whose chat it cannot show

agent.launch now reads the caller's capabilities by the rule tabs and restart offers
use (clientRendersStructuredAgent). A phone without registered-agents.v1 gets Grok as
a terminal again, as on main; the host's own callers and desktop clients are unchanged.

* fix(acp): strip every agent hook variable from the ACP child, from the shared list

ACP_CHILD_ENV_TO_DELETE was a second copy of the hook runtime keys that missed
ORCA_AGENT_HOOK_TRANSPORT; it now spreads AGENT_HOOK_RUNTIME_ENV_KEYS beside the pane
identity keys.

* refactor(native-chat): the mutation context carries the provider-wait registry itself

Keeps the host file within its line limit; one field instead of two closures over it.

* fix(agent-launch): agent.launch.v2 still vouches for Claude and Codex chats

The caller rule from the previous commit also turned Claude and Codex into terminals
for a client advertising only agent.launch.v2, whose contract says it opens a chat
(mobile retry-authority tests). Only an agent beyond those two now needs the client to
read it (clientRendersStructuredAgent); the test fixtures go back to what they were.

* refactor(native-chat): drop saved-history adoption from the timeline assembler

The common pattern discards the history a provider replays while loading a
session, so the assembler has no use for an input.history event.

* refactor(acp): drop session/load history adoption from the translator

The common pattern discards the history an agent replays during session/load,
keeping only what it says about the context window. Remove the adoption path
(acp-history-adoption.ts, the adopt option, and the historical background-task
liveness rewrite it fed) so load replay is always dropped except usage.

* refactor(native-chat): a pending input is only Orca's send now

Review follow-up to the adoption removal: drop the comment naming the
provider's saved message, and make requestedAt required since every pending
input comes from input.accepted.

* test(acp): keep the task-result status table on live frames

Review follow-up to the adoption removal: the result-status mapping was only
tested through adopted history, so run the same table on live frames, and
cover an unmarked task notice during a load being dropped.

* test(acp): a frame helper for a shell command Grok is running

* fix(acp): a Grok crash settles through the host's provider-exit batch, scoped to the turn it ended

A Grok crash ended the journal unverifiable before the adapter reported the exit, so the host's
provider-exit settlement found no running turn and wrote nothing: the adapter's failure (with
Grok's last words) never reached the journal, and a later stale-session pass wrote a bare,
thread-scoped cut-short row, so the partial reply was not folded as Claude's and Codex's are.

At a proven exit the ACP lane now ends its running turn interrupted at the exit instant, as the
host's exit contract expects of a child's own translator (Codex's does the same). When Grok's
stdout closed first (every POSIX crash), the turn is unverifiable only until the exit is proven:
the host's provider-exit settlement now takes the exit as proof naming the child's fence and
revises what that child left unverifiable in the same batch, with the turn-scoped row and the
adapter's failure. Claude and Codex write no unverifiable turn of a live child except a command
whose hand-off is in doubt; that turn is now revised at the exit instead of at the next open.

* test(acp): a crash seen first leaves the host no Grok turn to revise

* refactor(native-chat): what a gone generation left unfinished gets its own module

The settlement file passed 300 lines with the exit-proof revision. The unfinished-work reads
(capture, interrupted-by-the-exit, in-progress) are their own concept and move out unchanged,
apart from the exit proof they now take.

* refactor(native-chat): a watched exit revises what its child left unverifiable without reading Stop marks

An exit's own instant is the turn's end, so the revision needs only each row's fence: the
settlement's journal type gains itemFence alone, and the host test fakes say so.

* test(native-chat): drop the duplicate itemFence on the fake that already had one

* test(claude, codex): an exit whose stdout ended first still reports as it always did

The provider supervisor now ends Orca's stdout when the agent's ends, so on every crash EOF
arrives before the exit is seen. Claude's and Codex's connections report nothing at EOF and
report the exit, with its usual reason, once it is seen.

* fix(acp): reopen a chat with session/load, as the common pattern does

An agent that offers both now reloads its session instead of resuming it; the
reattach window still discards what it replays except context usage.

* fix(acp): drop the 60 s handshake bound; an abort fails the start's waits at once

Neither common design bounds an ACP handshake: Close, Stop and quit end a start
that never answers. The abort now also closes the connection, as a kill there
does, so the start settles even before the child's exit is proven. The
host-stopped start refusal only this bound produced goes with it; the idle
sweep keeps its words.

* fix(acp): a Stop naming an ended turn follows Claude's rule

It still stops nothing while another turn is live, but in the gap before a
follow-up's turn opens, which no client can name, it now stops what is in
flight and the session ends, as a Claude Stop does.

* fix(native-chat): a close no longer re-asks a failed start's unproven child

Neither common design retries that stop at Close, and Orca's Claude contract
re-asks only at the next start and at quit. The ACP adapter keeps the child
until its exit is proven and asks it again there, as Claude does.

* fix(acp): a message sent during a turn the agent began itself goes at once

Both common designs send it straight to the agent with no cancel; only Orca's
own running prompt is steered (cancelled, then re-prompted).

* test(native-chat): dismiss-all through a remote client's audience keeps a newer Orca's offer

Uses an audience production sends (one that cannot show every agent), per review.

* fix(acp): launch Grok as `grok agent stdio`, without the update and leader flags

The common pattern passes neither --no-auto-update, --no-leader nor
GROK_DISABLE_AUTOUPDATER; full access still adds --always-approve.

* fix(acp): an agent that ends its stdout, or answers unreadably, is not a lost connection

As in the common pattern, only a broken stdin (or Orca's own close) ends the
agent; one that closed its output but can still be written to stays until a
Stop, a close or its exit. The provider supervisor goes back to its base
content, so Claude and Codex no longer get the forwarded stdout end either.

* fix(native-chat): a person's close joining a failed one still binds the turn its child end cuts

On main every close of the chat writes its own Stop and settle. Here a later close joins the
first and writes no row, and the first's settle closed when its kill failed, so a turn that opened
in between and was cut by the next close read as failed. A person's close joining a person's close
whose Stop opened a settle now reopens that settle until its attempt is done.

Tests: a close whose kill failed still closes its settle; a turn opened between a failed close and
the next reads as the person's cancellation (each fails without its half of the fix).

* test(native-chat): Grok opens as a chat only behind the structured-chat setting

agent.launch and orchestration worker-start read the same setting as the
renderer route; pin both states for Grok on each. The setting's description no
longer names only Codex and Claude, in every catalog.

* docs(acp): generic ACP comments say what holds for every agent, not Grok

Stop ends the session for every ACP agent, as in the common pattern; the
adoption hook comment goes (adoption is not planned); a failed start's child is
retried at the next start or quit.

* test(claude, codex): type the EOF-before-exit test's streams; the supervisor no longer forwards EOF

The Claude test wrote to the child's stdout and stderr through their Readable
type, which the node typecheck rejects; it now holds its own PassThrough
streams. The comments no longer credit the reverted supervisor change.

* feat(acp): a steer's cancel asks once and never ends the agent

The runtime had one cancel: send session/cancel, wait at most 10 s for Orca's prompt to settle,
then close the connection, which ends the agent. A steer used it too, so a slow agent lost its
process just because the person added a message. requestSteerCancel() now sends session/cancel
once per prompt, cancels the agent's open requests and answers later permissions cancelled, and
never bounds or closes: the prompt's own reply ends it and the steer's prompt follows. cancel()
stays the Stop: bounded, then close. A Stop after a steer still bounds and closes. Both cancel
paths move into acp-prompt-cancel.ts over one cancel channel.

* chore(native-chat): keep the record store and recovery capsule under max-lines after the main merge

* fix(acp): a repeated steer shares the cancel in flight; say what the caller owns

Per review: a second steer before the first write lands returns that write instead of resolving
early. The steer's JSDoc says the wait for the prompt's reply is unbounded and that a prompt that
fails instead must not take the steer until the caller rebuilds the session; the Stop's says a
prompt that settles in time leaves the agent for the Stop's owner to end. The steer test now gives
the runtime a handler that would allow: the open permission's signal aborts and the late one never
reaches it.

* fix(native-chat): drop the stopDelivery the A3 merge doubled

* fix(acp): a steer's cancel asks Grok once and never ends it

A steer now uses D1's notify-only cancel. Two messages sent during a reply Grok began itself
cut that reply, as the common pattern does, and then both run; before, the queued first
message could not answer the bounded cancel and Orca ended Grok although Grok answered.
A Stop keeps the bounded cancel and its 4 s grace.

* fix(acp): a permission Grok asks with no prompt of Orca's running is declined

During a turn Grok began itself nobody asked it to act, so the request is answered
cancelled at once instead of opening a card that waits, as the common pattern does.

* fix(acp): a Grok that dies while starting is reported with its own last words

A dying process's stdout ends before its exit is seen, so the start failed as a closed
connection and Grok's stderr was lost. A start whose connection closed now waits, bounded by
the Stop grace (or a Close/Stop), for the exit before it is told.

* test: a Stop after a steer sends its own cancel; drop the import the A3 merge doubled

* test(native-chat): main's Stop-note test builds its turn context with the agent registry

* test(claude): say why the close test's fake child cast is safe

* test(native-chat): build the Stop-opened-turn test's identity and turn context the current way

The test (#25056) landed before the opaque provider handle (#24991), so main still built
the old {kind, threadId} handle; the turn context also needs this branch's agent registry.

* test(native-chat): build the Stop-opened-turn test's identity with the opaque handle

The test (#25056) landed before the opaque provider handle (#24991), so main still built
the old {kind, threadId} handle.

* test(ratchet): require src/main/provider-process now that it has landed

* test(native-chat): keep main's opaque-handle import in the Stop-opened-turn test

Main's #25706 and this branch both added the import at different lines; the merge kept both.

* test(native-chat): keep main's opaque-handle import in the Stop-opened-turn test

Main's #25706 made the same fix as this branch at a different line; the merge kept both imports.

* feat(native-chat): record a fresh provider conversation that replaced one the agent could not restore

A chat whose saved conversation the agent cannot reopen can now continue in a
fresh one: the handle chain records the new conversation as a creation that
replaces the lost one (which, why, and when), keeping every earlier link.

Rows keep a shape older builds read: the stored chain starts at the latest
replacement and carries the earlier links inside it.

* test(native-chat): build this stack's journal identities with main's opaque provider handle

Main's #24991 replaced the {kind, ...} handle with {transport, agent, nativeId}; three test
files from this stack still wrote the old shape. Same lines the downstream ACP branch uses.

* docs(acp): every reattach drops the agent's replay, not only for a chat the journal holds

* refactor(native-chat): store a replaced conversation flat; refuse it where older builds read the row

Older builds only read Claude and Codex records, so the nested stored form
protected rows no replacement can reach while adding a cap mismatch after a
downgrade. Store the chain as held, refuse a replacement in a Claude or Codex
chain until one has a stored shape older builds read, and refuse a supersession
key on a replacement that names no creation in the chain.

* refactor(native-chat): read hosts' structured agents from the app-shell services

Main grew the startup hydration hook to its line limit; the host agents sync is an app-lifetime subscription like the structured session tabs sync beside it, so it moves there.

* Use current provider handles in transition tests

* Use current provider handles in timeline fixtures

* test(native-chat): prove replacement rows survive downgrade and re-upgrade

* Require the ACP directory in the runtime import check

* test(ratchet): require src/main/acp now that this PR lands it

* feat(acp): a saved session the agent cannot reopen continues in a new one, with one warning row

When session/load (or session/resume) of a saved ACP session fails, the chat starts a new session and records it as a creation that replaces the lost one (#25747's 'replaces' link), and writes one warning row that the agent no longer remembers the earlier messages. A created session the agent reports missing is still superseded silently; a signed-out agent or a start that is over (Close, Stop, a lost agent) still fails the start.

* chore(acp): rewrap the acquire header comment

* test(acp): a start closed while the agent reopens fails without opening or announcing a new session

* Let ACP connections own their supervised agent process

* Preserve ACP cleanup evidence and isolate exit observers

* Expose ACP cleanup observations and type the permission fixture

* refactor(native-chat): the registered-agents capability lives in its own module

Main's growth put protocol-version.ts one counted line over its 300-line limit once the capability
was added; like main's other per-feature capabilities, it now has its own module, and importers read
it from there.

* refactor(acp): one connection owns the Grok process and its protocol

D3 now opens each ACP agent through createAcpAgentConnection (ACP-ALIGN #25810): one object spawns the
process on the execution host, owns its stdio and protocol, and reports its proven exit. It is built and
tracked before the handshake, so a start's abort (Close, Stop, quit) still reaches it, and a failed start
keeps that same connection for the next close to retry rather than spawning another process.

Deleted: the spawnAcpStructuredChild wrapper and its test, the raw-stream runtime assembly, the caller's
exit -> runtime.close wiring, the stdout-EOF heuristic (the connection no longer treats stdout EOF as
exit), and the 10 s steer/Stop cancel bound with requestSteerCancel. Reader control maps to
pauseReading/resumeReading; a close is connection.close after the host's existing 4 s Stop grace.

The adapter owns what the protocol no longer does: one session/cancel per running prompt however many
steers arrive (cleared with that send's settlement, retried after a failed write), and a Stop or steer
answers every open agent request the person has not already answered with the agent's own cancelled
reply. An answer already being saved when the Stop lands is sent.

Tests: blocked cancel write never holds Stop's grace, two quick steers send one cancel, a failed cancel
write is retried, a real process exiting while a child holds its stdout ends the session, and the
existing start-abort, retention, crash, connection-loss and reload-failure suites on the new rig.

* fix(acp): Grok signs in on its own machine with its API key or cached sign-in

When Grok reports that it needs authentication, Orca now names a sign-in method on the machine Grok runs
on, read from the same environment Grok was launched with: xai.api_key when XAI_API_KEY is set there and
Grok offers that method, else cached_token when Grok offers it, else none and the chat keeps the existing
not-signed-in refusal. The rule lives in Grok's launch spec; the adapter applies any agent's rule for new
and reopened sessions through the protocol client's caller-named method (authenticate, then retry once).
No new sign-in UI; interactive methods are never chosen.

* fix(acp): the adapter decides which of Grok's requests reach the person

The turn owner now admits every agent request, permission or question, from its own turn state: a
request reaches the person only while Orca's prompt runs and no steer or Stop is cutting it short (a
question may also come from a turn Grok began itself, until a Stop). Anything else gets the agent's
own cancelled reply and opens no card, so a question arriving after Stop or during a steer never
appears. A steer, like a Stop, withdraws the requests already open; an answer already being saved is
still sent. The protocol client's abort-on-cancel path is no longer used: after the connection
change its request signal aborts only when the connection closes.

* fix(acp): a plan Grok proposes shows as a plan, with no approval card

When Grok leaves plan mode it asks the client to approve its plan (x.ai/exit_plan_mode). Orca showed a
blocking 'Approve plan / Request changes' card for it; the common pattern has no such gate. Now the
plan goes into the chat's existing Plan row (the plan-document status row Codex and ACP plan updates
already use) and the request is answered at once with 'abandoned' plus feedback telling Grok to stop and
wait for the person's feedback or a request to implement it in a later turn, so nothing is approved on
the person's behalf. Dialects gain settleRequest for requests answered without asking anyone.

* fix(orchestration): worker-start opens a Grok worker in a terminal, as before

With the structured chat setting on, worker-start decided 'structured' for Grok and then the structured
worker factory (Claude and Codex only) refused it, so the start failed; main opened a terminal Grok
worker. Worker-start now decides with no registered agents beyond Claude and Codex, so Grok gets a
terminal worker as before. agent.launch and the app's own launches still open Grok as a structured
chat. Temporary until structured workers take registered agents.

* fix(acp): a prompt answer Orca can't read ends the turn instead of hanging it

A session/prompt rejection that was not the agent's own error answer (an answer that fails Orca's
schema, or one too large to read) left the turn running: the next message became a steer with nothing
to cancel and was never sent or settled, and Stop waited its full grace. As in the common pattern, any
prompt failure now ends the turn as failed (a failed-turn row without words, since none are the
agent's) and settles the send, so the next message goes. Only a closed connection keeps the send
running, for the connection-loss path to settle.

* fix(acp): send Grok's prompt-identity extension only to agents that echo it

session/prompt carried _meta {promptId, requestId} for every ACP agent, though only Grok's dialect
echoes it (injectedPromptIdentity). Now only an agent whose dialect declares it gets the extension;
other ACP agents get a plain prompt.

* refactor(native-chat): the registered-agents capability lives in protocol-version again, as on main

This reverts 0ef6d21815. That commit moved the capability to its own module only because main's
protocol-version.ts was then one counted line over its limit; main now defines it there itself within
the limit, and main's new restart test imports it from there. Main's test also reads the desktop's
capability list as an older client; on this branch the desktop advertises registered agents, so its
older client is that list without this one capability.

* test(acp): read the sign-in method with a schema, not a type assertion

* refactor(native-chat): composer transport and Stop control in their own modules

Main's rewind change (#19338) brought NativeChatStructuredSession.tsx and use-structured-agent-session.ts
to their line limits, leaving no room for this branch's image-acceptance and unpublished-Stop lines.
The composer's transport (sends, commands, options, image acceptance) moves to
use-native-chat-structured-composer-transport.ts, and whether Stop shows and what it does moves to
structured-agent-session-stop-control.ts. Behavior is unchanged; the runtime cast on the composer's
'local' | 'remote' is now a typed return.

* test(native-chat): read registered agents by id, as main's structuredAgentsReadBy now takes

Main's A3 squash changed structuredAgentsReadBy to take agent ids; this branch's test still passed
{ agent } objects (CI typecheck TS2322).

* fix(acp): a first reopen warns when Grok forgets a chat that exchanged turns

A Grok session the chat created was treated as one Grok never saved, so
when Grok reported it missing on the chat's first reopen, Orca swapped in
a fresh session silently, even after completed exchanges: the person saw
the old messages while Grok had forgotten them.

The launch now counts a created session as never saved only when the
chat's journal, read at the failed reopen, holds no turn of that Grok
session; anything else, an unreadable journal included, takes the normal
path: the fresh session is recorded as replacing the old one and the one
warning row is written.

* fix(acp): a start writes the warning row an earlier attach failure dropped

The row saying Grok forgot the chat was written only into the attach's
deferred sink, while the fresh session's link was saved earlier. An attach
failure, quit or crash in between dropped the row forever.

Every start now derives the owed rows: each conversation the chain says was
lost to a failed restore gets its row unless the chat's journal already holds
it. The row now names the lost conversation rather than the fresh session, so
a row an unused replacement wrote still counts after Grok supersedes it.

* fix(acp): a question during a turn the agent began itself gets its cancelled reply

A question or other card-opening request the agent sends while no prompt of
Orca's runs (a turn it began itself, as when a background task wakes it) now
gets the agent's own cancelled reply and opens no card, the same rule
permissions already follow there. Nobody is waiting on that turn. A plan the
agent shares in it is still shown.

* test(native-chat): import the unfinished-work capture from the module that owns it

Main's reasoning sweep test (#19221) imported it from the dead-generation settlement, which
this branch split it out of.

* refactor(native-chat): keep the session host within its line budget after main's Stop work

Main's #25949 left the host at exactly its 300-line budget, and this branch's acquire-abort wiring
adds one line. The reveal module now comes in as a namespace import, as the host already does for
its other helper modules, and the earlier reorder of two type imports is undone.

* test(codex): move the stdout-before-exit test into its own file

Main's connection test file is at its 800-line test budget, and this branch's exit-order test
pushed it over (CI lint, max-lines).

* fix(acp): cancel a running turn before a close, dispose or quit ends the agent

Closing a tab, disposing a session or quitting while an ACP agent's turn ran
killed the process without asking the agent to cancel first. A requested
close now does what a Stop does: withdraw the agent's open requests and held
steers, send session/cancel once, and wait for the turn to end, bounded by
the Stop's grace (4 s), before closing the process. A close that follows a
Stop sends no second cancel and waits only what is left of that Stop's grace.
An idle close, a lost connection and a sink-failure force close are
unchanged. The cancel-and-wait moves to acp-structured-stop.ts, shared by
Stop and close.

* test(codex): pass resolveLaunchArgs in the stopped-send-order test so typecheck passes

Main's typecheck fails here too: #25721 made resolveLaunchArgs required and
this test (#25051) predates it. Same line, same place as the open main fix
(#25977), so merging main after it lands is a no-op.

* test(native-chat): run the close-aborts-start test on the Node runtime, as its SQLite journal fixture requires

* refactor(native-chat): the Stop control reads the outbox itself, so the chat hook stays within its line limit

* Keep the session host under the line cap after main's two new delegates

Pass the lifetime's conversation opener to reveal directly; it is already passed unbound to the mutation context.
2026-10-06 22:09:26 -07:00
8464b151c7 Claude account profiles: dormant routing and consumers (Step 2 of 4) (#24351)
* feat(claude): add dormant profile setup and history sharing

* fix(claude): make profile setup one gated, typed, fail-safe entry

Review round 1 of the dormant profile setup found that the pieces could
be called without their safety checks, that one failed write or an
unreadable bookkeeping file could silently stop sharing for good, and
that Windows prompt history could bring back history the user cleared.

- One entry, provisionClaudeAccountProfile: the profile gate (namespace,
  no linked components, outside ~/.claude and ~/.config/claude, and an
  ownership marker beside the home naming the account and target) runs
  first and refuses before creating anything; then history sharing,
  config provisioning, and the hook install after the settings merge.
  Results come back per surface with closed warning codes instead of
  message text.
- The sharing ledger is keyed by surface name, records a value only
  after its write succeeded, and an unreadable ledger starts empty and
  is rewritten instead of blocking every surface.
- The profile state file goes through the same locked writer as folder
  trust (Claude's <file>.lock plus the in-process queue), generalized as
  updateClaudeGlobalConfig. Onboarding and trust are still applied when
  the personal state file is unreadable.
- WSL descriptors build guest POSIX paths; the state-file path style
  follows the injected platform.
- Orca's managed statusLine has one owner in a profile: the settings
  merge never shares it, a user's own statusLine is shared over it, and
  the profile installer follows the default home's slot so a default
  opt-out reaches every profile. remove() takes the same destination;
  the remote installer cannot accept one.
- Prompt history compares file identity (bigint dev+ino) on every
  platform, never drains the shared file into itself, drains retained
  copies in generation order, never reuses a stale cursor, and on Windows
  keeps a replaced default's old copy aside instead of replaying it.
  Directory merges keep going past a failed entry.

* fix(claude): share the user's own hooks and keep merged history whole

A user's own Claude hooks in ~/.claude (notifications, formatters) did
not run under a managed account, because the whole hooks key stayed
private. They are now shared like any other settings key: Orca's own
hook entries and its managed statusLine are stripped from both the
personal value and the profile's current value before the per-key
ledger comparison, so they never travel through the merge and never make
the key look user-owned. Orca entries already in the profile are kept on
write, and the profile hook installer adds them on top as before.

Prompt history: merged bytes that lack a final newline are terminated,
so Claude's next record no longer fuses onto the last merged line. When
a CLI rewrote the profile's history file (old records plus new), only
the lines past the part it shares with the default history are added,
instead of the whole file again.

* fix(claude): close review round 2 gaps in profile setup

Hooks and statusLine sharing:
- When ~/.claude holds only Orca's hook entries, the user's shared hooks
  now read as an empty value instead of a missing key. Removing the
  user's last own hook in ~/.claude therefore reaches profiles that
  never edited it, and deleting the only shared hook inside a profile
  stays deleted.
- A custom statusLine Orca shared, and the profile never edited, goes
  away when the default home drops it. When a shared custom line
  replaced Orca's line in a profile, the profile's statusline marker is
  dropped so Orca's line comes back once the default returns to it; a
  profile that opted out stays opted out. No other key gains deletion.
- install/remove/getStatus with a profile directory refuse when it is
  the default home, or its settings.json resolves to the default one,
  instead of editing System Default's hooks and opt-out state.
- The profile statusline rule reads the default settings under the
  userHome passed to the setup entry, not os.homedir().

Profile state and ownership:
- A malformed `projects` value skips only folder trust (new warning
  code trust-refused); onboarding and shared keys still apply.
- The ownership marker stores only host-local facts (account, runtime,
  distro). The execution host id is the caller's view of the host, so
  it stays in the in-memory descriptor and is not compared.

Prompt history interruption paths:
- With no cursor yet, a retained copy starts past the bytes it shares
  with the default history, so an interrupted share no longer replays
  the whole history.
- A retained name for the shared file itself is removed with its cursor
  instead of lingering until a later scrub makes it look new.
- The Windows link record is read three-state: unreadable stops the
  share instead of reading as "no link". If the record cannot be
  written after linking, the fresh link is undone.
- An unreadable retained copy is reported and no longer blocks linking.

* build(cli): list the new Claude hook modules in the CLI project

hook-service.ts and hook-settings.ts are compiled into the packaged CLI
project, which lists every file explicitly. The statusline policy and
profile destination modules they now import were missing, so the CLI
typecheck failed with TS6307. The CLI still loads hook-service through
the existing managed-agent-hook-controls build entry, which bundles
both modules; neither imports electron.

* fix(claude): close review round 3 regressions in profile setup

- A profile whose hooks hold only Orca's entries and that sharing never
  recorded is no longer treated as a user edit, so the user's first own
  hook in ~/.claude reaches it (for example when the profile was set up
  before ~/.claude had any hooks).
- A retained prompt-history file is removed as a second name for the
  shared file only when the default history does not itself link to it;
  otherwise it holds the only copy and is kept.
- Default-home checks compare file identity: the profile hook
  destination check uses device and inode, and the profile/default
  separation check resolves on-disk case, so a case-only alias of
  ~/.claude is refused on case-insensitive filesystems.
- A test pins that an unreadable leftover session tree no longer blocks
  linking.

* fix(claude): let shared keys leave a profile when ~/.claude drops them

QA found that removing a setting from ~/.claude never reached a managed
account: deleting the whole `hooks` block left the user's hook running
there. Only statusLine followed the default away.

Every shared key now follows the same rule through the existing per-key
ledger: when a key disappears from ~/.claude/settings.json (or
mcpServers/theme from the personal state file), it is removed from the
profile if the profile still holds exactly what Orca last shared. A
value changed inside the account is kept. Keys Orca never shared,
including denylisted ones, are never touched. Deleting the whole hooks
block removes the user's shared hooks and keeps Orca's own entries. A
missing source counts as empty; an unreadable source removes nothing.

* feat(claude): add dormant profile routing and account consumers

* fix(claude): drop the dormant profile selection RPC; clients negotiate by capability

Restores the inline mobile allowlist so its source-scan guard sees every
accounts.* method again, and the generated params catalog to generator order.

* fix(claude): guard the claude shell function and honour a hand-exported config dir

The function is defined only in a routed pane where claude is a real
executable (the codex function's guard), re-reads the pointer only while
CLAUDE_CONFIG_DIR is unset or still Orca's injected twin, accepts Git Bash
drive paths, and starts on its own line after the fish/PowerShell codex text.

* fix(claude): spawn-time profile env, total account listing, setup at lifecycle triggers

Round-1 review fixes for the dormant profile routing:
- Panes get the selected profile's CLAUDE_CONFIG_DIR plus an Orca twin at
  spawn, so nested shells and scripts inherit the account; System Default
  injects nothing and its home is the inherited CLAUDE_CONFIG_DIR.
- An absent routing owner is System Default, never a throw; AI Vault and
  session-search scans receive profile roots from their parent, and the
  capability is advertised only where an owner is installed.
- Account listing never throws: per-account readiness, a stale pointer is
  republished in the background and reported on the snapshot.
- Profiles are set up at select and startup; a launch only sets up one that
  never was, and a worker fault on a prepared profile is a warning. The
  Claude version probe is cached per binary identity.
- Pre-trust goes through the existing deadline- and realpath-guarded writer
  against the launch env's profile config.
- Skill discovery keeps a caller's Claude root and a broken Claude selection
  no longer fails other providers.
- The durable record carries a provider-neutral launchAccountHome, read
  through one helper by the launch fallback and the model catalog.

* test(claude): pin the version-probe cache, launch-account record and temp-home readers

* test(claude): pin dormant bash rc text alongside fish and PowerShell

* test(claude): read the fish launch init without a nullable index

* fix(claude): withdraw the profile pointer when a selection cannot be published

A pointer left naming the previous account would launch it silently; a
missing pointer makes the claude function refuse visibly. A newer selection
that raced the failed one keeps its pointer.

* fix(claude): read the fish profile pointer with read -z for fish older than 3.4

Shell tests skip system config and abort unless claude resolves to the fake.

* fix(claude): only the newest publish withdraws the pointer; total config dir lookup

- An overtaken publish that fails leaves the newer selection's pointer.
- The runtime config dir falls back to the legacy home for an unresolvable
  account or a WSL target, so skill roots never fail for other providers.
- WSL guest reader roots merge verbatim, never realpathed on this thread.
- History readers include ~/.claude, where step-1 setup pools profile history.
- System Default ignores a config dir an outer Orca injected (twin-marked).

* fix(claude): System Default launches and probes use the structured create resolver

A Claude agent-env CLAUDE_CONFIG_DIR the create path stored is now the home
the launch pins and the model probe accepts.

* test(claude): type the System Default launch record as an agent-session record

* fix(claude): install profile hook scripts under the setup job's home

A worker thread's os.homedir() ignores its own env, so the hook and
statusline scripts now go under the home the job names. The worker test pins
the process HOME to a sentinel, refuses to run unless the worker sees it, and
asserts nothing lands there.

* test(claude): skip shell cases whose shell the runner lacks

* fix(claude): remove env vars in the PowerShell claude function instead of setting null

On .NET 9+ (pwsh 7.5+) SetEnvironmentVariable with $null creates an empty
variable, so stripped auth vars reached claude as empty strings and the
restore left CLAUDE_CONFIG_DIR empty in the user's session.

* fix(claude): give plain fish tabs the claude function through the codex hand-off

Main now gives a plain fish tab Orca's codex function through a vendor_conf.d
snippet instead of a -C init. The claude function only rode the -C path, so a
plain fish tab would not re-read the account selection per invocation once
profiles are on. Define it at the first prompt beside codex; it stays empty
while the profile gate is off.

* fix(claude): share personal rules, themes, workflows and keybindings into account profiles

A managed account launches Claude with its own config folder, so user-level
rules/, custom themes/ (which a shared `custom:<slug>` theme points at),
personal workflows/ and keybindings.json silently stopped applying. Link the
three directories like skills and commands, and copy keybindings.json with the
same edit-preserving ledger as CLAUDE.md. routines/ stays unshared: routines
belong to the claude.ai account and the folder holds per-run state.

* test(claude): wait for the running child to read its account before switching

The test switched the selection after a fixed 20 ms, so under load the backgrounded claude
had not yet read the pointer and picked up the new account. The stand-in now marks when it has
started, and the test waits for that mark (bounded) before switching.

* fix(claude): import the personal CLAUDE.md into account profiles instead of copying it

Claude also loads ~/.claude/CLAUDE.md as a parent folder's memory for any project under home,
so a copied account CLAUDE.md made every such session read the user's instructions twice
(checked live with Claude 2.1.288). An @~/.claude/CLAUDE.md import resolves to the same real
file, which Claude loads once from home, from projects under home and from folders outside it.

* refactor(claude): simplify account profile setup toward the prior art

- Windows keeps each account's history private; drop the hardlink, link
  record and conflict-copy machinery that only Windows reached.
- Share hooks and statusLine as ordinary settings keys: Orca writes the
  same entries into every folder, so the installer finds them present.
  Drops the Orca-entry carve-out, the per-profile statusline follow
  logic and its marker.
- Unreadable ledger is just an empty ledger.
- Share from the user's own CLAUDE_CONFIG_DIR when they set one (marked
  so Orca's injected value is never mistaken for it), and refuse a
  profile at or around it.
- Pin the one canonical profile path spelling in a test.

* refactor(claude): route launches through one account router, superset-shaped

Replace the routing service, owner interface, setup worker thread, reader-root
merging, persisted launch account and capability string with one
ClaudeProfileRouter: the pointer is written first and setup runs best-effort
after it (superset's order); a missing pointer means System default.

The claude shell function re-reads the pointer on every launch, refuses only
a selected account whose folder is missing, and prints a note when the user's
own CLAUDE_CONFIG_DIR overrides the selected account in that terminal.

Still dormant: claudeProfileRoutingEnabled() is false.

* test(claude): type router test settings instead of casting

* fix(claude): run account setup on a worker thread, never Electron main

publish() writes the pointer and starts setup in the background, so neither
startup nor an account switch blocks on a history merge. Each setup runs in a
one-shot worker (the profile-state backup worker's pattern); one setup per
account at a time, reused by later requests. A launch waits only for a folder
that was never set up, and refuses with a clear message if that setup fails.

* fix(claude): do not await the synchronous pointer publish

* test(claude): give the routing launch test the merged resolver deps and handle shape

* fix(claude-accounts): dedupe merged prompt history, drop drained copies, link setup folders by path

- Prompt-history drain appends only lines the shared file lacks, so a purge never re-adds lines.
- A set-aside history copy whose saved offset reaches its end is deleted on the next run.
- Setup folders link to the default home's own entry, not its resolved target.
- The profile gate and folder creation run once, in provisionClaudeAccountProfile.
- installHooks receives only configDir; drop a duplicate test key that fails CI.

* fix(claude-accounts): refuse a routed resume whose transcript is in another account; zsh claude function; setup timeout

- With account routing, a chat resume checks its transcript is in the launch folder; a missing one
  with a stored leaf refuses with historyInOtherAccount instead of starting fresh.
- The launch folder of a selected account comes from prepareLaunch(); the resolver stays for System default.
- zsh panes get the claude function like bash, fish and PowerShell (empty while routing is off).
- The setup worker is terminated after 60 s so a later launch can retry.
- Document that the setup marker means setup started, not finished.

* fix(claude-accounts): refuse a routed resume only when the transcript is found in another folder

A transcript found in no known folder keeps the old stored-leaf resume.

* fix(claude-accounts): record installed hooks as Orca-shared; skip symlink tests on Windows

After Orca installs its hooks into an account, record the account's hooks in
the settings ledger so a later run can still bring the user's own hooks in.
Tests that create real symlinks now skip on Windows.

* fix(claude-accounts): trim the which-account file in the PowerShell claude function

Co-Authored-By: Claude <noreply@anthropic.com>

* test(claude-accounts): spell the user's own config folder as an absolute path on every platform

Co-Authored-By: Claude <noreply@anthropic.com>

* test(claude): skip the POSIX-only WSL profile test on Windows

A WSL profile's data root is a POSIX path, so building one from a Windows
temp dir fails the absolute-path check there.

* fix(claude): check the transcript before the account-switch recheck when no router is installed

Keeps the switched-off path identical to main: the recheck stays the last await.

---------

Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
2026-10-07 01:00:12 -04:00
Jinwoo Hong 7f83e0a53f fix(serve): keep every persisted terminal on the phone after serve restarts (#26022)
On a windowless host the first hydrate pass after a cold start only keeps
serve-/SSH-owned terminals, stores that partial list, and the full pass then
skips the worktree because it already has tabs. Terminals made by
terminal.create, agent launches, or splits never reappear on the phone.

With no renderer window and no snapshot yet, the runtime-owned pass now builds
the full snapshot (including a persisted split layout), since nothing else
publishes the remaining tabs.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-10-07 00:49:35 -04:00
Neil 76d480f808 Fix unsafe test fixtures and the Bun version pin (#26051)
* Keep test interruption signals within owned processes

* Pin Bun and add optional unit runner shutdown diagnostics

* Unblock CI lint without changing session host runtime

* Avoid duplicating runtime import-check dependency bundles

* Leave unit runner diagnostics disabled by default

* test: keep runner incident follow-up focused on durable guards

* test: apply transcript replacements as authoritative snapshots
2026-10-06 21:42:02 -07:00
Jinwoo Hong d1fe11c6ba refactor(terminal): carry pane placement on terminal start, report-only (#25677)
* refactor(terminal): carry pane placement on pty spawn, report-only

Adds a placement type (new tab, split, or root) that says which tab and leaf a
new PTY joins. pty:spawn accepts it as an optional field, main's runtime creates
and PTY-backed splits supply it, and the binding write records on its trace span
whether placement names the tab today's write picks. The write never branches on
it, so bindings and saved state are unchanged. applyPtyBinding becomes
copy-on-write and returns the new session.

* test(terminal): typed handler registration in the placement threading test

* refactor(terminal): keep the binding write in place; placement stays report-only

Reverts applyPtyBinding to main's in-place write: placement is read before the
write and does not need a new session, so this PR keeps identical behavior.

Follow-up: build rows and splits from placement, delete persistHeadlessTerminalSplit,
the provisional-graft strip, and mobile's follow-up props write.

* fix(terminal): a throwing placement check never fails the binding write

The agreement check is report-only, so a throw records check_threw on the span
and the binding proceeds exactly as without placement.

* refactor(terminal): carry only the placement fields this change reads

New-tab rows, sizes, split ratios and proposed trees come back with the change that builds tabs and splits from placement. Leaf ids are capped at 128 like the other session schemas.

* fix(lint): bring structured-agent-session-host back under max-lines

(cherry picked from commit 1a5f687244)
2026-10-07 00:18:34 -04:00
Jinwoo Hong 9d8e3638de chore(relay): move Asia cell c34 to the general same-cap lists after promotion (#25758)
* feat(relay): give Asia cell c34 a promotion wave so it can become a general cell

c34 launched on 2026-10-05 as a migration-only spare with no promotion path. This adds
it to the Asia admission promotion waves, the workflow's promote and canary cases, and
the canary evidence map, so the reviewed Asia workflow can promote it with the same
five-minute canary c30 and c31 ran.

The same-cap migration-only list is deliberately unchanged: a same-cap job reads a
cell's class from that list, and c34 must be rolled to the director's image as a
migration-only cell before promotion can run. The list moves after promotion, in its
own change.

Claude-Session: 1145a80d-dec4-4a9b-9373-bbbb876b9041

* chore(relay): move Asia cell c34 to the general same-cap lists after promotion

A same-cap job reads a cell's class from SAME_CAP_MIGRATION_ONLY_CELLS, not from the
selector. Once c34's Asia canary promotes it, leaving it there would make a same-cap
rollback demote it to migration-only. This moves c34 to the general same-cap list and
the shadow gate's fleet pool list (its pool is the Asia 16).

Merge only after c34's promotion succeeds. This file set is trusted evidence code, so
merging invalidates any sealed monitor or same-cap canary: merge outside a
monitor-to-enable window and before the next same-cap canary seals.

Claude-Session: 1145a80d-dec4-4a9b-9373-bbbb876b9041

* docs(relay): scope the c34 same-cap pause to the window after promotion

Claude-Session: 1145a80d-dec4-4a9b-9373-bbbb876b9041

* docs(relay): rewrap the c34 paragraph

Claude-Session: 1145a80d-dec4-4a9b-9373-bbbb876b9041
2026-10-07 00:10:52 -04:00
Jinwoo Hong bfcbcfd59c refactor(terminal): move a dragged-out pane in main before the window opens its tab (#25380)
* refactor(terminal): move the topology revision and leaf-lookup helpers into terminal-topology

advanceTerminalTopologyRevision, findTerminalTabIdForLeaf and
hasHostAuthoritativeTerminalMembership move verbatim from the renderer-save
membership rebase into persistence/terminal-topology, the home of the commit
boundary. Importers are repointed; no behavior change.

* test(terminal): test-only guard for topology writes outside the commit boundary

The three session sinks now publish through one commitWorkspaceSessionPartition
helper, which hands the prior and published partition to the topology write
guard. Production never arms the guard, so each sink pays one global lookup.

The unit suite arms it in report mode: a sink write that changes class-(a)
topology (membership, root, bindings, titles, incarnations, sleeping records,
remote session ids, tombstones, default-applied, revision) outside a commit
scope is attributed to its writer's file by stack, and fails the test unless
the writer is on the unrouted-writers allowlist that later routing PRs shrink.
Renderer saves and test seeding are exempt. Deep-freeze is available but stays
off suite-wide until the in-place writers return new sessions.

* refactor(terminal): add the topology commit module with bindLeaf, closeLeaf and closeTab

terminal-topology-commit.ts is the boundary for class-(a) terminal topology.
bindLeaf forwards to persistPtyBinding, whose write now runs in a commit
scope; the spawn commits and the relay reattach bind through it. closeLeaf and
closeTab wrap the existing close mutation in a commit scope and one
persistence.terminal-topology span (kind, outcome, refusal reason; no ids),
and the runtime close goes through them. Session output is unchanged: tests
compare it byte for byte with the old writers for local, ssh: and folder
workspaces.

* chore(terminal): drop an unused lint suppression from the topology write guard

* fix(terminal): attribute topology guard writers relative to the repo root

The guard read a frame's file through its last /src/ segment, so a test under
tests/ (folder-upgrade-identity-persistence.unit.test.ts) had no source frame
and failed as an unknown writer. Frames are now taken relative to the repo
root the setup passes in, tests/ counts as test seeding, Windows backslashes
are normalized before the node_modules skip, and nested src/ paths keep their
full path. The two runtime funnel files skip only their funnel function, so
another writer in them still shows. The R9 allowlist key names the file whose
frame actually writes. The class-(a) diff and the attribution move into their
own files; the stack limit is restored in a finally.

* refactor(terminal): drop bindLeaf until binding reaches a sink; add the boundary ratchet

bindLeaf and the commit scope inside persistPtyBinding changed nothing: the
binding write never reaches a session sink, and a scope inside the Store method
would have admitted every direct caller once it did. Both return in B1-4; the
spawn commits and the relay reattach call persistPtyBinding directly again.

The runtime close now calls one closeLeafOrTab entry, so its callbacks keep
their contextual types. A census test is the primary enforcement: only
persistence/terminal-topology and the callers it lists may call
persistPtyBinding, the session setters or the three sinks, and every
unrouted-writer allowlist entry must name an existing file.

* fix(test): resolve the topology guard's repo root without the global URL

Under happy-dom the global URL is not Node's, so fileURLToPath(new URL(...))
threw in the setup file and failed every happy-dom test file.

* feat(terminal): moveLeaf commits a pane detach in main before its new tab mounts

Ported from fix/terminal-topology-stage-a (1bafc87, 3941bd4, 2ea2070, 3e79f31,
a9d428e minus the bridge comment) and exposed as the commit module's moveLeaf,
which runs inside the topology commit scope and span and publishes every
rewritten partition through the session sink, so the test-only write guard sees
it. Store.moveTerminalLeafToNewTab only hands it the store's state.

Detach-to-new-tab now asks main first: the leaf and its binding move into the
new tab in every owner partition (local and ssh:), with the incarnation,
sleeping session, remote session id, UI marks and SSH lease, and the topology
revision bumps. pty:moveLeafToNewTab then aliases the agent-status pane key and
re-keys orchestration worker resources; the renderer's repeat transfer is
skipped. A refused or failed move toasts and the pane stays; a move the
renderer cannot apply is undone (STA-9259).

Stage A review-2 fixes:
- SF1: an undo whose source tab has closed retires the moved tab instead of
  refusing, so no ghost tab comes back on restart.
- SF2: an undo restores the split direction, ratio and position, the source
  tab's PTY id and its SSH remote session id, from origins main kept for the
  move; it never adds a second copy of a leaf the source holds again.
- N1: layouts of tabs that no longer exist do not refuse a move.
- N2: an undo into a tab that gained a pane is refused as target_changed,
  not a silent not_held.
- N3: a repeat drop of a pane whose move is still committing is a no-op.

Note: the move bumps the topology revision, so a revision-0 repo enters
rebased membership on its first detach.

* fix(terminal): tighten moveLeaf undo and move across disagreeing SSH copies

- An undo ignores a target layout whose tab row is gone, as the move does.
- A retired undo leaves UI marks and SSH leases on the moved pane key instead
  of re-keying them to the closed source tab.
- When an SSH respawn bound one partition before the other caught up, the move
  follows the renderer's live PTY id and drops the stale copy's incarnation
  instead of refusing with a toast.
- A test pins the late-spawn graft the renderer's in-flight guard exists for.

* fix(terminal): a pane move the renderer cannot finish never leaves main holding it

- A pane whose PTY spawn is still in flight is not dragged out (no toast):
  the late spawn result would bind it to the source tab and graft the leaf
  back beside its moved copy. The IPC transport reports a pending connect.
- A thrown or timed-out commit (10 s bound) sends an undo, since the write may
  have landed; undo answers not_held when nothing moved. The in-flight entry
  clears with it, so a stalled main cannot wedge the pane's drag.
- A throw while opening the new tab after the pane left its source undoes the
  move too.
- No "stays where it was" toast when the undo found the source tab closed or
  the source pane is gone; a failed undo says the pane may open in a new tab
  after a restart instead.

* fix(terminal): drop a half-opened move tab so the window agrees with main's undo

If opening the moved pane's tab throws after createTab ran, the renderer now
closes that tab without killing its PTY or telling main (main never had it),
and restores the source layout, while main undoes the move.

* refactor(terminal): drop the runtime topology write guard; the boundary ratchet enforces

The AST boundary ratchet is the enforcement for B1. The stack-attributed
runtime guard, its class-(a) diff, the unrouted-writer allowlist, the vitest
setup and the freeze option are removed; it saw two writers in the whole unit
suite and its real value starts only once binding reaches a sink (B1-4).

Also: one closeLeafOrTab wraps the close in the span (no per-kind copies or
narrowed types), the span has one finish like persistence.pty-binding, the
sink helper is publishWorkspaceSessionPartition (it publishes; the commit
boundary is the module), the ratchet drops the private publishSession row and
checks that every listed caller file exists, and the close comparison keeps
its two meaningful cases with span cases chosen by name.

* refactor(terminal): roll a pane move forward instead of undoing it

Once main commits a pane move, the renderer only rolls forward:
- A repeat of a committed move answers moved and writes nothing, so a thrown
  commit (outcome unknown) is retried, up to three attempts; the IPC skips
  the agent-status transfer it already made.
- The renderer applies the move by leaf id. If the user closed the pane or
  its tab meanwhile, it closes main's new tab through the ordinary
  session:close-terminal-surface path; if a sibling closed and the pane is
  the source's last, the source tab closes without killing the PTY.
- One toast, for a move main refused or never answered.

Deleted: the undo planner, origin ledger, undo flag, retired and
target_changed results, the commit timeout, the half-opened-tab cleanup and
the second toast.

Also: one request validation (distinct tabs, stable leaf id, tab ids a pane
key can carry, built through makePaneKey); a two-line PTY check where the
live PTY id wins; the move and close share one span wrapper; the moved tab's
layout always takes the live PTY id (a stale saved id was kept before);
fallbackPtyId is now livePtyId; tests are organized by behavior.

* refactor(terminal): trim the B1-1 commit module and ratchet to what they enforce

- Point the acknowledged-tab-retirement audit fixture at the moved
  advanceTerminalTopologyRevision; its old import no longer resolved.
- Drop publishWorkspaceSessionPartition: it was the removed guard's
  interception point, so the three session sinks return to origin/main.
- One traced(kind, mutate) wrapper in the commit file replaces the span
  factory; closeLeafOrTab is one call.
- The ratchet walks src/main with the shared scanSourceTree, drops the
  loading-store-internal rows and the redundant file-exists test; exact-set
  equality already fails on a missing file.
- The commit test is a pure unit test of the span outcomes: no Store
  harness, electron mock or self-comparing close.

* fix(terminal): keep a moved last pane alive and trim moveLeaf to one attempt

A sibling closed while main committed the move persists the source layout
down to the moved leaf, so the renderer took that for "pane gone" and
closed main's new tab while the pane was live: after a restart it was held
nowhere. The renderer now reads the store layout: when the moved leaf is the
only one left it opens the new tab with that layout, syncs PTY ownership,
and closes the source tab without killing the PTY. Main's new tab is closed
only when main moved the pane and it is gone here.

Also:
- One commit attempt; the repeat answer, the retry loop and the IPC alias
  guard never ran in production (a write with an unknown outcome faults the
  writer, so every retry rethrows before it reaches the move).
- traced takes a typed refusalOf; one write-and-restore helper and a shared
  partition assign replace the duplicated local/remote and rollback checks.
- The planner's holders carry their source tab and layout, so the partition
  move cannot fail; pane-keyed UI marks are re-keyed in one loop; the request
  validator calls makePaneKey directly.
- applyMove looks the pane up once and takes one PTY id (main's when it moved
  the pane); the move-commit file is inlined; duplicated renderer tests are
  removed and the commit-path tests live together.

* refactor(terminal): census the runtime session controller's write and drop stage ids from comments

The controller's setter was named set, which the boundary census could not
list without matching every Map.set, so a new OrcaRuntime mixin could write
sessions through it unseen. Rename it setForWorktree and census it.
Comments now describe state instead of citing plan stage ids.

* refactor(terminal): census writer references and trace refusals by callback

- traced() takes refusalOf instead of assuming an Error refusal, and only
  mutate() sits in the try, so a span outcome of threw means the write threw.
- The boundary ratchet counts references, not just direct calls: non-null
  calls, bracket keys, aliases, destructures, .call/.bind and parenthesized
  callees all count; declared names and type positions do not.
- Census terminalSurfaceCloseMutation (boundary-only) and the partition sinks
  setLocalWorkspaceSession / setHostWorkspaceSession.

* test(terminal): count writer uses in extends clauses and instantiations, skip type-only imports and local declarations

The census skipped ExpressionWithTypeArguments as a type, which also holds
`extends f(x)` and `x<T>` value expressions. Type-only import/export
specifiers and declared names (variables, parameters, accessors, enum
members) no longer count as uses. The audit fixture is listed in the table
instead of a separate exemption.

* test(terminal): count quoted and assignment-pattern destructures of layout writers

* test(terminal): count every mention of a layout writer except its definition

Telling definitions from uses per syntax kind kept missing nested and
for-of destructures. Exempt only the writer's own function or class-member
definition; any other mention (including object-literal keys) counts, so the
census errs toward a loud false alarm rather than a silent miss. Quoted names
count only in member-name position.

* test(terminal): count every string literal naming a layout writer

Member-name positions missed wrapped keys like store[('name')] and
store['name' as const]. Counting every string literal outside types is
shorter and errs toward a loud false alarm.

* test(terminal): exempt only class members and functions as writer definitions

Object-literal methods and accessors were exempt while equivalent arrow
properties counted; all object-literal keys now count alike.

* test(terminal): parse files with unicode escapes in the writer census

A name spelled with a \u escape never appears verbatim, so the text
prefilter skipped it.

* test(terminal): parse any file with an escape in the writer census

\x, identity and line-continuation escapes also decode to a writer name
without it appearing verbatim.

* test(agent-hooks): stub isPaneAuthorityTransferredTo on the hook server fake

* refactor(terminal): one move path for every client and an idempotent status transfer

The pane manager decides whether the dragged pane is the last one, and a
refused detach closes main's new tab like a vanished pane. Paired web
clients answer the move as not held, so every client takes the same
moved / not_held / refused branch. Repeating an agent-status transfer that
is already in place is now a no-op where it happens, replacing the IPC
special case.

* fix(lint): bring structured-agent-session-host back under max-lines

(cherry picked from commit 1a5f687244)
2026-10-06 23:51:42 -04:00
Brennan BensonandJinwoo-H dff65d55a3 feat(claude): prepare account profiles and shared history (Step 1 of 4) (#24300)
* feat(claude): add dormant profile setup and history sharing

* fix(claude): make profile setup one gated, typed, fail-safe entry

Review round 1 of the dormant profile setup found that the pieces could
be called without their safety checks, that one failed write or an
unreadable bookkeeping file could silently stop sharing for good, and
that Windows prompt history could bring back history the user cleared.

- One entry, provisionClaudeAccountProfile: the profile gate (namespace,
  no linked components, outside ~/.claude and ~/.config/claude, and an
  ownership marker beside the home naming the account and target) runs
  first and refuses before creating anything; then history sharing,
  config provisioning, and the hook install after the settings merge.
  Results come back per surface with closed warning codes instead of
  message text.
- The sharing ledger is keyed by surface name, records a value only
  after its write succeeded, and an unreadable ledger starts empty and
  is rewritten instead of blocking every surface.
- The profile state file goes through the same locked writer as folder
  trust (Claude's <file>.lock plus the in-process queue), generalized as
  updateClaudeGlobalConfig. Onboarding and trust are still applied when
  the personal state file is unreadable.
- WSL descriptors build guest POSIX paths; the state-file path style
  follows the injected platform.
- Orca's managed statusLine has one owner in a profile: the settings
  merge never shares it, a user's own statusLine is shared over it, and
  the profile installer follows the default home's slot so a default
  opt-out reaches every profile. remove() takes the same destination;
  the remote installer cannot accept one.
- Prompt history compares file identity (bigint dev+ino) on every
  platform, never drains the shared file into itself, drains retained
  copies in generation order, never reuses a stale cursor, and on Windows
  keeps a replaced default's old copy aside instead of replaying it.
  Directory merges keep going past a failed entry.

* fix(claude): share the user's own hooks and keep merged history whole

A user's own Claude hooks in ~/.claude (notifications, formatters) did
not run under a managed account, because the whole hooks key stayed
private. They are now shared like any other settings key: Orca's own
hook entries and its managed statusLine are stripped from both the
personal value and the profile's current value before the per-key
ledger comparison, so they never travel through the merge and never make
the key look user-owned. Orca entries already in the profile are kept on
write, and the profile hook installer adds them on top as before.

Prompt history: merged bytes that lack a final newline are terminated,
so Claude's next record no longer fuses onto the last merged line. When
a CLI rewrote the profile's history file (old records plus new), only
the lines past the part it shares with the default history are added,
instead of the whole file again.

* fix(claude): close review round 2 gaps in profile setup

Hooks and statusLine sharing:
- When ~/.claude holds only Orca's hook entries, the user's shared hooks
  now read as an empty value instead of a missing key. Removing the
  user's last own hook in ~/.claude therefore reaches profiles that
  never edited it, and deleting the only shared hook inside a profile
  stays deleted.
- A custom statusLine Orca shared, and the profile never edited, goes
  away when the default home drops it. When a shared custom line
  replaced Orca's line in a profile, the profile's statusline marker is
  dropped so Orca's line comes back once the default returns to it; a
  profile that opted out stays opted out. No other key gains deletion.
- install/remove/getStatus with a profile directory refuse when it is
  the default home, or its settings.json resolves to the default one,
  instead of editing System Default's hooks and opt-out state.
- The profile statusline rule reads the default settings under the
  userHome passed to the setup entry, not os.homedir().

Profile state and ownership:
- A malformed `projects` value skips only folder trust (new warning
  code trust-refused); onboarding and shared keys still apply.
- The ownership marker stores only host-local facts (account, runtime,
  distro). The execution host id is the caller's view of the host, so
  it stays in the in-memory descriptor and is not compared.

Prompt history interruption paths:
- With no cursor yet, a retained copy starts past the bytes it shares
  with the default history, so an interrupted share no longer replays
  the whole history.
- A retained name for the shared file itself is removed with its cursor
  instead of lingering until a later scrub makes it look new.
- The Windows link record is read three-state: unreadable stops the
  share instead of reading as "no link". If the record cannot be
  written after linking, the fresh link is undone.
- An unreadable retained copy is reported and no longer blocks linking.

* build(cli): list the new Claude hook modules in the CLI project

hook-service.ts and hook-settings.ts are compiled into the packaged CLI
project, which lists every file explicitly. The statusline policy and
profile destination modules they now import were missing, so the CLI
typecheck failed with TS6307. The CLI still loads hook-service through
the existing managed-agent-hook-controls build entry, which bundles
both modules; neither imports electron.

* fix(claude): close review round 3 regressions in profile setup

- A profile whose hooks hold only Orca's entries and that sharing never
  recorded is no longer treated as a user edit, so the user's first own
  hook in ~/.claude reaches it (for example when the profile was set up
  before ~/.claude had any hooks).
- A retained prompt-history file is removed as a second name for the
  shared file only when the default history does not itself link to it;
  otherwise it holds the only copy and is kept.
- Default-home checks compare file identity: the profile hook
  destination check uses device and inode, and the profile/default
  separation check resolves on-disk case, so a case-only alias of
  ~/.claude is refused on case-insensitive filesystems.
- A test pins that an unreadable leftover session tree no longer blocks
  linking.

* fix(claude): let shared keys leave a profile when ~/.claude drops them

QA found that removing a setting from ~/.claude never reached a managed
account: deleting the whole `hooks` block left the user's hook running
there. Only statusLine followed the default away.

Every shared key now follows the same rule through the existing per-key
ledger: when a key disappears from ~/.claude/settings.json (or
mcpServers/theme from the personal state file), it is removed from the
profile if the profile still holds exactly what Orca last shared. A
value changed inside the account is kept. Keys Orca never shared,
including denylisted ones, are never touched. Deleting the whole hooks
block removes the user's shared hooks and keeps Orca's own entries. A
missing source counts as empty; an unreadable source removes nothing.

* fix(claude): share personal rules, themes, workflows and keybindings into account profiles

A managed account launches Claude with its own config folder, so user-level
rules/, custom themes/ (which a shared `custom:<slug>` theme points at),
personal workflows/ and keybindings.json silently stopped applying. Link the
three directories like skills and commands, and copy keybindings.json with the
same edit-preserving ledger as CLAUDE.md. routines/ stays unshared: routines
belong to the claude.ai account and the folder holds per-run state.

* fix(claude): import the personal CLAUDE.md into account profiles instead of copying it

Claude also loads ~/.claude/CLAUDE.md as a parent folder's memory for any project under home,
so a copied account CLAUDE.md made every such session read the user's instructions twice
(checked live with Claude 2.1.288). An @~/.claude/CLAUDE.md import resolves to the same real
file, which Claude loads once from home, from projects under home and from folders outside it.

* refactor(claude): simplify account profile setup toward the prior art

- Windows keeps each account's history private; drop the hardlink, link
  record and conflict-copy machinery that only Windows reached.
- Share hooks and statusLine as ordinary settings keys: Orca writes the
  same entries into every folder, so the installer finds them present.
  Drops the Orca-entry carve-out, the per-profile statusline follow
  logic and its marker.
- Unreadable ledger is just an empty ledger.
- Share from the user's own CLAUDE_CONFIG_DIR when they set one (marked
  so Orca's injected value is never mistaken for it), and refuse a
  profile at or around it.
- Pin the one canonical profile path spelling in a test.

* fix(claude-accounts): dedupe merged prompt history, drop drained copies, link setup folders by path

- Prompt-history drain appends only lines the shared file lacks, so a purge never re-adds lines.
- A set-aside history copy whose saved offset reaches its end is deleted on the next run.
- Setup folders link to the default home's own entry, not its resolved target.
- The profile gate and folder creation run once, in provisionClaudeAccountProfile.
- installHooks receives only configDir; drop a duplicate test key that fails CI.

* fix(claude-accounts): record installed hooks as Orca-shared; skip symlink tests on Windows

After Orca installs its hooks into an account, record the account's hooks in
the settings ledger so a later run can still bring the user's own hooks in.
Tests that create real symlinks now skip on Windows.

---------

Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
2026-10-06 23:46:35 -04:00
Jinwoo Hong 55c86ceb91 refactor(terminal): report breaches of the two layout invariants; add an unrun load repair (#25673)
* refactor(terminal): report breaches of the two layout invariants; add an unrun load repair

The binding write now checks, after the fast lane, that a terminal is bound to at most one
leaf and a leaf id is in at most one tab, and records a breach on the persistence.pty-binding
span as binding.owner_conflict. It never refuses or changes the write (D16). The load repair
for saved duplicates is added and unit-tested but not called on load; the B2 core turns on
refusal and repair together.

* refactor(terminal): keep the binding write unchanged if the owner check throws; plain-language comments

* fix(lint): merge duplicate removal import in delete-worktree failure toast

Main (#25668) introduced a duplicate import that fails the focused code-quality lint.

* refactor(terminal): keep only the report-only owner check; one definition of same terminal

Defer the unrun load repair to the change that runs it. Repeating relay ids without both
incarnations are no longer one terminal, a leaf held on another host is reported apart, and the
refusal check takes host partitions directly.

* fix(lint): bring structured-agent-session-host back under max-lines

(cherry picked from commit 1a5f687244)
2026-10-06 23:31:02 -04:00
Neil d3e1494674 test: bound memory used by the runtime Electron audit (#26049) 2026-10-06 20:21:34 -07:00
Jinwoo Hong 7c9c3d431a fix(lint): bring structured-agent-session-host back under max-lines (#26050) 2026-10-06 23:17:02 -04:00
Jinwoo Hong 70c6876635 fix(rovo): detect and launch Rovo Dev through acli (#25981)
* fix(rovo): detect and launch Rovo Dev through acli

Rovo Dev ships as the `acli rovodev` subcommand, with no `rovo` binary on
PATH, so Orca never listed it as installed. Detect `acli`, launch
`acli rovodev run` (prompt still typed after start, since a positional
instruction is one-shot), and build resume as the launch command plus
`--restore <id>` so a full-command settings override composes too.

Fixes STA-9645

* test(skills): detect Rovo Dev through acli in the skills CLI fixture
2026-10-06 22:49:18 -04:00
Brennan Benson 923ef7cff1 fix(native-chat): refuse a Command that can't run instead of silently running the stock CLI (#25667)
* fix(native-chat): honor custom Claude and Codex launch commands

* fix(native-chat): sync launch failure localization catalogs

* feat(native-chat): run Settings → Agents → Command as the native chat program

Native chat ignored the Command setting, so a user who pointed it at a
custom Claude or Codex build still got the stock CLI. The host now reads
the Command per acquisition and per model-catalog probe, resolves it on the
execution host as a program (an absolute path, ~ path, or a name on PATH),
and spawns it with the normal structured arguments. A set Command that is
not a runnable program refuses the start with a new agentCommandNotRunnable
failure that names the setting; it never falls back to the stock CLI.

* fix(native-chat): find a configured Claude Command on the launch PATH; clearer copy

A Claude Command given as a bare name was looked up on Orca's own PATH,
not the PATH the Claude child launches with (login shell plus Settings →
Agents environment), so a name only that environment provides was refused.
Codex already resolved against its launch environment. Claude now reads
its launch env first and resolves a configured name against its PATH and
home; the stock lookup with no Command is unchanged.

The failure copy now states the rule (a program path or name, no
arguments or variables) and names the Settings control (Reset). The
catalog fingerprint doc notes that the program is not keyed, so a chat on
an older program keeps refreshing the shared entry until it ends.

* fix(native-chat): one PATH key for the Claude lookup and the child on Windows

On Windows, an inherited `Path` and a Settings → Agents env `PATH` both
reached the Claude child. The configured-program lookup read the first
inserted twin (inherited), while Node's child_process keeps the
lexicographically first (`PATH`, the user's), so a bare-name Command found
only on the Settings PATH was refused.

The Claude launch env now drops inherited case-twins of the overlay's
variables (the rule the structured shell env already applies), so the
lookup and the child read one PATH. pathEnvOf follows Node's win32 key
rule, and Codex's lookup uses it too, so both agents read PATH one way.
es/fr/ja copy now says the Command must be, as the English does.

* fix(native-chat): refuse a Command that can't run instead of running the stock CLI

A saved Settings → Agents → Command that did not resolve to a program (a
command line such as `codex --profile work`, `FOO=1 claude` or `npx …`, a
missing or relative path, a non-executable file) silently ran the stock
Claude/Codex in native chat and its model-catalog probe. The user's choice
was ignored with no sign of it.

A set Command that does not resolve now refuses the start with a new
agentCommandNotRunnable failure naming the setting, beside the existing
Retry; the catalog probe refuses the same way and never lists the stock
CLI. A blank Command keeps the stock lookup. On Windows a resolved file
without a spawnable extension (.exe/.com/.cmd/.bat) refuses too, instead
of failing to spawn with no word about the setting.

* chore(native-chat): correct comments that say native chat ignores the Command

After #25721 native chat runs a runnable Command, and this PR refuses one it can't run; four comments still said the Command applies to terminal launches only. The pre-spawn test fixture also drops the saved value from its error message, matching the resolver.
2026-10-06 19:47:31 -07:00
Brennan Benson 5b1bb78d14 [STA-6940] Prepare file drops for element owners (PR 2 of 6) (#25748)
* feat(file-drop): add element owner preparation plumbing

* fix(file-drop): preserve feedback acceptance and destination ordering

* fix(file-drop): keep path resolution in filesystem namespace
2026-10-06 19:32:15 -07:00
Brennan Benson 0bdcaf36ed fix(claude): start a Claude chat with its saved options and send the first message at once (#25152)
* fix(claude): end a Claude start that never answers initialize after 120 s

* Read the Claude startup deadline inside startup; fix a stale test comment

* fix(claude): start a Claude chat on its initialize answer, not on a frame only a SessionStart hook sends

Startup waited for system/init or a SessionStart hook frame as well as the initialize answer.
Before the first turn only a SessionStart hook sends one, and Orca adds that hook only through
its optional status hooks, so with them off the first message was held forever. Startup now
lands on the initialize answer; a start frame already seen is still checked, and one naming
another session ends a started session. The deadline drops to 90 s so it fires inside the
host's 120 s start wait.

* fix(claude): time the Claude start by silence, and fail it at once on another session's frame

Claude answers initialize only after its SessionStart hooks finish, so a total-time deadline
would fail every start behind a slow hook. Each start frame now restarts the clock. A frame
naming another session fails a start still waiting on initialize at once, as before.

* test(claude): a real Claude chat starts and answers with every hook disabled

* Say what the start-frame re-arm covers, and check only start frames in the hook-less real test

* fix(native-chat): a Claude chat starts with its saved options and takes its first message at once

Saved model, effort, Fast and permission mode are passed as launch options, checked against the
account's cached model catalog, instead of restored by control requests after initialize. With
nothing left to restore, the host no longer holds a message until the CLI answers initialize, and
the 90 s startup deadline is gone. A Stop on a start that never answers ends the child and settles
what it was handed as stopped. A failed result that repeats the turn's own API error reply writes
no second row.

* fix(native-chat): a host stop of a Claude start fails the message it was handed, with one row

With no start-hold the delivery loop no longer sees a host stop of a start it waited on. The
child's end now rejects what it handed over with the host-stopped words and writes the one row,
as an exit of its own would; an idle start the host stops still goes quietly.

* test(native-chat): a Claude chat's first message is written before initialize answers

Rewrites the tests that encoded the start-hold, the startup deadline and the option restore to
the new contract, and adds: saved options at launch (catalog checks, bypass, fresh-session Fast),
a message written before initialize answers (adapter and runtime), Stop on a start that never
answers (stopped, child closed, nothing working), and an API error said once.

* revert(native-chat): keep a failed Claude turn's error row

The shared turn fold already shows a failed turn's error once after it settles, as an error;
dropping the row left the CLI's synthetic reply looking like something Claude said.

* fix(native-chat): pass saved Claude options unchecked, heal a retired model on the CLI's word, and never leave an unrun message in doubt

- Saved model, effort and Fast are launched as picked; only values no Claude can parse are left
  out. The pre-spawn cache check is gone.
- A saved Fast on for a new conversation is applied once the settings readback shows no
  per-session opt-in (dropped when there is one, or when the model is listed without Fast), with
  nothing waiting on it; the record keeps the pick.
- Under an Agent Permissions bypass, a saved narrower mode launches with the allow flag so bypass
  stays reachable.
- A turn whose reply is the CLI's model_not_found for the launched model drops that model from
  the record; the launch's own row for it is kept out of the account model cache.
- A child that ends before it answered initialize, for any reason, settles every message it was
  handed as not sent (cancelled for a Stop).
- A launched effort the CLI reports only as `applied.effort` is confirmed from there.
- The untimed-initialize comment is back to main's text.
- A real-CLI test for a message written before initialize answers, under saved options.

* fix(native-chat): type the close's ended event and the start-exit test fixtures

The close's ended event is typed as the adapter event so its optional startupUnanswered spread
fits exactOptionalPropertyTypes; two tests guard the fixture's optional generation, and the
hung-start fixture records initialize on the fake connection it holds.

* fix(native-chat): a Claude model heal keeps a later pick, a refused Fast is dropped, no allow flag

- `options-skipped` carries the retired value; the record drops it only while it still holds it.
- `started` carries the values a heal retired, and the record does not take them back from the
  CLI's report of the same value.
- A saved Fast on a new conversation is applied before `started`: a refusal drops it and records
  it skipped, as main's refused restore did; silence keeps it wanted and unconfirmed.
- A saved narrower mode under an Agent Permissions bypass launches without any bypass flag again:
  the allow flag is one older CLIs reject at start. Kept as a known limit.
- The real-CLI test asserts the message was written before initialize answered.
- The fake reports a launch effort only under `applied`, and a misplaced doc comment moves back.

* fix(native-chat): a new Claude chat reports started before its saved Fast is applied

The Fast apply on a new conversation now runs after `started`, so a Stop interrupts a running
first turn and an option write is not refused while the round trip is out. A refusal drops the
pick through `options-skipped`, in order after `started`; silence keeps it unconfirmed. The
launch's unreachable skipped-model branch is gone.

* fix(native-chat): a healed Claude chat goes back to the default model live; comments match the no-hold design

When the CLI says the launched model does not exist, the live child is also put back on the CLI's
own default (set_model with no model, fire-and-forget), so later messages in the same chat run; a
user's pick sent after it wins, and a refused or unanswered reset only logs. Comments that still
described the start-hold or the option restore now describe the launch options and the
handed-over, never-echoed rule.

* fix(native-chat): a message handed to a Claude start that never answered is kept as main keeps an unsent one

A child that ended before it answered initialize ran nothing it was handed, the same fact as a
send accepted and never handed over. Its end now settles those sends exactly as the chat settles
a queued send for that end: a quit keeps a person's message as a held card (restart words), a
close keeps it as a held card (closed words), a person's Stop withdraws it as cancelled, and a host
stop fails the start with one row.

* fix(native-chat): a quit during a Claude start that never answered offers no resume for the message it keeps as a card

The restart snapshot now reads the same never-answered fact the exit does, so a message handed
to such a start counts as queued work, not as work to resume. The retired-model reset comment
names the default it really applies.

* refactor(native-chat): the saved permission-mode launch helpers live with the spawn options that use them

Keeps claude-structured-launch-resolution.ts within max-lines once merged with main, and names
the hung-start test envelope's field type.

* fix(native-chat): a Claude chat's saved options take precedence over the agent Arguments' own flags

Main now passes the saved agent Arguments to the Claude child, and the SDK writes them after its
own options. An Arguments --model or --effort therefore reached the CLI as a second flag after the
chat's saved pick (a commander CLI keeps the last), and a saved Fast's launch settings replaced an
Arguments --settings file outright. The saved model and effort now stand in for the Arguments'
flags, and a saved Fast beside an Arguments --settings is applied by the start instead of at
launch.

* test(native-chat): the hand-built Claude session in the options test carries fastModeAtStart

* test(native-chat): the queued rig's start spy carries a named Mock type

An unannotated vi.fn() inferred @vitest/spy's internal Procedure, which CI's typecheck cannot
name in the factories' inferred return types (TS2883).

* fix(ci): run the PR's SQLite-backed tests in the Node runtime project

Lists the hung-start Stop test, renames the send-during-startup entry from its old name, and
carries main's own two entries from #26010 so the boundary test passes before the next merge.
2026-10-06 19:19:47 -07:00
Brennan Benson 7a59aa39d0 fix(worktrees): decide setup before git worktree add on the host's create (#26006)
* fix(worktrees): decide setup before git worktree add on the host's create, so an undecided ask repo leaves nothing behind

The host's create (agent.launch, worktree.create, the phone and the CLI) only checked the setup policy
after adding the worktree. An ask repo with no setup decision then threw with the worktree already on
disk and no worktrees-changed event. It now refuses before the add, as the desktop create does, and a
setup hook the new branch adds (one nobody decided on) is skipped with a warning instead of failing a
create that already exists.

* fix(worktrees): read the host create's setup decision in one place, after recording its host

The pre-add refusal now runs after the create records its execution host, so a refused create's
failure telemetry still says where it ran; it still runs before any git work. Both setup checks
read the decision from one helper so they cannot disagree, and the post-add skip logs once.
The refusal test now also proves no base resolution, fetch or branch naming ran and that the
hooks came from the main checkout.

* test(worktrees): cover a setup hook only the new branch adds through the managed create

An ask repo with no decision, whose setup hook exists only in the new worktree's orca.yaml, now
creates, refreshes the worktree list, reports setup as skipped and returns the warning.

* fix(cli): name the --setup flag when a worktree create needs a setup decision

The host refuses an undecided create in an ask repo with "Setup decision required for this
repository", which desktop and phone answer in their own UI. The CLI now adds the next step:
pass --setup run or --setup skip.
2026-10-06 19:11:48 -07:00
Brennan Benson b3b6c5dc13 Give native chat names one source for tabs, sidebar and AI Vault (list and search) (#25986)
* Give native chat names one renderer source and drop Vault's name repair copies

The host's saved conversation name now rides the structured session status
feed, which already exists per host, is keyed by the durable session id, and
keeps a closed chat's summary. Tab strip, sidebar rows and AI Vault (list and
search) read it through one hook and one display order (tab alias, saved name,
host label). Vault no longer copies names into its cached results, so the
projection, recovery and pending-title modules and their tab-snapshot lanes are
removed. Indexed search hits now carry the native owner and saved name from the
host that indexed them.

* Type the sidebar name test fixture without an assertion

* Keep Vault search working when the chat host will not install

Naming and owning search hits is bookkeeping: if the native chat host fails to
install, return the plain hits instead of failing the search. The runtime RPC
only installs the host for clients that will receive the owners.

* Publish chat names to the feed independently of the tab retitle

A failed feed publication no longer skips retitling the open tab. The publish
now lives in the naming deps, where a test covers it.

* Note why the status feed must keep closed chats' summaries

* Bound names and owner ids that come from a paired host

Drop a published chat name the record store would refuse, and cap a search
hit's owner workspace id at the same length the list row and record use.

* Let native chat search hits from a paired host open their chat

A paired host's search hits carry no resume command, so the row disabled
every open action even for a native chat it can open through its owner, as
its list row does.

* Ignore workspace ids that name object members in tab lookups

A paired host's row or search hit could carry a workspace id such as
"constructor", which read an Object.prototype member as a tab list and broke
the render. The shared tab index now reads only own workspace entries.
2026-10-06 19:04:04 -07:00
Brennan Benson 9fdd90ffa7 feat(orchestration): a native chat can be a dispatch worker, like a terminal agent (#22972)
* feat(native-chat): a queued card can record a Dispatch's task as its source

A chat worker's task is held in the chat's queue like any agent message, so
the card needs to say which Dispatch it is from: the sender, run, task and
Dispatch ids. An older build reads an unknown kind as the person's card and
sends it as written.

* feat(orchestration): a chat can be a dispatch worker, like a terminal agent

`dispatch --to orca_session_id:<id>` and `worker-start --terminal
orca_session_id:<id>` now accept a chat on this host instead of refusing it.

- The chat is refused only where mail to it would be: unknown, a provider id,
  another host, or closed. A chat can't be its own coordinator's worker.
- Its task goes through sendAgentTurn as a queued send, as mail notices do:
  an idle chat starts a turn, a busy one holds a card naming the Dispatch.
  The operation id is derived from the Dispatch, so a resend replays.
- worker-start reads that outcome as the terminal path reads its write; a
  card held behind a running turn is handed over with its start unobserved.
- The Dispatch names the chat by its /clear root, with no pane or process,
  and records its Orca session id, so its own sub-dispatches nest under it.
- Mail to dispatch:<id> reaches the chat; worker-show/list/read read the
  chat's session records; stop and abandon never close the chat; closing the
  chat fails its Dispatch as closing a terminal does.
- A dispatch preamble for a structured session names the CLI the structured
  mail lane names.

* test(orchestration): a chat worker's reach, and its Dispatch across a /clear

At rest is live, a closed chat has exited, and another host or nothing to
read is unverifiable. A /clear keeps the chat's Dispatch, and the session
that continues the chat reports as it.

* fix(orchestration): chat worker review round 1

- dispatch --inject is a keepalive-backed wait, so a chat that takes a while
  to accept its task no longer reads as a dead runtime at the 30 s idle cut.
- An injected task whose delivery is unknown keeps its Dispatch open, as an
  unknown worker-start does, instead of failing it while the chat may run it.
- A busy chat's worker-start receipt says the task waits as a card in the
  chat's queue, and that worker-abandon does not remove that card.
- A close that puts the chat's tab back no longer fails its Dispatch: the
  closed-chat settle runs after the close's outcome, not at its hide.
- worker-start adopting a chat says it gave the task to the chat, instead of
  claiming it started a terminal agent; the mode value is unchanged.
- Test: a chat whose agent has not taken its task reads outcome_unknown.

* test(native-chat): a close's hide defers its hidden notice until the close settles

The rollback suite pinned the hide's arguments; it now expects the deferred
notice, and that the notice is sent once, after a rolled-back tab is back.

* fix(orchestration): a chat worker whose successor is unknown is unverifiable, not exited

A /clear successor this host has no record of is missing evidence, not proof
the chat is gone. Only a closed chat reads exited, so only a closed chat
settles its Dispatch; a lost successor reads unverifiable with that reason.

* fix(orchestration): chat worker review round 2

- A chat's close notice settles only that chat's Dispatch. A scan of every
  chat Dispatch read another chat whose close was still in flight (its tab
  hidden, maybe to be put back) as closed and failed its Dispatch.
- Only a dispatch --inject into a chat is a long poll; a terminal inject
  writes and returns, and keeps its short-RPC slot.
- Receipt wording: worker-start says it gave the task to the chat, and a
  queued task is sent when the chat's queue reaches it.

* fix(orchestration): chat worker wording, review round 3

- worker-start's mode sentence for a chat states the placement, which is
  true whether the task is delivered, queued or refused.
- Comments and a test title no longer claim a close notice re-derives every
  chat worker, or that a queued task waits for the current turn.

* fix(orchestration): build a structured worker's task sender from the start's own ids

The minted worker's preamble names who its task is from with the run, task
and Dispatch ids worker-start already holds, so it reads no Dispatch row.

* fix(orchestration): point a chat worker at its held Dispatch mail again

A chat worker's coordinator mail lands in its Dispatch's mailbox, but a
chat's idle edge re-derived the Dispatch mailbox only for a party with a
terminal handle. A pointer lost with the provider then waited for new mail
after a restart or a /clear. The re-derivation now looks the Dispatch up by
the party's address, which names a worker by its handle and a chat by its
root.

* fix(native-chat): open a task card's sender by address; a task carries no mail

Opening an agent message's sender looked up the mail it carried, which only a
mail notice has; with the task kind in the union that read no longer typed.
A task's coordinator is found by its address alone.
2026-10-06 18:36:21 -07:00
Jinwoo Hong 825d7bd5a9 test(vitest): run agent-launch-instant-tab in the SQLite runtime project (#26028)
#25430 added a test that opens a real agent-session record store, but not to the
SQLite runtime list, so vitest-sqlite-runtime-boundary fails on main.
2026-10-06 21:33:40 -04:00
Brennan Benson cdd0b7a491 fix(native-chat): stop flashing a reconnecting line on a stream drop (#24898)
* fix(native-chat): drop the per-chat reconnecting line on a stream drop

A single transcript stream drop flashed 'Reconnecting to this chat…' above the
composer for about a second. An unnamed read failure now adds nothing: the
transcript and composer stay, and the host's own status shows reachability.
Named or final failures keep their line.

* fix(native-chat): keep a loaded chat as it is when its stream drops naming nothing

A loaded chat with no messages yet still flipped to the full-pane "Could not
load conversation" on every stream blip: the read owner stored any read
failure as status 'error', and the pane shows that error whenever there is no
transcript. Once the chat has loaded, a failure that names no reason is now the
transport's own retry and is not stored; a named or final failure, and any
failure before the first load, are reported as before.

The "names a reason" check moves next to the final-refusal check so the owner
and the failure notice share it. The older-page read moves to its own module to
keep the read owner within its line budget.

* fix(native-chat): word a read failure beside a chat that never loaded

A chat that never loaded but already shows the user's own message (a launch
prompt or a queued send) said nothing when its first read kept failing with a
failure that names no reason. The loaded case is handled at the read owner, so
any read failure the pane sees beside messages is named, final, or from before
the first load; the status area now says its words in every such case.

* fix(native-chat): keep a loaded chat as it is only when contact is lost

The read owner skipped any loaded-chat failure without a named reason, but an
older host sends every refusal without details, so a refusal it really sent
(such as a damaged history) went unshown while the read retried in silence.
Skip only failures the host sent no refusal for, which is lost contact; any
refusal is the host's answer and is shown. The shared 'named' check is no
longer needed and is removed.

* docs(native-chat): say what a read failure without a refusal is treated as

Comment-only: the read owner treats a failure with no host refusal as lost
contact, which covers hosts too old to attach refusals; the test's reasonless
case is a reason this build doesn't know.

* Say a chat's host outage once, above the composer

A loaded remote chat looked live through a long host outage while sends sat
in the outbox. Derive a host-scoped notice from the host's connection state:
'<host> is reconnecting…' after a 2 s grace, '<host> is offline' at once with
the status bar's Connect as Reconnect, and a composer placeholder saying
sends go out when the host reconnects. A lost read beside the notice adds no
line of its own, and older history waits for the host.

* Drop the outage placeholder; no Reconnect for a refused host

A send while the host's transport is down fails and waits for its own Retry,
so the composer must not promise it will go out on reconnect. A host that
refused us (auth, protocol) still reads offline but offers no Reconnect,
which would be turned away the same way. Name the host with the shared
display-label selector, and keep the notice's live region mounted so it is
announced.

* test(native-chat): mock the font-size hook main renamed in the host-outage test
2026-10-06 18:31:31 -07:00