Commit Graph
3791 Commits
Author SHA1 Message Date
Neil 720c3299ba fix(ssh): require a host death certificate before recreating a pane, and unstick expired leases (#18013)
* fix(ssh): match an expired lease on where its leaf lives now, not its frozen tab

A lease freezes tabId at write time, but detachTerminalPaneToTab moves a live
pane, so the stored tab is the one the pane LEFT. getRecentExpiredSshLease
required lease.tabId === tabId, which is wrong in both directions: a viewer on a
stale mirror matched under the abandoned coordinates (and resolvePersistedStable
PaneOwner then reads an empty layout for that tab, so adoptStablePane is skipped
entirely and a fresh shell is spawned over a possibly-live one, binding the same
leaf in two tabs), while a viewer using the pane's real coordinates matched
nothing and got terminal_not_recoverable.

Resolve the leaf's current tab the way restoreReattachedPtyRuntime already does
and compare against that, falling back to the frozen tabId only when nothing can
say where the leaf lives. Both workspace partitions are read because SSH spawns
bind into ssh:<target> while reattach binds into local.

* fix(ssh): let a proven reattach take an expired lease back to attached

#17965 authorized reattach from `expired` but the state machine refused the
transition back, so a lease that reattached and proved itself alive stayed
`expired` forever. That silently exempted a demonstrably running remote shell
from `ssh:reset` (skips `expired`), from the SSH_TERMINATE_RECONNECT_REQUIRED
ownership fence in `ssh:terminateSessions` (marks it not-owned), and from the
quit-time `detached` sweep, and made it permanently ineligible to win
supersession so its own successors never retired their predecessors.

Only the id-qualified caller carries per-pty proof: markSshRemotePtyLeases
AttachedAsync is fed the relay's `attachedLeaseIds`, so an unqualified bulk mark
over a whole target still cannot revive `expired`. `terminated` stays absorbing.
Re-entering `attached` drops supersededBy/relayIdRecycled, since route
retirement belongs to the shell that lost the pane and this one just proved it
is not that shell — the same invariant upsertSshRemotePtyLease enforces.

* fix(ssh): make the pane-recovery liveness gate refuse without positive evidence of life

The gate refused only `live` and `unverifiable` and passed on `null` — but the
register is an in-memory Map, so `null` is equally what a fresh app start, a
never-asked host and a certified death look like. Absence of evidence was
reading as authorization to spawn a shell over a possibly-live remote process:
`!pty.connected` is cleared for every PTY a dropped relay owned, and `expired`
only ever says the CLIENT lost its route.

- `exited` is now RETAINED rather than deleted, so the register is three-valued
  in the map as well as in the type. Its one writer is a host-delivered exit
  frame — an exit with a real code, or an explicit `hostExitConfirmed` — which
  records the certificate instead of merely dropping the doubt.
- `recoverTerminalPane` refuses on `live` and `unverifiable`, and deliberately
  does NOT demand a positive `exited`. The only answer that ever reaches this
  gate is a reachable relay reporting no such id, and that is a union: pty.attach
  throws not-found for an unknown id with no liveness check, and a relay restart
  makes every previously minted id unknown (ids carry a per-start
  `ptyIdMintEpoch`). No writer of `exited` co-occurs with a reattachable
  `expired` lease either — a host-delivered exit frame tombstones the lease
  `terminated` — so requiring one would close the gate permanently.
- `handlePtyReattachFailure`'s not-found branch publishes `code: -1` to the
  renderer and does not call `runtime.onPtyExit`. The relay's not-found answer is
  not a death certificate, and #17963's ratchet on the same branch pins that.
- The inventory's `observed === false` hunk keeps dropping doubt rather than
  asserting a death: `pty.listProcesses` returns the relay's CURRENT session map,
  so a restarted relay omits every previously minted id whether or not those
  shells died — the same union, one hop away.

A live or unprovable pane refuses; a disowned one still recovers. No wire change.

The gate's ratchets live in terminal-pane-recovery-liveness-gate.test.ts:
config/vitest.config.ts — the config CI runs — matches only `*.test.ts`, so cases
placed under orca-runtime-tests/*.spec.ts would never execute.

* fix(ssh): gate paired-viewer pane recovery on the narrowed session-gone predicate

isSshSessionGoneError landed on the IPC transport, which never calls
terminal.recoverPane. The one caller that does — recoverExpiredHostPane in the
paired-viewer transport — still triggered on a bare SSH_SESSION_EXPIRED
substring, so the identity-mismatch reply (the relay found a LIVE PTY under that
id owned by another pane, which is evidence of presence) still asked the HUB to
replace the pane, putting a second agent on one transcript. Main already refuses
the respawn on that same reply; this makes the two agree.

A pane whose shell genuinely died is unaffected: plain SSH_SESSION_EXPIRED still
matches. The mismatch reply now surfaces as an error instead of a respawn.

* test(persistence): update the reattach ratchet for expired-lease reclaim

markSshRemotePtyLeasesAttachedAsync is id-qualified, so a named pty that
proved itself alive now returns to attached instead of staying expired.
2026-09-02 22:31:59 -07:00
Neil 08c7152ab6 fix(ssh): compare a lease's relay pty id against the pane's app id (#17969)
`getRecentExpiredSshLease` compared the stored lease ptyId (relay form,
written through `toStoredPtyId` -> `toRelaySshPtyId`) raw against the
runtime's app-form `pty.ptyId`, so `'pty-3' === 'ssh:target@@pty-3'` never
held and `recoverTerminalPane` refused every real SSH pane. Normalize with
the same tolerant helper the binding reader already uses, now shared as
`toComparableRelaySshPtyId`.

Switching the path on is only safe on top of #17957 (respawn gated on the
runtime liveness verdict), #17965 (`expired` no longer withdraws bindings)
and #17966 (supersession and id recycling carry their own marks).
`recoverTerminalPane` additionally refuses a lease those marks disqualify,
so it acts only on an `expired` lease that means "reattach gave up".

The path's outcome is a reattach, not a respawn: `createTerminal` calls
`adoptStablePane` first, which attaches attach-only to the retained
binding and only falls through to a fresh shell once the host itself
answers that the PTY is absent.
2026-09-02 21:48:17 -07:00
Neil f2e95e7860 fix(ssh): separate a superseded lease from an orphan so reattach can tell them apart (#17966)
`expired` was one word doing two unrelated jobs — "a newer lease won this pane"
and "reattach lost contact" — so `reattachKnownPtys` had to exclude all of them.
That kept the 2 -> 19 -> 20 fan-out fixed at the cost of never bulk-reattaching a
genuine orphan; those recovered only through the slower `adoptStablePane` path.

The blocker cited in #17965 does not apply. The STA-3077 note guards
`upsertSshRemotePtyLease`'s match against a RECYCLED `pty-N` after a relay
restart. `supersedeSiblingLeasesForPane` is a different path and already holds
`winner.ptyId` when it expires a predecessor, so recording which lease won needs
no relay-start identity.

`SshRemotePtyLease` gains two optional marks, each meaning exactly one thing:

- `supersededBy` — the winner's stored-form ptyId, written only by supersession.
- `relayIdRecycled` — written only by the pending-stop replay's
  `relay-id-recycled` retirement. That retirement wrote `expired` *purely* to
  keep the lease out of the reattach that runs one step later ("hands the user's
  old pane to whatever process now holds the recycled id"), and the reattach
  fences on paneKey/tabId, never on incarnation. Relaxing the filter without
  this would have silently reopened that hole.

Bulk reattach now skips a lease carrying either mark and re-adopts the rest, via
one shared `sshRemotePtyLeaseAllowsReattach`. `terminated` is untouched.

Recycled-id safety: both marks are dropped whenever the id is re-upserted
`attached`/`detached`, so a relay that renumbered onto a new shell cannot inherit
its predecessor's mark. Supersession also stamps an ALREADY-expired predecessor
for the same pane — same evidence, and it is what bounds the reattach set, since
otherwise every past orphan for that pane would stay reattachable forever. Its
`updatedAt` deliberately stays put: bumping it would make a stale lease look
recent to `getRecentExpiredSshLease`.

Persistence: the lease loader is a strict whitelist, so both fields are named in
`normalizeSshRemotePtyLease` or they would be stripped on every launch. Absence
reads as "orphan", which is the only thing an older build could have meant, and
an older build ignores keys it has never heard of (remote-wire Rule 1).
2026-09-02 21:39:08 -07:00
Neil e4c279fa60 fix(ssh): let an expired lease reattach its orphan instead of stranding it (#17965)
* fix(ssh): let an expired lease permit a reattach instead of unbinding the pane

`expired` never means the remote shell exited. Every writer records that the
CLIENT lost its route — a superseded sibling, a recycled relay id, a
persistPtyBinding refusal made *after* pty.attach proved the shell alive, a
failed reattach indistinguishable from a relay restart, a relay reset whose
kill may not have landed. docs/reference/ssh-execution-boundary.md grades all
of those `unverifiable`.

Three readers treated it as death, and together they made the pane unable to
reach a process that is still running:

- `isRestorablePtyBinding` / `hasRestorableSshRemotePtyLease` refused to replay
  a durable binding a renderer snapshot had omitted.
- `markSshRemotePtyLease(s)` wiped the persisted pane->pty binding, which is
  what makes `resolvePersistedStablePaneOwner` return null, `adoptStablePane`
  give up, and `createTerminal` cold-spawn a replacement. The user's terminal
  comes back empty and the running job is orphaned and invisible.

Only `terminated` now withdraws a binding: it is the operator-close state
(`ssh:terminateSessions`) and the one written after a host-acknowledged stop.

This authorizes a reattach ATTEMPT, never a respawn, so #17957's gates are
untouched and in fact fire less often — where the pane previously went straight
to a fresh spawn it now attaches first. A genuinely dead shell still converges:
`attachStablePaneOwner` retires the binding on `isPtyAlreadyGoneError` (the
relay's own absence answer, not a message match) and falls through to a fresh
spawn, so no pane retries forever.

Supersession keeps its own binding scrub in `supersedeSiblingLeasesForPane`,
where a NEWER lease for the same pane is the evidence — the 2 -> 19 -> 20
reattach fan-out stays fixed.

* test(persistence): split SSH remote PTY binding partition cases into their own file
2026-09-02 21:35:28 -07:00
Neil 8cf6e12009 fix(wsl): stop runWslProcess inheriting a removable spawn directory (#17837)
#17834 named an explicit Windows cwd for the wsl.exe spawns in
wsl-command-resolution and wsl.ts, but runWslProcess -- which 25 production
files route through, the bulk of WSL spawns -- still passed none, so #16463
survived on the majority path: an inherited cwd that is later deleted (the
worktree Orca launched from) fails every subsequent spawn for the session.

The test asserted `cwd` was undefined, so the fix turned it red. That
assertion was over-tight rather than a contract this violates. Its name and
the production comment both state the real invariant -- "that is a *Windows*
directory for wsl.exe", i.e. the GUEST path must never leak into it -- and
withGuestCwd still cds inside the guest, so the invariant holds. Undefined was
a proxy for it, and an inherited directory satisfies the proxy while being the
bug. Retargeted to assert what is actually meant: not the guest path, and
present.

Deliberately not silent: the salvage agent hit this, reverted rather than
overrule a documented contract in a module it was not sent to change, and
escalated. That was the right call to escalate; this is the answer.
2026-09-02 21:33:45 -07:00
Neil 57681ecd09 fix(remote): resolve the spawn cwd, the node manager dir, the vault host and the scrollback seed (#17952)
* fix(remote): resolve workspace cwd, mise Node, host scope, and TUI scrollback honestly

#15296 relay: a folder workspace id (`folder:<uuid>`) carries no path, so the
worktree-id split yielded nothing and $HOME silently won. Resolve the spawn cwd
through worktreeId -> ORCA_WORKSPACE_ROOT -> host default, and refuse an agent
spawn outright when a folder workspace names a root this host cannot resolve.

#11733 ssh: generalize the NVM dotfile scrape into `orca_dotfile_dirs` and drive
mise off `MISE_DATA_DIR` / `XDG_DATA_HOME` instead of a hardcoded
`$HOME/.local/share/mise`.

#13713 ai-vault: an unresolvable workspace host is `unverifiable`, not local.
Widen the default scope to every host rather than scanning the client's own
history and reporting "No agent sessions found".

#6106 terminal: hydration asked the renderer for `scrollback: 0` while an
alt-screen TUI was up, which drops the normal buffer's shell history rather than
the TUI bytes. Drop the flag; readers already split the two buffers apart.

* fix(remote): stop the relay answering host questions for a guest execution host

Three findings from review of the spawn-cwd resolver, all the same shape: a path
question answered against the wrong host, or with the wrong key.

- resolveRelaySpawnCwd refused an agent launch whenever a folder workspace named
  a root that did not stat on the relay. But relayHostDirectoryExists stats the
  relay's *own* filesystem, and the relay supports WSL shells, so a folder
  workspace on a Windows relay launching into WSL now threw where it previously
  spawned -- contradicting the function's own doc comment, which says an absent
  path for that exact host pair is a miss, not a refusal. Thread the shell's
  execution host in and demote the refusal to a miss when the spawn does not run
  on the relay's filesystem.

- requireRelaySpawnCwd's doc claims both call sites route through one resolver
  so the fence can never be keyed on a directory the spawn won't use, but the
  fence key was still computed with the non-stripping splitWorktreeId while the
  cwd used splitWorktreeIdForFilesystem. For a `::workspace:<uuid>` id those
  disagree by construction, in adjacent lines: the removal fence guarded a path
  no spawn ever enters. Same defect in shutdownForWorktreePath and the revive
  path; all three now use the filesystem split.

- The remote Node probe expanded `$HOME` and `~/` prefixes out of a dotfile
  assignment but not `$XDG_DATA_HOME`, so `MISE_DATA_DIR=$XDG_DATA_HOME/...`
  was used as a literal directory name. Add the case arm, defaulting to the
  POSIX `$HOME/.local/share` the seed value already uses -- sshd's exec channel
  usually has no XDG_DATA_HOME at all.
2026-09-02 21:33:41 -07:00
Neil 9e9b80cb37 perf(relay): stop two unbounded growth terms behind the long-session SSH slowdown (#17818)
Two costs grew for the life of an SSH session and never came back down.

1. The relay port scan walked every process in /proc and readlink'd every fd
   even after every listening socket already had an owner. Cost was
   O(host processes x fds) per scan, repeating for the session's life. Exit as
   soon as every inode is attributed.

2. SshPtyModelAdmission kept closed provider generations in a Set<number>.
   Provider generations are a process-global monotonic counter shared by every
   SSH target, so the set gained one entry per relay reconnect forever. After
   500k reconnects main retains ~10,234 KB / 500,000 entries; with this change,
   18 KB / 1 range.

Closed generations now live in SshPtyClosedGenerationRanges, which collapses
contiguous closed runs. Membership stays exact -- a generation below the
high-water mark can still be live on another host, so a high-water
approximation would reject a healthy target's output.

The range container's has()/add() were a linear scan; both are now binary
search. has() is on the per-output-chunk admission path, so a scan would have
traded a bounded Set lookup for one that degrades with fragmentation. This also
speeds up ssh-pty-output-generation-guard.ts, which already uses this container
on main.

Known limitation, deliberately not addressed here: the closed-generation set is
bounded in the healthy case (one range) but unbounded when generations leak,
since each leaked generation leaves a permanent gap. Sublinear is not bounded. A
live-generation set would be bounded by construction and is the better
long-term design; that is a follow-up.
2026-09-02 21:33:34 -07:00
Neil 64dac75d9b fix(ssh): stop respawning panes on client-side-only absence evidence (#17957)
* fix(ssh): stop respawning panes on client-side-only absence evidence

Three respawn gates acted on evidence weaker than host-attested exit.
Per docs/reference/ssh-execution-boundary.md, loss of contact, a failed
reattach, an identity mismatch and absence from a client map are all
`unverifiable`, never `exited`.

Gate 1 (ipc-pty-connect.ts): "belongs to SSH connection" is minted by the
id router from a pure client-side string compare, before any relay is
asked, and still returned `sessionExpired: true` -> fresh PTY + agent
resume. After an SSH target re-adoption the "other" connection is the
same machine, so that puts a second `claude --resume` on the transcript
the surviving PTY still owns. Now returns undefined with no error, which
routes the pane to recoverUnverifiableDirectSshReattach (remount +
reattach, no shell restart) and keeps #7661's no-red-toast outcome.

Gate 3 (ssh-reconnect-pane-retry.ts): `!tabPtyId` read `tab.ptyId`, which
is only the single-pane fallback for legacy attach. It diverges from the
real records deterministically: workspace-terminal-reconnect fills
ptyIdsByTabId from the leaf map but writes tab.ptyId only when a
tab-level id survives, and clearTransientTerminalState nulls tab.ptyId on
every hydrated row. Both leave live leaf PTYs with a null fallback field,
arming a generation bump onto the fresh-spawn path. Now consults
ptyIdsByTabId and the layout leaf map too; a tab with no PTY in any
record still retries.

Gate 2 (recoverTerminalPane): an `expired` lease plus `!pty.connected`
authorized createTerminal. Every writer of `expired` records that the
CLIENT lost its route, not that the shell died. Now also requires the
runtime's own liveness verdict to be neither `live` nor `unverifiable`,
and ssh-relay-session records markPtyLivenessLive at the persistPtyBinding
refusal, which is reached only after pty.attach succeeded. See the report
for why this branch is currently unreachable for SSH panes.

* fix(ssh): let the respawn gate see the relay's own absence answer

Gate 3 refused to respawn a pane whose records still named a PTY, which is
right for a transport drop and wrong for a killed relay: after the relay is
SIGKILLed and comes back, the leaf map still holds `pty2:<dead-epoch>:1` while
the new relay answers that it has no such id. #18017's "replaces the pane only
when the host proves the session is gone" regressed on exactly that.

The gap was not the predicate, it was its inputs. `handlePtyReattachFailure`
already distinguishes the three reattach outcomes and only its not-found branch
publishes anything — a lost link and an identity mismatch send nothing. But it
published `pty:exit { code: -1 }`, and `-1` is the sentinel every reader
resolves to `stop_unverified`, so the one branch holding positive host evidence
of absence arrived looking exactly like loss of contact. The renderer had no
host answer at all, which the gate's own comment conceded.

The exit now carries `livenessVerdict: 'exited'` beside the unchanged `-1`, so
the code keeps meaning "no provable status" for every existing reader while the
verdict rides its own field. A store bridge records those ids in
`hostAttestedAbsentPtyIds` regardless of whether a pane is mounted to hear it —
during reconnect none is — and the gate stops counting a recorded id the host
has disowned. Settled when a PTY answers to that id again, because a redeployed
relay renumbers from pty-1.

This narrows #17963, which pinned the same exit as unverified on the grounds
that not-found cannot separate "verified the pid is dead" from "my session map
never had this id". Everything #17963 protects is untouched: `-1` still fails
isProvenProcessExit, so the tab is not closed, the pane's leaf binding is not
dropped on exit, and markUnverifiedPtyLoss still fires. Only the reconnect
respawn gate reads the new field, and only for an id whose sole channel — the
relay that answered — has disowned it, which no client can reach again under
any verdict. That is the reading ssh-pty-relay-absence-verdict.test.ts already
pins for the spawn path; the reconnect path now agrees with it.

Rejected: parsing the relay's mint epoch out of `pty2:<epoch>:<n>`. It needs the
current epoch on the wire (a capability-negotiated relay change), it has no
answer for legacy `pty-N` ids, and a relay that comes back with zero PTYs gives
the client no epoch to compare against. Rejected: clearing the leaf record
outright, because the remote workspace snapshot re-hydrates those ids after the
clear and the gate would refuse again.

* refactor(ssh): name the relay-disowned signal for disownership, not exit
2026-09-02 21:21:54 -07:00
NeilandNeil 76836b30ea fix(ssh): stop respawning an agent onto a PTY the relay just proved alive (#17951)
* fix(ssh): stop reporting live relay PTYs as expired sessions

A `pty.attach` reply carrying `sourceRecovery: restoreRequired` is the relay
answering for a PTY it just found in its pool and proved alive with
`isProcessAlive`; only the stale output delivery was retired. Main converted
that into `SSH_SESSION_EXPIRED`, which is the token every caller uses to retire
the pane binding and cold-restore the agent, so a transient reconnect started a
second `claude --resume` over a running one's transcript and left the previous
remote PTY detached — one more per reconnect until the host refused to fork.

Retry the attach once (the relay retires the stale delivery as it answers, so
the next attach opens a fresh one with full replay), then fail with a
restore-required verdict that makes no claim about absence. Callers already
route anything short of absence to the unverifiable pane-recovery path.

Also tighten the renderer's expiry verdict, which was a bare substring test: an
identity mismatch names a LIVE PTY owned by another pane and observes nothing
about this one, and main's own gate already refuses to respawn on it.

Refs #11006, #9034

* fix(lint): merge the duplicate pty-connect-limits import

* test(ssh): stop pinning the expiry token on a restoreRequired refusal

The refusal now carries SSH_PTY_SOURCE_RESTORE_REQUIRED, so the ratchet
asserts the discriminating token instead of the one it no longer shares.

---------

Co-authored-by: Neil <neil@example.com>
2026-09-02 21:05:18 -07:00
Neil 946627f2ce fix(runtime): route runtime filesystem commands by resolved execution host (#18325)
`ResolvedRuntimeFileTarget` carried `connectionId?: string` and no host id, so
`undefined` spelled three different answers at once — "runtime: host", "unresolved"
and "genuinely local". Its sole resolver read `store.getRepo(worktree.repoId)?.connectionId`
and never looked at `worktree.hostId`, which outranks every repo row, so one
arbitrarily chosen row decided the execution host for ~30 filesystem dispatches.
This is #18307's defect in the same file family; it was deliberately left out of
that PR rather than doubling an already-36-site diff.

The target now carries `executionHostId: ExecutionHostId` (never null, never
optional), resolved through `resolveWorktreeHostRouting` — the same adapter #18307
added — and dispatched through #18296's `resolveFilesystemRouteForHost`. Dispatch
sites call `requireRuntimeFileProvider`, where `null` means exactly one thing: the
host is `local` and the read happens here.

Four answers that used to collapse into one:

- `ssh:x` with a rival row on `ssh:y` — routes to x. Previously the first row won.
- `local` with a surviving `connectionId` — a row contradicting itself; no SSH
  connection is handed out.
- `runtime:<env>` — throws `ExecutionHostNotDispatchableError`. Its repo row's
  connection names a target in the *server's* namespace; reading it here reaches a
  same-named target on this client.
- rival rows disagreeing with no worktree host — `worktree_execution_host_unresolved`,
  matching the launch and Git paths rather than guessing a row.

Two further reads stop degrading. `assertRuntimeFileMutationExpectation` recomputed
the host from `connectionId`, so a client's host expectation could pass against a
host the workspace never named; it now compares the resolved host. And the
cross-workspace terminal tap coalesced `knownWorkspaceTarget?.connectionId ??
connectionId`, so a sibling workspace resolved as `local` inherited the origin
worktree's SSH target and statted a local path on the remote box; a non-optional
host id replaces rather than coalesces.

An unreachable SSH host still throws `SSH_FILESYSTEM_PROVIDER_UNAVAILABLE_MESSAGE`;
loss of contact is never evidence of locality (docs/reference/ssh-execution-boundary.md).
Quick-open listing and path search keep degrading to empty for an unreachable host —
that is a false negative, not a local answer — and now do so only for a host that
really is remote.

The whole `runtime-file-commands-*` family carries `@ts-nocheck` from a mechanical
class split, so removing the field could not raise the compile errors that made
#18307 safe. `runtime-file-command-target.ts` is deliberately checked, and a ratchet
test stands in for the errors the family cannot produce.

No wire change: `ResolvedRuntimeFileTarget` is main-process internal, and the SSH
watcher-release and grant keys are byte-identical to before.
2026-09-02 20:47:09 -07:00
Brennan BensonandMerge Sim 21210aad34 fix(native-chat): make structured Codex launches race-resistant (#18251)
* fix(native-chat): cancel close-racing structured launches

* fix(native-chat): make structured launches observable and recoverable

* fix(native-chat): reconcile merged session tab publications

* refactor(native-chat): unify host snapshot versioning

* refactor(native-chat): complete launches from host snapshots

* fix(native-chat): replay unknown launches by intent

* fix(native-chat): guard duplicate launches and bound sync recovery

* test(native-chat): type owner fixture

* test(native-chat): type owner fixture

* fix(native-chat): back off structured session resubscribe

* fix(native-chat): fence delayed local session snapshots

* fix(native-chat): retry initial session sync safely

* fix(native-chat): refresh before sync retry

* test(native-chat): cover folder sync cursor cleanup

* fix(native-chat): retry failed structured session subscriptions

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-02 20:23:11 -07:00
Neil d5750648c2 fix(runtime): route runtime Git by resolved execution host, not repo connectionId (#18307)
`RuntimeGitTarget` carried `connectionId?: string` and no host id, so `undefined`
spelled three different answers at once — "runtime: host", "unresolved", and
"genuinely local". Its sole resolver read `store.getRepo(worktree.repoId)?.connectionId`
and never looked at `worktree.hostId`, which outranks every repo row, so one
arbitrarily chosen row decided the execution host for 36 downstream dispatches.

The target now carries `executionHostId: ExecutionHostId` (never null, never
optional), resolved through the shared rule that landed with #17909/#17919 and
dispatched through the host-keyed routes from #18296. Dispatch sites call
`requireRuntimeGitProvider`, where `null` means exactly one thing: the host is
`local` and the command runs here as free functions.

Four answers that used to collapse into one:

- `ssh:x` with a rival row on `ssh:y` — routes to x. Previously the first row won,
  which is the reproduced cross-host leak.
- `local` with a surviving `connectionId` — a row contradicting itself; no SSH
  connection is handed out.
- `runtime:<env>` — throws `ExecutionHostNotDispatchableError`. Its repo row's
  connection names a target in the *server's* namespace; dialling it here reaches a
  same-named target on this client.
- rival rows disagreeing with no worktree host — `worktree_execution_host_unresolved`,
  matching the launch path rather than guessing a row.

An unreachable SSH host still throws `SSH_GIT_PROVIDER_UNAVAILABLE_MESSAGE`; loss of
contact is never evidence of locality (docs/reference/ssh-execution-boundary.md).

`resolveWorktreeLaunchHost` keeps its exact signature and now delegates to
`resolveWorktreeHostRouting`, the same resolution answering "which host is this on"
rather than "what may this client dial" — the git target needs the first question
because `local` and `runtime:` are two different non-SSH answers.

No wire change: `RuntimeGitTarget` is main-process internal, and the SSH and local
model-discovery host keys are byte-identical to before.

`RuntimeFileTarget` has the same defect in ~30 filesystem dispatches and is
deliberately left for a follow-up.
2026-09-02 19:26:07 -07:00
Neil 9cda5a9dc0 fix(worktrees): stop a resolved-worktree snapshot answering for repos it never saw (#18295)
* fix(worktrees): stop a resolved-worktree snapshot answering for repos it never saw

`listResolvedWorktrees` caches one fleet-wide snapshot for
RESOLVED_WORKTREE_CACHE_TTL_MS (1s) and reuses it on time alone. Nothing
invalidates it when a repo is registered, so for up to a second after a repo
row lands, every caller reads a snapshot computed before that repo existed --
and reads the gap as a verdict.

The visible failure is the SSH skill install. `resolveSkillSshTarget` resolves
a workspace-scope destination through that snapshot, so installing into a
worktree on a host connected moments earlier threw
`skill-install-workspace-not-found`: the client asserting a remote workspace is
absent on the strength of client-side bookkeeping that had never looked at the
host. That is the shape `docs/reference/ssh-execution-boundary.md` rules out --
absence from a client-side set is not evidence about the execution host. It
made `tests/e2e/ssh-skill-installation.spec.ts:108` fail 3 runs in 4 locally
and deterministically in the Docker SSH lane, where connect-then-install lands
inside the one-second window every time.

The snapshot now carries the repo-registration revision it was computed under
and is only reused while that revision still holds. The counter is the one
`bumpLocalWorktreeScanGeneration` already advances on every repo add, removal
and update, so the check is O(1) and cannot drift from the mutation sites.

* fix(worktrees): key the snapshot on repo mutations only, not on generation reads

Two things the headless-reattach lane surfaced.

The revision I keyed the snapshot on was `generationSequence`, which
`getLocalWorktreeScanGeneration` also advances when it mints a key for a repo
id nothing has scanned yet. That is a read, not a mutation, so a read path
could discard a snapshot that was still perfectly valid -- the mirror image of
the staleness this fixes, and a way to make a lookup fail that would otherwise
have succeeded. The counter now advances only where the scan generation is
actually bumped: repo add, removal, update, and scan-cache invalidation.

Separately, `pty-restore-record-seeding.test.ts` primed the cache by writing
its private `resolved` field with a literal spelling out `worktrees`,
`platformByRepoId` and `expiresAt`. That literal is a second copy of the
cache's freshness contract, so adding a field to the real entry left the fake
one failing the check: the primed snapshot was rejected, resolution fell
through to a real scan, and the headless fixture -- which has no git -- got
`selector_not_found`. It now primes through `getSnapshot` so the cache stamps
its own entry and the two cannot drift again.

The revision never moved during that test (0 before and after), so nothing was
being invalidated; the fake entry simply never satisfied the contract.
2026-09-02 18:13:08 -07:00
Neil 953df47fc4 feat(providers): dispatch git and filesystem providers by execution host (#18296)
`const c = repo.connectionId; c ? sshProvider(c) : local()` overloads `null`
to mean both "resolved: local" and "could not resolve", so every path that
cannot determine the host silently runs remote work on the client (#11163).
It also cannot express a `runtime:` host at all.

Add a host-keyed dispatch whose input is an `ExecutionHostId` — never null —
with `local`, `ssh` and `runtime` as three symmetric entries, and which throws
on an id that names no host instead of degrading to this machine. `ssh` carries
`provider: null` for "remote, currently unreachable", which is now a different
answer from "local" rather than the same one.

`runtime:` is a distinct entry rather than a provider because main does not
execute runtime hosts at all: they are forwarded over the environment transport,
and a runtime row's `connectionId` names a target in the server's namespace.
Dialing it from this client's SSH table would trade a silent-local bug for a
silent-wrong-host one.

First migrations, both to rows resolved via `getRepoExecutionHostId`:
- repo-worktrees: an `executionHostId: 'ssh:*'`-only row no longer lists,
  root-matches, or strict-lists against a same-named local path.
- workspace-space-repo-scan: same for the size scan, and `isRemote` no longer
  contradicts the `executionHostId` emitted beside it.
2026-09-02 18:04:48 -07:00
Neil d084a2a36a fix(ssh): decide remote-vs-local from the resolved execution host, not a raw field (#18294)
`repoIsRemote` read `repo.connectionId` directly. That is one of four spellings
of host ownership, so the predicate was wrong in both directions: a row carrying
only `executionHostId: 'ssh:<target>'` read as local and got the Linux-only
`orca-ide` rename it cannot resolve through the relay shim, while a row that
declares itself `local` with a stale `connectionId` read as remote and lost the
rename it needs on a Linux desktop.

The predicate now resolves the host first and asks "does an SSH target hold this
row's files" via `getRepoSshConnectionId`. That keeps a `runtime:` host's nested
SSH target remote (that machine reaches the files through its own relay shim)
while a runtime with no nested target - a full Orca install - stays local, as do
WSL and local.

Its call sites did not all want that question:

- The four launch-scope sites in main already hold the resolved PTY route on
  `TerminalWorkspaceLaunchScope.connectionId`. `scope.repo` is documented display
  metadata and can be a row from a different host than the worktree names, so
  they now read the route they will actually spawn on. A launch shape that
  disagrees with its own route is the bug, not a second predicate.
- `launchAgentInNewTab` picked its repo row with a host-blind
  `store.repos.find`, so a worktree that names its own host could be shaped by
  another host's row. It now resolves through `getConnectionIdFromState`, the
  same rule the file already used for transcript readability.
- `resolveAgentBackgroundLaunchHost` derived the route, the trust write and the
  launch shape from three reads of the raw field; one resolution now feeds all
  three.

Also converts the raw `repo.connectionId` agent-detection probe eight lines above
`buildWorktreeStartupForDraft`'s launch shape, which #17919 deferred precisely
because converting it alone would have left that file internally inconsistent.

Tests cover two distinct SSH hosts (a single-host fixture passes even when the
answer comes off the wrong row, which is how the `ssh:m4air` -> openclaw leak
survived review) and a `runtime:` host carrying a nested SSH target.
2026-09-02 17:46:01 -07:00
Neil 7c94d12190 fix(ssh): route four host-blind seams through the resolved execution host (#17919)
* fix(host-routing): resolve the execution host before reading a connection

Three issues in one defect class: a resolver reads one spelling of one
arbitrarily chosen row instead of resolving the worktree's execution host,
so something local answers a question about a remote.

returned that row's connectionId. With duplicate repo rows for one repo id
it could pair a runtime owner with a client-owned SSH connection. It now
resolves through the same ambiguity-aware index getRuntimeEnvironmentIdForWorktree
uses, prefers the repo row for the host the worktree names, and derives the
connection from the resolved host. Conflicting rows return `undefined`
(this module's documented "cannot determine the host"), never `null`.

`store.getRepo(worktree.repoId)?.connectionId ?? null`. `getRepo` is
host-blind and the same repo id can exist on local, SSH and runtime hosts,
so a remote worktree could spawn its PTY on the client with the remote cwd.
resolveWorktreeLaunchHost picks the row for the worktree's host and reads
the connection off that host; conflicting rows are unresolved, not local.

session-partition owner maps that contradict each other. Both now compute
through one shared function whose argument records the divergence. No
behaviour change on either side: converging needs a read-both migration,
since both partitions hold real data written by shipping builds.

* fix(host-routing): keep nested SSH connections resolvable under a runtime host

getRepoSshConnectionId read only the resolved execution host, so a repo row
owned by a runtime that reaches a nested SSH target (connectionId: ssh-*,
executionHostId: runtime:*) resolved to no connection — answering 'local' for
a remote worktree, the same defect #17909 fixed in the other direction.

* fix(host-routing): resolve both sides of the execution host through one rule

The renderer resolver leaked between two different SSH hosts: a worktree on
`ssh:m4air` whose only indexed repo row belonged to `openclaw` answered
'openclaw', because the host-scoped lookup missing fell through to an id-only
one. Main's resolver, in the same change, answered 'm4air' — two resolvers, one
right and one wrong, on identical input.

Both sides now adapt one shared rule (`worktree-execution-host-resolution.ts`):
the worktree's own host outranks every repo row, and a row on a different host
is never evidence about this one. The renderer's WeakMap index becomes the
memoizing adapter it always was; `resolveWorktreeLaunchHost` becomes main's
mapping of unresolved onto its throw.

Settles the rule the change previously answered two ways.
`getRepoSshConnectionId` and `getSshTargetIdForExecutionHost` disagreed for a
runtime host carrying a nested `connectionId`; they now compose, so the
execution host is the single authority. On a `runtime:*` row that field is a
paired HUB's private SSH target, spread through by `repoWithFetchedOwner` and
unaddressable from this client — the project-first successor of the row nulls it
for exactly that reason. That also fixes the `kind !== 'ssh'` fallback, which
fired for `local`: a row declaring itself local handed out an SSH connection.

* fix(ssh): resolve the execution host in the worktree scan and managed create

The worktree scan and createManagedWorktree both picked remote-vs-local from
repo.connectionId, so a row stamped only executionHostId: 'ssh:*' was scanned
and created on the client against a remote path. The folder branch returns
before the check, so its agent-trust write landed locally too.

Refs #11163

* fix(ssh): stop over-rejecting and refusing SSH hosts the process owns

runtimeRepoMatchesExecutionHost rejected an unstamped SSH repo against its own
ssh:<connectionId>, so repo-add/clone dedupe could register a second row for a
path the host already owns. assertHostIsSupported made the CLI/runtime RPC
refuse --host ssh:* while the same process's IPC handler routed it correctly;
setupExistingFolder now shares that registration. Clone still refuses, because
nothing in this process clones onto an SSH host.

Refs #11163

* test(ssh): retarget the SSH host-setup guard spec at the substitution it prevents

setupProjectExistingFolder now registers the remote path through the same
addRemoteRepoFromPath the desktop IPC uses, so it fails on the host's terms
(connection not registered) rather than a categorical refusal. The local
clone/probe side effects it exists to catch are still asserted absent.

Refs #11163

* fix(cli): require an absolute path when setting a project up on an SSH host

Routing --host ssh:* to the remote registration made relative paths newly
reachable there, and they were resolved against the client cwd — registering a
path that names the wrong machine.

Refs #11163

* fix(repos): read the SSH registry directly so the runtime stays Node-bootable

Routing runtime project setup through addRemoteRepoFromPath dragged ipc/ssh --
and its 25-module electron graph -- into the runtime bundle. ssh-target-registry
already exists for exactly this; ipc/ssh only re-exports it.

* fix(ssh): close the agent-launch and session-export host-blind twins

Three sites left on the legacy spelling, all the same shape as the ones this
branch already fixed:

- `launchAgentTerminal` did `getRepo(worktree.repoId)` then wrote agent trust
  with that row's `connectionId`. Host-blind, so a repo id carried by two SSH
  hosts wrote a remote path into the *client's* Codex/Cursor/Copilot config and
  the agent on the host never saw the trust. Every sibling call site already
  passes the resolved `workspace.connectionId`; this was the last that did not.
- `targetForWorktree` (workspace-session export) fell back to the same
  host-blind read, so a session could be published to a machine that never
  owned the worktree. Unresolvable ownership now exports to nobody.
- `addRemoteRepoFromPath` minted `connectionId`-only rows while being the
  routing path this branch adds, so it kept creating rows in exactly the
  spelling the branch works around. It now stamps
  `toSshExecutionHostId(connectionId)` at creation; `reassignSshTargetId`
  already migrates both spellings, so target rename stays correct.

Tests cover two *different* SSH hosts throughout — the case none of the earlier
duplicate-row tests had, all of which were local-vs-ssh or runtime-vs-ssh.
2026-09-02 16:59:29 -07:00
Neil c61ca56a9b fix(ssh): resolve the worktree's execution host instead of guessing from one repo row (#17909)
* fix(host-routing): resolve the execution host before reading a connection

Three issues in one defect class: a resolver reads one spelling of one
arbitrarily chosen row instead of resolving the worktree's execution host,
so something local answers a question about a remote.

returned that row's connectionId. With duplicate repo rows for one repo id
it could pair a runtime owner with a client-owned SSH connection. It now
resolves through the same ambiguity-aware index getRuntimeEnvironmentIdForWorktree
uses, prefers the repo row for the host the worktree names, and derives the
connection from the resolved host. Conflicting rows return `undefined`
(this module's documented "cannot determine the host"), never `null`.

`store.getRepo(worktree.repoId)?.connectionId ?? null`. `getRepo` is
host-blind and the same repo id can exist on local, SSH and runtime hosts,
so a remote worktree could spawn its PTY on the client with the remote cwd.
resolveWorktreeLaunchHost picks the row for the worktree's host and reads
the connection off that host; conflicting rows are unresolved, not local.

session-partition owner maps that contradict each other. Both now compute
through one shared function whose argument records the divergence. No
behaviour change on either side: converging needs a read-both migration,
since both partitions hold real data written by shipping builds.

* fix(host-routing): keep nested SSH connections resolvable under a runtime host

getRepoSshConnectionId read only the resolved execution host, so a repo row
owned by a runtime that reaches a nested SSH target (connectionId: ssh-*,
executionHostId: runtime:*) resolved to no connection — answering 'local' for
a remote worktree, the same defect #17909 fixed in the other direction.

* fix(host-routing): resolve both sides of the execution host through one rule

The renderer resolver leaked between two different SSH hosts: a worktree on
`ssh:m4air` whose only indexed repo row belonged to `openclaw` answered
'openclaw', because the host-scoped lookup missing fell through to an id-only
one. Main's resolver, in the same change, answered 'm4air' — two resolvers, one
right and one wrong, on identical input.

Both sides now adapt one shared rule (`worktree-execution-host-resolution.ts`):
the worktree's own host outranks every repo row, and a row on a different host
is never evidence about this one. The renderer's WeakMap index becomes the
memoizing adapter it always was; `resolveWorktreeLaunchHost` becomes main's
mapping of unresolved onto its throw.

Settles the rule the change previously answered two ways.
`getRepoSshConnectionId` and `getSshTargetIdForExecutionHost` disagreed for a
runtime host carrying a nested `connectionId`; they now compose, so the
execution host is the single authority. On a `runtime:*` row that field is a
paired HUB's private SSH target, spread through by `repoWithFetchedOwner` and
unaddressable from this client — the project-first successor of the row nulls it
for exactly that reason. That also fixes the `kind !== 'ssh'` fallback, which
fired for `local`: a row declaring itself local handed out an SSH connection.
2026-09-02 16:41:13 -07:00
Neil 4b2d9b5aac fix(path): stop seeded user bin dirs from outranking the inherited PATH (#18265)
`patchPackagedProcessPath` prepends every seeded directory, so `~/bin` and
`~/.local/bin` land ahead of the PATH a GUI-launched Electron inherited.
That does more than make a tool findable, which is what seeding is for --
it re-ranks binaries the user already has, and those two directories are
user-writable and can hold a wrapper for any system tool.

On the #18234 reporter's box `~/.local/bin/gh` wraps `mise x gh -- gh`.
Seeded ahead of /usr/bin we ran the wrapper where their own shell ran the
real binary, and the wrapper's inner bare `gh` resolved back to itself.
Measured in an Ubuntu 24.04 container: with their shell's ordering the
chain exits in 22ms; with ours it never terminates and creates ~1,500
processes/second.

Seed order now follows the rule the WSL twin already documents in
posix-version-manager-bin-dirs.ts -- append, never prepend, because a
login PATH that did resolve is authoritative. Version-manager shim dirs
keep leading, since an nvm/mise/asdf user's runtime must still beat a
system install; the generic user bin dirs move behind the inherited PATH.
`getVersionManagerBinPaths` carries `~/bin` and `~/.local/bin` too (bun
and pnpm install there), so they are filtered out of the leading group by
name rather than by which list produced them.
2026-09-02 16:34:44 -07:00
Jinjing b8b7a6be9d fix(activity): persist the agents unread filter and grouping (#18255)
* fix(activity): persist the agents unread filter and grouping

The Agents view's "Show unread threads only" toggle and Group-by select were
plain component state in the sidebar and the Activity page, so both reset on
every mount — including app restart — while their neighbours in the same
toolbar (compact rows, show child agents) survived via the persisted UI store.

Promote both to `agentsReadFilter` / `agentsGroupBy` persisted UI preferences,
wired through the same seams as `agentsCompactMode`: shared type, default,
strict client RPC schema, pairing-local field census, web read pin, store
contract/actions, and hydration normalizers that reject unknown values. Both
consumers now read the store, so the sidebar and the Activity page share one
filter the way they already share compact mode.

* refactor: centralize thread filter value domains

Establish filter and groupby value domains as the single source of truth, with types derived from them to prevent drift between valid values and their normalizers. Extract common validation logic into a shared isMember helper to keep the two normalization functions in sync.

* refactor: centralize thread filter value domains

Consolidate filter value definitions in agents-view-thread-filters
and use them in Zod schema validation to ensure consistent,
persistent serialization of filter state.
2026-09-02 16:04:11 -07:00
Neil 510305e574 fix(relay): signal capacity loss instead of dropping, hanging, or truncating (#17870)
Three failures with one shape: a payload past a fixed capacity was met with
silence, with a wait that never ends, or with a prefix presented as a whole.

**The workspace snapshot was silently dropped.** `workspace.changed` carries the
tab/session list, and a snapshot past the producer frame capacity (12288 B on a
Node <=21 remote) was dropped with only a relay stderr line, so the client kept a
stale list forever. The relay now publishes per client and, for a client whose
sink refused the frame, sends a compact `workspace.stale` marker on the control
lane; the client re-reads through `workspace.get`, whose lane is budgeted in
megabytes rather than in one producer frame. A new JSON-RPC notification rather
than a new field on `workspace.changed`: `normalizeSnapshot(undefined, ns)` yields
revision 0 and an empty session, so a Rule-1 field would make an old client
replace its tab list with nothing — worse than the drop. An old client ignores the
unknown method and is exactly where it is today. The marker retention/retry
machinery is extracted from the `fs.changed` overflow path and shared by both.

**The Windows upload hung, and the fix for it could truncate.** `#16432` was
attributed to `[Console]::In.ReadToEnd()` materializing the base64 bundle. That is
not what the reporter measured: he also measured
`new IO.StreamReader([Console]::OpenStandardInput())` — an incremental reader —
hanging at 1 MB. The limit is in the stdin the host hands PowerShell over a
non-pty ssh exec, not in the string the script builds.

- `uploadFileViaSystemSsh` — the user file-import path — was piping a whole file
  into one Windows stdin, unchunked and untimed. That is the path large files
  take; it now chunks into 32 KB writes and bounds each wait.
- The Windows directory upload reuses that single-file path rather than repeating
  a weaker copy of chunk-read + write-buffer; the `ino`/`dev` TOCTOU verification
  comes with it.
- A Windows write needing more than one exec lands on a `.orca-partial` staging
  path and is published by rename, so a failed chunk cannot leave a truncated
  artifact under the real name. `exclusive` is enforced once at the rename, not on
  the first chunk, where a retry met its own leftovers.
- The mkdir batch reads stdin through the stream reader the reporter measured
  surviving 50 KB, not `[Console]::In`, which he measured wedging at that size.
- `waitForChannelClose` takes an optional bound. A wedged PowerShell stays alive
  at idle CPU and never closes, so without one the promise is simply never
  settled and the caller waits forever with no error to show.

**Quick Open showed a prefix as the whole workspace.** The mechanism "a full page
means there is more" only works if the caller named the cap, and the failing UI
named none — it hardcoded `truncated: false`. Quick Open now names
`QUICK_OPEN_LISTING_MAX_RESULTS` on both the Electron IPC hop and the runtime-RPC
hop (the field #17954 added to `files.listAll`), and reads a full page as
truncation. The local hop honours the cap too, which it previously ignored.

Rebase note on `fs.listFiles`: an earlier revision of this work also clamped the
host unconditionally, and #17934 escalated an uncapped request to an explicit
error. #17954 has since landed and made an oversized reply streamable, which
removes the premise — the host no longer has to choose between a prefix and a
refusal, so it returns the whole listing when no limit is named and only clamps a
limit it was given. Keeping either would have regressed #17954 and hard-failed
three in-tree callers that deliberately pass no options
(`runtime-file-commands-search-runtime-files.ts:81`,
`filesystem-read-handlers.ts:125`, `runtime-file-commands-constructor.ts:41`).
2026-09-02 15:42:08 -07:00
Brennan BensonandMerge Sim 7f8eb90ac3 Align worktree host labels across desktop and mobile (#18237)
* refactor: align worktree host labels across clients

* fix(mobile): expose safe host display labels

* fix(mobile): preserve legacy mixed-host labels

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-02 15:32:06 -07:00
Neil 278f9ee876 fix(ssh): answer every MFA stage, stop dialling an unclaimed alias, and say where a clone failed (#17946)
* fix(ssh): answer every MFA stage, not just the first

ssh2 walks one flat auth-method list exactly once, so keyboard-interactive
could only ever be offered a single time. A host running
`AuthenticationMethods keyboard-interactive,keyboard-interactive` (or any
ladder ending in a second challenge) partial-succeeds the first stage and
then finds the list exhausted, which the user sees as "All configured
authentication methods failed" — the reports in #8622 and #16820.

Orca's own auth handler now runs for every target instead of only multi-key
ones, and rebuilds its queue on each SSH_MSG_USERAUTH_FAILURE that carries
partial success, narrowed to the methods the host still offers. Narrowing
also stops keys being re-offered after the host has moved past publickey,
which is what exhausts MaxAuthTries before the challenge is ever shown.

Covered by a real ssh2 server fixture that stages partial success.

* fix(git): say where a failing clone ran and why nothing could prompt

Clones go through nonInteractiveGitEnv, so `ssh` runs with BatchMode=yes and an
emptied SSH_ASKPASS. On a remote or paired-runtime clone that produces
`fatal: Could not read from remote repository.` while the same `git clone`
typed by hand on that box succeeds — the divergence in #14533. Nothing in the
message said the clone ran on the other machine, under its keys, with the
prompt deliberately disabled.

getGitCloneFailureMessage now appends that fact, and names the two recognisable
shapes: a publickey refusal (load the key into an agent there) and a host-key
failure (record the key in that machine's known_hosts). Unrecognised SSH
failures still get the where-it-ran note; non-SSH failures are untouched.

One builder, so the SSH-target relay path and the runtime path both get it.

* fix(ssh): stop dialling a bare alias no ssh_config block claims

A wildcard `Host *` block supplies ProxyCommand/ProxyJump for every alias, so
shouldUseSystemSshTransport picks the system transport for an alias whose own
Host block was renamed or deleted, and buildSshArgs then dials that alias
verbatim: no -l, no -p, no Hostname. Orca connects as the wildcard's user to
the wildcard's host and discards the endpoint it stored (#11746).

The signal #11746 assumed (hostBlockMatch, from the still-open #11707) does not
exist, and `ssh -G` cannot supply it — it prints the merged config and answers
for unknown aliases too. The config file is the only source of truth, so:

- parseSshConfigAliasClaims retains raw Host patterns and flags Match blocks,
  which parseSshConfig discards because it mints importable targets.
- sshConfigMayClaimAlias is sound in the negative direction only: an unreadable
  file, any Match block, or any non-catch-all pattern that might match all
  answer "claimed", so absence of evidence is never read as evidence of
  absence. Only a proven-unclaimed alias licenses an override.
- buildSshArgs then states Hostname/Port/User, and only those: the wildcard is
  still the route, and -o Hostname does not change block selection, so the
  proxy keeps applying and %h expands to the host we mean.

The verdict is injected rather than read inside buildSshArgs, so an arg builder
does not answer differently per machine. Default is today's behaviour.

Scoped to the system-SSH transport and the connection's own command/transport
path. Port-forward processes and the ssh2 transport (#11707) are unchanged.

* fix(ssh): read a negated Host group as uncertainty, and gate clone SSH guidance

`Host * !prod` applies to every alias but `prod`, yet skipping both the catch-all
and the `!` pattern answered "unclaimed" for `stage` — which licences overriding
Hostname/Port/User against a block the user wrote. Any negation now makes the
whole group uncertain; the function is only sound in the negative direction.

Also require an ssh(1) diagnostic beside "could not read from remote repository"
before appending the SSH clone note: git prints that same line for the HTTP
remote helper, where advice about keys and agents is simply wrong.

* fix(i18n): restore the activity-options key the rebase dropped

* fix(i18n): union en.json with main so the rebase cannot drop keys
2026-09-02 15:14:53 -07:00
Neil 39330c5aca fix(relay): retire PTYs the host proves are gone, and stop two per-poll scan storms (#17832)
* fix(relay): stop three CPU growth terms in a long-running remote session

pty.resize gated only on `managed.disposed`, which is bookkeeping rather than
liveness. A shell that exits without node-pty's `onExit` leaves an undisposed
entry holding a closed master fd, and UnixTerminal.resize has no fd guard, so
the ioctl threw `ioctl(2) failed, EBADF` into the dispatcher's generic
parse-error catch. Nothing retired the entry, so it stayed advertised and kept
activePtyCount above zero -- which is what stops a relay with an unlimited
grace from reaching its idle-no-ptys exit (#12423). Probe liveness with the
same helper attach/listProcesses use, retire a provably dead pid, and contain
an ioctl failure over a live-or-unverifiable process.

processHasChildren forked `pgrep -P` per pane per inspection poll, uncached.
procps-ng opens six procfs files per process to resolve one ppid, so each call
cost O(host process count). Answer from the TTL-cached `ps` table the same RPC
already captured for the foreground lookup (#13537).

The remote AI Vault scanner had no parse cache at all, so every forced rescan
re-read and re-parsed the whole transcript corpus, including files untouched
for a month. Give it the mtime+size keyed memo the local scanner has (#13753).

* fix(pty): invalidate the descriptor when node-pty gives up the handle (#17930)

Carried forward from PR #17930, which merged into this branch. Rebased onto
current main; main's newer node-pty-fd-leak test is kept as-is.

* fix(ai-vault): refresh codex titles on the remote parse-cache reuse path

The remote cache keys on the transcript's (mtime, size, host), but codex
titles live in $CODEX_HOME/session_index.jsonl and are written after the
rollout — so a cache hit froze the fallback title forever. Mirrors the
local scanner's existing reuse-path refresh via a shared core.

* fix(relay): publish the exit a reap performs, and rescan for close decisions

Two review findings on the CPU work.

reapExitedPty told only the relay-internal exit listener, so a retirement left
the client's pane mounted against a session the relay had already forgotten --
the next attach answered `PTY "<id>" not found` with nothing before it to
explain why. Pre-existing on three probe paths; resize made it user-triggered.
Publish the same pending-exit the natural onExit path publishes, carrying -1
("gone, status unrecoverable"), and skip it when onExit already reported the
real code.

processHasChildren now answers from a 500ms TTL-cached table. That is right for
pty.inspectProcess, which every tracked pane polls, but pty.hasChildProcesses
gates the window-close confirmation and workspace cleanup's idle evidence --
one destructive decision per answer, where a child started inside the window
would be killed unasked. Give that RPC a fresh scan; pgrep used to.

* fix(relay): publish a reap's exit only on proven-exited evidence

The publication is a verdict the client acts on by retiring the pane, so it
must not be reachable from the disposed-record sweep, which retires off our own
bookkeeping rather than the host's process table. Only ESRCH earns it.

* fix(i18n): restore the activity-options key the rebase dropped

* fix(i18n): union en.json with main so the rebase cannot drop keys
2026-09-02 15:14:44 -07:00
Neil 7458c39181 fix(ssh): declare a wedged relay link lost, and stop reading silence as a verdict (#17817)
* fix(ssh): declare a wedged relay link lost instead of suppressing the dead-link check

* fix(ssh): make the Windows deps probe exit 0 on a real load failure, like its POSIX twin

* fix(relay): reap a client that has stopped answering instead of holding its leases forever

* test(relay): feed the primary before asserting the reaper exemption holds

* fix(ssh): keep a lost link's verdict unverifiable instead of reporting absence

* refactor(ssh): read the exec timeout from its typed code, not the message text

* fix(relay): bound a client that clears the handshake and then never frames anything

* fix(i18n): restore the activity-options key the rebase dropped

* fix(i18n): union en.json with main so the rebase cannot drop keys
2026-09-02 15:14:33 -07:00
Neil 07e1f953a5 fix(ssh): log an unanswered native-deps probe instead of launching silently (#18000)
* fix(ssh): log an unanswered native-deps probe instead of launching silently

The wrongful rebuild used to be the only visible symptom of a dropped exec
channel; #17979 removed it, so a real transport failure now leaves no trace.
Matches the install-path sibling, whose callers log the same class of failure.

* fix(i18n): restore the activity-options key the rebase dropped

* fix(i18n): union en.json with main so the rebase cannot drop keys
2026-09-02 15:14:25 -07:00
Neil 7104056984 fix(watcher): route relay watch-root capacity refusals off the fast ladder (#17950)
* fix(ssh): stop two unrecoverable relay refusal loops

A relay refusal that is a pure function of state the client cannot change was
being retried forever, on two different paths.

- pty.openClient: a superseded owner proof is refuted evidence, not a transient
  fault. The client kept re-presenting the identical proof, so every reconnect
  reproduced the same refusal until the relay was redeployed (#12895, #12931).
  It is now dropped exactly as a stale lease already is, and the claim re-asked
  without it.
- fs.watch: the relay's watch-root capacity refusal was classified 'unavailable'
  and retried at 1 Hz per root for 60s, re-armed indefinitely. A folder
  workspace with more repos than the cap turns that into a permanent install
  storm scaled by the excess root count (#11196). It is now its own 'capacity'
  result that goes straight to the existing dormant backoff, mirroring what the
  local watcher path already does.

* fix(watcher): route relay watch-root capacity refusals off the fast ladder

A full watch-root cap is a decision, not a fault, so a 1 Hz reinstall per refused
root only bills the relay the load that keeps the cap busy (#11196). Capacity
refusals now go straight to the dormant backoff.

The relay side no longer refuses on a slot it is about to hand back: an over-cap
caused by roots still unsubscribing waits once on the teardowns settling — the
release event, mirroring WatcherSupervisorCapacityWait — before it answers. A
parked waiter is excluded from the accounting so it cannot take a slot from the
root already reclaiming one.

Drops the SSH owner-recovery half of this branch. Its premise — that a -32043
SUPERSEDED refusal is permanent — is false: the refusal fires only while the
incumbent is 'active', and assertPtyConsumerOwnerRecovery explicitly admits the
identical lower-generation proof once the incumbent flips to 'disconnected'
(relay-pty-consumer-owner-displacement.test.ts proves it). The remedy could not
work either: the proofless re-ask routes into refuseHeldPtyConsumerOwner, which
is declared `: never` and, with sameClient true by construction, always throws.
It would have traded one refusal loop for another, minus the checkpoints and
minus the proof that resumes the claim once the relay reaps the incumbent.

* fix(i18n): restore the activity-options key the rebase dropped

* fix(i18n): union en.json with main so the rebase cannot drop keys
2026-09-02 15:14:21 -07:00
Neil 7c6c8ef85e fix(ssh): stop a late SFTP stream error crashing main, and keep the relay socket inside sun_path (#17862)
* fix(ssh): stop SFTP stream errors crashing main and bound the relay socket path

inside the protocol parser. Every transfer removed its listener on settle, so a
STATUS reply that arrived late - the normal case behind a jump host that chroots
its SFTP subsystem - threw synchronously out of Socket.emit('data') and killed
the main process. Keep one durable listener per stream, and report a sandboxed
SFTP namespace with an actionable message instead of a bare 'file does not
exist'.

104 macOS) and bind failed with a bare 'listen EINVAL'. Fall back to a per-uid
base whose length does not depend on $HOME, keeping the hashed socket name
intact.

* fix(ssh): validate the short socket dir before mutating it

* fix(ssh): keep the SFTP session guarded, scope the relocated socket, narrow the chroot verdict

Three review findings.

The CLI-launcher install ran writeStringViaSftp in a loop over a bare conn.sftp().
That helper removes its own session 'error' listener at each settle, so between
files and after the last one the emitter carried none -- and ssh2 raises a late
STATUS reply synchronously out of Protocol.parse, which is the uncaught exception
that kills main (#15479). The inline loop it replaced leaked one listener per file
and covered this by accident. Extract writeStringsViaSftp, which owns the session
latch, and share that latch with runSftpFallbackTransfer.

SSH_FX_PERMISSION_DENIED is a mode/ownership refusal on a path the subsystem can
see, not evidence of a chroot; sftp-namespace-resolution already treats only
NO_SUCH_FILE as conclusive. Narrow the predicate to code 2 so a read-only home
stops being reported as a bastion misconfiguration.

The relocated socket had no version dimension. relaySocketNameForInstanceId hashes
the target, not the build, and under $HOME the enclosing relay-<fullVersion> dir
supplied the rest -- so the short form made the path stable across updates. The
next build would bind the path the previous relay still holds, the handshake would
mismatch, and a relay holding live work would raise RelayEndpointHeldError with no
way through. Add a hashed version segment under the short base, mirroring the
relay-*/<sock> shape so one pattern serves both, and teach the superseded sweep and
force-stop about that base. The relocated tree now also gets reclaimed: nothing
else walks it.

* fix(i18n): restore the activity-options key the rebase dropped

* fix(i18n): union en.json with main so the rebase cannot drop keys
2026-09-02 15:14:17 -07:00
Neil 31007c0d86 fix(ssh): reclaim relay PTYs the client has provably lost, on host attestation only (#17831)
* fix(ssh): reclaim relay PTYs the host attests this client orphaned (#9819)

Orca could lose track of terminals running on an SSH relay until the
50-slot cap refused to open any more. This reclaims them, and the whole
design is built around the fact that getting it wrong destroys a user's
running process on their remote machine: the failure mode is leak, never
kill.

A stop requires all nine of:

1. the relay published an `ownerClientInstanceId` read from the live
   authenticated consumer grant of the connection that requested the
   spawn — never from a spawn parameter, since an echoed claim is no
   evidence; absent means skip
2. that id equals this client's persisted consumer identity
3. this connection holds the negotiated `session-owner` grant
4. `paneBound === true`, host-published
5. no `agentSessionOwners` — the host still advertises it as adoptable
6. `hostAgeMs >= 30s`, measured on the host's clock
7. this client has no route: not reattached, no lease outside
   terminated/expired, no pending kill, and no `expired` lease either —
   an expired lease is the record of a process deliberately left
   running, never a licence to kill it
8. every stop is fenced on the incarnation the same listing published,
   and on the owner identity, both re-checked by the host
9. a pass wanting to stop more than 8 refuses entirely

Absence from a client-side set is `unverifiable` by construction
(docs/reference/ssh-execution-boundary.md): a second machine attaches to
the same relay and displaces the session owner, and its live agents are
missing from this client's store for exactly the reason a genuine orphan
is. So the host has to attest ownership, and the host has to attest that
nothing is running.

That second attestation is measured over the pane's whole tty, not its
foreground process group. `tpgid == pgid` is foreground-only: on a real
`bash -i` on a real pty, a shell holding `sleep 300 &` and a shell
holding a Ctrl-Z'd job both read `pgid == tpgid`, `Ss+` — byte-identical
to an idle prompt, with only the job's own row differing. A
foreground-only gate therefore attests `pnpm build &` and a suspended
editor as idle, and the stop that follows SIGKILLs every process group
on the tty. `shellOwnsEveryTtyProcessGroup` is measured over that same
set of groups, so the evidence and the kill describe the same thing. No
new probe: `tpgid` already identifies the terminal, because a process
group belongs to one session and a session to at most one controlling
terminal.

The freshness field is real rather than decorative. `capturedAgeMs` is
stamped from when the capture was taken, deliberately as an upper bound
since the process table is TTL-shared, and the sweep refuses an
observation older than its own pass budget, counting its own elapsed
time since the listing arrived. Stale evidence degrades to "do not
sweep", never to "sweep". The display consumer of the same measurement
keeps no age budget, as a stated decision: a stale pane title costs a
redraw and self-corrects.

`pty.shutdown` is authorized on the host that owns the process.
`pty.spawn` and `pty.attach` both take a request context and check it;
the one irreversible call took none, so the rule above lived entirely on
the client that decided to make the call. It gains an optional
`expectedOwnerClientInstanceId` and refuses unless the connection still
authenticates as that identity AND this host recorded it at spawn.

Finally, a reattach refusal now says whether it observed the process.
Three refusals carry the same `SSH_SESSION_EXPIRED` text and only one is
absence; `restoreRequired` means the PTY is live and only its source
stream is not. Testing that text with `.includes()` expired the lease
and deleted ownership for a running process, erasing this client's only
record of it — and a PTY with no record is one the sweep may stop.

Wire compatibility: four new optional fields and one new optional param
on existing methods, no new method and no new stream opcode (Rule 1, and
Rule 2 does not apply). Rule 1's caveat is discharged explicitly — no
reader requires any of them, each absence is a named skip reason, and an
ordinary pane teardown must omit the owner fence because a revived PTY
carries no attested owner at all. New client plus old relay stops zero
PTYs; old client plus new relay never reads the fields. Windows relay
hosts publish no evidence and therefore never sweep.

Verified by joining the real publisher to the real client reader over
`ps` captured verbatim from a Linux container, and by driving a real
group-for-group SIGKILL against a real pty: backgrounded and suspended
jobs survive by pid, and an idle shell is still reclaimed, so the
narrowed predicate is not a silent no-op.

Squashed deliberately. The sweep is unsafe at every intermediate commit
of its own history — before the foreground gate it reaps a hand-launched
`claude`, and with a foreground-only gate it reaps a backgrounded build
— so this ships as one commit with no bisectable state that kills live
work.

Refs #9819. Folds in #17939.

* fix(i18n): restore the activity-options key the rebase dropped

* fix(i18n): union en.json with main so the rebase cannot drop keys
2026-09-02 15:14:14 -07:00
Neil fb48a9771b fix(gh): reap the whole gh/glab process tree at the deadline on POSIX (#18258)
`gh` and `glab` on PATH are routinely shims — mise, asdf, volta, or a
hand-written wrapper — so a timed-out invocation has a chain to stop, not
one process. `execFileCapture`'s POSIX kill path signals only the direct
child; the descendants are orphaned to init and keep running. #18234 is
exactly that shape: `bash ~/.local/bin/gh` -> `mise x gh` -> `gh`, where
the reporter found the tail reparented to `systemd --user` and still at
100% CPU nearly two hours later. The 15s deadline #18239 added bounds
Orca's semaphore slot and its promise; it does not bound the CPU burn.

Route both CLIs through `execFileCaptureToTermination`, the primitive
git's barrier path already uses: POSIX children spawn `detached`, the
deadline signals `-pgid` and escalates to SIGKILL, and the promise waits
for verified termination. Windows behaviour is unchanged (`taskkill /t`
either way).

Switching primitives also swapped execFile's hard maxBuffer failure for
`runProcess`'s silent clipping, which would have turned an oversized gh
response into a shorter valid-looking one. `ProcessResult` now reports
truncation and the capture rejects on it, restoring the old contract and
closing the same latent gap on git's barrier path.
2026-09-02 15:12:21 -07:00
Neil 02a417c04b perf(renderer): stop six timers from ticking behind a hidden window (#18134)
* perf(renderer): stop six timers from ticking behind a hidden window

IntensiveWakeUpThrottling is disabled in this app, so a renderer interval
really does fire at full rate with the window hidden. Six of them had
nothing to observe them:

- NativeChatWorkingStatus ran a 1s interval + setState per in-flight turn
  purely to advance an elapsed-seconds counter. Deleted the effect and
  derived elapsed during render from the shared, visibility-gated
  useNow(1_000) clock, so N turns collapse onto one tick.
- The chromium-error fallback poll (250ms) kept probing a stuck-loading
  guest to write a loadError nobody could see.
- The contextual-tour full-pass interval (500ms) woke twice a second to
  queue a rAF a hidden window never paints.
- Three feature-wall animation timers (3600/2400/2400ms) kept committing
  React renders for animations nobody was watching.
- The landing preflight poll (30s) kept forcing IPC refreshes.

All five gated timers reuse installWindowVisibilityInterval. Each either
resumes where it left off (animations) or re-derives from durable state on
the becoming-visible run, so hiding and re-showing is observationally
identical to never hiding.

* test(git): stop two empty commits in the divergence fixture from hashing alike

`counts drift in both directions` builds 100 empty commits, resets to the fork
point, then adds one more — expecting 100 ahead + 1 behind to clear the cap of
100. An empty commit's hash covers only parent, tree, message and a
one-second-granularity timestamp, and every commit in the fixture reuses
`commit ${index}` starting from 0. On a runner fast enough to finish the whole
build inside one wall-clock second (CI: 1059ms for the case, ~7ms per commit),
the post-reset `commit 0` hashed identically to the first `commit 0` of the
chain, so Git handed back that same object and left the branch 99/0 apart
instead of 100/1 — `within`, not `exceeded`.

Numbering the empty commits across calls makes the fixture build the 101
distinct commits it already claimed to. Reproduced deterministically by pinning
GIT_AUTHOR_DATE/GIT_COMMITTER_DATE, which forces the timestamp collision the
fast runner hits by chance: fails with the exact CI assertion before, passes
after.
2026-09-02 14:35:48 -07:00
Neil 0886db2b90 refactor(process-table): extract the correlation indexes into their own module (#18246)
`src/shared/process-table-snapshot.ts` is 308 code lines against the 300
cap for `**/*.ts`, so `static analysis` is red on `main` and every open PR
inherits it.

Neither PR that grew the file crossed the cap alone. #18151 took it to 427
raw lines; #18166 added ~35 more. #18166's branch predated #18151, so the
head CI linted was 428 raw lines and passed, while the squash onto main is
463 -> 308 code lines. The gate lints the PR head, not the merge result, so
nothing linted the sum until it was on main.

Pure move, no behaviour change: the generic index machinery
(ProcessIdentityRow, ProcessTableIndexOf, buildProcessTableIndex,
collectDescendantsFromIndex, lookupProcessTableIndex, getProcessTableIndex
and its WeakMap) moves to process-table-index.ts. `ProcessTableIndex` and
`scoreForegroundCandidateRow` stay behind because they need
`ProcessTableRow`, which keeps the new module free of any import back and
so introduces no cycle.
2026-09-02 14:01:24 -07:00
Brennan BensonandMerge Sim 623d58e386 fix(native-chat): show pasted images while they save, and make them previewable (#18118)
* fix(native-chat): show pasted images while they save, and make them previewable

Pasting an image into the native chat composer showed nothing until the
clipboard image finished being written to disk, and the resulting chip could
never render the image at all.

Preview was blocked by path authorization, not by rendering. Clipboard pastes
are written to the OS temp dir, which sits outside every allowed root, so the
composer's own `fs:readFile` of the file Orca had just written was denied.
`saveClipboardImageBufferAsTempFile` now authorizes the path it writes, the
same way other Orca-produced external files are handled.

The delay is the macOS paste route: Cmd+V is intercepted in main and delivered
through the app-menu paste channel, which has no clipboard blob in hand, so the
composer only learned an image existed after the save round-trip. A new
`clipboard:readImageThumbnail` probe reads the clipboard in memory and returns a
downscaled preview; it runs alongside the save rather than before it, so text
paste gains no latency. The DOM-paste route needs no probe — it mints a blob URL
from the clipboard file on the same tick.

Attachments now carry `pending` and `previewUrl`: the chip appears immediately
with the real image dimmed under a spinner, then settles in place on the saved
path. Send is blocked while anything is pending, because a pending chip has no
agent-readable path yet. Pending chips are kept out of the pane attachment cache
so a mid-save unmount cannot strand one, and blob previews are revoked on
remove/clear. SSH pastes now carry their connectionId onto the chip so remote
previews read over SFTP.

Verified in a real Codex native chat under an isolated dev instance: the chip
appears in 42-61ms with a spinner, settles at ~141ms, three rapid pastes produce
three independent chips with Send disabled throughout, and the lightbox opens the
full 5120x2880 image read from disk. Ablation confirms the authorization fix:
the written path reads back, an unauthorized sibling in the same temp dir does
not.

Claude-Session: https://claude.ai/code/session_01NnEfY8NpfFtVnboLKnmgdW

* fix(native-chat): avoid stale image attachments and preview cache growth

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-02 14:00:17 -07:00
Jinjing c2fce80289 Fix agent dashboard setting configure (#18245)
* Make agents activity always-on; toggle via bell icon

- Remove optional showAgentsSidebar setting
- Replace sidebar view-toggle with bell-button for activity access
- Agents activity now always accessible in sidebar
- Preserve migration flag for introduction to existing users
- Remove visibility inference utilities

* Simplify sidebar when agents view active: hide workspace options, add to

- Hide workspace options menu and add project button when agents view is
  active, reducing UI clutter in that mode
- Add tooltip to the activity bell button for better discoverability
- Localize sidebar search field text
- Move search and filter toggles to local state in SidebarAgentsList,
  removing unused callbacks from thread list components
- Manage search input focus properly when opening
2026-09-02 13:43:06 -07:00
Jinjing e3de6b2ce8 Add automation runs dashboard with pagination and filtering (#18226)
* Add automation runs dashboard with pagination and filtering

Adds a new Runs view in the Automations page that lets users browse all runs across automations with status/host filtering, search, and pagination support. Includes virtualized table rendering for efficient handling of large run histories and summary cards showing 24h/7d success/failure counts.

* Fix missing dependencies in useCallback hooks and imports

Missing dependencies in useCallback can cause stale closure bugs. This
adds missing state setters to dependency arrays and consolidates type
imports for consistency.

* Use keyset pagination for stable automation runs pages

Pagination now uses createdAt:id boundaries instead of offsets, so new
runs arriving between pages don't shift the window. Maintains backwards
compatibility with legacy offset cursors.

Move pagination to shared module, fix outcome counting for future-dated
runs, and improve hook state tracking on authority re-pairing or target
changes.

* Extract automation run details to top-level page view

Moves run display from detail pane to dedicated page, establishing
three-level navigation (Automations → Runs → Run Details) and simplifying
the detail pane component.

* Fix pagination stability when automation runs share createdAt

- Define a stable total order with createdAt and id tiebreaker to prevent runs tied on createdAt from being dropped when the boundary run is pruned between page requests
- Retain cursor on failed pagination so pages remain retryable
- Update ownerNotice type to AutomationActionNotice

* Extract automations list panel and worktree map logic

Split AutomationsPageSurface into smaller, focused modules for better maintainability and reusability. Move list panel UI rendering to AutomationsPageListPanel component and worktree map selection logic to a standalone utility function.

* Add i18n strings for automation runs dashboard

Adds localized strings for the automation runs dashboard view, including search, filtering by host and status, run counts for 24h/7d windows, and empty state messaging across all supported languages.

* fix missing translation

* fix missing translation
2026-09-02 13:42:14 -07:00
Jinwoo Hong a0de2fde0b fix(terminal): confirm an unrecognized foreground before downgrading agent prompts (#18238) 2026-09-02 13:39:56 -07:00
Neil 1910ab9c9c fix(github): bound and coalesce the Orca star check so gh children cannot pile up (#18239)
`checkOrcaStarred`, `starOrca` and `getAuthenticatedViewer` were the only gh
call sites that reached for the legacy `execFileAsync` instead of
`ghExecFileAsync`, so they ran with no deadline, no process-tree kill and no
coalescing. A `gh` that never exits therefore ran forever and never released
its slot in the 4-wide GitHub semaphore in gh-utils.

Route all three through `ghExecFileAsync`, coalesce concurrent star checks onto
one child, and hoist the Landing star-state effect out of the conditionally
rendered footer so a repo-catalog rewrite no longer remounts it and re-forks gh.

Adds a ratchet test asserting no file outside the command runner names `gh` as
a spawned program.

Fixes #18234
2026-09-02 13:29:12 -07:00
Neil 6b1cbe54a1 fix(process-table): fail a short ps capture loudly, and stop a resume spending 49 of them (#18166)
The POSIX process-table capture ran `execFile('ps', ...)` with no `maxBuffer`,
inheriting Node's 1MB default. Measured at 1,460 processes the capture is 326KB
with a 5,116-char longest row — ~3x headroom, which a busy host clears.

Two separate defects follow, fixed here:

1. `parseProcessTableRows` drops unparseable lines, so any short capture reads
   as a COMPLETE table whose missing processes simply are not running. Verified:
   a capture cut at 4KB parses to 59 of 1,463 rows, and an empty capture parses
   to `[]`, both with no error — and `resolveAgentForegroundProcessWithAvailability`
   then answers `available: true`. That is the `unverifiable` -> `exited` collapse
   the execution boundary forbids. The capture now rejects with
   `ProcessTableCaptureError` on a ceiling-length or row-less capture, so both the
   lenient and strict views fail loudly and callers report unavailable.

2. `maxBuffer` is now an explicit 32MB, matching the sibling reader in
   `pty-descendant-termination.ts` and its stated reasoning. Without it a 4,000-
   process host fails EVERY capture, degrading the whole subsystem permanently.

Separately, `readStructuredTuiProcessIdentity` polled a fresh whole-machine `ps`
every 50ms for up to 5s. Each capture costs ~0.065 CPU-s, and the 5s ceiling is
only reached when the child never appears — where the tight interval buys
nothing. The interval now holds at 50ms for the first second, then doubles to a
500ms cap. Identification latency is unchanged for any child appearing inside
that window, and the 5s ceiling is unchanged.
2026-09-02 13:22:25 -07:00
Neil 698584f6ec perf(pty): stop startup history GC freezing the main process for seconds (#18165)
`runHistoryGc` walked every terminal-history directory synchronously ten
seconds after launch: `readdirSync` on the root, then per directory a
`statSync`, a `readdirSync`, a `statSync` per file, an `existsSync`, a
`readFileSync` and a `JSON.parse`. On a real 613 MB root (2,776 dirs /
6,703 files) that is ~20,000 syscalls and 2,774 parses in one
uninterruptible pass — the main process was frozen for the whole of it,
7.5-10.6 s on the reporting machine.

Move the enumeration to `fs/promises` behind the existing
`forEachWithConcurrency` fixed-worker pool over an iterative frontier,
yielding through the shared `yieldToEventLoop()` every 32 entries. The
pass is cancellable and a second call joins the in-flight one rather
than racing its tombstone renames.

Max main-thread gap over the real root: 122-155 ms -> 1.1-1.9 ms idle,
522 ms -> 1.1 ms under load. Syscalls per pass 20,578 -> 17,803. Total
elapsed is lower too (88-108 ms vs 126-154 ms warm), so nothing is
smeared into a longer tail.

The prune decision logic and the tombstone path are untouched. A new
suite asserts the new walk removes exactly the set the synchronous walk
chose over a fixture covering every decision shape, and covers the races
async introduces: a directory removed mid-walk, a half-written
`meta.json`, and malformed/truncated/oversized metadata. All of those
resolve to "keep", matching what the sync version did on a read error.
2026-09-02 13:10:05 -07:00
Neil b00ec20731 perf(startup): stop an unreachable SSH host from gating local terminal restore (#18164)
* perf(startup): stop an unreachable SSH host from gating local terminal restore

An asleep or unreachable SSH target held the terminal-restoration gate for the
full 15s reconnect timeout, so no terminal restored — local ones included.
Startup now awaits only the target that owns the active workspace's tabs and
lets the rest connect in the background, folded into the existing deferred path
that reattaches their PTYs on tab focus.

Also splits the renderer's git-environment fence out of the first-window PTY
services barrier: worktree hydration needs shell-PATH generation and the managed
WSL CLI registration, not a daemon PTY spawn or a hook-server bind. Terminal
restoration still fences on the first-window services via
app:prepareTerminalStartupRestoration.

Measured with tests/tools/benchmarks/startup-time-bench.mjs (382 restored tabs,
28k-file profile, medians of 3):
  unreachable SSH host: 17.27s -> 1.34s to renderer-startup-hydration-done
  all-local:             1.98s -> 1.33s

* fix(startup): restore the startup-ordering oracle and keep a connected background SSH target undeferred

app-startup-routing.test.ts pinned the old step names, so the two ordering cases
went vacuous-then-red when the barrier split. Repoint them at the steps that now
carry the same fences: 'git-environment-barrier-await' (shell PATH + managed WSL,
the fence host Git needs) before hydration worktrees, and
'prepare-terminal-startup-restoration' (which awaits firstWindowStartupServicesReady
in main) before terminal reconnect. Both still fail against main's hydration source.

Also: the timed-out-eager rewrite of the deferred list re-added background targets
that had already connected, undoing removeDeferredSshReconnectTarget and sending
fresh panes on a reachable host down the cold-restore path.
2026-09-02 13:10:02 -07:00
Neil 53f105827b perf(windows): stop asking the process table for memory, and share one projection per snapshot (#18151)
Two costs on the Windows process-table hot path, plus the EDR doc that
described neither of them accurately.

1. The snapshot set `ProcessDataFlag.Memory` and surfaced `memoryBytes`,
   which nothing read. The addon serves that flag with a second
   `OpenProcess(PROCESS_QUERY_INFORMATION | PROCESS_VM_READ)` and a
   `GetProcessMemoryInfo` per process (process.cc:47-63), so the flag was
   one wasted handle per process per snapshot.

2. The shared TTL cache gave every pane the same native rows array, but
   each pane still ran `native.map(toProcessRow)` over the whole table,
   rebuilt a `childrenByPpid` Map from scratch, and did two linear scans.
   The `.map()` also handed `getProcessTableIndex` a new array each call,
   defeating the POSIX memo by construction. Both now cache per snapshot
   identity, and the POSIX resolver drops its duplicate descendant walk.

`getProcessTableIndex` / `buildProcessTableIndex` are generic over the row
shape so the Windows rows reuse the existing pass instead of a parallel one.

No behavior change: same rows in, same rows out, same descendant ordering
and same has-children answers.
2026-09-02 13:08:01 -07:00
Neil 084dbbc3b3 perf(persistence): build the state file once per save instead of seven times (#18161)
Every debounced save stringified the full persisted state, then ran two
`String.replace` passes per secret sentinel — one for the on-disk payload, one
for the guard hash. Each replace returns a rope the next one has to flatten
before it can search, so three sentinels cost seven flattened copies of a
4.65 MB state (a two-byte V8 string, ~8.9 MB each), and the state was then
UTF-8 encoded twice more: once inside `sha1.update(string)` and again inside
`handle.writeFile(payload, 'utf-8')`.

`applySecretSentinelSubstitutions` walks the state once with a single
alternation regex, encodes each literal run to a Buffer exactly once, and feeds
those same buffers to both the payload and the hash. Measured on the author's
4.65 MB store with three live secret slots: 48.8 MB -> 17.9 MB allocated per
save, 26.6 MB -> 0 of large_object_space churn, and 22.1 -> 15.1 ms (min) /
32.3 -> 16.9 ms (median) for build+hash+encode. Bytes on disk and the guard
hash are proven identical to the previous loop.

Separately, non-local host session partitions carried stale replicas of the
`browserUrlHistory` global — 589,807 bytes, 12.7% of the file — that neither
the split (which writes globals only to 'local') nor the merge (which reads
them only from 'local' unless local has none) can ever reach. The load path now
drops them when the local slice already holds the field. Only the two history
globals are dropped: the rest are read out of every partition by the worktree
ownership sweep or the mobile/runtime projections.
2026-09-02 13:05:16 -07:00
Neil 21a706a932 perf(terminal): stop shipping every agent spinner title frame to the renderer (#18155)
Main re-asserts a working OSC title per pane every 80ms (12.5/sec) while an
agent works, and every frame became its own pty:sideEffect IPC message. Both
renderer store writes already discard those frames via
isDecorativeAgentTitleFrameChange, and paired remote clients already never see
them (RuntimeClientEventBus's per-listener title gate). Only the local desktop
renderer was still paying for them.

Apply the same decorative gate main already computes for the mobile fan-out one
hop earlier, keeping a 500ms heartbeat so the renderer's 1500ms hook-done quiet
window still sees a working title and can cancel a Pi/OMP milestone 'done'.
2026-09-02 13:04:06 -07:00
Neil 104f9655e4 perf(git): answer remote-URL questions from one subprocess, not one per remote (#18158)
Four copies of the same loop ran `git remote` and then a serial
`git remote get-url <name>` per remote to answer "which remote has this
URL". On a repo with 58 remotes that is 59 subprocesses -- measured at
1083 ms -- for one question, and worktree create asks it several times.
`git remote -v` answers for every remote from one child, reporting the
same insteadOf-expanded first fetch URL `get-url` prints.

The batched `cat-file --batch-check` branch-conflict probe decides from
stdout, but its WSL route was unfenced, so a login-shell fallback printed
the distro banner onto the stream it parses. That broke the
one-line-per-ref contract, made every batch undecided, and fell straight
back to one `show-ref` per remote -- the cost the batch exists to remove.

Measured at 58 remotes / 4346 branches, spawns and wall time:
  push-target remote scan      59 -> 1  (1083 ms -> 8 ms)
  branch-conflict probe        60 -> 3  (984 ms -> 43 ms)
  configured push target      123 -> 6  (2707 ms -> 157 ms)
2026-09-02 12:53:48 -07:00
Brennan BensonandMerge Sim 1d94ebee3f fix(agents): stop a deeper vendor helper from stealing a pane's agent identity (#18062)
* fix(agents): keep outer agent identity over vendor helpers

* fix(agents): preserve outer identity across relay scans

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-02 12:49:15 -07:00
Brennan BensonandMerge Sim 616fa751da revert(native-chat): drop the speculative Fable model-switch detections (#18215)
Both changes shipped in #18055 were written against strings never observed
in a real session, and neither fixed a reported problem. Guessing at agent
output we have not seen is how the picker got a row that silently no-ops.

Fable consent detection is removed outright. It watched the session for
"Fable N uses usage credits and needs a one-time consent" and answered
`interaction-required`. No consent prompt appeared in any validation run —
the test account had already consented — so the matched wording was never
confirmed. With the detector gone nothing produces `interaction-required`,
so the outcome leaves the union and its unreachable handler goes with it.
A real consent prompt now reports the switch as unverified, which is the
honest failure mode for output we cannot recognize.

The weekly usage scope goes back to exact `display_name === 'fable'`. It
had been widened to `/^fable\b/` against a hypothetical rename of
Anthropic's own usage window; the API still reports "Fable", so the match
was insurance against a scenario with no evidence behind it.

Tests covering the removed behavior are deleted rather than rewritten,
including the two pre-existing `interaction-required` cases that asserted
the terminal is revealed.

The disabled-row filter from #18055 is deliberately untouched.

Claude-Session: https://claude.ai/code/session_01SJy4XGrdre6YaU1wYNKak4

Co-authored-by: Merge Sim <sim@local>
2026-09-02 11:20:53 -07:00
Jinjing 61e010079f New agent dashboard (#18222)
* more obvious toggle

* more obvious toggle

* feat(activity): redesign thread rows and add child agent filtering

- Emphasize task title and last activity in row layout over metadata
- Add child agent toggle; hide orchestration workers by default
- Support collapsible groups and ungrouped view mode
- Improve orchestration worker message handling to surface replies
- Add sidebar search and filter controls for agent activity

* periodic checkin

* feat(activity): add "Clear completed" action and performance improvement

- Add "Clear completed" action for activity threads with undo window; clears completed and interrupted rows from view, persists across restart
- Virtualize activity thread list to render only viewport-bounded rows
- Cache activity thread search text to prevent recomputation on every keystroke
- Cache dashboard bucket counts per-worktree for selective invalidation on unrelated changes
- Use useDeferredValue for activity search filtering to keep input responsive
- Make compact mode the default display for activity threads
- Add activity-cleared-at persisted state tracking (per-pane cutoff timestamps)

* improve style

* minor change

* feat(activity): add persisted host and project filters to agents view

Agents scope filters are deliberately separate from workspace-nav filters so a monitoring surface never inherits workspace context silently. Filters survive restarts and always display an active-filter chips row with hidden count, making filtering visible and reversible.

* Graduate Agents view from experimental, refine activity handling

- Agents Dashboard moves from experimental to standard feature with showAgentsSidebar setting controlling visibility
- Add identity-checked cache eviction (dropPersisted IPC) to prevent newer runs from being evicted when UI clears older status, fixing clear-completed safety
- Extract ActivityThreadHoverCardSummary and ActivityThreadListToolbar components for better organization and reusability
- Implement mark-thread-read as separate action from select with clickable bell icon
- Add hasActivityThreadWorkspace helper for checking workspace availability across hosts (SSH/runtime targets)
- Preserve scope filter array identity during hydration for memo optimization
- Track manually-unread turns in auto-ack to prevent re-acknowledgement
- Clean up activity cleared-at cutoffs on pane retirement
- Remove activity-thread-hover-card max-lines lint override (code refactored below threshold)

* Refactor agent cache identity to use timing fields only

- Simplify AgentStatusCacheIdentity: keep only paneKey, receivedAt, stateStartedAt
- This fixes silent no-ops where renderer-enriched fields diverged from main's cache
- Add worktree-jump-navigation for navigating activity to workspaces
- Add manual mark-unread protection separate from auto-ack
- Optimize activity owner resolution with per-build memoization
- Optimize detected worktree lookup with indexed search

* Remove sticky header, add scroll position persistence

Replace the floating sticky header overlay with scroll position memory via
a ref. This preserves the user's scroll location when switching between
threads or remounting the agents list, improving UX without requiring
React state.

* Implement sticky group headers in activity thread list

Keep group headers visible at the top while scrolling when threads are grouped. Headers stick to the viewport while their section is in view, then unstick as the next header approaches.

* add blue flash

* update settings appearnce

* Extracted activity acknowledgement/clearance actions from the oversized UI slice.
  - Removed dead sidebar search/menu props and the unused search ref.
  - Removed the unnecessary sidebar visibility bitmask.
  - Replaced hardcoded sidebar toggle colors with design-system tokens.
  - Removed duplicate “mark all read / clear completed” controls in the sidebar.
  - Preserved manual-unread state correctly across pane retire, transfer, and drop.
  - Made clear-completed cutoffs monotonic so clock skew cannot resurrect old activity.
  - Fixed blank workspace names in hover cards with the existing fallback helper.
  - Added missing localization entries and stabilized hydrated filter array identity.
  - Updated misleading Agents setting copy to describe both sidebar surfaces.

* add onboarding guide for the new agents panel

* Add activity clearance tracking and synced agent view settings

Agent view filters and presentation settings now sync across paired clients.
Preserves per-pane activity clearance cutoffs in persistent state. Improves
activity thread row accessibility with proper ARIA roles, and preserves
terminal host ownership after pane teardown via retained terminal handle.

* rm html

* Graduate Agents from experimental and improve activity visibility

- Migrate `showAgentsSidebar` setting from legacy experimental flags; default new profiles to the agents sidebar
- Replace scoped-thread filtering with visible-thread filtering so bulk actions (mark all read, clear completed) only affect rendered rows
- Rewrite child agent classification as a set of visible pane keys to fix orphan promotion and parent-cycle handling
- Improve activity cleared-at cutoff lifecycle: preserve on row dismissal (pane may still be live) but clear on pane removal
- Add pagehide flush for pending clear-completed evictions so quit/reload cannot replay cleared activity
- Polish agents sidebar: unread count badge, expand button, onboarding intro for migrated/new users
- Extract shared time-ago formatting to a library module
- Fix scroll restoration to defer until content can contain the saved offset
- Improve stable message hold for compact agent rows using state instead of refs
- Add worktree filter-visibility check to distinguish collapsed-but-unfiltered from filtered-hidden

* Graduate Agents from experimental and improve activity visibility

- Remove the deprecated full-page Agents view; fix settings navigation fallback
- Refactor bulk action bindings and separate mark-all-read from visible threads
- Preserve sidebar collapse state across remounts; fix child-agent badge filtering
- Add safety window for scroll-restore and improve worktree host-qualified filtering

* Graduate Agents from experimental and add manual unread tracking

- Move Agents sidebar from experimental settings to standard feature with intro flow
- Add persistent manual unread turn tracking for activity feed
- Consolidate workspace activation through activateAndRevealWorkspace dispatcher
- Improve sidebar view toggle with radio semantics and arrow-key navigation

* Graduate Agents sidebar and separate dashboard experiment

The Agents tab now has its own `showAgentsSidebar` setting (defaults on) independent from the dashboard popout experiment. Activity unread counting is simplified to count all events uniformly without mode-specific filtering. Dashboard visibility is now controlled solely by `experimentalAgentDashboardPopout`, with its own UI in the Experimental settings pane. Migration path updated: only `experimentalActivity=true` graduates to the sidebar; the dashboard experiment remains separate.

* Add agent-session tab support to activity tracking

Build activity event contexts from structured agent-session tabs and
worktree-attributed status entries. When activating a thread, try
agent-session tab activation before falling back to terminal pane.

* • The workspace sidebar tab is now a static Spaces
  label—no grouping-based “Projects” label or hidden
  width-reservation span.

* Show unread count badge and prioritize attention-needing agent threads

Activity group order now surfaces threads needing attention (blocked,
waiting, interrupted) before working/done so they're never buried. The
Agents tab shows an unread count badge while viewing Spaces, since the
open Agents list already highlights unread rows.

Also improves UX text ("Hide Agents" vs "Maybe later"), accessibility
with proper ARIA labels, and handles edge cases: preserves read state
for retained panes on SSH reconnect and handles deleted worktrees
gracefully in navigation.

* Batch agent-status evictions and optimize activity pane rebuilds

- Add dropPersistedStatusEntries batch API; consolidate evictions into one persist
- Implement fallback timeout in clear-completed for unseen toast callbacks
- Project only activity-relevant tabs; memoize terminal tab derivations
- Stabilize activity virtualizer key to prevent unnecessary item measurements

* Remove unread count badge from Agents sidebar tab

Simplify useActivityUnreadCount by removing the enabled parameter and
conditional logic, as the badge is no longer displayed in the UI.

* Deduplicate activity unread counts across source overlaps

Live pane status is the primary source; retained and migration entries
serve as fallback caches that may briefly overlap it during lifecycle
transitions. Count each pane only once by tracking seen keys, prioritizing
the live status as the canonical source.

Also fix monitoring state display: it's a distinct agent state, not a
tool-running row state, so exclude it from tool preview checks.

* Update activity pane tests to remove unread badge assertions

- Remove ActivityPaneVisibility type and readActivityPaneVisibility() helper
- Update agentsSidebarButton selector to match badge-less state
- Simplify assertions to check pane focus instead of visibility isolation
- Remove test for unread badge acknowledgement flow

* Fix activity pane workspace resolution and localization handling

- Thread defaultHostId through activity operations for correct host resolution
- Add language-aware caching for standalone terminal names with cache invalidation
- Fix scroll restoration bounds calculation for tall viewports
- Add focus management to sidebar radio group keyboard navigation
- Refresh localized sidebar content on language changes
- Preserve activity state across heartbeats to prevent history loss
- Improve host-id strictness in worktree jump navigation

* Preserve activity view when settings fetch fails

A failed window.api.settings.get() leaves settings null, which was
incorrectly treated as opt-out. Add the missing null check so the
activity-view gate only applies when settings are available.

Includes tests for this scenario and related edge cases in keyboard
navigation, worktree jumping, and session state handling.
2026-09-02 11:00:24 -07:00
Neil f737f3499f fix(relay): stream an oversized fs.listFiles reply instead of refusing it (#17954)
Opening Orca's own checkout over SSH cannot list its files in one response frame.
22,617 tracked paths average 58 characters, so the 20,001-row page the client asks
for serializes to 1,223,415 bytes — past `DISPATCHER_CONTROL_QUEUE_MAX_BYTES`, so
`sendResponse` demotes it to the `legacy-response` lane, where an unrelated
producer backlog can refuse it as an opaque `ResponseOverCapacity`. Break-even is
around 49 characters of average path; any `packages/<name>/src/...` monorepo is
over the line.

Picking a ceiling to refuse at does not fix that, it just moves where it shows up
and refuses listings that would have been delivered. `__streamResponse` already
exists for exactly this on the git methods, and it is its own negotiation in both
directions: an old client never sends it and gets the plain array on the
legacy-response lane as before, and an old relay ignores it and answers plainly,
which the client detects by the sentinel marker being absent. So fs.listFiles opts
into it — no new method, no new opcode, nothing to advertise — and the size of a
listing stops being a correctness question.

The response-stream registry becomes one per relay, shared by FsHandler and
GitHandler. A second registry is not an option and the header of
git-response-stream.ts says why: a client keys reassembly on `streamId` alone, so
two would hand out the same id and cross-feed chunks, and only the handler that
registers `git.responseAck` can credit the window a pump parks on.

Also declares `maxResults` on the runtime-RPC `files.listAll` and forwards it.
The mechanism "the client names its cap, so a full page reads as truncation" was
wired only on the Electron IPC hop; web and mobile were saved incidentally by
`remoteFileContentBudget` defaulting the cap inside `listRuntimeFiles`. A new
optional field is additive in both directions (wire rule 1).

The new Docker-gated spec is claimed by run-ssh-docker-e2e.mjs. The sharded e2e
lanes set no ORCA_E2E_SSH_DOCKER, so a Docker-gated spec that no runner names
self-skips everywhere and still reports green — pr-e2e-gate-contract enforces that.

Closes #12547
2026-09-02 05:36:54 -07:00
Neil f37d2fec97 fix(linux): land the reviewed Linux packaging stack on main (#18100)
* fix(linux): give the CLI one entrypoint by extracting the AppImage once

* refactor(linux): trim AppImage CLI registration seams

* test(cli): assert registration lock serialization

* fix(linux): fence AppImage terminal shim mounts

* fix(linux): accept extracted AppImage runtimes with APPDIR only

* docs(linux): make headless AppImage extraction runnable

* refactor(linux): import bundled launcher directly

* fix(linux): reclaim superseded AppImage payloads and packaged symlinks

Pruning removed 3215 of 3216 files from a superseded generation and always
stranded resources/app.asar, leaking ~105 MB per version update. Electron's
asar shim reports a *.asar file as a directory, so the recursive remove tried
to rmdir a real file and failed with ENOTEMPTY; the .catch(() => {}) hid it.
Reproduced end to end on Ubuntu 24.04: 519M -> 623M across one update, and
519M again once the payload is actually reclaimed.

removeExtractedAppImagePayload holds process.noAsar for the removal, counted
so overlapping removals cannot hand the shim back early, and the prune site
now warns with the path instead of swallowing the rejection. All three
removal sites use it -- staging cleanup and displaced roots leaked the same
way.

Also reclaim symlinks left by a packaged deb/rpm install, which the
extracted-cache-only rule turned into a hard conflict on a deb -> AppImage
migration, and name the remedy in the conflict error.

* fix(linux): bound the CLI registration lock wait

`retries: 1000` caps the attempt count, not elapsed time, so at up to 1s per
attempt an IPC-driven registration could hang ~16 minutes against a wedged
holder with no feedback.

A legitimate holder is bounded by the extraction timeout, so wait that plus
slack and then fail with a message naming the lock file, rather than hanging.
`maxRetryTime` is forwarded verbatim to the `retry` package by proper-lockfile.

* fix(linux): stop re-extracting the AppImage on inode metadata churn

The extracted-payload cache key hashed ctime alongside dev/ino/size/mtime.
ctime moves on any inode metadata write -- `chmod +x`, which every AppImage
user is told to run, plus `chown`, an ACL or SELinux relabel, and a backup
restore -- none of which alter a byte of the payload.

Measured on Ubuntu 24.04: `chmod +x` leaves dev, ino, size and mtime
identical and moves ctime alone, so the key changed and the next launch paid
a full ~519 MB re-extraction and a multi-second stall to rebuild a payload it
already had, then pruned the old generation.

Key on content identity instead. An in-place content change moves mtime and
almost always size; a replacement moves the inode. The existing
replace-in-place test still passes.

* fix(linux): stop CLI commands from falling through to Chromium startup

* refactor(cli): remove redundant command membership check

* test(cli): cover command-named project selectors

* fix(cli): redirect the open-url command before startup

* test(linux): cover AUR serve wrapper flags

* fix(linux): tighten CLI launch detection

* fix(linux): respect CLI flag value boundaries

* fix(linux): strip injected Chromium switches from CLI args

* fix(linux): report a missing display instead of dying in uv_close

* refactor(linux): read display locks without a preflight race

* fix(linux): preserve unverified external displays

* chore: format reliability gate manifest

* test(packaging): split runtime resource checks

* fix(linux): fail serve when no display is available

* fix(linux): do not treat a lockless X socket as a dead display

An X server writes its lock beside its socket and both survive a crash
(verified against Xvfb under SIGKILL), so a socket with no lock was never
left by a crashed server. It is an endpoint published from elsewhere: a
container bind-mounting only /tmp/.X11-unix, WSLg, or a foreign PID
namespace. Declaring those dead made the desktop gate exit(1) on displays
that work, with no workaround, and the serve gate refuse to start.

Liveness now splits by ownership. A foreign DISPLAY trusts a lockless
socket; Orca's own :99 does not, because removeStaleDisplayArtifacts
unlinks the lock before the socket and so manufactures that state itself --
adopting it would resurrect the orphan-socket bug and stop the cleanup from
self-healing. The stale-lock rejection is unchanged.

Also correct four doc statements this behaviour falsified.

* fix(linux): fail closed when a stale socket blocks the Xvfb rebind

Readiness only checked that /tmp/.X11-unix/X99 exists. A stale socket we
could not unlink still exists after our own Xvfb refused to bind, so Orca set
DISPLAY to a dead server and Chromium died in Ozone init.

Measured on Ubuntu 24.04 against the pre-fix build: with a leftover :99
socket and no lock, serve exits 139 (SIGSEGV), the socket inode is unchanged
before and after, and no lock is recreated -- it neither cleaned up nor
respawned. To a user that is a crash, not a misconfiguration.

This is reachable in the documented topology, where orca-xvfb.service has no
User= and runs as root while serve runs as User=orca: /tmp is sticky, so the
orca uid cannot unlink a root-owned socket, rmSync fails, and Xvfb exits with
the display already active.

Readiness now requires the display to actually be live -- our socket plus a
lock naming a running process -- so the same state reports an unusable
display and exits 1 with the existing diagnosis.

* fix(linux): recognise abstract X sockets and inherited Wayland fds

Two display setups this gate could not prove were refused outright, and on the
desktop path that is app.exit(1) with no workaround.

An X server may bind only the abstract namespace (`@/tmp/.X11-unix/X0`), which
leaves no filesystem socket to stat. Abstract addresses are kernel-owned and
vanish the moment the owner exits, so an entry in /proc/net/unix is proof of a
live server -- no lock file needed and no stale entry possible. Verified on
Ubuntu 24.04, where 139 such addresses were present.

WAYLAND_SOCKET is an already-connected fd handed over by the compositor, so
there is no path to stat and WAYLAND_DISPLAY may be unset entirely. Its
presence is the display.

Both are consulted only after the filesystem-socket check fails, so no
existing verdict changes.

* fix(linux): never treat Orca's own display number as a foreign endpoint

Recognising a lockless X socket as live is correct for an endpoint published
from elsewhere -- a container bind mount, WSLg -- because an X server writes
its lock beside its socket and both survive a crash. It is wrong for
VIRTUAL_DISPLAY_NUMBER, because Orca's own teardown unlinks the lock before
the socket and so manufactures that exact state.

The managed branch was already strict, but a caller that sets DISPLAY=:99
explicitly takes the foreign path and skipped it, accepting a dead display
left by Orca's own interrupted cleanup. Route the managed number through the
strict probe on both paths.

Found by an adversarial audit of the asymmetry introduced earlier in this
branch; the documented systemd topology is unaffected because its Xvfb writes
a real lock.

* test(linux): add a packaged-artifact contract for the CLI launch paths

* test(linux): avoid buffered serve readiness detection

* test(linux): signal AppImage serve owner directly

* test(linux): tolerate readiness timeout boundary

* test(linux): add startup margin to shutdown oracle

* ci(linux): give package contracts timeout headroom

* fix(ci): route all Linux packaging contract changes

* test(linux): poll shutdown readiness without tail leaks

* test(linux): bound shutdown cleanup grace

* test(linux): assert on CLI output, not the harness's own control lines

run-cli-case.sh echoes `RESULT status=N case=<name>`, and the two cases named
*-skills asserted `expectOutput: 'skills'`. That substring was satisfied by
the case name in the harness's own line, so 2 of 8 cases asserted nothing
about the command -- gutting `skills` entirely would still have gone green.

Control lines are now excluded before matching, and both cases assert the
rendered help header, which only real help output produces. Verified on an
Ubuntu 24.04 host: 8/8 still pass against a stack-tip AppImage.

Also register the gate in reliability-gates.jsonc, which #15085 added a CI
Docker gate without. Red/green is recorded from a stock release AppImage
failing 4 of 8, three of them at status 133 (SIGTRAP).

* fix(linux): require static AppImage runtimes (#17319)

* test(linux): reject a wrong-architecture native binary at packaging time

Cross-building the arm64 slice on an x64 host silently packed an x86-64
`pty.node` -- the rebuild logged "Forcing native rebuild for linux-arm64" and
shipped the host's binary anyway. Every gate here inspects symbol versions,
which are perfectly valid on the wrong architecture, so nothing noticed.

Observed on a Raspberry Pi 5: the packaged app loaded, then failed with
"Failed to load native module: pty.node", and the launch contract reported
3 of 8 cases crashed rather than naming the cause. Swapping in the aarch64
`pty.node` took the same build to 8/8.

Compare ELF `e_machine` against the slice being packaged and fail with the
offending path. Checked before the glibc pass, because a wrong-architecture
binary's symbol versions are valid but meaningless and would send the reader
down the wrong path.

Release CI builds arm64 on a native runner, so this guards local and future
cross-builds rather than a shipped artifact.

* test(linux): judge per-arch vendored binaries against their own path

The first CI run of the architecture gate failed the x64 package job on
`@parcel/watcher-linux-arm64-glibc/watcher.node`. That binary is arm64 on
purpose: the package ships every architecture and its loader picks the match,
so its presence in an x64 build is correct.

Judge a binary against the architecture its own path names, falling back to
the slice when the path names none. That keeps the case this gate exists for
-- `bin/linux-arm64-*/node-pty.node` holding an x86-64 binary, which is what
shipped to a Raspberry Pi 5 -- while letting multi-arch dependencies through.

Dry-run over the real dependency tree flags nothing for either target arch.

* fix(linux): move deb/rpm update installation outside Orca (#17318)

* fix(linux): complete deb/rpm package metadata

* fix(linux): preserve CLI link during package upgrades

* docs(linux): document local RPM build prerequisites

* fix(linux): move deb/rpm update installation outside Orca

* fix(updater): preserve Linux recovery across stale events

* fix(updater): fence stale downloaded events by active target

* fix(updater): preserve active Linux package recovery

* test(linux): keep workflow order assertion in scope

* test(updater): assert stale recovery stays silent

* fix(updater): preserve Linux package recovery after checks

* refactor(updater): keep Linux marker message with status

* fix(linux): describe the right manual update path for deb/rpm hosts

A remote host installed from .deb or .rpm now reports
manual-service-update-required, and the guidance told the operator to
"update through the service manager that starts this server" -- which is
correct for unsupported-headless-serve but wrong for a package install,
where nothing about the remedy involves the service manager.

Say both, keyed on how the host was installed.

* docs(linux): document orcad update restart safety

* docs(linux): scope restart census omissions

* docs(linux): use absolute service CLI launcher

* fix(serve): validate in-process serve options before startup (#17683)

* fix(linux): stop offering updates a distro-managed install cannot apply (#17918)

Closes #17702.

The resources/package-type marker is authoritative but never checked against
the host, so any repackager that unpacks Orca's .deb -- AUR, Nix, a container
rebuild -- inherits `deb` verbatim. Install feasibility was then computed
after a ~165 MB download, so those users got check -> download -> a card
promising an install command -> a dead end.

Validate the marker against the host: a deb/rpm marker with no matching
package manager in the trusted directories means a package manager owns this
install. This reuses the exact lists and resolver that
buildLinuxPackageInstallCommand already loops over, so a false positive is
impossible by construction -- any host flagged here would have failed with
no-package-manager after the download anyway. The gate only moves that
verdict earlier. Verified across Debian 12, Ubuntu 24.04, Arch, Fedora 40 and
openSUSE Leap: no false positive on a real deb host, correct on every
repackaging host.

The release is still reported, because the user does want to know 1.4.194
exists and to update through their distro; only the download path is closed.
`externallyManaged` is an additive optional field on the existing `available`
status, so older paired clients decode it unchanged. downloadUpdate() refuses
authoritatively, since main owns this verdict rather than the card, and
unwinds any pinned-build state first -- a Linux pinned jump resolves to
'release', and stranding isPinnedBuildActive would silently kill every
background check for the rest of the process.

Note the fix the issue suggests cannot work: electron-updater builds a
PacmanUpdater whose doDownloadUpdate looks for a .pacman asset Orca does not
publish, then dereferences undefined.

* style(cli): restore prettier wrapping on install error copy

* test(linux): re-pin the child-process ratchets and the batch-shim allowlist after the merge
2026-09-02 03:08:01 -07:00
Neil aa3ae6f56e fix(ssh): close the pty master fd leak on relay hosts too (#17920)
* fix(ssh): close the pty master fd leak on Linux relay hosts

The app gets the FD_CLOEXEC patch through pnpm patchedDependencies (#17914);
the relay installs stock node-pty from npm, where no pnpm patch reaches. Linux
is where that matters -- it is the only relay platform that takes forkpty()'s
no-atomic-O_CLOEXEC path, and it is also the only one that already compiles
node-pty at install time, so the fix costs a second compile rather than a first.

Ships the patch as a relay asset applied like the existing Windows console-list
one, and rebuilds only after the probe has proven node-pty loadable. The rebuild
is non-fatal by construction: the working build is moved aside first and moved
back on any failure, a failed attempt drops a skip marker so the compile is
attempted at most once per relay directory, and the caller swallows the whole
step. macOS and Windows relays never run it.

Measured on node:22 with a relay-style npm install: before, the master is
cloexec=false and shows up as `26 -> /dev/pts/ptmx` in both a later pty child
and a later child_process child; after, cloexec=true and neither child sees it.

Closes #17915.

* test(ssh): feed the cloexec patch exec to the hand-rolled namespace fixtures

These sequences are positional, so the new Linux-only patch exec swallowed the
READY slot and every install/repair case timed out waiting for the relay.

* fix(ssh): patch the pty master before publishing the shared native-deps tree

* fix(ssh): refuse to publish a native-deps tree whose cloexec patch did not take
2026-09-02 03:02:27 -07:00
Neil 34999e328e fix(orcad): stop demanding a spawn-helper only macOS builds (#18122)
node-pty declares the spawn-helper target inside binding.gyp's OS=="mac"
block and pty.cc execs it only under __APPLE__. Asserting it on
`!== 'win32'` made every Linux orcad boot degraded with
spawn_helper_missing while its terminals worked fine.

Route all four sites through one shared `usesNodePtySpawnHelper`
predicate: the precondition verdict, the prebuilt slot install, the
+x repair, and the prebuilds build script (which threw outright on a
Linux slot build).

Fixes #17844
2026-09-02 02:49:09 -07:00