Pullfrog: shortening-only jitter raised the mean rebind rate ~15%. Grant
55 min +/- 5 min instead; same cohort spread, unchanged steady-state load.
Auth expiry (5 min token) and the 75 s silence watchdog are enforced
separately, so a grant up to 60 min risks nothing.
CodeRabbit: a phone that hung up while resolveResume was in flight still
paid for resolveInviteForMove and assignments.resolve. Check after each
awaited lookup; regression test asserts neither later lookup runs.
Incident 2026-09-04 ~01:05Z: after a background/foreground cycle the phone's
relay dial timed out inside the cell's acceptClient DB phase while the fleet
was in a cell-inventory lock storm (55P03 retries ~7.5k/h vs a ~1k/h floor).
Relay (cloud/apps/relay)
- acceptClient checks socket.readyState after each serialized Postgres call
and abandons the accept once the phone has hung up, releasing the capacity
reservation, failing the credential reservation, and releasing the activity
lease it just acquired instead of leaking it to expiry cleanup and then
throwing host_data_reservation_already_bound at bind.
- New structured event orca_relay_client_accept_abandoned {stage, elapsedMs}
and runtime-metric fields clientAcceptsAbandonedByStageDelta /
clientAcceptAbandonedMsMax so the "phone gave up behind the lock" rate is
quantifiable per cell.
- Control lease grants are jittered: 55 min minus [0, 10 min). Every host
that (re)connected in the same minute rebound as one cohort every ~54 min
(c27 autoheal recreate at 23:23Z re-homed ~420 controls; ~1.1k 1006 +
~1k 4408 "control rebound" closes landed in a 3 s window at 00:50:14Z),
and each rebind is an activateControl transaction on the inventory lock.
No wire change: leaseExpiresAt was always a server-chosen absolute time.
Desktop (src/main/runtime/relay)
- Control rotation rebinds 1-6 min early instead of 1-2 min, so a re-homed
cohort spreads across cycles rather than pinning one phase for the life of
the process.
Phone (mobile/src/transport)
- openAuthenticatedDirectEndpoint treats 'reconnecting' as a failed probe. On
a dead LAN the foreground direct dial dies with an instant 1006 and the
direct client enters its own 500/1000/2000 ms backoff; the probe used to
wait out its full 12 s bound holding the supervisor mutex, so relay recovery
queued behind three doomed redials. The stage-aware bound from #18518 is
unaffected.
Not done here: rolling the 23 GCE cells onto the post-#18521 image (500 ms
lock_timeout) is a deploy owned by cloud-deploy-relay-production-same-cap.
* feat(agents): pane-identity canonical adapter, comparison telemetry, inventory ratchet phase 1
* fix(agents): preserve canonical coverage provenance
* refactor(agents): unify pane identity adapters for tranche 0
* fix(agents): keep title resolver cache-free after rebase
* Fix ladder tranche zero review findings
* fix(agents): restore title classifier memoization
* fix(agents): fence unknown canonical evidence sources
* docs: drop the ladder plan and decision table from the PR
Design docs stay out of the shipped tree; the code carries its own comments
and the decision table lives in the test fixture.
---------
Co-authored-by: Merge Sim <sim@local>
Patches @xterm/addon-search so one very long un-newlined line no longer overflows the stack, freezes the renderer, or goes unsearched. Submitted upstream as xtermjs/xterm.js#6149 (issue #6148); drop the patch once a release ships it. See the PR for measurements and the differential fuzz.
One SSH pane accumulated one extra reattachable lease per relay restart (2, 3, 4,
5, 6 across five), and every one of them costs a `pty.attach` round trip on every
later connect, forever. Nothing prunes `sshRemotePtyLeases`, so the fan-out only
grows.
`supersedeSiblingLeasesForPane` is fenced on the PTY the pane is durably bound to,
and `durablyBoundPtyIdForPane` read `state.workspaceSession` (local) before
`workspaceSessionsByHostId['ssh:<target>']`. But `persistPtyBinding(binding, hostId)`
updates ONLY the host partition:
AFTER-PERSIST local= ssh:t@@pty2:old:1 host= ssh:t@@pty2:new:1
So for the length of a reconnect the local copy still names the predecessor, the
fence resolved to it, supersession took an already-`expired` lease as its winner,
and returned having marked nothing. Both partitions agree again once the renderer
republishes its layout, which is why the settled store looks consistent and hid
this.
Read both partitions as an ordered list, target's own first, and test the fence by
membership rather than by equality with whichever was read first. Pick the winner
preferring a lease this client still has a route to, since the stale partition
names an expired one. Never retire a lease that is both bound and live, so a
partition disagreement can't strand a running remote process.
Superseded predecessors stay `expired` and are never `terminated`: losing a lease
is not evidence the shell died (docs/reference/ssh-execution-boundary.md). A pane
with no binding is skipped rather than pruned, so a genuine orphan stays askable.
Also re-runs supersession from the binding side after each spawn commit's binding
write, so the lease/binding order at a call site no longer decides, and reconciles
every pane for a target immediately before `reattachKnownPtys` reads the set it
feeds to `pty.attach` — that repairs stores which already accumulated these rows.
The guard suite could not catch this: every assertion bound the pane BEFORE
upserting the lease, an order no caller uses. Rewritten to the spawn commits' real
order (lease, then binding, then the binding-side trigger); it fails 8 assertions
without this change. Added a suite that drives the real `persistPtyIpcSpawnCommit`
rather than the store primitives, including the exact stale-partition state written
by production's own binding writer.
Verified on the Docker SSH lane: five `relay.js` SIGKILLs with recovery between
each, reattachable leases flat at one per pane.
Note: this bounds the reattach SET, not the store. `sshRemotePtyLeases` still has
no cap or TTL and rows still accumulate; pruning is left alone deliberately, since
an `expired` row without `supersededBy` is a genuine orphan and must not be dropped
on age.
* feat(native-chat): restore the terminal/chat switcher for bridge chat only
#16729 removed every user-facing terminal<->chat switching affordance as a
side effect of the structured Codex restructure ("renderer switching
affordances and their dead leftovers"). That was right for structured Codex
sessions, which render their own transcript with no live TUI underneath, but
it also took the switcher away from bridge native chat, which still reads the
terminal and has one to return to.
Restore all four surfaces, each gated so structured sessions keep the removal:
- pane header chat/terminal button (TerminalPaneHeaderOverlay)
- pane context-menu "Switch to chat/terminal view" (TerminalContextMenu)
- tab context-menu equivalent (SortableTabContextMenu)
- the keyboard chord, whose hook had survived uncalled since #16729
Gating is one rule in one place: `canSwitchNativeChatView` refuses whenever a
`structuredSessionId` is present, over the existing `canToggleNativeChat`
eligibility. Standalone structured tabs are already excluded by the
`contentType === 'terminal'` check; the new guard covers a terminal tab that
adopted a structured session. The shortcut hook applies the same rule.
The state plumbing (`viewMode`, `setTabViewMode`, `toggleTabViewMode`, host
mirroring, `native_chat_toggled` telemetry) was never removed, so this rewires
live actions rather than reintroducing logic.
SortableTab.tsx sat exactly at its 400-line cap, so its inline-rename state and
the window rename-request listener move to `use-sortable-tab-rename.ts` to make
room. No behavior change; its rename tests pass unmodified.
Two ratchets move for real, explained in place:
- store-subscription budget: per-pane listeners stay pinned at 17 (the folded
action bundle is still one listener); only the counterfactual pre-fold
constant grows 48 -> 49 for the added `toggleTabViewMode` key.
- hook-order parity: 204 -> 208 hooks for the four added `useCallback`s,
useMemo count unchanged at 8.
* fix(native-chat): restore bridge chat escape hatch
* test: update pane agent identity inventory
---------
Co-authored-by: Merge Sim <sim@local>
Two entry points create the same workspace and disagreed about how to read its
host. `orca-runtime-create-managed-worktree.ts:63` resolved through
`getRepoSshConnectionId` and then normalized the row; the `worktrees:create` IPC
handler branched on raw `repo.connectionId`
(`register-worktree-create-handlers.ts:66-69`). So a repo naming its owner only
as `executionHostId: 'ssh:<target>'` created remotely through the runtime and ran
`git worktree add` on the client against a remote path through IPC (#11163).
Same repo, two entry points, different answers.
Both now take one route, resolved through the existing layer
(`getRepoExecutionHostId` -> #18296's `resolveGitRouteForHost`). No new resolver.
The row normalization on the `ssh` variant is kept, and it is a **workaround, not
the pattern**. `createRemoteWorktree` and its callees re-read `repo.connectionId!`
at five depths in `ipc/worktree-remote.ts` (1627, 1847, 1848, 1865, 2029), so the
resolved connection has to reach them through the field they already read. It
travels only as far as that object does — anything downstream that re-reads the
row from the store still sees the unnormalized one, and it cannot express the
`runtime:` refusal on its own. Proper fix, deliberately not done here: give that
pipeline an explicit connection parameter and delete `repo.connectionId!` from it
so every reader becomes a compile error, the technique #18307/#18325 used. That
is a change inside a 2800-line module plus its callers, and it wants its own PR.
Three answers that used to collapse into one, now distinct at both entry points:
- `executionHostId: 'ssh:*'` with no `connectionId` -> that SSH host (IPC used to
create locally);
- `executionHostId: 'local'` with a surviving `connectionId` -> local, since a
local row cannot nest an SSH namespace. This is what `getRepoSshConnectionId`
and therefore the runtime sibling already answered; IPC used to go remote;
- `runtime:<env>` -> refused. Its worktree is created by that environment's own
server and the SSH target on its repo row is that server's nested one,
addressable only as (environmentId, targetId). The renderer already routes
runtime-environment creates over `worktree.create` RPC rather than this IPC
channel, so reaching either entry point with one is a routing mistake. Matches
`workspace-cleanup-git-route` and `runtime-git-command-target`.
Folder-workspace creation is untouched on both sides: it is a registration, not a
filesystem create, so the route is resolved after that branch on the IPC side, and
on the runtime side only the agent trust write consumes it — where a `runtime:`
host now yields `null` instead of the nested target, so the write stops going to a
same-named target in this client's table.
No wire or persistence change: the normalized row is a local value passed to the
create pipeline, never stored, and `CreateWorktreeResult` is untouched.
`removeManagedWorktree` resolved its host once — for the metadata prune
(`cleanupHostId ?? getRepoExecutionHostId(repo)`) — and then read raw
`repo.connectionId` for every step that touches the filesystem: the
`git worktree list` deciding whether the path is registered, the provider handed
to the unregistered-removal branch, the registered-remote-vs-local fork, and the
PTY/history teardown. One function, two spellings.
For a row naming its owner only as `executionHostId: 'ssh:<target>'` — the exact
class #18296 names — the list ran on the client against a remote path,
`removeRuntimeUnregisteredWorktree` was entered with `provider: null`, and the
metadata was pruned under `ssh:<target>` while a same-named *local* directory was
the one considered for deletion (#11163). #18358 made this reachable: it migrated
the cleanup scan, so `executionHostId`-only rows now surface as removable
candidates, but removal did not move with it.
Routing is now one answer for the whole removal, taken from the host the prune
already used, through #18296's host-keyed dispatch. The ambiguous
`provider: SshGitProvider | null` carrier is deleted from the callees rather than
supplemented, so every remaining reader is a compile error in the typed modules
that do the destructive work (`runtime-unregistered-worktree-removal`,
`runtime-registered-remote-worktree-removal`, `runtime-worktree-filesystem`).
The orchestrator itself carries `@ts-nocheck` from its mechanical split, so that
guarantee does not reach it — tests cover it instead.
Every change is in the refusing direction; nothing became more aggressive:
- an `ssh:` host with no registered provider throws instead of deleting a
client-side path (`requireSshGitProvider` already threw for rows that spelled
the same host as `connectionId`);
- `runtime:<env>` throws rather than dialling a same-named target in this
client's namespace, matching `workspace-cleanup-git-route` and
`runtime-git-command-target`;
- the folder-workspace teardown resolves its connection instead of reading the
raw field, so a `runtime:` row stops dialling the wrong namespace.
No wire or persistence change: `removeWorktreeMetadataAndHistory` already took
the resolved host, and the removal RPC result shape is untouched.
The relay asset from #17920 only rewrote the forkpty `default:` call site, which
sits in the `#else` arm of PtyFork's `#if defined(__APPLE__)`. macOS takes
`pty_posix_spawn`, so the asset had never patched anything a Mac executes -- and
`applyNodePtyMasterCloexecPatch` returned 'fixed' for any non-Linux host without
running the script at all, which is what publishes a tree to the shared
native-deps cache.
Stock `pty_posix_spawn` opens up to three throwaway ptys to push the real master
off fds 0-2 and never closes them: the cleanup loop is `for (; count > 0;
count--)`, but the first `posix_openpt()` in a running process already returns
>= 2, so it breaks with `count == 0` and the body never runs -- and where it does
run it closes `low_fds[count]`, never `low_fds[0]`. One orphaned /dev/ptmx fd per
terminal, for the life of the relay.
Ports the `low_fds` fix and the Apple-branch `pty_cloexec(master)` call from the
app's `config/patches/node-pty@1.1.0.patch`, byte-identical, and runs the gate on
darwin. macOS needs a different build layout than Linux: it has no `build/` at
all, so the fallback moved aside is `prebuilds/darwin-<arch>` -- which is also
what makes node-pty's install script fall through from "prebuild found" to
node-gyp -- and the compile writes a `build/Release` the loader checks first.
Verification is per-platform too: Linux's leak is inheritance (/proc), macOS's is
self-held (lsof).
Also corrects the asset's claim that "macOS re-opens the tty through uv_tty_init's
cloexec dup". Measured false: FD_CLOEXEC is not set on the master. What protects
it is POSIX_SPAWN_CLOEXEC_DEFAULT, one option away from gone since uid/gid drops
libuv back to fork()/exec() -- so the master is now marked there too.
Measured on darwin-arm64, one PTY per open/close cycle in a relay-shaped dir
running the relay's own commands:
before cycle:ptmx 1:1 2:2 3:3 ... 10:10 (10 after a settle)
after cycle:ptmx 1:0 2:0 3:0 ... 10:0 (0 after a settle)
Linux re-verified in docker node:22: inherited before, isolated after,
`already-patched` on the second run.
Refs #17915
Refs #8362
* fix(ssh): stop pane adoption certifying a death from the relay's not-found union
`attachStablePaneOwner` was the last reader that synthesised a runtime exit
from a reattach refusal, and it published code `0` — which
`orca-runtime-on-pty-exit` records as `rememberPtyLivenessVerdict(exited)`, a
death certificate whose only legitimate writer is a host-delivered exit frame.
The refusal it acted on is a union. `pty.attach` answers `PTY "<id>" not found`
both for a pid the relay probed with `isProcessAlive` and for an id its session
map simply never had — which, because ids carry a per-start mint epoch, is every
id minted before a relay restart, checked against nothing. So a relay restart
plus a reconnect certified a shell that was still running under the old daemon's
orphaned process tree, retired the pane binding, and cold-started a second agent
onto the same transcript. The sibling `handlePtyReattachFailure` has always
refused to certify from that union; this path did not.
- The relay marks the one refusal it backed with a liveness check
(`PTY_ATTACH_PROVEN_EXITED_MARKER`). The marker is additive, so an unmarked
answer — including an older relay's — stays ambiguous, which is the safe
direction.
- The client mints that half as `SshPtyProvenExitedOnRelayError`, a subclass so
every existing `isSshPtyAbsentFromRelayError` consumer is unchanged.
- Pane adoption publishes `UNVERIFIED_PROCESS_EXIT_CODE` (-1), the sentinel its
sibling publishes, and passes `hostExitConfirmed` only for evidence that
observed the process: the marked relay refusal, or `SessionNotFoundError` from
the registry that owns the PTY. The ambiguous half now records `unverifiable`
instead of `exited`.
- The gone-branch keys on the error type rather than the bare `PTY ".+" not
found` text, so an untyped string can no longer authorise abandoning a
binding — the discriminator `pty-connect-limits.ts` already documented.
Refs docs/reference/ssh-execution-boundary.md
* test(pty): make the pane-adoption fixtures throw what real providers throw
These four fixtures rejected with bare `new Error('Session not found: ...')` and
`new Error('PTY "..." not found')`. No provider produces either untyped:
`local-pty-spawn` and `decodeDaemonResponseError` both mint
`SessionNotFoundError`, and the SSH reattach path types the relay's wire text
before any pane sees it. Fixtures that skip the type were the reason a
message-shaped gate looked adequate.
The exit-code expectations move with it: the pane path now publishes the -1
stop sentinel plus `hostExitConfirmed`, so a certificate follows the evidence
rather than a synthesized zero.
The five `node ... | tee` steps in cloud-operate-relay-production-rehome-job.yml
reported tee's exit code, so a thrown inspect or apply passed green. The Aug 28
21:25Z and Aug 29 inspects and today's first inspect all printed
"director returned an invalid regional rehome control" (the durable control had
moved to generation 12 when the Aug 28 rehome aborted) and still succeeded.
`shell: bash` adds `-o pipefail`. A test pins the default and the tee count.
* feat(mobile): finalize structured native Codex chat
* fix(mobile): close structured chat lifecycle gaps
* wip(mobile): fence stale structured inventory and bound operation-id retention
Fence local structured-session inventory and subscription responses with a
sync generation so a toggle-off clear, reconnect restore, or retry cannot
apply a mirror from a superseded instance. Bound mobile ambiguous
operation-ID retention at 128 with unmount cleanup.
Staged on the reconcile branch only: the sync module is now 312 lines and
needs a real split before this can reach the PR head.
* fix(ci): split the structured session-tabs sync and give static analysis mobile types
The local structured session-tabs sync module outgrew the 300-line cap once it
took on generation fencing, so split it along its real seams instead of raising
the cap: the generation/cursor fence, snapshot projection, snapshot apply,
inventory refresh, and the subscription loop. The original path stays as a
barrel so no importer moves.
Repoint the host-session-mirror settle census at the apply module, which owns
two receipts now — the snapshot it mirrors in, and the toggle-off teardown that
retracts what it published. The teardown receipt is named rather than anonymous
so the pin says which direction it settles.
The changed-code quality gate lints mobile files and resolves their types from
mobile/node_modules, but mobile is a separate pnpm project that the root install
never populates, so every mobile type degraded to an `error` type and the gate
reported phantom findings. Install mobile dependencies in static analysis when
the diff touches mobile, gated on a new classifier output.
* fix(mobile): let a slow capability handshake still reach connected
The mobile capability update is an advisory whose result is discarded, yet an
unanswered one was fatal while an explicit rejection was tolerated. A 5s timeout
on the direct client force-closed the socket, and on the relay path it failed
`confirmResume` before `connected` was ever published, so a consistently slow
link redialled forever. Both paths now share one helper that settles every
ambiguous outcome (timeout, mid-flight drop) like a rejection and rejects only
when the frame never reached the wire — the one case nothing else recovers from,
since the socket's own desync force-close is gated on already being connected.
The generation guard still keeps a replaced session from connecting.
Retained structured-session operation ids were capped at 128 with oldest-first
eviction, but every retained id belongs to a send whose outcome is unknown, so
eviction turned a user's retry into a second message on the host. Bound the map
by expiry against the id's own embedded timestamp instead, mirroring the host's
operation ledger, so no id is released while the host would still honour it.
Also give the mobile CI install the root install's lockfile drift guard (mobile's
lockfile carries patchedDependencies a silent rewrite would drop), gate
mobile_dependencies on should_run, and key the pnpm store cache on both lockfiles.
* refactor(mobile): extract the relay pending-request registry
The merge composed two independently-sized changes — this branch's capability
handshake settle and main's dial-stage tracking — pushing the relay session file
to 304 lines against a 300 cap. Neither side broke it alone.
Move the in-flight request registry (id generation, tracking, settlement, and
reject-all with its delivery-ambiguity marking) into RelayPendingRequests,
matching the existing collaborator pattern alongside RelayDialStageTracker and
RpcSessionLivenessWatchdog. No behavior change.
---------
Co-authored-by: Merge Sim <sim@local>
* fix: scope workspace-creation-project tour target to project picker only
The tour target was previously applied to a container that included both
the project picker and the run target picker below it. Restructure the
layout to scope the target to only the project-related section, and add
a test to verify the tour target does not span into the run target picker.
* fix: scope workspace-creation-project tour target to project picker only
Move the tour target attribute from the outer project section to an inner
wrapper around just the combobox and its messages, excluding the header
label and "Add project" button. Update tests to verify the narrower scope.
The orphan-relay-PTY sweep authorizes `pty.shutdown { immediate: true }`, which
runs `forceKillPosixPtyProcessGroups`: collect every process group on the pane's
tty, then `killpg` each one. The blast radius is therefore (groups on the tty) x
(members of those groups, wherever they are). The idleness evidence measured only
the first factor, so three shapes read as idle and were SIGKILLed:
- with job control off (`set +m`) a background job keeps the SHELL's pgid, so the
tty carries exactly one process group and that group is running the user's build;
- a child that drops the controlling terminal (`ioctl(TIOCNOTTY)` without `setsid`)
keeps the pgid, reports `tpgid == -1`, and never appears in `ps -t <tty>`;
- a double-forked grandchild keeps the pgid and tty but reparents to pid 1, so the
`ppid` walk cannot reach it and the named-process backstop never fires.
`shellOwnsEveryTtyProcessGroup` now also requires the shell's own process group to
hold no other member anywhere in the table, indexed in the same single pass. A
pids-per-tty set would catch the first and third but not the second, which is why
the count is pgid-wide rather than tty-scoped. The wire field keeps its tty-shaped
name: the value only ever became stricter, so an old client skips more, never less.
Second, unrelated-in-mechanism but same file family: `foregroundSkipReason` summed
`capturedAgeMs + evidenceAgeSinceListingMs` without validating either. A non-numeric
`capturedAgeMs` makes the sum `NaN`, and `NaN > 5000` is false, so a malformed record
PASSED the freshness gate and proceeded toward the stop — the one place in the file
that defaulted toward kill. Nothing validated it on this path
(`mapSshPtyProcessList` checks the ownership fields and spreads the rest through;
`PtyProcessListAdmission` is not on the sweep path). It now runs
`isForegroundProcessEvidence` and fails closed.
Verified on real Linux, not only in mocks: a container drives `bash -i` on a real
pty, builds each construction, runs the real publisher and planner, and then calls
the real `forceKillPosixPtyProcessGroups`. Before, all three published
`shellOwnsEveryTtyProcessGroup: true`, planned SWEEP, and the planted pid was gone
after the signal. After, all three skip and survive, and an idle shell is still
reclaimed.
Residuals are written down at the predicate and in ssh-execution-boundary.md: the
capture is a snapshot (bounded by the evidence-age budget, not removed), and a
process the host's own `ps` cannot enumerate stays unobservable while `killpg`
still reaches it.
`orca worktree list` returned zero of 24 SSH worktrees at the default limit
(#18104). Rows are resolved repo by repo, so every SSH repo's rows land
contiguously at the end of the fleet order — the 24 remote rows sat at indices
496-520 of 521 and a plain `slice(0, 200)` never reached them.
The omission was not fully silent: text output printed `truncated: showing 200
of 521` and JSON carried `totalCount` / `truncated`. What was missing is that
the omission was *categorically every remote host* — no host column, no
`hostScope`, nothing to distinguish "200 of 521" from "one host is entirely
absent". Per docs/reference/ssh-execution-boundary.md, a listing that does not
name its scope reads as absolute.
Adopt the mechanism `terminal list` already has rather than inventing a second
one:
- `RuntimeTerminalListHostScope` becomes an alias of a shared
`RuntimeListingHostScope`, now also carried (optional, so old hosts are
unaffected) on `worktree.list` and `worktree.ps` results.
- `src/shared/host-balanced-listing-page.ts` round-robins the row cap across
hosts and returns the survivors in the caller's original relative order, so
the page stays a subsequence of the unbounded listing and nothing downstream
re-sorts. An uncapped listing is returned unchanged.
- `worktree list` / `worktree ps` text output gains a `host=` column and the
same trailing `scope:` line `terminal list` prints.
Third defect, same mechanism: `hostScope.omittedHostIds` is built from the
runtime's own bookkeeping, so it names `runtime:` ids for servers that are no
longer paired — 6 of 9 in the recorded QA run hard-error when queried. Since
`hostScope` is *the* documented way to complete a partial listing, that makes
the mechanism unreliable for its intended use.
Annotate rather than filter. Dropping an id would shrink what the listing
admits it did not cover, and the boundary doc requires a listing to name its
gaps — the gap is real whether or not this machine can name the host that owns
it. `src/cli/omitted-host-scope-selectors.ts` resolves each omitted id against
this machine's pairing store and the runtime's SSH-target registry and attaches
the exact flag that reaches it, or `null` marked "not selectable from this
machine". This is a client-side annotation: nothing new goes over the wire, it
answers "can I select it" and never "is it up", and the SSH round trip is only
paid when an `ssh:` host was actually omitted.
No `--host` filter was added; the host column plus scope line covers the
reported need without a new selector axis.
`orca host list --environment m4air` was not ignoring the flag — it was applying
it to half the answer. `shouldIgnoreRemoteSelection` never pinned the `host`
family, so the SSH-target lookup was routed to m4air while paired servers were
still read from this machine's own pairing store, and the handler stamped the
envelope `_meta.runtimeId: "local"` regardless. The result was one listing
describing two hosts: the openclaw row silently disappeared, which reads as
"m4air has no SSH targets". `environment list --environment X` had the pin but
no guard, so the flag vanished with no signal at all.
Reject rather than route. `host list` answers "what can this machine target and
with what flag"; its paired-server half comes from a client-local store and
cannot be routed at all, so any routed answer is necessarily half-substituted —
rule 1 of docs/reference/ssh-execution-boundary.md. `environment list` is
entirely client-local, so there is no other host to ask. This matches the
`account` and `artifacts` precedent, the only two pinned families that already
paired the pin with a rejection guard.
- pin the `host` family so an ambient ORCA_ENVIRONMENT cannot produce the same
two-machine listing with no flag to reject; `runtimeId: "local"` is now true
- extract the duplicated `rejectRemoteSelectionFlags` from account.ts and
artifacts.ts into src/cli/remote-selection-flag-rejection.ts
- `environment show` / `environment rm` / `environment add` are untouched: there
`--environment` and `--pairing-code` name the row to act on, not a route
Three reveal paths called bare setActiveWorktree + activateTabAndFocusPane,
skipping setActiveView('terminal'), ensureWorktreeHasInitialTerminal and
resumeSleepingAgentSessionsForWorktree. A parked SSH workspace has no resident
tab until those run, so the reveal landed on a workspace with no terminal.
Route all three through the incumbent activateAndRevealWorkspace dispatcher
(which the sidebar and "Jump to workspace" already use, and which also handles
folder workspaces). The Activity row-click additionally early-returned when the
thread's tab was absent from tabsByWorktree/unifiedTabsByWorktree, which made a
cold-parked remote thread a silent no-op; residency is now probed after
activation, so a revived tab is focused and a genuinely retained thread still
activates its workspace instead of doing nothing.
Also stop asserting `exited` from an absence of local state: SshPtyProvider
reports no authoritative buffer snapshot and the relay has no snapshot RPC, so
a null preview snapshot for a remote pty is loss of contact. The preview and
the no-pty dialog branch now say the remote preview is unavailable rather than
claiming the pane closed. Adding the relay snapshot RPC stays out of scope --
it needs capability negotiation.
Fixes#16731