Commit Graph
10082 Commits
Author SHA1 Message Date
Jinwoo-H efaefdc594 docs(cloud): PR #18606 opened; c7 two-hour evidence 2026-09-04 04:26:28 -04:00
Jinwoo-H ffb833a2ae docs(cloud): lock-removal PR review outcome 2026-09-04 04:25:16 -04:00
Jinwoo-H 9f162c151d docs(cloud): correct the rollout-speed plan after reading the cell job 2026-09-04 04:11:45 -04:00
Jinwoo-H 600b95921b docs(cloud): faster same-cap rollout design from measured limits 2026-09-04 04:09:55 -04:00
Jinwoo-H 21f69ded3b docs(cloud): lock-removal PR status 2026-09-04 04:06:24 -04:00
Jinwoo-H 6fffb33a4c docs(cloud): record the agreed lock-first plan 2026-09-04 03:59:51 -04:00
Jinwoo-H d0ff6f6141 docs(cloud): c7 post-restore observations 2026-09-04 02:27:24 -04:00
Jinwoo-H c4af2a21cf docs(cloud): c7 canary succeeded 2026-09-04 02:26:23 -04:00
Jinwoo-H f8dd8bd2ac docs(cloud): c7 up on target image 2026-09-04 02:25:15 -04:00
Jinwoo-H 60f60edf0b docs(cloud): attribute the 06:20Z 503 burst 2026-09-04 02:23:01 -04:00
Jinwoo-H f61a879c17 docs(cloud): record c27/c29 crash storm during c7 apply 2026-09-04 02:19:12 -04:00
Jinwoo-H 169d6ff816 docs(cloud): record c7 drain recovery numbers 2026-09-04 02:16:45 -04:00
Jinwoo-H 903da34b95 docs(cloud): record c7 drain effect on the director 2026-09-04 02:12:14 -04:00
Jinwoo-H 87047a2927 docs(cloud): record dry-run #6 pass and c7 canary dispatch 2026-09-04 02:08:02 -04:00
Jinwoo-H 4d80d24108 docs(cloud): record dry-run #6 dispatch and PR states 2026-09-04 01:47:45 -04:00
Jinwoo-H 5a96670d7f docs(cloud): record Asia-cell autoheal recreate amplifier 2026-09-04 01:43:09 -04:00
Jinwoo-H 262a18867d docs(cloud): record dry-run #5 freeze on c27 autoheal recreate 2026-09-04 01:42:28 -04:00
Jinwoo-H bad866ac8a docs(cloud): record dry-run #4 freeze (c27 crash) and dry-run #5 2026-09-04 01:40:33 -04:00
Jinwoo-H d85fe9e414 docs(cloud): record canary blast radius and dry-run #4 2026-09-04 01:27:30 -04:00
Jinwoo-H 7a8d1fa948 docs(cloud): record monitor dry-run dispatch inputs 2026-09-04 01:17:26 -04:00
Jinwoo-H 16c6aef916 docs(cloud): record retries-bar decision and PR #18580 basis 2026-09-04 01:15:27 -04:00
Jinwoo-H 8ebff89106 docs(cloud): move relay reconnect findings under cloud/docs
The root directory guard blocks new root-level files.
2026-09-04 01:13:22 -04:00
Jinwoo-H d338b344ec docs: relay reconnect investigation findings and roll status 2026-09-04 01:02:38 -04:00
Jinwoo-H aa3a7ef147 fix(relay): make the control lease jitter symmetric
Pullfrog: shortening-only jitter raised the mean rebind rate ~15%. Grant
55 min +/- 5 min instead; same cohort spread, unchanged steady-state load.
Auth expiry (5 min token) and the 75 s silence watchdog are enforced
separately, so a grant up to 60 min risks nothing.
2026-09-03 23:58:17 -04:00
Jinwoo-H e5ccd4a1a4 fix(relay): check client abandonment between each accept lookup
CodeRabbit: a phone that hung up while resolveResume was in flight still
paid for resolveInviteForMove and assignments.resolve. Check after each
awaited lookup; regression test asserts neither later lookup runs.
2026-09-03 23:33:08 -04:00
Jinwoo-H fdf6fb2f27 fix(relay): stop dead accept work, spread control rotations, fail direct probes fast
Incident 2026-09-04 ~01:05Z: after a background/foreground cycle the phone's
relay dial timed out inside the cell's acceptClient DB phase while the fleet
was in a cell-inventory lock storm (55P03 retries ~7.5k/h vs a ~1k/h floor).

Relay (cloud/apps/relay)
- acceptClient checks socket.readyState after each serialized Postgres call
  and abandons the accept once the phone has hung up, releasing the capacity
  reservation, failing the credential reservation, and releasing the activity
  lease it just acquired instead of leaking it to expiry cleanup and then
  throwing host_data_reservation_already_bound at bind.
- New structured event orca_relay_client_accept_abandoned {stage, elapsedMs}
  and runtime-metric fields clientAcceptsAbandonedByStageDelta /
  clientAcceptAbandonedMsMax so the "phone gave up behind the lock" rate is
  quantifiable per cell.
- Control lease grants are jittered: 55 min minus [0, 10 min). Every host
  that (re)connected in the same minute rebound as one cohort every ~54 min
  (c27 autoheal recreate at 23:23Z re-homed ~420 controls; ~1.1k 1006 +
  ~1k 4408 "control rebound" closes landed in a 3 s window at 00:50:14Z),
  and each rebind is an activateControl transaction on the inventory lock.
  No wire change: leaseExpiresAt was always a server-chosen absolute time.

Desktop (src/main/runtime/relay)
- Control rotation rebinds 1-6 min early instead of 1-2 min, so a re-homed
  cohort spreads across cycles rather than pinning one phase for the life of
  the process.

Phone (mobile/src/transport)
- openAuthenticatedDirectEndpoint treats 'reconnecting' as a failed probe. On
  a dead LAN the foreground direct dial dies with an instant 1006 and the
  direct client enters its own 500/1000/2000 ms backoff; the probe used to
  wait out its full 12 s bound holding the supervisor mutex, so relay recovery
  queued behind three doomed redials. The stage-aware bound from #18518 is
  unaffected.

Not done here: rolling the 23 GCE cells onto the post-#18521 image (500 ms
lock_timeout) is a deploy owned by cloud-deploy-relay-production-same-cap.
2026-09-03 23:07:52 -04:00
Brennan BensonandMerge Sim 90780acb85 refactor(agents): one pane-identity resolver behind six thin adapters (tranche 0) (#18243)
* feat(agents): pane-identity canonical adapter, comparison telemetry, inventory ratchet phase 1

* fix(agents): preserve canonical coverage provenance

* refactor(agents): unify pane identity adapters for tranche 0

* fix(agents): keep title resolver cache-free after rebase

* Fix ladder tranche zero review findings

* fix(agents): restore title classifier memoization

* fix(agents): fence unknown canonical evidence sources

* docs: drop the ladder plan and decision table from the PR

Design docs stay out of the shipped tree; the code carries its own comments
and the decision table lives in the test fixture.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-03 18:07:56 -07:00
Neil 8463dcb7b9 fix(terminal): make wrapped-line search rewind iterative and bound its scans (#18402)
Patches @xterm/addon-search so one very long un-newlined line no longer overflows the stack, freezes the renderer, or goes unsearched. Submitted upstream as xtermjs/xterm.js#6149 (issue #6148); drop the patch once a release ships it. See the PR for measurements and the differential fuzz.
2026-09-03 17:59:16 -07:00
Ihor 963839aa4f docs: add Ukrainian README translation
iho <4000375+iho@users.noreply.github.com>
2026-09-03 17:33:02 -07:00
foXaCe 49d6d35b16 feat(i18n): add French UI locale
foXaCe <290678+foXaCe@users.noreply.github.com>
2026-09-03 17:32:59 -07:00
hwantage 48cb575db5 feat(i18n): localize Orca Account settings and navigation to Korean
hwantage <82494320+hwantage@users.noreply.github.com>
2026-09-03 17:32:55 -07:00
Trevin Chow 912463c278 docs: document localization workflow
tmchow <517103+tmchow@users.noreply.github.com>
2026-09-03 17:32:51 -07:00
Jinwoo Hong dec1a1d788 fix(i18n): ship onboarding integration capability strings in boot catalog
Jinwoo-H <73622457+Jinwoo-H@users.noreply.github.com>
2026-09-03 17:32:48 -07:00
Neil e42c60e8a3 fix(ssh): resolve a pane's binding from the target partition, not the stale local copy (#18546)
One SSH pane accumulated one extra reattachable lease per relay restart (2, 3, 4,
5, 6 across five), and every one of them costs a `pty.attach` round trip on every
later connect, forever. Nothing prunes `sshRemotePtyLeases`, so the fan-out only
grows.

`supersedeSiblingLeasesForPane` is fenced on the PTY the pane is durably bound to,
and `durablyBoundPtyIdForPane` read `state.workspaceSession` (local) before
`workspaceSessionsByHostId['ssh:<target>']`. But `persistPtyBinding(binding, hostId)`
updates ONLY the host partition:

  AFTER-PERSIST  local= ssh:t@@pty2:old:1   host= ssh:t@@pty2:new:1

So for the length of a reconnect the local copy still names the predecessor, the
fence resolved to it, supersession took an already-`expired` lease as its winner,
and returned having marked nothing. Both partitions agree again once the renderer
republishes its layout, which is why the settled store looks consistent and hid
this.

Read both partitions as an ordered list, target's own first, and test the fence by
membership rather than by equality with whichever was read first. Pick the winner
preferring a lease this client still has a route to, since the stale partition
names an expired one. Never retire a lease that is both bound and live, so a
partition disagreement can't strand a running remote process.

Superseded predecessors stay `expired` and are never `terminated`: losing a lease
is not evidence the shell died (docs/reference/ssh-execution-boundary.md). A pane
with no binding is skipped rather than pruned, so a genuine orphan stays askable.

Also re-runs supersession from the binding side after each spawn commit's binding
write, so the lease/binding order at a call site no longer decides, and reconciles
every pane for a target immediately before `reattachKnownPtys` reads the set it
feeds to `pty.attach` — that repairs stores which already accumulated these rows.

The guard suite could not catch this: every assertion bound the pane BEFORE
upserting the lease, an order no caller uses. Rewritten to the spawn commits' real
order (lease, then binding, then the binding-side trigger); it fails 8 assertions
without this change. Added a suite that drives the real `persistPtyIpcSpawnCommit`
rather than the store primitives, including the exact stale-partition state written
by production's own binding writer.

Verified on the Docker SSH lane: five `relay.js` SIGKILLs with recovery between
each, reattachable leases flat at one per pane.

Note: this bounds the reattach SET, not the store. `sshRemotePtyLeases` still has
no cap or TTL and rows still accumulate; pruning is left alone deliberately, since
an `expired` row without `supersededBy` is a genuine orphan and must not be dropped
on age.
2026-09-03 16:47:32 -07:00
Brennan BensonandMerge Sim e85ebb0086 feat(native-chat): restore the terminal/chat switcher for bridge chat only (#18532)
* feat(native-chat): restore the terminal/chat switcher for bridge chat only

#16729 removed every user-facing terminal<->chat switching affordance as a
side effect of the structured Codex restructure ("renderer switching
affordances and their dead leftovers"). That was right for structured Codex
sessions, which render their own transcript with no live TUI underneath, but
it also took the switcher away from bridge native chat, which still reads the
terminal and has one to return to.

Restore all four surfaces, each gated so structured sessions keep the removal:

- pane header chat/terminal button (TerminalPaneHeaderOverlay)
- pane context-menu "Switch to chat/terminal view" (TerminalContextMenu)
- tab context-menu equivalent (SortableTabContextMenu)
- the keyboard chord, whose hook had survived uncalled since #16729

Gating is one rule in one place: `canSwitchNativeChatView` refuses whenever a
`structuredSessionId` is present, over the existing `canToggleNativeChat`
eligibility. Standalone structured tabs are already excluded by the
`contentType === 'terminal'` check; the new guard covers a terminal tab that
adopted a structured session. The shortcut hook applies the same rule.

The state plumbing (`viewMode`, `setTabViewMode`, `toggleTabViewMode`, host
mirroring, `native_chat_toggled` telemetry) was never removed, so this rewires
live actions rather than reintroducing logic.

SortableTab.tsx sat exactly at its 400-line cap, so its inline-rename state and
the window rename-request listener move to `use-sortable-tab-rename.ts` to make
room. No behavior change; its rename tests pass unmodified.

Two ratchets move for real, explained in place:
- store-subscription budget: per-pane listeners stay pinned at 17 (the folded
  action bundle is still one listener); only the counterfactual pre-fold
  constant grows 48 -> 49 for the added `toggleTabViewMode` key.
- hook-order parity: 204 -> 208 hooks for the four added `useCallback`s,
  useMemo count unchanged at 8.

* fix(native-chat): restore bridge chat escape hatch

* test: update pane agent identity inventory

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-03 16:28:45 -07:00
Neil 5d8532f6d3 fix(worktrees): resolve the execution host at both worktree-create entry points (#18545)
Two entry points create the same workspace and disagreed about how to read its
host. `orca-runtime-create-managed-worktree.ts:63` resolved through
`getRepoSshConnectionId` and then normalized the row; the `worktrees:create` IPC
handler branched on raw `repo.connectionId`
(`register-worktree-create-handlers.ts:66-69`). So a repo naming its owner only
as `executionHostId: 'ssh:<target>'` created remotely through the runtime and ran
`git worktree add` on the client against a remote path through IPC (#11163).
Same repo, two entry points, different answers.

Both now take one route, resolved through the existing layer
(`getRepoExecutionHostId` -> #18296's `resolveGitRouteForHost`). No new resolver.

The row normalization on the `ssh` variant is kept, and it is a **workaround, not
the pattern**. `createRemoteWorktree` and its callees re-read `repo.connectionId!`
at five depths in `ipc/worktree-remote.ts` (1627, 1847, 1848, 1865, 2029), so the
resolved connection has to reach them through the field they already read. It
travels only as far as that object does — anything downstream that re-reads the
row from the store still sees the unnormalized one, and it cannot express the
`runtime:` refusal on its own. Proper fix, deliberately not done here: give that
pipeline an explicit connection parameter and delete `repo.connectionId!` from it
so every reader becomes a compile error, the technique #18307/#18325 used. That
is a change inside a 2800-line module plus its callers, and it wants its own PR.

Three answers that used to collapse into one, now distinct at both entry points:

- `executionHostId: 'ssh:*'` with no `connectionId` -> that SSH host (IPC used to
  create locally);
- `executionHostId: 'local'` with a surviving `connectionId` -> local, since a
  local row cannot nest an SSH namespace. This is what `getRepoSshConnectionId`
  and therefore the runtime sibling already answered; IPC used to go remote;
- `runtime:<env>` -> refused. Its worktree is created by that environment's own
  server and the SSH target on its repo row is that server's nested one,
  addressable only as (environmentId, targetId). The renderer already routes
  runtime-environment creates over `worktree.create` RPC rather than this IPC
  channel, so reaching either entry point with one is a routing mistake. Matches
  `workspace-cleanup-git-route` and `runtime-git-command-target`.

Folder-workspace creation is untouched on both sides: it is a registration, not a
filesystem create, so the route is resolved after that branch on the IPC side, and
on the runtime side only the agent trust write consumes it — where a `runtime:`
host now yields `null` instead of the nested target, so the write stops going to a
same-named target in this client's table.

No wire or persistence change: the normalized row is a local value passed to the
create pipeline, never stored, and `CreateWorktreeResult` is untouched.
2026-09-03 16:27:13 -07:00
Neil 3c91404820 fix(worktrees): route managed worktree removal by resolved execution host (#18529)
`removeManagedWorktree` resolved its host once — for the metadata prune
(`cleanupHostId ?? getRepoExecutionHostId(repo)`) — and then read raw
`repo.connectionId` for every step that touches the filesystem: the
`git worktree list` deciding whether the path is registered, the provider handed
to the unregistered-removal branch, the registered-remote-vs-local fork, and the
PTY/history teardown. One function, two spellings.

For a row naming its owner only as `executionHostId: 'ssh:<target>'` — the exact
class #18296 names — the list ran on the client against a remote path,
`removeRuntimeUnregisteredWorktree` was entered with `provider: null`, and the
metadata was pruned under `ssh:<target>` while a same-named *local* directory was
the one considered for deletion (#11163). #18358 made this reachable: it migrated
the cleanup scan, so `executionHostId`-only rows now surface as removable
candidates, but removal did not move with it.

Routing is now one answer for the whole removal, taken from the host the prune
already used, through #18296's host-keyed dispatch. The ambiguous
`provider: SshGitProvider | null` carrier is deleted from the callees rather than
supplemented, so every remaining reader is a compile error in the typed modules
that do the destructive work (`runtime-unregistered-worktree-removal`,
`runtime-registered-remote-worktree-removal`, `runtime-worktree-filesystem`).
The orchestrator itself carries `@ts-nocheck` from its mechanical split, so that
guarantee does not reach it — tests cover it instead.

Every change is in the refusing direction; nothing became more aggressive:

- an `ssh:` host with no registered provider throws instead of deleting a
  client-side path (`requireSshGitProvider` already threw for rows that spelled
  the same host as `connectionId`);
- `runtime:<env>` throws rather than dialling a same-named target in this
  client's namespace, matching `workspace-cleanup-git-route` and
  `runtime-git-command-target`;
- the folder-workspace teardown resolves its connection instead of reading the
  raw field, so a `runtime:` row stops dialling the wrong namespace.

No wire or persistence change: `removeWorktreeMetadataAndHistory` already took
the resolved host, and the removal RPC result shape is untouched.
2026-09-03 16:21:52 -07:00
Neil 04ae62202a fix(ssh): close the macOS relay's per-terminal pty fd leak (#18534)
The relay asset from #17920 only rewrote the forkpty `default:` call site, which
sits in the `#else` arm of PtyFork's `#if defined(__APPLE__)`. macOS takes
`pty_posix_spawn`, so the asset had never patched anything a Mac executes -- and
`applyNodePtyMasterCloexecPatch` returned 'fixed' for any non-Linux host without
running the script at all, which is what publishes a tree to the shared
native-deps cache.

Stock `pty_posix_spawn` opens up to three throwaway ptys to push the real master
off fds 0-2 and never closes them: the cleanup loop is `for (; count > 0;
count--)`, but the first `posix_openpt()` in a running process already returns
>= 2, so it breaks with `count == 0` and the body never runs -- and where it does
run it closes `low_fds[count]`, never `low_fds[0]`. One orphaned /dev/ptmx fd per
terminal, for the life of the relay.

Ports the `low_fds` fix and the Apple-branch `pty_cloexec(master)` call from the
app's `config/patches/node-pty@1.1.0.patch`, byte-identical, and runs the gate on
darwin. macOS needs a different build layout than Linux: it has no `build/` at
all, so the fallback moved aside is `prebuilds/darwin-<arch>` -- which is also
what makes node-pty's install script fall through from "prebuild found" to
node-gyp -- and the compile writes a `build/Release` the loader checks first.
Verification is per-platform too: Linux's leak is inheritance (/proc), macOS's is
self-held (lsof).

Also corrects the asset's claim that "macOS re-opens the tty through uv_tty_init's
cloexec dup". Measured false: FD_CLOEXEC is not set on the master. What protects
it is POSIX_SPAWN_CLOEXEC_DEFAULT, one option away from gone since uid/gid drops
libuv back to fork()/exec() -- so the master is now marked there too.

Measured on darwin-arm64, one PTY per open/close cycle in a relay-shaped dir
running the relay's own commands:

  before  cycle:ptmx  1:1 2:2 3:3 ... 10:10   (10 after a settle)
  after   cycle:ptmx  1:0 2:0 3:0 ... 10:0    (0 after a settle)

Linux re-verified in docker node:22: inherited before, isolated after,
`already-patched` on the second run.

Refs #17915
Refs #8362
2026-09-03 16:13:07 -07:00
Neil a5d6114baf fix(ssh): stop pane adoption certifying a death from the relay's not-found union (#18531)
* fix(ssh): stop pane adoption certifying a death from the relay's not-found union

`attachStablePaneOwner` was the last reader that synthesised a runtime exit
from a reattach refusal, and it published code `0` — which
`orca-runtime-on-pty-exit` records as `rememberPtyLivenessVerdict(exited)`, a
death certificate whose only legitimate writer is a host-delivered exit frame.

The refusal it acted on is a union. `pty.attach` answers `PTY "<id>" not found`
both for a pid the relay probed with `isProcessAlive` and for an id its session
map simply never had — which, because ids carry a per-start mint epoch, is every
id minted before a relay restart, checked against nothing. So a relay restart
plus a reconnect certified a shell that was still running under the old daemon's
orphaned process tree, retired the pane binding, and cold-started a second agent
onto the same transcript. The sibling `handlePtyReattachFailure` has always
refused to certify from that union; this path did not.

- The relay marks the one refusal it backed with a liveness check
  (`PTY_ATTACH_PROVEN_EXITED_MARKER`). The marker is additive, so an unmarked
  answer — including an older relay's — stays ambiguous, which is the safe
  direction.
- The client mints that half as `SshPtyProvenExitedOnRelayError`, a subclass so
  every existing `isSshPtyAbsentFromRelayError` consumer is unchanged.
- Pane adoption publishes `UNVERIFIED_PROCESS_EXIT_CODE` (-1), the sentinel its
  sibling publishes, and passes `hostExitConfirmed` only for evidence that
  observed the process: the marked relay refusal, or `SessionNotFoundError` from
  the registry that owns the PTY. The ambiguous half now records `unverifiable`
  instead of `exited`.
- The gone-branch keys on the error type rather than the bare `PTY ".+" not
  found` text, so an untyped string can no longer authorise abandoning a
  binding — the discriminator `pty-connect-limits.ts` already documented.

Refs docs/reference/ssh-execution-boundary.md

* test(pty): make the pane-adoption fixtures throw what real providers throw

These four fixtures rejected with bare `new Error('Session not found: ...')` and
`new Error('PTY "..." not found')`. No provider produces either untyped:
`local-pty-spawn` and `decodeDaemonResponseError` both mint
`SessionNotFoundError`, and the SSH reattach path types the relay's wire text
before any pane sees it. Fixtures that skip the type were the reason a
message-shaped gate looked adequate.

The exit-code expectations move with it: the pane path now publishes the -1
stop sentinel plus `hostExitConfirmed`, so a certificate follows the evidence
rather than a synthesized zero.
2026-09-03 16:13:03 -07:00
Jinwoo Hong 7d27c841b4 fix(cloud): run the rehome control job under pipefail (#18537)
The five `node ... | tee` steps in cloud-operate-relay-production-rehome-job.yml
reported tee's exit code, so a thrown inspect or apply passed green. The Aug 28
21:25Z and Aug 29 inspects and today's first inspect all printed
"director returned an invalid regional rehome control" (the durable control had
moved to generation 12 when the Aug 28 rehome aborted) and still succeeded.
`shell: bash` adds `-o pipefail`. A test pins the default and the tee count.
2026-09-03 18:38:33 -04:00
Brennan BensonandMerge Sim 98e77ef1a7 feat(mobile): structured native Codex chat (#18074)
* feat(mobile): finalize structured native Codex chat

* fix(mobile): close structured chat lifecycle gaps

* wip(mobile): fence stale structured inventory and bound operation-id retention

Fence local structured-session inventory and subscription responses with a
sync generation so a toggle-off clear, reconnect restore, or retry cannot
apply a mirror from a superseded instance. Bound mobile ambiguous
operation-ID retention at 128 with unmount cleanup.

Staged on the reconcile branch only: the sync module is now 312 lines and
needs a real split before this can reach the PR head.

* fix(ci): split the structured session-tabs sync and give static analysis mobile types

The local structured session-tabs sync module outgrew the 300-line cap once it
took on generation fencing, so split it along its real seams instead of raising
the cap: the generation/cursor fence, snapshot projection, snapshot apply,
inventory refresh, and the subscription loop. The original path stays as a
barrel so no importer moves.

Repoint the host-session-mirror settle census at the apply module, which owns
two receipts now — the snapshot it mirrors in, and the toggle-off teardown that
retracts what it published. The teardown receipt is named rather than anonymous
so the pin says which direction it settles.

The changed-code quality gate lints mobile files and resolves their types from
mobile/node_modules, but mobile is a separate pnpm project that the root install
never populates, so every mobile type degraded to an `error` type and the gate
reported phantom findings. Install mobile dependencies in static analysis when
the diff touches mobile, gated on a new classifier output.

* fix(mobile): let a slow capability handshake still reach connected

The mobile capability update is an advisory whose result is discarded, yet an
unanswered one was fatal while an explicit rejection was tolerated. A 5s timeout
on the direct client force-closed the socket, and on the relay path it failed
`confirmResume` before `connected` was ever published, so a consistently slow
link redialled forever. Both paths now share one helper that settles every
ambiguous outcome (timeout, mid-flight drop) like a rejection and rejects only
when the frame never reached the wire — the one case nothing else recovers from,
since the socket's own desync force-close is gated on already being connected.
The generation guard still keeps a replaced session from connecting.

Retained structured-session operation ids were capped at 128 with oldest-first
eviction, but every retained id belongs to a send whose outcome is unknown, so
eviction turned a user's retry into a second message on the host. Bound the map
by expiry against the id's own embedded timestamp instead, mirroring the host's
operation ledger, so no id is released while the host would still honour it.

Also give the mobile CI install the root install's lockfile drift guard (mobile's
lockfile carries patchedDependencies a silent rewrite would drop), gate
mobile_dependencies on should_run, and key the pnpm store cache on both lockfiles.

* refactor(mobile): extract the relay pending-request registry

The merge composed two independently-sized changes — this branch's capability
handshake settle and main's dial-stage tracking — pushing the relay session file
to 304 lines against a 300 cap. Neither side broke it alone.

Move the in-flight request registry (id generation, tracking, settlement, and
reject-all with its delivery-ambiguity marking) into RelayPendingRequests,
matching the existing collaborator pattern alongside RelayDialStageTracker and
RpcSessionLivenessWatchdog. No behavior change.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-03 15:19:26 -07:00
mors c79e1c097b fix(i18n): localize remaining onboarding UI
Reviewed and approved by Codex.
2026-09-03 15:07:42 -07:00
韦编三绝 f13f2472c6 fix(i18n): distinguish Duplicate from Copy in Simplified Chinese
Reviewed and approved by Codex.
2026-09-03 15:07:38 -07:00
Ilya Gusev 8262fb147f fix(i18n): extract translateSearchKeyword calls so settings-search keywords reach en.json
Reviewed and approved by Codex.
2026-09-03 15:07:34 -07:00
Jinjing b1186c6beb Fix scope of workspace-creation-project tour target (#18502)
* fix: scope workspace-creation-project tour target to project picker only

The tour target was previously applied to a container that included both
the project picker and the run target picker below it. Restructure the
layout to scope the target to only the project-related section, and add
a test to verify the tour target does not span into the run target picker.

* fix: scope workspace-creation-project tour target to project picker only

Move the tour target attribute from the outer project section to an inner
wrapper around just the combobox and its messages, excluding the header
label and "Add project" button. Update tests to verify the narrower scope.
2026-09-03 14:45:58 -07:00
Neil f35015d0c8 fix(ssh): measure pane idleness in the unit the sweep's kill operates on (#18415)
The orphan-relay-PTY sweep authorizes `pty.shutdown { immediate: true }`, which
runs `forceKillPosixPtyProcessGroups`: collect every process group on the pane's
tty, then `killpg` each one. The blast radius is therefore (groups on the tty) x
(members of those groups, wherever they are). The idleness evidence measured only
the first factor, so three shapes read as idle and were SIGKILLed:

- with job control off (`set +m`) a background job keeps the SHELL's pgid, so the
  tty carries exactly one process group and that group is running the user's build;
- a child that drops the controlling terminal (`ioctl(TIOCNOTTY)` without `setsid`)
  keeps the pgid, reports `tpgid == -1`, and never appears in `ps -t <tty>`;
- a double-forked grandchild keeps the pgid and tty but reparents to pid 1, so the
  `ppid` walk cannot reach it and the named-process backstop never fires.

`shellOwnsEveryTtyProcessGroup` now also requires the shell's own process group to
hold no other member anywhere in the table, indexed in the same single pass. A
pids-per-tty set would catch the first and third but not the second, which is why
the count is pgid-wide rather than tty-scoped. The wire field keeps its tty-shaped
name: the value only ever became stricter, so an old client skips more, never less.

Second, unrelated-in-mechanism but same file family: `foregroundSkipReason` summed
`capturedAgeMs + evidenceAgeSinceListingMs` without validating either. A non-numeric
`capturedAgeMs` makes the sum `NaN`, and `NaN > 5000` is false, so a malformed record
PASSED the freshness gate and proceeded toward the stop — the one place in the file
that defaulted toward kill. Nothing validated it on this path
(`mapSshPtyProcessList` checks the ownership fields and spreads the rest through;
`PtyProcessListAdmission` is not on the sweep path). It now runs
`isForegroundProcessEvidence` and fails closed.

Verified on real Linux, not only in mocks: a container drives `bash -i` on a real
pty, builds each construction, runs the real publisher and planner, and then calls
the real `forceKillPosixPtyProcessGroups`. Before, all three published
`shellOwnsEveryTtyProcessGroup: true`, planned SWEEP, and the planted pid was gone
after the signal. After, all three skip and survive, and an idle shell is still
reclaimed.

Residuals are written down at the predicate and in ssh-execution-boundary.md: the
capture is a snapshot (bounded by the evidence-age budget, not removed), and a
process the host's own `ps` cannot enumerate stays unobservable while `killpg`
still reaches it.
2026-09-03 14:44:32 -07:00
Jinwoo Hong f974e98162 test(cloud): derive both reachability directions for the relay inventory census (#18524)
Mirrors stablyai/orca-cloud#472 (c82f98f), byte-identical under cloud/.
2026-09-03 17:43:40 -04:00
Neil 95eed52801 fix(cli): report which hosts a worktree listing covered, and stop the cap starving remote ones (#18417)
`orca worktree list` returned zero of 24 SSH worktrees at the default limit
(#18104). Rows are resolved repo by repo, so every SSH repo's rows land
contiguously at the end of the fleet order — the 24 remote rows sat at indices
496-520 of 521 and a plain `slice(0, 200)` never reached them.

The omission was not fully silent: text output printed `truncated: showing 200
of 521` and JSON carried `totalCount` / `truncated`. What was missing is that
the omission was *categorically every remote host* — no host column, no
`hostScope`, nothing to distinguish "200 of 521" from "one host is entirely
absent". Per docs/reference/ssh-execution-boundary.md, a listing that does not
name its scope reads as absolute.

Adopt the mechanism `terminal list` already has rather than inventing a second
one:

- `RuntimeTerminalListHostScope` becomes an alias of a shared
  `RuntimeListingHostScope`, now also carried (optional, so old hosts are
  unaffected) on `worktree.list` and `worktree.ps` results.
- `src/shared/host-balanced-listing-page.ts` round-robins the row cap across
  hosts and returns the survivors in the caller's original relative order, so
  the page stays a subsequence of the unbounded listing and nothing downstream
  re-sorts. An uncapped listing is returned unchanged.
- `worktree list` / `worktree ps` text output gains a `host=` column and the
  same trailing `scope:` line `terminal list` prints.

Third defect, same mechanism: `hostScope.omittedHostIds` is built from the
runtime's own bookkeeping, so it names `runtime:` ids for servers that are no
longer paired — 6 of 9 in the recorded QA run hard-error when queried. Since
`hostScope` is *the* documented way to complete a partial listing, that makes
the mechanism unreliable for its intended use.

Annotate rather than filter. Dropping an id would shrink what the listing
admits it did not cover, and the boundary doc requires a listing to name its
gaps — the gap is real whether or not this machine can name the host that owns
it. `src/cli/omitted-host-scope-selectors.ts` resolves each omitted id against
this machine's pairing store and the runtime's SSH-target registry and attaches
the exact flag that reaches it, or `null` marked "not selectable from this
machine". This is a client-side annotation: nothing new goes over the wire, it
answers "can I select it" and never "is it up", and the SSH round trip is only
paid when an `ssh:` host was actually omitted.

No `--host` filter was added; the host column plus scope line covers the
reported need without a new selector axis.
2026-09-03 14:43:09 -07:00
Neil 9bed758e36 fix(cli): reject runtime selectors on host list and environment list (#18405)
`orca host list --environment m4air` was not ignoring the flag — it was applying
it to half the answer. `shouldIgnoreRemoteSelection` never pinned the `host`
family, so the SSH-target lookup was routed to m4air while paired servers were
still read from this machine's own pairing store, and the handler stamped the
envelope `_meta.runtimeId: "local"` regardless. The result was one listing
describing two hosts: the openclaw row silently disappeared, which reads as
"m4air has no SSH targets". `environment list --environment X` had the pin but
no guard, so the flag vanished with no signal at all.

Reject rather than route. `host list` answers "what can this machine target and
with what flag"; its paired-server half comes from a client-local store and
cannot be routed at all, so any routed answer is necessarily half-substituted —
rule 1 of docs/reference/ssh-execution-boundary.md. `environment list` is
entirely client-local, so there is no other host to ask. This matches the
`account` and `artifacts` precedent, the only two pinned families that already
paired the pin with a rejection guard.

- pin the `host` family so an ambient ORCA_ENVIRONMENT cannot produce the same
  two-machine listing with no flag to reject; `runtimeId: "local"` is now true
- extract the duplicated `rejectRemoteSelectionFlags` from account.ts and
  artifacts.ts into src/cli/remote-selection-flag-rejection.ts
- `environment show` / `environment rm` / `environment add` are untouched: there
  `--environment` and `--pairing-code` name the row to act on, not a route
2026-09-03 14:43:05 -07:00
Neil 232d04f541 fix(dashboard): open remote sessions from every agent reveal path (#18403)
Three reveal paths called bare setActiveWorktree + activateTabAndFocusPane,
skipping setActiveView('terminal'), ensureWorktreeHasInitialTerminal and
resumeSleepingAgentSessionsForWorktree. A parked SSH workspace has no resident
tab until those run, so the reveal landed on a workspace with no terminal.

Route all three through the incumbent activateAndRevealWorkspace dispatcher
(which the sidebar and "Jump to workspace" already use, and which also handles
folder workspaces). The Activity row-click additionally early-returned when the
thread's tab was absent from tabsByWorktree/unifiedTabsByWorktree, which made a
cold-parked remote thread a silent no-op; residency is now probed after
activation, so a revived tab is focused and a genuinely retained thread still
activates its workspace instead of doing nothing.

Also stop asserting `exited` from an absence of local state: SshPtyProvider
reports no authoritative buffer snapshot and the relay has no snapshot RPC, so
a null preview snapshot for a remote pty is loss of contact. The preview and
the no-pty dialog branch now say the remote preview is unavailable rather than
claiming the pane closed. Adding the relay snapshot RPC stays out of scope --
it needs capability negotiation.

Fixes #16731
2026-09-03 14:43:01 -07:00