Commit Graph
10327 Commits
Author SHA1 Message Date
Merge Sim ab00a5a9f3 Run the Claude structured integration suite as a runtime client
The suite exercises agentSession.* for Claude, not the mobile surface: nothing
in it asserts anything mobile-specific and its sibling integration suites use
'runtime'. Mobile now additionally requires the experimental structured-chat
setting, which structured-agent-session.test.ts pins in both states, so the
stale 'mobile' fixture was claiming coverage it never had.
2026-09-03 19:50:22 -07:00
Merge Sim 8caba35704 Merge remote-tracking branch 'origin/main' into brennanb2025/claude-sweep-rederivation-fence
# Conflicts:
#	src/main/runtime/rpc/methods/session-tab-agent-status-projection.test.ts
#	src/main/runtime/rpc/methods/structured-agent-session.test.ts
#	src/main/runtime/rpc/methods/structured-agent-session.ts
2026-09-03 19:35:16 -07:00
Merge Sim bafc314bee fix(claude): fence the forced sweep on re-derivation, not on lstart's second
A descendant forked in the same wall-clock second as every walk that sees it was
signalled with SIGTERM and then never escalated, so a SIGTERM-resistant child
survived tab close, app quit and a full relaunch. Two children of one parent
96ms apart across a second boundary took opposite paths. The leak predates this
branch: it reproduces with the change reverted.

`ps lstart` has one-second resolution, so a walk landing inside a row's birth
second can never rule out a pid recycled later in that same second. But a walk
is not a match: a ppid walk only reaches what the root actually parents, and the
root is pinned by Node's own handle, so a row the walk re-derived is ours
whatever second it was born in -- a stranger would have to have been forked into
our tree, and then it is not a stranger. Fence the escalation on that.

Rows a merge retained from an earlier walk are not re-derived and still answer
to the start-time fence, which remains correct for them.

Scoped to callers that revalidate identity before signalling, which is the
Claude close path. Codex teardown reaches this same verifier and is unchanged;
the argument holds there too, but widening it is its own deliberate change.

Also reverts two changes from the previous attempt at this leak. Advancing the
capture boundary on a later walk is inert once the sweep fences on re-derivation
-- both key on the same set of rows, so the new term short-circuits for exactly
the rows whose boundary it advanced. The extra ladder refresh was a duplicate
full process-table read: close() already awaits tree.refresh() immediately
before proveClaudeChildExit, on the only path that reaches it.

Known property: the kill lands roughly a grace window after the walk that proved
membership, so a pid recycled inside that gap could in principle be signalled.
It is bounded -- matchingSnapshotRows already requires the live row to carry the
same start-second and pgid, so an impostor must be born in the remainder of that
one second, land on that exact pid, and sit in the same process group, and it
has already received the unfenced SIGTERM from the same loop.
2026-09-03 18:36:21 -07:00
Shahar MorandMerge Sim 7106101ed2 fix(mobile): restore terminal input when reopening worktrees (#16239)
* fix(mobile): restore terminal input when reopening worktrees

* test(mobile): update session parity facts

* refactor(mobile): split host client hooks

* chore: restore localization formatter scope

* fix(mobile): retain RpcClient type import

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-03 18:32:53 -07:00
Jinwoo Hong 11aace8dec fix(relay): reject malformed percent-escapes on upgrade instead of throwing (#18547)
decodeURIComponent on the /v1/connect/ and /v1/host/data/ path segments threw
URIError out of the http 'upgrade' listener, which is uncaught and kills the
relay process. Any client that sends GET /v1/connect/% could take down a cell
(and every connection on it) or a director instance. Pre-existing since the
splice landed (orca-cloud #20); not introduced by the import.

A malformed escape now takes the existing 4xx reject branch. The blackbox test
sends three malformed connect targets and one host-data target to the real
server and asserts no uncaughtException fires and a well-formed upgrade still
gets 101 afterwards; reverting either site fails it.
2026-09-03 21:31:53 -04:00
Neil 8c1a28d39c fix(i18n): repair French locale drift breaking static analysis (#18550)
The French UI locale landed with two catalog drifts that fail `static
analysis` on every PR in the repo:

- `fr.json` carried 15 keys absent from `en.json` (and from every other
  locale), so `verify:localization-catalog` rejected it. They are stale
  entries generated against an older `en.json` snapshot; none is
  referenced anywhere in the source.
- `settings.appearance.language.french` had no call site supplying a
  literal default, which promotes it to a boot-bundle-required entry
  that `en-runtime-required.json` does not ship, so
  `verify:localization-runtime-catalog` rejected it.

Registering the key in settings search alongside its siblings fixes the
runtime-catalog failure at its source and closes the real gap the drift
exposed: French was the only supported language not findable in settings
search.

`en-runtime-required.json` is deliberately untouched — the sync script
regenerates it wholesale and would drop 925 entries the check itself
documents as harmless.
2026-09-03 18:24:43 -07:00
Jinwoo Hong dd9eaa9585 fix(cloud): retry the committed-winner collision codes in relay schema startup (#18553)
* fix(cloud): retry the committed-winner collision codes in relay schema startup

`CREATE TABLE IF NOT EXISTS` only checks the name before the catalog inserts, so
the loser of a concurrent CREATE fails in one of two ways depending on timing:
on the catalog unique index (23505, which the startup retry already handled) or,
when the winner has committed by the time the loser reaches TypeCreate /
heap_create_with_catalog, on the name check those routines repeat (42710
duplicate type, 42P07 duplicate relation). The predicate treated the latter as
fatal, so a director could fail startup on a table it was about to find present.

This is what turned `postgres-schema-concurrency-postgres.test.ts` red on main
and on every relay PR (CI's shared runner loses the race more often than a dev
box): a throwaway diagnostic run in CI reported 42710 from TypeCreate and 42P07
from heap_create_with_catalog as the only rejection reasons.

Treat 42710/42P07 as retryable for `CREATE TABLE IF NOT EXISTS` and 42P07 for
`CREATE [UNIQUE] INDEX IF NOT EXISTS`; every other statement shape still fails
fast. The concurrency test now runs ten rounds and reports the loser's SQLSTATE
instead of a bare boolean.

* chore(cloud): allowlist the RFC 6455 example Sec-WebSocket-Key for upgrade tests

Cloud Verify's Secret scan runs gitleaks over --all refs, so the raw-socket
upgrade test on fix/relay-upgrade-malformed-uri (#18547) trips every cloud PR's
scan until its allowlist reaches main. Land the allowlist here first.
2026-09-03 21:20:27 -04:00
Brennan BensonandMerge Sim 90780acb85 refactor(agents): one pane-identity resolver behind six thin adapters (tranche 0) (#18243)
* feat(agents): pane-identity canonical adapter, comparison telemetry, inventory ratchet phase 1

* fix(agents): preserve canonical coverage provenance

* refactor(agents): unify pane identity adapters for tranche 0

* fix(agents): keep title resolver cache-free after rebase

* Fix ladder tranche zero review findings

* fix(agents): restore title classifier memoization

* fix(agents): fence unknown canonical evidence sources

* docs: drop the ladder plan and decision table from the PR

Design docs stay out of the shipped tree; the code carries its own comments
and the decision table lives in the test fixture.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-03 18:07:56 -07:00
Neil 8463dcb7b9 fix(terminal): make wrapped-line search rewind iterative and bound its scans (#18402)
Patches @xterm/addon-search so one very long un-newlined line no longer overflows the stack, freezes the renderer, or goes unsearched. Submitted upstream as xtermjs/xterm.js#6149 (issue #6148); drop the patch once a release ships it. See the PR for measurements and the differential fuzz.
2026-09-03 17:59:16 -07:00
Merge Sim 4108e951e5 Merge commit '6ade5578596e4a1d173f7da6763e35ae9e0e9950' into brennanb2025/claude-sdk-ready-mainsync 2026-09-03 17:42:38 -07:00
Merge Sim 6ade557859 Stop treating an unanswerable create-support probe as a refusal
A worktree is not resolvable for a beat after createWorktree resolves, so a
probe fired immediately after creation fails the RPC with selector_not_found
instead of answering. The catch collapsed that into `supported = false`, so the
composer refused and quietly built a terminal session — the gate never said no,
it was never asked successfully. Elapsed time was the only input that decided
whether a Claude launch went structured.

"Could not answer" and "answered no" are different states and only the second
is a verdict. Retry while the host cannot yet resolve the selector, with a
bounded backoff that covers the measured window with margin, and keep refusing
on the first ask for everything else. Fail-closed is unchanged: a probe that
still cannot be answered when the budget is spent refuses.

The retry is narrowed with the shared error-code matcher, which classifies a
token that transports re-wrap into a longer message without matching prose that
merely mentions it.

Codex never probes, so this race has never been able to refuse a Codex launch —
the race itself is identical for it. Recorded at the early return, because
whoever gives Codex a probe inherits the bug.
2026-09-03 17:41:12 -07:00
Ihor 963839aa4f docs: add Ukrainian README translation
iho <4000375+iho@users.noreply.github.com>
2026-09-03 17:33:02 -07:00
foXaCe 49d6d35b16 feat(i18n): add French UI locale
foXaCe <290678+foXaCe@users.noreply.github.com>
2026-09-03 17:32:59 -07:00
hwantage 48cb575db5 feat(i18n): localize Orca Account settings and navigation to Korean
hwantage <82494320+hwantage@users.noreply.github.com>
2026-09-03 17:32:55 -07:00
Trevin Chow 912463c278 docs: document localization workflow
tmchow <517103+tmchow@users.noreply.github.com>
2026-09-03 17:32:51 -07:00
Jinwoo Hong dec1a1d788 fix(i18n): ship onboarding integration capability strings in boot catalog
Jinwoo-H <73622457+Jinwoo-H@users.noreply.github.com>
2026-09-03 17:32:48 -07:00
Merge Sim c75a3f3ab5 Merge commit 'a3be41978d' into brennanb2025/claude-sdk-ready-mainsync 2026-09-03 17:25:30 -07:00
Merge Sim eea957b097 Merge commit 'd7e2c58549' into brennanb2025/claude-sdk-ready-mainsync 2026-09-03 17:25:30 -07:00
Merge Sim a3be41978d Support structured Claude when accounts are registered but none is selected
Registered-but-deselected Claude accounts were refused, which is behaviourally
identical to having no accounts at all: the auth policy does not strip, ambient
auth is the truth, and the UI names no host identity. A user who deselected
their accounts silently got legacy chat with nothing explaining why.

Nothing selected for the host runtime is two states the settings cannot tell
apart after the fact, because pruneInvalidClaudeRuntimeSelection empties the
host slot and persists null in the second one:

  honest deselection      -> ambient auth, UI names nothing   -> SUPPORTED
  the WSL-only steady state -> ambient auth, UI names the WSL account -> REFUSED

The presence of any WSL-bound account in the list decides. Simplifying this to
"none active -> supported" re-opens the auth-identity misrepresentation, so the
tests fail loudly on exactly that: five of them, across the unit rule and the
createSupport path.
2026-09-03 17:23:15 -07:00
Merge Sim 3b38af0c21 Treat an absent managed-account list as empty, not as unreadable
An empty claudeManagedAccounts array and a missing one are the same answer:
this user has no managed Claude accounts, so nothing claims an identity and
ambient auth is the truth. Refusing on absence strands any profile that simply
never wrote the key, and it disagrees with the auth policy, whose own predicate
takes `(accounts ?? [])` for exactly this reason.

Only settings that cannot be READ stay unknown, and those still refuse — as do
a WSL-bound active account and a selection naming an account the list does not
explain.

The earlier reasoning treated a missing field as settings we failed to parse.
That conflated "not present" with "not readable"; only the second is unknown.
2026-09-03 17:18:20 -07:00
Merge Sim d7e2c58549 fix(claude): let a re-walked descendant become eligible for the forced sweep
A descendant first observed by a capture inside its own birth second could never
be SIGKILLed: `ps lstart` is second-resolution, so that capture cannot rule out
a pid recycled later in the same second, and the merge pinned each retained row
to the boundary of the walk that first saw it. SIGTERM-resistant children forked
in that window were signalled and then never escalated -- they survived close,
quit and restart, reparented to init, and had to be killed by hand.

Advancing that boundary on any later capture would be unsound: a later capture
matching pid, pgid and start-second is exactly what an impostor would also show.
But a capture is not a match -- it is a fresh ppid walk from a root Node pins
through its own handle, so a row it re-derives is proved ours at that instant
without appealing to its start time. Chain the fence from there instead, and
take that walk at the close boundary while the root certainly still lives: the
root may leave inside the grace window, and the post-timeout refresh never runs.

A row absent from the later walk still keeps its earlier boundary, and a row no
walk has ever re-derived in a later second is still never escalated.
2026-09-03 17:10:43 -07:00
Neil e42c60e8a3 fix(ssh): resolve a pane's binding from the target partition, not the stale local copy (#18546)
One SSH pane accumulated one extra reattachable lease per relay restart (2, 3, 4,
5, 6 across five), and every one of them costs a `pty.attach` round trip on every
later connect, forever. Nothing prunes `sshRemotePtyLeases`, so the fan-out only
grows.

`supersedeSiblingLeasesForPane` is fenced on the PTY the pane is durably bound to,
and `durablyBoundPtyIdForPane` read `state.workspaceSession` (local) before
`workspaceSessionsByHostId['ssh:<target>']`. But `persistPtyBinding(binding, hostId)`
updates ONLY the host partition:

  AFTER-PERSIST  local= ssh:t@@pty2:old:1   host= ssh:t@@pty2:new:1

So for the length of a reconnect the local copy still names the predecessor, the
fence resolved to it, supersession took an already-`expired` lease as its winner,
and returned having marked nothing. Both partitions agree again once the renderer
republishes its layout, which is why the settled store looks consistent and hid
this.

Read both partitions as an ordered list, target's own first, and test the fence by
membership rather than by equality with whichever was read first. Pick the winner
preferring a lease this client still has a route to, since the stale partition
names an expired one. Never retire a lease that is both bound and live, so a
partition disagreement can't strand a running remote process.

Superseded predecessors stay `expired` and are never `terminated`: losing a lease
is not evidence the shell died (docs/reference/ssh-execution-boundary.md). A pane
with no binding is skipped rather than pruned, so a genuine orphan stays askable.

Also re-runs supersession from the binding side after each spawn commit's binding
write, so the lease/binding order at a call site no longer decides, and reconciles
every pane for a target immediately before `reattachKnownPtys` reads the set it
feeds to `pty.attach` — that repairs stores which already accumulated these rows.

The guard suite could not catch this: every assertion bound the pane BEFORE
upserting the lease, an order no caller uses. Rewritten to the spawn commits' real
order (lease, then binding, then the binding-side trigger); it fails 8 assertions
without this change. Added a suite that drives the real `persistPtyIpcSpawnCommit`
rather than the store primitives, including the exact stale-partition state written
by production's own binding writer.

Verified on the Docker SSH lane: five `relay.js` SIGKILLs with recovery between
each, reattachable leases flat at one per pane.

Note: this bounds the reattach SET, not the store. `sshRemotePtyLeases` still has
no cap or TTL and rows still accumulate; pruning is left alone deliberately, since
an `expired` row without `supersededBy` is a genuine orphan and must not be dropped
on age.
2026-09-03 16:47:32 -07:00
Merge Sim d4bc9448a9 fix(claude): keep command queue bookkeeping out of the transcript
Claude Code 2.1.258 emits a `command_lifecycle` frame for every uuid-stamped
command it starts, completes or cancels. The frame carries a command uuid and a
state and no content, and the CLI keeps it out of its own transcript -- but it
is absent from the SDK's SDKMessage union and so from Orca's frame catalogue,
where an uncatalogued kind defaults to a substantive row. Every structured turn
therefore painted raw JSON rows into the user-visible transcript.

Catalogue it and disposition it as status chrome. The unknown-kind default stays
`timeline-substantive`: a kind we have never seen is likelier to carry content
than to be chrome, and a visible row we can catalogue later beats content we
silently dropped. A lifecycle state that reads as a failure still surfaces,
because the payload error check in `classifyProviderFrame` outranks the
catalogue.
2026-09-03 16:31:54 -07:00
Brennan BensonandMerge Sim e85ebb0086 feat(native-chat): restore the terminal/chat switcher for bridge chat only (#18532)
* feat(native-chat): restore the terminal/chat switcher for bridge chat only

#16729 removed every user-facing terminal<->chat switching affordance as a
side effect of the structured Codex restructure ("renderer switching
affordances and their dead leftovers"). That was right for structured Codex
sessions, which render their own transcript with no live TUI underneath, but
it also took the switcher away from bridge native chat, which still reads the
terminal and has one to return to.

Restore all four surfaces, each gated so structured sessions keep the removal:

- pane header chat/terminal button (TerminalPaneHeaderOverlay)
- pane context-menu "Switch to chat/terminal view" (TerminalContextMenu)
- tab context-menu equivalent (SortableTabContextMenu)
- the keyboard chord, whose hook had survived uncalled since #16729

Gating is one rule in one place: `canSwitchNativeChatView` refuses whenever a
`structuredSessionId` is present, over the existing `canToggleNativeChat`
eligibility. Standalone structured tabs are already excluded by the
`contentType === 'terminal'` check; the new guard covers a terminal tab that
adopted a structured session. The shortcut hook applies the same rule.

The state plumbing (`viewMode`, `setTabViewMode`, `toggleTabViewMode`, host
mirroring, `native_chat_toggled` telemetry) was never removed, so this rewires
live actions rather than reintroducing logic.

SortableTab.tsx sat exactly at its 400-line cap, so its inline-rename state and
the window rename-request listener move to `use-sortable-tab-rename.ts` to make
room. No behavior change; its rename tests pass unmodified.

Two ratchets move for real, explained in place:
- store-subscription budget: per-pane listeners stay pinned at 17 (the folded
  action bundle is still one listener); only the counterfactual pre-fold
  constant grows 48 -> 49 for the added `toggleTabViewMode` key.
- hook-order parity: 204 -> 208 hooks for the four added `useCallback`s,
  useMemo count unchanged at 8.

* fix(native-chat): restore bridge chat escape hatch

* test: update pane agent identity inventory

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-03 16:28:45 -07:00
Neil 5d8532f6d3 fix(worktrees): resolve the execution host at both worktree-create entry points (#18545)
Two entry points create the same workspace and disagreed about how to read its
host. `orca-runtime-create-managed-worktree.ts:63` resolved through
`getRepoSshConnectionId` and then normalized the row; the `worktrees:create` IPC
handler branched on raw `repo.connectionId`
(`register-worktree-create-handlers.ts:66-69`). So a repo naming its owner only
as `executionHostId: 'ssh:<target>'` created remotely through the runtime and ran
`git worktree add` on the client against a remote path through IPC (#11163).
Same repo, two entry points, different answers.

Both now take one route, resolved through the existing layer
(`getRepoExecutionHostId` -> #18296's `resolveGitRouteForHost`). No new resolver.

The row normalization on the `ssh` variant is kept, and it is a **workaround, not
the pattern**. `createRemoteWorktree` and its callees re-read `repo.connectionId!`
at five depths in `ipc/worktree-remote.ts` (1627, 1847, 1848, 1865, 2029), so the
resolved connection has to reach them through the field they already read. It
travels only as far as that object does — anything downstream that re-reads the
row from the store still sees the unnormalized one, and it cannot express the
`runtime:` refusal on its own. Proper fix, deliberately not done here: give that
pipeline an explicit connection parameter and delete `repo.connectionId!` from it
so every reader becomes a compile error, the technique #18307/#18325 used. That
is a change inside a 2800-line module plus its callers, and it wants its own PR.

Three answers that used to collapse into one, now distinct at both entry points:

- `executionHostId: 'ssh:*'` with no `connectionId` -> that SSH host (IPC used to
  create locally);
- `executionHostId: 'local'` with a surviving `connectionId` -> local, since a
  local row cannot nest an SSH namespace. This is what `getRepoSshConnectionId`
  and therefore the runtime sibling already answered; IPC used to go remote;
- `runtime:<env>` -> refused. Its worktree is created by that environment's own
  server and the SSH target on its repo row is that server's nested one,
  addressable only as (environmentId, targetId). The renderer already routes
  runtime-environment creates over `worktree.create` RPC rather than this IPC
  channel, so reaching either entry point with one is a routing mistake. Matches
  `workspace-cleanup-git-route` and `runtime-git-command-target`.

Folder-workspace creation is untouched on both sides: it is a registration, not a
filesystem create, so the route is resolved after that branch on the IPC side, and
on the runtime side only the agent trust write consumes it — where a `runtime:`
host now yields `null` instead of the nested target, so the write stops going to a
same-named target in this client's table.

No wire or persistence change: the normalized row is a local value passed to the
create pipeline, never stored, and `CreateWorktreeResult` is untouched.
2026-09-03 16:27:13 -07:00
Neil 3c91404820 fix(worktrees): route managed worktree removal by resolved execution host (#18529)
`removeManagedWorktree` resolved its host once — for the metadata prune
(`cleanupHostId ?? getRepoExecutionHostId(repo)`) — and then read raw
`repo.connectionId` for every step that touches the filesystem: the
`git worktree list` deciding whether the path is registered, the provider handed
to the unregistered-removal branch, the registered-remote-vs-local fork, and the
PTY/history teardown. One function, two spellings.

For a row naming its owner only as `executionHostId: 'ssh:<target>'` — the exact
class #18296 names — the list ran on the client against a remote path,
`removeRuntimeUnregisteredWorktree` was entered with `provider: null`, and the
metadata was pruned under `ssh:<target>` while a same-named *local* directory was
the one considered for deletion (#11163). #18358 made this reachable: it migrated
the cleanup scan, so `executionHostId`-only rows now surface as removable
candidates, but removal did not move with it.

Routing is now one answer for the whole removal, taken from the host the prune
already used, through #18296's host-keyed dispatch. The ambiguous
`provider: SshGitProvider | null` carrier is deleted from the callees rather than
supplemented, so every remaining reader is a compile error in the typed modules
that do the destructive work (`runtime-unregistered-worktree-removal`,
`runtime-registered-remote-worktree-removal`, `runtime-worktree-filesystem`).
The orchestrator itself carries `@ts-nocheck` from its mechanical split, so that
guarantee does not reach it — tests cover it instead.

Every change is in the refusing direction; nothing became more aggressive:

- an `ssh:` host with no registered provider throws instead of deleting a
  client-side path (`requireSshGitProvider` already threw for rows that spelled
  the same host as `connectionId`);
- `runtime:<env>` throws rather than dialling a same-named target in this
  client's namespace, matching `workspace-cleanup-git-route` and
  `runtime-git-command-target`;
- the folder-workspace teardown resolves its connection instead of reading the
  raw field, so a `runtime:` row stops dialling the wrong namespace.

No wire or persistence change: `removeWorktreeMetadataAndHistory` already took
the resolved host, and the removal RPC result shape is untouched.
2026-09-03 16:21:52 -07:00
Neil 04ae62202a fix(ssh): close the macOS relay's per-terminal pty fd leak (#18534)
The relay asset from #17920 only rewrote the forkpty `default:` call site, which
sits in the `#else` arm of PtyFork's `#if defined(__APPLE__)`. macOS takes
`pty_posix_spawn`, so the asset had never patched anything a Mac executes -- and
`applyNodePtyMasterCloexecPatch` returned 'fixed' for any non-Linux host without
running the script at all, which is what publishes a tree to the shared
native-deps cache.

Stock `pty_posix_spawn` opens up to three throwaway ptys to push the real master
off fds 0-2 and never closes them: the cleanup loop is `for (; count > 0;
count--)`, but the first `posix_openpt()` in a running process already returns
>= 2, so it breaks with `count == 0` and the body never runs -- and where it does
run it closes `low_fds[count]`, never `low_fds[0]`. One orphaned /dev/ptmx fd per
terminal, for the life of the relay.

Ports the `low_fds` fix and the Apple-branch `pty_cloexec(master)` call from the
app's `config/patches/node-pty@1.1.0.patch`, byte-identical, and runs the gate on
darwin. macOS needs a different build layout than Linux: it has no `build/` at
all, so the fallback moved aside is `prebuilds/darwin-<arch>` -- which is also
what makes node-pty's install script fall through from "prebuild found" to
node-gyp -- and the compile writes a `build/Release` the loader checks first.
Verification is per-platform too: Linux's leak is inheritance (/proc), macOS's is
self-held (lsof).

Also corrects the asset's claim that "macOS re-opens the tty through uv_tty_init's
cloexec dup". Measured false: FD_CLOEXEC is not set on the master. What protects
it is POSIX_SPAWN_CLOEXEC_DEFAULT, one option away from gone since uid/gid drops
libuv back to fork()/exec() -- so the master is now marked there too.

Measured on darwin-arm64, one PTY per open/close cycle in a relay-shaped dir
running the relay's own commands:

  before  cycle:ptmx  1:1 2:2 3:3 ... 10:10   (10 after a settle)
  after   cycle:ptmx  1:0 2:0 3:0 ... 10:0    (0 after a settle)

Linux re-verified in docker node:22: inherited before, isolated after,
`already-patched` on the second run.

Refs #17915
Refs #8362
2026-09-03 16:13:07 -07:00
Neil a5d6114baf fix(ssh): stop pane adoption certifying a death from the relay's not-found union (#18531)
* fix(ssh): stop pane adoption certifying a death from the relay's not-found union

`attachStablePaneOwner` was the last reader that synthesised a runtime exit
from a reattach refusal, and it published code `0` — which
`orca-runtime-on-pty-exit` records as `rememberPtyLivenessVerdict(exited)`, a
death certificate whose only legitimate writer is a host-delivered exit frame.

The refusal it acted on is a union. `pty.attach` answers `PTY "<id>" not found`
both for a pid the relay probed with `isProcessAlive` and for an id its session
map simply never had — which, because ids carry a per-start mint epoch, is every
id minted before a relay restart, checked against nothing. So a relay restart
plus a reconnect certified a shell that was still running under the old daemon's
orphaned process tree, retired the pane binding, and cold-started a second agent
onto the same transcript. The sibling `handlePtyReattachFailure` has always
refused to certify from that union; this path did not.

- The relay marks the one refusal it backed with a liveness check
  (`PTY_ATTACH_PROVEN_EXITED_MARKER`). The marker is additive, so an unmarked
  answer — including an older relay's — stays ambiguous, which is the safe
  direction.
- The client mints that half as `SshPtyProvenExitedOnRelayError`, a subclass so
  every existing `isSshPtyAbsentFromRelayError` consumer is unchanged.
- Pane adoption publishes `UNVERIFIED_PROCESS_EXIT_CODE` (-1), the sentinel its
  sibling publishes, and passes `hostExitConfirmed` only for evidence that
  observed the process: the marked relay refusal, or `SessionNotFoundError` from
  the registry that owns the PTY. The ambiguous half now records `unverifiable`
  instead of `exited`.
- The gone-branch keys on the error type rather than the bare `PTY ".+" not
  found` text, so an untyped string can no longer authorise abandoning a
  binding — the discriminator `pty-connect-limits.ts` already documented.

Refs docs/reference/ssh-execution-boundary.md

* test(pty): make the pane-adoption fixtures throw what real providers throw

These four fixtures rejected with bare `new Error('Session not found: ...')` and
`new Error('PTY "..." not found')`. No provider produces either untyped:
`local-pty-spawn` and `decodeDaemonResponseError` both mint
`SessionNotFoundError`, and the SSH reattach path types the relay's wire text
before any pane sees it. Fixtures that skip the type were the reason a
message-shaped gate looked adequate.

The exit-code expectations move with it: the pane path now publishes the -1
stop sentinel plus `hostExitConfirmed`, so a certificate follows the evidence
rather than a synthesized zero.
2026-09-03 16:13:03 -07:00
Jinwoo Hong 7d27c841b4 fix(cloud): run the rehome control job under pipefail (#18537)
The five `node ... | tee` steps in cloud-operate-relay-production-rehome-job.yml
reported tee's exit code, so a thrown inspect or apply passed green. The Aug 28
21:25Z and Aug 29 inspects and today's first inspect all printed
"director returned an invalid regional rehome control" (the durable control had
moved to generation 12 when the Aug 28 rehome aborted) and still succeeded.
`shell: bash` adds `-o pipefail`. A test pins the default and the tee count.
2026-09-03 18:38:33 -04:00
Merge Sim 33f2293467 Merge commit '8cd63706e245f1450b5c5b31afb28f7fd75e1d2a' into brennanb2025/claude-sdk-ready-mainsync 2026-09-03 15:28:56 -07:00
Merge Sim 8cd63706e2 Pin the absent-vs-empty distinction in the managed-account gate
An empty claudeManagedAccounts array is a real answer: the user has no managed
accounts, nothing claims an identity, and the ambient path is legitimate. A
readable settings object with no such field is settings we failed to parse —
the same unknown as unreadable — so it refuses.

The two are one character apart in the code and the difference is invisible
without the reasoning, so record it at the branch and pin both sides. The test
fails under the obvious "consistency fix" of treating a missing field as empty.
2026-09-03 15:21:34 -07:00
Merge Sim 8a897b9911 Merge commit '6b038e114bd90f8e863ecdde56cb550dc50839a7' into brennanb2025/claude-sdk-ready-mainsync
# Conflicts:
#	src/main/runtime/structured-agent-session-runtime.ts
2026-09-03 15:20:27 -07:00
Brennan BensonandMerge Sim 98e77ef1a7 feat(mobile): structured native Codex chat (#18074)
* feat(mobile): finalize structured native Codex chat

* fix(mobile): close structured chat lifecycle gaps

* wip(mobile): fence stale structured inventory and bound operation-id retention

Fence local structured-session inventory and subscription responses with a
sync generation so a toggle-off clear, reconnect restore, or retry cannot
apply a mirror from a superseded instance. Bound mobile ambiguous
operation-ID retention at 128 with unmount cleanup.

Staged on the reconcile branch only: the sync module is now 312 lines and
needs a real split before this can reach the PR head.

* fix(ci): split the structured session-tabs sync and give static analysis mobile types

The local structured session-tabs sync module outgrew the 300-line cap once it
took on generation fencing, so split it along its real seams instead of raising
the cap: the generation/cursor fence, snapshot projection, snapshot apply,
inventory refresh, and the subscription loop. The original path stays as a
barrel so no importer moves.

Repoint the host-session-mirror settle census at the apply module, which owns
two receipts now — the snapshot it mirrors in, and the toggle-off teardown that
retracts what it published. The teardown receipt is named rather than anonymous
so the pin says which direction it settles.

The changed-code quality gate lints mobile files and resolves their types from
mobile/node_modules, but mobile is a separate pnpm project that the root install
never populates, so every mobile type degraded to an `error` type and the gate
reported phantom findings. Install mobile dependencies in static analysis when
the diff touches mobile, gated on a new classifier output.

* fix(mobile): let a slow capability handshake still reach connected

The mobile capability update is an advisory whose result is discarded, yet an
unanswered one was fatal while an explicit rejection was tolerated. A 5s timeout
on the direct client force-closed the socket, and on the relay path it failed
`confirmResume` before `connected` was ever published, so a consistently slow
link redialled forever. Both paths now share one helper that settles every
ambiguous outcome (timeout, mid-flight drop) like a rejection and rejects only
when the frame never reached the wire — the one case nothing else recovers from,
since the socket's own desync force-close is gated on already being connected.
The generation guard still keeps a replaced session from connecting.

Retained structured-session operation ids were capped at 128 with oldest-first
eviction, but every retained id belongs to a send whose outcome is unknown, so
eviction turned a user's retry into a second message on the host. Bound the map
by expiry against the id's own embedded timestamp instead, mirroring the host's
operation ledger, so no id is released while the host would still honour it.

Also give the mobile CI install the root install's lockfile drift guard (mobile's
lockfile carries patchedDependencies a silent rewrite would drop), gate
mobile_dependencies on should_run, and key the pnpm store cache on both lockfiles.

* refactor(mobile): extract the relay pending-request registry

The merge composed two independently-sized changes — this branch's capability
handshake settle and main's dial-stage tracking — pushing the relay session file
to 304 lines against a 300 cap. Neither side broke it alone.

Move the in-flight request registry (id generation, tracking, settlement, and
reject-all with its delivery-ambiguity marking) into RelayPendingRequests,
matching the existing collaborator pattern alongside RelayDialStageTracker and
RpcSessionLivenessWatchdog. No behavior change.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-03 15:19:26 -07:00
Merge Sim 585bf17a1f Derive the gate test's auth policy from the settings under test
A hardcoded stripAuthEnv asserts a gate/policy pairing production cannot
produce, and false additionally lets launch.env inherit the runner's real
process.env. Derive via claudeStructuredAuthPolicyForSettings instead: the
gate settings type is the same Pick the policy takes, and both resolve the
account through getSelectedClaudeAccountIdForTarget.
2026-09-03 15:18:32 -07:00
Merge Sim 6b038e114b Move the structured Claude gate out of the @ts-nocheck runtime files
Both call sites of the managed-account gate sat in files whose first line is
`// @ts-nocheck`, so neither was typechecked: three arguments to a one-argument
function plus an undeclared identifier compiled clean. New auth-identity
decision logic had no compiler behind it.

Move the verdict into a checked module that takes the two facts the runtime
owns — the adapter's answer and a settings getter — and decides. The runtime
class now only forwards. Move the gate reader's construction into the checked
installer too, so the nocheck file passes a plain settings closure and never
names a gate symbol.

Every reference to the gate predicate and its reader now lives in a checked
file, so the ablation that used to pass silently is a compile error at both the
create-support and reacquire sites.

Removing the file-level @ts-nocheck is a separate, larger job and is not
attempted here.
2026-09-03 15:14:24 -07:00
mors c79e1c097b fix(i18n): localize remaining onboarding UI
Reviewed and approved by Codex.
2026-09-03 15:07:42 -07:00
韦编三绝 f13f2472c6 fix(i18n): distinguish Duplicate from Copy in Simplified Chinese
Reviewed and approved by Codex.
2026-09-03 15:07:38 -07:00
Ilya Gusev 8262fb147f fix(i18n): extract translateSearchKeyword calls so settings-search keywords reach en.json
Reviewed and approved by Codex.
2026-09-03 15:07:34 -07:00
Jinjing b1186c6beb Fix scope of workspace-creation-project tour target (#18502)
* fix: scope workspace-creation-project tour target to project picker only

The tour target was previously applied to a container that included both
the project picker and the run target picker below it. Restructure the
layout to scope the target to only the project-related section, and add
a test to verify the tour target does not span into the run target picker.

* fix: scope workspace-creation-project tour target to project picker only

Move the tour target attribute from the outer project section to an inner
wrapper around just the combobox and its messages, excluding the header
label and "Add project" button. Update tests to verify the narrower scope.
2026-09-03 14:45:58 -07:00
Neil f35015d0c8 fix(ssh): measure pane idleness in the unit the sweep's kill operates on (#18415)
The orphan-relay-PTY sweep authorizes `pty.shutdown { immediate: true }`, which
runs `forceKillPosixPtyProcessGroups`: collect every process group on the pane's
tty, then `killpg` each one. The blast radius is therefore (groups on the tty) x
(members of those groups, wherever they are). The idleness evidence measured only
the first factor, so three shapes read as idle and were SIGKILLed:

- with job control off (`set +m`) a background job keeps the SHELL's pgid, so the
  tty carries exactly one process group and that group is running the user's build;
- a child that drops the controlling terminal (`ioctl(TIOCNOTTY)` without `setsid`)
  keeps the pgid, reports `tpgid == -1`, and never appears in `ps -t <tty>`;
- a double-forked grandchild keeps the pgid and tty but reparents to pid 1, so the
  `ppid` walk cannot reach it and the named-process backstop never fires.

`shellOwnsEveryTtyProcessGroup` now also requires the shell's own process group to
hold no other member anywhere in the table, indexed in the same single pass. A
pids-per-tty set would catch the first and third but not the second, which is why
the count is pgid-wide rather than tty-scoped. The wire field keeps its tty-shaped
name: the value only ever became stricter, so an old client skips more, never less.

Second, unrelated-in-mechanism but same file family: `foregroundSkipReason` summed
`capturedAgeMs + evidenceAgeSinceListingMs` without validating either. A non-numeric
`capturedAgeMs` makes the sum `NaN`, and `NaN > 5000` is false, so a malformed record
PASSED the freshness gate and proceeded toward the stop — the one place in the file
that defaulted toward kill. Nothing validated it on this path
(`mapSshPtyProcessList` checks the ownership fields and spreads the rest through;
`PtyProcessListAdmission` is not on the sweep path). It now runs
`isForegroundProcessEvidence` and fails closed.

Verified on real Linux, not only in mocks: a container drives `bash -i` on a real
pty, builds each construction, runs the real publisher and planner, and then calls
the real `forceKillPosixPtyProcessGroups`. Before, all three published
`shellOwnsEveryTtyProcessGroup: true`, planned SWEEP, and the planted pid was gone
after the signal. After, all three skip and survive, and an idle shell is still
reclaimed.

Residuals are written down at the predicate and in ssh-execution-boundary.md: the
capture is a snapshot (bounded by the evidence-age budget, not removed), and a
process the host's own `ps` cannot enumerate stays unobservable while `killpg`
still reaches it.
2026-09-03 14:44:32 -07:00
Jinwoo Hong f974e98162 test(cloud): derive both reachability directions for the relay inventory census (#18524)
Mirrors stablyai/orca-cloud#472 (c82f98f), byte-identical under cloud/.
2026-09-03 17:43:40 -04:00
Neil 95eed52801 fix(cli): report which hosts a worktree listing covered, and stop the cap starving remote ones (#18417)
`orca worktree list` returned zero of 24 SSH worktrees at the default limit
(#18104). Rows are resolved repo by repo, so every SSH repo's rows land
contiguously at the end of the fleet order — the 24 remote rows sat at indices
496-520 of 521 and a plain `slice(0, 200)` never reached them.

The omission was not fully silent: text output printed `truncated: showing 200
of 521` and JSON carried `totalCount` / `truncated`. What was missing is that
the omission was *categorically every remote host* — no host column, no
`hostScope`, nothing to distinguish "200 of 521" from "one host is entirely
absent". Per docs/reference/ssh-execution-boundary.md, a listing that does not
name its scope reads as absolute.

Adopt the mechanism `terminal list` already has rather than inventing a second
one:

- `RuntimeTerminalListHostScope` becomes an alias of a shared
  `RuntimeListingHostScope`, now also carried (optional, so old hosts are
  unaffected) on `worktree.list` and `worktree.ps` results.
- `src/shared/host-balanced-listing-page.ts` round-robins the row cap across
  hosts and returns the survivors in the caller's original relative order, so
  the page stays a subsequence of the unbounded listing and nothing downstream
  re-sorts. An uncapped listing is returned unchanged.
- `worktree list` / `worktree ps` text output gains a `host=` column and the
  same trailing `scope:` line `terminal list` prints.

Third defect, same mechanism: `hostScope.omittedHostIds` is built from the
runtime's own bookkeeping, so it names `runtime:` ids for servers that are no
longer paired — 6 of 9 in the recorded QA run hard-error when queried. Since
`hostScope` is *the* documented way to complete a partial listing, that makes
the mechanism unreliable for its intended use.

Annotate rather than filter. Dropping an id would shrink what the listing
admits it did not cover, and the boundary doc requires a listing to name its
gaps — the gap is real whether or not this machine can name the host that owns
it. `src/cli/omitted-host-scope-selectors.ts` resolves each omitted id against
this machine's pairing store and the runtime's SSH-target registry and attaches
the exact flag that reaches it, or `null` marked "not selectable from this
machine". This is a client-side annotation: nothing new goes over the wire, it
answers "can I select it" and never "is it up", and the SSH round trip is only
paid when an `ssh:` host was actually omitted.

No `--host` filter was added; the host column plus scope line covers the
reported need without a new selector axis.
2026-09-03 14:43:09 -07:00
Neil 9bed758e36 fix(cli): reject runtime selectors on host list and environment list (#18405)
`orca host list --environment m4air` was not ignoring the flag — it was applying
it to half the answer. `shouldIgnoreRemoteSelection` never pinned the `host`
family, so the SSH-target lookup was routed to m4air while paired servers were
still read from this machine's own pairing store, and the handler stamped the
envelope `_meta.runtimeId: "local"` regardless. The result was one listing
describing two hosts: the openclaw row silently disappeared, which reads as
"m4air has no SSH targets". `environment list --environment X` had the pin but
no guard, so the flag vanished with no signal at all.

Reject rather than route. `host list` answers "what can this machine target and
with what flag"; its paired-server half comes from a client-local store and
cannot be routed at all, so any routed answer is necessarily half-substituted —
rule 1 of docs/reference/ssh-execution-boundary.md. `environment list` is
entirely client-local, so there is no other host to ask. This matches the
`account` and `artifacts` precedent, the only two pinned families that already
paired the pin with a rejection guard.

- pin the `host` family so an ambient ORCA_ENVIRONMENT cannot produce the same
  two-machine listing with no flag to reject; `runtimeId: "local"` is now true
- extract the duplicated `rejectRemoteSelectionFlags` from account.ts and
  artifacts.ts into src/cli/remote-selection-flag-rejection.ts
- `environment show` / `environment rm` / `environment add` are untouched: there
  `--environment` and `--pairing-code` name the row to act on, not a route
2026-09-03 14:43:05 -07:00
Neil 232d04f541 fix(dashboard): open remote sessions from every agent reveal path (#18403)
Three reveal paths called bare setActiveWorktree + activateTabAndFocusPane,
skipping setActiveView('terminal'), ensureWorktreeHasInitialTerminal and
resumeSleepingAgentSessionsForWorktree. A parked SSH workspace has no resident
tab until those run, so the reveal landed on a workspace with no terminal.

Route all three through the incumbent activateAndRevealWorkspace dispatcher
(which the sidebar and "Jump to workspace" already use, and which also handles
folder workspaces). The Activity row-click additionally early-returned when the
thread's tab was absent from tabsByWorktree/unifiedTabsByWorktree, which made a
cold-parked remote thread a silent no-op; residency is now probed after
activation, so a revived tab is focused and a genuinely retained thread still
activates its workspace instead of doing nothing.

Also stop asserting `exited` from an absence of local state: SshPtyProvider
reports no authoritative buffer snapshot and the relay has no snapshot RPC, so
a null preview snapshot for a remote pty is loss of contact. The preview and
the no-pty dialog branch now say the remote preview is unavailable rather than
claiming the pane closed. Adding the relay snapshot RPC stays out of scope --
it needs capability negotiation.

Fixes #16731
2026-09-03 14:43:01 -07:00
Neil 5a626dcdf4 refactor(git): share push-target resolution between local and the SSH relay (#18406)
`src/relay/git-handler-push-target.ts` and `src/main/git/remote.ts` carried
identical ~160-line copies of the resolver that decides which remote a plain
`git push` hits. Identical today is exactly when to share it: the cost of a
future divergence is pushing to the wrong remote, which retrying does not undo.

Move the resolver to src/shared/git-push-target-resolution.ts, parameterized on
a `(args) => Promise<{ stdout }>` runner — the only thing the two hosts actually
differ in — and delete both copies. The relay entry point keeps only the work
that is genuinely relay-side: re-validating an explicit target that arrived over
the wire and running `check-ref-format` on it.

No behavior change on either path, and nothing new or different is published, so
this engages no rule in remote-wire-compatibility. No git command changes.

src/relay/git-push-target-local-parity.test.ts scripts one repository's config
and requires `git.push` over the real relay dispatcher and the desktop's
`gitPush` to emit the same push argv, plus the argv each case should produce.
2026-09-03 14:42:58 -07:00
Neil 53adf5e2e6 fix(git): share one failed-command error-text reader between local and the SSH relay (#18398)
* fix(git): share one error-text reader between the local and relay branch-delete fallbacks

The relay and the desktop each carried their own `getErrorText`, and they had
drifted: the relay read `message` + `stderr` + `stdout`, the desktop only
`message` + `stderr`. A `git branch -d` refusal arriving on `stdout` therefore
routed the SSH removal through prune-and-retry while the local removal gave up
and preserved the branch.

Against a real binary the two agree, because Git prints the refusal through
`error()` on every supported version — verified on 2.25.1, 2.38.1, 2.49.1 and
2.55.0, none of which put a byte of it on stdout. What the desktop copy actually
missed is that Orca classifies errors it built itself, with the Git output on
`.stdout`: `worktree remove`'s submodule retry attaches `git status --porcelain`
that way on both paths. The stdout-reading form is also already the shared
spelling — `isSubmoduleWorktreeRemovalRefusal` uses it for both hosts — so this
converges on it rather than on the shorter one.

Move the reader to src/shared/git-command-failure-text.ts and the predicate it
feeds to src/shared/git-branch-delete-refusal.ts, and delete all three copies.
The predicate carries both refusal wordings live in the supported range: Git
through 2.40 says "checked out at", 2.43+ says "used by worktree at".

The real-binary contract now pins that boundary: the refusal is recognized, it
lands on stderr, and stdout stays empty on every Git in the matrix.

* fix(test): consolidate the duplicate worktree import in the parity test
2026-09-03 14:42:54 -07:00
Merge Sim eba728d1d6 Merge commit '1d7a3938a1' into brennanb2025/claude-sdk-ready-mainsync
# Conflicts:
#	src/main/claude/claude-structured-launch-resolution.ts
#	src/main/runtime/orca-runtime-get-worktree-ps.ts
#	src/main/runtime/structured-agent-session-runtime.ts
#	src/main/runtime/structured-claude-runtime-adapter.ts
2026-09-03 14:38:21 -07:00
Merge Sim be45b351a8 Merge commit 'a632ea3353' into brennanb2025/claude-sdk-ready-mainsync 2026-09-03 14:34:07 -07:00
Merge Sim 1d7a3938a1 Run the managed-account gate on every Claude acquisition, not just create
createSupport gates the create path, but a session's account state can change
while it lives. A reacquire after an unexpected child exit re-resolves the
launch and re-derives auth, with nothing re-checking the gate — so a session
created while supported could come back up in the refused shape. With the strip
predicate keyed on there being an active non-WSL account, the WSL-only user's
normalized steady state (accounts exist, none active) does not strip, and that
reacquire reaches the child with ambient auth while the UI names the account.

Gate at resolveLaunch, the one choke point every acquisition passes through,
refusing with the pre-spawn error the caller already handles. Same predicate as
create-time, now sharing one settings reader so the two cannot drift.

Claude only; Codex resolves its account on a different path and is untouched.

The runtime class that wires this does not typecheck its own `this` calls — a
missing hookup compiles clean — so the wiring is pinned behaviourally rather
than trusted to the compiler.
2026-09-03 14:30:46 -07:00
Jinwoo Hong d66386bc82 fix(cloud): bound and yield the relay's global cell-inventory lock (#18521)
Mirrors stablyai/orca-cloud#471 (squash c3354e8), byte-identical under cloud/.

The relay's cell-inventory lock (SELECT ... FROM relay_cells FOR UPDATE over
all 23 rows) is one global critical section shared by the assignment hot path
and every director sweep; with the pool's 1s lock_timeout a blocked waiter held
a pooled client for a full second, producing ~690 55P03 retries per 5 minutes
in production. Request paths now bound the wait at 500ms with a SET LOCAL that
is restored to the pool default before the next statement; director-only sweeps
take the lock NOWAIT and skip the tick; sweep timers are jittered; hold time is
exported as additive runtime-metrics fields so the bound can be tuned.
2026-09-03 17:27:57 -04:00