Commit Graph
2 Commits
Author SHA1 Message Date
Neil 7c94d12190 fix(ssh): route four host-blind seams through the resolved execution host (#17919)
* fix(host-routing): resolve the execution host before reading a connection

Three issues in one defect class: a resolver reads one spelling of one
arbitrarily chosen row instead of resolving the worktree's execution host,
so something local answers a question about a remote.

returned that row's connectionId. With duplicate repo rows for one repo id
it could pair a runtime owner with a client-owned SSH connection. It now
resolves through the same ambiguity-aware index getRuntimeEnvironmentIdForWorktree
uses, prefers the repo row for the host the worktree names, and derives the
connection from the resolved host. Conflicting rows return `undefined`
(this module's documented "cannot determine the host"), never `null`.

`store.getRepo(worktree.repoId)?.connectionId ?? null`. `getRepo` is
host-blind and the same repo id can exist on local, SSH and runtime hosts,
so a remote worktree could spawn its PTY on the client with the remote cwd.
resolveWorktreeLaunchHost picks the row for the worktree's host and reads
the connection off that host; conflicting rows are unresolved, not local.

session-partition owner maps that contradict each other. Both now compute
through one shared function whose argument records the divergence. No
behaviour change on either side: converging needs a read-both migration,
since both partitions hold real data written by shipping builds.

* fix(host-routing): keep nested SSH connections resolvable under a runtime host

getRepoSshConnectionId read only the resolved execution host, so a repo row
owned by a runtime that reaches a nested SSH target (connectionId: ssh-*,
executionHostId: runtime:*) resolved to no connection — answering 'local' for
a remote worktree, the same defect #17909 fixed in the other direction.

* fix(host-routing): resolve both sides of the execution host through one rule

The renderer resolver leaked between two different SSH hosts: a worktree on
`ssh:m4air` whose only indexed repo row belonged to `openclaw` answered
'openclaw', because the host-scoped lookup missing fell through to an id-only
one. Main's resolver, in the same change, answered 'm4air' — two resolvers, one
right and one wrong, on identical input.

Both sides now adapt one shared rule (`worktree-execution-host-resolution.ts`):
the worktree's own host outranks every repo row, and a row on a different host
is never evidence about this one. The renderer's WeakMap index becomes the
memoizing adapter it always was; `resolveWorktreeLaunchHost` becomes main's
mapping of unresolved onto its throw.

Settles the rule the change previously answered two ways.
`getRepoSshConnectionId` and `getSshTargetIdForExecutionHost` disagreed for a
runtime host carrying a nested `connectionId`; they now compose, so the
execution host is the single authority. On a `runtime:*` row that field is a
paired HUB's private SSH target, spread through by `repoWithFetchedOwner` and
unaddressable from this client — the project-first successor of the row nulls it
for exactly that reason. That also fixes the `kind !== 'ssh'` fallback, which
fired for `local`: a row declaring itself local handed out an SSH connection.

* fix(ssh): resolve the execution host in the worktree scan and managed create

The worktree scan and createManagedWorktree both picked remote-vs-local from
repo.connectionId, so a row stamped only executionHostId: 'ssh:*' was scanned
and created on the client against a remote path. The folder branch returns
before the check, so its agent-trust write landed locally too.

Refs #11163

* fix(ssh): stop over-rejecting and refusing SSH hosts the process owns

runtimeRepoMatchesExecutionHost rejected an unstamped SSH repo against its own
ssh:<connectionId>, so repo-add/clone dedupe could register a second row for a
path the host already owns. assertHostIsSupported made the CLI/runtime RPC
refuse --host ssh:* while the same process's IPC handler routed it correctly;
setupExistingFolder now shares that registration. Clone still refuses, because
nothing in this process clones onto an SSH host.

Refs #11163

* test(ssh): retarget the SSH host-setup guard spec at the substitution it prevents

setupProjectExistingFolder now registers the remote path through the same
addRemoteRepoFromPath the desktop IPC uses, so it fails on the host's terms
(connection not registered) rather than a categorical refusal. The local
clone/probe side effects it exists to catch are still asserted absent.

Refs #11163

* fix(cli): require an absolute path when setting a project up on an SSH host

Routing --host ssh:* to the remote registration made relative paths newly
reachable there, and they were resolved against the client cwd — registering a
path that names the wrong machine.

Refs #11163

* fix(repos): read the SSH registry directly so the runtime stays Node-bootable

Routing runtime project setup through addRemoteRepoFromPath dragged ipc/ssh --
and its 25-module electron graph -- into the runtime bundle. ssh-target-registry
already exists for exactly this; ipc/ssh only re-exports it.

* fix(ssh): close the agent-launch and session-export host-blind twins

Three sites left on the legacy spelling, all the same shape as the ones this
branch already fixed:

- `launchAgentTerminal` did `getRepo(worktree.repoId)` then wrote agent trust
  with that row's `connectionId`. Host-blind, so a repo id carried by two SSH
  hosts wrote a remote path into the *client's* Codex/Cursor/Copilot config and
  the agent on the host never saw the trust. Every sibling call site already
  passes the resolved `workspace.connectionId`; this was the last that did not.
- `targetForWorktree` (workspace-session export) fell back to the same
  host-blind read, so a session could be published to a machine that never
  owned the worktree. Unresolvable ownership now exports to nobody.
- `addRemoteRepoFromPath` minted `connectionId`-only rows while being the
  routing path this branch adds, so it kept creating rows in exactly the
  spelling the branch works around. It now stamps
  `toSshExecutionHostId(connectionId)` at creation; `reassignSshTargetId`
  already migrates both spellings, so target rename stays correct.

Tests cover two *different* SSH hosts throughout — the case none of the earlier
duplicate-row tests had, all of which were local-vs-ssh or runtime-vs-ssh.
2026-09-02 16:59:29 -07:00
Neil a5796ec8eb refactor(runtime): split OrcaRuntimeService and compatibility tests (#17605)
* refactor(runtime): split OrcaRuntimeService into focused modules

* test(runtime): cover admission tiers and strict worktree reconciliation

* fix(runtime): preserve owner and structured session visibility

* fix(runtime): port post-extraction compatibility fixes

* fix(runtime): preserve skill-share cancellation barrier

* test(runtime): update identity inventory after extraction

* fix(runtime): preserve hook transport environment cleanup

* fix(runtime): consolidate idle probe imports

* test(runtime): retire split file process allowlist entry

* fix(runtime): route child process types through shared boundary

* test(runtime): preserve worktree host metadata precedence

* fix(runtime): update extracted test seams

* fix(runtime): gate the split's ts-nocheck set and restore the stop-confirmed contract

Audit follow-ups for the OrcaRuntimeService split:

- Freeze the 171 @ts-nocheck files behind a ratchet so no new file can disable
  type checking. The split's linear mixin chain cannot express forward
  references yet, so the existing suppressions are grandfathered; the baseline
  may only shrink.
- Drop the stray @ts-nocheck at the end of orca-runtime-get-status.ts. It sat
  after the first statement, where TypeScript ignores it, so the module was
  already checked.
- Restore `retireRejectedPty(ptyId, stopConfirmed: boolean)` as a required
  argument. The split widened it to optional and patched the resulting error
  with `stopConfirmed === true`; an omitted argument would have silently taken
  the unverified-stop path instead of failing to compile.
- Guard that every orca-runtime-tests fragment is imported by the compatibility
  entrypoint. The fragments are .spec.ts, which no Vitest include glob matches,
  so one left out of the list would silently stop running.

* fix(runtime): restore four behaviors the OrcaRuntimeService split dropped

Audit findings against the refactor's true base (ad5ba2572e):

- retirePtyAgentLaunchAuthority collected pane keys after deleting the
  restored-authority receipt instead of before it. collectPaneKeysForPty reads
  that receipt, so a receipt-only pane lost its key and never had its agent-hook
  compatibility authority retired. on-pty-exit.ts already carried a comment
  naming this exact invariant.
- The PTY-exit path kept orchestrationMailboxNotifications.retirePty but lost
  the loop that schedules a debounced mail-pointer repoint for the dead pty's
  terminal handle and any run bound to its panes. Restores the schedule call
  count to 7, matching base.
- subscribeToPtyExit lost isPtyKnownExited's leaf fallback and its
  post-registration lifecycle-generation recheck. leavesByPtyId is rebuilt from
  the renderer graph independently of ptysById, so a leaf can outlive its pty
  record; without the fallback a caller waiting on an already-dead pty never
  gets released.
- The chain root declared `[key: string]: unknown`, which base had nowhere. It
  leaked through the exported runtime type into every consumer, so any misspelled
  member access typechecked as unknown instead of erroring, and it accounted for
  957 of the suppressed errors. Removing it costs zero type errors.

* fix(runtime): restore escalation prose and unscoped automation publication

Two more behaviors the split dropped, each with a regression test that fails
against the pre-fix code:

- The worker-exit escalation stopped deriving its title through
  buildOrchestrationTaskDisplayMetadata and inlined `task.spec` instead. That
  ignored an explicit task_title, dropped the single-line normalization and the
  80-character bound, and turned the no-spec case into a quoted, duplicated id.
  A multi-paragraph spec landed verbatim in the coordinator's banner. The
  existing 11 tests all use short single-line specs, where the derived title and
  the raw spec are identical, so none of them could see it.
  Also reverts an added `if (!handle) return` guard: the dispatch lookup is
  deliberately keyed on the pane as well, because a reminted handle no longer
  matches the row while the pane identity outlives the remint.
- updateAutomation stopped going through automationChangePublications and
  published `source` unconditionally while gating the fallback on a non-null
  destination. A destination the store can no longer name then published only
  the stale source, so subscribers scoped elsewhere kept rendering a row that
  had left them — the exact case the helper documents. The helper had been left
  with zero callers; all three sites use it again.

* fix(skills): stop swallowing lookup errors and hard-erroring on non-ssh hosts

Follow-ups from auditing the skill install path against the refactor's base:

- resolveWorktree wrapped showManagedWorktree in `.catch(() => null)`, so a
  transient git or IO failure surfaced to the user as
  skill-install-workspace-not-found with the real cause discarded. Errors
  propagate again; a genuine id mismatch still returns null.
- resolveSkillSshTarget threw skill-install-workspace-host-unavailable when the
  execution host was neither local nor ssh, on both the repo and folder
  branches. Base gated these on connectionId, so a runtime-owned repo simply
  was not an SSH install and fell through to the local path. Both return null
  again, and the error code the split invented is now unreferenced.
- listManagedSkillInstalls awaited the receipt walk and the worktree resolve in
  sequence. They are independent and either can hit disk, WSL, or an SSH scan,
  so Promise.all is restored.

Deliberately unchanged: resolving the worktree through listResolvedWorktrees
rather than showManagedWorktree, which disambiguates a worktree id colliding
across hosts and is covered by its own test, and the SSH-folder
skill-install-ssh-dispatch-required throw, which matches the repo branch.

* fix(runtime): merge duplicate worktree-logic imports

The #17448 port added a third import from ../ipc/worktree-logic, which the
code-quality oxlint config rejects under --deny-warnings. Plain oxlint does not
flag it, so it only surfaced in CI's static analysis job.

* ci: run the ts-nocheck ratchet in PR checks

pr-workflow-lint-parity requires every leaf command in `pnpm lint` to have a
matching step in pr.yml. The ratchet was wired into lint but not the workflow,
so PR CI would not have enforced it.

* Merge remote-tracking branch 'origin/main' and retry the paired-host launch evaluate

main advanced 9 commits; none touch the orca-runtime.ts this branch splits, so
nothing needed porting.

CI failed twice on `Execution context was destroyed` thrown from
headless-paired-runtime-host's first `evaluate` after launch — a different spec
each run, which is the signature of the flake #17780 describes rather than a
regression. That commit added retryTransientMainEvaluate and adopted it in five
helpers but not this call site, even though its docblock names exactly this
case: the first evaluate after electron.launch() resolves, before the app is
ready. Wrapped it the same way.
2026-08-31 19:34:55 -07:00