Commit Graph
4 Commits
Author SHA1 Message Date
OrcaWinandm4air 4e8edc8872 feat(ssh): wire SshConnection through the work and transport close ledgers (#16741 T2 P2) (#24401)
* feat(ssh): wire SshConnection through the work and transport close ledgers (#16741 T2 P2)

Every operation SshConnection admits (exec, shell, sftp, file transfers, upload
sessions, forwarded channels and sockets, system-SSH commands) now runs through
the connection's work ledger, and every ssh2 client and proxy process it
allocates is tracked until it physically closes. Ordinary connect, reconnect
and disconnect behavior is unchanged.

Adds:
- subscribeTransportClosure: one-shot notice once the connection is disposed,
  every allocated transport has emitted 'close' and tracked work has drained.
  System-SSH startup is never proven closed from here.
- disconnectAndDrain(signal): for owned single-lifetime transports; fences new
  work, disconnects, and waits for physical close of the client, proxy, every
  allocated client and all fenced work. Refuses (after cleaning up) when the
  transport cannot be proven, e.g. system SSH or a connect still in flight.
- getExecutionDestination: the ssh2 endpoint, accepted host-key fingerprint and
  proxy-route digest proven by the current handshake (ssh-connection-destination).
- getTransportGeneration, prepareForwardRoute, openForwardSocket, forwardOut,
  forwardStreamLocal for later forwarding callers.
- An automaticReconnect constructor option (default on).

Channel close is local lifetime evidence only, never a remote-exit verdict.

Porting note (source: #16741 head a68b6f3531, merge-base 277c289bd4):
- Taken: the ledger hunks of ssh-connection.ts, ssh-connection-destination,
  ssh-forward-channel-lifetime, ssh-upload-session-lifetime, the system-SSH
  facade EOF hunk, and their tests.
- Adapted: operation bodies became private *Untracked methods called through
  the ledger instead of being re-indented; closure gating and the close drain
  moved to ssh-connection-transport-closure / ssh-connection-close-drain; the
  destination parser uses type guards instead of a cast. disconnectAndDrain
  fences through the ledger directly. Main's plain-SSH shell() goes through the
  ledger too. execCommand takes Pick<SshConnection, 'exec' |
  'usesSystemSshTransport'>; its string-stdin/maxOutputBytes hunk is not taken
  because main already streams stdin. Work-drain tests fence the private ledger
  until T3 adds the public fence; system-SSH drain cases split into their own
  file.
- Left for later slices: fenceWorkForReset (T3), isEphemeralRuntimeSshOwner
  (T6), assertProfileLifetimeAdmission (P8b), and the four manager drain cases
  in ssh-connection-disconnect-drain.test.ts (P3).

* refactor(ssh): shrink SshConnection below main and surface unhandled channel errors

ssh-connection.ts no longer grows under its max-lines exemption: it is 1872
lines, below main's 1917. The public API is unchanged.

- ssh-channel-open.ts: the channel-open waiter and session-limit retry.
- ssh-connection-file-transfers.ts: the uploadDirectory, downloadFile, upload
  session, writeFile and writeBuffer bodies, reading the connection through a
  small getter-based host so each read still sees the live transport.
- ssh-forward-channel-lifetime.ts: the forward client and stream-local checks.

The lifetime tracker's 'error' listener no longer hides errors. When it is a
channel's only error listener, the error is reported: SshConnection logs
"[ssh] Unhandled <kind> channel error for <target>: <message>", and other
callers fall back to a generic [ssh] warning. Nothing throws, so an orphaned
channel error still cannot crash the process.

* fix(ssh): name the forwarded local socket type and type the upload-session test stub

The forwarded local socket now takes SshConnectionWorkChannel, the event
surface the ledger tracks, instead of a broad object. The upload-session
lifetime test binds an EventEmitter rather than an untyped {}.

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 09:47:17 -07:00
fa85536f3a fix(ssh): repair unbuilt relay native deps (#8686)
Co-authored-by: Orca <help@stably.ai>
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
Co-authored-by: Jinwoo Hong <73622457+Jinwoo-H@users.noreply.github.com>
2026-07-16 21:32:13 -07:00
c2aa9c2ead feat(ssh): support file transfers over system ssh (#7804)
* feat(ssh): support file transfers over system ssh

* fix(ssh): harden system file transfers

* test(runtime): stub getRepo in headless mobile tab cwd test

The mobile-session selector validator (getValidatedExplicitWorktreeIdSelector,
from main) calls this.store?.getRepo to reject repo ids passed as worktree ids.
The store stub only implemented getWorkspaceSession, so the guarded call threw
'getRepo is not a function' once main merged into this branch. Add a getRepo
that returns null (wt-1 is a worktree, not a repo).

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com>
Co-authored-by: Orca <help@stably.ai>
2026-07-12 01:29:55 -07:00
4dbc9f3817 feat(ssh): add ControlMaster multiplexing for system SSH transport (#6922)
* feat(ssh): add ControlMaster multiplexing for system SSH transport

System SSH transport spawns a new OpenSSH process per exec command
(platform detect, relay install check, node resolution, relay launch,
socket probe). Each process pays the full SSH handshake cost — ~9s on
Uber devpods — making a typical relay connect take 54s+ and reliably
exceeding the 15s startup reconnect budget.

Add SSH ControlMaster multiplexing via a per-target socket in
$TMPDIR/orca-ssh-ctl/<hash>.sock. The first command establishes the
master; subsequent commands reuse it at ~100ms per exec instead of ~9s.
ControlPersist=300 keeps the master alive after commands exit so rapid
reconnects (e.g. on tab focus) also benefit. Windows is excluded since
OpenSSH's ControlMaster support there is limited.

* fix(ssh): address ControlMaster key collision and directory permission risks

- Use target.id in the socket key so distinct SSH targets can never
  collide even when configHost/port/user happen to match
- Switch from SHA1 to SHA256 and extend hash slice from 12 to 16 chars
- Stat the control-socket directory after mkdirSync to reject pre-existing
  dirs that are symlinks, foreign-owned, or have group/other write bits
  (mkdirSync mode is ignored on pre-existing dirs)
- Update two tests that used exact spawn-arg arrays; replace with
  ordering assertions (forward flags before --) that stay correct
  regardless of which extra ControlMaster options are injected

* fix(ssh): bind ControlPath identity to route and reject symlinked ctl dir

Fold proxyCommand/jumpHost/identity fields into the ControlPath hash so a
target whose route is edited no longer reuses a still-alive master built on
the old route. Switch the control-socket dir check from statSync to lstatSync
so a planted symlink fails the directory validation outright.

* test(ssh): drop tautological argv re-assertion in spawn checks

The toHaveBeenCalledWith re-passed the args array extracted from the same
mock call, making that argument position always pass. argv content is
already verified by the index-ordering assertions above; use expect.any(Array)
so the spawn check only claims what it actually verifies (binary path, stdio).

* fix(ssh): harden system ssh connection reuse

Co-authored-by: Orca <help@stably.ai>

* test(ssh): isolate control socket runtime dir

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Test <test@example.com>
Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
Co-authored-by: Orca <help@stably.ai>
2026-07-01 21:29:48 -07:00