* fix(ssh): probe npm via prepended PATH, not colocated with node (#9165)
The remote Node/npm toolchain gate invoked npm by its absolute path
<nodeBinDir>/npm (POSIX) / npm.cmd (Windows, behind a Test-Path
colocation check). But deploy (commandWithNodePath) runs bare `npm`
with nodeBinDir merely prepended to PATH, so npm can resolve from
anywhere on PATH.
A host whose only resolvable node has npm elsewhere on PATH (e.g. node
symlinked into a dir without npm) deployed fine on v1.4.144, but after
upgrade the candidate is rejected with no fallback → SSH/relay
connection fails to establish.
Make the probe resolve npm exactly the way deploy does — bare
`npm --version` under the same prepended PATH — so it still confirms npm
is runnable (the #8450 concern) without requiring colocation. Windows
now prepends the backslash-form dir (matching deploy) so bare-command
PATH lookup resolves reliably.
* test(ssh): cover split Node npm PATH resolution
Reconstruct remote POSIX commands with bounded printf arguments so non-POSIX SSH login shells can forward them without requiring remote base64. Preserve relay and system-SSH stdin, and centralize login-shell flag selection for csh/tcsh compatibility.\n\nValidated against real csh and tcsh OpenSSH targets with built-in and system SSH, including cold relay deployment, stdin upload, PTY I/O, file mutation, and reconnect.
* fix: address pr-bug-scan validated finding from #6952
throwNodeNotFound() now re-raises AbortError when the shared signal is aborted, so a signal-cancelled node probe no longer launders into 'Node.js not found'; sequential fallback runs.
* fix(ssh): make session-limited (MaxSessions=1) relay deploys actually succeed
Review of #7180 verified the parent fallback end-to-end against a real
MaxSessions=1 sshd and found the connect still failed. Four gaps, in order
of discovery:
- isSshSessionLimitError missed stock OpenSSH, which refuses session
channels over MaxSessions with SSH2_OPEN_CONNECT_FAILED (2) and 'open
failed' — reason 4 never matched, so the fallback never triggered.
- execCommand settled aborted commands before the channel finished
closing, so the sequential fallback reissued execs while sshd still
counted the old session.
- SshConnection.waitForSshCallback rejected aborts mid-channel-open
immediately, leaking a confirmed-late channel that held the only
session slot; it now settles after the late channel closes (bounded)
and drains its streams so ssh2 emits 'close'.
- Session channel opens now retry transient session-limit refusals
(sshd frees the slot only after processing our close-ack, which the
next open can beat by microseconds), and the remote orca CLI shim
install is non-fatal like the managed-hook install — after the relay
bridge occupies the sole session slot, raw-connection extras must
degrade instead of failing the connection.
Verified live against Docker sshd (OpenSSH 9.2, MaxSessions=1): fresh
deploy (upload + native deps + launch), reconnect cycles, and a PTY
round-trip all succeed; unrestricted-sshd regression run also passes.
Co-authored-by: Orca <help@stably.ai>
* Handle ssh execution aborts immediately during retry backoff or hangs
- Cancel the session-limit retry delay immediately if the operation is
aborted during backoff.
- Limit the wait time to a 5-second grace period when aborted during a
channel open that is hung and never invokes its callback, rather than
waiting for the full connection timeout.
---------
Co-authored-by: orca-bug-scan-bot <orca-bug-scan-bot@stably.ai>
Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com>
Co-authored-by: Orca <help@stably.ai>
Run independent SSH relay bootstrap probes concurrently when the connection can safely support overlapping execs. Preserve the old sequential path for system SSH without reusable ControlMaster and for remotes that reject concurrent session channels.
* Fix remote Node.js detection for nvm, mise, asdf, and volta
Remote Node resolution failed when node was installed via a version
manager (nvm with custom NVM_DIR, mise, asdf, volta) or when the user's
login shell was zsh/fish rather than bash.
Root cause: the SSH exec transport runs every command under /bin/sh,
which never sources shell init files. The only init-aware path was a
hardcoded `bash -lc` fallback that missed zsh/fish users and was never
reached for the newer version managers. nvm was handled by guessing
~/.nvm (breaking custom NVM_DIR), and mise/asdf/volta had no probes at
all. There was also no version gate, so nvm's highest-version glob
could return Node 8/10/12 and crash the relay on launch.
Fix: resolve via the user's own $SHELL as a login shell first (the only
path that runs nvm.sh / mise activate / asdf.sh init hooks), then fall
back to direct path probes for all major managers (nvm respecting
$NVM_DIR, fnm, mise, asdf, volta, n) plus system locations. Every
candidate is version-checked against the relay's Node 18+ requirement
before being accepted.
* Address CodeRabbit review: probes-first, no || short-circuit
- Reorder to path-probes first (deterministic, doesn't depend on shell
rc-file semantics where bash -lc skips .bashrc and zsh -lc skips
.zshrc — exactly where nvm/mise/asdf hooks live).
- Join probes with newlines instead of || so an empty
`ls | sort -V | tail -1` (exit 0) doesn't mask later probes.
- Deduplicate candidate paths before version-checking.
- Drop unreachable mock and fix misleading $SHELL-unset test name.
- Login shell is now a fallback for custom ~/.profile PATH setups.
* Fix remote Node path probing portability
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
Co-authored-by: Orca <help@stably.ai>
* feat: add windows ssh relay base support
* feat: support windows ssh relay runtime services
* fix: default windows ssh pty cwd to user profile
* fix: support windows hosts over system ssh
* fix: preserve degraded windows relay native deps
* fix: gate windows shell args by relay platform
* fix: preserve windows relay fallback pipes
* test: align windows native deps relay fixture
* fix: build valid windows install lock command
* fix: address windows SSH relay review findings
Resolve correctness, efficiency, and reuse issues found reviewing the
Windows SSH native-host support:
- GC liveness on Windows now probes the actual named pipe (via node
net.connect against markers + deterministic candidates) instead of
substring-matching Win32_Process command lines, which could remove a
live relay dir. Reports ALIVE conservatively only when there is no
liveness signal at all (no markers and no seed pipes).
- Resolve the remote node path once per deploy and thread it through
install/repair/launch instead of re-resolving 3-7x.
- Replace the 200ms node -e poll loop with a single long-lived remote
wait process during Windows relay startup.
- Skip the no-op executable command on Windows in uploadRelay.
- Make the Windows fallback pipe name deterministic and recoverable
(drop the global counter), with an extra reconnect attempt.
- Normalize the prepended node bin dir to backslashes on Windows PATH.
- Batch the system-SSH Windows directory upload into a single streamed
JSON package instead of one ssh process per file.
- Extract relay endpoint/marker helpers into ssh-relay-endpoints.ts and
consolidate the PowerShell EncodedCommand encoding into the shared
powershell-command-encoding module.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* Support cancellation and timeouts in Windows port scanning
- Propagate the request AbortSignal and a 5-second timeout to both
PowerShell and netstat child processes during Windows port scanning.
- Avoid spawning the netstat fallback process if the port scan has
already been aborted.
- Wrap the .NET OSArchitecture check in a try/catch block during SSH
Windows platform detection to robustly fall back to environment
variables if needed.
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com>