Commit Graph
10 Commits
Author SHA1 Message Date
Neil 57681ecd09 fix(remote): resolve the spawn cwd, the node manager dir, the vault host and the scrollback seed (#17952)
* fix(remote): resolve workspace cwd, mise Node, host scope, and TUI scrollback honestly

#15296 relay: a folder workspace id (`folder:<uuid>`) carries no path, so the
worktree-id split yielded nothing and $HOME silently won. Resolve the spawn cwd
through worktreeId -> ORCA_WORKSPACE_ROOT -> host default, and refuse an agent
spawn outright when a folder workspace names a root this host cannot resolve.

#11733 ssh: generalize the NVM dotfile scrape into `orca_dotfile_dirs` and drive
mise off `MISE_DATA_DIR` / `XDG_DATA_HOME` instead of a hardcoded
`$HOME/.local/share/mise`.

#13713 ai-vault: an unresolvable workspace host is `unverifiable`, not local.
Widen the default scope to every host rather than scanning the client's own
history and reporting "No agent sessions found".

#6106 terminal: hydration asked the renderer for `scrollback: 0` while an
alt-screen TUI was up, which drops the normal buffer's shell history rather than
the TUI bytes. Drop the flag; readers already split the two buffers apart.

* fix(remote): stop the relay answering host questions for a guest execution host

Three findings from review of the spawn-cwd resolver, all the same shape: a path
question answered against the wrong host, or with the wrong key.

- resolveRelaySpawnCwd refused an agent launch whenever a folder workspace named
  a root that did not stat on the relay. But relayHostDirectoryExists stats the
  relay's *own* filesystem, and the relay supports WSL shells, so a folder
  workspace on a Windows relay launching into WSL now threw where it previously
  spawned -- contradicting the function's own doc comment, which says an absent
  path for that exact host pair is a miss, not a refusal. Thread the shell's
  execution host in and demote the refusal to a miss when the spawn does not run
  on the relay's filesystem.

- requireRelaySpawnCwd's doc claims both call sites route through one resolver
  so the fence can never be keyed on a directory the spawn won't use, but the
  fence key was still computed with the non-stripping splitWorktreeId while the
  cwd used splitWorktreeIdForFilesystem. For a `::workspace:<uuid>` id those
  disagree by construction, in adjacent lines: the removal fence guarded a path
  no spawn ever enters. Same defect in shutdownForWorktreePath and the revive
  path; all three now use the filesystem split.

- The remote Node probe expanded `$HOME` and `~/` prefixes out of a dotfile
  assignment but not `$XDG_DATA_HOME`, so `MISE_DATA_DIR=$XDG_DATA_HOME/...`
  was used as a literal directory name. Add the case arm, defaulting to the
  POSIX `$HOME/.local/share` the seed value already uses -- sshd's exec channel
  usually has no XDG_DATA_HOME at all.
2026-09-02 21:33:41 -07:00
NeilandOrca aab112933e Revert "fix(memory): bound OOM-prone accumulators (#10179)" (#10255)
Co-authored-by: Orca <help@stably.ai>
2026-07-23 18:35:31 -07:00
Neil 8f40ddf328 fix(memory): bound OOM-prone accumulators (#10179) 2026-07-23 06:22:56 -07:00
Brennan Benson 3e276d78ba fix(ssh): probe npm via prepended PATH, not colocated with node (#9165) (#9255)
* fix(ssh): probe npm via prepended PATH, not colocated with node (#9165)

The remote Node/npm toolchain gate invoked npm by its absolute path
<nodeBinDir>/npm (POSIX) / npm.cmd (Windows, behind a Test-Path
colocation check). But deploy (commandWithNodePath) runs bare `npm`
with nodeBinDir merely prepended to PATH, so npm can resolve from
anywhere on PATH.

A host whose only resolvable node has npm elsewhere on PATH (e.g. node
symlinked into a dir without npm) deployed fine on v1.4.144, but after
upgrade the candidate is rejected with no fallback → SSH/relay
connection fails to establish.

Make the probe resolve npm exactly the way deploy does — bare
`npm --version` under the same prepended PATH — so it still confirms npm
is runnable (the #8450 concern) without requiring colocation. Windows
now prepends the backslash-form dir (matching deploy) so bare-command
PATH lookup resolves reliably.

* test(ssh): cover split Node npm PATH resolution
2026-07-17 18:51:10 -07:00
Brennan Benson f544820552 fix(ssh): require a coherent colocated Node/npm toolchain for the relay (#9165) 2026-07-17 14:27:31 -07:00
Haohan Lin 190ed7d4ed fix(ssh): make POSIX relay wrapper survive csh/tcsh login shells (#8709)
Reconstruct remote POSIX commands with bounded printf arguments so non-POSIX SSH login shells can forward them without requiring remote base64. Preserve relay and system-SSH stdin, and centralize login-shell flag selection for csh/tcsh compatibility.\n\nValidated against real csh and tcsh OpenSSH targets with built-in and system SSH, including cold relay deployment, stdin upload, PTY I/O, file mutation, and reconnect.
2026-07-15 02:38:40 -07:00
leter 6677de1807 Improve SSH Node.js prerequisite guidance
Improve SSH remote Node.js prerequisite failures by adding package-manager-specific POSIX install guidance, Windows Node 18+ candidate validation, fnm XDG probing, and regression coverage while preserving SSH abort and session-limit behavior.
2026-07-04 23:54:26 -07:00
f215a48064 fix: pr-bug-scan validated finding from #6952 (#7180)
* fix: address pr-bug-scan validated finding from #6952

throwNodeNotFound() now re-raises AbortError when the shared signal is aborted, so a signal-cancelled node probe no longer launders into 'Node.js not found'; sequential fallback runs.

* fix(ssh): make session-limited (MaxSessions=1) relay deploys actually succeed

Review of #7180 verified the parent fallback end-to-end against a real
MaxSessions=1 sshd and found the connect still failed. Four gaps, in order
of discovery:

- isSshSessionLimitError missed stock OpenSSH, which refuses session
  channels over MaxSessions with SSH2_OPEN_CONNECT_FAILED (2) and 'open
  failed' — reason 4 never matched, so the fallback never triggered.
- execCommand settled aborted commands before the channel finished
  closing, so the sequential fallback reissued execs while sshd still
  counted the old session.
- SshConnection.waitForSshCallback rejected aborts mid-channel-open
  immediately, leaking a confirmed-late channel that held the only
  session slot; it now settles after the late channel closes (bounded)
  and drains its streams so ssh2 emits 'close'.
- Session channel opens now retry transient session-limit refusals
  (sshd frees the slot only after processing our close-ack, which the
  next open can beat by microseconds), and the remote orca CLI shim
  install is non-fatal like the managed-hook install — after the relay
  bridge occupies the sole session slot, raw-connection extras must
  degrade instead of failing the connection.

Verified live against Docker sshd (OpenSSH 9.2, MaxSessions=1): fresh
deploy (upload + native deps + launch), reconnect cycles, and a PTY
round-trip all succeed; unrestricted-sshd regression run also passes.

Co-authored-by: Orca <help@stably.ai>

* Handle ssh execution aborts immediately during retry backoff or hangs

- Cancel the session-limit retry delay immediately if the operation is
  aborted during backoff.
- Limit the wait time to a 5-second grace period when aborted during a
  channel open that is hung and never invokes its callback, rather than
  waiting for the full connection timeout.

---------

Co-authored-by: orca-bug-scan-bot <orca-bug-scan-bot@stably.ai>
Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com>
Co-authored-by: Orca <help@stably.ai>
2026-07-04 01:26:55 -07:00
Brandon Barker e6ca3087c5 perf(ssh): resolve node path concurrently with remote install state (#6952)
Run independent SSH relay bootstrap probes concurrently when the connection can safely support overlapping execs. Preserve the old sequential path for system SSH without reusable ControlMaster and for remotes that reject concurrent session channels.
2026-07-02 13:55:36 -07:00
44935d4ccb Fix remote Node.js detection for nvm, mise, asdf, and volta (#6037)
* Fix remote Node.js detection for nvm, mise, asdf, and volta

Remote Node resolution failed when node was installed via a version
manager (nvm with custom NVM_DIR, mise, asdf, volta) or when the user's
login shell was zsh/fish rather than bash.

Root cause: the SSH exec transport runs every command under /bin/sh,
which never sources shell init files. The only init-aware path was a
hardcoded `bash -lc` fallback that missed zsh/fish users and was never
reached for the newer version managers. nvm was handled by guessing
~/.nvm (breaking custom NVM_DIR), and mise/asdf/volta had no probes at
all. There was also no version gate, so nvm's highest-version glob
could return Node 8/10/12 and crash the relay on launch.

Fix: resolve via the user's own $SHELL as a login shell first (the only
path that runs nvm.sh / mise activate / asdf.sh init hooks), then fall
back to direct path probes for all major managers (nvm respecting
$NVM_DIR, fnm, mise, asdf, volta, n) plus system locations. Every
candidate is version-checked against the relay's Node 18+ requirement
before being accepted.

* Address CodeRabbit review: probes-first, no || short-circuit

- Reorder to path-probes first (deterministic, doesn't depend on shell
  rc-file semantics where bash -lc skips .bashrc and zsh -lc skips
  .zshrc — exactly where nvm/mise/asdf hooks live).
- Join probes with newlines instead of || so an empty
  `ls | sort -V | tail -1` (exit 0) doesn't mask later probes.
- Deduplicate candidate paths before version-checking.
- Drop unreachable mock and fix misleading $SHELL-unset test name.
- Login shell is now a fallback for custom ~/.profile PATH setups.

* Fix remote Node path probing portability

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
Co-authored-by: Orca <help@stably.ai>
2026-06-22 16:46:47 -07:00