Commit Graph
17 Commits
Author SHA1 Message Date
Brennan Benson f23d0b166f fix(relay): mint PTY ids that carry the relay incarnation instead of a restarting counter (#16901)
* fix(relay): scope PTY ids to mint epochs

* test(relay): treat minted PTY ids as opaque

* test(relay): pin mint-epoch id shape and restore spawn-sequence assertions

The epoch escaping had no test: dropping encodeURIComponent left the whole
relay suite green. Pin the three-field id shape against an epoch that carries
both separators, and cover a colon-bearing relay id through the unchanged
app-side SSH id wrapper.

subprocess.test.ts had traded `pty-1`/`pty-2` for `expect.any(String)`, which
discarded the invariant those two cases exist to prove: an early node-pty load
failure burns no sequence, a late spawn failure burns one.

* test(relay): mirror production epoch escaping in testPtyId

The harness built the expected id without the encodeURIComponent production
applies at the mint site. A test epoch carrying a reserved character would
diverge silently across ~40 assertions in 11 files.
2026-08-30 14:49:18 -07:00
Neil b516300b8c refactor agent hook listener modules (#16187) 2026-08-24 20:45:38 -07:00
JinjingandOrca 5c0195af64 Bound remote watcher fan-out and defer File Explorer refreshes (#11908)
* batch remote watcher events and defer File Explorer refreshes

Remote filesystem watcher events now batch with the shared 150ms trailing and 500ms
max-wait window, coalescing per-path like local events. File Explorer tree and
directory refreshes are scheduled with debounce and transport-aware concurrency caps
(16 local, 8 runtime, 4 SSH). Stale directory cache tracking prevents trusting
collapsed listings skipped by full refresh; they are re-read on re-expansion. Relay
implements a 15-minute idle-only grace cap for zero-PTY relays via PTY pool
lifecycle tracking, independent of explicitly configured grace time.

* fix(watch/relay): bound remote watcher fan-out and read the live relay grace

Three P1 fixes from the SSH/remote freeze audit:

- Remote watchers now debounce on the same 150/500 window as local ones
  (finding D), and every teardown path drops the trailing flush timer
  instead of letting it fire into a dead watch. The deferred send is
  wrapped so a frame disposed mid-window can't escape as a fatal
  main-process exception.
- File Explorer refreshes are scheduled and concurrency-capped rather
  than fanned out unbounded over expanded dirs (finding C). Local
  transports use a zero window, since main already coalesced the burst.
- relay.startGrace reads ptyHandler.configuredGraceTimeMs instead of the
  launch-time argv closure, so a grace raised after launch is honored.
  The branch selection moves to relay-grace-branch.ts because relay.ts
  has no exports and calls main() at import, making it untestable.
  Consequence: a host-sleep relay holding zero PTYs now exits after the
  idle cap. Pinned by test and documented in
  docs/reference/relay-grace-time-reconfiguration.md.

Also drops the duplicated 150/500/5000 constants in the runtime-RPC
batcher in favor of the shared window module.

* docs(relay): correct grace-reconfiguration line numbers after the relay.ts edit

Co-authored-by: Orca <help@stably.ai>

* refactor(file-explorer): use useMemo for paths; remove relay reference

Replace manual ref-based caching with proper React hooks for content-stable path memoization. Remove outdated relay grace-time reference documentation from code review cycle.

* rm design doc

* fix(remote-watcher): prevent stranded timer after close

An in-flight provider receive can land after the batch is torn down.
Without a guard, pushing events to a closed batch would re-arm a timer
that would never be cleared, stranding the task indefinitely. Track the
closed state and skip pushes after close().

Relay.ts comment clarifies why pool watches remain registered during
grace-period shutdown deferral — the socket server stays listening so
a reconnecting client can cancel the grace and resume.

---------

Co-authored-by: Orca <help@stably.ai>
2026-08-01 11:55:58 -07:00
Brennan Benson 86a993aec4 fix(ssh): connect to Linux hosts that cannot compile node-pty (#10776)
* fix(ssh): connect to Linux hosts that cannot compile node-pty

node-pty ships no Linux prebuilt at any architecture, so it is compiled on
the remote. On a host without a C/C++ toolchain that build fails, and because
both native deps install in one npm command it also took down
@parcel/watcher — which does have a working Linux prebuilt — and failed the
whole connection. Every Linux image without build tools was unusable.

node-pty only backs remote terminals; files, git, and the editor do not need
it, and a missing native dep is already non-fatal further down the deploy. So
when the existing toolchain probe confirms the compiler is missing, reinstall
without node-pty instead of aborting. The manifest has to drop it too — npm
reconciles every dependency in package.json, not just the ones named on the
command line, so naming only @parcel/watcher still rebuilds node-pty.

If that reinstall also fails the actionable build-tools error is rethrown, so
a host broken for some other reason still reports the toolchain gap.

The relay's PTY error now names the fix rather than saying only that node-pty
is unavailable.

Verified on a stock Rocky Linux 10.2 aarch64 container (openssh-server, git,
nodejs, npm, no compiler): connect succeeds, /etc lists over SSH, node-pty is
absent while @parcel/watcher installs its linux-arm64-glibc prebuilt, and
spawning a terminal reports the install hint.

* fix(ssh): keep the node-pty skip path honest about platform and watcher

The PTY unavailable message named build tools unconditionally, but only Linux
compiles node-pty — the deploy-side skip is gated on linux and the toolchain
probe returns null on Windows. A Windows or macOS remote, where node-pty ships
prebuilds, was told to install make/g++/python3. Pick the remedy by the relay's
own platform.

The skip path returned before the install probe, so a @parcel/watcher that
installs but cannot require() (glibc below the floor) connected with dead file
watching and nothing logged. Probe before returning and warn; no rebuild, since
node-pty provably cannot compile on that host, and never fatal.

Also log the pty-less reinstall's own failure and attach it as cause — the
rethrown toolchain message is built from the original npm error, so an
unrelated retry failure (registry, ENOSPC, EACCES) was lost. The reinstall now
keeps the caller's resetDeps as well, so a repair reconnect still clears every
dep the probe found broken.

Tests: the skip-success fixture queued a chmod/probe/rebuild sequence
production never runs, and the surplus slots were absorbed by launchRelay's
readiness poll (1817ms vs 3-9ms for its peers). It now emits exactly the 12
execs production performs, and pins that no rebuild is issued. Adds the missing
negative case: a gyp-shaped failure on a host whose probe reports a complete
toolchain must still hard-fail rather than silently degrade.

* fix(ssh): hedge the node-pty remedy and keep repair resets on the skip path
2026-07-26 15:29:13 -07:00
fa85536f3a fix(ssh): repair unbuilt relay native deps (#8686)
Co-authored-by: Orca <help@stably.ai>
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
Co-authored-by: Jinwoo Hong <73622457+Jinwoo-H@users.noreply.github.com>
2026-07-16 21:32:13 -07:00
Brennan Benson 64be819790 fix(runtime): harden watcher and PTY teardown ownership (#8661)
* fix(runtime): retain watcher and PTY teardown ownership

* fix(runtime): restore watchers after interrupted cleanup

* fix(runtime): prevent stale watcher revival

* test(runtime): cover watcher shutdown ownership

* test(daemon): model physical PTY exit

* fix(daemon): keep shutdown terminating when disposal cannot prove exit

A rejecting host.dispose() (unreapable child past its exit deadline) left
the shutdown RPC without its process.nextTick(shutdown) and skipped socket
cleanup in shutdown(), stranding the daemon as an unreachable orphan after
the stale-daemon replacement flow unlinks its socket. Log and continue:
daemon exit reparents the child to init instead of blocking on it.

* fix(runtime): keep local watching alive after an idle-kill deadline miss

An idle child that outlived the exit deadline set shutdownRequested on the
shared desktop supervisor, which has no retire-and-replace path — every
later subscribe rejected supervisor_disposed and the roots were cached
unwatchable, silently ending local file watching for the session. The idle
path owns zero records, so there is no double-watch hazard; the zombie
keeps its capacity reservation until physical exit and the next subscribe
gets a fresh child.

* fix(renderer): resync replayed paired-web file watches

Transparent replay removed the implicit resync the old close-and-rebuild
path provided: a replayed files.watch only reports changes from its own
native setup, so changes during the reconnect gap were silently lost.
Deliver a conservative overflow to consumers once the replayed watch is
ready, matching the overflow-after-interruption contract everywhere else.

* fix(runtime): address teardown review findings

* fix(runtime): retry watches after teardown deadlines

* Fix PTY descendant leaks on forced teardown

* Fix jitter-sensitive terminal lifecycle test
2026-07-15 15:23:35 -07:00
Jinjing e3c47eff17 Fix ssh watcher isolation (#8463)
* fix(ssh): isolate relay filesystem watchers

* Fix relay watcher fault-harness pid file and in-process fallback isolati

- Use exclusive ('wx') creation for the fault-harness pid file so a leaked
  ORCA_WATCHER_CHILD_PID_FILE env var can't clobber an existing file, and
  have the harness remove the file after reading a replacement pid.
- Force useInProcessVitestFallback to false in the relay watcher pool so a
  leaked VITEST env var can never load the native watcher addon in-process
  on the relay; fail closed instead when the isolated child is missing.
- Thread an injectable RelayWatcherProcessPool into FsHandler/
  RelayFilesystemWatchRegistry for tests, and add coverage for both fixes.
2026-07-12 21:34:59 -07:00
NeilandOrca e33b2006f4 Remove stale max-lines lint disables from files under the limit (#7548)
110 files carried an eslint/oxlint-disable max-lines directive but are
already under the default max-lines budget (300 .ts / 400 .tsx / 600 .mjs
/ 800 test), so the suppression is dead. Removing it restores real
max-lines coverage on these files with zero behavior change.

Each removed directive had max-lines as its only rule; verified via a
full oxlint run (0 max-lines violations, 0 new errors). Diff is pure
deletions (200 lines, 0 additions) — no code touched.

Co-authored-by: Orca <help@stably.ai>
2026-07-06 02:12:32 -07:00
Jinwoo HongandOrca 782eb12688 Make remote SSH terminals persistent by default (#6955)
Co-authored-by: Orca <help@stably.ai>
2026-06-30 16:16:10 -07:00
NeilandOrca 46646d7ff1 chore(lint): upgrade oxlint to 1.71 + enable 7 new rules (autofixed backlog) (#6841)
* chore(lint): upgrade oxlint to 1.71 and enable 7 new rules

Upgrade oxlint 1.67.0 -> 1.71.0 (1.72 was blocked by the repo's 3-day
minimum-release-age supply-chain guard; nothing here needs it). The
bump is a no-op on the existing config.

Enable 3 error rules (backlog autofixed to zero in this commit) and
4 warn rules (surface signal without gating CI):

error (autofixed, behavior-preserving):
- unicorn/prefer-node-protocol        (~1531 sites: bare builtin -> node:)
- typescript/no-import-type-side-effects (~36: all-inline-type -> import type)
- unicorn/no-array-reverse            (19: copy-then-reverse -> toReversed)

warn (real signal, current fires are test-only/correct):
- unicorn/no-array-fill-with-reference-type  (aliasing footgun guard)
- typescript/no-unsafe-function-type         (bans bare Function type)
- unicorn/prefer-array-flat-map              (map().flat() -> flatMap())
- unicorn/prefer-regexp-test                 (.match() in bool ctx -> .test())

mobile/.oxlintrc.json extends root, so it inherits all 7; the autofix
ran from root and covered mobile/ too.

Verification (all green): oxlint 0 errors (root+mobile+aux configs),
oxfmt clean, typecheck (node+cli+web), vitest 22795 passed / 0 failed,
builds (electron-vite + web + cli) succeed. node: rewrites confirmed to
skip embedded SSH/CLI string payloads (AST-only); all toReversed sites
verified to operate on fresh copies or write-once locals.

* chore(lint): bump mobile oxlint to 1.71 so inherited rules parse

mobile/ is a standalone pnpm project pinning its own oxlint@1.67, which
lacks unicorn/no-array-fill-with-reference-type (needs >=1.70). Since
mobile/.oxlintrc.json extends the root config, mobile CI's 'cd mobile &&
oxlint' failed to parse the new rule. Bump mobile to match root (1.71).

Verified in mobile/: oxlint 0 errors, oxfmt --check clean, tsc --noEmit
pass, vitest 978 passed / 0 failed.

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Orca <help@stably.ai>
2026-06-29 22:38:29 -07:00
0ec3882cb8 Add project Windows runtime selection (#5519)
* Add project Windows runtime selection

* Fix project Windows runtime selection

Co-authored-by: Orca <help@stably.ai>

* fix: preserve WSL shell variables

---------

Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
Co-authored-by: Orca <help@stably.ai>
Co-authored-by: Neil <neil@stably.ai>
2026-06-17 16:08:14 -07:00
Jinjingandorca-bug-scan-bot be305c8e0a Fix stale relay socket startup recovery (#2563)
* fix: address pr-bug-scan validated finding from #2447

Restore self-recovery from a stale socket file (left behind when a prior relay was killed by SIGKILL/OOM/host crash) without unlinking a live duplicate's socket.

* fix: address review findings

---------

Co-authored-by: orca-bug-scan-bot <orca-bug-scan-bot@stably.ai>
2026-05-21 15:28:07 -07:00
c99633e612 fix(ssh-relay): prevent duplicate daemons from replacing active sockets (#2447)
* fix(ssh-relay): prevent duplicate detached relays

* review: harden relay socket ownership

- delay hook endpoint publication until relay socket ownership is proven

- verify socket identity before closing or unlinking owned relay sockets

- add regressions for endpoint poisoning, rebound sockets, and post-client grace

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
Co-authored-by: Orca <help@stably.ai>
2026-05-20 23:01:24 -04:00
Jinwoo HongandOrca ea9a718e29 fix(ssh): remove relay FS path allowlist to support symlinks outside workspace (#1661) (#1672)
When a remote SSH workspace contains a symlink whose target lies outside
the registered repo/worktree roots, file reads failed with 'Path outside
authorized workspace'. This silently broke common workflows: HPC dataset
mounts, multi-checkout repos, dotfile editing, and any cross-mount
symlink.

Drop `RelayContext.authorizedRoots`, `validatePath`, and
`validatePathResolved` along with all ~33 call sites in fs-handler.ts
and git-handler.ts. The relay's threat model becomes 'the relay runs as
the SSH user and trusts the renderer.'

Why this is acceptable: `pty.spawn` and `git.exec` already concede the
same threat. A renderer that wants to reach `/etc/passwd` can spawn a
shell or run `git -C /etc cat-file`; the FS allowlist was friction, not
a security boundary. Intra-worktree path checks in `getDiff` and
`discard` are intentionally preserved.

Back-compat preserved: `session.registerRoot` (notification + request)
remains a valid RPC, retained as no-ops on new relays. Old main + new
relay and new main + old relay both keep working through the upgrade
window. `registerRelayRoots` is also kept for the same reason. A
narrowed error-translation block in `worktree-remote.ts` handles old
relays still surfacing the legacy error string to users.

Tests: removed two negative-allowlist tests; added a positive control
('reads files outside any registered root') and a direct regression
test for #1661 ('reads files via symlinks resolving outside the
workspace'). All 469 relay/SSH/IPC tests pass.

See docs/relay-fs-allowlist-removal.md for the full rationale,
back-compat matrix, alternatives considered, and follow-up cleanup
plan.

Closes #1661

Co-authored-by: Orca <help@stably.ai>
2026-05-10 15:59:54 -07:00
Jinwoo HongandOrca 451d7cab42 fix(ssh): await root registration before remote worktree creation (#1120)
Co-authored-by: Orca <help@stably.ai>
2026-04-26 14:14:30 -07:00
Jinwoo Hong 68e658e082 fix(ssh): use SSH target label for home directory remote repos (#1031) 2026-04-24 13:58:48 -07:00
Jinwoo Hong cc66e120eb feat: Add SSH remote support (beta) (#590) 2026-04-13 19:23:09 -07:00