Commit Graph
5 Commits
Author SHA1 Message Date
Neil 510305e574 fix(relay): signal capacity loss instead of dropping, hanging, or truncating (#17870)
Three failures with one shape: a payload past a fixed capacity was met with
silence, with a wait that never ends, or with a prefix presented as a whole.

**The workspace snapshot was silently dropped.** `workspace.changed` carries the
tab/session list, and a snapshot past the producer frame capacity (12288 B on a
Node <=21 remote) was dropped with only a relay stderr line, so the client kept a
stale list forever. The relay now publishes per client and, for a client whose
sink refused the frame, sends a compact `workspace.stale` marker on the control
lane; the client re-reads through `workspace.get`, whose lane is budgeted in
megabytes rather than in one producer frame. A new JSON-RPC notification rather
than a new field on `workspace.changed`: `normalizeSnapshot(undefined, ns)` yields
revision 0 and an empty session, so a Rule-1 field would make an old client
replace its tab list with nothing — worse than the drop. An old client ignores the
unknown method and is exactly where it is today. The marker retention/retry
machinery is extracted from the `fs.changed` overflow path and shared by both.

**The Windows upload hung, and the fix for it could truncate.** `#16432` was
attributed to `[Console]::In.ReadToEnd()` materializing the base64 bundle. That is
not what the reporter measured: he also measured
`new IO.StreamReader([Console]::OpenStandardInput())` — an incremental reader —
hanging at 1 MB. The limit is in the stdin the host hands PowerShell over a
non-pty ssh exec, not in the string the script builds.

- `uploadFileViaSystemSsh` — the user file-import path — was piping a whole file
  into one Windows stdin, unchunked and untimed. That is the path large files
  take; it now chunks into 32 KB writes and bounds each wait.
- The Windows directory upload reuses that single-file path rather than repeating
  a weaker copy of chunk-read + write-buffer; the `ino`/`dev` TOCTOU verification
  comes with it.
- A Windows write needing more than one exec lands on a `.orca-partial` staging
  path and is published by rename, so a failed chunk cannot leave a truncated
  artifact under the real name. `exclusive` is enforced once at the rename, not on
  the first chunk, where a retry met its own leftovers.
- The mkdir batch reads stdin through the stream reader the reporter measured
  surviving 50 KB, not `[Console]::In`, which he measured wedging at that size.
- `waitForChannelClose` takes an optional bound. A wedged PowerShell stays alive
  at idle CPU and never closes, so without one the promise is simply never
  settled and the caller waits forever with no error to show.

**Quick Open showed a prefix as the whole workspace.** The mechanism "a full page
means there is more" only works if the caller named the cap, and the failing UI
named none — it hardcoded `truncated: false`. Quick Open now names
`QUICK_OPEN_LISTING_MAX_RESULTS` on both the Electron IPC hop and the runtime-RPC
hop (the field #17954 added to `files.listAll`), and reads a full page as
truncation. The local hop honours the cap too, which it previously ignored.

Rebase note on `fs.listFiles`: an earlier revision of this work also clamped the
host unconditionally, and #17934 escalated an uncapped request to an explicit
error. #17954 has since landed and made an oversized reply streamable, which
removes the premise — the host no longer has to choose between a prefix and a
refusal, so it returns the whole listing when no limit is named and only clamps a
limit it was given. Keeping either would have regressed #17954 and hard-failed
three in-tree callers that deliberately pass no options
(`runtime-file-commands-search-runtime-files.ts:81`,
`filesystem-read-handlers.ts:125`, `runtime-file-commands-constructor.ts:41`).
2026-09-02 15:42:08 -07:00
NeilandOrca aab112933e Revert "fix(memory): bound OOM-prone accumulators (#10179)" (#10255)
Co-authored-by: Orca <help@stably.ai>
2026-07-23 18:35:31 -07:00
Neil 8f40ddf328 fix(memory): bound OOM-prone accumulators (#10179) 2026-07-23 06:22:56 -07:00
fa85536f3a fix(ssh): repair unbuilt relay native deps (#8686)
Co-authored-by: Orca <help@stably.ai>
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
Co-authored-by: Jinwoo Hong <73622457+Jinwoo-H@users.noreply.github.com>
2026-07-16 21:32:13 -07:00
4dbc9f3817 feat(ssh): add ControlMaster multiplexing for system SSH transport (#6922)
* feat(ssh): add ControlMaster multiplexing for system SSH transport

System SSH transport spawns a new OpenSSH process per exec command
(platform detect, relay install check, node resolution, relay launch,
socket probe). Each process pays the full SSH handshake cost — ~9s on
Uber devpods — making a typical relay connect take 54s+ and reliably
exceeding the 15s startup reconnect budget.

Add SSH ControlMaster multiplexing via a per-target socket in
$TMPDIR/orca-ssh-ctl/<hash>.sock. The first command establishes the
master; subsequent commands reuse it at ~100ms per exec instead of ~9s.
ControlPersist=300 keeps the master alive after commands exit so rapid
reconnects (e.g. on tab focus) also benefit. Windows is excluded since
OpenSSH's ControlMaster support there is limited.

* fix(ssh): address ControlMaster key collision and directory permission risks

- Use target.id in the socket key so distinct SSH targets can never
  collide even when configHost/port/user happen to match
- Switch from SHA1 to SHA256 and extend hash slice from 12 to 16 chars
- Stat the control-socket directory after mkdirSync to reject pre-existing
  dirs that are symlinks, foreign-owned, or have group/other write bits
  (mkdirSync mode is ignored on pre-existing dirs)
- Update two tests that used exact spawn-arg arrays; replace with
  ordering assertions (forward flags before --) that stay correct
  regardless of which extra ControlMaster options are injected

* fix(ssh): bind ControlPath identity to route and reject symlinked ctl dir

Fold proxyCommand/jumpHost/identity fields into the ControlPath hash so a
target whose route is edited no longer reuses a still-alive master built on
the old route. Switch the control-socket dir check from statSync to lstatSync
so a planted symlink fails the directory validation outright.

* test(ssh): drop tautological argv re-assertion in spawn checks

The toHaveBeenCalledWith re-passed the args array extracted from the same
mock call, making that argument position always pass. argv content is
already verified by the index-ordering assertions above; use expect.any(Array)
so the spawn check only claims what it actually verifies (binary path, stdio).

* fix(ssh): harden system ssh connection reuse

Co-authored-by: Orca <help@stably.ai>

* test(ssh): isolate control socket runtime dir

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Test <test@example.com>
Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
Co-authored-by: Orca <help@stably.ai>
2026-07-01 21:29:48 -07:00