Commit Graph
174 Commits
Author SHA1 Message Date
Neil 673d7ca926 refactor(relay): collapse the duplicated FrameDecoder into one shared module (#12078)
src/relay/relay-frame-decoder.ts and src/main/ssh/relay-frame-decoder.ts
were 264 identical lines apart from one default: the relay logs decode
faults to stderr when no handler is supplied, the SSH side stays silent.
Two copies of framing logic is exactly where a wire-format fix lands in one
and not the other.

The decoder's contract and buffer already live in src/shared, so the class
joins them there. The relay keeps a thin subclass that supplies its stderr
default, preserving behaviour for the call sites that omit onError. The
SSH copy is deleted and relay-protocol.ts points at shared directly.

Verified: pnpm typecheck, 102 tests across the 9 framing/backpressure/
handshake suites, and `pnpm build:relay` for all six platform targets plus
the WSL hook relay — the standalone bundle has no new dependencies.
2026-08-02 00:58:10 -07:00
NeilandOrca 73c5009b82 chore(dead-code): drop ~2k lines of unreachable exports and orphan modules (#12077)
* chore(dead-code): drop 2k lines of unreachable exports and orphan modules

Ran knip across every build entry (main, preload, renderer, popout, web,
cli, relay, workers, forked sidecars, config scripts) and removed what no
entry graph can reach.

- 11 orphan modules nothing imported, plus one test that only covered them
- 159 unused exports/types, with their now-dead helpers, imports and tests

Each candidate was verified against dynamic references before deletion.
42 knip hits were false positives and are kept: shared modules consumed by
the mobile/ workspace, the src/shared/plugins/** public API, vendored
shadcn primitives, and relay wire-protocol constants held for compatibility.

Adds knip.json + `pnpm audit:dead-code` so this stays measurable.

Verified: pnpm typecheck, pnpm lint, and 2081 tests across the 73 affected
test files all pass.

* chore(dead-code): move knip config under config/

Root-level additions are blocked by the root directory guard.

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Orca <help@stably.ai>
2026-08-02 00:33:57 -07:00
Jinjing 25fefa4072 fix(P1-A): async SSH consumer-recovery persistence and detach on failed connect (#12026)
* fix(P1-A): persist SSH consumer recovery without a sync store flush

rememberPtyConsumerRecovery ran on the live establish/reconnect path and
called flushOrThrow -> writeToDiskSync, parking the Electron main thread on
the profile-directory write. On a stalled or slow profile mount that freezes
the whole app during SSH recovery and reconnect.

Add Store.flushAsync(): same debounce-cancel and write serialization as
flushOrThrow, but awaits writeToDiskAsync instead of blocking. The consumer
recovery upsert/remove pair is now async and awaits it, and the SSH callers
await through to establish()/reconnect() so ownership is still durable before
relay setup continues. In-memory state still mutates synchronously (before the
first await), so no caller can observe a torn record and dispose() stays
synchronous.

* fix(P1-A): detach the SSH session when a connect attempt fails

Both failure exits in doConnect dropped the session from activeSessions
without calling detach(). claimSshPtyConsumerRecovery only reuses an existing
in-memory entry when detached === true, so the next connect attempt fell
through to minting a fresh clientInstanceId, discarding the remembered owner
lease and its resume identity.

Route both exits through abandonFailedSshSession(), which detaches (keeping
PTY ownership, unlike dispose()) before removing the session, and tolerates a
teardown throw so it can't mask the connect error being rethrown.

* fix(P1-A): await async lease persistence in SSH relay teardown

Failed connect attempts now wait for 'detached' leases to persist before
throwing, preventing reconnects from claiming them before cleanup completes.
Detach and dispose operations are now async and await store durability.

* fix(ssh): make session detach lease writes retryable on failure

Separate in-memory detach (identity recovery, provider cleanup) from lease
write persistence so rejected writes can be re-issued without re-running
provider teardown or re-minting the session identity. Introduce
flushDurableStateOrThrowAsync to flush only SSH-recovery state on the
live establish/reconnect path, avoiding snapshot writes of sidecars that
belong to quit/startup. Use Promise.allSettled in test reset to prevent
one rejected disposal from leaking state into the next test.

* fix(ssh): dispose mux on failed establish and propagate sync errors

- Dispose mux when session is disposed during establish to prevent resource leak
- Propagate synchronous errors in teardown via the completion promise instead of leaving completion undefined
- Add test coverage for terminated PTYs that exit mid-reattach and must stay dead
2026-08-01 23:16:24 -07:00
Jinjing ce5b639e03 fix(P1-B): recover SSH targets and remote file watchers after a network drop (#12032)
* fix(P1-B): recover system-SSH targets after a network drop

Two defects stopped a remote workspace auto-recovering after a blip.

runReconnectAttempt classified failures with isTransientError, which only
matches ETIMEDOUT/ECONNREFUSED/ECONNRESET by errno code or literal
substring. The system-SSH transport — the only transport FIDO2 and
ProxyUseFdpass targets can use — reports network failures as OpenSSH
prose ("System SSH connection timed out"), so the ladder published a
permanent 'error' on the first timeout and the target never came back
without a manual reconnect. isTransientReconnectError adds a
network-shaped prose table on top of isTransientError and is used only on
the reconnect path: connect() keeps the narrow classifier so an
unreachable host still fails fast instead of burning five 30s attempts
and five security-key touch prompts. Auth and passphrase failures stay
permanent on both paths.

runReconnectAttempt also had no generation fence, so a superseded attempt
published its cancellation as a permanent error over the winner's live
connection — reachable when a system-transport proc.onExit schedules a
reconnect while an attempt is still in flight. Cancellation now carries a
stable error name, and both connect() and runReconnectAttempt claim their
connectGeneration and stay silent when a newer attempt owns the state.

* fix(P1-B): retry a dropped watcher overflow marker on real capacity

emitWatcherOverflowToClient published the {kind:'overflow'} resync marker
with controlOverflow:'reject'. A full control queue rejects at admission
with no settlement callback, so the marker was silently discarded and the
remote File Explorer stayed stale until some later watcher event happened
to produce another one — for a quiet tree, possibly never.

The emitter now retains a rejected marker per (client, root) and
republishes it when the sink actually frees up. The existing
onLegacyPtyCapacity signal cannot drive that: it is gated on producer
retention, so it stays silent exactly under the dual-queue pressure that
caused the rejection. RelayDispatcher.onClientCapacity is an ungated
per-client capacity signal that fires on every writer settlement and
drain. It lives on the dispatcher rather than the writer so a retained
marker survives setWrite() replacing the primary sink, and setWrite
notifies capacity once afterwards so the marker does not wait on traffic
that may never arrive.

Retention is bounded to one marker per (client, root), released on
settlement and purged on client detach.

* fix(P1-B): address all review findings on SSH network recovery

Fix four issues from code review:

1. **Bug — admitted overflow markers lost on setWrite**: Retain markers when
   settlement fails `ok: false`, not just on admission rejection. Prevents
   desynced filesystem trees after SSH sink replacement.

2. **SSH error classification expanded**: Add missing OpenSSH patterns
   (`ssh_exchange_identification`, `connection closed by remote`) and new
   `isDefiniteSystemSshHostFailure()` classifier.

3. **ControlMaster retry optimization**: Skip second probe when first failure is
   already definite host-level (network timeout, refused, unreachable). Saves
   ~30s per reconnect ladder step.

4. **Overflow flush under dual-queue pressure**: Gate pending marker retries on
   control-lane headroom instead of re-attempting on every capacity notification.
   Reduces thrash proportional to producer traffic.

Add regression tests for marker republish on sink replacement and validate auth
error detection against live OpenSSH credential rejection messages.

* rm random doc

* fix(P1-B): skip credential-failure retries and recover watcher markers o

- Auth and passphrase errors fail immediately without retry attempts
- Bare "System SSH probe failed (exit 255)" is transient only for reconnect
- Watcher markers survive client invalidation when switching SSH connections
- Add network error patterns: "lost connection", "remote end closed"
2026-08-01 22:21:37 -07:00
NeilandOrca a20d82294b fix(agent-status): preserve Claude background work (#11838)
* fix(agent-status): preserve Claude background work

* fix(agent-status): harden background task lifecycle

* fix(agent-status): narrow interruption retention

* fix(agent-status): scope background task authority

* fix(agent-status): isolate lifecycle inventories

* fix(agent-status): harden background evidence recovery

* fix(agent-status): reject ambiguous child authority

* perf(agent-status): skip lifecycle inventory scans

* refactor(agent-status): isolate task inventory parsing

* fix(agent-status): clear stale background evidence

* fix(agent-status): gate accepted remote evidence

* test(agent-status): pin session cron interrupts

* fix: harden Claude inventory tracking

* test: pin Claude cron drain authority

* refactor(agent-status): unify Claude turn-boundary predicate

Collapse the five inline copies of the Stop/StopFailure test into a single
isTurnBoundary constant and drop the reportedStateName/stateName alias, so a
future edit can't move one copy and leave the others behind.

Pin the two behaviors that unification now depends on: a non-interrupted
StopFailure keeps gating on live background work, and interrupted state does
not survive a mid-turn lead event that has no prompt submit.

Co-authored-by: Orca <help@stably.ai>

* fix(agent-hooks): gate local Claude background evidence

---------

Co-authored-by: Orca <help@stably.ai>
2026-08-01 21:10:08 -07:00
Jinjing de75003df9 fix(P1-C): gate FIDO2 system-SSH transport on an OpenSSH binary (#12029)
* fix(ssh): gate FIDO2 system-transport on an OpenSSH binary

`ssh -G` echoes OpenSSH's built-in default identity list for every host, so
`usesDefaultPaths` was almost never true and the security-key gate returned
`!usesDefaultPaths || findSystemSsh() !== null` — forcing system transport
without checking that an `ssh` binary exists. `spawnSystemSsh()` then throws
`No system ssh binary found`, hard-failing connections that worked on ssh2.

The same flag also stopped the default scan at the first existing normal
private key, so a host that only accepts a FIDO2 key never reached system
OpenSSH when `~/.ssh/id_rsa` happened to exist.

Both decisions are independent of where an identity path came from: always
require `findSystemSsh() !== null` before forcing system transport, and scan
every candidate identity instead of stopping on the first normal key.
`shouldUseSystemSshTransport()` is untouched, so ProxyCommand / ProxyJump /
ProxyUseFdpass keep their intentional system transport.

* test(ssh): isolate connection tests from the developer's own FIDO2 keys

Transport selection now scans every default identity instead of stopping at
the first normal key, so a `~/.ssh/id_ed25519_sk` on the machine running the
suite would decide which transport the default-target tests take. Mock
`findSystemSsh` to null by default and opt the two security-key tests in.
2026-08-01 18:45:37 -07:00
JinjingandOrca a07427e970 fix(ssh, relay): keep remote sessions alive through reconnects and backpressure (#11999)
* fix(ssh,relay): stop remote connections from being killed by backoff and frame caps

Three independent connection killers found in the SSH/remote freeze audit.

FINDING A - the reconnect ladder never escalated for post-handshake drops.
scheduleReconnect() used the single published state.reconnectAttempt for both
the delay index and the give-up test, and runReconnectAttempt() zeroed it
before connecting (ssh.ts gates the relay redeploy on 0-at-connected). Every
post-handshake drop therefore re-entered at 1000ms forever, ~3600 relay
redeploys/hour, and 'reconnection-failed' was unreachable for a flapping host.
New SshReconnectLadder splits the delay index (advanced by every retry) from
the failure streak (advanced only by a failed handshake), so flaps back off
while give-up semantics stay byte-identical to shipped.

FINDING B - notify() closed the client whenever a frame exceeded the producer
frame capacity, conflating a permanently un-sendable frame with transient
backpressure. A 5000-event fs.changed is 425KB against a 49KB cap, so the
watcher flood killed the link and re-killed on every reattach+replay. notify()
now drops and logs once per generation; fs.changed is chunked to each sink's
capacity with a control-lane overflow marker as the resync fallback; agent-hook
envelopes shed lastAssistantMessage/interactivePrompt/subagents to fit.

FINDING B2 - sendResponse routed >1MB responses to a lane whose admission
ignores the frame cap and closed the client on rejection, so a large
fs.listFiles dropped the SSH host. It now substitutes a JSON-RPC error so the
request fails instead of the connection.

Also moves fs.streamEnd/fs.streamError to the control lane so a terminal frame
cannot be dropped by the producer-lane check.

Co-authored-by: Orca <help@stably.ai>

* fix(relay): stop the overflow marker from re-killing the link it protects

Round-1 review fixes on the P0 freeze work.

The control-lane overflow marker could reinstate the exact failure this P0
removes: dispatcher-client-writer closes the client when control-lane
admission fails, and admitControl is the only lane that returns an error, so
one marker per failing batch accumulated to the 256-frame/1MB bound and
dropped the link. Markers are now deduped to one outstanding per
(client, root), cleared on settle.

Chunking also defeated the renderer's per-payload directory dedupe -- events
are now stable-grouped by parent directory so one directory lands in one
chunk -- and the halving walk overshot the byte minimum ~1.7x while the fast
path paid three JSON encodes; both are fixed by publishing first and sizing
from a measured bytes-per-event estimate.

Agent-hook shedding now surrenders the blocking interactive prompt LAST
rather than first, so a degraded envelope cannot strand a pane at
state=waiting with no answerable question card.

The dropped-notification log now distinguishes over-capacity from producer
queue backpressure and no longer lets the first dropped method silence every
other producer for the life of the connection.

* fix(relay,ssh): keep status delivery and terminal frames from trading one freeze for another

Round-2 review fixes.

The round-0 change from close-on-rejection to silent drop removed the only
redelivery path for agent.hook envelopes: they are fire-and-forget and the
per-pane cache only replays on handler install, so a saturated link stranded
a pane on a stale Working spinner until reconnect. Closing used to guarantee
delivery by forcing that replay. Envelopes now publish per client and pend
for bounded latest-wins redelivery when the producer queue rejects them.

Shed fields are now named on the wire. The subagent roster is not cosmetic --
the renderer replaces rather than merges it, and hibernation gates on its
length -- so an unmarked shed could sleep a live pane.

fs.streamEnd rode the control lane because it must not be dropped, but that
lane kills rather than drops. The stream's concurrency slot is now held until
the terminal frame settles rather than until the fd closes, capping queued
terminal frames well under the control budget; overflow costs one refused
read instead of the connection.

The watcher chunk walk now stops while producer retention sits past its
reserve and degrades to a resync, so a 5000-event flood cannot fill the queue
that interactive PTY traffic shares and stall every remote terminal.

The reconnect ladder caps its flap-path delay so delay plus handshake timeout
cannot cross the relay grace floor and let the remote daemon kill live PTYs.

Also: the suppression key no longer embeds a NUL byte, which had made the
file binary to git and grep; producerEnvelopeBudget no longer reports
infinite capacity for a departed client; the drop logger no longer encodes a
frame it will not log; and an over-capacity response substitution no longer
settles as if the result had been delivered.

* fix(relay,ssh): restore relay-shed status fields and scope backpressure per client

Round 3 + 4 review fixes.

Watcher chunking is now gated on the *client's* retention reserve rather than
the dispatcher-wide one, so one stalled peer no longer forces a healthy client
into a full file-tree resync. The relay-lost redeploy ladder no longer burns its
6-attempt budget while the SSH transport itself is down: it holds at the 15s step
with a non-terminal status and rearms, so a laptop that slept past the ladder
comes back instead of landing on a terminal "give up" banner.

The shedFields wire marker had no consumer, so an agent-hook envelope whose
subagent roster was dropped to fit the frame read as "roster cleared" on the Orca
side: live child rows blanked and a done pane became hibernation-eligible while
its teammates were still running. ingestRemote now restores shed fields from the
cached payload (interactivePrompt deliberately excluded — a stale answerable
question card is worse than none).

Also: stream terminal-frame slots are counted per client, since the control queue
they protect is per client; the chunking fast path no longer logs a drop for a
batch it goes on to deliver in full; -32010 is now RelayErrorCode.ResponseOverCapacity.

Test debt from the review: pending-pane eviction, per-client stream isolation, and
the reconnect budget are now asserted rather than assumed; four fragile exact-byte
pins dropped in favour of the tier comparisons that carry the requirement.

* fix(relay,ssh): restore relay-shed status fields and scope backpressure

- Oversized relay responses now fail their request instead of closing the connection,
  preventing one frame from killing every pane on the host
- Restore subagent state for correct hibernation; don't resurrect stale prose
  across turns
- Account for relay re-establishment and PTY reattach time in SSH flap delay caps
- Only log drops of final unsendable envelopes, not temporary rejections during
  measurement probes
- Fix watcher overflow marker release race when notification admission rejects
  without settlement; use precise byte counting for event batching

* Restore relay-shed fields with digest validation and scoped backpressure

Validate that shed subagent rosters match their wire digest and turn identity before
restoration, preventing stale roster resurrection. Compact interactive prompts for waiting
states instead of dropping them. Demote control-queue overflow to non-fatal rejection so
clients can retry on capacity recovery, keeping the link alive during transient backpressure.

* fix(relay): correct ResponseOverCapacity error code

ResponseOverCapacity should use -33008 to stay in the -33xxx range
for relay protocol errors, not -32010.

* fix(relay): close client when pty.replay overflows control queue

Replay is never retried, so it uses the control lane where overflow
is fatal — the writer closes the client and reconnect reloads history
rather than stranding a short buffer.

* fix(relay): prevent infinite redeploy on flapping SSH transports

Charge reconnect attempts when connection restores mid-backoff, preventing
infinite loop on transports that flap between states. Refactor control overflow
handling to use entry property instead of WeakSet marker for clarity.

---------

Co-authored-by: Orca <help@stably.ai>
2026-08-01 14:27:16 -07:00
Neil 33c14bc716 fix(ssh): fall back to OpenSSH for FIDO2 keys (#11913)
Closes #11645
2026-08-01 03:27:42 -07:00
Brennan Benson ed00ab0f34 fix(ssh): restore relay ownership after app restart (#11860) 2026-07-31 20:45:52 -07:00
Rod BoevandOrcaWin f56e6ade80 fix(ssh): recover orphaned relay install locks (#9828) (#10207)
* fix(ssh): recover orphaned relay install locks (#9828)

* test(ssh): split staged upload relay specs (#9828)

* fix(ssh): verify staged relay upload namespace

* fix(ssh): bound stale relay stage cleanup

* fix(ssh): complete bounded stage recovery

* fix(ssh): generate valid PowerShell stage scripts

* fix(ssh): make staged upload cancellation safe

* fix(ssh): fence staged relay recovery

* test(ssh): align deploy timeout oracle

---------

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-07-31 16:17:37 -07:00
NeilandOrca 390ae08232 [P1] fix(ssh): key the PTY model-migration fence by app pty id (#11617)
* fix(ssh): key the PTY model-migration fence by app pty id

Co-authored-by: Orca <help@stably.ai>

* test(ssh): pin the post-recovery checkpoint rekey

The finishSourceRecovery app-id rekey had no coverage: reverting it left every
suite green while reconnects silently resumed from the stale migration-era
checkpoint.

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Orca <help@stably.ai>
2026-07-30 17:49:08 -07:00
Henry SuandOrcaWin 5fe3aaf2b7 fix(ssh): preserve first config directive value (#11297)
* fix(ssh): preserve first config directive value

* test(ssh): cover false-first config booleans

* fix(ssh): trust fresh OpenSSH config authority

* fix(ssh): preserve ordered config identities

---------

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-07-30 14:23:30 -07:00
Brennan Benson 53430e34d6 fix(test): stabilize system SSH transport integration (#11597)
* fix(test): stabilize system SSH transport integration

* fix(lint): extract terminal display mode predicate

* fix(test): exercise fake relay socket bridge
2026-07-30 13:59:50 -07:00
650dd48ec9 feat(cli): add orca account add / account list for headless hosts (Claude + Codex) (#9177)
* feat(cli): add `orca account add` / `account list` for headless hosts

The desktop "Add account" UI is disabled when the renderer drives a remote
runtime (isRemoteAccountScope === kind:'environment'), so a headless server
reached from a remote desktop/web client has no way to register managed
Claude accounts. Add a host-local CLI path that reuses the existing capture
logic:

- ClaudeAccountService.addAccountFromConfigDir(): register a managed account by
  capturing credentials from an already-authenticated CLAUDE_CONFIG_DIR instead
  of spawning the interactive browser login (extracted persist/rollback helpers
  shared with the existing add flow)
- RPC accounts.addClaudeFromConfigDir, bridged via OrcaRuntime; rejected for
  mobile device tokens (host-local only)
- `orca account add` runs `claude login` in the user's own terminal into a temp
  CLAUDE_CONFIG_DIR, then registers it via the local runtime; `orca account list`
  lists managed accounts

Switching (select) already works from a remote client; only adding was blocked.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(cli): support Codex in `orca account add` / `account list`

Mirror the Claude headless-account CLI for Codex:

- CodexAccountService.addAccountFromHome(): register a managed Codex account by
  importing auth.json from an already-authenticated CODEX_HOME, reusing a shared
  persist helper extracted from doAddAccount (no interactive login spawned here)
- RPC accounts.addCodexFromHome + OrcaRuntime.addCodexAccountFromHome bridge,
  rejected for mobile device tokens (host-local only)
- `orca account add --agent claude|codex` (default claude); `orca account list`
  now renders both Claude and Codex managed-account blocks

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test: cover headless account-add capture paths (Claude + Codex)

- ClaudeAccountService.addAccountFromConfigDir: registers a managed account by
  capturing an authenticated CLAUDE_CONFIG_DIR; rejects and rolls back when the
  dir has no .credentials.json
- CodexAccountService.addAccountFromHome: imports auth.json from an
  authenticated CODEX_HOME into a managed account; rejects when auth.json is
  missing

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix: address CodeRabbit review on headless account-add flows

- CLI login spawn uses a shell on Windows so `.cmd` agent shims resolve without
  ENOENT (args are fixed literals, no injection risk)
- Claude capture skips the `.credentials.json` precheck on macOS, where creds
  live in the Keychain and captureAuthFromConfigDir reads them
- Claude add rollback is best-effort: a failed rematerialization no longer skips
  managed-auth cleanup or masks the original add error
- Codex persist restores the prior account/selection if a post-write sync or
  rate-limit refresh fails, so a failure can't leave a dangling managed account
- Codex sync passes the account's selection target (correct runtime for WSL)
- Add JSDoc to the new public service methods and CLI functions

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(cli): harden headless account capture

* fix(cli): correct account command flag surface and interrupt cleanup

- `account` commands no longer accept or advertise the browser `--page`
  flag; `supportsBrowserPageFlag` allow-listed them by omission, so
  `orca account list --page x` was silently accepted and `--help`
  rendered a browser-only option
- account specs declare GLOBAL_FLAGS, so `--help`/`--json` render in the
  Options block like every other command
- `--agent` on `account add` documents the account provider instead of
  the terminal TUI-agent meaning inherited from the shared flag table
- a SIGINT/SIGTERM during the interactive login now removes the temp
  login dir (and restores the macOS Keychain item) before exiting 130;
  Node terminates without unwinding `finally`, which stranded live OAuth
  credentials on disk

* perf(cli): stop `account list` forcing a provider usage refresh

`accounts.list` awaited refreshAccountsForMobile(), which runs
fetchAll({ force: true }) — bypassing both the poll throttle and the
per-provider Retry-After gate — then O(N) serial per-account round
trips. `orca account list` renders only emails and the active ids, so
all of that work was discarded. The RPC now takes `refreshUsage`
(default true, so mobile and web keep the forced lane) and the CLI opts
out. Older hosts declare `params: null` and ignore the field, so a newer
CLI degrades to the previous behavior rather than failing.

Also documents on `account list` that `--environment` does not retarget
it, matching the host-local behavior of shouldIgnoreRemoteSelection.

* fix(cli): survive repeated and hangup signals during account add

withInterruptCleanup latched cleanup behind a boolean, so a second signal
got an already-resolved promise and its process.exit fired while the first
cleanup was still inside a Keychain call (3s each) — the temp dir's OAuth
credentials and the swapped macOS Keychain item both survived. Memoize the
cleanup promise so every signal awaits the same run, and register with
`on` instead of `once` so a second Ctrl-C cannot fall through to Node's
terminate-immediately default mid-cleanup.

Handle SIGHUP too. This flow exists for headless/SSH hosts, where the most
likely interrupt is the connection dropping, which hangs up the login's
terminal and previously ran no cleanup at all.

Warn when the interrupt lands after sign-in completed: the runtime finishes
the add independently of this process, so exiting 130 silently would tell
the user it was cancelled when the account may exist.

Reject a valueless `--agent`; the parser turns it into boolean true, which
silently ran a full OAuth login for Claude when the user asked for another
provider.

Also lock two behaviors the refactor changed but left uncovered: a WSL Codex
add must sync the WSL runtime lane rather than the default host lane, and
rename the account-spec help test to describe the Options block it actually
asserts rather than the usage string it never reads.

* fix(build): bundle the main modules the account CLI imports

electron-vite cleans out/main and emits only its declared entries, and
`build:desktop` runs it after `build:cli`, so the tsc-emitted copies of
`claude-accounts/keychain`, `codex-cli/command` and `win32-utils` were
deleted before packaging. Both `orca account add` and `orca account list`
then died at require time with "Cannot find module
'../../main/claude-accounts/keychain'" — reproduced against a real
`--serve` host. `agent-hooks/managed-agent-hook-controls` already carried
an entry for exactly this reason; these three were missing.

Adds a parity test so any future CLI import of a `src/main` module fails
in CI rather than at a user's shell after packaging.

* test: cover the desktop add-path behavior this PR changes

Both changes ride in the persist/rollback helpers the existing GUI add
flow shares with the new headless path, and neither had coverage:

- Claude: rollbackAddAccount now guards forceMaterializeCurrentSelection-
  ForRollback, so a rejecting rematerialization no longer replaces the
  real add error nor skips safeRemoveManagedAuth. Asserts the original
  error surfaces and the throwaway auth dir is gone.
- Codex: the desktop add now passes the account's selection target to
  syncForCurrentSelection, matching reauthenticate and select. Asserts
  the host target alongside the existing WSL assertion.

Both fail when the corresponding change is reverted.

* fix(cli): close the remaining account-add interrupt and preflight gaps

The round-1 interrupt fix detached the signal handlers before running the
finally-path cleanup, so the very window it was meant to protect — the two
serial 3s `security` calls plus rmSync on the success/error path — was
still covered only by Node's terminate-immediately default. Both review
lanes reproduced it independently. Await cleanup first, detach in a nested
finally, and stop a cleanup failure from replacing the error that actually
explains why the add failed.

Do not burn the interactive login when the runtime is unreachable. The
RuntimeClient is lazily constructed and the first call was the registration
RPC itself, so "Requires the Orca runtime to be running" was discovered
only after the user completed a full OAuth round trip. Preflight with the
now-cheap `accounts.list { refreshUsage: false }`.

Reject `--environment` / `--pairing-code` on `account add`.
shouldIgnoreRemoteSelection pins account commands to the local runtime, so
`orca account add --environment homelab` silently registered the account on
the laptop instead of the headless host it names.

Survive a daemon that cannot spawn `claude`. `allowFailure` is honored in
onClose but not onError, and unlike the GUI flow nothing has run `claude` in
the daemon before this point — so a launchd/systemd daemon with a minimal
PATH hard-failed an add the user had already signed in for, even though
identity resolves fine from the config dir's oauthAccount.

Also align the `--agent` help description with the global flag column.

* fix(cli): reject runtime selectors on `account list` too

`orca account list --environment homelab` was accepted and silently
listed the LOCAL machine's accounts, because shouldIgnoreRemoteSelection
pins account commands to the local runtime. Documenting that in --help
does not reach someone who already typed the flag, and answering with the
wrong host's accounts is the specific wrong answer they would act on.

`account add` already errors; this makes the new command group internally
consistent. The other groups in shouldIgnoreRemoteSelection keep their
existing silent-ignore behavior — changing those is not this PR's job.

* test: harden account-add signal tests and cover cleanup failure

- Identify the handler under test by set difference instead of
  `process.listeners(sig).at(-1)`. Vitest installs its own once-wrapped
  SIGINT teardown, so the positional lookup could grab the wrong listener;
  the helper also asserts exactly one new listener was added.
- Mock rmSync while keeping the real implementation by default, so the
  temp-dir assertions elsewhere stay honest.
- Cover that a cleanup failure in the `finally` does not replace the error
  explaining why the add failed. Fails when that guard is removed.

Completes the review loop's final round; the loop died on an API error
before it could commit this, and its `import()` type annotation would
have failed oxlint.

* fix(cli): harden interactive account add

* test(cli): make account cancellation coverage portable

* fix(cli): preserve merged skills runtime modules

---------

Co-authored-by: Dominik <marketing@gavaplast.sk>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
2026-07-30 12:50:07 -07:00
Brennan Benson f8b553b7d5 fix(agent-hooks): skip unavailable agent homes (#11442)
* fix(agent-hooks): skip unavailable agent homes

* refactor(agent-hooks): separate Pi and OMP home fix

* test(agent-hooks): update merged protocol harnesses

* fix(agent-hooks): avoid redundant reconciliation

* fix(agent-hooks): harden reconciliation and detection

* test(agent-hooks): cover settings reconciliation

* fix(agent-hooks): hydrate PATH for paired clients
2026-07-29 20:19:18 -07:00
JinjingandOrcaWin 5f7807497e feat(ssh): bound relay PTY output end to end (#11005)
* docs: design SSH relay PTY backpressure

* fix(ssh): bound relay frame decoding

* fix(relay): bound PTY output publication

* fix(ssh): bound PTY model admission

* fix(ssh): settle closed model admissions

* feat(ssh): negotiate bounded PTY consumer sessions

* fix(ssh): fence exit on renderer settlement

* feat(ssh): track PTY source credit end to end

* fix(ssh): recover bounded PTY output across reconnect

* feat(ssh): complete relay PTY output backpressure

* fix(ssh): close final PTY source credit races

* docs(ssh): record final backpressure validation

* feat(ssh): complete relay PTY source-credit lifecycle

* test(ssh): complete provider notification fixture

* fix(ssh): preserve terminal source credit across rotation

* fix(ssh): fail closed on recovery cancellation

* fix(ssh): prioritize mux control writes after drain

* fix(ssh): retire canceled relay restore deliveries

* fix(ssh): order exit cancellation cleanup

* fix(ssh): gate provisional source activation

* test(ssh): register mux drain-priority coverage

* fix(ssh): type stale owner recovery mismatches

* fix(ssh): close projection replacement races

* fix(relay): contain streaming edge failures

* fix(ssh): secure relay endpoint credentials

* docs(ssh): reconcile final backpressure lifecycle

* fix(ssh): bound main IPC output lifecycle

* fix(ssh): close recovery ownership gaps

* docs(ssh): record exact artifact validation

* fix(ssh): reject reclaimed snapshot replacements

* fix(ssh): fence model admission across reconnect

* fix(ssh): contain migration failure per PTY

* docs(ssh): record final exact-head validation

* test(ssh): align deploy fixtures with credential publication

* feat(ssh): add per-target bounded output setting

* fix(ssh): close source recovery review gaps

* fix(ssh): latch source credit environment override

* feat(ssh): make PTY source credit the default

* docs(ssh): record always-on relay validation

* docs(ssh): bind validation to current main

* test(ssh): grant source credit in IPC fixture

* test(ssh): grant source credit in fake relay

---------

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-07-29 17:03:15 -07:00
OrcaWinandOrcaWin 363e478909 fix(orchestration): preserve active workers across updates (#11271)
* fix(orchestration): preserve active workers across updates

* test(ssh): model absent legacy adoption

* test(orchestration): align compatibility contracts

* fix(windows): escape updater PowerShell booleans

* fix(windows): restore stock uninstall process check

* fix(orchestration): keep recovery off renderer startup barrier

* fix(orchestration): harden legacy recovery migration

* fix(orchestration): close recovery review gaps

* fix(orchestration): complete legacy worker cutover recovery

* fix(orchestration): preserve legacy workers across updates

---------

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-07-29 11:31:35 -07:00
JinjingandOrcaWin a40183389b feat: bound direct SSH reconnect fan-out and recovery (#11003)
* docs: design for direct SSH reconnect fan-out

Capture the implementation-ready plan for host-qualified, epoch-fenced
SSH reconnect recovery after two rounds of multi-model LLM counsel review.

* docs: reconcile SSH reconnect fan-out design

* docs: close reconnect design consistency gaps

* feat: implement bounded direct SSH reconnect recovery

* fix: bound direct SSH retry settlement

* fix: harden direct SSH reconnect authority

* fix: preserve split SSH retry ownership

* fix: preserve SSH split continuation authority

* docs: record final SSH reconnect validation

* fix: preserve SSH authority through retained and detached state

* fix: retain SSH authority across delayed split mounts

* fix: close SSH authority recovery gaps

* fix: fence stale SSH transport replacement

* fix: serialize SSH target teardown

* fix: settle SSH teardown failures before reconnect

* fix: retire failed SSH reset sessions

* test: reconcile current main E2E contracts

* fix: close direct SSH reconnect review gaps

* fix: fence stale SSH reconnect side effects

* fix: close final SSH reconnect lifecycle gaps

* test: stabilize current-main reliability gates

* test: prove plugin navigation containment

* test: make plugin navigation oracle authoritative

* test: make plugin navigation oracle deterministic

* ci: allow sharded e2e suite to finish

* test: wait for runtime pane publication

* test: classify pane readiness by error code

* test: select close persistence terminal by tab identity

* docs: mark reconnect implementation validated

---------

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-07-28 12:33:17 -07:00
Jinwoo HongandOrcaWin 0d6f9195d8 fix(orchestration): reveal worker terminals reliably (#11142)
Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-07-28 01:59:49 -07:00
Neil badf91101b fix(quality): enforce performance-safe lint baseline (#11074)
* fix(quality): clear safe existing lint findings

* fix(quality): keep lint cleanup allocation-free

* fix(quality): enforce performance-safe baseline

* test(terminal): drain deferred confirmation cleanup
2026-07-27 20:54:02 -07:00
Brennan Benson 72875bda24 fix(ci): stabilize flaky terminal and SFTP tests (#11018) 2026-07-27 17:40:46 -07:00
OrcaWin cd05f2ff93 Implement robust orchestration primitives and connected-server workers (#9925) 2026-07-27 12:31:37 -07:00
Brennan Benson 86a993aec4 fix(ssh): connect to Linux hosts that cannot compile node-pty (#10776)
* fix(ssh): connect to Linux hosts that cannot compile node-pty

node-pty ships no Linux prebuilt at any architecture, so it is compiled on
the remote. On a host without a C/C++ toolchain that build fails, and because
both native deps install in one npm command it also took down
@parcel/watcher — which does have a working Linux prebuilt — and failed the
whole connection. Every Linux image without build tools was unusable.

node-pty only backs remote terminals; files, git, and the editor do not need
it, and a missing native dep is already non-fatal further down the deploy. So
when the existing toolchain probe confirms the compiler is missing, reinstall
without node-pty instead of aborting. The manifest has to drop it too — npm
reconciles every dependency in package.json, not just the ones named on the
command line, so naming only @parcel/watcher still rebuilds node-pty.

If that reinstall also fails the actionable build-tools error is rethrown, so
a host broken for some other reason still reports the toolchain gap.

The relay's PTY error now names the fix rather than saying only that node-pty
is unavailable.

Verified on a stock Rocky Linux 10.2 aarch64 container (openssh-server, git,
nodejs, npm, no compiler): connect succeeds, /etc lists over SSH, node-pty is
absent while @parcel/watcher installs its linux-arm64-glibc prebuilt, and
spawning a terminal reports the install hint.

* fix(ssh): keep the node-pty skip path honest about platform and watcher

The PTY unavailable message named build tools unconditionally, but only Linux
compiles node-pty — the deploy-side skip is gated on linux and the toolchain
probe returns null on Windows. A Windows or macOS remote, where node-pty ships
prebuilds, was told to install make/g++/python3. Pick the remedy by the relay's
own platform.

The skip path returned before the install probe, so a @parcel/watcher that
installs but cannot require() (glibc below the floor) connected with dead file
watching and nothing logged. Probe before returning and warn; no rebuild, since
node-pty provably cannot compile on that host, and never fatal.

Also log the pty-less reinstall's own failure and attach it as cause — the
rethrown toolchain message is built from the original npm error, so an
unrelated retry failure (registry, ENOSPC, EACCES) was lost. The reinstall now
keeps the caller's resetDeps as well, so a repair reconnect still clears every
dep the probe found broken.

Tests: the skip-success fixture queued a chmod/probe/rebuild sequence
production never runs, and the surplus slots were absorbed by launchRelay's
readiness poll (1817ms vs 3-9ms for its peers). It now emits exactly the 12
execs production performs, and pins that no rebuild is issued. Adds the missing
negative case: a gyp-shaped failure on a host whose probe reports a complete
toolchain must still hard-fail rather than silently degrade.

* fix(ssh): hedge the node-pty remedy and keep repair resets on the skip path
2026-07-26 15:29:13 -07:00
Jinjing 76b2a3b44d fix(cli): bound orchestration ask timeouts (#10689)
* fix(cli): bound orchestration ask timeouts

* fix(cli): harden remote timeout boundaries
2026-07-26 12:50:05 -07:00
Jinjing 9f81f97a0c Fix SSH relay installs on split shell/SFTP namespaces (#10645)
* Fix ssh-relay install on hosts with split shell/SFTP namespaces

On Synology DSM and similar hosts, the SSH shell and SFTP subsystem expose
different absolute paths for the same directory (e.g., /var/services/homes/alice
vs /homes/alice). The relay installer silently picked the wrong path and failed
discovery. This fix implements SFTP namespace detection: each install creates an
unguessable ownership marker and probes both namespaces to detect divergence.
When paths differ, SFTP writes redirect to the candidate namespace while shell
commands keep the canonical path. Markers are random tokens redacted from logs.

* Fix ssh-relay install on hosts with split shell/SFTP namespaces

Strengthen path validation to catch traversal and empty segments in
absolute POSIX paths, preventing security issues. Improve split-namespace
handling with comprehensive wire tests for uploads and file writes.
Ensure system SSH connections bypass namespace mapping entirely rather
than attempting incorrect retargeting.
2026-07-25 19:20:17 -07:00
6d39e49480 fix(ssh): accept GitHub restricted-shell SSH probes (#6988) (#7659)
* fix(ssh): accept GitHub restricted-shell SSH probes (#6988)

* fix: match first stderr line for GitHub restricted-shell probe (bug-bash takeover)

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
Co-authored-by: Orca <help@stably.ai>
2026-07-24 00:24:53 -07:00
NeilandOrca aab112933e Revert "fix(memory): bound OOM-prone accumulators (#10179)" (#10255)
Co-authored-by: Orca <help@stably.ai>
2026-07-23 18:35:31 -07:00
Neil 8f40ddf328 fix(memory): bound OOM-prone accumulators (#10179) 2026-07-23 06:22:56 -07:00
Jinjing fc05769edf fix(ssh): allow blank user in VS Code authority (#10072) 2026-07-22 21:10:05 -07:00
OrcaWin 41751dd90d fix(runtime): route HUB-owned SSH worktrees through owning runtime (#9994) 2026-07-22 18:25:05 -07:00
Jinjing adc020393e feat(ssh): open SSH workspaces in VS Code Remote-SSH (#10005)
Allow the existing "Open in" entries to launch a configured VS Code
launcher against an SSH-backed worktree via Remote-SSH:

    code --remote ssh-remote+<authority> <remote-path>

- Split the blanket SSH/runtime block into a capability model: file
  managers and non-VS Code launchers stay local-only (disabled with
  "Local only" metadata); a recognized VS Code command is enabled and
  forwarded with connectionId over a typed object IPC.
- Main process stays authoritative: rejects active/owned runtimes,
  resolves the SshTarget from the persisted Store, derives the authority
  (config alias, or username@host on port 22, or ssh-alias-required on a
  non-default port), validates POSIX/Windows absolute remote paths without
  local stat/normalize, and rejects non-VS Code and compound commands
  before spawn.
- Authority and remote path are passed as separate argv; getSpawnArgsForWindows
  remains the cmd/bat shim boundary and fails closed on metacharacters.
- Same capability rules across the worktree menu, Explorer overflow, and
  the source-control entry context menu.

Refs STA-2386
Closes #9999
2026-07-22 17:17:59 -07:00
NeilandOrca 95ae346b87 perf(ssh): send a single keepalive after wake, not two (#9880)
Co-authored-by: Orca <help@stably.ai>
2026-07-22 03:31:44 -07:00
OrcaWin b232df732b fix(terminal): make remote agent sessions host-authoritative (#9687) 2026-07-21 20:51:28 -07:00
Brennan Benson f1c84d3858 refactor(cli): split oversized command modules (#9775) 2026-07-21 13:50:17 -07:00
Brennan Benson a10a2ba53c feat(linear): add MCP-style save issue (#9670)
* feat(linear): add MCP-style save issue

* fix(linear): harden save issue parity

* fix(linear): close save issue contract gaps

* docs(linear): bundle project discovery with save issue
2026-07-21 13:25:22 -07:00
Brennan Benson 87af1c8673 feat(linear): add complete issue relations (#9674)
* feat(linear): add complete issue relations

* fix(linear): harden relation reads and writes

* fix(linear): classify ambiguous relation writes
2026-07-21 13:21:19 -07:00
Brennan Benson 42a4f017b4 feat(linear): add MCP-compatible issue listing (#9672) 2026-07-21 13:16:44 -07:00
Brennan Benson be066fe8e9 feat(linear): expose issue activity history (#9667) 2026-07-21 13:13:05 -07:00
OrcaWinandOrcaWin 88c78611b7 fix(ssh): patch node-pty helper in Windows relay (#9638)
Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-07-20 19:09:18 -07:00
4f81dfc128 perf(ssh): cut warm high-latency connects from 88.7s to 7.7s (#9015)
Move managed agent-hook filesystem work behind one relay RPC so high-latency SSH connects pay one WAN round trip instead of hundreds. Keep installers serial, lock shared account config across relay processes, and fence cancelled connection generations from replacement state.

Co-authored-by: nasagong <zinho2000@gachon.ac.kr>
Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-07-20 12:18:57 -07:00
NeilandOrca 190de8223e refactor(comments): slim verbose comments in main integrations (git/providers/…) (#9543)
Collapse multi-line explanatory comment blocks into single-line "why" statements
per AGENTS.md ("Document the Why, Briefly"): drop restatements of the code and
mechanism narration; keep the non-obvious reason, external refs, and directives.

Comments-only — verified no code changed via a Babel/esbuild comment-strip
token-equality gate against origin/main; typecheck and oxlint clean.

Area: main — git, source-control, providers & integrations. 40 files changed, 1432 insertions(+), 4473 deletions(-).

Co-authored-by: Orca <help@stably.ai>
2026-07-20 03:18:28 -07:00
Jinwoo Hong c0f0810dd9 Fix native Windows PTY startup query handling (#9500)
* Fix native Windows PTY startup query handling

* Fix daemon boot smoke protocol lookup

* Fix Windows daemon repro protocol lookup
2026-07-19 20:25:52 -07:00
Jinwoo Hong 3a847bfac9 fix(ssh): clear stamped agent status on disconnect (#9484)
* fix(ssh): clear stamped agent status on disconnect

Batch transient cleanup by accepted SSH connection authority and use a monotonic cutoff so reconnect replay wins over delayed clears. Preserve pane launch, resume, acknowledgement, and retention metadata.

Caveat: legacy or renderer-owned rows without an accepted connection stamp are intentionally left to existing pane/PTY teardown; clearing them by host would be ambiguous.

* docs(ssh): explain stale status watermark

* fix(ssh): preserve status ordering after restart
2026-07-19 22:14:15 -04:00
gatsby74andJinjing 8ed8f0d109 feat(ssh): download folders from remote explorer (#7793)
* feat(ssh): download folders from remote explorer

* fix(ssh): harden remote folder downloads

* Add missing getRepo stub to worktree cwd test mock

Restoring headless mobile tabs looks up the repo for the active
worktree id; the mock lacked getRepo, so the test only passed
incidentally. Add it explicitly and return undefined since wt-1
is a worktree id, not a registered repo.

* Enable SSH folder downloads, gated for system-SSH connections

- Folder downloads require SFTP, unavailable on system SSH (which
  offers only raw file operations). Add supportsFolderDownload flag
  to gate the feature in the UI layer.
- Reject symlinks at directory-entry level, preventing tree escapes
  and eliminating unnecessary stat calls.
- Check abort signal before opening dialog for better responsiveness
  when renderer closes.
- Log cleanup errors without re-throwing to preserve underlying
  transfer failures.

* Gate SSH folder downloads to SFTP-capable connections

Enforce fail-closed gating and add Windows path traversal validation to
ensure downloads are only available when explicitly supported and safe.

* Gate SSH folder downloads to SFTP-capable connections

Enforce fail-closed gating and add Windows path traversal validation to
ensure downloads are only available when explicitly supported and safe.

* fix(ssh): keep provider types under max-lines after main merge

Move FolderDownloadOptions next to the SFTP download implementation and
narrow IFilesystemProvider.downloadFolder options to AbortSignal only so
types.ts stays within the 300-line oxlint budget when merged with main.

---------

Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com>
2026-07-18 23:01:29 -07:00
Brennan Benson 3e276d78ba fix(ssh): probe npm via prepended PATH, not colocated with node (#9165) (#9255)
* fix(ssh): probe npm via prepended PATH, not colocated with node (#9165)

The remote Node/npm toolchain gate invoked npm by its absolute path
<nodeBinDir>/npm (POSIX) / npm.cmd (Windows, behind a Test-Path
colocation check). But deploy (commandWithNodePath) runs bare `npm`
with nodeBinDir merely prepended to PATH, so npm can resolve from
anywhere on PATH.

A host whose only resolvable node has npm elsewhere on PATH (e.g. node
symlinked into a dir without npm) deployed fine on v1.4.144, but after
upgrade the candidate is rejected with no fallback → SSH/relay
connection fails to establish.

Make the probe resolve npm exactly the way deploy does — bare
`npm --version` under the same prepended PATH — so it still confirms npm
is runnable (the #8450 concern) without requiring colocation. Windows
now prepends the backslash-form dir (matching deploy) so bare-command
PATH lookup resolves reliably.

* test(ssh): cover split Node npm PATH resolution
2026-07-17 18:51:10 -07:00
eisen0419andJinjing 3f335efdb9 feat(agents): pi session resume support (#8876)
* feat(agents): pi session resume support

* fix(pi): require persisted session files for resume

* test(sleeping-agent): use non-resumable sentinel in malformed-record fixture

The 'drops malformed sleeping agent resume records' test used agent:'pi' as
its example of an unknown/non-resumable agent, expecting the record to be
dropped. This PR added 'pi' to RESUMABLE_TUI_AGENTS, making that fixture
valid and retained, so the toBeUndefined assertion broke. Switch the
malformed-case fixture to a genuinely non-resumable sentinel
('definitely-not-an-agent') so the drop-malformed path is still exercised;
no other assertions changed.

* Add durable resume identity for Pi sessions without fabricating turn sta

Pi's `session_start` hook now carries the session file needed to resume
a sleeping pane, but until now Orca either discarded it or treated it
as a fake status transition. Thread a `providerSessionOnly` envelope
through the hook listener, relay, main-process server, and renderer
store so resume identity (and its session-file-scoped equality/claim
key) can be persisted and replayed without emitting prompt telemetry
or a visible working/done row.

* Add durable resume identity for completed Pi sessions

Pi's agent_end hook marks a turn done, but the underlying TUI session
stays alive and resumable. Previously a `done` status wiped sleeping
records and launch config as if the session ended, so hibernation,
manual worktree sleep, and quit-capture all lost Pi's resume identity.

- Track a "live recovery" record for done-but-still-resumable Pi
  sessions, exempting it from the usual done-state cleanup paths in
  agent-status.ts and agent-hibernation-planner.ts
- Gate providerSessionOnly rows and sleeping-agent schema records on
  actual resumability (getAgentResumeArgv) instead of trusting the
  presence of a provider session
- Wait for Pi to persist its session file before advertising resume
  metadata, and treat `/reload` as a non-terminal event so it doesn't
  clobber visible status
- Extend SSH relay envelopes to carry providerSessionOnly so remote
  hosts get the same behavior

* Add explicit periodic/quit mode to sleeping-agent session capture

Split captureAllSleepingAgentSessions into 'periodic' and 'quit' modes
so a background checkpoint can no longer downgrade a confirmed-quit
record or promote a completed Pi session without an authoritative
transcript path. Updates all call sites and tests accordingly.

* Use normalizeAgentStatusPayload for default pi status

Remove unnecessary JSON.stringify wrapper and call the appropriate normalization function directly.

---------

Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com>
2026-07-17 17:29:04 -07:00
Brennan Benson f544820552 fix(ssh): require a coherent colocated Node/npm toolchain for the relay (#9165) 2026-07-17 14:27:31 -07:00
fa85536f3a fix(ssh): repair unbuilt relay native deps (#8686)
Co-authored-by: Orca <help@stably.ai>
Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com>
Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
Co-authored-by: Jinwoo Hong <73622457+Jinwoo-H@users.noreply.github.com>
2026-07-16 21:32:13 -07:00
JinjingandOrca b0def3b130 fix(ssh): keep Windows CLI when the launcher compiler is missing (#8918)
The remote CLI installer deleted the legacy orca.cmd bridge before it
confirmed csc.exe existed, so a minimal Windows host without the .NET
compiler lost its existing CLI command and got nothing back.

Move the legacy-shim removal to run only after the compile and the
launcher-existence guard both pass. Missing-compiler and compile-failure
paths now exit non-zero while leaving the existing orca.cmd untouched; a
successful upgrade still clears the unsafe %* bridge (orca.exe already
shadows orca.cmd via PATHEXT). No %* forwarding is restored.

Co-authored-by: Orca <help@stably.ai>
2026-07-15 18:49:34 -07:00
e7f785f0ac fix(ssh): AND lsof selectors so relay reset can't kill unrelated processes (#8849)
* fix(ssh): AND lsof selectors so relay reset can't kill unrelated processes

`forceStopRelayForTarget` resolved PIDs with `lsof -t -U "$sock"`.
lsof ORs its selectors by default, so this selects every process holding
ANY unix socket in addition to the socket's owner. On hosts where lsof
cannot match AF_UNIX sockets by path, the path term matches nothing and
the sweep degrades to all unix-socket holders — including systemd --user
— which the TERM+KILL loop then takes down (#8762).

Add `-a` to AND the unix-socket and path selectors: where lsof supports
path matching the result is exactly the socket's holders (verified: 1
PID with -a vs 97 without on macOS); where it doesn't, the result is
empty and the existing pgrep fallback — scoped to command lines that
reference the relay's per-instance socket name — takes over.

Fixes #8762

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* review: add relay reset process-safety coverage

Exercise the generated POSIX reset script with controlled lsof and pgrep commands plus a mocked kill, and keep the selector rationale concise.

Co-authored-by: Orca <help@stably.ai>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com>
Co-authored-by: Orca <help@stably.ai>
2026-07-15 13:55:48 -07:00