* fix: address pr-bug-scan validated finding from #2447
Restore self-recovery from a stale socket file (left behind when a prior relay was killed by SIGKILL/OOM/host crash) without unlinking a live duplicate's socket.
* fix: address review findings
---------
Co-authored-by: orca-bug-scan-bot <orca-bug-scan-bot@stably.ai>
- Add a narrow SSH relay RPC for refreshing remote-tracking refs without
reopening generic fetch execution
- Resolve SSH connection context from composite worktree IDs during startup
before worktree discovery completes
- Make the sleeping workspace filter negative-form and reset to the new visible
default
- Tolerate transient xterm scroll restoration failures during layout
- Restore macOS Electron framework symlinks after copying the dev app
- Resolve an effective upstream so legacy branches tracking origin/main use
origin/<branch> when that remote branch exists.
- Pull, sync, and ahead/behind status now operate on the same branch the UI
reports, including after non-fast-forward push rejections.
* Create PRs directly from Source Control
- Replace the modal flow with an inline PR composer in the sidebar
- Keep PR creation state and validation scoped per worktree
- Rename the recovery action to clarify it only pushes before creating PRs
* Clean up fork PR remotes after worktree deletion
- Track Orca-created push target remotes in worktree metadata
- Reuse ownership markers when later worktrees share the same fork remote
- Fetch only the selected PR base instead of every remote before drafting PRs
- Mirror local branch cleanup for SSH worktree deletion
* Stabilize pull request creation flow
- Keep PR actions and composer fields locked while generation or creation is in flight
- Refresh git status, branch comparison, and history after remote actions settle
- Disable push-only actions on diverged branches so users sync first
* Make PR context generation read-only
- Stop rebasing or probing HEAD before collecting PR draft context
- Allow git operations on known repo roots without refreshing worktree cache
Co-authored-by: Orca <help@stably.ai>
* fix: address review findings
---------
Co-authored-by: Orca <help@stably.ai>
Fix Linux/bash agent status cleanup by emitting OSC 133 command lifecycle markers from bash shell wrappers, including local, daemon, and SSH relay paths.
- Treat new remote worktrees without a HEAD commit as an empty compare when
the base ref exists, avoiding a broken source-control compare state
- Keep the existing unborn-head error for cases where the base cannot resolve
* Squashed commits
- WIP: uncommitted changes before rebase
- ci
- Show inline PR check details in task drawer
- Add a Checks tab that opens from the PR checks cell and expands runs inline
- Fetch check output, annotations, and workflow job steps through IPC/RPC
- Improve markdown/comment wrapping so long PR content stays within the drawer
* Use app-styled confirmations for PR actions (#2324)
Co-authored-by: Orca <help@stably.ai>
* fix: pr-bug-scan validated finding from #2274 (#2296)
Co-authored-by: orca-bug-scan-bot <orca-bug-scan-bot@stably.ai>
* fix: avoid optional git locks during status checks
---------
Co-authored-by: Jinwoo Hong <73622457+Jinwoo-H@users.noreply.github.com>
Co-authored-by: Orca <help@stably.ai>
Co-authored-by: buf0-bot[bot] <252831055+buf0-bot[bot]@users.noreply.github.com>
Co-authored-by: orca-bug-scan-bot <orca-bug-scan-bot@stably.ai>
* Strip Grok user query wrapper from status prompts
* Document PR evidence image handling
* Surface Grok final responses in agent status
* Harden Grok status result extraction
* feat(file-explorer): show gitignored files with dimmed italic decoration
Surfaces `.gitignore`d files in the right-sidebar file explorer with an
italicised, dimmed filename and a CircleSlash icon in the same trailing
slot used by the git status letter. A tracked change always wins — the
ignored decoration only applies when no other git status is present.
Gated behind a new `showGitIgnoredFiles` global setting (default on) so
heavy SSH workspaces can keep the smaller payload by skipping
`--ignored=matching` on `git status`.
`ignoredPaths` lives as a peer field on GitStatusResult rather than an
extension of GitFileStatus/GitStagingArea, so Source Control's
staging-area grouping is untouched.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* chore(file-explorer): trim redundant comments from gitignored decoration
Removes duplicated "Why:" explanations that ended up restating the same
backward-compat rationale across five files (relay, ssh provider, runtime
git commands, RPC handler, renderer git client) plus a few comments that
narrated the mechanism the code already shows.
Net: -39 lines of comment across 10 files; no behavior change.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* chore(file-explorer): drop remaining comments from gitignored decoration
The code reads well without them.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* review: harden gitignored file decorations
- clear ignored decoration cache when ignored status is disabled or omitted
- keep ignored decoration state scoped across worktree and runtime cleanup
- add coverage for local, SSH, runtime, relay, and Explorer precedence
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com>
* fix: address pr-bug-scan validated finding from #1865
SIGKILL fallback in graceful pty.shutdown now emits pty.exit before notifyExitListener; renderer sees the exit even though onExit short-circuits on disposed.
* test: cover relay pty kill fallback exit
---------
Co-authored-by: orca-bug-scan-bot <orca-bug-scan-bot@stably.ai>
Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com>
* feat(ssh): stream fs.readFile to lift 10MB SSH preview cap (#1095)
Replaces the single-shot fs.readFile path on the SSH relay with a
push-style stream protocol modeled on VS Code's readFileStream.
Wire shape:
- fs.readFileStream request returns metadata (streamId, totalSize,
isBinary, mimeType, chunkEncoding, resultEncoding, optional empty)
- Relay pumps fs.streamChunk notifications (256 KB base64 chunks) and
ends with fs.streamEnd or fs.streamError
- Client cancels via fs.cancelStream notification
Invariants:
- Max 16 concurrent streams per FsHandler (TooManyStreams)
- Client clamps totalSize against caps before allocating
- Sequence-number defense against out-of-order/missing chunks
- Subscribe-before-await with frame queueing until streamId is known
- Pump cleans up registry+handle in finally; disposeAll aborts before
release so in-flight reads exit cleanly instead of EBADF
- Empty files short-circuit (no streamId, no handle open)
Compat:
- New client tries fs.readFileStream first, falls back to legacy
fs.readFile on JSON-RPC -32601 (with once-per-session warn log)
- Bumps MAX_PREVIEWABLE_BINARY_SIZE 10 MB to 50 MB to match local
Tests: 91 streaming tests across relay, client, mux, integration.
Co-authored-by: Orca <help@stably.ai>
* test(ssh): wait for streamEnd instead of fixed flush() in stream test
Why: the binary-streaming test relied on 5 setImmediate ticks to drain
the pump, which is racy on slower CI runners (each handle.read is async
I/O). Swap to a deadline-bounded waitFor(streamEnd) so the test is
deterministic regardless of scheduler latency.
Co-authored-by: Orca <help@stably.ai>
* fix(ssh): preserve small binary detection in streamed reads
Co-authored-by: Orca <help@stably.ai>
* fix(ssh): rebind file watcher when connection id hydrates
Co-authored-by: Orca <help@stably.ai>
* fix(ssh): refresh explorer for update-only file creates
Co-authored-by: Orca <help@stably.ai>
* fix(ssh): recompute file watches when repo connection changes
Co-authored-by: Orca <help@stably.ai>
* Revert "fix(ssh): refresh explorer for update-only file creates"
This reverts commit 7c3c683cd0.
* fix(ssh): install relay watcher dependency
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: Orca <help@stably.ai>
* fix(ssh): enable TCP_NODELAY on ssh2 client to eliminate per-keystroke typing lag (#1660)
ssh2 leaves Nagle's algorithm on by default. For single-byte keystrokes
through a remote PTY, Nagle interacts with the kernel's delayed-ACK timer
and adds up to ~40 ms per keystroke — visible as the typing lag reported
in #1660. OpenSSH's `ssh` client sets TCP_NODELAY whenever a PTY is
allocated; this change mirrors that on the ssh2 client right after the
`ready` event in doSsh2Connect, covering both initial connect and
auto-reconnect.
Proxy-command / proxy-jump connections (where ssh2's underlying socket
is a custom Duplex over a child-process pipe) are a no-op by design,
gated by the public Client.setNoDelay()'s own type guard. A discriminating
log line records which path each connect took.
Tests cover initial connect and a full reconnect cycle to guard against
the regression class "Nagle is re-enabled because someone refactored
only the initial connect path."
Co-authored-by: Orca <help@stably.ai>
* fix(ssh): bound relay-lost reconnect with exponential backoff
When the relay exec channel keeps dying (e.g. a remote-side bug closes
every fresh --connect channel right after handshake, or a stale bridge
keeps being replaced), the unguarded _onRelayLost handler reconnects as
fast as the network allows — spawning relay deploy attempts in a tight
loop until the user force-quits. Each iteration spawns a fresh ssh2 exec
channel, hammers sshd's MaxSessions counter, and floods the renderer
with state churn.
Add per-target exponential backoff (500ms → 15s, capped at 6 attempts)
so the loop terminates instead of running forever. After the cap the
session goes to 'error' state with a 'Relay channel kept dropping.
Please reconnect.' message — visible in the renderer instead of an
invisible failure where typing in remote terminals just stops working.
Successful 'ready' resets the attempt counter only if the session
stabilized for >= 5s; faster flaps preserve the counter so a flaky
remote backs off rather than retrying indefinitely on every brief
ready→lost cycle.
Backoff state is cleared on explicit disconnect, on session replacement
during reconnect, and on connect failures, so a real reconnect attempt
after backoff exhaustion always starts from zero.
Co-authored-by: Orca <help@stably.ai>
* fix(ssh): detect stale relay daemons via running-version marker
The on-disk relay version check compares local .version against the
remote .version file in the relay dir. A daemon launched by an earlier
deploy keeps running its in-memory copy of the OLD relay code, so when
the client later rewrites relay.js + .version on disk and bridges in
via --connect, the new bridge process drives a stale daemon. Protocol
or behavior changes between the two versions then tear down the
channel in a tight reconnect loop (observed against PR #1672 on a
daemon predating that change).
The daemon now writes its running version into a .running-version
sidecar at startup, anchored to the relay-script directory rather than
process.cwd() so test spawns cannot pollute the repo root. Before
attaching to an existing socket, the client probes that marker and,
on mismatch with the locally-deployed .version, kills the stale
daemon (TERM only, never KILL) and falls through to a fresh launch.
Conservative defaults: when either marker is unreadable, attach so
older builds keep their live PTYs.
Co-authored-by: Orca <help@stably.ai>
* Revert "fix(ssh): detect stale relay daemons via running-version marker"
This reverts commit e58acf07c0.
* fix(ssh): isolate relay versions via per-version install dirs and wire handshake
The relay's previous single-dir layout (~/.orca-remote/relay-v0.1.0/) let
the deploy step rewrite relay.js in place while a daemon was still loaded
in memory at the previous version. New clients then drove that stale
daemon, surfacing as a reconnect loop (issue #1660 follow-up) and the
field failure observed against an 8-day-old daemon on openclaw.
Switch to a VS Code-style versioned layout where each (RELAY_VERSION +
content-hash) bundle installs into its own directory and is never
mutated after install. A v2 client's --connect socket path is rooted in
relay-${v2-hash}/ and structurally cannot reach a v1 daemon's socket.
Defense-in-depth: the daemon now reads exactly one Handshake-typed frame
on each newly-accepted Unix socket before attaching the JSON-RPC
dispatcher (mirrors VS Code's remoteExtensionHostAgentServer.ts:340).
Mismatch closes the socket; the bridge exits with code 42; client maps
that to a typed RelayVersionMismatchError and skips the relay-lost
backoff loop instead of retrying through 6 attempts.
Other deploy hardening:
- atomic mkdir-based install lock with stale-lock recovery serialises
concurrent first-installs of the same version
- .install-complete sentinel distinguishes a finished install from a
crashed-mid-install partial that should be retried
- gcOldRelayVersions removes unreferenced sibling dirs (allowlist regex,
skips locked or incomplete dirs, skips dirs with a live socket)
- readLocalFullVersion fails fast on a missing/empty local .version
rather than silently falling back to a path where a daemon from a
different code generation may already be running
Includes a cross-version isolation test that fails any future refactor
which collapses the per-version layout back to a shared dir.
Co-authored-by: Orca <help@stably.ai>
* fix(ssh): harden relay versioning per review feedback
Address must-fix and should-fix findings from the parallel triple review of
26d1666e:
- Surface RelayVersionMismatchError to ssh.ts on initial establish() (not
just reconnect), so the user sees the typed terminal error instead of
silent retry on first connect (#13).
- Give the sentinel timeout a 500ms grace window for the close handler to
deliver exit-42, so a slow remote does not misclassify a wire-handshake
mismatch as a generic timeout (#D11).
- Drain the handshake decoder's residue at the handshake -> dispatcher
transition on both daemon and --connect sides; pipelined frames that
were coalesced with the handshake are now forwarded into the dispatcher
/ stdout instead of silently dropped (#A1, #A2).
- Reset the install-lock acquire timer after a stale-lock recovery so a
single post-recovery race does not immediately exhaust the budget (#E14).
- Treat a stale install-lock as recoverable in the GC pass when
.install-complete is present (covers an interrupted finalize where the
rm-lock failed) (#E15).
- GC legacy relay-v\d+\.\d+\.\d+ install dirs whose daemons have died,
now that .install-complete is no longer required for them (#12).
- Resolve symlinks in readLaunchVersion() so a daemon launched via a
symlinked entry script still reads .version next to the real file (#G21).
- Flush stderr before exit-42 in --connect handshake mismatch path so the
diagnostic line reaches the client before the process tears down (#C8).
Tests:
- Round-trip handshake over a real Socket pair: matching version, mismatch
exit-42, leftover bytes preserved on both sides when frames are
coalesced with the handshake.
- waitForSentinel exit-42 -> RelayVersionMismatchError, exit-1 -> generic.
- SshRelaySession terminal-error callback fires on both establish() and
reconnect() when deployAndLaunchRelay throws RelayVersionMismatchError.
- acquireInstallLock concurrent BUSY -> OK polling, stale-lock recovery
with reset timeout window, and fresh-lock timeout failure path.
- gcOldRelayVersions stale-lock-with-complete branch, legacy-dead path,
legacy-alive path; existing locked-test asserts fresh-lock now keeps.
- Cross-version isolation test now asserts a blanket invariant that every
v1-referencing command from a v2 deploy is a read-only liveness probe.
Lint and typecheck clean across all 3 tsconfigs; 426 SSH/relay tests pass.
Co-authored-by: Orca <help@stably.ai>
* fix(ssh): bypass npm init for content-hashed relay dirs and harden install probe
The versioned-install dirs land at `relay-${version}+${hash}/` (e.g.
`relay-0.1.0+07994a7870e1`). npm 11 / Node 26 reject the `+` in derived
package names and `npm init -y` exits 1 — silently, since both stderr
and the failure landed inside the `2>/dev/null && ...` chain. The catch
swallowed the throw, `.install-complete` was written anyway, and every
reconnect surfaced 'node-pty is not available' at first pty.spawn.
Sidestep `npm init` entirely: SFTP-write a hardcoded minimal
package.json (`name: orca-relay`, `type: commonjs`) and run
`npm install node-pty` directly. `type: commonjs` pins the module
system against future Node default flips or remote-side .npmrc overrides.
Also harden the install path against the same class of silent failure:
- npm install errors now propagate (no more `.install-complete` on hard
fail; future reconnects retry instead of stranding the user)
- Replace the weak `test -d node-pty` post-install probe with
`node -e 'require("node-pty")'` so built-but-unloadable installs
(missing prebuild, wrong arch, broken native binding) surface clearly
- Add a session-level error handler on the SFTP write so a torn-down
session rejects the promise instead of hanging until enclosing timeout
Separate fix: add `for-each-ref` to the relay's git subcommand allowlist.
Client code (`src/main/git/repo.ts` ref-search and worktree-listing)
calls `git for-each-ref` over SSH; the relay was rejecting it. The
`--shell`/`--python`/`--perl`/`--tcl` format flags only control output
quoting (no eval) and the relay invokes git via execFileAsync (no shell),
so the read-only allowlist treatment matches `rev-parse`, `log`, etc.
Co-authored-by: Orca <help@stably.ai>
* comment(ssh-relay): TODO link to #1693 for VS Code-style pre-bundled node-pty
Co-authored-by: Orca <help@stably.ai>
* fix(ssh): harden node-pty install probe and tighten review-fix tests
Round-3 review fixes on top of 963f56d7.
deploy.ts:
- Replace endsWith('OK') with includes('ORCA-NPTY-PROBE-OK'). Node can emit
deprecation/experimental warnings to stderr after our stdout 'OK' write,
and 2>&1 would push them past 'OK' producing false NPTY-MISSING warnings.
A unique sentinel survives any trailing stderr noise.
- Switch sftpPkg/ws .on -> .once for error/close. A late session 'error'
after the promise had already settled would otherwise become an unhandled
EventEmitter error and crash main.
- Trim per-block comments to 1-2 lines per AGENTS.md (was 7-9).
Tests:
- Pin the BEFORE-ordering contract: SftpWriteCapture now records the count
of execCommand calls observed at the moment ws.end() ran for each path,
and the test asserts that count <= the index of npm install. Catches a
future Promise.all-style refactor that would still pass final-state checks.
- Strengthen the SSH-channel-failure test: assert the rejection actually
came from the probe call (not an earlier exec) by finding the probe
invocation in mock.calls. Also assert NPTY-INSTALL-FAIL is NOT logged
(channel failure must not be conflated with install failure) and that
abandonInstall was called so the lock is released.
- Fix misleading clearAllMocks comment: it claims to wipe mockReturnValue,
but actually clearAllMocks only resets .mock.calls. Re-priming was
defense-in-depth, not a correctness requirement.
Validator:
- Add for-each-ref negative cases (--git-dir, --output, --work-tree) to the
global-denied-flags it.each. The first round of for-each-ref enablement
trusted that the post-subcommand GLOBAL_DENIED_FLAGS check applied; this
pins it so a future allowlist refactor that bypasses the global check
fails loudly.
Co-authored-by: Orca <help@stably.ai>
* fix(ssh): split node-pty probe into test-d guard + load-test
Round-4 review fixes for the install probe in installNativeDeps:
(1) test -d guard runs before the load-test. If the install dir vanished
between npm install and probe (concurrent rm, fs unmount, permission
flip), the deploy now throws and the next reconnect retries fresh —
previously the cd failure flowed into '|| echo MISSING' and we'd
write .install-complete, stranding the user in degraded mode.
(2) Load-test discards stderr (2>/dev/null) so customized .bashrc
output (NVM init, conda greetings, etc.) can't pollute the sentinel
match. The shell-level '|| echo MISSING' is preserved so SSH-channel
rejections still propagate as exec errors, distinct from require
failures which exit the node process nonzero.
(3) PROBE_OK is passed via process.argv[1] so the JS literal stays
trivial regardless of future sentinel characters.
Test changes:
- New 'dir-gone' probe mode in makeExecResponses
- New test pinning that vanished-dir throws (not silent MISSING)
- SSH-channel test now asserts probeCallIdx > npmInstallIdx
- cross-version-isolation feeds an extra '' for the test -d slot
Co-authored-by: Orca <help@stably.ai>
* fix(ssh): simplify node-pty probe and harden test ordering pins
Round-5 review fixes for installNativeDeps:
Production:
- Drop redundant test -d guard. `cd ${dir} && (...)` short-circuits on
cd-failure (dir-vanished) and propagates as exec reject already; the
separate guard added a round trip without preventing anything.
- Capture probe stderr to a per-deploy file rather than 2>&1 or 2>/dev/null.
.bashrc noise can't pollute the sentinel match, but the require() error
message is preserved in the [NPTY-MISSING] log breadcrumb so bug reports
point at the real cause (e.g. GLIBC version mismatch).
- Mirror the install command's PATH (export PATH=${binDir}:$PATH) so any
future require-time child_process call resolves the same node binary
used during install.
- Add platform tuple to [NPTY-MISSING] and [NPTY-INSTALL-FAIL] logs for
triageable bug reports without asking users to dig out their arch.
- Trim probe comment per AGENTS.md (why-only, no mechanism narration).
Tests:
- Pin full installNativeDeps ordering: npm install < chmod prebuilds <
probe. Catches refactors that probe before install or move chmod after.
- Pressure-test .includes(PROBE_OK) survives bashrc/MOTD noise prefixed
to probe stdout (corporate banner / NVM init / conda greeting case).
- Pressure-test MISSING detection survives Node deprecation warnings
prepended to the MISSING token.
- Pin platform tuple appears in [NPTY-MISSING] log.
- Pin finalizeInstall called exactly once + abandonInstall not called
on happy paths; reverse on failure paths.
- Strengthen dir-gone test: assert probeIdx > npmInstallIdx so a refactor
that swaps order doesn't silently let the test pass on its own injected
error string.
- New probeStdoutOverride option in makeExecResponses for shell-noise
injection tests.
Cross-version-isolation: dropped obsolete test -d slot, added rm-stderr
cleanup slot to match the new probe shape.
eslint-disable max-lines on both files with rationale (pattern used widely
in this repo for cohesive single-responsibility modules).
441/441 tests pass; lint clean; typecheck clean. Probe shape verified
end-to-end on real remote.
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: Orca <help@stably.ai>
* WIP: Changes before auto-review fixes
Co-authored-by: Orca <help@stably.ai>
* WIP: Changes before auto-review fixes
Co-authored-by: Orca <help@stably.ai>
* fix(worktree): preserve user push.autoSetupRemote, include path in warn
- Probe push.autoSetupRemote with `git config --get` before writing so a
deliberate user value at any scope (local/global/system) is preserved.
- Include worktree path in the warn log for failed config writes.
- Add test pinning the preserve-existing-value behavior.
- Remove stray 00-review-context.md committed during review tooling.
Co-authored-by: Orca <help@stably.ai>
* WIP: Changes before auto-review fixes
Co-authored-by: Orca <help@stably.ai>
* fix(worktree): narrow config --get error handling, tighten test asserts
Treat only exit code 1 from `git config --get push.autoSetupRemote`
as "key unset". Other read failures (corrupt config, locked file,
parse error) now re-throw to the outer warn handler instead of being
silently treated as unset and overwriting whatever value the user
actually has.
Also: add test for the non-unset read-error path; convert the
"preserves existing value" test from `.some()` predicates to a
full-array `toEqual` matching sibling-test style; explicitly mock
`config --get` (with code: 1) in the sparse-failure rollback test
so it exercises the intended branch instead of the helper's empty-
stdout fallthrough; document in the design notes that
addSparseWorktree's rollback intentionally does not unset
push.autoSetupRemote.
Co-authored-by: Orca <help@stably.ai>
* test(worktree): pin --get-empty-stdout and worktree-add-fail invariants
Why: addWorktree's post-create config probe has two ordering
invariants worth pinning so a future refactor can't silently
regress them: (1) `git config --get` succeeding with empty stdout
still counts as "already set" so we don't overwrite an explicit
empty value, and (2) the entire config block is skipped when
`worktree add` itself rejects.
Co-authored-by: Orca <help@stably.ai>
* WIP: Changes before auto-review fixes
Co-authored-by: Orca <help@stably.ai>
* docs(worktree): cross-ref local↔SSH addWorktree, clarify SSH-host git version, add empty-stdout parity test
JSDoc on local addWorktree now flags the push.autoSetupRemote side
effect; both paths cross-reference each other so the next change keeps
them in lockstep. Relay comment clarifies that the git version that
matters is the SSH host's, not the client's. Adds the missing
empty-stdout-as-already-set parity test on the relay side.
Co-authored-by: Orca <help@stably.ai>
* chore: remove 00-review-context.md from PR
Stray file from local review workflow; should not ship in this PR.
Co-authored-by: Orca <help@stably.ai>
* chore: remove worktree-ssh-no-track-parity.md from PR
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: Orca <help@stably.ai>
When the editor is disposed during a parent render, the dispose
listener's setState re-runs this effect and triggers a synchronous
root.unmount() inside React's commit work loop, producing React 19's
"Attempted to synchronously unmount a root while React was already
rendering" warning. Snapshot the roots and clear bookkeeping
synchronously, then unmount via queueMicrotask — matches the
deferred-unmount pattern already used in the diff-pass effect.
Co-authored-by: Orca <help@stably.ai>
* feat(agent-hooks): introduce relay wire envelope + connectionId stamping
Adds the shared `agent-hook-relay.ts` module with the `agent.hook` JSON-RPC
notification envelope, the `agent_hook.requestReplay` /
`agent_hook.installPlugins` method names, and the
`ORCA_FEATURE_REMOTE_AGENT_HOOKS` flag helper. Promotes `AgentHookSource` to
`shared/` so the relay can import it without dragging Electron in.
Threads a `connectionId: string | null` field through `AgentHookEventPayload`,
the `agentStatus:set` IPC contract, and the renderer-bound preload listener.
Local hook posts stamp `null`; the relay-forwarded path will stamp from `mux`
identity in a later commit. Renderer uses the stamp for stale-event filtering
when an SSH connection tears down with notifications still in flight.
See docs/design/agent-status-over-ssh.md §1, §5, §8 (commit #1).
Co-authored-by: Orca <help@stably.ai>
* refactor(agent-hooks): extract shared listener; add relay-side adapter
Extracts the listener internals (request parsing, payload normalization,
endpoint-file writing, per-CLI extractors, warn-once Sets, slowloris timer
helper, request size cap, paneKey caches) from `src/main/agent-hooks/server.ts`
into a new transport-agnostic `src/shared/agent-hook-listener.ts`. The shared
module uses only Node builtins (no Electron) so it is safe to import from
`src/relay/`.
Adds `src/relay/agent-hook-server.ts` — a thin HTTP-loopback adapter that
wires the shared listener to a `forward(envelope)` callback so `relay.ts` can
re-emit each parsed payload as an `agent.hook` JSON-RPC notification on the
existing SshChannelMultiplexer. The adapter owns:
- 127.0.0.1:0 socket + bearer-token auth, identical shape to the local server
- per-paneKey last-payload cache + replayCachedPayloadsForPanes() for the
request-driven replay path used after `--connect` reattach (see §5 Path 3)
- clearPaneState(paneKey) for PTY-exit eviction (symmetric with local server)
- buildPtyEnv() / endpoint-file writing for relay-spawned PTYs
Orca's `AgentHookServer` is now a ~200-LoC adapter over the shared listener
that owns the IPC fanout, listener replay, and `ingestRemote(envelope, connId)`
entry point that bypasses the HTTP path for relay-forwarded events.
See docs/design/agent-status-over-ssh.md §3, §8 (commit #2).
Co-authored-by: Orca <help@stably.ai>
* fix(preload): expose connectionId on agentStatus.onSet type
src/preload/index.ts already passes through `connectionId?: string | null`
from main, but the PreloadApi declaration in api-types.ts was missing the
field. Align the type with the runtime contract so renderer call sites
can read connectionId without an `as` cast.
Co-authored-by: Orca <help@stably.ai>
* fix(agent-hooks): harden ingestRemote + relay replay; review-driven cleanup
- ingestRemote: re-run normalizeAgentStatusPayload at trust boundary;
trim+validate connectionId/paneKey/tabId/worktreeId
- relay: preserve source/env/version through replay via sidecar map;
drop sourceFromAgentType fallback that mis-tagged unknown agents
- shared listener: exhaustive switch+never on AgentHookSource dispatch
chains; extractPromptText returns trimmed values; export MAX_PANE_KEY_LEN
- preload: tighten connectionId from optional to required (always sent)
- main IPC: reorder spread so explicit envelope fields win on collision
Co-authored-by: Orca <help@stably.ai>
* chore(docs): drop agent-status-over-ssh design doc from PR
The design RFC was useful for authoring this PR series but doesn't belong
in-tree — keeping it here would freeze line-number references and design
prose against future churn. Folding it into the PR description instead.
Co-authored-by: Orca <help@stably.ai>
* chore(agent-hooks): widen ingestRemote type for env/version (PR2 prep)
Declares `env?: string` and `version?: string` on the `ingestRemote` envelope
parameter so PR2 only needs to add the `warnOnHookEnvOrVersionMismatch`
callsite, not also widen the type. The fields are forwarded verbatim from
the agent CLI POST body on the remote and let Orca's warn-once cross-build
/ dev-vs-prod diagnostics fire identically on remote-sourced events.
Type-only addition; no runtime consumer in this PR.
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: Orca <help@stably.ai>
When a remote SSH workspace contains a symlink whose target lies outside
the registered repo/worktree roots, file reads failed with 'Path outside
authorized workspace'. This silently broke common workflows: HPC dataset
mounts, multi-checkout repos, dotfile editing, and any cross-mount
symlink.
Drop `RelayContext.authorizedRoots`, `validatePath`, and
`validatePathResolved` along with all ~33 call sites in fs-handler.ts
and git-handler.ts. The relay's threat model becomes 'the relay runs as
the SSH user and trusts the renderer.'
Why this is acceptable: `pty.spawn` and `git.exec` already concede the
same threat. A renderer that wants to reach `/etc/passwd` can spawn a
shell or run `git -C /etc cat-file`; the FS allowlist was friction, not
a security boundary. Intra-worktree path checks in `getDiff` and
`discard` are intentionally preserved.
Back-compat preserved: `session.registerRoot` (notification + request)
remains a valid RPC, retained as no-ops on new relays. Old main + new
relay and new main + old relay both keep working through the upgrade
window. `registerRelayRoots` is also kept for the same reason. A
narrowed error-translation block in `worktree-remote.ts` handles old
relays still surfacing the legacy error string to users.
Tests: removed two negative-allowlist tests; added a positive control
('reads files outside any registered root') and a direct regression
test for #1661 ('reads files via symlinks resolving outside the
workspace'). All 469 relay/SSH/IPC tests pass.
See docs/relay-fs-allowlist-removal.md for the full rationale,
back-compat matrix, alternatives considered, and follow-up cleanup
plan.
Closes#1661
Co-authored-by: Orca <help@stably.ai>