* perf(runtime): gate terminal.list visual layouts and stop the false writable claim
visualLayouts is ~31% of a large terminal.list payload (44,208 B of 137,412 B on a live 134-terminal remote runtime) and has exactly one consumer: the human-readable CLI formatter. Gate it behind an includeVisualLayouts request param that defaults to included, so pre-flag clients are unaffected, and have every --json/internal caller opt out.
Also drop the record-backed builder's writable, which was a verbatim copy of connected. terminal.show now states writability explicitly as exactly what terminal.send's PTY gate enforces.
* test(runtime): type the payload-size fixture arrays for tsc
* fix(runtime): preserve terminal list compatibility
* test(runtime): guard terminal list optimization
* fix(cli): preserve agent access to terminal layouts
* fix(ssh): handle owner displacement and graceful shutdown
SSH connections can reconnect with valid session proof after network loss or
device sleep. When the incumbent owner is still half-open, allow the
reconnecting client to displace it outright rather than wait for socket closure
— a window that may never close. Retain displaced deliveries for the new owner
to rotate. During app shutdown, drain SSH sessions without terminating recovery
operations, and retry pending owner grants in case a replacement commits mid-drain.
* fix(ssh): handle owner displacement and graceful shutdown
Make QuitTeardownStartGate a shared singleton so SSH connects use the
same shutdown fence as the main quit path. Track test-connection probes
to ensure they complete before final teardown. Guard owner displacement
to prevent stale owners from clearing recovery state claimed by newer
owners.
* fix(ssh): fix flaky test sync and add error code safety check
Test was using tick-based Promise.resolve() loops which don't guarantee
the async operation has started. Replace with signal-based synchronization
that waits for the actual lease flush. Also add nullish-coalescing to
error code check to prevent crashes if error is null or undefined.
* fix(ssh): fence reset transport opens during shutdown
* fix(ssh): keep recovery leases stable across reconnects
* fix(ssh): close transports owned by cancelled connect attempts
When a connect is cancelled after its transport has opened, that cancelled
attempt still owns the transport and must close it — otherwise it leaks. Add
disconnectConnection() to close by identity (not by target ID) so a cancelled
attempt closes only the transport it minted, without tearing down its
replacement's live transport. Track priorConnection to detect whether this
attempt opened a new transport or reused an existing one, and close only on
abandonment if this attempt owns the session.
* fix(ssh): fence old owner proofs and close superseded transports
When an owner reconnects with a new proof while an old one is still live,
the old proof is now fenced with SUPERSEDED_ERROR instead of retrying
indefinitely. The relay also closes stale transports to signal that their
recovery generation has been overtaken by a newer one.
This ensures overlapping reconnect scenarios complete with the newest proof
rather than getting blocked by stale recovery attempts.
* fix(relay): re-pin stdin/stdout fds after closing to prevent recycling
When the relay closes stdin/stdout to signal EOF to the SSH peer, the OS
can recycle those fds (0 and 1) for new sockets or files. If Node still
treats process.stdin/stdout as those numbers, subsequent operations
corrupt socket clients and trigger shutdown errors. Re-pin the fds by
opening /dev/null to keep them occupied and prevent recycling.
* feat(ssh): add SSH config host picker for add-host form
Users can now click 'Fill from ~/.ssh/config…' to browse available SSH
config hosts in a picker, select one, and have the form automatically
prefill with resolved connection details (hostname, port, username, auth).
Previously, an 'import' button provided bulk sync on this form—confusing
and unhelpful when everything was already synced. That action is now
available as a secondary 'Add all' option in the picker.
* fix(ssh): import filter preservation and label fallback
- Reuse search loader on import completion to preserve active filter inside generation guard
- Fall back to hostname when manual host has no label, not empty string
- Make alias duplicate detection case-insensitive to match config picker behavior
- Validate host availability when restoring project group selection
- Add aria-selected attribute to picker options for accessibility
* fix(ssh): harden config picker import, alias folding, and host targeting
Review findings on the ~/.ssh/config picker + bulk add:
- Guard config-host resolution with a generation counter so a late resolve
cannot overwrite a later pick or a form the user backed out of; freeze the
other rows while a pick resolves.
- Stop "Add all N" from re-adopting deleted hosts — it now imports without
reAdopt, matching the new-host count it advertises. Settings → Import keeps
the explicit re-adopt path.
- Fold SSH aliases through a shared normalizeSshConfigAlias for import
ownership, delete tombstones, reclaim, picker search, and the save-time
duplicate check, which now occupies configHost *and* label like the picker.
- Persist GSSAPIAuthentication only when a parsed Host entry asks for it, not
when `ssh -G` merely echoes the /etc/ssh system default.
- Fail closed with unavailable/setup-not-found when an explicit
projectHostSetupId names a non-actionable host instead of silently creating
the workspace on a sibling host.
- Cache the parsed config for the picker session (refresh on open/retry) so
filter keystrokes no longer reparse and Include-expand the file, keep the
filter usable during loads, add a Retry on load errors, explain an empty
Identity file after a config fill, and drop the always-false aria-selected.
* refactor(ssh): centralize host result limit and extract folder group val
Move SSH_CONFIG_HOST_RESULT_LIMIT to shared types so the renderer's limit message
cannot drift from the host's query limit. Extract findActionableFolderProjectGroup
to avoid repeating the folder-host-availability check across the composer hook.
* fix(ssh): pass -F to ssh -G when HOME differs from passwd home
In E2E tests and sandboxes, isolated HOME can differ from the system
passwd home. OpenSSH resolves the default config via getpwuid (passwd),
while Node's loadUserSshConfig uses os.homedir() (HOME-aware). Pass -F
to explicitly specify the config path when they diverge, so ssh -G and
the picker resolve the same file.
* fix(ssh): verify config host exists before resolving with ssh -G
When a user edits ~/.ssh/config and removes a host, the import picker
should not fall back to ssh -G's echoed response (which treats any alias
as valid). Check the reloaded config file before resolving.
- Force reload config on each resolve to catch user edits post-open
- Reject aliases not in the current config before calling ssh -G
- Add test for deleted alias edge case
- Fix workspace-target fallback to honor explicit host selection
* fix(ssh): let tombstoned aliases be re-picked in the config picker
Allow users to reclaim a deleted SSH host by re-picking it from ~/.ssh/config. Tombstoned aliases now appear in the picker with a "Removed from Orca" badge and remain pickable, but don't count toward "Add all" operations — ensuring passive import never resurrects a deleted alias while still giving the user a recovery path.
* fix(ai-vault): support session scanning in SSH worktrees
Add relay-native aiVault.listSessions scanning that discovers agent
sessions on SSH hosts. Includes fallback to filesystem crawl for
legacy relays, full cancellation support, result validation, and
scan coalescing to reduce redundant work.
* fix(ai-vault): scan sessions in SSH worktrees with coordinated cancellat
- Extract batching logic to `mapRemoteScanBatches` for reuse and proper cancellation checkpoints
- Move `AiVaultScanCoordinator` from relay to main to handle concurrent same-key requests with individual cancellation signals
- Report scope path truncation consistently across relay and SSH fallback paths
- Gracefully degrade relay handler on unsupported platforms instead of aborting startup
- Refactor issue display to separate blocking errors, scope notices, and skipped transcript counts
* fix(ai-vault): stabilize SSH session scan CI
Swallow async WSL relay stdin EPIPE so the live hook-relay shard no longer
fails after all tests pass. Merge main, resolve scan/relay conflicts, and
align cancellation/host-issue reporting with IPC expectations.
* fix(ai-vault): harden session scan cancellation, relay timeouts, and preemption
Thread the abort signal through every scan and parse path so superseded or
cancelled scans stop promptly instead of parsing every remaining transcript
for a caller that already left. Replace the fragile message-text relay
timeout check with a typed error code so unrelated errors carrying the
phrase "timed out after" no longer suppress the filesystem fallback. Fix
scan coordinator preemption so a forced Refresh in one window no longer
re-enters as a spurious cancellation in another. Add a host-leg cache for
the all-hosts view and cap filesystem concurrency so a single slow remote
home cannot stall the whole merge.
Co-authored-by: Orca <help@stably.ai>
* fix(ai-vault): use stable React keys for scan issue banners
Drop array-index keys so react-doctor/no-array-index-as-key passes.
Uniqueness comes from host, kind, agent, path, and message.
* fix(ai-vault): SSH session scanning with configurable depth limits
Implement depth-aware caching and proper scan boundaries to make SSH session
scanning reliable in worktrees. Users can now select between faster (250
sessions) and comprehensive (unlimited) history scans. The scanner:
- Deduplicates scans across relay, host leg, runtime, and renderer layers
- Reuses larger scans to serve smaller depth requests
- Properly bounds in-scope discovery per-limit
- Fixes timeout enforcement when SSH providers ignore abort signals
* Move sessionLimit ref update to useLayoutEffect
Keep render pure for React Doctor by deferring ref updates to
a layout effect, which still executes before render-dependent
effects that consume the ref.
* fix(adhoc): stamp version prefix from main, not the feature branch
Adhoc builds check out arbitrary refs whose package.json often lags
version bumps (e.g. 1.4.165-rc.0 while main is 1.4.168-rc.1). Hourly
always builds main so it already tracks the product line; adhoc now
resolves the base version from origin/main (or ORCA_ADHOC_BASE_VERSION)
so branch builds share that prefix.
* Revert "fix(adhoc): stamp version prefix from main, not the feature branch"
This reverts commit a26a18eb3fd83f7e7d2db9a6a7c3e02e0f79089a.
* fix(ai-vault): fix scoped backfill and coordinator race conditions
Resolve race where the last waiter leaving could abort an already-settled scan (add `settled` flag). Redesign scoped session backfill to keep searching through newer files until the scope reaches its requested session quota instead of stopping at the candidate limit; out-of-scope files no longer consume the scope budget. Centralize scan limit normalization and fix error classification for cancelled scans using the proper helper instead of checking Error.name. Disambiguate cache keys using JSON and add cancellation check after scope discovery phase.
---------
Co-authored-by: Orca <help@stably.ai>
* Add first user prompt to AI Vault session history rows
Re-parse transcripts on demand to extract and display the untruncated first
user prompt for copy/reuse. List scans omit the body (payload/perf); UI loads
it when session details expand. Grok sessions extract the typed ask from
<user_query> envelope, skipping injected <user_info> bootstrap rows. Supports
Claude, Codex, Grok, and OpenCode agents.
* fix(ai-vault): split SessionTime out to pass max-lines lint
AiVaultSessionDetails exceeded the 400-line oxlint limit after adding
first-prompt UI; move SessionTime into its own module.
* fix(ai-vault): handle corrupt transcripts and fix OpenCode prompt captur
Corrupt transcripts now resolve null instead of rejecting the IPC call, matching behavior for other unavailable cases. OpenCode SQLite parsing now correctly captures all text parts from the earliest user message only, fixing truncation of large prompts and padding of small ones. Add stale-response guard in the UI to prevent late results from overwriting the current session when tabs switch. Consolidate text slicing via `sliceAtCodeUnitLimit` to avoid surrogate-pair splits across all callers.
* test(ai-vault): add first-user-prompt UTF-16 safety tests
Ensure truncation at safety limits doesn't split UTF-16 surrogate pairs,
preventing corruption of astral characters in captured prompts.
* fix(ai-vault): key first-prompt-card by session.id
Remounting the card on session switches prevents late responses from
a previous load from writing stale data into the component's refs.
Also improves conversation-turn key stability.
* fix(ai-vault): preserve first prompt after preview truncation
* refactor(ai-vault): improve first user prompt capture robustness and per
- Add 15s timeout to full-prompt load to prevent indefinite loading states
- Extract seedFullFirstUserPrompt helper for reuse across parsers
- Prevent AI-generated summaries from becoming the copyable first prompt
- Fix truncation detection in OpenCode SQLite by probing for N+1 rows
- Optimize text bounding to apply safety limit before toLowerCase
- Gate synthetic OpenCode path detection on agent type, not just # presence
- Add test coverage for remote execution host handling
* Fix FirstPromptCard loading state stranded by stale promise reuse
Clears loadPromiseRef during cleanup to prevent the dedupe handle from
causing StrictMode remounts to await stale in-flight requests. Stops loading
when session becomes non-loadable mid-request. Adds tests for StrictMode
double-invoke resolution and main-process timeout scenarios.
* refactor(ai-vault): split session parsers into modular files
Split secondary-parsers into individual files per agent type (copilot,
cursor, hermes, opencode) for improved modularity. Add test coverage
for first-user-prompt envelope handling: unwrap user_query tags and
reject bare user_info dumps.
* fix(ci): clear max-lines and flaky portal readiness check
Collapse an accidental multi-line regex wrap in ssh-connection-utils that
pushed counted lines to 301. Harden the latched-readiness test's ready
transition so CI load can re-observe attach after MutationObserver gaps.
* fix(ssh): extract proxy command helpers to pass max-lines
Move resolveEffectiveProxy/spawnProxyCommand out of ssh-connection-utils
so oxfmt line wrapping cannot push that file over the 300-line lint cap.
* capture first user prompt by ordering OpenCode messages by creation time
- Add `readOpenCodeMessagesInOrder` to rebuild transcript by timestamp, handling
corrupt/partial files gracefully instead of discarding sessions
- Extract SSH proxy command tests to dedicated file; add backpressure handling
and stderr draining to prevent proxy process stalls
- On Windows, reject unsafe characters in ProxyCommand values instead of
pretending to escape them; properly format cmd.exe invocation with verbatim
arguments
- Expand ProxyJump chains into -J plus final hop, mirroring OpenSSH behavior
- Decouple portal readiness reapply budget from flip-count budget via explicit
constant
src/relay/relay-frame-decoder.ts and src/main/ssh/relay-frame-decoder.ts
were 264 identical lines apart from one default: the relay logs decode
faults to stderr when no handler is supplied, the SSH side stays silent.
Two copies of framing logic is exactly where a wire-format fix lands in one
and not the other.
The decoder's contract and buffer already live in src/shared, so the class
joins them there. The relay keeps a thin subclass that supplies its stderr
default, preserving behaviour for the call sites that omit onError. The
SSH copy is deleted and relay-protocol.ts points at shared directly.
Verified: pnpm typecheck, 102 tests across the 9 framing/backpressure/
handshake suites, and `pnpm build:relay` for all six platform targets plus
the WSL hook relay — the standalone bundle has no new dependencies.
* chore(dead-code): drop 2k lines of unreachable exports and orphan modules
Ran knip across every build entry (main, preload, renderer, popout, web,
cli, relay, workers, forked sidecars, config scripts) and removed what no
entry graph can reach.
- 11 orphan modules nothing imported, plus one test that only covered them
- 159 unused exports/types, with their now-dead helpers, imports and tests
Each candidate was verified against dynamic references before deletion.
42 knip hits were false positives and are kept: shared modules consumed by
the mobile/ workspace, the src/shared/plugins/** public API, vendored
shadcn primitives, and relay wire-protocol constants held for compatibility.
Adds knip.json + `pnpm audit:dead-code` so this stays measurable.
Verified: pnpm typecheck, pnpm lint, and 2081 tests across the 73 affected
test files all pass.
* chore(dead-code): move knip config under config/
Root-level additions are blocked by the root directory guard.
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: Orca <help@stably.ai>
* fix(P1-A): persist SSH consumer recovery without a sync store flush
rememberPtyConsumerRecovery ran on the live establish/reconnect path and
called flushOrThrow -> writeToDiskSync, parking the Electron main thread on
the profile-directory write. On a stalled or slow profile mount that freezes
the whole app during SSH recovery and reconnect.
Add Store.flushAsync(): same debounce-cancel and write serialization as
flushOrThrow, but awaits writeToDiskAsync instead of blocking. The consumer
recovery upsert/remove pair is now async and awaits it, and the SSH callers
await through to establish()/reconnect() so ownership is still durable before
relay setup continues. In-memory state still mutates synchronously (before the
first await), so no caller can observe a torn record and dispose() stays
synchronous.
* fix(P1-A): detach the SSH session when a connect attempt fails
Both failure exits in doConnect dropped the session from activeSessions
without calling detach(). claimSshPtyConsumerRecovery only reuses an existing
in-memory entry when detached === true, so the next connect attempt fell
through to minting a fresh clientInstanceId, discarding the remembered owner
lease and its resume identity.
Route both exits through abandonFailedSshSession(), which detaches (keeping
PTY ownership, unlike dispose()) before removing the session, and tolerates a
teardown throw so it can't mask the connect error being rethrown.
* fix(P1-A): await async lease persistence in SSH relay teardown
Failed connect attempts now wait for 'detached' leases to persist before
throwing, preventing reconnects from claiming them before cleanup completes.
Detach and dispose operations are now async and await store durability.
* fix(ssh): make session detach lease writes retryable on failure
Separate in-memory detach (identity recovery, provider cleanup) from lease
write persistence so rejected writes can be re-issued without re-running
provider teardown or re-minting the session identity. Introduce
flushDurableStateOrThrowAsync to flush only SSH-recovery state on the
live establish/reconnect path, avoiding snapshot writes of sidecars that
belong to quit/startup. Use Promise.allSettled in test reset to prevent
one rejected disposal from leaking state into the next test.
* fix(ssh): dispose mux on failed establish and propagate sync errors
- Dispose mux when session is disposed during establish to prevent resource leak
- Propagate synchronous errors in teardown via the completion promise instead of leaving completion undefined
- Add test coverage for terminated PTYs that exit mid-reattach and must stay dead
* fix(P1-B): recover system-SSH targets after a network drop
Two defects stopped a remote workspace auto-recovering after a blip.
runReconnectAttempt classified failures with isTransientError, which only
matches ETIMEDOUT/ECONNREFUSED/ECONNRESET by errno code or literal
substring. The system-SSH transport — the only transport FIDO2 and
ProxyUseFdpass targets can use — reports network failures as OpenSSH
prose ("System SSH connection timed out"), so the ladder published a
permanent 'error' on the first timeout and the target never came back
without a manual reconnect. isTransientReconnectError adds a
network-shaped prose table on top of isTransientError and is used only on
the reconnect path: connect() keeps the narrow classifier so an
unreachable host still fails fast instead of burning five 30s attempts
and five security-key touch prompts. Auth and passphrase failures stay
permanent on both paths.
runReconnectAttempt also had no generation fence, so a superseded attempt
published its cancellation as a permanent error over the winner's live
connection — reachable when a system-transport proc.onExit schedules a
reconnect while an attempt is still in flight. Cancellation now carries a
stable error name, and both connect() and runReconnectAttempt claim their
connectGeneration and stay silent when a newer attempt owns the state.
* fix(P1-B): retry a dropped watcher overflow marker on real capacity
emitWatcherOverflowToClient published the {kind:'overflow'} resync marker
with controlOverflow:'reject'. A full control queue rejects at admission
with no settlement callback, so the marker was silently discarded and the
remote File Explorer stayed stale until some later watcher event happened
to produce another one — for a quiet tree, possibly never.
The emitter now retains a rejected marker per (client, root) and
republishes it when the sink actually frees up. The existing
onLegacyPtyCapacity signal cannot drive that: it is gated on producer
retention, so it stays silent exactly under the dual-queue pressure that
caused the rejection. RelayDispatcher.onClientCapacity is an ungated
per-client capacity signal that fires on every writer settlement and
drain. It lives on the dispatcher rather than the writer so a retained
marker survives setWrite() replacing the primary sink, and setWrite
notifies capacity once afterwards so the marker does not wait on traffic
that may never arrive.
Retention is bounded to one marker per (client, root), released on
settlement and purged on client detach.
* fix(P1-B): address all review findings on SSH network recovery
Fix four issues from code review:
1. **Bug — admitted overflow markers lost on setWrite**: Retain markers when
settlement fails `ok: false`, not just on admission rejection. Prevents
desynced filesystem trees after SSH sink replacement.
2. **SSH error classification expanded**: Add missing OpenSSH patterns
(`ssh_exchange_identification`, `connection closed by remote`) and new
`isDefiniteSystemSshHostFailure()` classifier.
3. **ControlMaster retry optimization**: Skip second probe when first failure is
already definite host-level (network timeout, refused, unreachable). Saves
~30s per reconnect ladder step.
4. **Overflow flush under dual-queue pressure**: Gate pending marker retries on
control-lane headroom instead of re-attempting on every capacity notification.
Reduces thrash proportional to producer traffic.
Add regression tests for marker republish on sink replacement and validate auth
error detection against live OpenSSH credential rejection messages.
* rm random doc
* fix(P1-B): skip credential-failure retries and recover watcher markers o
- Auth and passphrase errors fail immediately without retry attempts
- Bare "System SSH probe failed (exit 255)" is transient only for reconnect
- Watcher markers survive client invalidation when switching SSH connections
- Add network error patterns: "lost connection", "remote end closed"
* fix(agent-status): preserve Claude background work
* fix(agent-status): harden background task lifecycle
* fix(agent-status): narrow interruption retention
* fix(agent-status): scope background task authority
* fix(agent-status): isolate lifecycle inventories
* fix(agent-status): harden background evidence recovery
* fix(agent-status): reject ambiguous child authority
* perf(agent-status): skip lifecycle inventory scans
* refactor(agent-status): isolate task inventory parsing
* fix(agent-status): clear stale background evidence
* fix(agent-status): gate accepted remote evidence
* test(agent-status): pin session cron interrupts
* fix: harden Claude inventory tracking
* test: pin Claude cron drain authority
* refactor(agent-status): unify Claude turn-boundary predicate
Collapse the five inline copies of the Stop/StopFailure test into a single
isTurnBoundary constant and drop the reportedStateName/stateName alias, so a
future edit can't move one copy and leave the others behind.
Pin the two behaviors that unification now depends on: a non-interrupted
StopFailure keeps gating on live background work, and interrupted state does
not survive a mid-turn lead event that has no prompt submit.
Co-authored-by: Orca <help@stably.ai>
* fix(agent-hooks): gate local Claude background evidence
---------
Co-authored-by: Orca <help@stably.ai>
* fix(ssh): gate FIDO2 system-transport on an OpenSSH binary
`ssh -G` echoes OpenSSH's built-in default identity list for every host, so
`usesDefaultPaths` was almost never true and the security-key gate returned
`!usesDefaultPaths || findSystemSsh() !== null` — forcing system transport
without checking that an `ssh` binary exists. `spawnSystemSsh()` then throws
`No system ssh binary found`, hard-failing connections that worked on ssh2.
The same flag also stopped the default scan at the first existing normal
private key, so a host that only accepts a FIDO2 key never reached system
OpenSSH when `~/.ssh/id_rsa` happened to exist.
Both decisions are independent of where an identity path came from: always
require `findSystemSsh() !== null` before forcing system transport, and scan
every candidate identity instead of stopping on the first normal key.
`shouldUseSystemSshTransport()` is untouched, so ProxyCommand / ProxyJump /
ProxyUseFdpass keep their intentional system transport.
* test(ssh): isolate connection tests from the developer's own FIDO2 keys
Transport selection now scans every default identity instead of stopping at
the first normal key, so a `~/.ssh/id_ed25519_sk` on the machine running the
suite would decide which transport the default-target tests take. Mock
`findSystemSsh` to null by default and opt the two security-key tests in.
* fix(ssh,relay): stop remote connections from being killed by backoff and frame caps
Three independent connection killers found in the SSH/remote freeze audit.
FINDING A - the reconnect ladder never escalated for post-handshake drops.
scheduleReconnect() used the single published state.reconnectAttempt for both
the delay index and the give-up test, and runReconnectAttempt() zeroed it
before connecting (ssh.ts gates the relay redeploy on 0-at-connected). Every
post-handshake drop therefore re-entered at 1000ms forever, ~3600 relay
redeploys/hour, and 'reconnection-failed' was unreachable for a flapping host.
New SshReconnectLadder splits the delay index (advanced by every retry) from
the failure streak (advanced only by a failed handshake), so flaps back off
while give-up semantics stay byte-identical to shipped.
FINDING B - notify() closed the client whenever a frame exceeded the producer
frame capacity, conflating a permanently un-sendable frame with transient
backpressure. A 5000-event fs.changed is 425KB against a 49KB cap, so the
watcher flood killed the link and re-killed on every reattach+replay. notify()
now drops and logs once per generation; fs.changed is chunked to each sink's
capacity with a control-lane overflow marker as the resync fallback; agent-hook
envelopes shed lastAssistantMessage/interactivePrompt/subagents to fit.
FINDING B2 - sendResponse routed >1MB responses to a lane whose admission
ignores the frame cap and closed the client on rejection, so a large
fs.listFiles dropped the SSH host. It now substitutes a JSON-RPC error so the
request fails instead of the connection.
Also moves fs.streamEnd/fs.streamError to the control lane so a terminal frame
cannot be dropped by the producer-lane check.
Co-authored-by: Orca <help@stably.ai>
* fix(relay): stop the overflow marker from re-killing the link it protects
Round-1 review fixes on the P0 freeze work.
The control-lane overflow marker could reinstate the exact failure this P0
removes: dispatcher-client-writer closes the client when control-lane
admission fails, and admitControl is the only lane that returns an error, so
one marker per failing batch accumulated to the 256-frame/1MB bound and
dropped the link. Markers are now deduped to one outstanding per
(client, root), cleared on settle.
Chunking also defeated the renderer's per-payload directory dedupe -- events
are now stable-grouped by parent directory so one directory lands in one
chunk -- and the halving walk overshot the byte minimum ~1.7x while the fast
path paid three JSON encodes; both are fixed by publishing first and sizing
from a measured bytes-per-event estimate.
Agent-hook shedding now surrenders the blocking interactive prompt LAST
rather than first, so a degraded envelope cannot strand a pane at
state=waiting with no answerable question card.
The dropped-notification log now distinguishes over-capacity from producer
queue backpressure and no longer lets the first dropped method silence every
other producer for the life of the connection.
* fix(relay,ssh): keep status delivery and terminal frames from trading one freeze for another
Round-2 review fixes.
The round-0 change from close-on-rejection to silent drop removed the only
redelivery path for agent.hook envelopes: they are fire-and-forget and the
per-pane cache only replays on handler install, so a saturated link stranded
a pane on a stale Working spinner until reconnect. Closing used to guarantee
delivery by forcing that replay. Envelopes now publish per client and pend
for bounded latest-wins redelivery when the producer queue rejects them.
Shed fields are now named on the wire. The subagent roster is not cosmetic --
the renderer replaces rather than merges it, and hibernation gates on its
length -- so an unmarked shed could sleep a live pane.
fs.streamEnd rode the control lane because it must not be dropped, but that
lane kills rather than drops. The stream's concurrency slot is now held until
the terminal frame settles rather than until the fd closes, capping queued
terminal frames well under the control budget; overflow costs one refused
read instead of the connection.
The watcher chunk walk now stops while producer retention sits past its
reserve and degrades to a resync, so a 5000-event flood cannot fill the queue
that interactive PTY traffic shares and stall every remote terminal.
The reconnect ladder caps its flap-path delay so delay plus handshake timeout
cannot cross the relay grace floor and let the remote daemon kill live PTYs.
Also: the suppression key no longer embeds a NUL byte, which had made the
file binary to git and grep; producerEnvelopeBudget no longer reports
infinite capacity for a departed client; the drop logger no longer encodes a
frame it will not log; and an over-capacity response substitution no longer
settles as if the result had been delivered.
* fix(relay,ssh): restore relay-shed status fields and scope backpressure per client
Round 3 + 4 review fixes.
Watcher chunking is now gated on the *client's* retention reserve rather than
the dispatcher-wide one, so one stalled peer no longer forces a healthy client
into a full file-tree resync. The relay-lost redeploy ladder no longer burns its
6-attempt budget while the SSH transport itself is down: it holds at the 15s step
with a non-terminal status and rearms, so a laptop that slept past the ladder
comes back instead of landing on a terminal "give up" banner.
The shedFields wire marker had no consumer, so an agent-hook envelope whose
subagent roster was dropped to fit the frame read as "roster cleared" on the Orca
side: live child rows blanked and a done pane became hibernation-eligible while
its teammates were still running. ingestRemote now restores shed fields from the
cached payload (interactivePrompt deliberately excluded — a stale answerable
question card is worse than none).
Also: stream terminal-frame slots are counted per client, since the control queue
they protect is per client; the chunking fast path no longer logs a drop for a
batch it goes on to deliver in full; -32010 is now RelayErrorCode.ResponseOverCapacity.
Test debt from the review: pending-pane eviction, per-client stream isolation, and
the reconnect budget are now asserted rather than assumed; four fragile exact-byte
pins dropped in favour of the tier comparisons that carry the requirement.
* fix(relay,ssh): restore relay-shed status fields and scope backpressure
- Oversized relay responses now fail their request instead of closing the connection,
preventing one frame from killing every pane on the host
- Restore subagent state for correct hibernation; don't resurrect stale prose
across turns
- Account for relay re-establishment and PTY reattach time in SSH flap delay caps
- Only log drops of final unsendable envelopes, not temporary rejections during
measurement probes
- Fix watcher overflow marker release race when notification admission rejects
without settlement; use precise byte counting for event batching
* Restore relay-shed fields with digest validation and scoped backpressure
Validate that shed subagent rosters match their wire digest and turn identity before
restoration, preventing stale roster resurrection. Compact interactive prompts for waiting
states instead of dropping them. Demote control-queue overflow to non-fatal rejection so
clients can retry on capacity recovery, keeping the link alive during transient backpressure.
* fix(relay): correct ResponseOverCapacity error code
ResponseOverCapacity should use -33008 to stay in the -33xxx range
for relay protocol errors, not -32010.
* fix(relay): close client when pty.replay overflows control queue
Replay is never retried, so it uses the control lane where overflow
is fatal — the writer closes the client and reconnect reloads history
rather than stranding a short buffer.
* fix(relay): prevent infinite redeploy on flapping SSH transports
Charge reconnect attempts when connection restores mid-backoff, preventing
infinite loop on transports that flap between states. Refactor control overflow
handling to use entry property instead of WeakSet marker for clarity.
---------
Co-authored-by: Orca <help@stably.ai>
* fix(ssh): key the PTY model-migration fence by app pty id
Co-authored-by: Orca <help@stably.ai>
* test(ssh): pin the post-recovery checkpoint rekey
The finishSourceRecovery app-id rekey had no coverage: reverting it left every
suite green while reconnects silently resumed from the stale migration-era
checkpoint.
Co-authored-by: Orca <help@stably.ai>
---------
Co-authored-by: Orca <help@stably.ai>
* feat(cli): add `orca account add` / `account list` for headless hosts
The desktop "Add account" UI is disabled when the renderer drives a remote
runtime (isRemoteAccountScope === kind:'environment'), so a headless server
reached from a remote desktop/web client has no way to register managed
Claude accounts. Add a host-local CLI path that reuses the existing capture
logic:
- ClaudeAccountService.addAccountFromConfigDir(): register a managed account by
capturing credentials from an already-authenticated CLAUDE_CONFIG_DIR instead
of spawning the interactive browser login (extracted persist/rollback helpers
shared with the existing add flow)
- RPC accounts.addClaudeFromConfigDir, bridged via OrcaRuntime; rejected for
mobile device tokens (host-local only)
- `orca account add` runs `claude login` in the user's own terminal into a temp
CLAUDE_CONFIG_DIR, then registers it via the local runtime; `orca account list`
lists managed accounts
Switching (select) already works from a remote client; only adding was blocked.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(cli): support Codex in `orca account add` / `account list`
Mirror the Claude headless-account CLI for Codex:
- CodexAccountService.addAccountFromHome(): register a managed Codex account by
importing auth.json from an already-authenticated CODEX_HOME, reusing a shared
persist helper extracted from doAddAccount (no interactive login spawned here)
- RPC accounts.addCodexFromHome + OrcaRuntime.addCodexAccountFromHome bridge,
rejected for mobile device tokens (host-local only)
- `orca account add --agent claude|codex` (default claude); `orca account list`
now renders both Claude and Codex managed-account blocks
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test: cover headless account-add capture paths (Claude + Codex)
- ClaudeAccountService.addAccountFromConfigDir: registers a managed account by
capturing an authenticated CLAUDE_CONFIG_DIR; rejects and rolls back when the
dir has no .credentials.json
- CodexAccountService.addAccountFromHome: imports auth.json from an
authenticated CODEX_HOME into a managed account; rejects when auth.json is
missing
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix: address CodeRabbit review on headless account-add flows
- CLI login spawn uses a shell on Windows so `.cmd` agent shims resolve without
ENOENT (args are fixed literals, no injection risk)
- Claude capture skips the `.credentials.json` precheck on macOS, where creds
live in the Keychain and captureAuthFromConfigDir reads them
- Claude add rollback is best-effort: a failed rematerialization no longer skips
managed-auth cleanup or masks the original add error
- Codex persist restores the prior account/selection if a post-write sync or
rate-limit refresh fails, so a failure can't leave a dangling managed account
- Codex sync passes the account's selection target (correct runtime for WSL)
- Add JSDoc to the new public service methods and CLI functions
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cli): harden headless account capture
* fix(cli): correct account command flag surface and interrupt cleanup
- `account` commands no longer accept or advertise the browser `--page`
flag; `supportsBrowserPageFlag` allow-listed them by omission, so
`orca account list --page x` was silently accepted and `--help`
rendered a browser-only option
- account specs declare GLOBAL_FLAGS, so `--help`/`--json` render in the
Options block like every other command
- `--agent` on `account add` documents the account provider instead of
the terminal TUI-agent meaning inherited from the shared flag table
- a SIGINT/SIGTERM during the interactive login now removes the temp
login dir (and restores the macOS Keychain item) before exiting 130;
Node terminates without unwinding `finally`, which stranded live OAuth
credentials on disk
* perf(cli): stop `account list` forcing a provider usage refresh
`accounts.list` awaited refreshAccountsForMobile(), which runs
fetchAll({ force: true }) — bypassing both the poll throttle and the
per-provider Retry-After gate — then O(N) serial per-account round
trips. `orca account list` renders only emails and the active ids, so
all of that work was discarded. The RPC now takes `refreshUsage`
(default true, so mobile and web keep the forced lane) and the CLI opts
out. Older hosts declare `params: null` and ignore the field, so a newer
CLI degrades to the previous behavior rather than failing.
Also documents on `account list` that `--environment` does not retarget
it, matching the host-local behavior of shouldIgnoreRemoteSelection.
* fix(cli): survive repeated and hangup signals during account add
withInterruptCleanup latched cleanup behind a boolean, so a second signal
got an already-resolved promise and its process.exit fired while the first
cleanup was still inside a Keychain call (3s each) — the temp dir's OAuth
credentials and the swapped macOS Keychain item both survived. Memoize the
cleanup promise so every signal awaits the same run, and register with
`on` instead of `once` so a second Ctrl-C cannot fall through to Node's
terminate-immediately default mid-cleanup.
Handle SIGHUP too. This flow exists for headless/SSH hosts, where the most
likely interrupt is the connection dropping, which hangs up the login's
terminal and previously ran no cleanup at all.
Warn when the interrupt lands after sign-in completed: the runtime finishes
the add independently of this process, so exiting 130 silently would tell
the user it was cancelled when the account may exist.
Reject a valueless `--agent`; the parser turns it into boolean true, which
silently ran a full OAuth login for Claude when the user asked for another
provider.
Also lock two behaviors the refactor changed but left uncovered: a WSL Codex
add must sync the WSL runtime lane rather than the default host lane, and
rename the account-spec help test to describe the Options block it actually
asserts rather than the usage string it never reads.
* fix(build): bundle the main modules the account CLI imports
electron-vite cleans out/main and emits only its declared entries, and
`build:desktop` runs it after `build:cli`, so the tsc-emitted copies of
`claude-accounts/keychain`, `codex-cli/command` and `win32-utils` were
deleted before packaging. Both `orca account add` and `orca account list`
then died at require time with "Cannot find module
'../../main/claude-accounts/keychain'" — reproduced against a real
`--serve` host. `agent-hooks/managed-agent-hook-controls` already carried
an entry for exactly this reason; these three were missing.
Adds a parity test so any future CLI import of a `src/main` module fails
in CI rather than at a user's shell after packaging.
* test: cover the desktop add-path behavior this PR changes
Both changes ride in the persist/rollback helpers the existing GUI add
flow shares with the new headless path, and neither had coverage:
- Claude: rollbackAddAccount now guards forceMaterializeCurrentSelection-
ForRollback, so a rejecting rematerialization no longer replaces the
real add error nor skips safeRemoveManagedAuth. Asserts the original
error surfaces and the throwaway auth dir is gone.
- Codex: the desktop add now passes the account's selection target to
syncForCurrentSelection, matching reauthenticate and select. Asserts
the host target alongside the existing WSL assertion.
Both fail when the corresponding change is reverted.
* fix(cli): close the remaining account-add interrupt and preflight gaps
The round-1 interrupt fix detached the signal handlers before running the
finally-path cleanup, so the very window it was meant to protect — the two
serial 3s `security` calls plus rmSync on the success/error path — was
still covered only by Node's terminate-immediately default. Both review
lanes reproduced it independently. Await cleanup first, detach in a nested
finally, and stop a cleanup failure from replacing the error that actually
explains why the add failed.
Do not burn the interactive login when the runtime is unreachable. The
RuntimeClient is lazily constructed and the first call was the registration
RPC itself, so "Requires the Orca runtime to be running" was discovered
only after the user completed a full OAuth round trip. Preflight with the
now-cheap `accounts.list { refreshUsage: false }`.
Reject `--environment` / `--pairing-code` on `account add`.
shouldIgnoreRemoteSelection pins account commands to the local runtime, so
`orca account add --environment homelab` silently registered the account on
the laptop instead of the headless host it names.
Survive a daemon that cannot spawn `claude`. `allowFailure` is honored in
onClose but not onError, and unlike the GUI flow nothing has run `claude` in
the daemon before this point — so a launchd/systemd daemon with a minimal
PATH hard-failed an add the user had already signed in for, even though
identity resolves fine from the config dir's oauthAccount.
Also align the `--agent` help description with the global flag column.
* fix(cli): reject runtime selectors on `account list` too
`orca account list --environment homelab` was accepted and silently
listed the LOCAL machine's accounts, because shouldIgnoreRemoteSelection
pins account commands to the local runtime. Documenting that in --help
does not reach someone who already typed the flag, and answering with the
wrong host's accounts is the specific wrong answer they would act on.
`account add` already errors; this makes the new command group internally
consistent. The other groups in shouldIgnoreRemoteSelection keep their
existing silent-ignore behavior — changing those is not this PR's job.
* test: harden account-add signal tests and cover cleanup failure
- Identify the handler under test by set difference instead of
`process.listeners(sig).at(-1)`. Vitest installs its own once-wrapped
SIGINT teardown, so the positional lookup could grab the wrong listener;
the helper also asserts exactly one new listener was added.
- Mock rmSync while keeping the real implementation by default, so the
temp-dir assertions elsewhere stay honest.
- Cover that a cleanup failure in the `finally` does not replace the error
explaining why the add failed. Fails when that guard is removed.
Completes the review loop's final round; the loop died on an API error
before it could commit this, and its `import()` type annotation would
have failed oxlint.
* fix(cli): harden interactive account add
* test(cli): make account cancellation coverage portable
* fix(cli): preserve merged skills runtime modules
---------
Co-authored-by: Dominik <marketing@gavaplast.sk>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
* fix(ssh): connect to Linux hosts that cannot compile node-pty
node-pty ships no Linux prebuilt at any architecture, so it is compiled on
the remote. On a host without a C/C++ toolchain that build fails, and because
both native deps install in one npm command it also took down
@parcel/watcher — which does have a working Linux prebuilt — and failed the
whole connection. Every Linux image without build tools was unusable.
node-pty only backs remote terminals; files, git, and the editor do not need
it, and a missing native dep is already non-fatal further down the deploy. So
when the existing toolchain probe confirms the compiler is missing, reinstall
without node-pty instead of aborting. The manifest has to drop it too — npm
reconciles every dependency in package.json, not just the ones named on the
command line, so naming only @parcel/watcher still rebuilds node-pty.
If that reinstall also fails the actionable build-tools error is rethrown, so
a host broken for some other reason still reports the toolchain gap.
The relay's PTY error now names the fix rather than saying only that node-pty
is unavailable.
Verified on a stock Rocky Linux 10.2 aarch64 container (openssh-server, git,
nodejs, npm, no compiler): connect succeeds, /etc lists over SSH, node-pty is
absent while @parcel/watcher installs its linux-arm64-glibc prebuilt, and
spawning a terminal reports the install hint.
* fix(ssh): keep the node-pty skip path honest about platform and watcher
The PTY unavailable message named build tools unconditionally, but only Linux
compiles node-pty — the deploy-side skip is gated on linux and the toolchain
probe returns null on Windows. A Windows or macOS remote, where node-pty ships
prebuilds, was told to install make/g++/python3. Pick the remedy by the relay's
own platform.
The skip path returned before the install probe, so a @parcel/watcher that
installs but cannot require() (glibc below the floor) connected with dead file
watching and nothing logged. Probe before returning and warn; no rebuild, since
node-pty provably cannot compile on that host, and never fatal.
Also log the pty-less reinstall's own failure and attach it as cause — the
rethrown toolchain message is built from the original npm error, so an
unrelated retry failure (registry, ENOSPC, EACCES) was lost. The reinstall now
keeps the caller's resetDeps as well, so a repair reconnect still clears every
dep the probe found broken.
Tests: the skip-success fixture queued a chmod/probe/rebuild sequence
production never runs, and the surplus slots were absorbed by launchRelay's
readiness poll (1817ms vs 3-9ms for its peers). It now emits exactly the 12
execs production performs, and pins that no rebuild is issued. Adds the missing
negative case: a gyp-shaped failure on a host whose probe reports a complete
toolchain must still hard-fail rather than silently degrade.
* fix(ssh): hedge the node-pty remedy and keep repair resets on the skip path
* Fix ssh-relay install on hosts with split shell/SFTP namespaces
On Synology DSM and similar hosts, the SSH shell and SFTP subsystem expose
different absolute paths for the same directory (e.g., /var/services/homes/alice
vs /homes/alice). The relay installer silently picked the wrong path and failed
discovery. This fix implements SFTP namespace detection: each install creates an
unguessable ownership marker and probes both namespaces to detect divergence.
When paths differ, SFTP writes redirect to the candidate namespace while shell
commands keep the canonical path. Markers are random tokens redacted from logs.
* Fix ssh-relay install on hosts with split shell/SFTP namespaces
Strengthen path validation to catch traversal and empty segments in
absolute POSIX paths, preventing security issues. Improve split-namespace
handling with comprehensive wire tests for uploads and file writes.
Ensure system SSH connections bypass namespace mapping entirely rather
than attempting incorrect retargeting.
Allow the existing "Open in" entries to launch a configured VS Code
launcher against an SSH-backed worktree via Remote-SSH:
code --remote ssh-remote+<authority> <remote-path>
- Split the blanket SSH/runtime block into a capability model: file
managers and non-VS Code launchers stay local-only (disabled with
"Local only" metadata); a recognized VS Code command is enabled and
forwarded with connectionId over a typed object IPC.
- Main process stays authoritative: rejects active/owned runtimes,
resolves the SshTarget from the persisted Store, derives the authority
(config alias, or username@host on port 22, or ssh-alias-required on a
non-default port), validates POSIX/Windows absolute remote paths without
local stat/normalize, and rejects non-VS Code and compound commands
before spawn.
- Authority and remote path are passed as separate argv; getSpawnArgsForWindows
remains the cmd/bat shim boundary and fails closed on metacharacters.
- Same capability rules across the worktree menu, Explorer overflow, and
the source-control entry context menu.
Refs STA-2386
Closes#9999
* feat(linear): add MCP-style save issue
* fix(linear): harden save issue parity
* fix(linear): close save issue contract gaps
* docs(linear): bundle project discovery with save issue
Move managed agent-hook filesystem work behind one relay RPC so high-latency SSH connects pay one WAN round trip instead of hundreds. Keep installers serial, lock shared account config across relay processes, and fence cancelled connection generations from replacement state.
Co-authored-by: nasagong <zinho2000@gachon.ac.kr>
Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
Collapse multi-line explanatory comment blocks into single-line "why" statements
per AGENTS.md ("Document the Why, Briefly"): drop restatements of the code and
mechanism narration; keep the non-obvious reason, external refs, and directives.
Comments-only — verified no code changed via a Babel/esbuild comment-strip
token-equality gate against origin/main; typecheck and oxlint clean.
Area: main — git, source-control, providers & integrations. 40 files changed, 1432 insertions(+), 4473 deletions(-).
Co-authored-by: Orca <help@stably.ai>