mirror of
https://github.com/stablyai/orca.git
synced 2026-10-08 08:02:32 +00:00
cf71ae4cb6f8202a7cc8a424aa84d127c845cbe2
42
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
93d8b1f042 |
fix(ssh): complete keyboard-interactive MFA prompt handling (#15588)
Honor SSH keyboard-interactive prompt echo and empty responses, reuse login passwords without replaying rejected values, and stop cancelled or stale credential requests from continuing authentication or restoring the cache. Original implementation: Junho Kim (#8750). Port and follow-up work: Allen (#15588). Verified with 2,773 SSH/credential tests, full typecheck, changed-code quality, a production Electron build, real-socket MFA fixtures, and rendered UI checks. Fixes #8622 Co-authored-by: Junho Kim <arkimjh@illinois.edu> Co-authored-by: microdaery <microdaery@gapp.nthu.edu.tw> |
||
|
|
fbe94ceff6 |
fix: close readiness gaps found by merged-change audit (#17159)
* fix(ssh): fence stale kills and retired pane replay * fix(ssh): support cancellable interactive authentication * fix(ssh): await remote catalog before snapshot adoption * fix(pty): contain Windows ConPTY input failures * fix(power): avoid redundant macOS display blocking * perf(editor): narrow markdown override subscriptions * fix(quick-open): close directory handles after reads * refactor(linux): remove unused proc socket scanner * fix(usage): apply flat Sonnet 4.6 pricing * ci: prime Node next native test cache * docs(skills): resolve snapshot cleanup data path * fix(ssh): recover install locks after host reboot * test(ssh): recognize boot-aware install locks * test(ssh): prove previous-boot lock recovery live * test(wire): pin pre-metadata release coverage * fix(terminal): preserve remote tab ownership through recovery races * test(runtime): fence replaced terminal handles in agent guard * fix(ssh): preserve remote snapshot authority across polls * fix(pty): contain late ConPTY output EPIPE * test(pty): register Windows exit watcher before kill * fix: close SSH and tab readiness race gaps * fix(tabs): retain headless order and placeholder titles * fix(build): avoid parallel electron-vite config race * test(windows): avoid MSYS temp path rewriting * test(windows): avoid killing exited PTY * fix(pty): avoid late ConPTY input teardown race * fix(terminal): sync reconnect error ownership after commit * fix(runtime): use canonical worktree identity comparison * test(ssh): assert complete cold-hydration baseline * test(windows): invoke quoted retention fixture via PowerShell * test(windows): read ConPTY grid through mode con * fix(terminal): publish PTY replacements atomically * fix(terminal): infer stale identity on reattach * fix(terminal): fence stale pane PTY callbacks * fix(terminal): fence stale pane binds after rebind * fix(terminal): reject stale pane transport callbacks * fix(terminal): fence mirrored reattach spawn callbacks * fix(terminal): replace stale pane PTYs on remount * fix(ci): size the Windows launcher-compile test budget from measurement `native-smoke (windows-latest)` fails ~4.5% of runs on `preserves a multiline argument through the compiled remote launcher` with "Test timed out in 15000ms" — on unrelated PRs, for reasons that have nothing to do with them. Across 176 sampled attempts it is the only red that job produced, and it hit seven different PRs in two days: #16900, #16904, #16915, #16955 (twice), #16979, #17014, #17085. The test is six process creations: powershell.exe forks csc.exe, then the freshly compiled orca.exe forks node.exe, twice. Hosted Windows runners periodically slow process creation down, and this test amplifies that far harder than anything else in the job. Comparing the 80 attempts where it ran under 3s against the 12 where it ran over 12s, its own median goes 2198ms -> 15917ms (7.2x) while the same file's powershell-only test moves 556 -> 686ms (1.2x), the cmd.exe and Git Bash process tests in the neighbouring file move 1.4x, and the other 35 files put together move 1.5x. Measured across those 176 attempts: 1881ms to 35438ms, p50 4264ms, correlation +0.881 with the job's total Vitest duration. 8 of 176 (4.5%) exceeded the 15s cap; 2 of 176 (1.1%) also exceeded the shared 30s testTimeout, so deleting the override and inheriting the config is not enough on its own. 60s clears all 176 with 1.7x headroom on the worst. This is slow, not hung. Every body here is synchronous spawnSync, so Vitest cannot interrupt one — the timer fires only after the body returns and the reported duration is real elapsed time. That is why a failure reads `× ... 22464ms` under `Test timed out in 15000ms`. The work finished; the stopwatch was short. Seven reruns at one identical head measured 2053 / 4680 / 5551 / 8732 / 13506 / 14868 / 21937ms — the last of those would have been red on code that had not changed. The 15s came from #8897, which raised this test off Vitest's built-in 5s default because the job then ran bare `pnpm vitest run`. #8909 landed 3h27m later and pointed the job at config/vitest.config.ts, which is the real fix for that. The constant stayed behind and has been the binding budget ever since. * fix(terminal): fence stale remount reattach ownership * fix(terminal): reconcile mounted pane identity after replacement * fix(terminal): fence stale reattach fallback ownership * fix(terminal): fence deferred SSH reattach ownership * fix(terminal): fence stale split pane ownership callbacks * fix(terminal): keep stale spawns from consuming startup --------- Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com> |
||
|
|
24e662adc1 |
feat(ssh): verify host keys, and restore panes correctly across a reconnect (#14844)
* docs(ssh): design for real host key verification (STA-4319)
Today's ssh2 verifier records a fingerprint and returns true — every host key is
accepted, with no known_hosts consult and no change detection anywhere in
src/main/ssh/. Scope is per-connection, so exec, SFTP, port forwarding, the
watcher and relay deploy all ride that one unverified handshake, and the
ProxyJump path puts the final hop — the topology most likely to cross untrusted
network — on ssh2 specifically.
Decisions worth calling out:
- Read the user's known_hosts as a trust source but NEVER write to it. That file
is shared with every other SSH tool on the machine; appending means line
endings, permissions, concurrent writers and a corruption blast radius well
beyond us. Accepted keys go to our own per-target store. Reading theirs is also
the entire migration story: most developers already have their hosts there.
- Mismatch is scoped to the SAME key type. A host with only an RSA entry that
presents ed25519 is unknown, not changed. ssh2 negotiates ed25519 first, so
without this we would fire a change-of-key alarm at nearly every existing user
on their first upgraded connect — training them to dismiss the one warning that
is supposed to mean something. Flagged in review as the decision I am least
sure of; a downgrade-vector argument against it is being tested.
- Changed key hard-fails with no override button; recovery is a separate explicit
action, offered only when OUR store is what disagreed, because forgetting our
record cannot unblock a known_hosts conflict.
- Background reconnects deny rather than prompt. A dialog the user cannot place
in context only teaches click-through.
Two traps are documented because either would make the fix silently do nothing:
an async verifier returns a Promise, which ssh2 reads as truthy and accepts
immediately; and the existing test mock invokes hostVerifier with one argument
and ignores the return, so it would pass against a verifier that never decides.
Design only — no behaviour change. The doc is added to the tracked-reference
allowlist in .gitignore alongside the other docs/reference entries.
* docs(ssh): revise the host key design after security and migration review
Three things the reviews changed, kept visible rather than quietly edited out.
THREAT MODEL WAS WRONG IN THREE PLACES. Jump hosts are not the worst case — they
are already safe: shouldUseSystemSshTransport branches on exactly the inputs
resolveEffectiveProxy does, and attemptConnect returns after the system probe, so
ProxyJump goes through OpenSSH and is verified. Agent forwarding was overstated
(gated on the user's ForwardAgent). Credential theft was understated: any auth
error counts as agent fallback, so a MITM walks the user to the password AND
private-key passphrase prompts, and cachedPassword replays without prompting. The
relay claim was backwards — the attacker owns their own machine; the real impact
is the return direction, where they become the host our workspace trusts.
TYPE SCOPING IS A DOWNGRADE VECTOR WITHOUT ALGORITHM ORDERING. This was the
decision I flagged as least certain and asked to have argued both ways. OpenSSH
is safe only because order_hostkeyalgs() puts known types first and RFC 4253
gives the client's order priority. ssh2 negotiates ed25519 first regardless, so
an attacker who cannot forge the RSA key on file just presents ed25519 and gets a
friendly first-contact prompt instead of a hard failure. Keep scoping, but set
algorithms.serverHostKey to lead with the types on file — and add a sixth
outcome for 'unknown type, known host', which must never read as first contact.
SHIP THE DEFENCE BEFORE THE DIALOG. Startup restore fires eager connects for all
targets in parallel with a 15s timeout while a prompt would live 120s; ephemeral
VM targets present a new key every launch; paired-web connects run on the host
desktop, so the dialog opens on someone else's screen. Phase 1 is therefore no
modal at all: consult known_hosts and our store, match connects, unknown persists
with accept-new semantics, mismatch and revoked hard-fail. That is the whole MITM
defence with none of the migration risk.
Also folded in, verified live against OpenSSH 10.2p1: the without-port fallback
(bracketed lookup first, then bare, where the second pass can only yield match or
unknown — otherwise a bare line plus a non-default port produces a spurious
prompt); hashed entries hash the candidate form; multiple files union; a
cert-authority line does not match a plain key. IPv6 and bracket parsing moved
INTO scope — that is a parser requirement, not a scope call, and getting it wrong
produces the prompt-training harm the design exists to avoid.
* feat(ssh): parse and match OpenSSH known_hosts
The matcher half of STA-4319. No behaviour change yet — nothing calls this.
Hand-rolled because no maintained JS implementation exists, and written against
behaviour observed from OpenSSH 10.2p1 rather than inferred from the man page.
Three of those behaviours a reasonable reading gets wrong:
- A non-default port is TWO ordered lookups, not one candidate set: '[host]:port'
first, then bare host ('checking without port identifier' in ssh -v). The
fallback pass can only yield match or unknown — OpenSSH downgrades a wrong key
there rather than reporting a change. Collapse them and anyone holding a bare
line who connects off-port gets a spurious first-contact result; treat the
fallback as authoritative and they get a false change-of-key alarm.
- Revocation resolves in its own pass so the verdict cannot depend on line order.
Verified both orderings.
- A cert-authority line never matches a plain host key; it only validates
certificates. A normal line alongside it still decides.
Mismatch is scoped to the same key type, and a host known by a DIFFERENT type
returns unknown-type-known-host rather than plain unknown — an attacker who
cannot forge the key on file must not get a friendly first-contact result by
presenting another type. That outcome is only half the defence; the other half
(leading serverHostKey with known types) lands with the wiring.
47 tests from vectors executed against real sshd, including ssh-keygen -H hashed
entries. Each of six mutations reddens it: collapsing the passes, letting the
fallback report mismatch, dropping type scoping, resolving revocation in line
order, honouring an unrecognised marker, and skipping the blob/type agreement
check.
* feat(ssh): decide what to do with a presented host key
The policy half of STA-4319, kept separate from the ssh2 wiring so it is testable
without a handshake and injected rather than importing its sources, so a test
states its own trust state instead of writing files.
Phase 1 ships no dialog — a test asserts the decision is never 'prompt'. Startup
restore opens every previously-active target at once, ephemeral VM targets would
ask every launch, and paired-web connects run on the host desktop where the
dialog would appear on someone else's screen.
Ordering that matters: revocation outranks everything including
StrictHostKeyChecking=no, because a revoked key is a statement that this key is
known-bad rather than merely unrecognised. known_hosts is named before our own
store on a change, because its remedy (ssh-keygen -R) is the one that also
unblocks ssh and git — pointing at a remedy that cannot work is worse than none.
Two carve-outs with reasons: an ephemeral runtime target accepts WITHOUT
recording, since a fresh VM presents a new key every launch and a stored record
would accumulate per launch and eventually read as a spurious change; and when
ssh -G ran on the HOME-divergent path that suppresses /etc/ssh/ssh_config, an
unknown host is denied, because a site-wide policy may forbid it and being laxer
than ssh is the one outcome that is never acceptable.
Rejection text deliberately avoids 'authentication failed' and 'permission
denied': the reconnect ladder classifies on those substrings, so a denial phrased
that way is retried forever against a decision that will never change. Pinned by
a test.
* feat(ssh): build the host key verifier and the algorithm order that makes it safe
Still not wired into the handshake — that lands next. This is the piece that
turns a decision into an ssh2 callback, plus the half of the design that is easy
to forget because it lives in a different config field.
The verifier MUST be a plain function returning undefined. ssh2 does
'const ret = verifier(key, verify); if (ret !== undefined) verify(ret)', so an
async function returns a Promise — neither undefined nor falsy — and ssh2 accepts
the key immediately while ignoring whatever the callback later decides. Making
this async would silently restore exactly the accept-everything behaviour the
module exists to remove, so a test asserts the return value is undefined.
orderServerHostKeyAlgorithms is what makes type-scoped matching safe rather than
a downgrade. RFC 4253 gives the client's algorithm order priority, so leading
with the types we already hold for a host denies a server the choice of
presenting some other type to convert a hard failure into first contact. Without
it, an attacker who cannot forge the key on file just offers a different
algorithm. Revoked entries never contribute to that order.
Also fails closed on two paths that would otherwise hang or over-trust: a key
whose own length-prefixed header cannot be read is refused rather than reasoned
about, and a throw from any dependency denies, because ssh2 may not catch an
exception raised inside the verifier and the handshake would hang instead of
failing.
18 tests. Includes the two negative cases that matter — first-contact keys are
recorded, but keys we already know, rejected keys, ephemeral runtime targets and
a lax StrictHostKeyChecking are not.
* fix(ssh): promote every RSA signature algorithm for a known ssh-rsa key
A known_hosts entry names the KEY type, which is not the negotiated ALGORITHM
name. One ssh-rsa key is offered as rsa-sha2-512, rsa-sha2-256 or ssh-rsa
depending on the signature algorithm, so matching the literal name only would
leave a host we know by RSA ordered behind ed25519 — precisely the ordering this
function exists to prevent, and precisely the population (RSA-era known_hosts
entries) it was written for.
Verified from ssh2's own negotiation while wiring this: kex.js iterates the
CLIENT list and takes the first entry the server also offers, so client order
does decide, as RFC 4253 says. ssh2's default order leads with ed25519 and places
the RSA algorithms fifth through seventh.
* fix(ssh): verify host keys instead of accepting every one (STA-4319)
The actual fix. ssh-connection's verifier recorded a fingerprint and returned
true, so every ssh2 connection accepted every host key — no known_hosts consult,
no change detection. It now consults the user's known_hosts plus our own store
and refuses a changed, revoked or unverifiable key.
Phase 1 by design: no dialog. Unknown hosts are accepted and recorded
(accept-new semantics), because startup restore opens every previously-active
target at once, ephemeral VM targets present a new key each launch, and
paired-web connects run on the host desktop where a prompt would appear on
someone else's screen. The MITM defence lands now; the prompt is Phase 2.
Also sets algorithms.serverHostKey to lead with the types already known for the
host. Without it the type-scoped matching is a downgrade — an attacker who cannot
forge the key on file just presents another type and turns a hard failure into
first contact. Verified from ssh2's kex.js that the client list decides.
Denial replaces ssh2's generic handshake error with the specific reason, because
the reconnect ladder cannot distinguish a generic failure from a transient fault
and would retry forever against a decision that will never change.
An unreadable trust store degrades to known_hosts only rather than failing the
connect: a changed key is still refused, and a host trusted only by us falls back
to first contact and is re-recorded, reaching the same decision.
The ssh2 mock now uses the callback form and aborts the handshake on denial. As
written it called hostVerifier(key) with one argument and ignored the result, so
it would have passed against a verifier that never decides — flagged in the
design as a mock that had to change, not a test to quietly rewrite. Two new tests
pin the wiring rather than the module: an unidentifiable blob is refused, and a
well-formed key is accepted.
Note for review: commit
|
||
|
|
9367169888 |
refactor(tests): split every oversized test file off the max-lines suppression list (#14728)
* refactor(tests): split oversized test files off the max-lines suppression list Every `*.test.ts`/`*.spec.ts` that carried an `eslint/oxlint-disable max-lines` directive is now split into focused, behavior-scoped suites that fit the 800-line test budget, with shared setup extracted into co-located `*-test-harness.ts` / `*-test-fixtures.ts` modules (300-line budget). 83 files became ~930; the largest output is 797 effective lines. `orca-runtime.test.ts` is intentionally untouched. Test bodies were moved by scripted line-range slicing rather than retyped, so assertions are byte-identical. The only permitted body edits were mechanical rebinding where a shared value moved into a harness (e.g. `tmpHome` -> `homes.tmpHome`). Registries that enumerate test files were updated in lockstep: - config/max-lines-baseline.txt: pruned 341 -> 258 entries (all 83 removed). - config/reliability-gates.jsonc: 33 gates repointed at the split files, with assertionRefs split per file where a gate's coverage now spans several. - .github/workflows/pr.yml: the real-zsh lane now lists the 4 split files that actually exercise zsh, so they keep running in the dedicated shell lane. Also renamed agent-hooks `server-test-fixtures.ts` to `server.test-fixtures.ts` so the global-fetch call-site audit keeps skipping it, and added `.js` extensions to the CLI suites' dynamic harness imports (node16 resolution) to unbreak `build:cli`. Verification: full suite 52,449 passing vs 52,448 at baseline with zero assertions lost; `pnpm lint`, `pnpm typecheck`, and `pnpm build:cli` all exit 0; the terminal-pane e2e spec runs 31/31 headless. * refactor(tests): split hook-idle arbitration suite that oxfmt pushed over budget The pre-commit oxfmt pass reflowed pty-connection-hook-idle-arbitration.test.ts to 811 effective lines, 11 over the test budget. Split the hook-completion side effect and replacement-agent veto cases into their own suite; both files now sit well under the cap and the 15 tests are unchanged. * test: port upstream test changes into the split files after rebase Rebasing onto main surfaced 27 tests that main had added to files this branch deleted, plus edits to tests that had already moved. Taking the deletion side of those modify/delete conflicts would have dropped that coverage silently, so each upstream change is ported into the split file that now owns the behavior — for example main's six orchestration mailbox tests land across orchestration-runs, -send, and -check. Also repoints `orchestration.notification-mailbox-consistency`, a gate main added after this branch's gate remap, at those same three split files, and re-prunes the max-lines baseline against main's (257 entries). Verified: all 27 upstream test titles present; full suite 52,761 passing with the only diff vs baseline being 12 tests main itself removed and 3 that moved from skipped to passing; lint and typecheck exit 0. * fix(test): flush pending continuations before tearing down terminal test globals CI shard 5/16 failed on both Node 24 and 26 with `ReferenceError: window is not defined` from pty-connection.ts, surfacing through pty-connection-daemon-snapshot-replay.test.ts. The reattach/settle chains `await` a real promise and then touch `window.api`. Under fake timers those continuations cannot run, so they only become schedulable once restoreTerminalTestGlobals() switches back to real timers — which previously happened immediately before `delete globalThis.window`, so a late continuation threw and failed the whole file. Flush async ticks in that window instead. This is latent in the source rather than new: the pre-split 25k-line file kept running other tests after these, which gave the chains time to settle before teardown. Splitting the file moved teardown directly behind them. * fix(test): keep an inert window after terminal test teardown instead of deleting it The async-tick flush was not enough: the reattach/settle chain can resolve after teardown regardless of how long we drain, so CI shard 5/16 still failed with `ReferenceError: window is not defined` from pty-connection.ts. A real renderer never loses `window`, so deleting it was the artificial part. Swap in an inert proxy whose properties resolve to callables and whose calls resolve to undefined, making a late `window.api.pty.*` call a harmless no-op. The next test replaces it wholesale via installTerminalTestGlobals(), and no test asserts that `window` is absent. |
||
|
|
637c7e94c9 |
Add SSH config host picker to add-host dialog (#12334)
* feat(ssh): add SSH config host picker for add-host form Users can now click 'Fill from ~/.ssh/config…' to browse available SSH config hosts in a picker, select one, and have the form automatically prefill with resolved connection details (hostname, port, username, auth). Previously, an 'import' button provided bulk sync on this form—confusing and unhelpful when everything was already synced. That action is now available as a secondary 'Add all' option in the picker. * fix(ssh): import filter preservation and label fallback - Reuse search loader on import completion to preserve active filter inside generation guard - Fall back to hostname when manual host has no label, not empty string - Make alias duplicate detection case-insensitive to match config picker behavior - Validate host availability when restoring project group selection - Add aria-selected attribute to picker options for accessibility * fix(ssh): harden config picker import, alias folding, and host targeting Review findings on the ~/.ssh/config picker + bulk add: - Guard config-host resolution with a generation counter so a late resolve cannot overwrite a later pick or a form the user backed out of; freeze the other rows while a pick resolves. - Stop "Add all N" from re-adopting deleted hosts — it now imports without reAdopt, matching the new-host count it advertises. Settings → Import keeps the explicit re-adopt path. - Fold SSH aliases through a shared normalizeSshConfigAlias for import ownership, delete tombstones, reclaim, picker search, and the save-time duplicate check, which now occupies configHost *and* label like the picker. - Persist GSSAPIAuthentication only when a parsed Host entry asks for it, not when `ssh -G` merely echoes the /etc/ssh system default. - Fail closed with unavailable/setup-not-found when an explicit projectHostSetupId names a non-actionable host instead of silently creating the workspace on a sibling host. - Cache the parsed config for the picker session (refresh on open/retry) so filter keystrokes no longer reparse and Include-expand the file, keep the filter usable during loads, add a Retry on load errors, explain an empty Identity file after a config fill, and drop the always-false aria-selected. * refactor(ssh): centralize host result limit and extract folder group val Move SSH_CONFIG_HOST_RESULT_LIMIT to shared types so the renderer's limit message cannot drift from the host's query limit. Extract findActionableFolderProjectGroup to avoid repeating the folder-host-availability check across the composer hook. * fix(ssh): pass -F to ssh -G when HOME differs from passwd home In E2E tests and sandboxes, isolated HOME can differ from the system passwd home. OpenSSH resolves the default config via getpwuid (passwd), while Node's loadUserSshConfig uses os.homedir() (HOME-aware). Pass -F to explicitly specify the config path when they diverge, so ssh -G and the picker resolve the same file. * fix(ssh): verify config host exists before resolving with ssh -G When a user edits ~/.ssh/config and removes a host, the import picker should not fall back to ssh -G's echoed response (which treats any alias as valid). Check the reloaded config file before resolving. - Force reload config on each resolve to catch user edits post-open - Reject aliases not in the current config before calling ssh -G - Add test for deleted alias edge case - Fix workspace-target fallback to honor explicit host selection * fix(ssh): let tombstoned aliases be re-picked in the config picker Allow users to reclaim a deleted SSH host by re-picking it from ~/.ssh/config. Tombstoned aliases now appear in the picker with a "Removed from Orca" badge and remain pickable, but don't count toward "Add all" operations — ensuring passive import never resurrects a deleted alias while still giving the user a recovery path. |
||
|
|
a000839465 |
Add first prompt to agent session history rows (#12085)
* Add first user prompt to AI Vault session history rows Re-parse transcripts on demand to extract and display the untruncated first user prompt for copy/reuse. List scans omit the body (payload/perf); UI loads it when session details expand. Grok sessions extract the typed ask from <user_query> envelope, skipping injected <user_info> bootstrap rows. Supports Claude, Codex, Grok, and OpenCode agents. * fix(ai-vault): split SessionTime out to pass max-lines lint AiVaultSessionDetails exceeded the 400-line oxlint limit after adding first-prompt UI; move SessionTime into its own module. * fix(ai-vault): handle corrupt transcripts and fix OpenCode prompt captur Corrupt transcripts now resolve null instead of rejecting the IPC call, matching behavior for other unavailable cases. OpenCode SQLite parsing now correctly captures all text parts from the earliest user message only, fixing truncation of large prompts and padding of small ones. Add stale-response guard in the UI to prevent late results from overwriting the current session when tabs switch. Consolidate text slicing via `sliceAtCodeUnitLimit` to avoid surrogate-pair splits across all callers. * test(ai-vault): add first-user-prompt UTF-16 safety tests Ensure truncation at safety limits doesn't split UTF-16 surrogate pairs, preventing corruption of astral characters in captured prompts. * fix(ai-vault): key first-prompt-card by session.id Remounting the card on session switches prevents late responses from a previous load from writing stale data into the component's refs. Also improves conversation-turn key stability. * fix(ai-vault): preserve first prompt after preview truncation * refactor(ai-vault): improve first user prompt capture robustness and per - Add 15s timeout to full-prompt load to prevent indefinite loading states - Extract seedFullFirstUserPrompt helper for reuse across parsers - Prevent AI-generated summaries from becoming the copyable first prompt - Fix truncation detection in OpenCode SQLite by probing for N+1 rows - Optimize text bounding to apply safety limit before toLowerCase - Gate synthetic OpenCode path detection on agent type, not just # presence - Add test coverage for remote execution host handling * Fix FirstPromptCard loading state stranded by stale promise reuse Clears loadPromiseRef during cleanup to prevent the dedupe handle from causing StrictMode remounts to await stale in-flight requests. Stops loading when session becomes non-loadable mid-request. Adds tests for StrictMode double-invoke resolution and main-process timeout scenarios. * refactor(ai-vault): split session parsers into modular files Split secondary-parsers into individual files per agent type (copilot, cursor, hermes, opencode) for improved modularity. Add test coverage for first-user-prompt envelope handling: unwrap user_query tags and reject bare user_info dumps. * fix(ci): clear max-lines and flaky portal readiness check Collapse an accidental multi-line regex wrap in ssh-connection-utils that pushed counted lines to 301. Harden the latched-readiness test's ready transition so CI load can re-observe attach after MutationObserver gaps. * fix(ssh): extract proxy command helpers to pass max-lines Move resolveEffectiveProxy/spawnProxyCommand out of ssh-connection-utils so oxfmt line wrapping cannot push that file over the 300-line lint cap. * capture first user prompt by ordering OpenCode messages by creation time - Add `readOpenCodeMessagesInOrder` to rebuild transcript by timestamp, handling corrupt/partial files gracefully instead of discarding sessions - Extract SSH proxy command tests to dedicated file; add backpressure handling and stderr draining to prevent proxy process stalls - On Windows, reject unsafe characters in ProxyCommand values instead of pretending to escape them; properly format cmd.exe invocation with verbatim arguments - Expand ProxyJump chains into -J plus final hop, mirroring OpenSSH behavior - Decouple portal readiness reapply budget from flip-count budget via explicit constant |
||
|
|
ce5b639e03 |
fix(P1-B): recover SSH targets and remote file watchers after a network drop (#12032)
* fix(P1-B): recover system-SSH targets after a network drop
Two defects stopped a remote workspace auto-recovering after a blip.
runReconnectAttempt classified failures with isTransientError, which only
matches ETIMEDOUT/ECONNREFUSED/ECONNRESET by errno code or literal
substring. The system-SSH transport — the only transport FIDO2 and
ProxyUseFdpass targets can use — reports network failures as OpenSSH
prose ("System SSH connection timed out"), so the ladder published a
permanent 'error' on the first timeout and the target never came back
without a manual reconnect. isTransientReconnectError adds a
network-shaped prose table on top of isTransientError and is used only on
the reconnect path: connect() keeps the narrow classifier so an
unreachable host still fails fast instead of burning five 30s attempts
and five security-key touch prompts. Auth and passphrase failures stay
permanent on both paths.
runReconnectAttempt also had no generation fence, so a superseded attempt
published its cancellation as a permanent error over the winner's live
connection — reachable when a system-transport proc.onExit schedules a
reconnect while an attempt is still in flight. Cancellation now carries a
stable error name, and both connect() and runReconnectAttempt claim their
connectGeneration and stay silent when a newer attempt owns the state.
* fix(P1-B): retry a dropped watcher overflow marker on real capacity
emitWatcherOverflowToClient published the {kind:'overflow'} resync marker
with controlOverflow:'reject'. A full control queue rejects at admission
with no settlement callback, so the marker was silently discarded and the
remote File Explorer stayed stale until some later watcher event happened
to produce another one — for a quiet tree, possibly never.
The emitter now retains a rejected marker per (client, root) and
republishes it when the sink actually frees up. The existing
onLegacyPtyCapacity signal cannot drive that: it is gated on producer
retention, so it stays silent exactly under the dual-queue pressure that
caused the rejection. RelayDispatcher.onClientCapacity is an ungated
per-client capacity signal that fires on every writer settlement and
drain. It lives on the dispatcher rather than the writer so a retained
marker survives setWrite() replacing the primary sink, and setWrite
notifies capacity once afterwards so the marker does not wait on traffic
that may never arrive.
Retention is bounded to one marker per (client, root), released on
settlement and purged on client detach.
* fix(P1-B): address all review findings on SSH network recovery
Fix four issues from code review:
1. **Bug — admitted overflow markers lost on setWrite**: Retain markers when
settlement fails `ok: false`, not just on admission rejection. Prevents
desynced filesystem trees after SSH sink replacement.
2. **SSH error classification expanded**: Add missing OpenSSH patterns
(`ssh_exchange_identification`, `connection closed by remote`) and new
`isDefiniteSystemSshHostFailure()` classifier.
3. **ControlMaster retry optimization**: Skip second probe when first failure is
already definite host-level (network timeout, refused, unreachable). Saves
~30s per reconnect ladder step.
4. **Overflow flush under dual-queue pressure**: Gate pending marker retries on
control-lane headroom instead of re-attempting on every capacity notification.
Reduces thrash proportional to producer traffic.
Add regression tests for marker republish on sink replacement and validate auth
error detection against live OpenSSH credential rejection messages.
* rm random doc
* fix(P1-B): skip credential-failure retries and recover watcher markers o
- Auth and passphrase errors fail immediately without retry attempts
- Bare "System SSH probe failed (exit 255)" is transient only for reconnect
- Watcher markers survive client invalidation when switching SSH connections
- Add network error patterns: "lost connection", "remote end closed"
|
||
|
|
de75003df9 |
fix(P1-C): gate FIDO2 system-SSH transport on an OpenSSH binary (#12029)
* fix(ssh): gate FIDO2 system-transport on an OpenSSH binary `ssh -G` echoes OpenSSH's built-in default identity list for every host, so `usesDefaultPaths` was almost never true and the security-key gate returned `!usesDefaultPaths || findSystemSsh() !== null` — forcing system transport without checking that an `ssh` binary exists. `spawnSystemSsh()` then throws `No system ssh binary found`, hard-failing connections that worked on ssh2. The same flag also stopped the default scan at the first existing normal private key, so a host that only accepts a FIDO2 key never reached system OpenSSH when `~/.ssh/id_rsa` happened to exist. Both decisions are independent of where an identity path came from: always require `findSystemSsh() !== null` before forcing system transport, and scan every candidate identity instead of stopping on the first normal key. `shouldUseSystemSshTransport()` is untouched, so ProxyCommand / ProxyJump / ProxyUseFdpass keep their intentional system transport. * test(ssh): isolate connection tests from the developer's own FIDO2 keys Transport selection now scans every default identity instead of stopping at the first normal key, so a `~/.ssh/id_ed25519_sk` on the machine running the suite would decide which transport the default-target tests take. Mock `findSystemSsh` to null by default and opt the two security-key tests in. |
||
|
|
a07427e970 |
fix(ssh, relay): keep remote sessions alive through reconnects and backpressure (#11999)
* fix(ssh,relay): stop remote connections from being killed by backoff and frame caps Three independent connection killers found in the SSH/remote freeze audit. FINDING A - the reconnect ladder never escalated for post-handshake drops. scheduleReconnect() used the single published state.reconnectAttempt for both the delay index and the give-up test, and runReconnectAttempt() zeroed it before connecting (ssh.ts gates the relay redeploy on 0-at-connected). Every post-handshake drop therefore re-entered at 1000ms forever, ~3600 relay redeploys/hour, and 'reconnection-failed' was unreachable for a flapping host. New SshReconnectLadder splits the delay index (advanced by every retry) from the failure streak (advanced only by a failed handshake), so flaps back off while give-up semantics stay byte-identical to shipped. FINDING B - notify() closed the client whenever a frame exceeded the producer frame capacity, conflating a permanently un-sendable frame with transient backpressure. A 5000-event fs.changed is 425KB against a 49KB cap, so the watcher flood killed the link and re-killed on every reattach+replay. notify() now drops and logs once per generation; fs.changed is chunked to each sink's capacity with a control-lane overflow marker as the resync fallback; agent-hook envelopes shed lastAssistantMessage/interactivePrompt/subagents to fit. FINDING B2 - sendResponse routed >1MB responses to a lane whose admission ignores the frame cap and closed the client on rejection, so a large fs.listFiles dropped the SSH host. It now substitutes a JSON-RPC error so the request fails instead of the connection. Also moves fs.streamEnd/fs.streamError to the control lane so a terminal frame cannot be dropped by the producer-lane check. Co-authored-by: Orca <help@stably.ai> * fix(relay): stop the overflow marker from re-killing the link it protects Round-1 review fixes on the P0 freeze work. The control-lane overflow marker could reinstate the exact failure this P0 removes: dispatcher-client-writer closes the client when control-lane admission fails, and admitControl is the only lane that returns an error, so one marker per failing batch accumulated to the 256-frame/1MB bound and dropped the link. Markers are now deduped to one outstanding per (client, root), cleared on settle. Chunking also defeated the renderer's per-payload directory dedupe -- events are now stable-grouped by parent directory so one directory lands in one chunk -- and the halving walk overshot the byte minimum ~1.7x while the fast path paid three JSON encodes; both are fixed by publishing first and sizing from a measured bytes-per-event estimate. Agent-hook shedding now surrenders the blocking interactive prompt LAST rather than first, so a degraded envelope cannot strand a pane at state=waiting with no answerable question card. The dropped-notification log now distinguishes over-capacity from producer queue backpressure and no longer lets the first dropped method silence every other producer for the life of the connection. * fix(relay,ssh): keep status delivery and terminal frames from trading one freeze for another Round-2 review fixes. The round-0 change from close-on-rejection to silent drop removed the only redelivery path for agent.hook envelopes: they are fire-and-forget and the per-pane cache only replays on handler install, so a saturated link stranded a pane on a stale Working spinner until reconnect. Closing used to guarantee delivery by forcing that replay. Envelopes now publish per client and pend for bounded latest-wins redelivery when the producer queue rejects them. Shed fields are now named on the wire. The subagent roster is not cosmetic -- the renderer replaces rather than merges it, and hibernation gates on its length -- so an unmarked shed could sleep a live pane. fs.streamEnd rode the control lane because it must not be dropped, but that lane kills rather than drops. The stream's concurrency slot is now held until the terminal frame settles rather than until the fd closes, capping queued terminal frames well under the control budget; overflow costs one refused read instead of the connection. The watcher chunk walk now stops while producer retention sits past its reserve and degrades to a resync, so a 5000-event flood cannot fill the queue that interactive PTY traffic shares and stall every remote terminal. The reconnect ladder caps its flap-path delay so delay plus handshake timeout cannot cross the relay grace floor and let the remote daemon kill live PTYs. Also: the suppression key no longer embeds a NUL byte, which had made the file binary to git and grep; producerEnvelopeBudget no longer reports infinite capacity for a departed client; the drop logger no longer encodes a frame it will not log; and an over-capacity response substitution no longer settles as if the result had been delivered. * fix(relay,ssh): restore relay-shed status fields and scope backpressure per client Round 3 + 4 review fixes. Watcher chunking is now gated on the *client's* retention reserve rather than the dispatcher-wide one, so one stalled peer no longer forces a healthy client into a full file-tree resync. The relay-lost redeploy ladder no longer burns its 6-attempt budget while the SSH transport itself is down: it holds at the 15s step with a non-terminal status and rearms, so a laptop that slept past the ladder comes back instead of landing on a terminal "give up" banner. The shedFields wire marker had no consumer, so an agent-hook envelope whose subagent roster was dropped to fit the frame read as "roster cleared" on the Orca side: live child rows blanked and a done pane became hibernation-eligible while its teammates were still running. ingestRemote now restores shed fields from the cached payload (interactivePrompt deliberately excluded — a stale answerable question card is worse than none). Also: stream terminal-frame slots are counted per client, since the control queue they protect is per client; the chunking fast path no longer logs a drop for a batch it goes on to deliver in full; -32010 is now RelayErrorCode.ResponseOverCapacity. Test debt from the review: pending-pane eviction, per-client stream isolation, and the reconnect budget are now asserted rather than assumed; four fragile exact-byte pins dropped in favour of the tier comparisons that carry the requirement. * fix(relay,ssh): restore relay-shed status fields and scope backpressure - Oversized relay responses now fail their request instead of closing the connection, preventing one frame from killing every pane on the host - Restore subagent state for correct hibernation; don't resurrect stale prose across turns - Account for relay re-establishment and PTY reattach time in SSH flap delay caps - Only log drops of final unsendable envelopes, not temporary rejections during measurement probes - Fix watcher overflow marker release race when notification admission rejects without settlement; use precise byte counting for event batching * Restore relay-shed fields with digest validation and scoped backpressure Validate that shed subagent rosters match their wire digest and turn identity before restoration, preventing stale roster resurrection. Compact interactive prompts for waiting states instead of dropping them. Demote control-queue overflow to non-fatal rejection so clients can retry on capacity recovery, keeping the link alive during transient backpressure. * fix(relay): correct ResponseOverCapacity error code ResponseOverCapacity should use -33008 to stay in the -33xxx range for relay protocol errors, not -32010. * fix(relay): close client when pty.replay overflows control queue Replay is never retried, so it uses the control lane where overflow is fatal — the writer closes the client and reconnect reloads history rather than stranding a short buffer. * fix(relay): prevent infinite redeploy on flapping SSH transports Charge reconnect attempts when connection restores mid-backoff, preventing infinite loop on transports that flap between states. Refactor control overflow handling to use entry property instead of WeakSet marker for clarity. --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
33c14bc716 |
fix(ssh): fall back to OpenSSH for FIDO2 keys (#11913)
Closes #11645 |
||
|
|
5fe3aaf2b7 |
fix(ssh): preserve first config directive value (#11297)
* fix(ssh): preserve first config directive value * test(ssh): cover false-first config booleans * fix(ssh): trust fresh OpenSSH config authority * fix(ssh): preserve ordered config identities --------- Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
a40183389b |
feat: bound direct SSH reconnect fan-out and recovery (#11003)
* docs: design for direct SSH reconnect fan-out Capture the implementation-ready plan for host-qualified, epoch-fenced SSH reconnect recovery after two rounds of multi-model LLM counsel review. * docs: reconcile SSH reconnect fan-out design * docs: close reconnect design consistency gaps * feat: implement bounded direct SSH reconnect recovery * fix: bound direct SSH retry settlement * fix: harden direct SSH reconnect authority * fix: preserve split SSH retry ownership * fix: preserve SSH split continuation authority * docs: record final SSH reconnect validation * fix: preserve SSH authority through retained and detached state * fix: retain SSH authority across delayed split mounts * fix: close SSH authority recovery gaps * fix: fence stale SSH transport replacement * fix: serialize SSH target teardown * fix: settle SSH teardown failures before reconnect * fix: retire failed SSH reset sessions * test: reconcile current main E2E contracts * fix: close direct SSH reconnect review gaps * fix: fence stale SSH reconnect side effects * fix: close final SSH reconnect lifecycle gaps * test: stabilize current-main reliability gates * test: prove plugin navigation containment * test: make plugin navigation oracle authoritative * test: make plugin navigation oracle deterministic * ci: allow sharded e2e suite to finish * test: wait for runtime pane publication * test: classify pane readiness by error code * test: select close persistence terminal by tab identity * docs: mark reconnect implementation validated --------- Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
badf91101b |
fix(quality): enforce performance-safe lint baseline (#11074)
* fix(quality): clear safe existing lint findings * fix(quality): keep lint cleanup allocation-free * fix(quality): enforce performance-safe baseline * test(terminal): drain deferred confirmation cleanup |
||
|
|
6d39e49480 |
fix(ssh): accept GitHub restricted-shell SSH probes (#6988) (#7659)
* fix(ssh): accept GitHub restricted-shell SSH probes (#6988) * fix: match first stderr line for GitHub restricted-shell probe (bug-bash takeover) Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com> Co-authored-by: Orca <help@stably.ai> |
||
|
|
aab112933e |
Revert "fix(memory): bound OOM-prone accumulators (#10179)" (#10255)
Co-authored-by: Orca <help@stably.ai> |
||
|
|
8f40ddf328 | fix(memory): bound OOM-prone accumulators (#10179) | ||
|
|
4f81dfc128 |
perf(ssh): cut warm high-latency connects from 88.7s to 7.7s (#9015)
Move managed agent-hook filesystem work behind one relay RPC so high-latency SSH connects pay one WAN round trip instead of hundreds. Keep installers serial, lock shared account config across relay processes, and fence cancelled connection generations from replacement state. Co-authored-by: nasagong <zinho2000@gachon.ac.kr> Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
8ed8f0d109 |
feat(ssh): download folders from remote explorer (#7793)
* feat(ssh): download folders from remote explorer * fix(ssh): harden remote folder downloads * Add missing getRepo stub to worktree cwd test mock Restoring headless mobile tabs looks up the repo for the active worktree id; the mock lacked getRepo, so the test only passed incidentally. Add it explicitly and return undefined since wt-1 is a worktree id, not a registered repo. * Enable SSH folder downloads, gated for system-SSH connections - Folder downloads require SFTP, unavailable on system SSH (which offers only raw file operations). Add supportsFolderDownload flag to gate the feature in the UI layer. - Reject symlinks at directory-entry level, preventing tree escapes and eliminating unnecessary stat calls. - Check abort signal before opening dialog for better responsiveness when renderer closes. - Log cleanup errors without re-throwing to preserve underlying transfer failures. * Gate SSH folder downloads to SFTP-capable connections Enforce fail-closed gating and add Windows path traversal validation to ensure downloads are only available when explicitly supported and safe. * Gate SSH folder downloads to SFTP-capable connections Enforce fail-closed gating and add Windows path traversal validation to ensure downloads are only available when explicitly supported and safe. * fix(ssh): keep provider types under max-lines after main merge Move FolderDownloadOptions next to the SFTP download implementation and narrow IFilesystemProvider.downloadFolder options to AbortSignal only so types.ts stays within the 300-line oxlint budget when merged with main. --------- Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com> |
||
|
|
fa85536f3a |
fix(ssh): repair unbuilt relay native deps (#8686)
Co-authored-by: Orca <help@stably.ai> Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com> Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com> Co-authored-by: Jinwoo Hong <73622457+Jinwoo-H@users.noreply.github.com> |
||
|
|
190ed7d4ed |
fix(ssh): make POSIX relay wrapper survive csh/tcsh login shells (#8709)
Reconstruct remote POSIX commands with bounded printf arguments so non-POSIX SSH login shells can forward them without requiring remote base64. Preserve relay and system-SSH stdin, and centralize login-shell flag selection for csh/tcsh compatibility.\n\nValidated against real csh and tcsh OpenSSH targets with built-in and system SSH, including cold relay deployment, stdin upload, PTY I/O, file mutation, and reconnect. |
||
|
|
0302ae86b8 |
feat(ssh): support Kerberos/GSSAPI hosts via the system OpenSSH transport (#7507)
* feat(ssh): support Kerberos/GSSAPI hosts via the system OpenSSH transport ssh2 has no gssapi-with-mic support, and adding it would mean forking its protocol layer plus packaging the kerberos native module for three platforms. Instead, route GSSAPI hosts through the existing system-OpenSSH transport, which delegates Kerberos (tickets, SSPI on Windows) to the platform ssh binary. Two tiers, because RHEL-family distros enable GSSAPIAuthentication globally in /etc/ssh/ssh_config and ssh -G therefore reports it for every host: - Targets whose ~/.ssh/config Host block explicitly sets GSSAPIAuthentication yes (imported as target.gssapiAuthentication) try system ssh first, falling through to ssh2 so key auth and credential prompts still work when no ticket is available. - When ssh2 exhausts key/agent auth and the ssh -G-resolved config enables GSSAPI, retry over system ssh before prompting for credentials, so Kerberos-only hosts on distro-default configs connect without a password prompt. Hosts where keys work never leave the ssh2 path. Manual targets flagged for GSSAPI pass -o GSSAPIAuthentication=yes explicitly since they bypass ssh_config. Both tiers work headless (no credential callbacks required). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ssh): harden GSSAPI transport selection (review fixes for PR #7507) Review fixes on top of the Kerberos/GSSAPI feature branch (s546126/kerberos-ssh): - HIGH: reset useSystemSshTransport on the ssh2 fall-through. doSystemSshProbe sets the flag before spawnSystemSshCommand, which throws synchronously when no system ssh binary is on PATH (outside the probe try/catch). The proactive fall-through previously reset only 2 of 3 transport fields, so exec/sftp kept routing through the failed transport - breaking GSSAPI on Windows-with-Git-ssh and headless Linux. - MEDIUM: throw a cancellation error (not the stale ssh2 authError) when a disconnect supersedes the reactive probe mid-flight, and guard connect()'s catch on disposed, so a deliberate disconnect is not overwritten with auth-failed. - MEDIUM: skip the encrypted-key passphrase prompt when the GSSAPI fallback applies, so a Kerberos ticket is tried before prompting; the general prompt still fires if the probe fails. Adds 3 mutation-verified regression tests and hardens two existing tests to assert the probe actually ran. Not connected to any PR remote. Co-authored-by: Orca <help@stably.ai> * fix(ssh): isolate GSSAPI system transport Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: s546126 <268420947+s546126@users.noreply.github.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Neil <4138956+nwparker@users.noreply.github.com> Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com> Co-authored-by: Orca <help@stably.ai> |
||
|
|
c2aa9c2ead |
feat(ssh): support file transfers over system ssh (#7804)
* feat(ssh): support file transfers over system ssh * fix(ssh): harden system file transfers * test(runtime): stub getRepo in headless mobile tab cwd test The mobile-session selector validator (getValidatedExplicitWorktreeIdSelector, from main) calls this.store?.getRepo to reject repo ids passed as worktree ids. The store stub only implemented getWorkspaceSession, so the guarded call threw 'getRepo is not a function' once main merged into this branch. Add a getRepo that returns null (wt-1 is a worktree, not a repo). Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com> Co-authored-by: Orca <help@stably.ai> |
||
|
|
f215a48064 |
fix: pr-bug-scan validated finding from #6952 (#7180)
* fix: address pr-bug-scan validated finding from #6952 throwNodeNotFound() now re-raises AbortError when the shared signal is aborted, so a signal-cancelled node probe no longer launders into 'Node.js not found'; sequential fallback runs. * fix(ssh): make session-limited (MaxSessions=1) relay deploys actually succeed Review of #7180 verified the parent fallback end-to-end against a real MaxSessions=1 sshd and found the connect still failed. Four gaps, in order of discovery: - isSshSessionLimitError missed stock OpenSSH, which refuses session channels over MaxSessions with SSH2_OPEN_CONNECT_FAILED (2) and 'open failed' — reason 4 never matched, so the fallback never triggered. - execCommand settled aborted commands before the channel finished closing, so the sequential fallback reissued execs while sshd still counted the old session. - SshConnection.waitForSshCallback rejected aborts mid-channel-open immediately, leaking a confirmed-late channel that held the only session slot; it now settles after the late channel closes (bounded) and drains its streams so ssh2 emits 'close'. - Session channel opens now retry transient session-limit refusals (sshd frees the slot only after processing our close-ack, which the next open can beat by microseconds), and the remote orca CLI shim install is non-fatal like the managed-hook install — after the relay bridge occupies the sole session slot, raw-connection extras must degrade instead of failing the connection. Verified live against Docker sshd (OpenSSH 9.2, MaxSessions=1): fresh deploy (upload + native deps + launch), reconnect cycles, and a PTY round-trip all succeed; unrestricted-sshd regression run also passes. Co-authored-by: Orca <help@stably.ai> * Handle ssh execution aborts immediately during retry backoff or hangs - Cancel the session-limit retry delay immediately if the operation is aborted during backoff. - Limit the wait time to a 5-second grace period when aborted during a channel open that is hung and never invokes its callback, rather than waiting for the full connection timeout. --------- Co-authored-by: orca-bug-scan-bot <orca-bug-scan-bot@stably.ai> Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com> Co-authored-by: Orca <help@stably.ai> |
||
|
|
e6ca3087c5 |
perf(ssh): resolve node path concurrently with remote install state (#6952)
Run independent SSH relay bootstrap probes concurrently when the connection can safely support overlapping execs. Preserve the old sequential path for system SSH without reusable ControlMaster and for remotes that reject concurrent session channels. |
||
|
|
4dbc9f3817 |
feat(ssh): add ControlMaster multiplexing for system SSH transport (#6922)
* feat(ssh): add ControlMaster multiplexing for system SSH transport System SSH transport spawns a new OpenSSH process per exec command (platform detect, relay install check, node resolution, relay launch, socket probe). Each process pays the full SSH handshake cost — ~9s on Uber devpods — making a typical relay connect take 54s+ and reliably exceeding the 15s startup reconnect budget. Add SSH ControlMaster multiplexing via a per-target socket in $TMPDIR/orca-ssh-ctl/<hash>.sock. The first command establishes the master; subsequent commands reuse it at ~100ms per exec instead of ~9s. ControlPersist=300 keeps the master alive after commands exit so rapid reconnects (e.g. on tab focus) also benefit. Windows is excluded since OpenSSH's ControlMaster support there is limited. * fix(ssh): address ControlMaster key collision and directory permission risks - Use target.id in the socket key so distinct SSH targets can never collide even when configHost/port/user happen to match - Switch from SHA1 to SHA256 and extend hash slice from 12 to 16 chars - Stat the control-socket directory after mkdirSync to reject pre-existing dirs that are symlinks, foreign-owned, or have group/other write bits (mkdirSync mode is ignored on pre-existing dirs) - Update two tests that used exact spawn-arg arrays; replace with ordering assertions (forward flags before --) that stay correct regardless of which extra ControlMaster options are injected * fix(ssh): bind ControlPath identity to route and reject symlinked ctl dir Fold proxyCommand/jumpHost/identity fields into the ControlPath hash so a target whose route is edited no longer reuses a still-alive master built on the old route. Switch the control-socket dir check from statSync to lstatSync so a planted symlink fails the directory validation outright. * test(ssh): drop tautological argv re-assertion in spawn checks The toHaveBeenCalledWith re-passed the args array extracted from the same mock call, making that argument position always pass. argv content is already verified by the index-ordering assertions above; use expect.any(Array) so the spawn check only claims what it actually verifies (binary path, stdio). * fix(ssh): harden system ssh connection reuse Co-authored-by: Orca <help@stably.ai> * test(ssh): isolate control socket runtime dir Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Test <test@example.com> Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com> Co-authored-by: Orca <help@stably.ai> |
||
|
|
46646d7ff1 |
chore(lint): upgrade oxlint to 1.71 + enable 7 new rules (autofixed backlog) (#6841)
* chore(lint): upgrade oxlint to 1.71 and enable 7 new rules Upgrade oxlint 1.67.0 -> 1.71.0 (1.72 was blocked by the repo's 3-day minimum-release-age supply-chain guard; nothing here needs it). The bump is a no-op on the existing config. Enable 3 error rules (backlog autofixed to zero in this commit) and 4 warn rules (surface signal without gating CI): error (autofixed, behavior-preserving): - unicorn/prefer-node-protocol (~1531 sites: bare builtin -> node:) - typescript/no-import-type-side-effects (~36: all-inline-type -> import type) - unicorn/no-array-reverse (19: copy-then-reverse -> toReversed) warn (real signal, current fires are test-only/correct): - unicorn/no-array-fill-with-reference-type (aliasing footgun guard) - typescript/no-unsafe-function-type (bans bare Function type) - unicorn/prefer-array-flat-map (map().flat() -> flatMap()) - unicorn/prefer-regexp-test (.match() in bool ctx -> .test()) mobile/.oxlintrc.json extends root, so it inherits all 7; the autofix ran from root and covered mobile/ too. Verification (all green): oxlint 0 errors (root+mobile+aux configs), oxfmt clean, typecheck (node+cli+web), vitest 22795 passed / 0 failed, builds (electron-vite + web + cli) succeed. node: rewrites confirmed to skip embedded SSH/CLI string payloads (AST-only); all toReversed sites verified to operate on fresh copies or write-once locals. * chore(lint): bump mobile oxlint to 1.71 so inherited rules parse mobile/ is a standalone pnpm project pinning its own oxlint@1.67, which lacks unicorn/no-array-fill-with-reference-type (needs >=1.70). Since mobile/.oxlintrc.json extends the root config, mobile CI's 'cd mobile && oxlint' failed to parse the new rule. Bump mobile to match root (1.71). Verified in mobile/: oxlint 0 errors, oxfmt --check clean, tsc --noEmit pass, vitest 978 passed / 0 failed. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
c455c576ae |
Fallback to system SSH on reachability errors (#5303)
Co-authored-by: Orca <help@stably.ai> |
||
|
|
98d02bca47 |
fix: support windows ssh hosts (#5004)
* feat: add windows ssh relay base support * feat: support windows ssh relay runtime services * fix: default windows ssh pty cwd to user profile * fix: support windows hosts over system ssh * fix: preserve degraded windows relay native deps * fix: gate windows shell args by relay platform * fix: preserve windows relay fallback pipes * test: align windows native deps relay fixture * fix: build valid windows install lock command * fix: address windows SSH relay review findings Resolve correctness, efficiency, and reuse issues found reviewing the Windows SSH native-host support: - GC liveness on Windows now probes the actual named pipe (via node net.connect against markers + deterministic candidates) instead of substring-matching Win32_Process command lines, which could remove a live relay dir. Reports ALIVE conservatively only when there is no liveness signal at all (no markers and no seed pipes). - Resolve the remote node path once per deploy and thread it through install/repair/launch instead of re-resolving 3-7x. - Replace the 200ms node -e poll loop with a single long-lived remote wait process during Windows relay startup. - Skip the no-op executable command on Windows in uploadRelay. - Make the Windows fallback pipe name deterministic and recoverable (drop the global counter), with an extra reconnect attempt. - Normalize the prepended node bin dir to backslashes on Windows PATH. - Batch the system-SSH Windows directory upload into a single streamed JSON package instead of one ssh process per file. - Extract relay endpoint/marker helpers into ssh-relay-endpoints.ts and consolidate the PowerShell EncodedCommand encoding into the shared powershell-command-encoding module. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Support cancellation and timeouts in Windows port scanning - Propagate the request AbortSignal and a 5-second timeout to both PowerShell and netstat child processes during Windows port scanning. - Avoid spawning the netstat fallback process if the port scan has already been aborted. - Wrap the .NET OSArchitecture check in a try/catch block during SSH Windows platform detection to robustly fall back to environment variables if needed. --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: Jinjing <6427696+AmethystLiang@users.noreply.github.com> |
||
|
|
9c2cec3d6c |
Use system ssh for proxy-based targets (#4911)
Co-authored-by: Orca <help@stably.ai> |
||
|
|
c7c2330723 | fix: close late ssh channel callbacks (#4116) | ||
|
|
ccfc1a9b97 |
fix: time out ssh channel opens (#3742)
* fix: time out ssh channel opens * test: cover ssh channel-open timeout edge cases Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com> Co-authored-by: Orca <help@stably.ai> |
||
|
|
8044be7cc7 | fix: clean up system ssh probe listeners (#3824) | ||
|
|
41f0bc6394 |
test: preserve ssh mock event listeners (#3015)
* test: preserve ssh mock event listeners * test: reuse runtime file command helper |
||
|
|
186327af08 | fix: guard ssh startup destroy errors (#3013) | ||
|
|
9d396e115c | fix: remove SSH startup listeners after connect (#2882) | ||
|
|
eb579a0e54 | Fix SSH sessions after host sleep | ||
|
|
ffec74c7ba |
Fix SSH agent and identity file auth ordering (#2681)
* fix: address review findings * fix: address CI failures |
||
|
|
9c67557a0c |
Support system SSH transport for fdpass proxies (#2279)
Co-authored-by: Orca <help@stably.ai> |
||
|
|
fdc38bfc33 | Fix SSH relay connect when default shell is fish (#2239) | ||
|
|
9a39b1345a |
fix(ssh): enable TCP_NODELAY on ssh2 client to eliminate per-keystroke typing lag (#1660) (#1679)
* fix(ssh): enable TCP_NODELAY on ssh2 client to eliminate per-keystroke typing lag (#1660) ssh2 leaves Nagle's algorithm on by default. For single-byte keystrokes through a remote PTY, Nagle interacts with the kernel's delayed-ACK timer and adds up to ~40 ms per keystroke — visible as the typing lag reported in #1660. OpenSSH's `ssh` client sets TCP_NODELAY whenever a PTY is allocated; this change mirrors that on the ssh2 client right after the `ready` event in doSsh2Connect, covering both initial connect and auto-reconnect. Proxy-command / proxy-jump connections (where ssh2's underlying socket is a custom Duplex over a child-process pipe) are a no-op by design, gated by the public Client.setNoDelay()'s own type guard. A discriminating log line records which path each connect took. Tests cover initial connect and a full reconnect cycle to guard against the regression class "Nagle is re-enabled because someone refactored only the initial connect path." Co-authored-by: Orca <help@stably.ai> * fix(ssh): bound relay-lost reconnect with exponential backoff When the relay exec channel keeps dying (e.g. a remote-side bug closes every fresh --connect channel right after handshake, or a stale bridge keeps being replaced), the unguarded _onRelayLost handler reconnects as fast as the network allows — spawning relay deploy attempts in a tight loop until the user force-quits. Each iteration spawns a fresh ssh2 exec channel, hammers sshd's MaxSessions counter, and floods the renderer with state churn. Add per-target exponential backoff (500ms → 15s, capped at 6 attempts) so the loop terminates instead of running forever. After the cap the session goes to 'error' state with a 'Relay channel kept dropping. Please reconnect.' message — visible in the renderer instead of an invisible failure where typing in remote terminals just stops working. Successful 'ready' resets the attempt counter only if the session stabilized for >= 5s; faster flaps preserve the counter so a flaky remote backs off rather than retrying indefinitely on every brief ready→lost cycle. Backoff state is cleared on explicit disconnect, on session replacement during reconnect, and on connect failures, so a real reconnect attempt after backoff exhaustion always starts from zero. Co-authored-by: Orca <help@stably.ai> * fix(ssh): detect stale relay daemons via running-version marker The on-disk relay version check compares local .version against the remote .version file in the relay dir. A daemon launched by an earlier deploy keeps running its in-memory copy of the OLD relay code, so when the client later rewrites relay.js + .version on disk and bridges in via --connect, the new bridge process drives a stale daemon. Protocol or behavior changes between the two versions then tear down the channel in a tight reconnect loop (observed against PR #1672 on a daemon predating that change). The daemon now writes its running version into a .running-version sidecar at startup, anchored to the relay-script directory rather than process.cwd() so test spawns cannot pollute the repo root. Before attaching to an existing socket, the client probes that marker and, on mismatch with the locally-deployed .version, kills the stale daemon (TERM only, never KILL) and falls through to a fresh launch. Conservative defaults: when either marker is unreadable, attach so older builds keep their live PTYs. Co-authored-by: Orca <help@stably.ai> * Revert "fix(ssh): detect stale relay daemons via running-version marker" This reverts commit |
||
|
|
42e04268fb | fix: harden SSH connection reliability and add passphrase prompt (#636) | ||
|
|
cc66e120eb | feat: Add SSH remote support (beta) (#590) |