mirror of
https://github.com/stablyai/orca.git
synced 2026-09-30 16:02:56 +00:00
fix/execution-host-resolution
377
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
510305e574 |
fix(relay): signal capacity loss instead of dropping, hanging, or truncating (#17870)
Three failures with one shape: a payload past a fixed capacity was met with silence, with a wait that never ends, or with a prefix presented as a whole. **The workspace snapshot was silently dropped.** `workspace.changed` carries the tab/session list, and a snapshot past the producer frame capacity (12288 B on a Node <=21 remote) was dropped with only a relay stderr line, so the client kept a stale list forever. The relay now publishes per client and, for a client whose sink refused the frame, sends a compact `workspace.stale` marker on the control lane; the client re-reads through `workspace.get`, whose lane is budgeted in megabytes rather than in one producer frame. A new JSON-RPC notification rather than a new field on `workspace.changed`: `normalizeSnapshot(undefined, ns)` yields revision 0 and an empty session, so a Rule-1 field would make an old client replace its tab list with nothing — worse than the drop. An old client ignores the unknown method and is exactly where it is today. The marker retention/retry machinery is extracted from the `fs.changed` overflow path and shared by both. **The Windows upload hung, and the fix for it could truncate.** `#16432` was attributed to `[Console]::In.ReadToEnd()` materializing the base64 bundle. That is not what the reporter measured: he also measured `new IO.StreamReader([Console]::OpenStandardInput())` — an incremental reader — hanging at 1 MB. The limit is in the stdin the host hands PowerShell over a non-pty ssh exec, not in the string the script builds. - `uploadFileViaSystemSsh` — the user file-import path — was piping a whole file into one Windows stdin, unchunked and untimed. That is the path large files take; it now chunks into 32 KB writes and bounds each wait. - The Windows directory upload reuses that single-file path rather than repeating a weaker copy of chunk-read + write-buffer; the `ino`/`dev` TOCTOU verification comes with it. - A Windows write needing more than one exec lands on a `.orca-partial` staging path and is published by rename, so a failed chunk cannot leave a truncated artifact under the real name. `exclusive` is enforced once at the rename, not on the first chunk, where a retry met its own leftovers. - The mkdir batch reads stdin through the stream reader the reporter measured surviving 50 KB, not `[Console]::In`, which he measured wedging at that size. - `waitForChannelClose` takes an optional bound. A wedged PowerShell stays alive at idle CPU and never closes, so without one the promise is simply never settled and the caller waits forever with no error to show. **Quick Open showed a prefix as the whole workspace.** The mechanism "a full page means there is more" only works if the caller named the cap, and the failing UI named none — it hardcoded `truncated: false`. Quick Open now names `QUICK_OPEN_LISTING_MAX_RESULTS` on both the Electron IPC hop and the runtime-RPC hop (the field #17954 added to `files.listAll`), and reads a full page as truncation. The local hop honours the cap too, which it previously ignored. Rebase note on `fs.listFiles`: an earlier revision of this work also clamped the host unconditionally, and #17934 escalated an uncapped request to an explicit error. #17954 has since landed and made an oversized reply streamable, which removes the premise — the host no longer has to choose between a prefix and a refusal, so it returns the whole listing when no limit is named and only clamps a limit it was given. Keeping either would have regressed #17954 and hard-failed three in-tree callers that deliberately pass no options (`runtime-file-commands-search-runtime-files.ts:81`, `filesystem-read-handlers.ts:125`, `runtime-file-commands-constructor.ts:41`). |
||
|
|
39330c5aca |
fix(relay): retire PTYs the host proves are gone, and stop two per-poll scan storms (#17832)
* fix(relay): stop three CPU growth terms in a long-running remote session pty.resize gated only on `managed.disposed`, which is bookkeeping rather than liveness. A shell that exits without node-pty's `onExit` leaves an undisposed entry holding a closed master fd, and UnixTerminal.resize has no fd guard, so the ioctl threw `ioctl(2) failed, EBADF` into the dispatcher's generic parse-error catch. Nothing retired the entry, so it stayed advertised and kept activePtyCount above zero -- which is what stops a relay with an unlimited grace from reaching its idle-no-ptys exit (#12423). Probe liveness with the same helper attach/listProcesses use, retire a provably dead pid, and contain an ioctl failure over a live-or-unverifiable process. processHasChildren forked `pgrep -P` per pane per inspection poll, uncached. procps-ng opens six procfs files per process to resolve one ppid, so each call cost O(host process count). Answer from the TTL-cached `ps` table the same RPC already captured for the foreground lookup (#13537). The remote AI Vault scanner had no parse cache at all, so every forced rescan re-read and re-parsed the whole transcript corpus, including files untouched for a month. Give it the mtime+size keyed memo the local scanner has (#13753). * fix(pty): invalidate the descriptor when node-pty gives up the handle (#17930) Carried forward from PR #17930, which merged into this branch. Rebased onto current main; main's newer node-pty-fd-leak test is kept as-is. * fix(ai-vault): refresh codex titles on the remote parse-cache reuse path The remote cache keys on the transcript's (mtime, size, host), but codex titles live in $CODEX_HOME/session_index.jsonl and are written after the rollout — so a cache hit froze the fallback title forever. Mirrors the local scanner's existing reuse-path refresh via a shared core. * fix(relay): publish the exit a reap performs, and rescan for close decisions Two review findings on the CPU work. reapExitedPty told only the relay-internal exit listener, so a retirement left the client's pane mounted against a session the relay had already forgotten -- the next attach answered `PTY "<id>" not found` with nothing before it to explain why. Pre-existing on three probe paths; resize made it user-triggered. Publish the same pending-exit the natural onExit path publishes, carrying -1 ("gone, status unrecoverable"), and skip it when onExit already reported the real code. processHasChildren now answers from a 500ms TTL-cached table. That is right for pty.inspectProcess, which every tracked pane polls, but pty.hasChildProcesses gates the window-close confirmation and workspace cleanup's idle evidence -- one destructive decision per answer, where a child started inside the window would be killed unasked. Give that RPC a fresh scan; pgrep used to. * fix(relay): publish a reap's exit only on proven-exited evidence The publication is a verdict the client acts on by retiring the pane, so it must not be reachable from the disposed-record sweep, which retires off our own bookkeeping rather than the host's process table. Only ESRCH earns it. * fix(i18n): restore the activity-options key the rebase dropped * fix(i18n): union en.json with main so the rebase cannot drop keys |
||
|
|
7458c39181 |
fix(ssh): declare a wedged relay link lost, and stop reading silence as a verdict (#17817)
* fix(ssh): declare a wedged relay link lost instead of suppressing the dead-link check * fix(ssh): make the Windows deps probe exit 0 on a real load failure, like its POSIX twin * fix(relay): reap a client that has stopped answering instead of holding its leases forever * test(relay): feed the primary before asserting the reaper exemption holds * fix(ssh): keep a lost link's verdict unverifiable instead of reporting absence * refactor(ssh): read the exec timeout from its typed code, not the message text * fix(relay): bound a client that clears the handshake and then never frames anything * fix(i18n): restore the activity-options key the rebase dropped * fix(i18n): union en.json with main so the rebase cannot drop keys |
||
|
|
7104056984 |
fix(watcher): route relay watch-root capacity refusals off the fast ladder (#17950)
* fix(ssh): stop two unrecoverable relay refusal loops A relay refusal that is a pure function of state the client cannot change was being retried forever, on two different paths. - pty.openClient: a superseded owner proof is refuted evidence, not a transient fault. The client kept re-presenting the identical proof, so every reconnect reproduced the same refusal until the relay was redeployed (#12895, #12931). It is now dropped exactly as a stale lease already is, and the claim re-asked without it. - fs.watch: the relay's watch-root capacity refusal was classified 'unavailable' and retried at 1 Hz per root for 60s, re-armed indefinitely. A folder workspace with more repos than the cap turns that into a permanent install storm scaled by the excess root count (#11196). It is now its own 'capacity' result that goes straight to the existing dormant backoff, mirroring what the local watcher path already does. * fix(watcher): route relay watch-root capacity refusals off the fast ladder A full watch-root cap is a decision, not a fault, so a 1 Hz reinstall per refused root only bills the relay the load that keeps the cap busy (#11196). Capacity refusals now go straight to the dormant backoff. The relay side no longer refuses on a slot it is about to hand back: an over-cap caused by roots still unsubscribing waits once on the teardowns settling — the release event, mirroring WatcherSupervisorCapacityWait — before it answers. A parked waiter is excluded from the accounting so it cannot take a slot from the root already reclaiming one. Drops the SSH owner-recovery half of this branch. Its premise — that a -32043 SUPERSEDED refusal is permanent — is false: the refusal fires only while the incumbent is 'active', and assertPtyConsumerOwnerRecovery explicitly admits the identical lower-generation proof once the incumbent flips to 'disconnected' (relay-pty-consumer-owner-displacement.test.ts proves it). The remedy could not work either: the proofless re-ask routes into refuseHeldPtyConsumerOwner, which is declared `: never` and, with sameClient true by construction, always throws. It would have traded one refusal loop for another, minus the checkpoints and minus the proof that resumes the claim once the relay reaps the incumbent. * fix(i18n): restore the activity-options key the rebase dropped * fix(i18n): union en.json with main so the rebase cannot drop keys |
||
|
|
31007c0d86 |
fix(ssh): reclaim relay PTYs the client has provably lost, on host attestation only (#17831)
* fix(ssh): reclaim relay PTYs the host attests this client orphaned (#9819) Orca could lose track of terminals running on an SSH relay until the 50-slot cap refused to open any more. This reclaims them, and the whole design is built around the fact that getting it wrong destroys a user's running process on their remote machine: the failure mode is leak, never kill. A stop requires all nine of: 1. the relay published an `ownerClientInstanceId` read from the live authenticated consumer grant of the connection that requested the spawn — never from a spawn parameter, since an echoed claim is no evidence; absent means skip 2. that id equals this client's persisted consumer identity 3. this connection holds the negotiated `session-owner` grant 4. `paneBound === true`, host-published 5. no `agentSessionOwners` — the host still advertises it as adoptable 6. `hostAgeMs >= 30s`, measured on the host's clock 7. this client has no route: not reattached, no lease outside terminated/expired, no pending kill, and no `expired` lease either — an expired lease is the record of a process deliberately left running, never a licence to kill it 8. every stop is fenced on the incarnation the same listing published, and on the owner identity, both re-checked by the host 9. a pass wanting to stop more than 8 refuses entirely Absence from a client-side set is `unverifiable` by construction (docs/reference/ssh-execution-boundary.md): a second machine attaches to the same relay and displaces the session owner, and its live agents are missing from this client's store for exactly the reason a genuine orphan is. So the host has to attest ownership, and the host has to attest that nothing is running. That second attestation is measured over the pane's whole tty, not its foreground process group. `tpgid == pgid` is foreground-only: on a real `bash -i` on a real pty, a shell holding `sleep 300 &` and a shell holding a Ctrl-Z'd job both read `pgid == tpgid`, `Ss+` — byte-identical to an idle prompt, with only the job's own row differing. A foreground-only gate therefore attests `pnpm build &` and a suspended editor as idle, and the stop that follows SIGKILLs every process group on the tty. `shellOwnsEveryTtyProcessGroup` is measured over that same set of groups, so the evidence and the kill describe the same thing. No new probe: `tpgid` already identifies the terminal, because a process group belongs to one session and a session to at most one controlling terminal. The freshness field is real rather than decorative. `capturedAgeMs` is stamped from when the capture was taken, deliberately as an upper bound since the process table is TTL-shared, and the sweep refuses an observation older than its own pass budget, counting its own elapsed time since the listing arrived. Stale evidence degrades to "do not sweep", never to "sweep". The display consumer of the same measurement keeps no age budget, as a stated decision: a stale pane title costs a redraw and self-corrects. `pty.shutdown` is authorized on the host that owns the process. `pty.spawn` and `pty.attach` both take a request context and check it; the one irreversible call took none, so the rule above lived entirely on the client that decided to make the call. It gains an optional `expectedOwnerClientInstanceId` and refuses unless the connection still authenticates as that identity AND this host recorded it at spawn. Finally, a reattach refusal now says whether it observed the process. Three refusals carry the same `SSH_SESSION_EXPIRED` text and only one is absence; `restoreRequired` means the PTY is live and only its source stream is not. Testing that text with `.includes()` expired the lease and deleted ownership for a running process, erasing this client's only record of it — and a PTY with no record is one the sweep may stop. Wire compatibility: four new optional fields and one new optional param on existing methods, no new method and no new stream opcode (Rule 1, and Rule 2 does not apply). Rule 1's caveat is discharged explicitly — no reader requires any of them, each absence is a named skip reason, and an ordinary pane teardown must omit the owner fence because a revived PTY carries no attested owner at all. New client plus old relay stops zero PTYs; old client plus new relay never reads the fields. Windows relay hosts publish no evidence and therefore never sweep. Verified by joining the real publisher to the real client reader over `ps` captured verbatim from a Linux container, and by driving a real group-for-group SIGKILL against a real pty: backgrounded and suspended jobs survive by pid, and an idle shell is still reclaimed, so the narrowed predicate is not a silent no-op. Squashed deliberately. The sweep is unsafe at every intermediate commit of its own history — before the foreground gate it reaps a hand-launched `claude`, and with a foreground-only gate it reaps a backgrounded build — so this ships as one commit with no bisectable state that kills live work. Refs #9819. Folds in #17939. * fix(i18n): restore the activity-options key the rebase dropped * fix(i18n): union en.json with main so the rebase cannot drop keys |
||
|
|
0886db2b90 |
refactor(process-table): extract the correlation indexes into their own module (#18246)
`src/shared/process-table-snapshot.ts` is 308 code lines against the 300 cap for `**/*.ts`, so `static analysis` is red on `main` and every open PR inherits it. Neither PR that grew the file crossed the cap alone. #18151 took it to 427 raw lines; #18166 added ~35 more. #18166's branch predated #18151, so the head CI linted was 428 raw lines and passed, while the squash onto main is 463 -> 308 code lines. The gate lints the PR head, not the merge result, so nothing linted the sum until it was on main. Pure move, no behaviour change: the generic index machinery (ProcessIdentityRow, ProcessTableIndexOf, buildProcessTableIndex, collectDescendantsFromIndex, lookupProcessTableIndex, getProcessTableIndex and its WeakMap) moves to process-table-index.ts. `ProcessTableIndex` and `scoreForegroundCandidateRow` stay behind because they need `ProcessTableRow`, which keeps the new module free of any import back and so introduces no cycle. |
||
|
|
a0de2fde0b | fix(terminal): confirm an unrecognized foreground before downgrading agent prompts (#18238) | ||
|
|
104f9655e4 |
perf(git): answer remote-URL questions from one subprocess, not one per remote (#18158)
Four copies of the same loop ran `git remote` and then a serial `git remote get-url <name>` per remote to answer "which remote has this URL". On a repo with 58 remotes that is 59 subprocesses -- measured at 1083 ms -- for one question, and worktree create asks it several times. `git remote -v` answers for every remote from one child, reporting the same insteadOf-expanded first fetch URL `get-url` prints. The batched `cat-file --batch-check` branch-conflict probe decides from stdout, but its WSL route was unfenced, so a login-shell fallback printed the distro banner onto the stream it parses. That broke the one-line-per-ref contract, made every batch undecided, and fell straight back to one `show-ref` per remote -- the cost the batch exists to remove. Measured at 58 remotes / 4346 branches, spawns and wall time: push-target remote scan 59 -> 1 (1083 ms -> 8 ms) branch-conflict probe 60 -> 3 (984 ms -> 43 ms) configured push target 123 -> 6 (2707 ms -> 157 ms) |
||
|
|
1d94ebee3f |
fix(agents): stop a deeper vendor helper from stealing a pane's agent identity (#18062)
* fix(agents): keep outer agent identity over vendor helpers * fix(agents): preserve outer identity across relay scans --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
f737f3499f |
fix(relay): stream an oversized fs.listFiles reply instead of refusing it (#17954)
Opening Orca's own checkout over SSH cannot list its files in one response frame. 22,617 tracked paths average 58 characters, so the 20,001-row page the client asks for serializes to 1,223,415 bytes — past `DISPATCHER_CONTROL_QUEUE_MAX_BYTES`, so `sendResponse` demotes it to the `legacy-response` lane, where an unrelated producer backlog can refuse it as an opaque `ResponseOverCapacity`. Break-even is around 49 characters of average path; any `packages/<name>/src/...` monorepo is over the line. Picking a ceiling to refuse at does not fix that, it just moves where it shows up and refuses listings that would have been delivered. `__streamResponse` already exists for exactly this on the git methods, and it is its own negotiation in both directions: an old client never sends it and gets the plain array on the legacy-response lane as before, and an old relay ignores it and answers plainly, which the client detects by the sentinel marker being absent. So fs.listFiles opts into it — no new method, no new opcode, nothing to advertise — and the size of a listing stops being a correctness question. The response-stream registry becomes one per relay, shared by FsHandler and GitHandler. A second registry is not an option and the header of git-response-stream.ts says why: a client keys reassembly on `streamId` alone, so two would hand out the same id and cross-feed chunks, and only the handler that registers `git.responseAck` can credit the window a pump parks on. Also declares `maxResults` on the runtime-RPC `files.listAll` and forwards it. The mechanism "the client names its cap, so a full page reads as truncation" was wired only on the Electron IPC hop; web and mobile were saved incidentally by `remoteFileContentBudget` defaulting the cap inside `listRuntimeFiles`. A new optional field is additive in both directions (wire rule 1). The new Docker-gated spec is claimed by run-ssh-docker-e2e.mjs. The sharded e2e lanes set no ORCA_E2E_SSH_DOCKER, so a Docker-gated spec that no runner names self-skips everywhere and still reports green — pr-e2e-gate-contract enforces that. Closes #12547 |
||
|
|
99d9111653 |
fix(relay): fail an over-budget RPC response, not the connection (#17968)
The relay's control lane is a shared 1 MiB budget, and `sendResponse` admitted responses onto it with the fatal default: once the lane was full, admission closed the client. A ~900 KB `fs.listFiles` reply from a large remote workspace therefore took down the whole remote session -- every terminal on it -- rather than failing the one Quick Open request. The substitute `ResponseOverCapacity` frame already there only covered the `legacy-response` lane, because the fatal close beat it to the client. A JSON-RPC response is the droppable class of control frame: it carries an id, so one caller can be told and can retry. `pty.replay` and `notifyControl` keep the fatal default -- they are never re-sent, and a silent drop there desyncs the client with nothing to retry. Both response enqueues now pass `controlOverflow: 'reject'`, so the substitute error is what the caller sees; in the corner where even ~150 bytes will not fit, the caller's own 30s request timeout settles it and the session survives. Old clients are unaffected: they already decode this error code and message generically (`ssh-channel-multiplexer.handleResponse` rejects the pending promise with both), and the frame shape is unchanged. What changes is that a listing which used to drop the connection now returns an error on it. |
||
|
|
3d3b4f9053 |
fix(ssh): scope every activate() release path to the record its caller owns (#18038)
Two of the three release/cancel sites in RelayPtySourcePublication.activate() acted on `current` unconditionally. A superseded transport re-entering activate() therefore released — or cancelled and deleted — the delivery its own replacement had just opened: releasing the fence resumes a send the replacement is still rotating, and retiring it blanks the pane that owns it. The live path is the unadmitted/subscriber branch, so guarding only the first site leaves the defect exactly as it was; all three now act only on a record the caller still owns. Also give the restore-required token its own toast copy. It must not join UNREATTACHABLE_SESSION_SOURCES: that copy says "Open a new terminal", which here abandons a running agent on a PTY the relay has just proven alive (docs/reference/ssh-execution-boundary.md). |
||
|
|
b4ba3e97ff |
perf(worktree): defer fork-PR remote creation from create-time to first use (#17922)
* perf(worktree): defer fork-PR remote creation from create-time to first use Fork-PR review worktrees eagerly ran `git remote add` + `git fetch` for the contributor's fork (and pinned branch.<x>.remote) at create time, even for a read-only review. That grows remote count unboundedly with review volume and pays a network fetch nobody asked for yet. Defer prepareWorktreePushTarget(Ssh) and the --set-upstream-to configure step at create time (local + SSH, IPC + runtime create paths); persist the pushTarget metadata untouched. Materialize the remote on demand the first time push/pull/fetch/fast-forward actually needs it, via two shared functions (materializeWorktreePushTargetRemote(Ssh)) reused across the legacy IPC handlers and the RPC runtime sync commands. A cheap `remote get-url <name>` probe keeps steady-state calls down to one extra subprocess once materialized, instead of repeating the O(remotes) scan. Add repo-local `remote.<name>.orca-created` config provenance, written when the remote is added, so cleanup can recognize ownership of a remote that was lazily materialized (and therefore never round-tripped through the store's `remoteCreated` flag). Refs #17828 * perf(worktree): materialize a deferred fork-PR remote on terminal spawn An agent running raw git in a freshly opened fork-PR review terminal has no usable upstream until an Orca-driven sync happens -- "sync through Orca first" isn't available mid-task, and git pull/log @{u}.. hard-fail without one (verified against real git). Fire the same on-demand materialization used by push/pull/fetch/fast-forward from the single terminal-spawn resolver (resolveTerminalWorkspaceLaunchTarget), fire-and-forget, so a newly opened terminal gets a working upstream without blocking spawn. * fix(worktree): retest deferred fork-remote CI failures, fix SSH provenance-marker RPC Rewrites the 5 CI failures on the deferred fork-remote change (#17828) as evidence, not fixtures: the SSH relay-upgrade/rollback/sibling-ownership tests move to materializeWorktreePushTargetRemoteSsh, where that unchanged logic now actually runs (create defers it to first sync). While writing a stricter test that routes its mock exec through the relay's real validateGitExecArgs, found that the SSH provenance-marker write (`git config remote.<name>.orca-created true`) was unconditionally rejected by the relay's generic git.exec (it blocks all non-read-only config writes) -- a real bug that would break every SSH fork-remote materialization against a live relay. Fixes it with a narrow git.markRemoteOrcaCreated RPC, mirroring renameCurrentBranch, with a graceful no-op fallback for relays that predate it. * fix(worktree): scope post-#17887 test assertions past narrow-refspec config calls Rebasing onto #17887's narrow-refspec `remote add` broke two broad `['config']` call-filters into false positives/negatives, and the local materialize test still asserted the pre-#17887 wide `remote add`/fetch-refspec forms. * fix(worktree): restructure upstream restore, persist provenance, widen short-circuit refspec (#17828 review) - Move upstream restoration to the materializer level so it runs on both the remoteAlreadyMatchesUrl short-circuit and the full-prepare path, not just buried inside prepare*. - Persist {remoteCreated, remoteName} to the store on materialize so #17842's orphan sweep can see a lazily-created remote, including via desktop IPC, terminal-spawn, and the RPC host-callback paths. - Widen the refspec on the local short-circuit path too (SSH's bare `remote add` refspec gap remains a documented, pre-existing limitation). - Fetch the branch's tracking ref before restoring upstream when the short-circuit widens onto a *new* branch on an already-existing remote -- a bare refspec-config widen never itself imports anything, so `branch --set-upstream-to` was hard-failing for a sibling worktree's first materialize (found via a real-git fixture, not just mocked unit tests). Skipped when the ref already exists so the common repeat-call case stays a local-only probe with no network round-trip. * fix(worktree): merge duplicate shared/worktree/types import oxlint --deny-warnings flags the split import as no-duplicates; full pnpm lint was failing on it after the #17828 review restructuring. * fix(worktree): scope the deferred fetch timeout to fetch calls, retarget stale create-time assertions CI on the previous push failed 3 shards, all argument-shape mismatches: - worktrees-wsl-runtime-routing.test.ts: the "restructure upstream restore" commit wrapped every call `prepareWorktreePushTarget` makes (remote, remote add, config, fetch) with DEFERRED_PUSH_TARGET_FETCH_TIMEOUT_MS, not just the network fetch. Local git subprocesses never need a timeout; scope it to `args[0] === 'fetch'` only, matching the short-circuit path's existing pattern. Updated the test to expect the timeout on the fetch call specifically (point 5 legitimately adds it there), while every other call stays untimed. - worktrees-create-metadata-persistence.test.ts (2 tests): stale from before this session -- create no longer mints a fork remote at all (#17828 deferred that to first sync), so asserting `remote add`/`fetch`/`remoteCreated: true` at create time no longer matches reality. Retargeted both tests to assert the deferred contract (no remote add at create, pushTarget persisted unmaterialized); minting itself stays covered by worktree-remote-push-target-materialization.test.ts and worktree-push-target-setup.test.ts. Re-verified all 5 fixture points (mint upstream, store persistence, single-flight, short-circuit refspec widen + fetch-missing-ref for local and SSH, finite timeout) against a real git fixture after this fix -- all still pass. * fix(worktree): hook pty:spawn into deferred push-target materialization (#17828) triggerTerminalSpawnPushTargetMaterialization only fired for agent/background/ mobile terminals; the desktop GUI's own pty:spawn path (new tab, split, reattach) never materialized a deferred fork-PR remote before raw git commands could run there. Add a small wrapper that resolves the worktree's push target and owning repo from args.worktreeId via the store, and fire-and-forget delegates to the existing materializer, wired as the first statement of runPtyIpcSpawn. Degrades silently (optional chaining + catch) so a partial/fake Store in existing spawn tests can't turn this into a spawn-blocking throw. * test(worktree): retarget stale editor-remote-branch assertions for worktreeId threading runtime-git-sync-client's local-path fetch/pull/fastForward/push calls now forward context.worktreeId (needed by the main-process handlers to key deferred push-target materialization). Update the 17 call-site mocks across 15 tests in editor-remote-branch-actions.test.ts to expect worktreeId: 'wt-1', matching the already-correct source behavior -- no assertion was loosened. * fix(worktree): give a materialize joiner its own branch wiring The materialize single flight is keyed on the remote, but everything after the remote add is per-branch. A sibling worktree joining an in-flight mint for a different branch received the minter's target and skipped its own refspec widen, tracking-ref fetch, and upstream link, so its branch ended with no upstream at all. Wait for the remote, then run the per-branch work against the joiner's own target -- the same path the already-exists short-circuit takes, now shared rather than duplicated. Adopting a remote a sibling minted also stamps ownership, so removing the minter cannot strand the survivor's metadata outside the orphan sweep's reach. * fix(worktree): stop a failed mint from leaving a config-only fork remote Review of the joiner fix found it made things worse in three ways. Swallowing the mint's rejection let a joiner adopt a remote the rollback had already removed, writing remote.<name>.fetch with no URL. Verified on real git: that ghost section breaks `git fetch --all`, forces every later mint to a `-2` name, and cannot be removed by `git remote remove`. Propagate instead; the in-flight map is already cleared, so a retry re-mints. The SSH twin still returned the minter's target to a joiner, so the original per-branch bug survived there. It now adopts against its own target through a twin helper. The ownership stamp was unreachable: it required both a store and a repo id, and no caller passes both. Derive the repo id from the worktree id. Adopters also write remote config, and concurrent `git config --add` has no lock retry -- 135 of 160 writes failed at 8-way concurrency, and equal values duplicate the refspec. Chain adoptions per remote. |
||
|
|
b552bcb91f |
fix(relay): diagnose why node-pty will not load instead of hedging (#17891)
The relay could only say "terminals are unavailable" and then list three remedies for four different faults, none of which the user could verify (#17830). Two things were destroying the evidence: - `loadPtyUncached` caught the load error into bare `catch {}` blocks (pty-handler.ts:539, :551) and returned null. The only cause anyone had was discarded on the spot. - node-pty's own loader walks three directories and rethrows only the LAST failure, so even an uncaught error arrives as `Cannot find module '../prebuilds/...'` — the GLIBC/ABI/arch sentence is already gone. The relay now keeps the load error, recovers the real dlopen message with an out-of-process load of the file node-pty would have opened, reads what node-gyp configured the binding for (`build/config.gypi`), captures the host's Node ABI, arch and glibc, and probes the toolchain only when nothing was compiled. Each fault gets its own message naming values the user can check: toolchain_missing, dependency_missing, abi_mismatch, arch_mismatch, libc_floor, shared_library_missing, load_crashed, and load_failed which quotes the loader verbatim. A probe that did not answer stays `unverifiable` and prescribes nothing. The classification is now also structured data on the error, so a client can repair the host instead of printing a paragraph: an additive, schema-validated `data` field on an existing JSON-RPC error, with `repairable` true only for a proved fault that recompiling on the host actually fixes. Reuses orcad's loader-message parsers and out-of-process probe rather than adding a second copy; `classifyLoaderMessage` moves to a shared module and gains architecture and missing-shared-library cases, which the orcad boot precondition picks up too. |
||
|
|
058e618bb4 |
fix(ssh): stop a failed worktree scan from publishing authoritative emptiness (#17833)
* fix(ssh): keep an unreadable worktree catalog from authorizing teardown #14004: the relay's worktree-list fallback caught every failure and returned `[]`, so `SshGitProvider.listWorktrees` resolved as a success with an empty list. Downstream reconciliation treats a resolved listing as authoritative, which reaches `teardownMissingWorktreeTerminalsBestEffort` and the unregistered-worktree removal paths — a data-loss path from a failed scan. - relay: the `-z`-unsupported fallback lane propagates its failure instead of swallowing it to `[]`. - provider: an empty or malformed `git.listWorktrees` response is refused as `WorktreeCatalogUnavailableError`. A Git repo always lists its own checkout, so a zero-row listing can only be a scan that never answered — this is the mixed-version guard against relays that still swallow. - `listRepoWorktrees`: an unreachable SSH host reports unavailable instead of an empty catalog. #12661: `ssh:terminateSessions` now returns `{ terminated, unverifiable }`, so an offline sweep that only tore down local transport cannot be mistaken for a remote kill. The Manage-hosts toast warns instead of claiming success. * chore(i18n): register the unreachable-terminal terminate message |
||
|
|
406bd0e378 |
perf(relay): cache process-table descendant indexes (#17646)
* perf(relay): cache process-table descendant indexes * fix(relay): keep the process-table index first-wins and narrow Two defects in the memoized index this PR introduced. - Restore the first-wins duplicate-pid tie-break the relay had as `rows.find()`. A process whose argv contains a newline makes `ps` print a continuation line that the lenient parser can accept as a spurious row duplicating a real pid; that row always FOLLOWS the real one, so last-wins let it capture the pane's foreground. The rule now lives in `buildProcessTableIndex`, so the batched evidence resolver's `byPid.get(rootPid)` root lookup gets the same semantics the subsystem had before indexing. - Build only the two indexes a resolver reads. `byPgid`/`byTpgid` have no readers repo-wide, and delegating to a four-map build made a one-pane relay pay more per 500ms capture than the single `childrenByParent` map it replaced -- a regression in the majority topology, in a PR whose point is relay CPU. Matches the same deletion in #17763 line for line so whichever merges second resolves trivially. |
||
|
|
704167197a |
perf(relay): serve one ps capture per window and pin the batched inventory path (#17763)
Follow-up defect fixes for the batched PTY-inventory evidence path (#17525), now on main. - One memoized `ps` capture serves both the lenient and strict views. The two readers ran byte-identical argv behind separate caches, so a relay serving both forked `ps` twice per 500ms window — the doubling issue #6288 removed. - Drop the `byPgid`/`byTpgid` indexes no resolver reads, plus the zero-caller `parseProcessTableRowsStrict` and `getFreshStrictProcessTableSnapshot`; the batch resolver now reuses the shared index lookup and candidate score instead of private copies. - Restore `getForegroundProcessName`'s ladder contract: the extracted table scan answers null again, so an unconfirmed wrapper fallback publishes the recognized (normalized) name rather than node-pty's raw one. - Pin the SHIPPED `pty.listProcesses` path: one capture and one linear row pass for N panes, and node-pty's own name (never "shell") when the capture cannot disambiguate a `node`/`python` wrapper. - Pin the hidden-pane cadence gate in the production option shape, and move the strict-parser coverage next to the parser it tests. |
||
|
|
8ac1c6e2ac |
perf(git): bound ref and worktree scans (#17655)
* perf(git): bound ref and worktree scans * fix(repo-search): clamp oversized ref limits * fix(worktree): keep strict worktree listing unshared The shared-scan re-export flipped every `listWorktreesStrict` caller from an isolated subprocess to the coalesced scan. `git worktree prune` in the removal recovery path does not bump the scan generation, so a post-prune verification could join a pre-prune scan, see the stale row, and report a successful removal as a stale registration. The same gap defeats the post-archive-hook rechecks that exist to catch an external Git client locking the row. Restore the unshared export and make coalescing opt-in via `listWorktreesSharedStrict`, which existing callers already use deliberately. * fix(git): separate a proven absent ref from a failed probe `show-ref --verify --quiet` exits 1 for a missing ref, but so does `wsl.exe` when its own launch fails, so reading any exit 1 as absence collapsed `unverifiable` into `exited`. A genuine miss prints nothing while a wrapper failure always explains itself, so require empty stderr alongside the exit code; a runner that reports no stderr at all keeps its exit-code contract. That same signal removes a spawn regression: `show-ref` is a direct-git read under WSL, and the runner retried any numeric exit through the user's interactive login shell. The replaced `for-each-ref` exited 0 on a miss, so absence never retried; every absent probe now would. Treat a quiet exit 1 as Git control flow and skip the fallback. Also narrow the hosted-review suffix fallback: the replaced `refs/remotes/*/<base>` could not cross a slash, but `show-ref -- <base>` matches at any depth, so `origin/feature/main` answered a query for `main` and submitted a review against a base the provider rejects. Refresh the real-binary compatibility contract to the shipped excludes, and assert exact probe concurrency rather than an upper bound so a regression to serial probing fails. |
||
|
|
9477b5fcbb |
feat(ssh): batch process evidence in PTY inventory (#17525)
* feat(ssh): batch process evidence in PTY inventory * fix(ssh): accept Linux kernel process rows and make no-evidence polling push-driven * fix(ssh): preserve process evidence polling semantics --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
fbe94ceff6 |
fix: close readiness gaps found by merged-change audit (#17159)
* fix(ssh): fence stale kills and retired pane replay * fix(ssh): support cancellable interactive authentication * fix(ssh): await remote catalog before snapshot adoption * fix(pty): contain Windows ConPTY input failures * fix(power): avoid redundant macOS display blocking * perf(editor): narrow markdown override subscriptions * fix(quick-open): close directory handles after reads * refactor(linux): remove unused proc socket scanner * fix(usage): apply flat Sonnet 4.6 pricing * ci: prime Node next native test cache * docs(skills): resolve snapshot cleanup data path * fix(ssh): recover install locks after host reboot * test(ssh): recognize boot-aware install locks * test(ssh): prove previous-boot lock recovery live * test(wire): pin pre-metadata release coverage * fix(terminal): preserve remote tab ownership through recovery races * test(runtime): fence replaced terminal handles in agent guard * fix(ssh): preserve remote snapshot authority across polls * fix(pty): contain late ConPTY output EPIPE * test(pty): register Windows exit watcher before kill * fix: close SSH and tab readiness race gaps * fix(tabs): retain headless order and placeholder titles * fix(build): avoid parallel electron-vite config race * test(windows): avoid MSYS temp path rewriting * test(windows): avoid killing exited PTY * fix(pty): avoid late ConPTY input teardown race * fix(terminal): sync reconnect error ownership after commit * fix(runtime): use canonical worktree identity comparison * test(ssh): assert complete cold-hydration baseline * test(windows): invoke quoted retention fixture via PowerShell * test(windows): read ConPTY grid through mode con * fix(terminal): publish PTY replacements atomically * fix(terminal): infer stale identity on reattach * fix(terminal): fence stale pane PTY callbacks * fix(terminal): fence stale pane binds after rebind * fix(terminal): reject stale pane transport callbacks * fix(terminal): fence mirrored reattach spawn callbacks * fix(terminal): replace stale pane PTYs on remount * fix(ci): size the Windows launcher-compile test budget from measurement `native-smoke (windows-latest)` fails ~4.5% of runs on `preserves a multiline argument through the compiled remote launcher` with "Test timed out in 15000ms" — on unrelated PRs, for reasons that have nothing to do with them. Across 176 sampled attempts it is the only red that job produced, and it hit seven different PRs in two days: #16900, #16904, #16915, #16955 (twice), #16979, #17014, #17085. The test is six process creations: powershell.exe forks csc.exe, then the freshly compiled orca.exe forks node.exe, twice. Hosted Windows runners periodically slow process creation down, and this test amplifies that far harder than anything else in the job. Comparing the 80 attempts where it ran under 3s against the 12 where it ran over 12s, its own median goes 2198ms -> 15917ms (7.2x) while the same file's powershell-only test moves 556 -> 686ms (1.2x), the cmd.exe and Git Bash process tests in the neighbouring file move 1.4x, and the other 35 files put together move 1.5x. Measured across those 176 attempts: 1881ms to 35438ms, p50 4264ms, correlation +0.881 with the job's total Vitest duration. 8 of 176 (4.5%) exceeded the 15s cap; 2 of 176 (1.1%) also exceeded the shared 30s testTimeout, so deleting the override and inheriting the config is not enough on its own. 60s clears all 176 with 1.7x headroom on the worst. This is slow, not hung. Every body here is synchronous spawnSync, so Vitest cannot interrupt one — the timer fires only after the body returns and the reported duration is real elapsed time. That is why a failure reads `× ... 22464ms` under `Test timed out in 15000ms`. The work finished; the stopwatch was short. Seven reruns at one identical head measured 2053 / 4680 / 5551 / 8732 / 13506 / 14868 / 21937ms — the last of those would have been red on code that had not changed. The 15s came from #8897, which raised this test off Vitest's built-in 5s default because the job then ran bare `pnpm vitest run`. #8909 landed 3h27m later and pointed the job at config/vitest.config.ts, which is the real fix for that. The constant stayed behind and has been the binding budget ever since. * fix(terminal): fence stale remount reattach ownership * fix(terminal): reconcile mounted pane identity after replacement * fix(terminal): fence stale reattach fallback ownership * fix(terminal): fence deferred SSH reattach ownership * fix(terminal): fence stale split pane ownership callbacks * fix(terminal): keep stale spawns from consuming startup --------- Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com> |
||
|
|
75e5c996c1 |
perf(relay): stop ACK boundary scans at first pending boundary (#17491)
* perf(relay): stop ACK boundary scans at first pending boundary * test(relay): pin PTY source boundary cleanup and guard ascending sends The early-`break` in advanceCredit is only correct while sentBoundaries is inserted in ascending sentEndSu order. Turn that implicit invariant into a throw at the sole live write site (commitPtySourceSend), and assert the post-state directly instead of inferring it from an iteration budget: - assert the surviving boundary set after the 1,023-ACK benchmark - cover the jump-ahead cumulative ACK that must delete many boundaries in one pass (the case an over-eager `break` would get wrong) - cover the settleReservedPtySourceAck -> advanceCredit entry point - drop an arithmetically-implied assertion and CI benchmark log noise * perf(relay): reclaim ACK boundaries with a monotone cursor The early-break Set scan still rebuilt a Set iterator per ACK, so V8 walked delete tombstones and the drain stayed superlinear; the visit-count test could not see it because it stubbed sentBoundaries with a generator over a private Set. Replace the Set with an ascending boundary list plus a monotone cursor, assert the real structure, and add a benchmark over the shipped code. * test(relay): enforce ascending sent-boundary inserts in the collection Move the ascending-order precondition into PtySourceSentBoundaries.add so both insert sites are covered, and assert per-ACK span reclamation in the drain. * test(relay): collapse ledger test record accessors into getDeliveryRecord Rebase onto #17490 left two structurally identical internals accessors (getCursorRecord, getBoundaryRecord); one typed accessor covers both. |
||
|
|
ae2eeff55d |
perf(relay): index PTY source-credit send spans (#17490)
* perf(relay): index PTY source-credit send spans * perf(relay): maintain PTY source-credit retention totals * test(relay): pin PTY send-cursor rebase across ACK reclaim Cover the Math.max clamp branch in reclaimCreditedSpans where reclaim removes spans at or past the send cursor, and widen the seeded fuzz case to 20 spans per seed so the cursor actually traverses spans; assert the cursor never overshoots the span containing sentEndSu. * refactor(relay): drop dead retained-total helpers and pin retention counters The incremental PtySourceCreditRetention counters replaced the recompute-from-records helpers; delete the now-unreferenced exports and recompute the totals from the live records inside the ledger tests so the counters have an independent oracle. * test(relay): bound send-span reads instead of pinning the read pattern Address review feedback on the send-span cursor coverage: - replace the exact indexed-read pin and the tautological naive-visit assertion with a linear bound that still fails on the old Array.find path - drop the per-run bench console.log - assert retention totals immediately after rotate(), the only path that removes and re-adds a record in one call Also count the replacement delivery in retention as it enters the delivery map so the "in deliveries <=> counted" invariant never has a hole. |
||
|
|
b5746724d4 | perf(relay): account pending PTY output incrementally (#17639) | ||
|
|
1e4c56baa1 | perf(git): classify status line-stat inputs once (#17461) | ||
|
|
c9ef3aab5a | perf(filesystem): avoid per-entry directory promises (#17452) | ||
|
|
3d65466d99 |
perf(agent-hooks): coalesce Codex transcript poll timers
Coalesce per-pane Codex transcript polling onto a shared deadline scheduler while preserving cancellation and stale-callback fencing. |
||
|
|
1f20a53d22 |
Fix duplicate Codex startup command echo
Deliver local POSIX Codex startup commands through the shell wrapper at shell initialization, preventing duplicate PTY echo. |
||
|
|
a085c28e1b | fix(relay): propagate worktree listing failures | ||
|
|
f23d0b166f |
fix(relay): mint PTY ids that carry the relay incarnation instead of a restarting counter (#16901)
* fix(relay): scope PTY ids to mint epochs * test(relay): treat minted PTY ids as opaque * test(relay): pin mint-epoch id shape and restore spawn-sequence assertions The epoch escaping had no test: dropping encodeURIComponent left the whole relay suite green. Pin the three-field id shape against an epoch that carries both separators, and cover a colon-bearing relay id through the unchanged app-side SSH id wrapper. subprocess.test.ts had traded `pty-1`/`pty-2` for `expect.any(String)`, which discarded the invariant those two cases exist to prove: an early node-pty load failure burns no sequence, a late spawn failure burns one. * test(relay): mirror production epoch escaping in testPtyId The harness built the expected id without the encodeURIComponent production applies at the mint site. A test epoch carrying a reserved character would diverge silently across ~40 assertions in 11 files. |
||
|
|
b5a85890ac |
perf(git): bound git subprocess execution with an atomic admission scheduler (#16874)
* perf(git): bound git subprocess execution with an atomic admission scheduler Field traces (#16038, #11363) show Windows freeze storms driven by unbounded concurrent git children (12+ at once, 50-65s status convoys for 25+ minutes). Admit every main-process git child against atomic per-budget base+headroom counters (general / network / per-route), with reserved interactive capacity, ordering-only aging, close-bound permit release, a 120s fail-safe read timeout that feeds scheduler backoff, tier plumbing through every option carrier, and coalesced+jittered visibility pollers. Killswitch: ORCA_GIT_ADMISSION_DISABLED=1. Storm harness A/B: max concurrent children 65 -> 6, interactive p95 791ms -> 88ms; output-parity battery byte-identical with admission on vs off. * test(git): run the admission output-parity battery on every platform Parity needs real git, not the storm harness's PATH stub, so it must not share that file's POSIX gate - Windows is the platform where parity evidence matters. * fix(git): preserve interactive admission invariants * perf(git): keep admission queue drains linear * fix(git): close final admission gaps * perf(git): bound eligible route selection * fix(merge): remove unrelated stale snapshot changes * fix(git): preserve refresh lifecycle authority * test(git): align admission lifetime contracts * fix(git): harden admission across runtime paths * fix(git): restore freshness for bulk status reads * test(git): repoint delete-dialog source pins after admission plumbing The hydration effect now orders its targets through orderDeleteWorktreeStatusHydrationTargets and passes includeLineStats alongside the abort signal, so both literal anchors stopped matching. The invariants are unchanged and still pinned: dropping the signal, the main-worktree/folder filter, or getState-instead-of-subscribe each still reddens this test. * Fix git admission tier propagation and lock ordering Decode optional Git status tiers permissively and default runtime RPC status reads to the status lane while preserving renderer caller intent. Acquire the FETCH_HEAD mutex before atomic admission so same-repository fetch waiters hold no global or route permits. Preserve automatic pull-request refresh reasons, keep explicit hosted-review refreshes interactive, remove the dead candidate tier, and keep relay scheduling unchanged. Use tier-aware status lease keys because a shared lease cannot be safely promoted after its admission request is queued or granted. * test: align expectations with admission plumbing * refactor(child-process): move the process contract types to process-spec run-process.ts crossed its line cap after gaining the termination observer; the public types and defaults move out with re-exports so no caller changes. * chore: restore pnpm-lock.yaml to main (unintended local drift) --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
ba5f33402f |
fix(relay): reap owned PTYs when the daemon dies on an uncaught exception (STA-5697) (#16894)
* fix(relay): reap PTY jobs on fatal exit (STA-5697) * test(relay): cover the POSIX fatal reap and make a failed reap observable The fatal reap had no POSIX coverage at all -- every case forced win32 -- and the daemon discarded the rethrown reap error in an empty catch, so a reap that failed on a remote host left no trace in the only log a crash produces. Collapse the job-terminated branch onto the forceKillSent flag it already sets: the flag is what suppresses the redundant signal, so the separate "continue" was a second expression of one intent, and the two could only be caught together -- reverting either one alone left the suite green. |
||
|
|
91a35b9257 |
Split relay Git handler responsibilities (#17217)
* Split speech session lifecycle * Split terminal output scheduler pipeline * Split mobile browser pane modules * Prune resolved max-lines suppressions * Split pane tree equalization logic * Extract mobile troubleshoot screen styles * Split external automation manager * Split main window service attachments * Split hosted review creation checks * Split automation dispatch event handling * Split settings navigation metadata * Split daemon initialization lifecycle * Split GitLab item dialog * Split relay dispatcher layers * Split mobile host screen * Retarget mobile view settings source test * Split runtime file client layers * Split ports panel layers * Split runtime environments pane layers * Split local PTY provider responsibilities * Split CDP bridge responsibilities * Split relay Git handler responsibilities * Track moved relay Git fetch audit * Fix F3-speech for #17123 * Fix F1-cycle for #17131 * Fix F4-navtest for #17157 * Fix F2-allowlist for #17161 |
||
|
|
c6641152f1 |
Split relay dispatcher layers (#17174)
* Split speech session lifecycle * Split terminal output scheduler pipeline * Split mobile browser pane modules * Prune resolved max-lines suppressions * Split pane tree equalization logic * Extract mobile troubleshoot screen styles * Split external automation manager * Split main window service attachments * Split hosted review creation checks * Split automation dispatch event handling * Split settings navigation metadata * Split daemon initialization lifecycle * Split GitLab item dialog * Split relay dispatcher layers * Fix F3-speech for #17123 * Fix F1-cycle for #17131 * Fix F4-navtest for #17157 * Fix F2-allowlist for #17161 |
||
|
|
871601b6fe |
Prune resolved max-lines suppressions (#17142)
* Split speech session lifecycle * Split terminal output scheduler pipeline * Split mobile browser pane modules * Prune resolved max-lines suppressions * Fix F3-speech for #17123 * Fix F1-cycle for #17131 |
||
|
|
2dfaa676d8 | chore: update oxlint and oxfmt (#17150) | ||
|
|
94e7586665 |
fix(relay): stop advertising an agent session for a pane whose tab is gone (#12447) (#17012)
`pty.shutdown` was the only signal the relay ever got that a tab had closed, and it retired nothing: it requested a kill and returned. Every retirement path in the relay keys on proof of process death, so when the kill did not reap the pane shell the relay kept forwarding the orphaned agent's hook events as a live agent pane with no tab, and kept publishing its `agentSessionOwners` from `pty.listProcesses` with no liveness check at all. The relay now records the retirement the client stated, drops the pane's cached agent status with it, refuses to forward or replay posts from a retired pane, verifies the kill actually landed instead of assuming it, and reaps any PTY whose pid it can prove is gone before listing it. No wire change: no new RPC, field or stream opcode. |
||
|
|
fc8c981103 | fix(browser-preview): enforce canonical runtime grants (#16975) | ||
|
|
cc384c5a3d |
fix(agent-hooks): post posix payloads as json (#11292)
* fix(agent-hooks): post posix payloads as json * fix(agent-hooks): mark header merged envelopes * docs(agent-hooks): describe header merge envelope * fix(agent-hooks): encode posix metadata headers * test(agent-hooks): update WSL JSON hook assertions * fix(agent-hooks): negotiate raw JSON transport * fix(agent-hooks): preserve packed metadata in POSIX shells * test(agent-hooks): include hook envelope in relay boundary inventory --------- Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com> |
||
|
|
aaef5e8c9f |
fix(agent-hooks): deliver hook events that fire while Orca is restarting (STA-5329) (#16685)
* fix(agent-hooks): correct durable spool delivery * fix(agent-hooks): spool curl failures after retries * fix(agent-hooks): keep replay out of runtime observations * test(agent-hooks): pin managed hooks inert outside an Orca terminal * fix(agent-hooks): address review findings on the durable spool - claude: pass the literal source; options.agent does not exist (typecheck) - kimi: the windows-local ordering runs its guard pre-stdin and before the function exists, so it no longer spools there (printed command-not-found) - writer: require a readable endpoint file before creating a spool tree - antigravity: carry its out-of-band event name into the record and filter on it - drain: truncate only the bytes consumed, preserving concurrent appends and a torn trailing line * fix(agent-hooks): ignore spool events without pane attribution * fix(agent-hooks): make spool replay and appends robust * test(agent-hooks): type spool replay records * fix(agent-hooks): defer unterminated spool records * fix(agent-hooks): replay spool events through relays * fix(agent-hooks): preserve Codex prompt across child replay * fix(relay): keep startup alive when spool replay fails * fix(relay): simplify spool replay startup guard |
||
|
|
933345d347 |
Clarify upstream divergence stats for rebased branches (#16358)
* Clarify upstream divergence stats for rebased branches When a branch is rebased, it still tracks the pre-rebase upstream while comparing against the new base. Move upstream arrows to the head line to prevent them being confused with compare-base counts. * Show upstream divergence stats independent of compare base Measure HEAD against upstream regardless of compare-base state, so divergence indicators stay visible even when comparison is missing, loading, or failed. Also use cross-platform temp paths in tests. * Show commit counts against compare base, not upstream Upstream divergence (↑/↓ against tracking branch) was confusing for rebased branches — the counts appeared beside the base ref but measured against the upstream branch. Show only the compare base count instead, on the line that names it. * Report branch divergence in both directions Rebased branches are typically ahead AND behind their base; a single count hides this case. Use symmetric range with --left-right --count to capture both directions efficiently, then expose commitsBehind in the UI alongside commitsAhead. * Use semantic names for i18n keys and template variables Rename hash-based translation keys to descriptive identifiers and replace generic value0/value1 placeholders with semantic variable names like `count` and `ref`. Improves code maintainability and makes translation strings self-documenting. |
||
|
|
a9781a4118 |
STA-4150: client-hosted remote browser (consolidated) (#15448)
Co-authored-by: Jinwoo-H <jinwoo@stably.ai> |
||
|
|
b516300b8c | refactor agent hook listener modules (#16187) | ||
|
|
2960fe9193 | Split relay entrypoints into focused modules (#16151) | ||
|
|
ec4687c434 |
feat(agents): distinguish Claude background monitoring (takes over #14205) (#16201)
* feat(agents): distinguish Claude background monitoring Adds an optional `workingMode: 'monitoring'` discriminator for a Claude session whose lead turn finished but which still has background shell tasks or session crons registered. The wire state stays `working`, so older peers that never read the field keep rendering Working. (cherry picked from commit |
||
|
|
1921ba2250 |
fix(agent-status): clear the pane when a Claude compact finishes (STA-2915, STA-4613) (#15202)
* fix(agent-status): clear the pane when a Claude compact finishes (STA-2915, STA-4613) A manual /compact ends at an idle prompt without emitting Stop, so nothing in the compact window could ever clear the pane. A worktree that entered the compact `working` stayed `working` until the 30-minute stale sweep -- and the summarizer's start-less SubagentStop kept republishing the row, resetting that clock each time. The correlation added by #12332 was supposed to own this, but it could never run: PreCompact and PostCompact were never added to CLAUDE_EVENTS, so they were never registered with Claude. compactTrigger was always undefined, and the transition guard, the ownership cache, the relay wire field and the ingest branch were all unreachable. Five test files exercised the logic by injecting events past the registration boundary, so the suite stayed green over code that could not execute. Register PostCompact -- and deliberately NOT PreCompact. Measured on Claude Code 2.1.227, a successful manual compact emits PreCompact, a start-less SubagentStop, SessionStart(source=compact), then PostCompact; an ABORTED compact ("Not enough messages to compact") emits PreCompact ALONE. Mapping PreCompact to `working` would strand the pane on every aborted compact, which is the bug being fixed, so the abort guard is structural: Orca never subscribes to the pre-validation event. PostCompact carries its own trigger, so no anchor is needed to tell manual from auto and the correlation machinery is deleted rather than repaired. Manual becomes a `done` with sessionBoundary set -- a finished compact is a session-shaped boundary, not a completed turn, so completion notifications, unread counts and automation-run evidence stay out of it. Auto claims nothing: it runs inside a turn that resumes and emits its own Stop. The source-blind early return that dropped compact events for EVERY provider before its normalizer ran is narrowed to Claude, so it keeps failing closed on a malformed payload without pre-empting other providers. Ownership is kept where the deleted guard had it: a valid provider prompt id is required, a completion clears a row but never creates one (a retired pane must not be resurrected), and a hydrated row is matched on provider session only -- it carries the previous session's connectionId, and older rows carry no session at all, so a strict check would reject the restart case this fixes. A consumed prompt id keeps relay duplicates from refreshing the row. Mixed versions: no new wire field and no new opcode. An older relay normalizes with its own shipped mapping and forwards the event, so ingest drops `auto` envelopes and stamps the boundary on `manual` ones; its replay strips the trigger entirely, so payload state stands in for it while ownership is still enforced. The relay now caches a completion with its compact identity removed, so a client that was offline during the compact still receives the clearing row on reconnect. Tests go red before this change and green after: 6 of 12 in the new registration-gated suite and 5 of 8 in the relay/ingest suite. The harness delivers only events present in CLAUDE_EVENTS, so a fix that is never registered cannot pass -- the failure mode that let the original correlation ship unreachable. * test(agent-status): restate the compact reliability gate around the new invariant The gate pinned a test file this change deletes, so the manifest check failed. Repointing the path alone would have left the gate describing an invariant that no longer exists: it required a manual PostCompact to match its exact PreCompact generation, and PreCompact is no longer consumed at all. Restate it. The invariant is now that PreCompact never moves a pane, that only a manual PostCompact marks done and does so as a session boundary, that a completion clears an existing row but never creates one, and that a relay predating the contract has its automatic envelopes dropped and its trigger- stripped replays classified by payload state under the same ownership checks. Evidence runs are the real ones: the 105-test suite from this branch, and the Claude Code 2.1.227 PTY capture that measured PreCompact arriving alone on an aborted compact. * fix(agent-status): clear the restart-stuck pane a compact was meant to clear Review found the completion did not clear the pane STA-2915 actually reports, and that republishing it was a strict regression. - A manual completion now retires a subagent that exists only as a disk snapshot: a /compact only completes at an idle prompt, so a restored child is proof of nothing. Live evidence -- a child observed in this runtime, an unclassifiable running background task, a registered session cron -- still holds the pane. - A completion that cannot clear now publishes nothing instead of restating the row, which was stripping restoredUnconfirmed off a hydrated row and restarting the staleness clock for work the compact never observed. - The relay defers compact ownership to the client that owns pane identity, so a cold relay cache can no longer swallow the one event that clears a remote pane. - claudeConsumedCompactPromptIdByPaneKey joins all three pane-scoped teardown routes, and an auto compact no longer spends the pane's consumed-compact slot. - The promptless completion keeps the summarized turn's label with or without a trigger on the envelope. Tests: the two restart cases now deliver the completion while the hydrated row is still cached, so they exercise the restored-row branch instead of passing through the strict one; the triggerless working replay is asserted from a FINISHED pane so it can fail. Reverting the four source files turns 12 of 21 registration-gated and 8 of 12 relay/ingest tests red, and 18 of 18 targeted mutations are caught. * fix(agent-hooks): preserve compact identity across relay replay * docs(reliability): describe compact replay ownership |
||
|
|
7e76bb3aec |
Fix rebase race by fetching to private ref before rebasing (#15990)
* Fix rebase race by fetching to private ref before rebasing
`git pull --rebase` is vulnerable to concurrent fetches modifying remote-tracking refs during execution. Fetch to a temporary private ref (refs/orca/rebase/*) first, then rebase from that stable ref to avoid the race condition.
* Fix rebase race by fetching to private ref with timeout
Concurrent fetches can interfere with remote-tracking refs between
fetch and rebase. Use a unique private ref and 60-second timeout to
isolate each rebase operation and prevent hangs on stalled remotes.
Extract gitPullRebaseFromBase to a dedicated module.
* fix rebase race by fetching to private ref with timeouts
Concurrent fetches can replace FETCH_HEAD and remote-tracking refs between
fetch and rebase, causing the rebase to fail. Fetch to a temporary private
ref instead, use --no-write-fetch-head when available (Git 2.29+), and
serialize FETCH_HEAD access for older versions. Add process termination
barriers to ensure proper cleanup and extend timeouts for SSH operations.
* Fix rebase race by fetching to both private and tracking refs
Concurrent fetches between source and rebase can replace remote-tracking refs,
causing rebases to use stale bases. Now fetch to both a private ref and the
remote-tracking ref simultaneously, ensuring the tracking ref stays current.
Also improves process termination for WSL guests with process-group tracking,
fixes process-tree termination timeouts on POSIX, and serializes FETCH_HEAD
operations for linked worktrees through their shared Git directory.
* Add WSL setsid --wait probe and barrier termination timeout
Probe for `setsid --wait` support and fall back to unwrapped execution for BusyBox compatibility. Add a deadline for process termination barriers to prevent hanging when tree termination cannot be verified. Update tests for cross-platform compatibility.
* Add wsl-process-group-termination to WSL invocation allowlist
* Serialize per-worktree git mutations to fix rebase race
Introduce operation locking for each worktree to prevent concurrent
mutations (like rebase) from interfering with each other. Ensures
rebasing a linked worktree doesn't affect the source worktree state.
Add SIGKILL fallback if process termination barriers cannot verify
tree termination.
* Serialize pull and fastForward operations per-worktree
- Extract generic git operation lock to reuse locking pattern
- Refactor existing locks to use the generic implementation
- Apply per-worktree serialization to pull and fastForward to prevent races
* Route WSL group termination through runWslProcess
|
||
|
|
202d74a8a4 |
fix(git): enable Windows long paths for worktree creation (local, sparse, and SSH hosts) (#15866)
Co-authored-by: hwantage <hwantagexsw2@gmail.com> |
||
|
|
15fd723bc4 | fix(terminal): drop conda's orphaned CONDA_SHLVL sentinel (#15885) | ||
|
|
02bee48e1d |
Retry transient ripgrep spawn failures instead of demanding install (#15983)
* retry transient ripgrep spawn failures instead of missing binary errors Fork/exec pressure (EAGAIN, EMFILE, ENFILE, ENOMEM, ETXTBSY) should not trigger ripgrep-not-found guidance. Add bounded retries (max 2x) for transient spawn failures in Quick Open and file listing, respecting cancellation signals. Introduce RipgrepLaunchFailureError to distinguish fork/exec pressure from unavailable ripgrep installations. * Handle cancellation during transient spawn failure retry window When a query is cancelled after a transient ripgrep spawn failure but before the retry decision resumes, the cancellation must be reported to the caller rather than proceeding with a retry attempt. |
||
|
|
5651662494 |
fix(wsl): migrate 21 call sites onto the WSL runner (#15923)
* fix(wsl): migrate 21 call sites onto the runner, after five review rounds Rebased onto main now that the runner (#15903) has landed. 21 sites across 15 files move off ad-hoc `execFile('wsl.exe', ...)`. Allowlist 23 -> 16 on the WSL guard; 163 -> 152 on the W1 child_process guard, which moved as a consequence. Five review rounds, each finding real defects -- several introduced by the previous round's fixes: 1. Hooks ran user orca.yaml scripts under dash; probe failure fell back to the login shell, reintroducing the ~/.profile stall the runner exists to remove. 2. An unparseable probe was cached permanently, disabling every WSL feature on the distro; hooks regressed from "runs degraded" to "fails". 3. Exit 127 had no expiry; a starved 5s probe hard-failed the 10s scan behind it; a joiner burned its budget on someone else's probe. 4. The comment stripper blanked live code, so the windowsHide guard walked past a real unguarded spawn and reported the file clean; an ownership-probe timeout silently deselected the user's Claude account. 5. Verification of the guards themselves. The recurring finding -- a call answering "is this installed?" on a degraded PATH -- was eventually fixed structurally rather than per-caller: the runner refuses an unresolved guest PATH unless the caller opts in. Per-site vigilance was demonstrably not holding; 3 of 8 sites had already forgotten the analogous exit-code check. Remaining 16 files need a runner mode that does not exist: a long-lived streaming child (OAuth logins, hook relay), a synchronous caller, or a host-level flag like --status that the guest-command API cannot express. * fix(wsl): close round 5's P1s -- degrade where PATH was never needed Round 5 measured the guards by re-executing their algorithms standalone rather than reading them, and found four things. P1 -- four skill/plugin paths gained a hard dependency on the login-shell probe that they never had. They ran under a plain non-login `sh -c` on main, so a probe failure now breaks WSL skill discovery and install on exactly the distro the runner was built for: one with a slow `~/.profile`. Worse, the throw escapes before each site's own error mapping, so the UI gets a raw internal string. They degrade now, per the rule this branch already wrote down in `wsl-fish-history-cleanup.ts`. P1 -- Codex and Claude were asymmetric. Claude's five credential sites degrade; Codex's were strict, so adding a WSL Codex account failed where adding a Claude one succeeded. Three of the four are byte-equivalent to Claude sites, and their scripts read `$HOME`/`$WSL_DISTRO_NAME`, which wsl.exe supplies without a login shell. `assertWslCodexCliAvailable` stays strict on purpose -- that one really does answer "is this installed?" (#9725). P1 -- the ownership-probe timeout fix did not survive the rebase onto main. A timeout still returned "not owned", which the caller *persists*, clearing the user's account selection. P1 -- `blankStringContents` desynced on a nested template literal (`` `${`x`}` ``), leaving 116 lines of a child_process importer outside the ratchet, with 27 importers structurally at risk. Now tracks template depth. Regenerating against the fixed blanker: 70 -> 68 offenders. Also: the windowsHide vacuity check could not fail while the allowlist alone exceeded its bound -- the exact defect the sibling guard documents avoiding. It now names a file that definitely offends. * fix(wsl): close round 6 -- my blanker fix had traded a false positive for a miss Round 6 re-derived the guard's answer from a TypeScript AST instead of trusting the regex, and caught two things. P1 -- the nested-template fix I shipped in round 5 introduced a worse bug than the one it closed. Switching to "code mode" inside `${...}` without also resetting the quote at a newline meant an apostrophe in a regex literal -- `` `'${value.replace(/'/g, "'\\''")}'` `` , which is exactly the shellQuote shape all over this codebase -- inverted the lexer for the rest of the file. `claude-accounts/service.ts` went blind from line 96, hiding a REAL unguarded `spawn` at :1097: the WSL Claude managed-login path, which opens a console and steals foreground on Windows. Round 5 traded one false positive for one false negative and I did not notice, because the offender count went down. The blanker now resets non-backtick quotes at a newline (the rule stripComments already had) and tracks brace depth per interpolation. The spawn is fixed rather than allowlisted, and the count is 69 -- the number the AST predicted. P1 -- the ownership-timeout guard was dead code: it threw into its own `catch` three lines below, which returned null, which the caller persists as "not owned" and clears the user's account selection. Now a typed sentinel the catch rethrows. P2 -- `WslGuestEnvironmentUnavailableError` reached the UI verbatim from the CLI installer and the Codex availability check. Both mapped. Method note: I had been regenerating the allowlist with a Python transcription of the scanner, and the two drifted -- the same two-implementations problem this workstream keeps finding. The allowlist is now generated by running the shipped test with an empty list and taking what it reports. * fix(guards): stop patching the lexer -- make the scanner fail closed instead Round 7 proved my round-6 fix also did not work, by planting a plainly-named unguarded `spawn` in `claude-accounts/service.ts` and watching the guard pass 3/3. That is three consecutive attempts at an exact lexer, each shipping a desync that hid real calls, and each time the offender count went DOWN, which I read as progress. Round 6's diagnosis was wrong too: the culprit is the `templates` brace-depth stack, which nothing resets, not quote state. So stop trying to be exact. `blankStringContentsDesynced` reports when the lexer lost its bearings, and the guard treats that as an offender. Over-reporting is a nuisance; under-reporting is a false clean, and a false clean is what let a real console-flash spawn out of the ratchet twice. The allowlist goes 69 -> 82: the 13 extra are files whose scan cannot be trusted, now named rather than assumed fine. The planted violation is now caught. Also from round 7: - `SPAWN_CALL` missed promisified and renamed bindings, so `exec('where gemini')` (a real Windows cmd.exe spawn) and a detached `shell: true` in `cli/runtime/launch.ts` were invisible. Added execAsync/execFileAsync/ execFileCb/spawnDetached. - `BASHISM` matched `set -o pipefail` but not `set -euo pipefail`, which is the only spelling this tree uses -- so the check could not have caught the #14292 signature it exists for. Fixed, and it immediately flagged a file; that one turned out to be a comment, so the bashism scan now strips comments too. - The CLI installer error mapping my round-6 commit claimed was "both mapped" was never applied -- only the Codex side had been. Now actually mapped. * fix(guards): close the four holes round 8 found by planting violations Round 8 stopped reasoning about the guard and planted spawns into it. Four holes, none of which reading had found: - `windowsHide: false` **passed**. The check was `args.includes('windowsHide')`, a substring test. Now matches `windowsHide: true`. - A ternary first argument was silently skipped: the method-declaration filter `/^\(\s*\w+\s*[:?]/` also matches `exec(useAlt ? 'a' : 'b', …)`. Now requires a type after the colon. - Renamed bindings were not covered, despite the comment I wrote saying they were -- I had hardcoded three names. Aliases are now resolved from the import. Each is verified closed by planting it and watching the guard fail. `fork` is deliberately still unscanned. Round 8 is right that Node forwards the option, but `ForkOptions` does not declare it, so the two live sites cannot be fixed without a cast. Recorded in the verification doc rather than left as a silent gap, along with two others worth knowing: the allowlist is file-granular, so its ~18 false-positive entries carry a standing pre-approval for real regressions in those files and cannot be retired by fixing code; and `stripComments` has no desync report, so the fail-closed check is only half applied. The doc now also says how to verify a guard change: plant a violation. Every guard fix here that was verified by reading was wrong. * fix(wsl): stop preflight reporting installed CLIs as absent on a slow distro Round 9's merge blocker, and the sharpest finding of the whole workstream: the branch built to close #9725 had reopened it from the other side. `preflight-wsl-command.ts` was one of five sites without `allowDegradedEnvironment`, so a guest-PATH probe failure threw. Every consumer collapses a throw into a verdict: `isCommandAvailable` and `isCommandOnPath` catch to `false` ("not installed"), `isGhAuthenticated` and `isGlabAuthenticated` read an empty payload as "not authenticated". So a slow distro made WSL git, gh and glab read as missing. Two things made it likely rather than theoretical. The probe took two thirds of a 5s budget, leaving the command ~1667ms where main gave it the full 5s inside its own login shell -- a cold WSL VM start routinely lands in that band. And a probe timeout is cached for 30s with a re-probe threshold of 1.5x the failed budget, which a 5s caller can never clear, so every preflight command short-circuited without spawning wsl.exe at all -- and Re-check does not invalidate the cache. Fixes: preflight degrades instead of refusing, and the probe is capped at half the caller's budget and at 4s, so no caller ends up with less time than it had before the runner existed. Also fixes a real console flash found on the way: `preflight-command-exec.ts` spawns git/gh/node through `promisify(execFile)` with no `windowsHide`. Round 9 also confirmed the credential paths are now *safer* than main: all 11 account sites degrade, every destructive guest operation is still marker-gated, and main's `getOwnedManagedAuthPath` could disown an account on a 5s timeout -- which this branch turns into a failed launch instead of a destroyed selection. * fix(wsl): make "Try again" able to succeed, and test the round-9 fix Round 10 returned MERGE with one residual worth closing first. A transient probe failure left the null-resolving promise in `inFlight`, so the only way back was `retryAfter` -- and the 4s probe cap made the 1.5x budget escape unreachable, because no caller can pass more than 4s. For the full 30s window the four non-degrading sites returned their error *without spawning wsl.exe at all*, and each of those errors says "Try again". The advice was guaranteed to fail. The entry is now dropped on a transient outcome and an explicit cooldown gate replaces it, so the window alone decides. The window drops 30s -> 5s: long enough to stop a stampede, short enough that the user's next click reaches a distro that has since warmed up. Round 10 also noted the round-9 fix shipped untested, which was fair. Added: the probe-budget floor for 5s/8s/10s callers, and preflight's degrade opt-in plus its stdout/stderr-carrying rejection, which isGhAuthenticated reads off the caught error as an auth-success fallback. * test(wsl): make the probe-budget guard actually guard Round 11 caught that the regression test I added for the probe cap did not bind: it seeded the guest environment, so the probe resolved in ~0ms and the assertion read the command leg's timeout instead. Reverting the cap to the old 2/3 split left all three cases green. Dropping the seed and asserting on the probe leg fixes it -- verified by reverting the cap and watching all three fail. A regression guard that cannot fail is the shape that has cost the most in this workstream: the windowsHide guard silently passed a real unguarded spawn twice for the same reason. |