mirror of
https://github.com/stablyai/orca.git
synced 2026-09-30 00:03:15 +00:00
orchestration-dispatch-error-codes
378
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
4efc86a33c |
feat(app): open Markdown files from the OS in the floating workspace (#17906)
* feat(app): open Markdown files from the OS in the floating workspace Registers Orca as a Markdown handler on macOS, Windows and Linux, and opens an OS-handed .md/.markdown/.mdx file as a floating-workspace editor tab — the one editor surface that needs no project. Works cold-start and when Orca is already running. Main buffers the paths and both pushes to a live renderer and answers a pull on renderer mount, mirroring SkillShareDeepLinkState. The buffer is only released once delivery is possible: the renderer's pull is what proves its ui:openMarkdownFiles listener is attached, because a push into a window whose renderer has not subscribed is dropped by Electron with no error. Both the push and the pull restore an undelivered batch, and a renderer reload clears the latch so the fresh renderer re-proves itself. Paths are stat'd and proven to be files before authorizeExternalPath sees them. Windows association is registered by hand in the NSIS include rather than through electron-builder's `fileAssociations`: app-builder-lib emits APP_ASSOCIATE, whose first line overwrites Software\Classes\.md's default value with no backup — silently taking .md from whichever editor owns it, for every existing user on their next update — and APP_UNASSOCIATE never restores it. The hand-rolled registration is additive (ProgID + OpenWithProgids + SupportedTypes) and leaves the user's default alone; verified end to end on a real Windows 11 host. Co-authored-by: Wooseong Kim <innocarpe@users.noreply.github.com> Co-authored-by: Jaydev <java-jaydev@users.noreply.github.com> Closes #10138 * fix(os-open): register the new listener in the IPC inventory, and guard a non-array payload CI caught two things the local run did not. useIpcEvents-lifecycle.test.ts is an inventory of every App-lifetime IPC listener and the exact order they register in; ui.onOpenMarkdownFiles now appears there, positioned after the workspace-shortcut bridge's last listener, which is where it actually registers. Chasing that failure surfaced a real gap: the pending-open payload crosses the preload boundary, so a stale or mismatched preload can resolve with something that is not an array, and reading .length off it threw inside the promise chain instead of failing at the boundary. Array.isArray now gates it, with a regression test. |
||
|
|
c31e7b0d9a | docs(main): restore rationale comments lost in the startup and cookie splits | ||
|
|
5b4e7edb50 |
refactor(main): split backend services and startup
(cherry picked from commit
|
||
|
|
28373fcea7 |
fix(windows): repair the install-dir package ACL that blanks the window (#17740)
* fix(windows): repair the install-dir package ACL that blanks the window An install tree carrying an orphan AppContainer ACE (S-1-15-2-<x>) with no ALL RESTRICTED APPLICATION PACKAGES grant denies Chromium's LPAC children read on the shipped modules; they die at init with 0x80000003 and the window stays blank forever (electron/electron#51761). - Tighten the probe verdict to require the S-1-15-2-2 grant specifically: an ALL APPLICATION PACKAGES (S-1-15-2-1) ACE, the Program Files default, does not appear in an LPAC token and cannot satisfy the orphan. - Drop BUILTIN from the English-locale heuristic (fr-FR/es-ES print it verbatim) so a localized icacls is correctly reported as un-name-checkable. - Add an additive, marker-guarded icacls self-repair: an inheritable root grant plus a flagless (RX) /T pass, never /grant:r. - Route the crash-loop dialog through a testable prompt module that names the permission cause, offers Copy Commands without dismissing itself, and keeps the graphics-driver hint. The repair only runs on win32, off serve mode, and only on the exact probe verdict that reproduced the crash. * docs(windows): correct the install-tree ACL walk cost model * fix(windows): keep the install-ACL poison gate at the reproduced shape An orphan package ACE alongside the Program Files ALL APPLICATION PACKAGES default launches clean on win32 10.0.26200 / Electron 43.4.1, so requiring S-1-15-2-2 specifically declared poison on healthy installs - and this branch acts on that verdict with a tree-wide icacls write and the crash-recovery dialog's primary cause. hasRestrictedPackageGrant stays reported for triage; only the verdict reverts. |
||
|
|
4ac8a8912c |
fix(startup): install the app environment with the userData decision (#17755)
src/main/index.ts decided where userData lives at module scope, then installed the AppEnvironment port ~180 lines later inside the single-instance-lock block. Every statement in that gap was a latent failure: a path resolve there either threw 'AppEnvironment not initialized' and killed the process, or — with the accessor installed but the decision not yet run — would have memoized the pre-override directory in getCanonicalUserDataPath() for the whole session. The first outcome shipped. #16761/#16698/#17509 were one statement landing in that gap and killing every macOS `orca serve` across 1.4.190-1.4.192; #16762 moved that call but left the gap. Install the port and capture the canonical path immediately after the two calls that decide them, so the window is zero rather than small. Both are inert at this point — ElectronAppEnvironment holds no state and calls `app` lazily per accessor, and initDataPath only joins strings — so nothing that depended on the old position moves with them. The secret store stays where its pre-ready Keychain note applies. The throw is kept and still covers the case it should: resolving a path before the decision has run. Guarded by a source-level assertion that the decision, the install and the capture stay adjacent. Fixes #17750 |
||
|
|
8f15f217a2 |
Preserve user-set workspace names across branch changes (#17448)
* fix(worktrees): preserve user workspace names across branch changes * test(worktrees): cover pinned rename metadata * fix(workspaces): address display-name review edge cases * fix(workspaces): keep automatic names fresh across refreshes * fix(workspaces): preserve legacy CLI labels * fix(workspaces): preserve display-name provenance across hosts * fix(workspaces): honor legacy display-name provenance * fix(workspaces): fence display-name refresh races * fix(workspaces): accept peer renames from provenance-less hosts The old-host preserve fence kept a pinned local label on every refresh, which also suppressed a legitimate rename another client persisted through the same host until app restart. Narrow it to labels the host re-derived itself (branch short name, or path basename when detached); any other changed label in a mode-less response is explicit meta a peer wrote there. Stale prior-label responses stay covered by the downstream staleness fence, in-flight writes by the pending fence. * refactor(workspaces): unify display-name pin derivation Three call sites (renderer optimistic update, local IPC updateMeta handler, remote worktree.set handler) each restated the same formula; a future edit to one would silently skew provenance between paths. |
||
|
|
20a12a6a46 |
perf(codex): share one launch-prep hook install across a spawn burst (#17669)
* perf(codex): share one launch-prep hook install across a spawn burst Codex launch prep runs a full managed-hook install on every local PTY spawn, and both install lanes serialize globally per Codex home. Opening a multi-pane worktree therefore paid N full installs back to back, and a resumed Codex pane prepares twice. Concurrent spawns for the same runtime home now share one run; the promise is dropped as soon as it settles, so the next launch still re-reads hooks.json and the user's trust state. Also split the `host_env` spawn-timing phase, which spanned the entire Codex preamble and pinned that cost on the env builder that ran last. * refactor(codex): unify the two hook-install single-flight lanes Both the WSL and launch-prep lanes now share one generic in-flight helper instead of duplicating the map bookkeeping. Also routes the WSL launch-prep install through the serialized variant, which closes the same per-spawn serialization gap on WSL that the native lane just got. * refactor: extract the shared in-flight run dedupe The codex hook service and the GitHub conflict-summary cache had grown near-identical private copies of the same single-flight helper. Both now use one module, which also keeps the hook service clear of the 300-line budget. The shared copy keeps the identity check on clear so a late settle cannot evict a newer entry for the same key. |
||
|
|
18e8fe4770 |
fix(serve): install supervisor disconnect quit after app environment init (#16762)
Moves installServeSupervisorDisconnectQuit(isServeMode) out of module scope in src/main/index.ts to just after setAppEnvironment() and initDataPath(). The call resolves the serve update handoff path through getCanonicalUserDataPath(), which throws by design until the app environment accessor is installed. At module scope that throw was unconditional on macOS whenever the CLI set ORCA_SERVE_UPDATE_HANDOFF_PATH — which it does by default — so every `orca serve` process died at startup before it could listen, and the supervising service manager restarted it into the same crash. Reported in #16761, #16698 and #17509; shipped in 1.4.190 through 1.4.192. Guards added so it cannot drift back: a source-level ordering assertion that also pins the call synchronous and inside the single-instance block, and a runtime test that keeps the real path resolver, since the existing suite mocks it and therefore could never have caught this. Fixes #16761 Fixes #16698 Fixes #17509 |
||
|
|
5ff1aa540e |
fix(codex): re-land WSL direct-home cutover with counsel findings fixed (#16854)
* fix(codex): safely re-land WSL direct homes * fix(codex): finish WSL direct-home cutover * fix(codex): coalesce WSL launch hook installs * perf(codex): avoid duplicate retired WSL session scan * fix(codex): retain canonical WSL retired-home path * fix(codex): fail closed before retiring WSL auth * fix(codex): reopen WSL drain after rollback * fix(codex): preserve WSL source on unknown panes * fix(codex): harden repeated WSL runtime drains * perf(codex): bound pending WSL session scans * fix(codex): recover invalid WSL session watermarks * fix(codex): validate retained WSL scan state * fix(codex): accept durable WSL scan state * test(codex): cover the drain's inode-identity guard against destination replacement Removing the four `target_auth -ef temporary_destination_auth` assertions left all 33 apply-script tests passing, so a regression deleting them would have shipped silently. Reproduced before writing this. A hash check cannot catch the case. The pinned hard link keeps the original inode, so it still hashes correctly after another writer atomically renames a different file over the destination path; only inode identity sees it. Without the guard the script exits 0 and retires the source, leaving the user holding bytes nothing validated. The new case asserts the source survives. The harness is split by responsibility so no file exceeds its max-lines budget: fixtures, the coreutils interference shims, the run types, the apply runner, and the recovery/absent runners. The atomic-rename hook is deliberately separate from the in-place rewrite shim because different guards catch them. * fix(codex): keep the split drain harness inside the child-process boundaries Extracting the harness into non-test modules moved it out of the exemptions the single test file had: three new files import child_process, and two spawned without windowsHide. Adds the three to the import allowlist, and sets windowsHide on the spawns rather than exempting them - the flag is correct for these calls regardless of the ratchet, and they are skipped on win32 anyway. --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
e54cfc1901 |
fix(omp): read Pi/OMP static state-title markers and retire stale spinners (#14602)
* fix(omp): read Pi/OMP static state-title markers and retire stale spinners OMP 17.2.12 replaced its animated braille title frames with static markers on WSL/ConPTY (`π : working`, `π > idle`, `π ! needs input`). Orca read all three as idle, so a working OMP pane lost its status, and a synthetic title spinner started by an earlier hook kept rotating after its status row was gone. Classify the markers from one shared table so a later upstream punctuation change is a row, not a reparse, and stop the spinner when the hook row it stands in for is cleared or dismissed. Fixes #13890 * test(omp): preserve static state titles during normalization --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
d9870c6c75 |
fix(browser): apply the app-wide HTTP proxy to embedded browser sessions (#15536)
* fix(browser): apply the app-wide HTTP proxy to embedded browser sessions The proxy setting was only ever written to `session.defaultSession`, but browser guests run on their own `persist:orca-*` partitions. Any host reachable only via the configured proxy failed to load in an embedded tab, landing on `chrome-error://chromewebdata/`, while the same setting worked everywhere else. Adds a per-session applier alongside the existing defaultSession path, keyed by a WeakMap so one session's applied config can't suppress another's, and applies it to every browser partition through the single installer they all pass through. Startup awaits an explicit sweep so the first guest navigation can't race the installer's fire-and-forget write, and a settings change re-sweeps so toggling the proxy takes effect without a restart. Env-var fallback and the system-proxy probe mirror the defaultSession behaviour, so a browser partition resolves the proxy the same way the rest of the app does. Fixes STA-4779 * fix(browser): await per-session proxy readiness * fix(proxy): preserve loopback and authenticate * fix(proxy): settle browser partition update races * fix(proxy): close partition policy races * fix(proxy): order settings and release removed sessions * test(browser): await partition proxy readiness * fix(proxy): cancel removed partition retries * refactor(proxy): keep OpenCode rate limits out of scope * fix(proxy): preserve sessionless host policy * fix(proxy): gate requests on policy readiness * fix(proxy): retire deleted browser sessions * fix(proxy): close retired browser guests * fix(proxy): retain retired session guards * fix(proxy): retain retired partition policies * fix(browser): retry transient proxy application failures * fix(browser): release deleted partition installer state * fix(proxy): retry delayed transient failures * fix(proxy): preserve route session authority after rebase * fix(proxy): clear retired session credentials * fix(proxy): retire failed browser profiles * fix(proxy): harden failed session cleanup --------- Co-authored-by: Jinwoo-H <jinwoo0825@gmail.com> |
||
|
|
5ea9daba97 | fix(window): keep automated Electron launches out of the foreground (#17347) | ||
|
|
3457acb647 |
fix(memory): hydrate retained PTYs before diagnostics (#17308)
Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
fb6c2800ee |
fix(gpu): capture hardware identity in crash reports (#16973)
* fix(gpu): capture hardware identity in crash reports * fix(gpu): order bounded crash diagnostics before fallback * fix(gpu): keep fallback persistence ahead of diagnostics |
||
|
|
c5591d0893 |
fix(gpu): persist the safe-graphics marker before the restart prompt (#16945)
* wip: gpu-startup-recovery * fix(gpu): offer hardware retry after safe recovery |
||
|
|
fd9125ea8c |
feat(native-chat): Codex structured native chat restructure (#16729)
* feat(native-chat): port structured Codex sessions from restructure-recovery Rebuilds the desktop structured native-chat implementation from brennanb2025/native-chat-restructure-recovery (tip 4e31c08db3) on top of current main as a single commit, scoped to the local Codex path. Ported: - Structured agent-session core: durable record store + single-writer lease, canonical journal, agent-session wire host/attach/eviction/subscribers, `agentSession.*` RPC surface (registered via ALL_RPC_METHODS; host-side mobile allowlist included for wire compat), pty write gate, transcript additions, and the Codex app-server adapter/launch resolution. - Renderer: NativeChatStructuredSession view/composer stack, structured launch path with the single-flight guard, local structured session tabs sync, activation gate + structured inventory (read-only `agentSession.handoffStatus` probe), agent-session tabs in the tab strip, AI-vault structured session activation, and the settings pane with the parent Experimental Chat UI toggle plus the nested "Use updated structured native chat" toggle. New sessions require both flags, agent codex, no prompt, and a local non-WSL, non-Windows-host execution host (structured-native-chat-availability). - Fixes 72c013cea6 (verified Codex launch recovery), 8ddbaf5e3d (defer native terminal view switching affordances), and 4e31c08db3 (release the launch gate after a visibility retry) with their regression tests, including the third-launch-after-retry guard case. - Cross-version agent-session wire test + CI lane, packaging entries (proper-lockfile, agent-tooling asar excludes), and the wire-compat doc section. Deliberately not ported: mobile/ changes, the Claude structured runtime (only the claude-transcript-branch-proof and claude-structured-owner-identity leaf modules remain, backing the kept TUI-recovery arms), the terminal↔chat adoption/handoff flow (`agentSession.adoptTerminal`/`requestHandoff`, the handoff request engine, TUI adoption machinery, orca-runtime adoption methods), renderer switching affordances and their dead leftovers, the hook/subagent-status refactor cluster, and unrelated branch changes. The crash-during-acquisition recovery path (restart handoff adjudication, restore/reverse re-acquire, lease schema handoff keys) is kept because every plain direct launch depends on it; a trimmed handoff coordinator exposes only status/restore/close. Branch edits that targeted files main has since split (ipc/pty.ts, worktrees.ts, rpc/methods/terminal.ts, useIpcEvents, pty-connection, store/slices/terminals.ts, runtime-types, web preload) were re-applied to the split modules, preserving main's newer logic (Windows CIM fallback, browser tab close rework, cold-restore resume flow, dispatcher threading). Known seam: the mobile clipboard image-provenance CONSUMER gate ships (agentSession.send refuses unproven mobile image refs with agent_session_image_untrusted) but the producer hunk in rpc/methods/clipboard.ts stays with the unported mobile cluster, so mobile image sends into structured chat fail closed until that side ports. * fix(native-chat): trust only authenticated local image uploads * fix(build): preserve Windows process-tree patch application * test(windows): include process creation time in addon fixture * fix(build): run windows-process-tree node-gyp from the physical package dir gyp expands the node-addon-api dependency by probing node, whose cwd resolves to the package's physical directory in the store, so the emitted target is a store-relative ../../../../node-addon-api@... hop. gyp then resolves that hop against the rebuild cwd; from the node_modules symlink/junction it escapes the store and configure fails with "node_addon_api.gyp not found" (run 32999886072). Rebuild from realpath(package dir) so both bases agree, matching how the package manager itself runs native install scripts. The regression test replays gyp's expansion+resolution against the planned cwd and fails without the fix. * fix(native-chat): keep chat tabs visible through terminal closes and empty-worktree launches Two proven blockers in the native Codex tab contract: closeTerminalTab pre-empted the canonical unified close. With one terminal left it deactivated the worktree on a terminal/editor/browser-only check, blanking a workspace that still held a renderable agent-session tab; with two or more it pre-picked a successor from terminal entities only, re-stamping the group active before closeUnifiedTab's MRU/neighbor repair could land on the chat tab. Successor choice now defers to the unified contract whenever the terminal has a unified row, and deactivation is gated on the unified renderable count (matching leaveWorktreeIfEmpty), with the legacy pre-pick kept only for terminals without a unified row. A structured session created on an empty worktree was published into the host's headless group while preserveLocalLayout froze the local layout, leaving the tab in store but permanently off screen. A preserveLocalLayout owner now always takes client-owned placement — repairing a rendered leaf whose group record is missing, or materializing a rendered group on a truly empty worktree — and applies the client-derived layout repair while still rejecting host-authored layout. Regression tests drive the real store through closeTerminalTab (git worktree and folder workspace) and the real snapshot applier for the empty-worktree adoption states; all fail without the fixes. * fix(native-chat): close stale turns and retry rejected sends * fix(native-chat): retire hosted rows on structured tab activation * fix(native-chat): preserve rpc defaults across main merge * chore: format remote wire compatibility guide * test(native-chat): cover retry after unconfirmed send * fix(native-chat): reload outbox on session switch * docs(settings): disclose structured chat platform limits * fix(native-chat): await Codex launch-home preparation * fix(codex): align child-process allowlist with async trust bridge * test(identity): update inventory for tab surface refactor * fix(windows): preserve process-tree CRLF patch sources * fix(native-chat): anchor an unmatched chat echo where it was sent (#16117) * fix(native-chat): anchor an unmatched chat echo where it was sent The reported symptom was old user messages replaying below every new turn, so the conversation read as scrambled. The cause was not that the echo failed to match a transcript row. Claude consumes a mid-turn send through a `queued_command` attachment and writes no `type:"user"` record for it, so some echoes can never match, and no amount of matching will change that. The cause was WHERE an unmatched echo rendered: buildMobileNativeChatTransientData appended every pending item after the entire transcript, so it re-read below each turn that landed afterwards. Render each echo directly after the transcript row it was sent against, using the baseline the send already captures. An unmatched echo is then at worst a duplicate in the right position rather than a scrambled one, and it stays visible. Echoes sharing an anchor keep send order; a send with no baseline, or one whose anchor folding dropped, still falls back to the tail. Deliberately NOT fixed by deleting the echo. Inferring from send ordering that an echo can never match, then removing it, loses the user's own text for a message the agent did receive, and it cannot fire in the common case anyway - measured drain groups are 1,017 of size 1 against 55 larger. It also escalates an existing gap: the count pass has no baseline-tail guard, unlike the glue pass, while `messages` is a 40-row window that head-trims, resets on reconnect and grows at the front on loadEarlier, so a false landing there would license deleting a DIFFERENT outstanding message. That count-pass gap is real and left for a separate change; anchoring makes its worst case a duplicate in place rather than a scrambled conversation. * fix(native-chat): preserve folded echo anchors * fix(native-chat): preserve forward-folded echo anchors * fix(native-chat): keep leading folded echoes in place * fix(workspace-cleanup): show git status for every row (#16690) * fix(native-chat): refuse structured chat on every Windows execution path canUseStructuredNativeChat only refused win32 when a project runtime resolved, so folder-workspace keys (and other keys with no project runtime) failed open into structured chat on Windows. Fail closed on win32 unconditionally after the host check, matching the settings copy: local macOS/Linux only; Windows/WSL/SSH stay on terminal chat. * fix(native-chat): restore runtime refusals behind the win32 gate |
||
|
|
ca0a6ec9be |
Stop the OS keyring probe from gating the first window on Linux (#16912)
* Stop the OS keyring probe from gating the first window on Linux 1.4.190 added an at-rest secret protection report and called it from the `app.whenReady()` startup path, before the first window is created. `describeProtectionGap()` asks Electron `safeStorage` whether the OS keyring is usable, and on Linux that is a blocking D-Bus round trip to `org.freedesktop.secrets`. A keyring that is present but locked with no unlock prompter never answers, so the call sits until D-Bus times it out and the app shows no window for over a minute. Measured on Ubuntu 24.04 against a Secret Service that accepts the connection and never replies, time to first window: 1.4.188 1.06s (never contacts the keyring) 1.4.190 76.05s 1.4.190 --password-store=basic 1.06s (probe bypassed) Nothing on the startup path consumes the report, so it now waits for the first window's `ready-to-show`, with a timer fallback because that event can fail to fire when the GPU cannot present and headless serve has no window at all. Same build under the same hanging keyring: first window 80.48s -> 5.31s, with the report still delivered. STA-5765 * Pin the deferral the keyring-probe test exists to protect The suite passed with the probe fired on browser-window-created instead of ready-to-show, with setImmediate dropped, and with the fallback stretched to 10 minutes — every one of which reintroduces the STA-5765 stall. Drain the queue before asserting and bracket the fallback so those mutations fail. Also reformats the file to oxfmt. * Report the keyring gap inline in headless serve Deferring the probe to the first window is right for the desktop app, but serve never opens one, so the fallback timer became its only path. That moved the stall to after `printServeReady`: the runtime advertises itself, a relay or mobile client pairs, and only then does the main thread freeze on the keyring — stalling pings and PTY pumps, which a client reads as a dead host. Serve now reports inline, which is the timing it already had, and blocks before anything is advertised rather than under a live client. STA-5765 * test(secrets): pin the once-guard against a late window reveal The fallback can report first and the window reveal arrive after it; without the guard that probes the keyring a second time, blocking the main thread just as the user starts interacting. No existing case covered that order — removing the guard left all six tests green. * fix(secrets): keep the deferred protection report non-fatal, and pin the wiring Deferring the report moved it off `whenReady`'s promise chain. A throw there was an unhandled rejection the app survives; inside `setImmediate` it is an uncaught exception, and `installUncaughtPipeErrorGuard` re-throws those fatally — so a diagnostic the module documents as deliberately not fatal could kill the app. Wrap the deferred call so it degrades to a warn. Serve keeps the inline posture. Nothing outside index.ts referenced the scheduler, so reverting the call site, or flipping `deferUntilFirstWindow`, left the whole suite green — including the headless-serve regression an earlier review already caught once. Pin the wiring as source text, following the host-port-bootstrap-wiring idiom with every anchor bounded, and pin the two module gates that were only jointly covered. * test(secrets): make the deferral wiring pin resist an inert call site Round 2 of the review gamed the pin it had just added. Both anchors were bounded against -1 but not against overshoot, and the marker matched anywhere in the file — so nesting the call in a block, prefixing it with a guard, or commenting it out all left three green tests standing over a call that never runs. Bound the slice length, and anchor the marker to a statement at whenReady's own indent. Commenting the call out, wrapping it in `if (...) schedule(...)` with or without a block, and flipping the flag each redden now; previously only the unused-import typecheck error caught the first. |
||
|
|
2c86d2a3bd |
fix(agent-hooks): stop test runs and secondary profiles deleting the user's agent hooks (STA-5679) (#16980)
* fix(agent-hooks): stop startup from deleting another instance's managed hooks (STA-5679) Startup reconciliation removed the managed agent hooks whenever THIS profile had the agent-status-hooks off switch set. The hook files it removes are user-global (~/.claude/settings.json, ~/.cursor/hooks.json), so a second Orca profile with the switch off deleted the hooks every other running instance depends on. Cursor is the only agent with no title-derived status fallback: its native title is deliberately parsed as status-less, so a hookless Cursor pane is floored at 'idle' rather than showing a spinner. A global hook wipe therefore surfaces as "Cursor loading status missing from the sidebar" while Claude and Codex still paint status from their own titles, which is why this reads as a Cursor-only bug. Codex is unaffected either way because its hooks live in an Orca-owned runtime home. Honoring the off switch only requires skipping the install; removal stays on the explicit Settings toggle, which is the user-initiated path that should own it. Regression from #2778, which restored the destructive startup branch. * fix(cli-tests): stop the deferral suite deleting the developer's real agent hooks runtime-client-deferral.test.ts runs the REAL `main()` and feeds it `agent hooks off`. It mocks only ./runtime/environments and ./runtime-client, so the production handler ran end to end: updateEnabledOnDisk() wrote its state file and applyAgentStatusHooksEnabled(false) called removeManagedAgentHooks() against the developer's OWN ~/.claude/settings.json and ~/.cursor/hooks.json. A green test run therefore deleted every Orca-managed hook on the machine. Agent status then stopped reporting until the next Orca restart reinstalled them — silently, because the hook POSTs still return 204 and Cursor has no title-derived status fallback at all. The byte-for-byte equivalence twin already refuses these exact tokens, commented "MUTATING — writes outside ORCA_USER_DATA_PATH (`agent hooks off` parks the real ~/.claude hooks)". The vitest twin never got that guard. Stub the hook-controls module rather than dropping the row: `agent hooks off` is the only case in the table that reads ctx.client, so it carries the null-vs-undefined coverage the other four cannot. All 23 tests still pass, and a sandboxed HOME now keeps its hooks (5 -> 5) where it previously lost them (5 -> 0). * fix(cli-tests): ratchet agent hook deferral safety * fix(agent-hooks): keep startup reconciliation install-only |
||
|
|
b19a397d3e |
feat(browser-preview): reland remote HTML document previews (STA-5758) (#16920)
Reapply the reverted remote HTML document preview implementation so remote workspace files render locally over the orca-preview scheme. |
||
|
|
551fbb9ac7 |
Revert "feat(browser-preview): render remote HTML docs locally over an orca-preview scheme (STA-5557) (#16679)"
This reverts commit
|
||
|
|
249d93bc5d | feat(browser-preview): render remote HTML docs locally over an orca-preview scheme (STA-5557) (#16679) | ||
|
|
f400f8fd5f |
fix(macos): opt out of press-and-hold so held keys repeat (#14746) (#15589)
* fix(macos): opt out of press-and-hold so held keys repeat (#14746)
macOS routes press-and-hold to the accent picker unless an app sets
ApplePressAndHoldEnabled=false for its own bundle, so holding j in vim
inserted one character instead of repeating. Orca never set it.
Written at most once, and never over an explicit value: `defaults read`
is domain-scoped and exits 1 when the key is absent, which is the only
way to tell "unset" from a deliberate false — Electron's
systemPreferences.getUserDefault reports false for both. A recorded
decision in userData keeps a later launch from re-clobbering a user who
deletes the key to get the accent picker back.
* docs(macos): record the revert hazard and CI's macOS test gap
Two things a reader of this module cannot otherwise know.
A revert leaves the key written in every user's domain forever. AppKit reads
the plist, not this file, so removing the code alone keeps press-and-hold
disabled for everyone who ran an affected build. The sibling period-substitution
module carries the same warning because that fix was already lost once this way.
And the real-binary test file that pins the defaults(1) exit-code semantics this
design rests on never runs in CI: the e2e workflow and both unit-test jobs are
ubuntu and windows, and the only macOS runners in the repo are build and
packaging jobs that run no tests. Those six tests plus the real-bundle e2e case
pass on a developer Mac and execute zero times in a green PR, so the comment
should not imply enforcement that is not there.
Refs #14746
* feat(macos): let users turn the accent menu back on (#14746)
Orca disables press-and-hold for its own preferences domain so held keys
repeat. That is the right default, but the way back was a `defaults write`
buried in a source comment: nothing in docs/ or the README mentioned it, and
the preference is per-application, so it silently takes the accent picker
away from the Markdown editor and every other text field too.
Terminal -> Advanced now carries a "Character Accent Menu" switch, macOS and
desktop only. A web client cannot write a macOS preference for the machine the
user is looking at, so the control and its search-index entry are both gated on
that, not on the client's platform alone.
Precedence, which is the part that is easy to get wrong: the setting is
`undefined` until the user touches it, which is what keeps a hand-run `defaults
write` in charge for everyone who never opens the toggle. Once used, Orca owns
the key and writes exactly what the switch asks for -- `ApplePressAndHoldEnabled`
*is* the accent-menu switch, so it maps straight through with no inversion. The
choice is compared against `appliedSetting` in the existing decision record
rather than against the domain, so a `defaults write` made *after* using the
toggle is still the newer choice and survives the next launch. Re-asserting the
value every launch would have reintroduced the clobbering the record exists to
prevent.
The write lands for the next launch, since AppKit reads the preference as the
process starts, so the toggle shows the same restart banner the window-blur
setting uses. That banner is now a shared component, keeping its original
translation keys.
docs/reference/macos-press-and-hold.md records the precedence rules, the
`defaults read` rationale, the revert hazard, and the fact that none of this
executes in CI: every macOS job builds or packages and runs no tests, so the
real-binary and e2e coverage here passes only on a developer Mac.
* docs(macos): stop asserting when AppKit re-reads the press-and-hold key
Five places stated "AppKit reads the preference as the process starts" as
fact. That is the reason given for requiring a relaunch, and it is not
something this change ever measured.
Evidence points the other way: terminal emulators that register this key
after their process has started get key repeat in that same launch, which a
read-once-at-startup model cannot explain.
The relaunch requirement itself still looks right, but for a different and
verifiable reason: the write goes out through a separate `defaults` process,
so this app's own cached copy need not observe it. That is what the comments
now say, with the AppKit question left open rather than answered.
Refs #14746
* docs(macos): correct the startup comment's launch-timing claim
The comment said this call site is "the last point that can still matter for
this launch", which contradicts the rest of the module: the write is assumed
to land for the next launch because it goes out through a separate `defaults`
process. Reported on the PR by @innocarpe, who also supplied the replacement
wording.
Co-authored-by: innocarpe <innocarpe@users.noreply.github.com>
* refactor(macos): probe press-and-hold through the shared spawn chokepoint
`src/shared/child-process/child-process-import-boundary.test.ts` forbids a
direct `node:child_process` import outside its allowlist, and the allowlist only
shrinks — so this module moves to `runProcessSync`, which exists for callers
that genuinely cannot await. This one runs before `app.whenReady()`.
`runProcessSync` returns a non-zero exit instead of throwing it, so the
three-way read decision is re-expressed against `ProcessResult`: exit 0 is an
explicit value, exit 1 is a missing key, and a timeout, a signal kill, any other
exit, or a child that never started all stay 'unknown'. The throw path is now
inside `interpretDefaultsRead` so a spawn failure is reachable from a test
rather than hidden in an untested catch, and the write checks the exit code —
a refused `defaults write` no longer looks like success.
Both boundary-test failures were the same import: with it gone the offender
count returns to 155, so no ratchet baseline is bumped.
* Revert "feat(macos): let users turn the accent menu back on (#14746)"
This reverts commit
|
||
|
|
88ecc739e3 |
fix(browser): bound the agent-browser daemon's life instead of hoping teardown runs (#16367) (#16588)
* fix(browser): bound the agent-browser daemon lifetime (#16367) `agent-browser` is a client/daemon CLI. Orca only ever spawns the short-lived client; that client forks a daemon Orca holds no handle on, which reparents to pid 1 immediately. Nothing in Orca reclaimed it, so a crashed or SIGKILL'd run left one daemon per browser tab alive forever — two of them at ~25.6 GiB and ~7.2 GiB RSS saturated a 64 GiB cgroup under headless `orca serve`. Three fixes, in order of how much they cover: 1. Set `AGENT_BROWSER_IDLE_TIMEOUT_MS` on both spawn paths (the bundled-binary bridge and the orcad external-Chromium provider). This is the only bound that survives every way Orca can die, including SIGKILL, where no teardown code ever runs. 10 minutes: >6x the bridge's 90s `EXEC_TIMEOUT_MS`, so it can never cut a command, a retry chain, or an ordinary gap between two user commands, while capping an abandoned daemon at minutes instead of days. Verified against agent-browser 0.27.0: an idle daemon exits and takes its Chromium tree and socket sidecar files with it, and a daemon attached over `--cdp` (the bridge's case) leaves the attached browser running, so an Orca tab is never closed by its daemon idling out. 2. Await `destroyAllSessions()` in the will-quit teardown barrier. It was fire-and-forget and the only browser member missing from `settleTeardownWithinDeadline`; each session's close is its own agent-browser child taking hundreds of ms, so `app.quit()` won. 3. Sweep daemons a previous run left behind, using agent-browser's own `session list` / `close` rather than a pid walk (see `windows-pty-job.ts` for why walking your own orphans is guesswork). `closeStaleAgentBrowserSession` only ever reset the one name a new tab was about to reuse. The sweep runs only when `AGENT_BROWSER_SOCKET_DIR` is set, because that private per-profile directory is what proves the enumeration can only see this Orca profile's daemons; it is never set on Windows, so Windows gets no enumeration rather than a machine-wide sweep that could close a daemon Orca does not own. Windows stays bounded by the idle timeout, which needs no ownership proof. orcad's session name is stable across runs, so it closes that one name at start instead — a killed orcad's daemon would otherwise be reused while still holding the previous run's Chromium on a dead serve port. Where the 25 GiB went is inference from code, not a measurement: `captureStart` sets `activeCapture` and only an explicit `captureStop` ends it, so a HAR capture in a daemon living for days is unbounded. Not claimed as proven; the idle bound caps it either way. Not re-landed: the queue bounds from #10179 (reverted by #10255) bound Orca's own main-process heap, not the daemon's RSS, so they do not address this report. Also true but left alone: the 3-strike breaker's `destroySession` is an unawaited call whose `close` is `catch {}`-swallowed, and it only fires while a command is in flight — an idle-but-bloated daemon is never noticed. The idle timeout now bounds that case. `getOffscreenBrowserBackend()?.destroyAll?.()` is declared `void` and fully synchronous, so unlike `destroyAllSessions` it has no promise to lose and needs no barrier entry. * fix(browser): scope the daemon idle bound and close every daemon Orca owns Review follow-ups on the agent-browser orphan fix. - Never idle-bound the orcad external-Chromium daemon: it owns the user's remote browser, so the 10-minute bound closed a live session and every tab in it. Per-tab helper daemons keep the bound; the stable session name plus the `close` in start() is what reclaims a killed orcad's Chromium tree. - Retire a page's daemon from the headless offscreen backend, which is the only place `orca serve` closes a page and never reached the bridge. Credit to @Jinwoo-H (#16564) for identifying this owner-boundary gap. - Bound the teardown close at 5s so the will-quit barrier member cannot inherit the 90s exec timeout, and close sessions still being created. - Gate the startup sweep on a socket directory Orca derived itself; an inherited AGENT_BROWSER_SOCKET_DIR is no proof of per-profile ownership. - Replay a session's network routes when the daemon idled out between two commands, instead of silently serving unstubbed requests. * fix(browser): give the orphan sweep a kill switch Of the three behaviours this PR adds, two are already recoverable in the field without a build: the idle bound is an env passthrough an operator can raise, and the quit close is bounded by its own timeout inside the teardown deadline. The startup sweep was the exception — it fires unconditionally, and if it closes a daemon it should not, or spawns one process per stale name on a profile holding hundreds, the only remedy was a revert. ORCA_DISABLE_AGENT_BROWSER_SWEEP=1 turns it off, matching the existing ORCA_DISABLE_CODEX_TRUST_RPC / ORCA_DISABLE_HTTP2 convention. Note for anyone reaching for it on macOS: a Finder-launched Orca does not see shell env, so it needs launchctl setenv or a terminal launch. * fix(orcad): reuse a surviving browser session instead of closing the user's start() closed the daemon before every open, killing the Chromium tree with it. That runs on every provider start, not just after a crash — so an `orca serve` restart took the remote user's browser and every tab in it. The justification was borrowed from the pane bridge, which passes --cdp and so really does hold a port that dies with its Orca. This session passes only --session and --profile: nothing binds it to the old process, and the daemon owns its Chromium independently. A survivor is reusable as-is. start() now probes for an active tab first and returns it untouched. Only a name that answers nothing gets closed and reopened — which is still the killed-orcad case the stable session name exists to recover. This matters more as orcad becomes the backend the remote host runs on: the browser it manages belongs to a user, not to the process that happens to be driving it this minute. It is also the same principle that already exempts this path from AGENT_BROWSER_IDLE_TIMEOUT_MS. |
||
|
|
26721bd632 |
fix(codex): stop blocking the main thread on trust grants (#16441) (#16594)
* fix(codex): stop blocking the main thread on trust grants (#16441) Codex hook trust was granted by blocking the Electron main thread on `spawnSync` of a bundled ELECTRON_RUN_AS_NODE entry for the whole app-server deadline: 15s native, 35s WSL, ~45s on the real-home path (rebase inspect + repair + grant). Cold start and every Codex pane launch showed "Not Responding"; the reported event-loop gap was 15,049 ms. The subprocess only ever existed to donate an event loop to a deliberately blocked parent — `runCodexHookTrustGrantSession` was already the real async implementation. Make the callers async and the fork is unnecessary, so the bridge, the forked entry and its envelope are deleted along with their build/knip/tsconfig registrations. The CLI `agent hooks prepare-codex` handler is already async, so it awaits the in-process session and saves a process spawn per managed-home shell. `resolveCodexTrustGrantHost` is async too; the WSL identity probe moves from `execFileSync` to `runProcess`, dropping that file from the child-process import allowlist. Status reads keep a synchronous native-only stamp path. Two invariants that held only because the lane blocked: - Overlapping capability probes were impossible by construction. `GitCapabilityCache`'s dedupe engine is extracted to a shared `CapabilityProbeCache` and `CodexAppServerCapabilityCache` now inherits it, so concurrent launches against a cold host share one app-server session instead of one each. - Two grants on one `config.toml` could not interleave capture and restore. A reentrant per-file lane now serializes the whole install sequence (managed, WSL runtime, real-home ensure, legacy sweep) and the grant and rebase inside it. Cold-start work moves off the critical path: retained-home reconciliation (N sequential sessions) is fire-and-forget behind the daemon provider, and the startup real-home ensure chains into managed hook reconciliation instead of blocking app init. Every preserved semantic is unchanged: never throws, the ORCA_DISABLE_CODEX_TRUST_RPC kill switch, ledger hits, backfill-pending and cooldown fallbacks, config rollback on every failure path, pre-grant self-computed trust removal, the verify-failure taxonomy, diagnostics and telemetry. * fix(codex): widen the trust-config lane to every config.toml writer Review follow-ups on #16441's async trust grant: - `markCodexProjectTrusted` now runs inside the runtime+system config.toml lanes, so a project-trust write can no longer land inside a hook grant's capture->restore window and be silently reverted. Its callers await it. - `install`/`refreshRuntimeUserHooks`/`remove` hold the system config.toml lane as well as the runtime one — they promote approvals into ~/.codex/config.toml and mirror it back. Lock order is runtime-before-system everywhere. - The real-home ensure chain resumes after a rejection instead of returning the same rejected promise to every later pane launch, and resolving the real home is now inside the module's never-throws boundary. - `buildSpawnEnv` awaits inside a cancelable pending-spawn registration, so shutdown during the (now long) env build stops the PTY from launching. `prepareLocalPtySpawn` generalizes into `awaitCancelableLocalPtySpawn`. - CapabilityProbeCache drops the test-only `nowMs` passthrough; its probe backstop comment now describes what it actually guards. - Preflight is a plain async function; the trust dispatch in orca-runtime collapses into one `markWorkspaceTrustedForAgent`. * test(codex): exercise the trust-config lane under real concurrency The async grant makes two pane launches overlap for the first time. These drive the real modules end to end on real files: a rollback swallowing a sibling's grant, a markCodexProjectTrusted write landing inside a capture -> restore window, shared capability-probe dedupe on a cold host, the host-scoped transient cooldown, and reentrancy from inside an installer. Each was verified to fail against a deliberately broken implementation (lane removed, dedupe disabled, cooldown made global, reentrancy pass- through disabled). * test(codex): stop hook-service suites spawning the developer's real codex The forked grant bundle never existed under vitest, so the RPC lane was unreachable in tests on main. Running it in-process makes these suites spawn a real `codex app-server` when one is installed: 38 spawns and two failures in hook-service-runtime-trust-repair on a machine with codex, green in CI where there is none. Stand in for the missing binary so both environments exercise the same fallback lane. * docs(codex): scope the trust-RPC kill switch comment to what it actually gates The comment read as though the flag forces the fallback lane everywhere. It gates the managed grant only: the real-home rebase still runs its own inspect/repair app-server sessions when Orca's insertion shifts a user's hook positions, and never reads the flag. Verified by exercise, not by reading — with the flag set, both inspect-user-hook-trust and repair-user-hook-trust still ran. Pre-existing: main has no check there either, it just blocked the main thread while doing it. Widening the flag to cover the rebase is a follow-up; this only stops the comment promising something the constant does not do. |
||
|
|
96565fe370 |
perf(source-control): stop re-running every git read on each file selection (#15036) (#16600)
* perf(source-control): stop blocking main on four sync git-dir probes per status poll detectConflictOperation ran four existsSync calls against the git dir on every status poll. On a `\\wsl.localhost\...` worktree each one is a 9p round trip, and being synchronous they landed on the Electron main thread back to back. Replace them with concurrent fs/promises access probes: same "any failure reads as absent" semantics existsSync had, one wave instead of four serialized blocking calls. The outer try/catch went with them -- neither resolveGitDir nor the probes can throw now, so it was unreachable. Part of #15036 (source-control latency). * perf(wsl): let git reads take the shell-free route from a cwd-derived distro shouldAttemptWslDirectGit required options.wslDistro, so a `\\wsl.localhost\...` worktree without a resolved WSL project runtime never qualified -- even though the distro is right there in the cwd and wslDistroForCommand already knew how to read it. Every `git show` behind a diff therefore ran through the user's login shell, executing their rc once per blob read. Three changes: - Derive the distro from the cwd when no override was supplied. This is the fix; the routing decision now depends on where the repo actually lives. - Wait, bounded, for a cold read-environment probe instead of resolving without it. The probe is one wsl.exe call shared per distro, so the wait is paid at most once, and past WSL_GIT_READ_ENVIRONMENT_WAIT_MS the shell route runs exactly as before. It returns null rather than a settled promise when there is nothing to wait for, so a non-WSL git call is not pushed into a later microtask. - Opt the blob reads into preferWslDirectGit via gitReadOptionsForWorktree (renamed from gitStatusReadOptionsForWorktree; it was never status-specific). Belt-and- braces only: `show`, `config --get-regexp`, `ls-files` and `rev-parse` were all already matched by isWslDirectGitReadCommand, so this changes no routing today -- it just stops the diff path depending on a heuristic it knows the answer to. git-blob-read also gains a `failed` flag distinguishing "git ran and reported the path absent" (exit 128) from "the read never got an answer"; nothing consumes it yet, the settled diff cache does. Part of #15036 (source-control latency). * perf(source-control): give diff reads a settled cache keyed on stamped git state gitDiffReadDedupe coalesces only while a read is in flight, so every file selection re-ran the whole read: a `git config --file .gitmodules` spawn, one or two `git show` spawns, and a working-tree stat+read. On a WSL/UNC worktree each git spawn is a wsl.exe invocation, which is the ">3s Loading diff..." in #15036. Correctness first -- a stale diff is worse than a slow one. The cache never expires on a clock and there is no TTL to tune. Instead: - worktree-diff-stamp.ts takes a subprocess-free stamp of exactly the inputs a file diff is built from: HEAD (by resolved tip *content*, so a commit is visible even though HEAD's own bytes never move), `.git/index` (mtime+size), `.gitmodules` (submodule routing), and the working-tree file. A linked worktree's commondir and the packed-refs/reftable fallback are handled; an unborn branch is caught by recording "no loose ref" rather than only the packed stamps. - The stamp is captured BEFORE the read and stored with the result. Anything that moves during or after the read leaves the stored stamp behind, so the next lookup misses. That, not a freshness window, is why a stale diff cannot be served. - A store is refused unless the stamp was taken a full mtime bucket (2s, FAT's granularity) after its newest component. Below that, a second write inside the same bucket would be invisible -- git's own racy-index rule. - `null` stamp means "cannot prove" and never caches: a folder workspace, a repo whose layout cannot be read, or a filesystem reporting no usable mtime. - Submodule routes and reads that failed rather than proved absence are not reusable. A wsl.exe hiccup produces the same empty left side a new file does, and pinning that would persist a wrong diff. - invalidateGitReadCaches clears it and bumps a generation, so a read that started pre-mutation cannot store its result post-mutation. `ino` is deliberately optional in the working-tree component: Windows reports 0 for it on the redirector behind `\\wsl.localhost`, and requiring an unstable 0 to match would make the cache silently never hit on the exact host it exists for. Cache counters are exposed for the same reason -- a miss storm and a cold start otherwise look identical. Also drops gitDiffReadDedupe.clear() from getStatus. A status poll is a read; all it did was destroy a live coalescing entry so a concurrent identical request started duplicate git work. Mutations still invalidate through the shared point. Memory is bounded by retained characters, not entry count -- one diff result can legitimately hold megabytes. Fixes the source-control half of #15036. * perf(source-control): reuse BoundedMap and stop the WSL probe wait from outliving its answer Review follow-ups on the settled-diff-cache work: - SettledDiffCache now sits on the shared BoundedMap instead of hand-rolling the same Map + character ledger + evict-oldest loop. - pendingWslDirectGitReadEnvironment returns null once the probe has settled either way, so a distro whose direct route was disabled no longer pays for a 1.5s timer and two microtask hops on every git read. - That wait now honours the read's abort signal and goes through withTimeout, so an aborted read is not held for the full bound and a probe rejection can never surface as a read failure. - The settled-cache generation fence is taken before the stamp read, so a mutation that lands entirely inside the stamp's stats can no longer store an entry whose stamp is torn across it. - The cache counters are folded into the main-thread churn probe report, which is what tells a permanently-cold cache apart from a cold start in the field. * fix(source-control): tell WSL clock skew apart from a genuinely fresh write The racy-write margin compares two clocks: capturedAtMs is this host's, while the component mtimes come from whatever wrote the files. On a \\wsl.localhost worktree the guest sets them, so a guest running ahead pushes every recently-touched file past the margin and the cache refuses to store — for as long as the skew lasts, on exactly the platform this cache exists for. Nothing was wrong with the refusal; it was invisible. racyWrites alone cannot distinguish "the repo was just edited" from "the clocks disagree and this will never resolve on its own", so a permanently cold cache looked like a cold start. isDiffStampClockSkewed flags the one thing no local write can produce — an mtime in this host's future — and the cache counts those separately as clockSkewedWrites. A nonzero count is the signal that the cache is off for a reason idling will not fix. Found by review of #16600; behavior is unchanged, only observability. |
||
|
|
015f904fca |
fix(codex): stop re-scanning all Codex session history on every launch (#16251) (#16593)
* fix(codex): stop re-scanning all Codex session history on every launch (#16251) A launch deleted the backfill completion marker, and a marker could never be written while a Codex pane was open, so every launch re-derived "needs full scan" and walked the entire .codex/sessions tree — on Windows with a large history that read as a hung window. - v4 marker keeps a durable full-history baseline plus a bounded set of pending dates. v3 is read as a baseline, so upgrades pay no full scan. - A launch now marks dates pending instead of deleting the marker, and a full pass certifies the baseline even while a pane is still running; the live pane's own date just stays pending. - Pending dates are persisted, so an abnormal exit or a cross-midnight pane recovers a bounded window instead of a full walk. - A date-limited pass can only extend an existing baseline, never create one, so it can no longer certify history it never looked at. - Marker and index-heal target roots compare through normalizeRuntimePathForComparison, so Windows spellings of one directory stop invalidating each other. - Both append-only ledgers stream instead of readFileSync + whole-file JSON.parse, keeping the main thread responsive on large histories. * fix(codex): keep the backfill marker's full-scan demand durable Review follow-ups on the v4 backfill marker: - markCodexSessionBackfillMarkerPending no longer erases a persisted needsFullScan; the demand survives until a generation-current full walk retires it, and the function now reports it so the launch path folds it into its own in-memory flag (as @rumoii's #16252 does). - A full pass settles the whole pending set instead of subtracting the empty set, so a date a full walk provably covered stops forcing an extra bounded pass on every startup. - isCodexSessionBackfillDate does a real calendar check, so a corrupted marker cannot carry 2026/99/99. No age or future bound: the same guard gates rollout publication and a clock-skewed directory holds real sessions. - 'scans only the current date once a baseline exists' now has a second date directory, so it fails on a full walk instead of passing either way. |
||
|
|
0e10fc5925 | fix(browser): retire helpers with page owners (#16564) | ||
|
|
5a59bc5bc4 |
fix(grok): stop Orca's Grok hooks from costing anything outside Orca (#16666)
* fix(grok): stop Orca's Grok hooks from costing anything outside Orca Orca registers Grok agent-status hooks in the global $GROK_HOME/hooks. Grok loads that directory on every session, so a Grok run that Orca did not launch still paid for the hook on every event, and Orca rewrote the file even after a user had emptied it to opt out (#15518). The registered POSIX command now guards on ORCA_PANE_KEY before doing anything. That variable is part of the pane identity Orca injects into terminals it launches, and unlike the port and token it never comes from the endpoint file, so it is present exactly when the session belongs to Orca. A standalone session short-circuits without spawning a shell for the managed script at all. The same guard is applied to the remote install, because a remote host runs standalone Grok sessions too. PreToolUse is no longer registered. It is a blocking hook, so Orca sat on the critical path of every tool call and doubled the per-tool spawns, for a transition PostToolUse already reports. Windows cannot use the guard: the command there must be a single spawnable token, so it is a bare script path with no shell to evaluate a test. For that case the hooks are removed when Orca quits -- locally, on WSL guests, and on connected SSH hosts -- and reinstalled on the next launch. A config the user has emptied is left alone on startup; turning the setting back on in Settings is an explicit and later choice, so that path reinstalls. Removal is careful about what it is deleting. It strips only Orca's own entries, keeps user-authored ones, and deletes the file only when no hook entries remain -- keying that off the whole object would leave a stray non-hook key behind, and the emptied-config check would then read that remnant as a deliberate opt-out and never reinstall. A config the user has symlinked into a dotfiles repo is written through rather than unlinked, and is exempt from the emptied-config check for the same reason: after a quit it is a file Orca emptied, not one the user did. Writes go through temp+rename. Grok refuses to build a sandbox profile for a hook JSON with more than one hard link, so publishing by hard link would fail any session that started during the write. Install and removal on remote hosts now read the platform from the same field. They did not, so a Windows remote whose bridge env was incomplete had hooks installed and never removed. Co-authored-by: Siddiqui Qamar <137684575+siddqamar@users.noreply.github.com> * fix(grok): preserve hook state outside Orca --------- Co-authored-by: Siddiqui Qamar <137684575+siddqamar@users.noreply.github.com> |
||
|
|
cda2280d63 |
Show all automations (#16532)
* Add all-host automations with scoped ownership and multi-authority suppo
Enable automations to run on multiple hosts (SSH targets and local) with
owner-fenced mutations, scoped list queries per host, and conflict
resolution. Introduces desktop and runtime authorities as distinct
automation storage owners, with per-host caching, invalidation, and
retry scheduling on the renderer. Captures registration generations for
SSH hosts to survive re-adoption. Adds CLI support for destination
selection and conflict recovery.
* Filter automation create projects by destination host
Only offer projects available on the selected destination, preventing
the mismatches that would fail at submit time. Auto-adjust the project
selection if it becomes unavailable when the destination changes.
* Add runtime storage authority support for automations
- Support both runtime and desktop as automation storage authorities
- Make owner preconditions optional for legacy-client compatibility
- Cache automation list projections to improve performance
- Add per-row repo/worktree resolution for cross-authority collisions
- Extend automation.list RPC to always include owner metadata
* Replace child_process.execFile with runProcess for external automations
- Migrate external-manager to use cross-platform runProcess wrapper per child-process safety policy
- Abstract electron app/ipcMain APIs in orca-runtime via environment accessors
- Install fake app environment in automation tests for consistent setup
- Reorganize imports to use specific module paths (ssh-target-registry, agent-detection, browser-error)
- Remove external-manager from child-process import allowlists (no longer violates direct import)
* Unify desktop automation CRUD onto the local runtime RPC surface
The desktop authority now speaks the same automation.* RPC contract as
remote runtimes, via callRuntimeRpc({kind:'local'}) -> runtime:call ->
the shared RpcDispatcher. The automations:list/listRuns/create/update/
delete/runNow IPC arms, their preload members, and every renderer
desktop-vs-runtime transport fork are retired; the runtime methods are
the single implementation of scoped lists, owner fencing, and change
publication for both transports (mobile clients already exercised them).
The desktop probe scheduler's priority lease survives the move as an
AutomationService hook the IPC registration installs and the runtime
methods take, so Orca's own automation traffic still parks queued
external-manager probes.
External-manager scope arms and dispatch-loop plumbing stay on IPC by
design; automation change events keep their existing channels (renderer
ingestion already converges them by authority).
* Remove automation ghost SSH tombstone scanning
This functionality for synthesizing tombstones for automation-referenced SSH
targets is no longer needed as part of the automation system refactoring.
* Refuse orphan automations at dispatch time, not migration time
Remove migration-time disabling of orphan automations and the `enabledDecidedBy` field. Dispatch now refuses orphans at runtime instead, simplifying state management and UI. Orphans are left unstamped and enabled; dispatch refuses to run them via `resolveAutomationRunTarget`.
* Show all automations in flat table with unified filter menu
- Replace host picker component with comprehensive Filters menu supporting status, last run, agent, and host filters
- Flatten automation list layout to single table instead of host-grouped sections
- Add Host column to display execution host for each automation
- Display active filters as removable pills below toolbar
- Delete unused AutomationHostPicker* components
* Add automation owner fencing and destination validation
- New AUTOMATION_OWNER_FENCING_RUNTIME_CAPABILITY for owner preconditions; legacy clients get owner metadata snapshotted at RPC boundary for compatibility
- Editor captures and revalidates automation destination before save, preventing silent retargeting if SSH infrastructure changes mid-edit
- SSH target types now isolate renderer-authored fields; generation is server-owned and stripped by IPC handlers
* Route automation recovery actions to the origin host
When an automation action fails due to owner fencing, recovery verbs
("Update server", "Reconnect") must run on the host where the refusal
originated: the row's captured owner for row operations, or the
destination the create dialog captured, not the list's filtered host.
* Remove external manager scope limitation notices
Consolidate create destination eligibility checks with a unified predicate
and fix the bug where desktop repo IDs could be sent to runtime hosts where
they cannot resolve.
* Persist only store-derived automation contexts, not client-perspective o
Store contexts must never be based on client-provided runContext or sourceContext
values—clients speak a different perspective (e.g., 'runtime:<id>' for host IDs
they assign), and persisting those makes the store projection orphan automations
it actually owns. Derived contexts now take precedence in create and update paths,
with explicit null still honored to clear a value. Tests verify this by simulating
drift after storage and confirming that moves re-derive while toggles preserve.
|
||
|
|
a9781a4118 |
STA-4150: client-hosted remote browser (consolidated) (#15448)
Co-authored-by: Jinwoo-H <jinwoo@stably.ai> |
||
|
|
e361da7fb7 |
Deleting skill (#16357)
* Add skill deletion with cross-platform transaction safety Implements end-to-end skill removal with placement enumeration, dependency guards, and transactional recovery. Covers native, WSL, and remote hosts; users can delete canonical directories and alias placements (symlinked directories or files) in a single atomic batch. Includes UI selection flow, preview, confirmation, and results band. Block reasons (bundled, plugin, unowned, stale) gate deletions that would fail or contradict user intent. * Organize IPC handlers into module subdirectories Move register-core-handlers and skill-delete-ipc-handlers into dedicated subdirectories for improved code organization and to reduce the flat structure in src/main/ipc/. * Make skill deletion recovery transactions idempotent Defer journal cleanup until both staging removal and receipt cleanup succeed, leaving the journal in place for startup to retry if either operation fails. This ensures the recovery process is safe to run multiple times without leaving partially-deleted skills. * Consolidate skill-delete files into dedicated module Reorganize skill deletion functionality into a modular structure under `src/main/skills/skill-delete/` with simplified file names. Remove the redundant `skill-delete-` prefix from file names since they now live in the dedicated directory. Update all import paths throughout the codebase to reflect the new structure, including imports from IPC handlers and RPC methods. * Fix broken import paths and add deletion robustness improvements Import paths using `..//'` were invalid and broken. Replace with explicit module names (`skill-discovery-sources`, `skill-install-filesystem`, etc.) to clarify dependencies. - Bind WSL filesystem methods to preserve `this` context - Keep recovery journal when rollback rename fails, so startup can retry - Skip symlink-based tests on Windows where they cannot run - Only treat ENOENT/ENOTDIR as empty directories; propagate other errors - Fix cross-platform path parent calculation to handle drive roots - Replace shared constant with localized string for user-facing message - Use `runProcess` for WSL integration test instead of bare `execFile` * Add batch limit for skill deletion and improve host availability checkin - Limit concurrent deletions to prevent remote host overload - Add retry logic for capability probing to handle transient unavailability - Add reprobe() method to recheck capability after errors or user refresh - Fix status logic: receipt cleanup is best-effort, completion depends only on content removal - Improve error message for unreachable hosts |
||
|
|
4218d5068e |
fix(cli): seed nvm's default version, not the newest install (#16420)
* fix(cli): seed nvm's default version, not the newest install #16314 stopped the login-shell probe inheriting the seeded PATH, but left the seed itself picking the newest installed nvm version. That ordering decides which node a CLI runs under whenever the probe does not land — a timeout, or a login shell whose rc never initializes nvm — and newest is precisely the wrong guess: it is usually the version the user just added and has installed nothing into. That is the root cause reported in #10932. Resolve `alias/default` instead, mirroring nvm: follow the alias chain (`default` -> `lts/*` -> `lts/krypton` -> a version), resolve a partial version like `24` to the highest matching install, and treat `system`/`node`/`stable` as no preference. The chain is bounded and cycle-guarded because nvm's own resolver tracks seen aliases and hand-edited files can point at each other. Ordering is a preference, not a restriction: the remaining versions stay behind the default, so a CLI installed outside it is still reachable. Measured on a real machine with nvm default=24 and a bare v26.7.0 installed: the old resolver seeds v26.7.0/bin (no CLIs), the new one seeds v24.18.0/bin (every CLI). Tests were written first and verified to fail on the three bug cases against main before the fix existed. Also raise the probe budget from 5s to 10s. The old value was never measured against a real profile: a bash -ilc loading nvm, rvm, conda and gcloud takes ~1s idle but 6-7s on a loaded machine, so a cold start under load silently fell back to the seed. Startup does not block on the probe, and the one awaited consumer is agent detection, which is better served by a probe that finishes late than one that gives up early. * fix(cli): reject non-version alias tokens instead of matching v0.x Review finding, and a real bug I introduced. parseVersionSegment coerces every unparseable segment to 0, so an unresolvable default alias — `garbage`, `iojs`, `lts/nonexistent`, any hand-named alias — became [0] and prefix-matched a `v0.12.x` install, or any stray non-version directory. Orca would then seed a decade-old node as the preferred runtime. Real nvm answers N/A for all of them. The `wanted.length === 0` bail could never have caught this: ''.split('.') is [''], never empty. Replaced with a shape check that still admits legitimate numeric prefixes — verified against nvm itself, which resolves `24` to v24.18.0 and `0` to an installed v0.x while answering N/A for the rest. Also corrects two comments that no longer described the code: the seed is no longer "newest install", and the probe budget note claimed startup never blocks on hydration, which is false on packaged Windows where it gates terminal services and git. The traversal-guard comment claimed a containment join() already normalizes away; the real guarantee is that matchNvmVersion can only return an entry of the versions directory. * fix(cli): match nvm's version-token grammar, not just its first character Round-2 review finding, and the same bug one layer down. The previous guard anchored only the first character, but parseInt stops at the first non-digit, so `0x18`, `00` and `0abc` still parsed to [0] and prefix-matched a v0.12.x install — the decade-old-node seed the earlier fix was supposed to close. Reachable: `nvm alias default 0x18` warns that the version does not exist and writes the alias anyway, then resolves it to N/A. Use nvm's actual grammar, leading zeros included — nvm calls `00` and `024` N/A while parseInt reads them as 0 and 24. Verified by executing 17 tokens against a five-version fixture: every one now agrees with nvm, including the legitimate prefixes `0`, `0.12`, `24` and `v24.18.0`. Also drops a dead disjunct (the hop bound already caps the loop, so seen.size can never exceed it) and corrects the log comment in index.ts, which still told the reader a failed probe leaves the newest install in front. It leaves the default version in front now, which is usually survivable but still not what the shell would have resolved. * test(cli): skip the lts/* chain fixture on Windows Round-3 review finding. makeNvmHome materializes each alias as a real file, and the chain case uses nvm's actual `lts/*` alias — `*` is a reserved Win32 filename character, so writeFileSync fails with EINVAL. PR CI runs a Windows allowlist that excludes this file, so the breakage only reaches a Windows developer running the suite locally. Skipped rather than renamed: `lts/*` is the alias nvm really ships, and the assertion pins platform: 'darwin' anyway, so the real name costs no coverage. Matches the skipIf convention already used across src/shared. Also reflows a comment line that a previous edit ran to 143 characters; oxfmt does not reflow comments, so nothing would have caught it. |
||
|
|
1f39c93b01 | refactor(ipc): split ssh.ts into focused modules (#16394) | ||
|
|
5c116b6ec2 |
fix(startup): stop the PATH seed pinning nvm to its newest install (#16314)
* fix(startup): stop the PATH seed pinning nvm to its newest install patchPackagedProcessPath prepends the newest nvm version dir to process.env.PATH, then hydrateShellPath probes the login shell with that same env. nvm's startup `use` honors whatever node is already on PATH instead of the user's `default` alias, so the probe returns a PATH pinned to the newest install and every terminal pane inherits it. A user whose newest nvm node is a bare install then loses every global CLI (codex, claude, gemini, vercel...) inside Orca while they still resolve in Ghostty/Terminal, which start from the bare GUI PATH and fall through to `default`. Probe with the PATH the process launched with. Windows already gets this through WindowsShellPathOwnership. * test(startup): pin the platform in the probe-env test shellProbeEnv short-circuits on win32, so the POSIX-only assertion failed for anyone running the suite on Windows. Matches the convention in hydrate-shell-path.windows.test.ts. Also narrows the win32 exemption comment: WindowsShellPathOwnership snapshots its baseline after the seeds land, so it does not unwind them. The exemption holds because Windows keys PATH as `Path` and no Windows seed pins a node version. * fix(startup): give win32 the same probe insulation, keyed by Path The previous commit exempted win32 on the stated grounds that no Windows seed pins a node version. That is wrong: getVersionManagerDirectories calls getNvmVersionDirectories on every platform, so a Git Bash user whose nvm uses the POSIX ~/.nvm/versions/node layout gets the newest version dir seeded on Windows too, and the -ilc Git Bash probe inherits it. Record the PATH key alongside the value and overwrite that entry in place, so Windows never carries both `Path` and `PATH` — which was the only real reason to skip win32. * fix(startup): snapshot the launch PATH at module load, log probe failures Three loose ends from the review, folded in rather than deferred. The probe's clean PATH was handed over by an explicit recordLaunchPath call from the seeding site, so the invariant lived across three files and a refactor that moved the seed call would silently re-pin nvm. Snapshot PATH during module init instead: that runs while the import graph is evaluated, strictly before any statement in main's body, so it cannot observe the seeds and there is no call ordering left to break. All seven importers are static, so no lazy import can defeat it. A test asserts the probe ignores later process.env mutation, and fails if the live read is reintroduced. A failed startup probe leaves the seeded newest-nvm dir in front and said nothing, so the population whose rc files blow the 5s budget hit the original symptom with no diagnosable trace. Log the failureReason. The probe is an interactive login shell, so rc files that exec into a multiplexer or start a heavy prompt can outrun that budget with no way to opt out. Set ORCA_SHELL_PATH_PROBE=1 so they can take a fast path. * fix(startup): drop the other-cased PATH key from the Windows probe env Caught by running the suite on a real Windows machine, not a mocked platform. The spread of process.env is a plain, case-sensitive object, while Windows resolves env names case-insensitively. Writing the captured `Path` back onto it left the seeded value still live under `PATH`, so the probe shell could read either one — the exact duplicate-key hazard the win32 branch was supposed to prevent. Drop any other-cased variant of the key before writing. * fix(startup): preserve launch PATH across app restarts |
||
|
|
fba910f2ea | fix(crash-reporting): scope renderer crash evidence (#16313) | ||
|
|
94f231737d |
fix(agent-status): retire panes whose agent process is gone (STA-4612) (#15212)
* fix(agent-status): retire panes whose agent process is gone (STA-4612) Agent status can hold `working` on a pane where no work is outstanding, and nothing closes the gap. A pane's Claude state is a join of a lead turn and three latches — the subagent roster, the background-task gate and the session-cron gate — and each is set by a hook and cleared only by another hook. Claude Code emits no terminating hook on `/exit`, `/clear`, Ctrl+C, crash, SIGKILL or terminal close, so every one of those latches is a claim with no owner and no expiry. The join is also materialised at ingest time and persisted, so a stale `working` survives restart and blocks hibernation, which requires `done`. Registering `SessionEnd` is not the fix: it covers roughly a third of exit paths (measured on 2.1.231/2.1.233; upstream anthropics/claude-code#17885 and #6428 are both closed as not planned). Nor is a TTL — `AGENT_STATUS_STALE_AFTER_MS` only decays the sidebar dot at read time while the stored row stays non-terminal. So the backstop is built from evidence Orca already owns. A session id that changes means the conversation was replaced. On the first hook of the new session — whatever that hook is — the previous session's own claims are void: its session crons and its one-shot subagents. Deliberately not voided: the background-task gate (a background shell is an OS process that survives `/clear`, and the previous inventory is positive evidence it was running), and `confirmedTeammate` rows (persistent in-process teammates a lead swap cannot end). The lead record is left to the incoming event's own fold. A certified process exit retires the pane. Orca already does this on every attributable PTY exit — `clearProviderPtyState` resolves the pane key and calls `clearPaneState` — but that resolution depends on the spawn-time `ptyPaneKey` mapping, which a restored or reattached PTY may never rebuild. Those panes keep their row and latches for good. `onPtyExit` knows the keys teardown could not resolve, so it reconciles them from its own records. The certificate is `exitCode >= 0 || hostExitConfirmed || providerExitObserved`: a synthetic `-1` from a failed stop is not a death (the PTY can have survived it), while a real exit can also report `-1`, so neither the code nor the SSH surface predicate is sufficient alone. `providerExitObserved` is additive and separate from `hostExitConfirmed`, which also drives the liveness verdict and the SSH surface decision. A confirmed shell foreground is the `/exit` case: the agent died, the shell lived. That already dropped the row, but through `agentStatus:drop`, which by its own contract preserves a live pane's caches — so every latch survived and the next event resolved the pane back to `working`. It now routes through the reconciler instead, gated on a per-pane accepted-status generation rather than row identity: the confirming process read can take seconds, and `updatedAt` cannot order two writes inside one millisecond (the store deliberately admits equal timestamps). Cold start generalises the same way. The startup sweep required a restored subagent roster, so a stranded lead row, background-task gate or cron gate — the shapes with no child event left to reap them — were never candidates. Hibernation needs no change: with the above, those rows become genuinely `done` and the lockout resolves through the front door. A `restoredUnconfirmed` bypass in the planner would let it reclaim the heap of an agent that may be working. Not included: folding `background_tasks` from a child-attributed `SubagentStop`. Writing its test surfaced #11838's deliberate assertion that child inventories are not authoritative for lead-owned background work, and the listener says the same — "background_tasks is trusted only where unambiguous". An empty list on a `SubagentStop` does not prove the lead's shell ended, so the fold would have cleared a gate on evidence that establishes nothing. STA-4119's live-side question — whether a genuinely live background shell should hold the lead row after the lead turn ends — is untouched. This change extends gate-clearing to zero new triggers. * fix(agent-status): make the confirmed-shell reconcile survive its own drop The /exit leg never fired. `settleDeferredCommandFinishedStatusDrop` runs the paired drop before the reconcile, and `dropAgentStatus` cleared the per-pane accepted-status counter the reconcile's guard then read — so the guard compared a live anchor against a zeroed counter and skipped itself on every pane that had a status row, which is every pane worth reconciling. The existing test passed only because it used a pane with no row, where the drop early-returns and both sides read 0. Stop keying the guard on a counter a sibling teardown path can reset: the ordinal is now stamped on the row itself, derived from the row it replaces, so there is no side table to clear and a batched burst lands the same ordinals as the equivalent sequential writes. A removed row means "nothing reported", which is exactly what the paired drop leaves behind. Also: - Keep the `providerSessionOnly` resume identity that the paired dismissal mints when the shell outlived the agent; a certified PTY exit still takes it, since there is no pane left to resume into. - De-vacuum two guard tests. The confirmed-teammate pin never anchored a session owner, so the void it claimed to survive never ran; the unavailable-inspection pin asserted before the confirm ladder settled. Both now fail when their guard is removed. - Derive `hasLiveClaimsForPaneKey` from a predicate that lives beside `clearPaneCacheState`, so a new latch cannot be added to the teardown and silently missed by the claim check. - Drop the unreachable compact-`trigger` clauses; SessionStart is the whole guard. - Cover the connectionId arm of the exit certificate, where a provider-observed death and a preserved SSH surface are deliberately independent. * fix(agent-status): keep agent-status-types under its line cap main already sits exactly at the 300-line max-lines cap for this file, so the single `acceptedStatusSeq` field this branch adds pushed it to 301 once main's observation facet merged in. Declared the field as a mixin beside the observation facet instead. Both are per-write facets mixed into `AgentStatusEntry` rather than fields a reporter supplies, so they belong together — and the capped file loses a line rather than gaining one, since it already imports from that module. No lint suppression. * fix(agent-status): collapse the entry facets into one intersection The previous attempt still tripped max-lines: two mixins on one intersection wrap across two lines under oxfmt, so removing the field line bought nothing. Expose a single AgentStatusRowFacets that already includes the observation facet, so the entry intersects one short name on one line. The payload keeps intersecting the observation facet alone — it must not carry the renderer-local ordinal. Verified by formatting first and then linting, which is the order that catches this. * fix(agent-status): retire resume authority with dead panes |
||
|
|
1375c57b16 |
fix(secrets): report the protection gap on change, not on every launch (#16044)
The gap warning fired every startup with no way to stop it. It usually needs a keyring installed and unlocked to fix, so repeating it every launch is nagging the user cannot act on and will learn to ignore. It now reports when the answer changes: once when the gap starts being true, again if it becomes true for a different reason, and once when it is fixed — because silence after a "your secrets are not protected" warning would leave the user assuming that is still the case. State lives beside the profile data file, which is why the call moved out of the port bootstrap: that state has nowhere to live until the profile exists. A corrupt state file re-reports rather than trusting it, and a failed write logs instead of failing startup, since re-reporting next launch is the safe direction. ORCA_ALWAYS_REPORT_SECRET_PROTECTION=1 forces a re-report for support without disturbing the stored state. |
||
|
|
838f5bfb75 |
fix(secrets): tell Linux users when their secrets are only obfuscated (#16033)
On Linux with no keyring, Electron falls back to the `basic_text` backend, which "encrypts" with a hardcoded password. `isEncryptionAvailable()` returns true for it, so Orca reported those secrets as sealed. They are not. The obvious fix — returning false for basic_text — is wrong and would have been a credential regression: `decryptWithStatus()` skips decryption entirely when encryption is unavailable, so every already-stored secret would read back empty. Sealing genuinely works on basic_text and must keep working. So capability and trust are now separate questions. `isEncryptionAvailable()` still answers "can this host seal and unseal", and `describeProtectionGap()` (renamed from `describeUnavailable`) answers "is my data actually protected", covering both no-sealing and weak-sealing. That method had no production caller — the port documented a promise nothing kept. `reportSecretProtectionGap()` now reads it at startup. A user-visible surface is follow-up; this at least stops the silence. Adds a bootstrap wiring guard over all nine host port installs. The no-op defaults are correct for a renderer-less host and silently wrong for the desktop, and a dropped or reordered install fails no existing test. Verified in both directions: it fails when an install is removed, and when one moves after the runtime is constructed. |
||
|
|
03fcfdfb92 |
feat(orcad): boot the Orca runtime on plain Node (#15968)
* refactor(host): resolve the app root through the port in fork-reachable modules
`parcel-watcher-entry-path.ts` and `session-scanner-service-entry-path.ts` read the
app root via `require('electron').app` inside a try/catch that already returns null
when Electron is absent. They were therefore correct under plain Node at runtime and
only failed the *static* text check — which is real, not pedantic: the comment in
`ports/port-scan-command-client.ts:19` records that the plain-node-entry-guard fails
on that literal text, try/catch or not.
`hasAppEnvironment() ? getAppEnvironment() : null` gives the identical "no app root
here" answer without the text. That restores `hasAppEnvironment`, which an earlier
commit in this stack deleted as unused — it now has the caller it was waiting for.
Ratchet baseline 27 → 25.
Verified: 74 files / 458 tests; `pnpm typecheck` clean; `oxlint` clean.
* feat(orcad): boot the Orca runtime on plain Node
Closes the last two Electron couplings and makes `orcad` a working artifact:
a 4.43 MB Node bundle that boots, pairs, registers a repo, creates a real git
worktree and round-trips a PTY — with zero `require("electron")`.
Ratchet 2 -> 0, so `config/runtime-electron-baseline.txt` is now empty and its
test asserts exactly that: any reachable electron import is a regression.
- speech: inject the service factories, so importing ModelManager for its type
no longer drags Electron's streaming net.request into the graph
- filesystem-watcher: add a WorktreeWatcherRemoval port. Every entry in those
maps arrives through an ipcMain handler carrying a renderer sender, so a host
with no renderer has nothing to close, restore or forget — the inert default
is what the desktop code does against empty maps, not a stub hiding work
- user-data-path / profile-storage-paths: resolve userData through
AppEnvironment. These surfaced only once orcad pulled the store in
Both host ports now anchor to a realm-global symbol. `vi.resetModules()` gives
the re-imported graph a fresh module copy, so a binding installed before the
reset silently read back as uninstalled.
The acceptance smoke drives both hosts through one code path (`--target
orcad|electron`) and seeds its own git repo, so it is hermetic and asserts the
same contract of each. Wired into PR CI.
* test(smoke): remove the seeded workspace container, not just the worktree
* test(smoke): surface the server's stderr when it dies before ready
* fix(smoke): build node-pty for Node before booting orcad in CI
* fix(smoke): drive the CLI built from this checkout, not one on PATH
* docs(ratchet): say the baseline must stay empty, not merely shrink
* build(orcad): externalize only the native modules actually in the graph
|
||
|
|
f975035809 |
refactor(ipc): split preflight and SSH registry out of the ipcMain modules (#15927)
* refactor(preflight): split agent detection out of the ipcMain registration
First of the IPC extractions the revised design requires. `src/main/ipc/preflight.ts`
mixed 285 lines of agent/tool detection with 35 lines of `ipcMain.handle`
registration, and the runtime calls that detection during normal operation
(`orca-runtime.ts:573`, plus the preflight RPC methods). So the runtime dragged
`ipcMain` into its graph to reach pure logic.
Detection moves to `src/main/preflight/agent-detection.ts` — named for what it
contains, per AGENTS.md. `ipc/preflight.ts` keeps only the handler registration and
re-exports the domain module so existing importers are unaffected. The runtime and
its RPC methods now import the domain module directly.
Ratchet baseline 36 → 35: `src/main/ipc/preflight.ts` is no longer reachable from
the runtime. The gate detected the improvement and refused to pass until the
baseline tightened, which is the behaviour it was built for.
Verified: 2 files / 1,187 tests pass across every suite touching preflight;
`pnpm typecheck` clean; `oxlint` clean.
* refactor(ssh): split the SSH target registry out of the ipcMain module
Second IPC extraction, and by far the biggest win: this removes **eight** modules
from the runtime's Electron graph, taking the ratchet baseline 35 → 27.
The runtime needed five thin accessors from `src/main/ipc/ssh.ts` —
`connectRegisteredSshTarget`, `getRegisteredSshState`, `listRegisteredSshTargets`,
`listRegisteredRemovedSshTargetLabels`, `getActiveMultiplexer`. Each is a one-line
read over module-level state. Importing them dragged in `ipcMain`, `powerMonitor`
and a `BrowserWindow` accessor — and, transitively, `ipc/pty.ts` (8,031 lines),
`ssh-browse`, `ssh-passphrase`, `ssh-relay-deploy`, `ssh-remote-cli-host-passthrough`,
`wsl-hook-relay-launch` and `user-data-path`.
`src/main/ssh/ssh-target-registry.ts` now holds that state plus its accessors.
`registerSshHandlers` populates it; the runtime reads it. The indirection is kept
deliberately: SSH providers register after construction and may reconnect, so
callers must resolve the current generation rather than freeze one.
`ipc/ssh.ts` re-exports all five, so non-test importers are unaffected.
`connectRegisteredSshTarget` still throws `ssh_handlers_not_registered` when no
handler layer registered — a headless host must fail loudly rather than report a
target as unreachable, which would read as `exited` (see ssh-execution-boundary.md).
Verified: 9 files / 59 tests across the ssh, automations and trust-preset suites;
orca-runtime.test.ts 1,183 pass; `pnpm typecheck` clean; `oxlint` clean.
* refactor(host): resolve the app root through the port in fork-reachable modules
`parcel-watcher-entry-path.ts` and `session-scanner-service-entry-path.ts` read the
app root via `require('electron').app` inside a try/catch that already returns null
when Electron is absent. They were therefore correct under plain Node at runtime and
only failed the *static* text check — which is real, not pedantic: the comment in
`ports/port-scan-command-client.ts:19` records that the plain-node-entry-guard fails
on that literal text, try/catch or not.
`hasAppEnvironment() ? getAppEnvironment() : null` gives the identical "no app root
here" answer without the text. That restores `hasAppEnvironment`, which an earlier
commit in this stack deleted as unused — it now has the caller it was waiting for.
Ratchet baseline 27 → 25.
Verified: 74 files / 458 tests; `pnpm typecheck` clean; `oxlint` clean.
* test(ssh): mock the SSH target registry alongside the ipc/ssh mock
Thirty-eight suites mocked `vi.mock('./ssh')` for `getActiveMultiplexer`. That
factory went inert when production started importing the accessor from
`../ssh/ssh-target-registry`, so the real module loaded and the assertions drifted.
Adds a companion registry mock returning the same stub, plus a
`sshTargetRegistryModuleMock` builder beside the existing `sshModuleMock` so the
shared harness stays one place. No assertion changed.
Found by a full-suite run: the targeted ssh/runtime suites were green while
30 tests in ipc/worktrees and ipc/repos were not.
* refactor(runtime): read app paths and the packaged flag through the port
`orca-runtime.ts` is the last module in its own graph that imports `electron`
directly. Nineteen of its uses were `app.getPath` (12) and `app.isPackaged` (7) —
exactly what the AppEnvironment port already covers.
Also removes a dead `const { app } = require('electron')` inside
`getOrchestrationDb`. It was left unused once the path came from the port, and it
is precisely the dynamic-require pattern `plain-node-entry-guard.ts` exists to
catch, sitting in the runtime's own constructor path.
What still binds `orca-runtime.ts` to Electron is now three sites, not nineteen:
`new Notification(...)` (one), `BrowserWindow.fromId` (one), and the
`ipcMain.on('terminal:tabCreateReply')` renderer round-trip — which is the browser
tab path, and the same one that would hang a headless host for ten seconds.
Two suites drove `electronMocks.app.isPackaged` directly; they now install a fake
AppEnvironment reading the same mutable field, so their per-test toggles work
unchanged and no assertion moved.
Verified: 376 files / 4,717 tests across src/main/runtime; typecheck and oxlint clean.
* test(serve): add the built-artifact terminal round-trip acceptance smoke
"The server started" proves almost nothing. Terminal creation dispatches into
OrcaRuntimeService, and without an installed headless PTY controller that path
falls through to a renderer reply that never arrives and times out after ten
seconds. A boot probe, a port bind, and a `host.platform` call all pass against a
server whose terminals are dead — which is exactly the gap the design doc's own
boot proof was retracted for.
This boots the BUILT `out/main/index.js --serve`, parses its ready payload, pairs a
real client over the advertised endpoint, lists worktrees, creates a terminal, runs
a command through the PTY, asserts the output comes back, and asserts clean
shutdown. It drives nothing but the public pairing + RPC surface, so the same
script is the acceptance gate a future Node-only backend must pass unchanged.
The sentinel invokes `process.execPath` rather than `echo`, because the shell
differs per platform and node does not.
Verified both directions: passes against the real server, and fails with an
actionable message when the command produces no output — a smoke that cannot fail
is worthless.
* fix(ssh): fail loudly when the multiplexer resolver was never installed
`getActiveMultiplexer` resolves through a resolver that `ipc/ssh.ts` installs at
module scope. A process that never loads the SSH layer — which is the whole point
of the Node-only backend — would get `undefined` from every call.
`undefined` already means something specific here: "not connected". So a missing
resolver and a disconnected target were indistinguishable, and a host with no SSH
layer would quietly report every target as not connected. That is the
unverifiable-reported-as-exited conflation `docs/reference/ssh-execution-boundary.md`
exists to prevent — the doc is explicit that absence of contact is never evidence
of absence of the thing.
A missing resolver is a wiring error, not a connection state, so it throws, matching
what `connectRegisteredSshTarget` already does for unregistered handlers.
Verified: 432 files / 4,759 tests across ipc, ssh, preflight, automations and trust
presets; typecheck and oxlint clean.
* refactor(pty): stop faking a BrowserWindow for the headless PTY path
`registerHeadlessPtyRuntime` passed `registerPtyHandlers` a stub object cast to
`BrowserWindow` whose `isDestroyed()` returned true and whose `webContents.send`
was a no-op — a window-shaped thing that lied about being a window, purely to
satisfy the type. Adversarial review named it as the same "looks fine, silently
returns a lie" pattern this codebase rejects elsewhere, and it is the shape that
keeps `electron` on a path that otherwise needs none.
`registerPtyHandlers` now takes `BrowserWindow | null`. An absent renderer is
semantically identical to a destroyed one — all 42 call sites already guarded on
`isDestroyed()` and skipped — so `src/main/ipc/pty-renderer-surface.ts` states that
directly: `isRendererGone`, `sendToRenderer`, `rendererWebContents`. The compound
`isDestroyed() || webContents.isDestroyed()` guards collapse into one predicate.
`isPtyWriteEventFromMainWindow` becomes null-tolerant and fails closed: with no
renderer no sender can legitimately match, so every write is rejected. Those
handlers cannot fire headless today, but failing closed is the right answer if that
ever changes.
This is the precondition for installing a PTY controller without Electron, which is
what a Node-only backend needs and what `terminal.create` actually calls.
Verified: 129 files / 2,473 tests across ipc/pty, providers and orca-runtime; the
built-artifact acceptance smoke still passes end-to-end (boot → pair →
terminal.create → sentinel → close), which is the check that matters most here
since this changes the headless PTY path itself; typecheck and oxlint clean.
* refactor(pty): read app paths and the packaged flag through the port
Follows the fake-window removal. `ipc/pty.ts` had nine `app.*` reads — all
`getPath`, `getVersion` or `isPackaged` — which the AppEnvironment port already
covers. The `BrowserWindow` import was also dead after the null-window change.
What still binds this file to Electron is now `ipcMain` (75 uses, all handler
registration) and `powerMonitor` (2). That is a clean statement of the remaining
job: split logic from registration, the same shape already applied to preflight
and the SSH registry.
Test wiring: the shared `pty-ipc-suite-environment` beforeEach installs a fake
AppEnvironment that reads through the existing `vi.mock('electron')` app object
rather than freezing values — suites toggle `app.isPackaged` mid-test to exercise
dev-mode spawn paths, so the port has to observe the same mutable field. One edit
in the shared harness covers every pty suite.
Verified: 128 files / 1,290 tests across ipc/pty and providers; the built-artifact
acceptance smoke passes; typecheck and oxlint clean; ratchet unchanged at 25.
* refactor(pty): inject the ipcMain surface so the PTY module loads without Electron
This closes the round-3 blocker: "the doc never says how orcad installs
setPtyController without Electron."
`registerPtyHandlers` owns the `RuntimePtyController` that `terminal.create`
actually spawns through — the thing a Node backend needs and cannot get from the
provider thunks. The module was otherwise host-agnostic already; the only thing
pinning 8,031 lines to Electron was a static `ipcMain` / `powerMonitor` import used
purely to register renderer handlers that no headless host will ever receive.
`src/main/ipc/pty-host-bindings.ts` makes those surfaces settable, defaulting to
no-ops. Unlike AppEnvironment and SecretStore, the default does NOT throw: a host
with no renderer legitimately has nothing to register against, so not registering
handlers nobody can call is correct rather than a hidden downgrade. The desktop
installs the real objects in `attach-main-window-services` before its handlers run.
Also converts the remaining electron import to a top-level `import type`. oxlint's
`no-import-type-side-effects` caught that inline `type` specifiers still leave a
side-effect import — precisely the "type-only is not enough if esbuild still emits
require('electron')" trap a reviewer flagged.
**`src/main/ipc/pty.ts` now bundles with zero `require("electron")`.** A Node entry
can call `registerPtyHandlers(null, runtime, …)` and get a working PTY controller.
Verified: 128 files / 1,290 tests across ipc/pty and providers; the built-artifact
acceptance smoke passes end-to-end — which is the check that matters, since this
changes how every PTY handler registers; typecheck and oxlint clean.
* fix(pty-bindings): drop two unused eslint-disable directives
CI runs oxlint with unused-disable reporting; the two
`@typescript-eslint/no-explicit-any` suppressions I added were never triggered by
any enabled rule, so they failed static analysis as dead directives. The `any[]`
rest args stay — they mirror electron's own IpcMain signature, and narrowing them
would reject the real object at the desktop call site.
Verified with the exact CI invocation: `oxlint --format github` reports 0 warnings,
0 errors across the repo.
* fix(pty): install the host bindings per process, not per window
A real regression my own change introduced, caught by the SSH docker E2E
(`paired-startup-exec-readiness` — "recovers startup exec through a headed paired
desktop owner"). It reproduced on rerun, so it was not a flake.
`setPtyHostBindings` was called inside `attachMainWindowServices`, i.e. when a
window attaches. But `registerHeadlessPtyRuntime` (index.ts:3163) calls
`registerPtyHandlers` on the serve path *before* any window exists — so those
handlers registered against the no-op default and never reached the real `ipcMain`.
A paired desktop owner then attached to a runtime whose PTY handlers were wired to
nothing.
The bindings describe the *host*, not the *window*: an Electron main process always
has `ipcMain`, whether or not a window is open. Installing them beside
`setAppEnvironment`/`setSecretStore` at the top of bootstrap fixes both paths.
Verified: 128 files / 1,290 tests; the built-artifact acceptance smoke passes;
typecheck clean; `oxlint --format github` (the exact CI invocation) reports 0/0.
* feat(orcad): de-electron the runtime core and add the Node entry + build gate
**`src/main/runtime/orca-runtime.ts` — 41,048 lines — no longer imports electron.**
Its last three sites go through `runtime-desktop-surface.ts`: a native notification,
the authoritative-window lookup, and the one `ipcMain` channel used by the
renderer-backed tab-create fallback. All three are unreachable without a renderer —
`createTerminal` already takes the background branch when no window exists (#10333) —
so a Node host installs none and the runtime relays notifications to paired clients,
which is the better destination anyway. Ratchet 25 → 24.
Adds `src/main/orcad/orcad-entry.ts`: Node host adapters plus a `startOrcad` that
constructs the runtime, installs the PTY controller via `registerPtyHandlers(null, …)`,
and serves RPC. It sets two defaults the constructor gets wrong for a headless host —
`canRecoverPersistentLocalPtys: false` (no daemon here) and
`getDesktopWindowStatus: 'blocked'` (a Node host can never be promoted to a desktop
window, which is what `'openable'` claims).
Adds `config/scripts/build-orcad.mjs`, which **currently fails, on purpose**: 25
modules still import electron (browser and speech clusters, plugins, jira/proxy,
filesystem-watcher, and four `require('electron').app` one-liners). It names them.
Two bugs found while building it, both worth recording:
- The first bundle looked clean and was not. `electron` was bundleable, so esbuild
rewrote the metafile `path` to the resolved file under node_modules and a check for
`path === 'electron'` passed while the package was in the bundle — it failed at
runtime with electron's own installer message. The check now reads `original`, and
electron is marked external so a residual import fails loudly instead.
- `jsonc-parser`'s UMD build breaks the bundle at load; aliased to its ESM entry, the
same fix `build-relay.mjs` already carries.
Verified: desktop unchanged — the built-artifact acceptance smoke passes, runtime/pty/
provider suites green, typecheck clean, `oxlint --format github` 0/0.
* refactor(host): drop the last two require('electron') app lookups
`computer/sidecar-client.ts` and `ports/port-scan-command-client.ts` read the app
root through `require('electron').app` inside a try/catch. Both were already correct
under plain Node at runtime — they return null when it throws — but the literal text
fails the plain-Node entry guard regardless, which is why port-scan carried a comment
warning it must never become reachable from a fork entry.
Reading the AppEnvironment port gives the identical "no app root here" answer without
the text, so that warning is now obsolete and the comment says so.
Ratchet 24 → 22. Every remaining entry is a real coupling: the browser cluster (15,
which variant B does not ship), speech (2), plugins (2), and jira/proxy-settings (2,
needing an HttpClient port for Chromium session partitions).
Verified: 25 files / 209 tests; acceptance smoke passes; typecheck and
`oxlint --format github` clean.
* docs(orcad): record that the ratchet under-counts orcad's graph
The ratchet reports 22 electron importers; the orcad build reports 23. The extra is
agent-hooks/wsl-hook-relay-launch.ts, and the cause is a gap in the gate rather than
a rounding error: the ratchet measures what orca-runtime + runtime-rpc reach, while
orcad's entry also imports ipc/pty directly to install the PTY controller.
Once orcad ships it must become a ratchet entry point, or the two numbers drift and
the gate quietly stops covering the artifact it exists for.
* refactor(runtime): inject the browser commands factory
Drops 14 modules from the runtime's Electron graph in one change — the whole Chromium
browser cluster. Ratchet 22 → 8.
`OrcaRuntimeService` constructed `RuntimeBrowserCommands` as a field initializer, and
that construction is what pulled in `BrowserWindow`, `session`, `webContents` and the
cookie jars. Importing the class for its *type* is free; only building it costs.
So the class import becomes `import type`, and the instance comes from
`runtime-browser-commands-factory.ts`. The desktop installs the real factory at the
Electron entry. **All ~80 existing `this.browserCommands.*.bind(...)` delegations are
untouched** — a review round specifically warned that rewriting those was the
expensive, risky part, and this avoids it entirely.
With no factory installed, browser commands reject per call with `browser_unavailable`
rather than resolving to a stub that silently succeeds. The runtime already filters
browser capabilities out of `getStatus()` when no backend exists, so clients do not
offer the affordance in the first place.
Also corrects a stale comment in `pty-renderer-surface.ts` that still described the
fake window as present tense; it was deleted two commits ago.
Verified: 451 files / 5,513 tests across `src/main/browser` and `src/main/runtime` —
the entire browser automation suite; the built-artifact acceptance smoke passes;
`pnpm typecheck` and `oxlint --format github` clean.
* refactor(host): extract the plugin client list and port two app lookups
Ratchet 8 → 5.
- `listPluginsForClients` moves to `src/main/plugins/plugin-client-list.ts`. It needed
only three `plugins/*` helpers, none of them Electron — it was colocated with
`ipcMain.handle` registrations, so the runtime's `plugins.list` RPC dragged all of
Electron in to call a function that reads a lockfile. Same shape as preflight.
Dropping it also releases `ipc/plugin-marketplaces.ts`.
- `agent-hooks/wsl-hook-relay-launch.ts` and `speech/stt-service.ts` read `getAppPath`
and `isPackaged` through the AppEnvironment port.
The five that remain are all genuinely Chromium and need the HttpClient port or a
watcher split, not another mechanical swap: `browser/cdp-bridge` (webContents),
`ipc/filesystem-watcher` (ipcMain), `jira/authenticated-request` and
`network/proxy-settings` (net + session partitions), `speech/model-manager`
(`net.request`, which honors app proxy settings that Node https does not — replacing
it is a behaviour change, not a rename).
Verified: 219 files / 1,922 tests across plugins, speech, agent-hooks and the runtime
RPC methods; the built-artifact acceptance smoke passes; typecheck and
`oxlint --format github` clean.
* refactor(network): resolve the default proxy session lazily
Ratchet 5 → 4.
`proxy-settings.ts` needed exactly one Electron value: `session.defaultSession`, as
the fallback when a caller does not pass `options.proxySession`. Callers could already
inject a session; only the default was hard-wired. It now comes from a settable
resolver, so the module loads under plain Node.
**A resolver rather than a Session, because a Session eagerly throws.** The first
attempt installed `session.defaultSession` directly in pre-ready bootstrap and broke
startup outright — `TypeError: Session can only be received when app is ready`. The
acceptance smoke caught it before commit. Deferring to first use is always after ready.
Behaviour with no session is not a degradation: there is no Chromium proxy config to
discover, so `resolveProxy` is skipped and the environment variables become the whole
answer rather than a fallback. Applying rules to a session that does not exist is
likewise skipped; settings are still honoured because outbound requests read the env.
This reaches past Jira — a review round noted `ensureElectronProxyFromEnvironment` is
also on the Claude HTTP path via `oauth-refresh.ts` and `rate-limits/claude-fetcher.ts`.
Verified: 48 files / 526 tests across network, jira and rate-limits; the
built-artifact acceptance smoke passes; typecheck and `oxlint --format github` clean.
* fix(index): merge the duplicate proxy-settings import
CI's code-quality lint (`oxlint --config config/oxlint-code-quality-native-plugins.json
--deny-warnings`) flags a module imported twice in one file. My earlier insertion added
a second `./network/proxy-settings` import beside the existing one.
Verified with CI's exact invocation: exit 0.
* refactor(network): add the HttpClient port and lift BrowserError out of cdp-bridge
Ratchet 4 → 2.
Two unrelated couplings, both of the same shape — a small thing living inside a
Chromium-heavy file.
`BrowserError` is a seven-line error class with no dependencies, but it lived in
`browser/cdp-bridge.ts`, which imports `webContents`. The runtime catches that type on
paths with nothing to do with CDP, so one import kept a Node host from loading the
runtime at all. Moved to `browser/browser-error.ts`; cdp-bridge re-exports it.
`jira/authenticated-request.ts` fetches through `net.fetch` and reads
`session.defaultSession`. `network/http-client.ts` makes both settable. This one is a
**named port rather than a silent fallback, because the fallback is not transparent**:
Electron's net follows Chromium session/proxy state, avoids undici's stale keep-alive
sockets after a VPN path change, and sends a Chrome user agent that Jira's XSRF check
depends on. A Node host gets `globalThis.fetch`, reads proxy config from the
environment, and sends Node's user agent. That difference is documented at the port.
`session.defaultSession` is read per call, not captured at install — it throws before
the app is ready, which is the mistake the previous commit made and the acceptance
smoke caught.
Test wiring: `jira/client.test.ts` installs the port *inside* `loadClientModule`, after
its `vi.resetModules()`, since the reset gives the module a fresh singleton.
Verified: 461 files / 5,616 tests across jira, browser, network and runtime; the
built-artifact acceptance smoke passes; typecheck, `oxlint --format github` and the
code-quality lint with `--deny-warnings` all clean.
* fix(http-client): register the Node fetch fallback with the call-site audit
`global-fetch-call-site-audit.test.ts` guards every global-fetch use, because the
global runs on undici where an unread response body can crash the whole process
(orca#8695). The HttpClient port's Node fallback is a new such call site and was
unregistered — the guard caught it in a full-suite run.
Registered with the reasoning, and the port's doc comment now states the body-safety
contract explicitly: it hands the Response straight to its caller and never inspects
it, so the consume/cancel obligation stays exactly where it already was — with the
caller, unchanged from when they called Electron's net directly.
Two comments elsewhere mentioned the global by name and tripped the line scan as false
positives; reworded to describe the behaviour rather than name the API.
Verified: audit passes; typecheck and `oxlint --format github` clean.
* fix(app-environment): read hasAppEnvironment through the realm slot
|
||
|
|
0bbc6c80e8 |
refactor(host): route app paths and version through an AppEnvironment port (#16019)
* refactor(host): route app paths and version through an AppEnvironment port
`app.getPath('userData')` is the single largest Electron coupling in the main
process — 37 call sites — and it is one of the things stopping the Orca runtime
from booting on plain Node. Give it the same treatment as SecretStore.
- `src/shared/app-environment.ts` — the port plus a settable registry, covering
the members the runtime's module graph actually reads: paths, app path,
version, packaged flag, shutdown hook, exit, and Chromium process metrics.
`getAppEnvironment()` throws until installed, for the same reason the secret
store does: a silent default resolves `userData` to the wrong directory and the
caller writes real state there before anyone notices. No `node:` imports,
because `src/shared/**` is in the web build graph.
- `src/main/host/electron-app-environment.ts` — the desktop adapter, a
pass-through to `electron.app`.
- 9 modules migrated: telemetry, opencode/mimo/pi hook services,
terminal-history-paths, terminal-scrollback-snapshots, cli-installer,
clipboard-image-temp-file, memory/collector.
Deliberately NOT migrated: `src/main/browser/**`. That cluster is Chromium-
adjacent by nature — cookie jars, download destinations, offscreen pages — and a
Node backend does not ship it at all, so porting it buys nothing and churns
heavily-mocked suites. Also left alone for now: the call sites that additionally
touch `app.asar` path literals or `app.setName`, which need more than a
mechanical swap.
`getAppMetrics` stays on the port rather than being injected because
memory/collector.ts is its only caller and reads it from module scope; a Node
host returns [], having no Chromium processes to measure.
Test wiring: the secret-store setup file becomes `vitest-host-ports-setup.ts` and
installs both ports, exporting `fakeAppEnvironment`/`installFakeAppEnvironment`
so suites needing one specific member state only that instead of restating all
seven — which is boilerplate, and had pushed one suite past the max-lines budget.
Verified: 159 files / 1651 tests pass across every touched area; `tsc` clean on
both the node and web projects; `oxlint` clean.
* fix(typecheck): list the vitest host-ports setup in the node project
Three suites import `installFakeAppEnvironment` from config/scripts, but that
directory is outside tsconfig.node.json's include list, so composite typecheck
failed with TS6307. Listing the one file matches how this config already pins
individual files it needs.
Local `tsc --composite false` does not reproduce this — only `pnpm typecheck`
does, which is what CI runs.
* refactor(host): drop two unused AppEnvironment exports
hasAppEnvironment() and resetAppEnvironmentForTests() had zero callers. The
secret-store equivalents are used, so these were mirror-symmetry rather than
need; add them back when something actually needs them.
* test(terminal-history): install the AppEnvironment fake instead of mocking electron
These three suites mocked `electron.app.getPath` to point at a fixture dir. The
production module now reads the port, so the mock was inert and the global test
default's temp dir won — which broke the WSL path assertions and every deletion
count.
Found by a full-suite run, not by the targeted checks around the migrated modules,
which is the argument for running the whole suite on a refactor this wide.
* test(host-ports): remove the per-environment temp dir on teardown
The setup allocated a mkdtemp directory at module scope, which vitest evaluates
once per test *environment* — one per test file, not one per worker. Nothing
removed them, so a full 6,000-file run left thousands behind.
Proven: with an isolated TMPDIR, a three-file run previously added directories and
now leaves zero.
* fix(app-environment): anchor the installed environment to a realm global
Same reason as the SecretStore: vi.resetModules() rebuilds the module registry,
and an environment installed before the reset read back as uninstalled.
|
||
|
|
d07ce15cff | refactor(host): route secret storage through a SecretStore port (#15916) | ||
|
|
1ce2e562b3 | fix(skills): isolate concurrent upload staging (#15693) | ||
|
|
26bdfc0fe4 |
feat(agent-status): stamp observation provenance at every status ingress (STA-4293) (#14706)
Add an optional `observation` facet to agent status rows recording the origin (hook | osc | title | process | launch | orchestration), the authority that sequenced it, a per-pane incarnation, a monotonic revision, and the authority's own clock. Stamp it at every ingress; no consumer reads it. Boundary is stamped from the hook listener's existing per-provider `isNewTurnEvent`, not a second list of event-name literals. Identity-only (`providerSessionOnly`) rows are tagged `kind: 'identity-only'` so future consumers do not each rediscover that they are not turn transitions. The staleness-decay contract is documented at the type: staleness must be computed against the same authority clock that stamped `observedAt`, or replicas must decay on local receipt time. Not fixed here. Behavior-neutral: optional field on existing JSON, never persisted, never published to paired clients, and never inherited across writes. |
||
|
|
8e13485c9b |
fix(stats): count agent sessions from hook transitions, not OSC titles (STA-2445) (#14657)
* test(stats): dual-record the OSC-title detector against canonical hook transitions
Adds an AgentSessionTransitionRecorder that derives agent-session start/stop
boundaries from agent-hook status transitions, and a side-by-side comparison
that feeds both pipelines into a real StatsCollector.
Nothing is rewired yet — this commit only measures the delta:
title detector canonical
hook-only agent 0 1
braille-spinner non-agent TUI 1 0
one agent, one reconnect 2 1
totals 3 2
Refs #10201, STA-2445.
* fix(stats): count agent sessions from hook transitions and delete the title detector
Switches StatsCollector off AgentDetector and onto the canonical agent-hook
status stream, then removes the detector and its raw-PTY invocation.
- main/index.ts subscribes the recorder to subscribeEnrichedStatus and
subscribePaneStatusClear, next to where StatsCollector is constructed.
- orca-runtime.ts no longer feeds raw PTY bytes to a stats detector.
- StatsCollector keys sessions on a stable pane key, not a per-spawn ptyId.
Fixes #10201, refs STA-2445.
|
||
|
|
c40b0ab96b | fix(dev): stop macOS Keychain password prompts on pnpm dev (#15183) | ||
|
|
9a41119a99 |
feat(crash-reporting): read-only Windows install-dir DACL probe breadcrumb (#15107)
* feat(crash-reporting): read-only Windows install-dir DACL probe Records whether the install tree carries an orphan S-1-15-2-* package ACE with no S-1-15-2-1/-2 to satisfy it — the state that reproduces the 0x80000003 GPU/renderer init crash 10/10 (electron/electron#51761). Diagnostic only: never writes an ACL, never changes behavior. * fix(crash-reporting): evaluate the ACL signature per target and flag locale risk Three readiness-review findings: - the signature was merged across targets, so a grant on the directory masked its absence on the module file - the exact per-file state the probe exists to detect - the well-known package name check is English-only and icacls localizes it, so a non-English box could false-positive silently; report whether the check could be trusted - the serve-mode test omitted platform, so the gate was never exercised Also switch to the durable recorder (this runs after initObservability, so the span lands in the diagnostics bundle) and correct two comments that misstated where the probe runs. |
||
|
|
24e662adc1 |
feat(ssh): verify host keys, and restore panes correctly across a reconnect (#14844)
* docs(ssh): design for real host key verification (STA-4319)
Today's ssh2 verifier records a fingerprint and returns true — every host key is
accepted, with no known_hosts consult and no change detection anywhere in
src/main/ssh/. Scope is per-connection, so exec, SFTP, port forwarding, the
watcher and relay deploy all ride that one unverified handshake, and the
ProxyJump path puts the final hop — the topology most likely to cross untrusted
network — on ssh2 specifically.
Decisions worth calling out:
- Read the user's known_hosts as a trust source but NEVER write to it. That file
is shared with every other SSH tool on the machine; appending means line
endings, permissions, concurrent writers and a corruption blast radius well
beyond us. Accepted keys go to our own per-target store. Reading theirs is also
the entire migration story: most developers already have their hosts there.
- Mismatch is scoped to the SAME key type. A host with only an RSA entry that
presents ed25519 is unknown, not changed. ssh2 negotiates ed25519 first, so
without this we would fire a change-of-key alarm at nearly every existing user
on their first upgraded connect — training them to dismiss the one warning that
is supposed to mean something. Flagged in review as the decision I am least
sure of; a downgrade-vector argument against it is being tested.
- Changed key hard-fails with no override button; recovery is a separate explicit
action, offered only when OUR store is what disagreed, because forgetting our
record cannot unblock a known_hosts conflict.
- Background reconnects deny rather than prompt. A dialog the user cannot place
in context only teaches click-through.
Two traps are documented because either would make the fix silently do nothing:
an async verifier returns a Promise, which ssh2 reads as truthy and accepts
immediately; and the existing test mock invokes hostVerifier with one argument
and ignores the return, so it would pass against a verifier that never decides.
Design only — no behaviour change. The doc is added to the tracked-reference
allowlist in .gitignore alongside the other docs/reference entries.
* docs(ssh): revise the host key design after security and migration review
Three things the reviews changed, kept visible rather than quietly edited out.
THREAT MODEL WAS WRONG IN THREE PLACES. Jump hosts are not the worst case — they
are already safe: shouldUseSystemSshTransport branches on exactly the inputs
resolveEffectiveProxy does, and attemptConnect returns after the system probe, so
ProxyJump goes through OpenSSH and is verified. Agent forwarding was overstated
(gated on the user's ForwardAgent). Credential theft was understated: any auth
error counts as agent fallback, so a MITM walks the user to the password AND
private-key passphrase prompts, and cachedPassword replays without prompting. The
relay claim was backwards — the attacker owns their own machine; the real impact
is the return direction, where they become the host our workspace trusts.
TYPE SCOPING IS A DOWNGRADE VECTOR WITHOUT ALGORITHM ORDERING. This was the
decision I flagged as least certain and asked to have argued both ways. OpenSSH
is safe only because order_hostkeyalgs() puts known types first and RFC 4253
gives the client's order priority. ssh2 negotiates ed25519 first regardless, so
an attacker who cannot forge the RSA key on file just presents ed25519 and gets a
friendly first-contact prompt instead of a hard failure. Keep scoping, but set
algorithms.serverHostKey to lead with the types on file — and add a sixth
outcome for 'unknown type, known host', which must never read as first contact.
SHIP THE DEFENCE BEFORE THE DIALOG. Startup restore fires eager connects for all
targets in parallel with a 15s timeout while a prompt would live 120s; ephemeral
VM targets present a new key every launch; paired-web connects run on the host
desktop, so the dialog opens on someone else's screen. Phase 1 is therefore no
modal at all: consult known_hosts and our store, match connects, unknown persists
with accept-new semantics, mismatch and revoked hard-fail. That is the whole MITM
defence with none of the migration risk.
Also folded in, verified live against OpenSSH 10.2p1: the without-port fallback
(bracketed lookup first, then bare, where the second pass can only yield match or
unknown — otherwise a bare line plus a non-default port produces a spurious
prompt); hashed entries hash the candidate form; multiple files union; a
cert-authority line does not match a plain key. IPv6 and bracket parsing moved
INTO scope — that is a parser requirement, not a scope call, and getting it wrong
produces the prompt-training harm the design exists to avoid.
* feat(ssh): parse and match OpenSSH known_hosts
The matcher half of STA-4319. No behaviour change yet — nothing calls this.
Hand-rolled because no maintained JS implementation exists, and written against
behaviour observed from OpenSSH 10.2p1 rather than inferred from the man page.
Three of those behaviours a reasonable reading gets wrong:
- A non-default port is TWO ordered lookups, not one candidate set: '[host]:port'
first, then bare host ('checking without port identifier' in ssh -v). The
fallback pass can only yield match or unknown — OpenSSH downgrades a wrong key
there rather than reporting a change. Collapse them and anyone holding a bare
line who connects off-port gets a spurious first-contact result; treat the
fallback as authoritative and they get a false change-of-key alarm.
- Revocation resolves in its own pass so the verdict cannot depend on line order.
Verified both orderings.
- A cert-authority line never matches a plain host key; it only validates
certificates. A normal line alongside it still decides.
Mismatch is scoped to the same key type, and a host known by a DIFFERENT type
returns unknown-type-known-host rather than plain unknown — an attacker who
cannot forge the key on file must not get a friendly first-contact result by
presenting another type. That outcome is only half the defence; the other half
(leading serverHostKey with known types) lands with the wiring.
47 tests from vectors executed against real sshd, including ssh-keygen -H hashed
entries. Each of six mutations reddens it: collapsing the passes, letting the
fallback report mismatch, dropping type scoping, resolving revocation in line
order, honouring an unrecognised marker, and skipping the blob/type agreement
check.
* feat(ssh): decide what to do with a presented host key
The policy half of STA-4319, kept separate from the ssh2 wiring so it is testable
without a handshake and injected rather than importing its sources, so a test
states its own trust state instead of writing files.
Phase 1 ships no dialog — a test asserts the decision is never 'prompt'. Startup
restore opens every previously-active target at once, ephemeral VM targets would
ask every launch, and paired-web connects run on the host desktop where the
dialog would appear on someone else's screen.
Ordering that matters: revocation outranks everything including
StrictHostKeyChecking=no, because a revoked key is a statement that this key is
known-bad rather than merely unrecognised. known_hosts is named before our own
store on a change, because its remedy (ssh-keygen -R) is the one that also
unblocks ssh and git — pointing at a remedy that cannot work is worse than none.
Two carve-outs with reasons: an ephemeral runtime target accepts WITHOUT
recording, since a fresh VM presents a new key every launch and a stored record
would accumulate per launch and eventually read as a spurious change; and when
ssh -G ran on the HOME-divergent path that suppresses /etc/ssh/ssh_config, an
unknown host is denied, because a site-wide policy may forbid it and being laxer
than ssh is the one outcome that is never acceptable.
Rejection text deliberately avoids 'authentication failed' and 'permission
denied': the reconnect ladder classifies on those substrings, so a denial phrased
that way is retried forever against a decision that will never change. Pinned by
a test.
* feat(ssh): build the host key verifier and the algorithm order that makes it safe
Still not wired into the handshake — that lands next. This is the piece that
turns a decision into an ssh2 callback, plus the half of the design that is easy
to forget because it lives in a different config field.
The verifier MUST be a plain function returning undefined. ssh2 does
'const ret = verifier(key, verify); if (ret !== undefined) verify(ret)', so an
async function returns a Promise — neither undefined nor falsy — and ssh2 accepts
the key immediately while ignoring whatever the callback later decides. Making
this async would silently restore exactly the accept-everything behaviour the
module exists to remove, so a test asserts the return value is undefined.
orderServerHostKeyAlgorithms is what makes type-scoped matching safe rather than
a downgrade. RFC 4253 gives the client's algorithm order priority, so leading
with the types we already hold for a host denies a server the choice of
presenting some other type to convert a hard failure into first contact. Without
it, an attacker who cannot forge the key on file just offers a different
algorithm. Revoked entries never contribute to that order.
Also fails closed on two paths that would otherwise hang or over-trust: a key
whose own length-prefixed header cannot be read is refused rather than reasoned
about, and a throw from any dependency denies, because ssh2 may not catch an
exception raised inside the verifier and the handshake would hang instead of
failing.
18 tests. Includes the two negative cases that matter — first-contact keys are
recorded, but keys we already know, rejected keys, ephemeral runtime targets and
a lax StrictHostKeyChecking are not.
* fix(ssh): promote every RSA signature algorithm for a known ssh-rsa key
A known_hosts entry names the KEY type, which is not the negotiated ALGORITHM
name. One ssh-rsa key is offered as rsa-sha2-512, rsa-sha2-256 or ssh-rsa
depending on the signature algorithm, so matching the literal name only would
leave a host we know by RSA ordered behind ed25519 — precisely the ordering this
function exists to prevent, and precisely the population (RSA-era known_hosts
entries) it was written for.
Verified from ssh2's own negotiation while wiring this: kex.js iterates the
CLIENT list and takes the first entry the server also offers, so client order
does decide, as RFC 4253 says. ssh2's default order leads with ed25519 and places
the RSA algorithms fifth through seventh.
* fix(ssh): verify host keys instead of accepting every one (STA-4319)
The actual fix. ssh-connection's verifier recorded a fingerprint and returned
true, so every ssh2 connection accepted every host key — no known_hosts consult,
no change detection. It now consults the user's known_hosts plus our own store
and refuses a changed, revoked or unverifiable key.
Phase 1 by design: no dialog. Unknown hosts are accepted and recorded
(accept-new semantics), because startup restore opens every previously-active
target at once, ephemeral VM targets present a new key each launch, and
paired-web connects run on the host desktop where a prompt would appear on
someone else's screen. The MITM defence lands now; the prompt is Phase 2.
Also sets algorithms.serverHostKey to lead with the types already known for the
host. Without it the type-scoped matching is a downgrade — an attacker who cannot
forge the key on file just presents another type and turns a hard failure into
first contact. Verified from ssh2's kex.js that the client list decides.
Denial replaces ssh2's generic handshake error with the specific reason, because
the reconnect ladder cannot distinguish a generic failure from a transient fault
and would retry forever against a decision that will never change.
An unreadable trust store degrades to known_hosts only rather than failing the
connect: a changed key is still refused, and a host trusted only by us falls back
to first contact and is re-recorded, reaching the same decision.
The ssh2 mock now uses the callback form and aborts the handshake on denial. As
written it called hostVerifier(key) with one argument and ignored the result, so
it would have passed against a verifier that never decides — flagged in the
design as a mock that had to change, not a test to quietly rewrite. Two new tests
pin the wiring rather than the module: an unidentifiable blob is refused, and a
well-formed key is accepted.
Note for review: commit
|
||
|
|
619ee2cc90 | fix(agent-hooks): detect IDS-truncated hook POSTs instead of failing open silently (STA-2870) (#14625) | ||
|
|
5652fb7469 |
fix(codex): stop resuming a session under the wrong account when a sessions tree is locked (#15093)
* fix(codex): stop resuming a session under the wrong account when a sessions tree is locked Two probes reported "this rollout is not bridged here" for any filesystem error, not just a genuine absence: - codex-session-resume-home.ts used existsSync on each ranked home's sessions directory. existsSync returns false on EBUSY/EPERM, so a briefly locked tree made the scan continue to the next ranked home — and the winning home becomes the resumed pane's CODEX_HOME, so it picks the account. - codex-legacy-session-resume.ts caught every lstat failure for the selected account's candidate rollout and returned null, after which the caller kept the source per-account home. Either way the session resumed under a different account's credentials while the UI still showed the selected one. Only a definitive ENOENT/ENOTDIR now means "not bridged here". Any other error raises the typed temporary-unavailability refusal the ownership gate already uses, which both PTY paths convert into a clean abort before spawn. The refusal is scoped to the SELECTED account's home. An unreadable home that is not the selected account cannot cause a wrong-account resume, so it is still skipped rather than stranding the user. These are pre-existing and independent of the STA-4422 ownership-marker failure: they route to another account today with the gate uninvolved. Fixes STA-4607 * fix(codex): refuse a resume when the selected sessions tree is locked mid-listing Review found the first pass incomplete in two places, both the same category error one layer further down. The preliminary statSync on the selected sessions root was guarded, but the real directory read happens later and listCodexSessionRolloutFilesIncrementally swallows every opendir error. A lock held during enumeration — where nearly all the I/O is, and so the far more likely case — still yielded nothing for the selected account and fell through to another one. The listing now reports directory errors through its existing onDirectoryError hook, and a non-definitive error anywhere under the selected sessions root raises the typed refusal. Non-selected homes and definitive absence still skip. Separately, index.ts wrapped prepareLegacySharedCodexSessionResume in a blanket catch and fell back to the source home, so the typed refusal from the candidate lstat was swallowed and the resume still ran under the peer account's credentials. That catch now rethrows ManagedCodexHomeTemporarilyUnavailableError while ordinary migration failures keep warning and falling back, since a genuine migration failure legitimately should not block a resume. A typed refusal is only as strong as the narrowest catch between the throw and the spawn. The frames between both throw sites and the PTY spawn were audited: findTrustedCodexSessionResume, resolveCodexSessionResumeProvenance and prepareCodexSessionResume have no catches, and the PTY layer already maps the typed error before spawn. Both fixes are mutation-checked. Disabling the listing hook makes the resume resolve to the other account again; the index.ts rethrow is covered only by typecheck, because src/main/index.ts has no unit-test entry point in this repo. * test(codex): pin nested-directory lock coverage; document the resume repin contract Review flagged the listing guard as matching only the exact sessions root, so a nested dated directory would leak. It does not — the guard keys on the root being listed, not the failing directory — but nothing pinned that. Added a test that faults only sessions/2026/07/20 while the root stats fine; it fails under mutation alongside the root case. Also documented why the index.ts rethrow cannot fire today. That launch path pins CODEX_HOME to the account that owns the rollout and deliberately refuses to repin onto whichever account is selected now (#10793), so it does not wire the selected-home resolver. The branch stays as a contract guard so the blanket catch below can never silently swallow a typed refusal if that changes. |