mirror of
https://github.com/stablyai/orca.git
synced 2026-10-03 08:02:12 +00:00
fix/execution-host-resolution
1877
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
642607bfa7 |
feat(diagnostics): name the code driving a React commit cascade (#16730)
* feat(diagnostics): name the code driving a React commit cascade React #185 reports blame whichever component dispatched after the root-global counter tripped. react-update-depth-attribution already tells the report that boundary_id names a bystander; nothing recorded what the real driver was. Count commits through react-dom's devtools commit hook — the only per-commit seam that survives minification. Profiler's onRender is compiled out of the production bundle, and a dependency-less root layout effect fires per render of its own component, not per commit (measured: a root effect saw 1 of 11 commits a leaf drove). Mirror React's own reset rule rather than a time window: a commit that leaves no sync lanes pending ends the cascade, and a different root restarts it. The steady-state cost is a mask, a compare and an increment, with no clock read and no allocation. Stack sampling arms only once a cascade is already deep, so ordinary work never pays for it. * fix(diagnostics): remove the install-order trap and guard the write path Adversarial and perf review of the cascade diagnostic: The install-order ratchet guarded the wrong thing. The observer self-installs at the bottom of its own module, so it only ran after its transitive graph evaluated — one new import reaching react-dom would have killed the diagnostic in production with every test green. The entries now import the import-free shim instead, which only has to make the global exist; wrapping the callback is timing-independent because react-dom re-reads it per commit. The store write probe called the sampler unguarded, so a throw there dropped the write on the app's universal write path. Guarded; the try/catch measured free at +0.005ns. Report the frames that name the driver instead of capturing eight and reporting one, arm the self-check on the paths where install fails, bind the sample cap to the write count rather than a V8-only API, and stop defining the devtools global for every test file to serve one. The cascadeRoot comment claimed a strong reference cannot retain; a WeakRef probe disproved it. It is still not a leak — the next non-cascading commit clears the slot — so the comment now says that instead. * test(diagnostics): close the ratchet holes guarding the cascade hook Adversarial review loop 2: The install-order ratchet only saw imports whose `from` shared a line with the keyword, so a multi-line `import { createRoot } from 'react-dom/client'` in the shim passed it — and that is the one edit that kills the diagnostic in production. 43% of files in this directory use the multi-line form. Scan the shim source directly as well as walking the graph. The 4000-char budget for the driver frames is bought by the key ending in `stack`, but the only test asserting that emitted its own literal key, so renaming the real one truncated the frames with the suite green. Assert the name the renderer actually emits. Also correct the comment on the `installed` placement: the self-check never reads that flag, it arms because it sits outside the try. * test(diagnostics): stop the shim ratchet firing on prose Adversarial review loop 3 caught two flaws in the guards added last commit. The source-scan regex used an unbounded `[\s\S]*?` after an anchor that also matched the shim's own `export type`, so it degenerated to "does the word `from` appear later in the file" — rewriting a doc comment to say "reads the hook from the global" failed the ratchet. A guard that fails on prose is a guard someone deletes, and this one is what stands between a reshuffled import and a silently dead diagnostic. Require a quote after `from`, tolerate comment obfuscation, and catch `await import(...)`, which makes the shim async so react-dom evaluates before the hook is installed. The 4000-char budget assertion matched `/stack$/i` against the raw key, but the real rule camel-splits first — so `driverstack` would pass while shipping truncated frames. Assert through sanitizeCrashReportDetails, resolving the key from the payload rather than hard-coding it. |
||
|
|
92536346bd |
Add terminal unavailability exports to parity test (#16810)
New runtime exports for handling terminal unavailability: RuntimeTerminalUnavailableReason type and related error codes and messaging constants. |
||
|
|
5631aa00dd |
feat(orcad): items 2–7 — degradation, natives, daemon, ops, deploy (#16398)
* fix(ports): stop joining an undefined resourcesPath on a non-Electron host `resolveWorkerEntryPath` branched on `isPackaged` alone and joined `process.resourcesPath`. orcad reports `isPackaged` true — correctly, it is a production build, and ~15 consumers read it that way to gate HTTPS-only skill downloads and the real CLI name — but `process.resourcesPath` is Electron-only and `undefined` under plain Node. So the packaged branch threw `TypeError [ERR_INVALID_ARG_TYPE]: The "path" argument must be of type string` where a clean "worker unavailable" was the honest outcome. The type said `resourcesPath: string`, which is how it went unnoticed; it is now `string | undefined`, so the compiler carries the fact. A host with no Electron resources tree has no asar to look in, so it falls back to the module directory and lets the caller report a missing worker. Found by the item 1 agent while auditing the same `isPackaged` defect class in the watcher. Verified in both directions: reverting the guard reproduces the TypeError. * feat(orcad): prove node-pty loads before anything requires it Of the two ways node-pty fails, only one is catchable. A missing module throws MODULE_NOT_FOUND. A module built against the wrong libc or Node ABI is refused by the dynamic loader, and in the worst case takes the process down before any handler exists — that is #9902, which crashed the desktop app on Ubuntu 20.04 before a window appeared. There was no libc or ABI precondition anywhere in the tree. So orcad now proves the load in a CHILD process, from main.ts, before anything requires node-pty. Whatever the child does — throw, abort, die on a signal — is data rather than our own death, and the operator gets a sentence naming the host's libc, Node ABI and prebuild slot plus the command to run. Proven-unloadable exits 78 (EX_CONFIG), so a supervisor does not restart an unequippable host forever. A probe that never answered is unverifiable, not blocked: refusing to boot on an inconclusive signal would take down hosts that work. The child dlopens the file node-pty would have chosen, before requiring the package. node-pty's loader walks several directories and rethrows only the LAST error, so a refused binary reads as "Cannot find module ./prebuilds/..." — which sends the operator to install a module that is already there. It also reports through stdout: node echoes the whole -e source above a stack trace, and matching tokens against stderr made the probe's own source text answer for the verdict. Verdicts reach clients as a terminal_unavailable degradation alongside the existing browser_unavailable one, through the same cause-registry shape. degradations[].code is now an open vocabulary; clients already render only `message`. Prebuilds are compiled from PATCHED sources — the patch IS the glibc-floor fix, so an upstream tarball reproduces #9902 — into linux-{x64,arm64}-{glibc,musl} and darwin-{x64,arm64} slots. libc is in the slot name because node-pty's loader falls back to prebuilds/<platform>-<arch> and cannot tell glibc from musl. orcad installs the matching slot at boot, so a host with no compiler serves terminals. The relay's five pure toolchain-diagnosis functions moved to a transport-free module so the Node bundle can reuse them without dragging ssh2 in behind them; the relay keeps its API by re-export. macOS gets `xcode-select --install` rather than the cross-distro apt/dnf/pacman/apk menu, every line of which is wrong there. * test(orcad): pin the node-pty precondition to ground truth, not a prepared host CI's test shard runs `vitest` directly, so `ensure-native-runtime --runtime=node` never prepares node-pty for the Node ABI — `degraded` is the correct verdict there, and asserting 'ok' encoded an environment the shard does not have. Asserting whatever it returned would be vacuous, so the expectation is now derived from an independent require() of node-pty. Verified it still bites: forcing the precondition to always report 'ok' fails the suite. * feat(orcad): run the terminal daemon, and the ops contract around it orcad declared `canRecoverPersistentLocalPtys: () => false` because it did not run the terminal daemon, so every restart, update and rollback SIGKILLed every running terminal — on the host whose selling point is that work survives the client going away. That is the one property `ssh-execution-boundary.md` recommends the peer model for. Item 4 — the daemon: - Port the launch path off electron: `daemon-init.ts`, `daemon-host-relocation.ts` and `observability/logs-directory.ts` now read the `AppEnvironment` port. Relocation additionally asks whether the app root is an asar archive rather than whether the build is packaged, so a Node host answering `isPackaged() === true` no longer walks into an Electron-only NSIS-escape path (same precedent as `parcel-watcher-entry-path.ts`). - `build-orcad.mjs` emits `daemon-entry.js` beside `orcad.js`, scans the forked children's metafiles for electron/node:sqlite, and load-checks the child under plain Node. - orcad spawns and adopts the daemon; shutdown disconnects and never kills it. `canRecoverPersistentLocalPtys` now reads the live provider and is false under degraded routing, where fresh terminals would die with the process. Item 3 — the ops contract (docs/reference/orcad-operations.md): - Bind policy: `--bind`, default loopback, pinned so neither `orca serve`'s wide default nor the connected-device widen can override it, and so a paired client cannot rebind the listener from outside. - Instance lock on the data root before profile load, scoped to the runtime role so it never refuses a restart that a live daemon makes worthwhile. - Supervision: exit codes a supervisor can act on (78 = do not retry), second-signal escalation, a shutdown deadline, and crash-loop containment on daemon respawn. - Health in the readiness payload: build hash, Node ABI, and a PTY self-test that spans both processes — the daemon spawns a real PTY in its own process and the verdict crosses its socket. Both bundle load-checks now assert on exit codes: these bundles are minified onto one line, so Node's uncaught-exception report echoes every string literal in the bundle and the previous message match passed against a bundle that never loaded. * feat(orcad): deploy, activate and roll back a versioned orcad install Plan items 6 and 7 from docs/design/shipping-orcad.html. Install reuses the relay's transaction verbatim — per-version lock, staged SFTP write, .install-complete sentinel, stale-lock recovery — under a parameterized namespace, so orcad-<v>/ sits beside relay-<v>/ permanently (§06). Parameterizing GC is the trap that creates: each model now collects only its own directories, enforced twice (prefix-scoped remote listing plus a local ownership re-check), and a client picks its model from how the host is registered, never from what it finds on disk. Activation is separate from installation, because a versioned directory selects nothing. A candidate is launched, publishes orca_server_ready, and only becomes active if its cross-process health payload passes: right build hash, listening, daemon live, PTY self-test green. A rejected candidate is stopped and the incumbent restarted, so a careful deploy cannot cause the outage it was being careful about. Update and rollback are shaped by the daemon. An update restarts orcad, the daemon outlives it, and the surviving daemon was forked from the outgoing bundle — so live terminals defer the update rather than proceed, and GC pins the active version, the rollback target and the live daemon's bundle. Orca's persisted state carries no schema version, so rollback restores a pre-activation snapshot rather than trusting backward-readability; the point past which it is unsafe is the first terminal created after activation, which the snapshot cannot describe and the surviving daemon still owns. Running the generated shell for real found two bugs the text assertions missed: tar members re-quoted inside a shell variable captured nothing, and kill -0 reports a zombie as alive. * test(orcad): assert the precondition is self-consistent, not environment-shaped The real-host case cannot predict a status: CI's shard runs vitest directly, so node-pty is never built for the Node ABI and 'degraded' is correct there, while a prepared checkout gives 'ok'. The previous attempt used require('node-pty') as ground truth, which resolves the JS wrapper while the native binding loads lazily — it proved strictly less than the precondition checks, and failed CI for exactly that reason. What is invariant on a host with node-pty installed: never 'blocked', and never a degraded verdict carrying an unestablished reason. The injected-input tests keep the logic coverage. * fix(orcad): drop an eslint-disable the rule no longer needs * test(orcad): separate slot placement from the load verdict Both remaining CI failures were the same shape: tests reaching into node_modules for a pty.node that only exists after `ensure-native-runtime --runtime=node`, which CI's shard never runs because it invokes vitest directly. Slot *placement* is the logic worth checking on every host, so it now uses a synthetic payload and asserts the verdict stays honest about not loading. The three assertions that genuinely need a Node-ABI binding are gated on it existing. Verified: breaking slot installation fails both placement tests; with the real pty.node hidden the file is 17 passed / 3 skipped instead of ENOENT. * test(orcad): gate the load-dependent cases on a real load, not on the file existing CI ships a pty.node built for Electron's ABI, so existsSync was true while require still failed — the gate ran exactly the tests that host can never satisfy. It now probes the binding in a child process, so a bad one cannot take the runner down. The self-consistency assertion also allowed too little: 'blocked' is the honest verdict for a corrupt binding, alongside 'ok' on a prepared host and 'degraded' on an unprepared one. What stays invariant is that anything other than 'ok' names an established cause, so a terminal is never declined for a reason nobody worked out. Verified against all three host states: prepared (19 passed), unprepared, and a corrupt binding (17 passed / 3 skipped, no failures). * test(orcad): gate on the whole premise — binding AND spawn-helper CI has a loadable pty.node but no spawn-helper, and a slot without the helper is legitimately 'degraded'. So the previous gate let a test run whose premise ('a complete slot yields ok') that host cannot satisfy. Verified in both states: with the helper present 19 pass; with it removed the load-dependent cases skip (17 passed / 3 skipped) instead of failing. * fix(orcad): preserve degradation types after rebase |
||
|
|
de6fe8b7ea |
fix(worktrees): resolve id: worktree selectors by path equivalence (#16243) (#16494)
* test(worktrees): cover id: selector path-spelling parity with path: (#16243) The renderer can only address a workspace by id (toRuntimeWorktreeSelector always emits id:<repoId>::<path>), and the runtime matches that id byte for byte while a path: selector has always compared through normalizeRuntimePathForComparison. A stored id that spells its path differently from `git worktree list` therefore resolves for the CLI and answers selector_not_found for the UI, which reads that as a stale local mirror, calls forgetLocal, reports success, and lets the row return on the next catalog refresh: a silent delete. These tests fail on both resolution sites -- the fleet `id:` branch of resolveWorktreeSelector and the scoped resolveScopedWorktreeIdRow a host-qualified removal takes -- and pin what must stay closed: an exact repo id (STA-4343), host qualification, dot segments neither selector canonicalizes, and a refusal rather than a guess when two rows spell one path. 13 failing, 35 passing. * fix(worktrees): resolve id: worktree selectors by path equivalence (#16243) worktreeIdComparisonKey names one repo, one filesystem location, and one folder-workspace instance, folding exactly the path spellings normalizeRuntimePathForComparison already folds for a path: selector -- and nothing more, so dot segments stay unresolved for both shapes. Both id: resolution sites consult it only after an exact match finds nothing: the fleet branch of resolveWorktreeSelector and resolveScopedWorktreeIdRow, which a host-qualified removal takes. runtimeWorktreeIdsEqual now derives from the same key so the runtime has one normalizer rather than a parallel one. Not a pure refactor at that last site: runtimeWorktreeIdsEqual used to normalize-compare ids that parse but carry an empty repoId or an empty path ('::/p', or 'repo::' against 'repo::/'), and worktreeIdComparisonKey returns null for those, so across its call sites (PTY identity, refresh, mutation queue) such ids now compare byte-exact instead. That narrows matching rather than widening it, no real worktree carries such an id, and it is the behavior #15616 guarantees for malformed ids -- but it is a behavior delta, not just a tidy-up. Perf (#14399): the exact match is still tried first and still wins outright, so a resolvable id costs exactly what it did before. Neither site adds a scan -- the fleet branch re-filters the array it had already listed, the scoped lookup re-filters the single owning repo's projected rows -- so an explicit id still never scans every repo. Fail-closed behavior is unchanged: the repo id compares exactly (STA-4343), host qualification is untouched, the folder-workspace instance suffix stays part of the path, and a scoped lookup with two equivalent rows refuses instead of guessing. The bare unprefixed selector branch keeps byte-exact id matching, since only the id: shape reaches a renderer caller. Shares src/shared/worktree/id.ts with the open #15616, which introduces worktreeIdComparisonKey for the same divergence in lineage pruning and authoritative-scan purging; this adopts that helper rather than adding a second one. Complementary to the open #16295, which makes the miss visible; this removes the miss. * chore(worktrees): satisfy oxfmt and oxlint on #16243 tests oxfmt --check flagged both new test files and oxlint's unicorn/no-useless-fallback-in-spread flagged the store mock; the full lint and format gates now match the pre-change baseline. * test(worktrees): pin Windows spellings and malformed-id exactness (#16243) Review found two axes the first pass left unproven at the two id: resolution sites. Both are the invariants the open #15616 guarantees for the shared worktreeIdComparisonKey it introduces for #15598, so violating either here would break a contract a sibling PR depends on. Windows: #15598's whole defect is that one checkout is recorded under both `D:\Agentic\game2` and `D:/Agentic/game2`. The fleet branch, the scoped removal lookup, and the key itself now each resolve the backslash spelling against the forward-slash spelling git reports, and fold drive-letter case -- while a backslash inside a POSIX path stays a filename character and a POSIX root stays case-sensitive, exactly as normalizeRuntimePathForComparison already decides for a path: selector. Malformed ids keep exact matching at both sites: an id with no repo boundary or an empty path still refuses, and the scoped lookup still refuses it without scanning. Four of these fail without the production change (three fleet/removal Windows cases and the scoped one); the malformed-id and POSIX-backslash cases are invariant guards that hold either way. Verified: 58 passed in the three files; 18 fail with the production hunks reverted; orca-runtime.test.ts and worktree-teardown-unstopped-pty.test.ts green (1270 passed | 1 skipped); pnpm tc:node clean. * test(worktrees): pin Windows id: spelling folds and fleet ambiguity refusal (#16243) The Windows backslash spelling now rides the ID_SPELLINGS rows, so it is driven through both id: sites -- resolveWorktreeSelector and the scoped removal target -- and compared against what the same workspace's path: selector resolves, rather than only through worktreeIdComparisonKey. That is the spelling #15598/#15616 found in the wild and the one the owner's Windows client produces. The fleet path's ambiguity refusal had no test: two same-repo rows spelling one directory, an id: matching neither exactly, must reject selector_ambiguous. It is the fail-closed guard on a delete-capable resolver, and the property a later refactor is most likely to turn into a silent pick. Also records two limits at the source instead of leaving them to be rediscovered: a UNC or WSL root never folds into a drive-letter location (while Windows' two WSL UNC aliases do name one location), and a folder-workspace id keeps a trailing slash placed before the ::workspace:<uuid> suffix, so that spelling stays exact-match-only. Neither behavior changes here. The file docblock overclaimed parity. path: collapses duplicate same-host registrations to the first row while a folded id: refuses them; the contract this file pins is path-spelling parity, not dedup parity, and the divergence is deliberate because this resolver also serves delete. Non-vacuity, verified by temporarily reverting the production hunks: neutralizing both id: fallbacks turns 10 of these tests red, including both new Windows rows and the ambiguity refusal (it degrades to selector_not_found). Making the fleet fallback pick the first folded match instead of collecting all of them turns the ambiguity test red on its own. The remaining cases -- malformed ids, dot segments, the POSIX backslash, the folder-workspace slash -- pass against the pre-fix code too: they guard against future widening rather than proving this fix. Drops the two Windows cases the ID_SPELLINGS row subsumes. The drive-letter case test asserted only the Windows half its name promised; it now also pins that a POSIX root does NOT fold case, since an unconditional lowercase would merge /data/Foo with /data/foo on the platform CI runs on. Fixture paths use the upstream-attested /srv/projects prefix (and a neutral plugin-host leaf) instead of a local install root; the spelling variations the tests exist to pin -- doubled separator, dot segment, trailing slash, uppercase POSIX, cafe NFC/NFD, and the Windows D: rows -- are unchanged in form. * docs(worktrees): trim the id: selector test header and document the comparison key (#16243) |
||
|
|
aaef5e8c9f |
fix(agent-hooks): deliver hook events that fire while Orca is restarting (STA-5329) (#16685)
* fix(agent-hooks): correct durable spool delivery * fix(agent-hooks): spool curl failures after retries * fix(agent-hooks): keep replay out of runtime observations * test(agent-hooks): pin managed hooks inert outside an Orca terminal * fix(agent-hooks): address review findings on the durable spool - claude: pass the literal source; options.agent does not exist (typecheck) - kimi: the windows-local ordering runs its guard pre-stdin and before the function exists, so it no longer spools there (printed command-not-found) - writer: require a readable endpoint file before creating a spool tree - antigravity: carry its out-of-band event name into the record and filter on it - drain: truncate only the bytes consumed, preserving concurrent appends and a torn trailing line * fix(agent-hooks): ignore spool events without pane attribution * fix(agent-hooks): make spool replay and appends robust * test(agent-hooks): type spool replay records * fix(agent-hooks): defer unterminated spool records * fix(agent-hooks): replay spool events through relays * fix(agent-hooks): preserve Codex prompt across child replay * fix(relay): keep startup alive when spool replay fails * fix(relay): simplify spool replay startup guard |
||
|
|
9c01e09ecc |
Revert "fix(codex): launch WSL accounts from direct homes" and "refactor(codex): remove WSL runtime mirror machinery" (#16722)
This reverts commit |
||
|
|
762fb05caa |
fix(agents): one-shot submit-retry Enter for codex (STA-5379) (#16689)
Codex silently discards Enter for ~75-150ms after its composer glyph first renders, and the boundary widens with prompt size and machine load, so no fixed first-Enter delay is provably safe on slow hosts. Submit success is not verifiable from PTY output, but a redundant Enter is a measured no-op on codex in both post-submit states, so send one blind retry after the first Enter. - tui-agent-config: new optional submitRetryDelayMs knob, set to 1200 on codex only; every other agent is byte-identical to today. - agent-paste-draft: after the post-paste '\r', wait the configured gap and send exactly one more '\r' inside the same PTY input transaction, so a concurrent paste cannot interleave. The retry is best-effort and never downgrades the first Enter's result. - Retry tests live in a new file to keep agent-paste-draft.test.ts under the max-lines budget. active-agent-note-send is deliberately exempt: its Enter rides the terminal.send RPC (different transport, server-side sendable guard), has no local agent identity to read the config from, and only fires on an already-running agent, where the codex cold-boot submit gate cannot occur. |
||
|
|
0f522c35e5 | fix(remote): gate empty session inventory on host authority (#16546) | ||
|
|
673842db35 |
refactor(codex): remove WSL runtime mirror machinery (#16505)
* refactor(codex): remove WSL runtime mirror machinery * test(codex): drop allowlist entries the mirror removal made stale runtime-home-service.ts no longer spawns wsl.exe or imports child_process; both boundary guards fail closed on a stale entry so the goalpost keeps moving. * fix(codex): drain legacy WSL auth before restart * fix(codex): await WSL auth drain before restart |
||
|
|
d60a3c900b | Reset stale terminal modes after dead TUI replay (#16379) | ||
|
|
26721bd632 |
fix(codex): stop blocking the main thread on trust grants (#16441) (#16594)
* fix(codex): stop blocking the main thread on trust grants (#16441) Codex hook trust was granted by blocking the Electron main thread on `spawnSync` of a bundled ELECTRON_RUN_AS_NODE entry for the whole app-server deadline: 15s native, 35s WSL, ~45s on the real-home path (rebase inspect + repair + grant). Cold start and every Codex pane launch showed "Not Responding"; the reported event-loop gap was 15,049 ms. The subprocess only ever existed to donate an event loop to a deliberately blocked parent — `runCodexHookTrustGrantSession` was already the real async implementation. Make the callers async and the fork is unnecessary, so the bridge, the forked entry and its envelope are deleted along with their build/knip/tsconfig registrations. The CLI `agent hooks prepare-codex` handler is already async, so it awaits the in-process session and saves a process spawn per managed-home shell. `resolveCodexTrustGrantHost` is async too; the WSL identity probe moves from `execFileSync` to `runProcess`, dropping that file from the child-process import allowlist. Status reads keep a synchronous native-only stamp path. Two invariants that held only because the lane blocked: - Overlapping capability probes were impossible by construction. `GitCapabilityCache`'s dedupe engine is extracted to a shared `CapabilityProbeCache` and `CodexAppServerCapabilityCache` now inherits it, so concurrent launches against a cold host share one app-server session instead of one each. - Two grants on one `config.toml` could not interleave capture and restore. A reentrant per-file lane now serializes the whole install sequence (managed, WSL runtime, real-home ensure, legacy sweep) and the grant and rebase inside it. Cold-start work moves off the critical path: retained-home reconciliation (N sequential sessions) is fire-and-forget behind the daemon provider, and the startup real-home ensure chains into managed hook reconciliation instead of blocking app init. Every preserved semantic is unchanged: never throws, the ORCA_DISABLE_CODEX_TRUST_RPC kill switch, ledger hits, backfill-pending and cooldown fallbacks, config rollback on every failure path, pre-grant self-computed trust removal, the verify-failure taxonomy, diagnostics and telemetry. * fix(codex): widen the trust-config lane to every config.toml writer Review follow-ups on #16441's async trust grant: - `markCodexProjectTrusted` now runs inside the runtime+system config.toml lanes, so a project-trust write can no longer land inside a hook grant's capture->restore window and be silently reverted. Its callers await it. - `install`/`refreshRuntimeUserHooks`/`remove` hold the system config.toml lane as well as the runtime one — they promote approvals into ~/.codex/config.toml and mirror it back. Lock order is runtime-before-system everywhere. - The real-home ensure chain resumes after a rejection instead of returning the same rejected promise to every later pane launch, and resolving the real home is now inside the module's never-throws boundary. - `buildSpawnEnv` awaits inside a cancelable pending-spawn registration, so shutdown during the (now long) env build stops the PTY from launching. `prepareLocalPtySpawn` generalizes into `awaitCancelableLocalPtySpawn`. - CapabilityProbeCache drops the test-only `nowMs` passthrough; its probe backstop comment now describes what it actually guards. - Preflight is a plain async function; the trust dispatch in orca-runtime collapses into one `markWorkspaceTrustedForAgent`. * test(codex): exercise the trust-config lane under real concurrency The async grant makes two pane launches overlap for the first time. These drive the real modules end to end on real files: a rollback swallowing a sibling's grant, a markCodexProjectTrusted write landing inside a capture -> restore window, shared capability-probe dedupe on a cold host, the host-scoped transient cooldown, and reentrancy from inside an installer. Each was verified to fail against a deliberately broken implementation (lane removed, dedupe disabled, cooldown made global, reentrancy pass- through disabled). * test(codex): stop hook-service suites spawning the developer's real codex The forked grant bundle never existed under vitest, so the RPC lane was unreachable in tests on main. Running it in-process makes these suites spawn a real `codex app-server` when one is installed: 38 spawns and two failures in hook-service-runtime-trust-repair on a machine with codex, green in CI where there is none. Stand in for the missing binary so both environments exercise the same fallback lane. * docs(codex): scope the trust-RPC kill switch comment to what it actually gates The comment read as though the flag forces the fallback lane everywhere. It gates the managed grant only: the real-home rebase still runs its own inspect/repair app-server sessions when Orca's insertion shifts a user's hook positions, and never reads the flag. Verified by exercise, not by reading — with the flag set, both inspect-user-hook-trust and repair-user-hook-trust still ran. Pre-existing: main has no check there either, it just blocked the main thread while doing it. Widening the flag to cover the rebase is a follow-up; this only stops the comment promising something the constant does not do. |
||
|
|
08a447dfb2 |
fix(terminal): size the pre-Enter wait to what the host actually ingests (#15925) (#16586)
* fix(terminal): size the pre-Enter wait to what the host actually ingests The Windows agent-prompt submit delay was a flat 1_500 ms frozen from the client's process.platform at import. Measured on two real Win11 hosts, ConPTY ingests a bracketed paste linearly at ~0.009-0.010 ms/byte, so the constant was both far too long for a 2-8 KB prompt (14-89 ms of real cost) and too short past ~145 KB — at 160 KB one host took 1_499 ms, meaning Enter landed mid-paste, exactly the corruption the delay exists to prevent, up to the 16 MB input ceiling. Replace it with getTerminalPasteIngestMs(platform, byteLength) and derive every pre-Enter wait from it: - open-loop fallback = 500 ms settle + ingest bound, uncapped - claude/codex render gate cannot start its quiet window before the ingest bound elapses (an agent that repaints mid-ingest could otherwise satisfy marker-then-quiet while ConPTY was still feeding the paste), and its 8 s hard cap now sits on top of the ingest bound instead of standing in for it - the plain terminal.send suffix path, which had an undocumented flat 500 ms The rate follows the host that owns the pty transport, not the client: a WSL pane is spawned as wsl.exe behind the Windows pseudoconsole so it still pays ConPTY, while an SSH pane follows the relay's reported remotePlatform. Also swap the inter-chunk setTimeout(0) for setImmediate. It cost a full ~15 ms Windows timer tick per 16 KiB chunk (~0.95 s/MB) while pacing ~1.07 MB/s — 11x above ConPTY's drain rate — so it never provided backpressure; the event-loop yield it did provide is preserved. * fix(terminal): stop double-charging paste ingest in the render gate The render gate's hard cap is armed twice -- once at arm() and again when the show-cursor marker arrives -- but it re-added the whole ingest window each time while the ingest clock itself runs once from gate construction. A marker seen mid-ingest pushed the cap out by a second full ingest term (~34 s instead of ~24 s for a 1 MB prompt on ConPTY). Capture the ingest deadline absolutely and arm with what is left of it. Also thread the request AbortSignal through terminal.send so the now payload-scaled suffix wait can be cancelled: at 16 MB it runs ~262 s, well past the CLI's 60 s request budget, and previously nothing stopped the eventual Enter. Cleanups: a pty record's connectionId is only ever an SSH target id, so the wsl: relay-id guard in getPtyWriteHostPlatform was dead; and hoisting action.text removes both non-null assertions in writeTerminalAction. |
||
|
|
9135b6f004 |
feat(orchestration): surface nested worker depth and propagate it across hosts (#16669)
* feat(orchestration): surface nested worker depth and propagate it across hosts Builds on the depth enforcement in the previous commit, which shipped with the setting reachable only by editing settings.json and with workers never told they could nest. Adds the Settings -> Agents control (a 1/2/3 select rather than a free-form number, which bounds the value without inventing a numeric input primitive). The key stays absent from the SettingsUpdate RPC schema, matching agentSkillSharingEnabled: settings.update is reachable from the CLI, so an RPC-writable depth would let a worker raise its own cap. Adds a SUB-DISPATCH block to the dispatch preamble, emitted only when the worker actually has budget left. A worker told it "usually cannot" delegate still tries and then reports the refusal as a blocker, so the section is omitted entirely rather than softened. Propagates depth to federated worker hosts. Previously the home side computed and stored a depth the remote host never received, so a remote attachment always read as depth 1. That is correct at the default cap and wrong as soon as the cap is raised — precisely when someone starts relying on nesting. The field is optional, so an older Run home simply omits it and the attachment's NOT NULL DEFAULT 1 keeps the fail-closed behaviour. Enforcement still runs on the executing host against that host's own cap, consistent with the SSH execution boundary. * fix(orchestration): close nested depth readiness gaps * fix(settings): defer nested depth translations * fix(orchestration): drop federated depth keys that main already landed The enforcement PR's review pass added the same federated depth propagation before it merged, so replaying this branch onto main produced duplicate object keys. Keep main's versions -- its schema entry validates an integer >= 1 rather than any finite number. * fix(settings): label nested worker depth select * fix(settings): move nested depth to orchestration * fix(settings): refine nested depth placement |
||
|
|
64c992cd56 |
fix(memory): report the Windows number that predicts paging, not just resident pages (#16211) (#16589)
* fix(memory): report Windows commit charge, not just working set (#16211) On Windows the per-process figure was working set — resident pages only. An agent whose pages Windows has trimmed to the pagefile shrinks its working set while still holding the commit that pushes the host into paging, so Resource Manager and `orca diagnostics memory` understated an owned tree by 10-40x (9 codex.exe: 1.4 GB working set, 13.4 GB private) and could not warn before the host was already thrashing. Add committed private bytes as a second, separately-labelled quantity rather than redefining the existing one: - CIM sweep gains one property (PageFileUsage, UInt32 KB); the typeperf fallback gains one counter (\Process(*)\Private Bytes). Both ride the sweep that already runs. - MemorySnapshot gains optional `privateMemory` per app/worktree/session plus `processCommitMetric` and `totalPrivateMemory`. Rule 1 additive optional fields: old clients ignore them, and absence reads as "not measured", never as zero — Unix hosts and older hosts send nothing. - `totalMemory` and `processMemoryMetric` keep their exact meaning, so the "shared pages may repeat" copy stays true; the working-set copy now also says paged-out memory is not counted. - Resource Manager shows "Σ Private" beside "Σ WS", and tints the badge yellow/red once tracked commit passes 60/80% of physical RAM — the same thresholds `usageTextColorClass` already uses for host usage. Tint and tooltip only; no toast, and the badge number is unchanged. The parsers move to windows-process-sample-parsing.ts and the Windows sweep tests to their own file to stay under max-lines. Not migrating the collector to windows-process-table.ts: the native snapshot exposes no commit figure and no CPU times, and truncates WorkingSetSize through a DWORD. Documented in the enumeration reference. * fix(memory): derive the typeperf field cap from the counter list The fallback parser's 8192-field cap was sized for three `\Process(*)` counters. Adding `Private Bytes` cut the parsable process count from ~2730 to ~2047, and overrun is a blackout (`parseTypeperfCsvLine` returns `[]`, so the whole sweep reports nothing) rather than a truncation. The counter list now lives beside the decoder that reads those names back out of the PDH header, and the cap is derived from it. Also collapses the four spellings of "omit privateMemory when unmeasured" in collector.ts onto one `commitField` helper, drops the unread parameter and the never-rendered `columnLabel` from `getResourceCommitMetricCopy`, folds `getCommitPressurePercent` into the only function that called it, and reverts unrelated Prettier churn in the Windows enumeration doc. The commit tint's doc comment no longer claims to predict host paging: it measures Orca's own share of physical RAM. Host commit charge / commit limit stays a follow-up (#16211). |
||
|
|
0e10fc5925 | fix(browser): retire helpers with page owners (#16564) | ||
|
|
ef0d5931bc |
fix(source-control): budget WSL bulk git command lines by bytes, not path count (#16634)
Selecting ~100 changed files in a WSL worktree and hitting Stage All did nothing: the files stayed unstaged and the operation reported a failure. Bulk stage/unstage/discard chunked pathspecs 100 at a time, a count picked against a raw argv. A WSL-routed write is not a raw argv -- it is folded into one login-shell command line that shell-quotes every pathspec, quotes the result again, and embeds it three times (one branch per guest shell), so the finished line runs ~3.4x the raw pathspec bytes. Realistic project paths blew past the 32767-character CreateProcess cap at 100 paths and wsl.exe refused to spawn, with nothing staged. Chunking now measures the finished command line through the real resolver, so the wrapper's quoting rules live in one place and native, WSL and SSH hosts each get the budget of the host that actually spawns. A pathspec too long to fit alone still ships alone rather than being dropped, and no chunk is ever emitted empty -- a pathspec-free `clean -ffdx` would have swept the whole worktree. The tracked-path listing behind that discard also fences the WSL login shell now. Its stdout was parsed NUL-delimited without a fence, so Ubuntu's interactive rc banner glued itself onto the first record: that path failed to match anything git reported and was treated as untracked, sending a tracked file to `git clean` instead of `git restore`. Not observing a path in ls-files output is not evidence the path is untracked. The Windows command-line cap and its libuv-aware length estimate move out of the WSL runner into src/shared/windows-command-line-budget.ts, shared by both callers. |
||
|
|
d1a11b3299 |
fix(pty): match the echo shapes a real tty actually produces (#16542)
Reply echo suppression modelled two echo shapes from the spec rather than from
a tty. Captured under node-pty against real bash, at a readline prompt and
under `read`:
- Readline mangles CSI replies, not just OSC: `ESC [ ?` becomes BEL and the
residue echoes. The projection was gated on an OSC introducer, so a private
DSR echo was never matched at a readline prompt. This is the reachable one:
a mode-2031 theme push (`CSI ?997;1n`) left latched by an exited TUI paints
`997;1n` on a bash prompt (#9993's scenario).
- ECHOCTL carets EVERY control, not just ESC. A BEL-terminated OSC reply
echoes as `^G`, but the needle kept a literal BEL — a string no tty
produces. Hardening only: every in-tree OSC reply is ST-terminated
(terminal-osc-color-reply.ts:112, xterm's own reply), so the changed byte
is unreachable except from a foreign or older emulator.
Why this is not the CSI projection #13160 review dropped: that one was the
identity (`replaceAll('\x1b]', …)` is a no-op on a CSI reply), so it was
ESC-led and 500ms-held bare-ESC tails away from the query parser. This one is
BEL-led. The rule is now asserted for every shape rather than implied by the
gate: holdPartial iff the needle does not start with ESC.
The readline branch is keyed on the private-DSR grammar with a non-empty
parameter list, plus a floor on needle length. The containment grammar admits
`CSI ? n`, and `answerLiveQueryReply` takes client-supplied bytes on the relay
path, so a peer could otherwise arm a two-byte `BEL n` needle and delete the
first bell-then-`n` in ordinary output. #61c65151129 proved this system can eat
real output when a needle outlives its budget; a length floor is cheap.
Live coverage: pty-reply-echo-shapes.node-pty.test.ts writes a reply to a real
bash master and feeds back what it echoes, so a shell or libc change fails the
suite instead of silently disarming suppression. Registered in the
shell-contracts lane. The transcript tests and the caretEcho helpers that
encoded the same ESC-only assumption are corrected alongside.
Suppression is display-only. This does not change what reaches the child's
stdin — the reply is written to the master either way, in call order.
|
||
|
|
a624e7cd5d |
test(agent-status): inventory legacy pane identity surfaces (#16575)
* feat(agent-status): measure identity evidence before migrating any consumer PR 1 of the identity migration. It changes no displayed or routed identity — it only measures. Why measure first: the hierarchy shipped in #16148/#16157 has zero consumers, while ~31 sites still derive identity independently. Every migration decision after this is currently a guess, including the one that matters most — how often a real pane has no evidence at all. A live P0 reports "No Claude status shown", and this design trades toward showing nothing when uncertain, so the blank rate has to be a number before any surface moves. - `pane-agent-identity-evidence.ts` — one assembler that gathers a pane's evidence, so consumers stop each inventing their own ladder. - `pane-agent-identity-census.ts` — shadow-only counters keyed by host kind (native / wsl-host / wsl-distro / ssh / relay) and launch mode (typed / orca-launch / resume). Records a bitmask of which sources were present and whether the resolver returned null or ambiguous. No titles, prompts, paths, handles, or agent text. - `pane-agent-identity-inventory.test.ts` — a ratchet that fails when a legacy identity helper gains a new production caller, so the surface cannot grow while the migration runs. Three review findings are encoded rather than deferred: launch stays above run-key-less completed hooks (promoting the hook lets a stale record hijack a pane); OMP/Pi evidence is owner-normalized before assembly, since OMP emits Pi-compatible frames and a wrapper's hook would otherwise be read as the agent it wraps; and Windows-side `wsl.exe` is rejected as process evidence, because the host observes the distro wrapper rather than the agent inside it. The census cannot be completed from a worktree. It needs representative native, SSH, WSL and relay cohorts collected from real use, and that review is the gate on PR 3 — not this PR. * test(agent-status): keep identity migration inventory-only * test(agent-status): reuse reliable source scanner * test(agent-status): bound inventory scan work * test(agent-status): avoid inventory path false negatives * test(agent-status): refresh identity inventory after base repair * test(agent-status): correct inventory classifications * test(agent-status): correct action boundary inventory * test(agent-status): pin inventory occurrence counts * test(agent-status): fail closed on scanner desync |
||
|
|
8d61cb8b77 |
fix(relay): survivable mobile pairing recovery + desktop assign rate gate (#16659)
* fix(mobile): retry the stored assignment when the director reports no newer move A director answering /v1/connect can only reply relay-moved with the stored assignment; it has no 'assignment unchanged' verb, and sticky assignments make equal-epoch replies the steady state. Treating every non-newer move as fatal made pairing recovery unwinnable for any transient cell dial failure (DRAINING, 1006), which bricked off-LAN pairing on Android 0.0.44. A non-newer move now confirms the stored assignment: the candidate re-dials it with a 250ms floor instead of abandoning the relay path. The move is never adopted or persisted, so the anti-rollback contract (requireStrictlyNewerEpoch for persisted moves) is unchanged. 4429 stays out of director recovery: each cell dial burns an invite attempt server-side and a director hop cannot relieve cell load. * fix(mobile): honor relay director Retry-After when pacing recovery Mobile /v1/resolve collapsed every non-OK status into a generic error and discarded Retry-After, so overloaded windows produced hammering instead of paced retries. RelayDirectorHttpError now carries status and retryAfterMs (clamped to 120s), and the reconnect controller floors its existing transport delay with it — no new timers or retry state. The Retry-After parser is extracted from the desktop relay client into src/shared and reused by both. * fix(mobile): attribute pairing log lines to their candidate path The pairing race interleaves the direct LAN and relay candidates into one PAIRING LOG pane; direct lines (WebSocket closed, Reconnecting 10.x.x.x:6768) carried no path label and repeatedly read as Relay retrying a private IP — misleading users and two investigations. The coordinator now wraps each candidate's sink with an idempotent Direct:/Relay: prefix at the one seam where both paths are known. * fix(relay): gate desktop /v1/assign at the per-host rate limit The director rate-limits /v1/assign per host at 5s, but every desktop retry path could fire immediately: both schedulers draw full jitter from [0, cap] (floor 0), first attempts after a drain are undelayed, the 400-fallbacks issue up to 3 assigns per round trip, and reconcile() cancels the armed Retry-After timer from ~8 refreshDemand callers. Production shows hosts permanently rejected at ~100-200 rejects per success. A shared per-host gate now lives inside requestRelayAssignment — the single assign call site — so every path books a >=5s (+jitter) slot. Retry-After raises the gate persistently, surviving the coordinator's timer cancellation. Concurrent callers serialize through a per-key chain. Callers with staleness fencing pass isCurrent; a superseded caller aborts after the wait instead of spending the host's slot. Internal 400-fallback retries stay one logical attempt and do not re-enter the gate. * refactor(mobile): rename the log-only assignment-echo predicate isCurrentAssignmentMove no longer gates control flow — every non-newer move retries the stored assignment — so the name overstated its role. * fix(relay): honor mid-wait raises and cap the assign gate's inline wait Review findings on the per-host assign gate: the deadline was read once before sleeping, so a sibling's Retry-After landing mid-wait was ignored (the exact storm the gate exists for), and the sleep was uncancellable — a booked five-minute Retry-After could park pairing IPC, which awaits reconcile inline, for its full duration. The wait now runs in 1s slices, re-reading the deadline and the caller's isCurrent fence each slice. Remaining waits beyond 15s fail fast as a RelayHttpError 429 carrying the remainder, so the existing schedulers pace with it while the gate keeps the deadline. Staleness aborts are classified non-retryable. Also from review: the broker's isCurrent wiring and the shared-gate default are now pinned by tests, the 4429 comment states the reservation-order rationale precisely, the mobile Retry-After ceiling is renamed to avoid colliding with the desktop's 5-minute one, and a past-HTTP-date header case is covered. * fix(relay): tag locally paced assigns and warn about frozen test clocks Review polish: the synthesized 429 for a beyond-cap local wait now carries a distinct message (relay_assignment_locally_paced_429) so log censuses can tell it from a real director 429, with the comment stating the invariant that makes the translation honest (local booking alone never exceeds ~5.5s). The gate's sleep option documents that test fakes must advance the clock — the slice loop re-reads it and never terminates against a frozen one. * fix(relay): fence superseded callers at the assign send boundary reserve() checks staleness while waiting, but a caller superseded after booking — or between the 400 field-fallback retries — could still spend one to two requests on an assignment nobody consumes. Re-check isCurrent at the top of sendRelayAssignment so the fallback recursion is fenced too. |
||
|
|
8a07bbd8cf |
fix(orchestration): enforce nested worker depth instead of an accidental fence (#16668)
* fix(orchestration): enforce nested worker depth instead of an accidental fence Orca documented that "dispatched workers cannot spawn their own sub-workers (worker-start is coordinator-fenced)". No such check existed. What existed was a single Run-binding check in the workerStart RPC: a worker's terminal is not bound to a Run, so worker-start happened to fail. The rule was emergent, asserted by no test, and written in no doc — and it leaked. A worker could run-create its own Run, task-create, and worker-start: now bound, the check passed. Replace it with a real, configurable depth cap. Depth is derived from the caller's own active Dispatch rather than from Run binding, which is what dissolves the run-create bypass: creating a Run does not stop you being a worker. Enforcement lives in a single dispatch-row writer that owns all three INSERTs that mint a live worker — the generic claim, the supervised worker-start path (including every retry), and the remote attachment. Two of those were missed by earlier drafts of this change, so `creator` and `maxDepth` are required parameters: a new spawn path cannot compile without deciding, and a boundary test refuses the SQL anywhere else. Schema v30 adds depth to dispatch_contexts and remote_dispatch_attachments, NOT NULL DEFAULT 1 and backfilled to 1 so an unstamped or pre-upgrade row fails closed rather than reading as a root coordinator. The attachment pane indexes widen to the five states in which a remote worker may still be running: loss of contact is not evidence of process death, so an unverifiable worker still counts as a nesting parent. Also adds the caller-evidence assertion that workerStart was the only Run-scoped verb to skip, so a declared --from cannot name another terminal's pane and inherit its depth. Default is 1, so behaviour is unchanged unless the new setting is raised. Two limitations are deliberate and documented rather than papered over: this is a guardrail and not a security boundary, since a caller whose launch evidence is unverifiable (any ordinary restored terminal) can declare another handle; and it is enforced at supervised dispatch creation, so a settled worker whose process is still alive counts as a root again. * fix(orchestration): share caller resolution and pin worker gaps * refactor(orchestration): make the caller resolver's pane contract explicit Overloads so requireStablePane callers get a non-null string instead of casting, and rename the attestation opt-out to say what it means: the caller asserts it itself. A flag called assertEvidence:false reads as "attestation optional", which is the hole this helper exists to close. * fix(orchestration): propagate dispatch depth to federated workers * chore(cli): refresh bundled orchestration guide |
||
|
|
588eec68b4 |
fix(native-chat): stop rendering a tool result whose call is outside the window (#15653)
* fix(native-chat): stop rendering a tool result whose call is outside the window A tool result carries no call id, so it can only be attributed to a tool call loaded alongside it. Both chat views read a windowed transcript tail (mobile 40 messages, desktop 300), and the window regularly opens between an assistant's `tool_use` record and the user-role record that answers it. Claude also re-emits already-answered `tool_result` records at a `/compact` boundary, long after their call scrolled out of the window. `foldToolMessages` had no rule for those: with no assistant predecessor in the output they were pushed through as standalone messages and rendered as a bare, unowned block of raw tool output with no tool name — reading as a message from nowhere mid-conversation. Sampling real Claude transcripts, 176 of 400 sessions (44%) produced one in a mobile-sized first page. Drop a result no loaded call can own, before folding. It is not lost: it comes back attached to its call as soon as the owning turn pages in. * fix(native-chat): scope tool result attribution to folded turns * fix(native-chat): preserve harness-attributed tool results * fix(native-chat): keep interruption boundaries |
||
|
|
cda2280d63 |
Show all automations (#16532)
* Add all-host automations with scoped ownership and multi-authority suppo
Enable automations to run on multiple hosts (SSH targets and local) with
owner-fenced mutations, scoped list queries per host, and conflict
resolution. Introduces desktop and runtime authorities as distinct
automation storage owners, with per-host caching, invalidation, and
retry scheduling on the renderer. Captures registration generations for
SSH hosts to survive re-adoption. Adds CLI support for destination
selection and conflict recovery.
* Filter automation create projects by destination host
Only offer projects available on the selected destination, preventing
the mismatches that would fail at submit time. Auto-adjust the project
selection if it becomes unavailable when the destination changes.
* Add runtime storage authority support for automations
- Support both runtime and desktop as automation storage authorities
- Make owner preconditions optional for legacy-client compatibility
- Cache automation list projections to improve performance
- Add per-row repo/worktree resolution for cross-authority collisions
- Extend automation.list RPC to always include owner metadata
* Replace child_process.execFile with runProcess for external automations
- Migrate external-manager to use cross-platform runProcess wrapper per child-process safety policy
- Abstract electron app/ipcMain APIs in orca-runtime via environment accessors
- Install fake app environment in automation tests for consistent setup
- Reorganize imports to use specific module paths (ssh-target-registry, agent-detection, browser-error)
- Remove external-manager from child-process import allowlists (no longer violates direct import)
* Unify desktop automation CRUD onto the local runtime RPC surface
The desktop authority now speaks the same automation.* RPC contract as
remote runtimes, via callRuntimeRpc({kind:'local'}) -> runtime:call ->
the shared RpcDispatcher. The automations:list/listRuns/create/update/
delete/runNow IPC arms, their preload members, and every renderer
desktop-vs-runtime transport fork are retired; the runtime methods are
the single implementation of scoped lists, owner fencing, and change
publication for both transports (mobile clients already exercised them).
The desktop probe scheduler's priority lease survives the move as an
AutomationService hook the IPC registration installs and the runtime
methods take, so Orca's own automation traffic still parks queued
external-manager probes.
External-manager scope arms and dispatch-loop plumbing stay on IPC by
design; automation change events keep their existing channels (renderer
ingestion already converges them by authority).
* Remove automation ghost SSH tombstone scanning
This functionality for synthesizing tombstones for automation-referenced SSH
targets is no longer needed as part of the automation system refactoring.
* Refuse orphan automations at dispatch time, not migration time
Remove migration-time disabling of orphan automations and the `enabledDecidedBy` field. Dispatch now refuses orphans at runtime instead, simplifying state management and UI. Orphans are left unstamped and enabled; dispatch refuses to run them via `resolveAutomationRunTarget`.
* Show all automations in flat table with unified filter menu
- Replace host picker component with comprehensive Filters menu supporting status, last run, agent, and host filters
- Flatten automation list layout to single table instead of host-grouped sections
- Add Host column to display execution host for each automation
- Display active filters as removable pills below toolbar
- Delete unused AutomationHostPicker* components
* Add automation owner fencing and destination validation
- New AUTOMATION_OWNER_FENCING_RUNTIME_CAPABILITY for owner preconditions; legacy clients get owner metadata snapshotted at RPC boundary for compatibility
- Editor captures and revalidates automation destination before save, preventing silent retargeting if SSH infrastructure changes mid-edit
- SSH target types now isolate renderer-authored fields; generation is server-owned and stripped by IPC handlers
* Route automation recovery actions to the origin host
When an automation action fails due to owner fencing, recovery verbs
("Update server", "Reconnect") must run on the host where the refusal
originated: the row's captured owner for row operations, or the
destination the create dialog captured, not the list's filtered host.
* Remove external manager scope limitation notices
Consolidate create destination eligibility checks with a unified predicate
and fix the bug where desktop repo IDs could be sent to runtime hosts where
they cannot resolve.
* Persist only store-derived automation contexts, not client-perspective o
Store contexts must never be based on client-provided runContext or sourceContext
values—clients speak a different perspective (e.g., 'runtime:<id>' for host IDs
they assign), and persisting those makes the store projection orphan automations
it actually owns. Derived contexts now take precedence in create and update paths,
with explicit null still honored to clear a value. Tests verify this by simulating
drift after storage and confirming that moves re-derive while toggles preserve.
|
||
|
|
19e9ec695b |
perf(windows): ship the native process table to Windows relay hosts (#16598)
* feat(windows): let a relay host bind the native process table directly The CIM fallback from #16550 answers on relay hosts, but it costs a powershell.exe and ~1.4s per scan where the native reader costs ~57ms. It is a parachute, not the destination. Teach the loader a second source: the desktop app keeps resolving the npm package, and a relay host -- which has none of our node_modules -- binds a bare `windows-process-tree.node` staged beside the bundle. The CIM scan stays as the last resort, so a host with neither is unchanged. Bind the addon directly rather than its package wrapper. lib/index.js adds only a queue over getProcessList, and that queue is the wedge this module already defends against: it latches a module-global requestInProgress with no try/catch. We hold our own single-flight and deadline, so going straight to the addon drops the duplicate. Measured on a Windows 11 SSH host with ~1490 processes, running the relay-externals bundle from the deployed relay directory: no addon staged nativeAvailable=false 1247ms (CIM) addon staged nativeAvailable=true 57ms memory restored Degradation was exercised on that host, not just in fakes: a truncated upload, a text file, and a foreign-arch ELF each fall through to the scan rather than throwing, and restoring a good addon recovers. A file that loads but lacks getProcessList is rejected by shape, because binding to it would reject every read forever where falling through still answers. No artifact is staged yet, so this is inert until the packaging change lands: today every relay takes the same CIM path it does now. * build(relay): ship the Windows process-table addon to relay hosts The CIM scan restored correctness on Windows SSH hosts, but it costs a powershell.exe and ~1.4s per read where the native addon costs ~57ms. It was always the floor, not the destination. The addon cannot be npm-installed on a relay host: it carries a binding.gyp, so npm rebuilds from source and the build wants Spectre-mitigated libraries even where MSVC is already present. The binary inside the published tarball loads, but predates our patch and still caps enumeration at 1024 processes -- on a 1486-process host it returned exactly 1024 rows with the querying process among the missing, which reads as unavailable only under load. No published alternative clears the bar either; the one fork with a working prebuild story still carries the same cap. So build it where a compiler exists and ship the result. The build script refuses unpatched source -- checking the source rather than trusting the install, because the Spectre hunk fails loudly while the 1024 hunk fails silently -- and verifies the PE machine field so a cross-build cannot emit host arch for another target. The artifact is optional: hashed when present so a relay carrying it never shares an immutable directory with one that does not, and never probed, since requiring a file only a Windows build machine can produce would make a correct relay read as MISSING and redeploy forever. Builds on any other OS keep using the scan, unchanged. arm64 cross-compiles from the x64 runner but needs the optional MSVC ARM64 toolset, so it stays best-effort: a runner image without that component should cost arm64 relays the fast path, not fail the release the x64 relay is riding on. ORCA_REQUIRE_RELAY_NATIVE_ADDONS is a per-arch list rather than a flag for exactly that reason. * build(relay): require the arm64 process-table addon too The arm64 cross-compile is no longer unproven. On a Windows x64 machine with the MSVC v143 ARM64 build tools component installed, node-gyp --arch=arm64 produces a genuine ARM64 image: x64 machine=0x8664 152064 bytes arm64 machine=0xaa64 139776 bytes So arm64 stops being best-effort and joins x64 in the required list. It was only best-effort because the component is optional and I had not seen it succeed; a runner image without it now fails the build with MSB8020 naming the missing component, and that step runs before the long packaging step so the failure costs seconds rather than twenty minutes. The env var stays a per-arch list rather than reverting to a flag, so a future arch can land best-effort before being promoted the same way. |
||
|
|
1fafccb26b |
fix(settings): use Workspace Directory for the Create-project default path (#14767) (#16583)
* fix(settings): use Workspace Directory for the Create-project default path `repos:getDefaultCreateProjectParent` hardcoded `join(homedir(), 'orca', 'projects')` and never consulted the settings store, so Settings -> General -> Workspace Directory had no effect on the Location field of "Create new project". Users had to retype the path every time, or fake it with an NTFS junction. Resolve the parent from the store instead, through the same rule the rest of the app uses for a host preference: `host override ?? client default`, i.e. `getEffectiveHostSetting(settings, LOCAL_EXECUTION_HOST_ID, 'defaultWorktreeLocation', settings.workspaceDir)`. This handler only ever answers for the local host, and a local-host override previously could not win either. A seeded value is not a user choice. `workspaceDir` is never blank -- new installs seed it with `~/orca/workspaces` -- so treating any non-blank value as configured would silently relocate every existing user's new projects into the worktree root. Worktrees nest at `<workspaceDir>/<repoName>/<branch>`, so such a project would then host its own worktrees inside its own working tree. Compare against `getDefaultWorkspaceDir(homedir())` (now exported) via `normalizeRuntimePathForComparison`, and keep `~/orca/projects` for blank, whitespace-only, and untouched-default values. Also scope the `~/orca/projects` shorthand in `formatCreateProjectParentSummary` to the fallback path itself. Otherwise a user with Workspace Directory set to `J:\PROJECTS` saw the summary line claim `~/orca/projects` while the field held `J:\PROJECTS`. Fixes #14767 * fix(settings): keep configured orca/projects paths verbatim in the create summary The collapsed Location summary used a tail match on orca/projects, so a configured directory like /data/orca/projects rendered as ~/orca/projects. Scope the shorthand to usual home layouts and pin the lookalike cases. |
||
|
|
933345d347 |
Clarify upstream divergence stats for rebased branches (#16358)
* Clarify upstream divergence stats for rebased branches When a branch is rebased, it still tracks the pre-rebase upstream while comparing against the new base. Move upstream arrows to the head line to prevent them being confused with compare-base counts. * Show upstream divergence stats independent of compare base Measure HEAD against upstream regardless of compare-base state, so divergence indicators stay visible even when comparison is missing, loading, or failed. Also use cross-platform temp paths in tests. * Show commit counts against compare base, not upstream Upstream divergence (↑/↓ against tracking branch) was confusing for rebased branches — the counts appeared beside the base ref but measured against the upstream branch. Show only the compare base count instead, on the line that names it. * Report branch divergence in both directions Rebased branches are typically ahead AND behind their base; a single count hides this case. Use symmetric range with --left-right --count to capture both directions efficiently, then expose commitsBehind in the UI alongside commitsAhead. * Use semantic names for i18n keys and template variables Rename hash-based translation keys to descriptive identifiers and replace generic value0/value1 placeholders with semantic variable names like `count` and `ref`. Improves code maintainability and makes translation strings self-documenting. |
||
|
|
290f192d84 |
fix(updater): surface and degrade renderer shutdown checkpoint failures (STA-5505) (#16497)
* fix(updater): surface and degrade renderer shutdown checkpoint failures The in-app updater could refuse to install with 'Renderer shutdown checkpoint was not completed.' while the actual persist() error was swallowed unlogged, leaving users stranded on old builds (STA-5505). - report the swallowed persist error: console, crash breadcrumb, and a cross-world DOM attribute so the thrown error (and the Update Error dialog) names the underlying cause - stop failing the checkpoint on sleeping-agent quit-capture errors; the periodic capture bounds the loss to one minute - extend the existing durable-session degradation to full-session staging failures during an intentional restart, preserving the dirty-draft guard * fix(quit): degrade and surface checkpoint-vetoed app quits (#15352) Cmd+Q walked the same shutdown checkpoint as the updater: a persist() throw preventDefault()ed the synthetic beforeunload and confirmNativeWindowClose returned silently — quit accepted, nothing logged, SIGKILL the only exit. - run the quit checkpoint inside a window-close scope so full-session staging failures degrade to the durable tier for app-level closes too (dirty editor drafts still hard-block) - when the checkpoint still vetoes the quit, toast the published failure reason instead of dying silently * fix(updater): retry-then-degrade staging and honest capture-loss accounting Review findings on the first pass: - a first full-session staging failure now stays a visible, retryable error; only a repeat failure degrades to durable-only staging, so a transient IPC failure keeps its retry instead of silently dropping just-captured scrollback - the sleeping-capture comment no longer overstates periodic coverage (periodic mode skips done panes and never stamps quit origin); the swallowed failure records a crash breadcrumb - pin the exact degradable-shutdown gate expression in the source-shape test so rewiring it cannot pass silently * fix(updater): arm the staging-retry flag only for degradable shutdowns An unrelated unload's staging failure must not burn the visible first retry of a later restart or quit. * fix(updater): isolate shutdown checkpoint retries Reset full-session staging retry state when a shutdown attempt is abandoned, and route Terminal-less closes through the same scoped synthetic checkpoint as mounted workspaces. Keep arbitrary thrown-value diagnostics non-throwing and localize the quit failure toast. * fix(updater): preserve checkpoint retry lifecycle * fix(updater): preserve empty checkpoint failure reason |
||
|
|
a1ec0479e2 |
fix(windows): revalidate PTY liveness from the job object, not a forked helper (#16419)
* fix(windows): answer console membership from the job object, not a forked helper
node-pty answers "which processes are attached to this pane's console?" by
FORKING a helper, because GetConsoleProcessList must run from a process
attached to that console. Orca asked on a foreground poll, per pane, so each
read spawned a conpty_console_list_agent -- hundreds of hidden processes
exhausting RAM within minutes, respawning as fast as they were killed (#10857).
QueryInformationJobObject has no console-attachment constraint: any process
holding the job handle can ask. Orca already creates that job per PTY, and
listPtyJobProcessIds has exposed it since the W1/W2 work with zero callers.
One syscall, no children.
Semantics the three call sites rely on are preserved: a root-only set still
proves the shell is alone (so a stale agent can be retired), and size > 1 still
proves something is running under it. The single difference is that a
descendant detached from the console stays in the job -- which widens the set,
the conservative direction for every caller.
Also fixes the third call site, which returned { available: false } whenever
membership was unavailable AND a recognized agent existed -- i.e. exactly while
an agent was running. Membership only ever narrowed the candidate list, so an
unavailable answer now leaves it unfiltered instead of failing the whole
resolution.
The no-fork test is asserted through a module-level vi.mock of
node:child_process. A vi.spyOn of a require()'d child_process does not
intercept the module's own import binding: the first version of that test
passed with a fork() deliberately reintroduced.
* fix(windows): keep console attachment for the candidate filter
Readiness review caught that this PR changed two different questions as if they
were one, and the repo's own plan doc had already said so:
"The job is the wrong set here -- it would re-admit precisely the detached
process the filter exists to drop." (windows-wsl-root-cause-plan.html, Use B)
The two uses:
- Use A, `size > 1` at local-pty-provider and the daemon tracker -- "is anything
in this pane besides the shell?". The job answers this, in-process and with no
fork. Unchanged from the previous commit.
- Use B, the candidate filter -- "which of these are ATTACHED TO THIS CONSOLE?".
Its whole job is dropping a descendant that detached, and the job object keeps
those, so answering it from the job makes the filter a no-op in its motivating
case: a detached `Start-Process droid` would be granted byte authority, and a
detached sibling would make an attached agent look ambiguous.
Use B goes back to GetConsoleProcessList, in its own module named for what it
answers, with its fail-closed null restored. That path is not the #10857 storm:
it runs only when a recognized agent candidate already exists, not on every
foreground poll. Bounding it to one pooled supervised helper is the remaining
half, and per the plan doc either half alone takes #10857 from unbounded to one.
My earlier claim that widening membership is "the conservative direction for
every caller" was wrong -- true for Use A, backwards for Use B. The hardware run
did not catch it because I measured a WSL pane, where the superset is harmless,
and never a detached GUI child, which is the divergence.
* fix: restore the coverage and ratchets the module split dropped
Round 2 of review. Two blockers, both from moving the forking code to a new
file without moving what guarded it.
- The child_process import ratchet was RED: windows-console-attached-processes.ts
imports node:child_process and was unlisted, and the old entry was stale. I
never ran that suite -- lint and the providers/daemon tests both pass without
it, which is exactly the gap the ratchet exists to close. Entry repointed;
count unchanged at 159.
- The forking module had ZERO tests. Its 11 assertions -- bounded timeout,
single kill, spawn error, malformed message, helper-pid removal -- were in the
file that now answers a different question, so the module that actually caused
#10857 was shipping untested. Moved with the code.
Also: nothing pinned the round-1 fix itself. No test drove console attachment to
null and asserted the fail-closed result, so re-deleting that branch would have
gone green. Now covered, and verified to fail when the branch is removed.
Cleanups the split left behind: `consoleMembershipUnavailable`/`consoleProcessIds`
renamed to `pane*` where they now hold job membership, the duplicated
`WindowsConptyMembershipDeps` type name, comments still describing the console
on the job path, and eight reliability-gate paths pointing at the moved tests.
* fix(windows): let a superset job answer expire instead of vetoing retirement
Round 3. The job read had reintroduced #9258's bug by a new mechanism.
`size > 1` returned unconditionally, so any pane holding a console-detached
descendant never retired its cached agent. A WSL pane always holds some: the
measurement in this PR's own test recorded job [40980,104068,4888,69908] against
console [69908,40980], i.e. console said "shell alone, retire" while the job said
"three others alive, keep". #9258's third commit describes the identical failure
from the other direction -- a bare shell reading as [helper, shell] "looked like
it still had a child ... the foreground refresh held the exited agent's identity
indefinitely" -- and that is what came back.
It bites because the read branch that serves the cached name across a Windows
shell fallback is deliberately untimed: #9258 made it so on the stated assumption
that "the background refresh authoritatively retires it". Removing the retire
authority left the identity with no bound at all. Second-order: a non-null cache
makes idleNoEvidenceShell false, which pins the refresh at the 1s TTL, so an idle
WSL pane also scanned the process table every second forever.
A TTL on the read would have been the wrong fix -- untimed is deliberate, because
on Windows the fallback name is structurally uninformative. Instead the job answer
is treated as what it is: a SUPERSET of the console, which cannot tell a working
agent from a leftover. Proof of absence retires immediately (size 1, unchanged);
an inconclusive answer ages out at 30s; unverifiable (null) still holds forever
per ssh-execution-boundary.md. Only successful scans that found no agent advance
the clock -- a degraded scan returns before this -- so the fix cannot expire an
agent it simply failed to see.
Also from review:
- Restore the root requirement the forked probe had. Without it a set of one
non-root pid -- shell gone, descendant alive -- read as "shell alone, retire",
inverting the truth.
- Rename to windows-pty-job-membership.ts / readWindowsPtyJobProcessIds. The old
name still said ConPTY console while reading the job, and conflating those two
sets is precisely the bug
|
||
|
|
91a500712c |
fix(crash-reporting): see the renderer memory the heap counters never report (#16449)
* fix(crash-reporting): see the renderer memory the heap counters never report Windows renderer crash 36048e26 arrived with 618MB of private renderer memory and a `renderer_memory` breadcrumb reporting a 150MB V8 heap. Both numbers were right: xterm scrollback lives in `Uint32Array` backing stores and glyph atlases live in GPU transfer buffers, and neither is counted by `usedHeapSize`, `mallocedMemory`, or Blink's allocator. That made the report unanalyzable. `renderer_memory_highwater` is the crumb carrying the subsystem census that names what grew, and it is armed on `usedHeapSize / heapSizeLimit`. At 150MB of a 4192MB limit that ratio is 3.6% — nowhere near the 60% mark — so the census never reached a single one of these reports. Measured on Windows (6 worktrees x 4 terminal tabs, 8000 lines each, this app at 4218d505): filling 24 mounted panes moved the renderer working set from 210MB to 656MB while `usedJSHeapSize` stayed at 43MB for the whole run. Sample the renderer's own OS footprint through `process.getProcessMemoryInfo()` (available in the sandboxed preload) and: - report `privateMB`, `residentMB`, and `outsideHeapMB` — the footprint minus everything V8 and Blink admit to holding — on every `renderer_memory` crumb; - arm the highwater census on private-footprint marks (600MB / 1000MB) as well as the heap ratio, so growth outside the JS heap now carries the pane and store census that names it. The footprint read is async, so a sample annotates with the previous read and refreshes in the background: one interval of staleness is irrelevant to a footprint trend, and awaiting it would make every sample reentrant. A shell without the bridge, or a runtime that withholds the read, keeps sampling exactly as before. Retained-breadcrumb keys now distinguish the two threshold ladders; keying only on `thresholdPct` collapsed every footprint crumb onto one slot. crash-diagnostics.ts split at the max-lines budget: memory sampling moves to renderer-memory-sampling.ts and the shared payload shaping to crash-breadcrumb-data.ts. * fix(crash-reporting): retain all renderer memory marks |
||
|
|
efa3b972c2 | fix(native-chat): prevent duplicate mobile prompt echoes (#15656) | ||
|
|
a9781a4118 |
STA-4150: client-hosted remote browser (consolidated) (#15448)
Co-authored-by: Jinwoo-H <jinwoo@stably.ai> |
||
|
|
e361da7fb7 |
Deleting skill (#16357)
* Add skill deletion with cross-platform transaction safety Implements end-to-end skill removal with placement enumeration, dependency guards, and transactional recovery. Covers native, WSL, and remote hosts; users can delete canonical directories and alias placements (symlinked directories or files) in a single atomic batch. Includes UI selection flow, preview, confirmation, and results band. Block reasons (bundled, plugin, unowned, stale) gate deletions that would fail or contradict user intent. * Organize IPC handlers into module subdirectories Move register-core-handlers and skill-delete-ipc-handlers into dedicated subdirectories for improved code organization and to reduce the flat structure in src/main/ipc/. * Make skill deletion recovery transactions idempotent Defer journal cleanup until both staging removal and receipt cleanup succeed, leaving the journal in place for startup to retry if either operation fails. This ensures the recovery process is safe to run multiple times without leaving partially-deleted skills. * Consolidate skill-delete files into dedicated module Reorganize skill deletion functionality into a modular structure under `src/main/skills/skill-delete/` with simplified file names. Remove the redundant `skill-delete-` prefix from file names since they now live in the dedicated directory. Update all import paths throughout the codebase to reflect the new structure, including imports from IPC handlers and RPC methods. * Fix broken import paths and add deletion robustness improvements Import paths using `..//'` were invalid and broken. Replace with explicit module names (`skill-discovery-sources`, `skill-install-filesystem`, etc.) to clarify dependencies. - Bind WSL filesystem methods to preserve `this` context - Keep recovery journal when rollback rename fails, so startup can retry - Skip symlink-based tests on Windows where they cannot run - Only treat ENOENT/ENOTDIR as empty directories; propagate other errors - Fix cross-platform path parent calculation to handle drive roots - Replace shared constant with localized string for user-facing message - Use `runProcess` for WSL integration test instead of bare `execFile` * Add batch limit for skill deletion and improve host availability checkin - Limit concurrent deletions to prevent remote host overload - Add retry logic for capability probing to handle transient unavailability - Add reprobe() method to recheck capability after errors or user refresh - Fix status logic: receipt cleanup is best-effort, completion depends only on content removal - Improve error message for unreachable hosts |
||
|
|
4218d5068e |
fix(cli): seed nvm's default version, not the newest install (#16420)
* fix(cli): seed nvm's default version, not the newest install #16314 stopped the login-shell probe inheriting the seeded PATH, but left the seed itself picking the newest installed nvm version. That ordering decides which node a CLI runs under whenever the probe does not land — a timeout, or a login shell whose rc never initializes nvm — and newest is precisely the wrong guess: it is usually the version the user just added and has installed nothing into. That is the root cause reported in #10932. Resolve `alias/default` instead, mirroring nvm: follow the alias chain (`default` -> `lts/*` -> `lts/krypton` -> a version), resolve a partial version like `24` to the highest matching install, and treat `system`/`node`/`stable` as no preference. The chain is bounded and cycle-guarded because nvm's own resolver tracks seen aliases and hand-edited files can point at each other. Ordering is a preference, not a restriction: the remaining versions stay behind the default, so a CLI installed outside it is still reachable. Measured on a real machine with nvm default=24 and a bare v26.7.0 installed: the old resolver seeds v26.7.0/bin (no CLIs), the new one seeds v24.18.0/bin (every CLI). Tests were written first and verified to fail on the three bug cases against main before the fix existed. Also raise the probe budget from 5s to 10s. The old value was never measured against a real profile: a bash -ilc loading nvm, rvm, conda and gcloud takes ~1s idle but 6-7s on a loaded machine, so a cold start under load silently fell back to the seed. Startup does not block on the probe, and the one awaited consumer is agent detection, which is better served by a probe that finishes late than one that gives up early. * fix(cli): reject non-version alias tokens instead of matching v0.x Review finding, and a real bug I introduced. parseVersionSegment coerces every unparseable segment to 0, so an unresolvable default alias — `garbage`, `iojs`, `lts/nonexistent`, any hand-named alias — became [0] and prefix-matched a `v0.12.x` install, or any stray non-version directory. Orca would then seed a decade-old node as the preferred runtime. Real nvm answers N/A for all of them. The `wanted.length === 0` bail could never have caught this: ''.split('.') is [''], never empty. Replaced with a shape check that still admits legitimate numeric prefixes — verified against nvm itself, which resolves `24` to v24.18.0 and `0` to an installed v0.x while answering N/A for the rest. Also corrects two comments that no longer described the code: the seed is no longer "newest install", and the probe budget note claimed startup never blocks on hydration, which is false on packaged Windows where it gates terminal services and git. The traversal-guard comment claimed a containment join() already normalizes away; the real guarantee is that matchNvmVersion can only return an entry of the versions directory. * fix(cli): match nvm's version-token grammar, not just its first character Round-2 review finding, and the same bug one layer down. The previous guard anchored only the first character, but parseInt stops at the first non-digit, so `0x18`, `00` and `0abc` still parsed to [0] and prefix-matched a v0.12.x install — the decade-old-node seed the earlier fix was supposed to close. Reachable: `nvm alias default 0x18` warns that the version does not exist and writes the alias anyway, then resolves it to N/A. Use nvm's actual grammar, leading zeros included — nvm calls `00` and `024` N/A while parseInt reads them as 0 and 24. Verified by executing 17 tokens against a five-version fixture: every one now agrees with nvm, including the legitimate prefixes `0`, `0.12`, `24` and `v24.18.0`. Also drops a dead disjunct (the hop bound already caps the loop, so seen.size can never exceed it) and corrects the log comment in index.ts, which still told the reader a failed probe leaves the newest install in front. It leaves the default version in front now, which is usually survivable but still not what the shell would have resolved. * test(cli): skip the lts/* chain fixture on Windows Round-3 review finding. makeNvmHome materializes each alias as a real file, and the chain case uses nvm's actual `lts/*` alias — `*` is a reserved Win32 filename character, so writeFileSync fails with EINVAL. PR CI runs a Windows allowlist that excludes this file, so the breakage only reaches a Windows developer running the suite locally. Skipped rather than renamed: `lts/*` is the alias nvm really ships, and the assertion pins platform: 'darwin' anyway, so the real name costs no coverage. Matches the skipIf convention already used across src/shared. Also reflows a comment line that a previous edit ran to 143 characters; oxfmt does not reflow comments, so nothing would have caught it. |
||
|
|
822087c8ec |
refactor(git): split runner.ts into focused command-runner modules (#16395)
* refactor(git): split runner.ts into focused command-runner modules * chore(ratchets): repoint child_process and wsl.exe allowlists at the split modules --------- Co-authored-by: Neil <n@example.com> |
||
|
|
fcf55f2d68 |
fix(terminal): stop Orca mangling the OMP/Pi title it writes itself (#16381)
* fix(terminal): collapse identity group in the title churn signature Replaces the ingest-time title rewrite from #16373 with a non-destructive fix at the actual cause. The churn suppressor `isDecorativeAgentTitleFrameChange` keyed on the literal label, so `working:OMP` and `working:Pi` compared unequal and every alternating frame from a wrapped harness committed a store patch. #16373 made the labels agree by rewriting the stored title to the tab's launch owner — but `runtimePaneTitlesByTabId` is also the Windows Shift+Enter byte-encoding input, so normalizing at ingest destroyed evidence other consumers read (fixed separately in #16376). Collapse the identity group inside the signature instead. Which member of a group a frame names is decoration, exactly like the spinner glyph the signature already strips, so frames compare equal without touching what is stored. Suppression now changes only WHETHER a frame commits, never WHAT it says. Also fixes the flap under a multiplexer (#8032): the collapse runs over wrapper segments, so "zsh | ⠋ Pi" and "zsh | ⠙ OMP" compare equal, which the anchored owner-relabel in #16373 never matched. Reverts the store changes from #16373 and drops the helper it added. Co-authored-by: Seongho.Bak <49228032+psh4607@users.noreply.github.com> * fix(terminal): fold only bare identity frames into the group token A legacy "π - <session> - <cwd>" title is Pi-compatible too, so folding every profile match collapsed two different sessions to the same signature and suppressed the change outright — reintroducing #16093 through the churn signature. Fold only exact bare identity frames, matched per wrapper segment, so semantic session titles keep comparing on their own text. Co-authored-by: Seongho.Bak <49228032+psh4607@users.noreply.github.com> * docs(terminal): correct the flap diagnosis in the repro header Verified against the OMP source: it emits only π-glyph frames (`DEFAULT_TERMINAL_TITLE = "π"`, title-generator.ts:25), and on an Orca-hosted pane its native titler cedes to Orca's own injected extension, which writes `⠋ π - <session> - <cwd>`. So OMP emits neither "OMP" nor "Pi". Both flap sides are Orca's: "OMP" from driveSyntheticTitleFromHook, "Pi" from normalizeTerminalTitle collapsing our own extension's output to a hardcoded literal. The prior header credited the wrapped harness for frames it never sends, which is the same wrong narrative that produced eight fixes at eight layers. No behavior change. * fix(terminal): stop Orca mangling the OMP/Pi title it writes itself Verified against the OMP source: it emits only π-branded frames (`DEFAULT_TERMINAL_TITLE = "π"`, title-generator.ts:25), and on an Orca-hosted pane its native titler cedes to Orca's OWN injected extension, which writes `π - <session> - <cwd>` / `⠋ π - <session> - <cwd>` at 80ms. So neither flapping string came from OMP. Orca made both: "Pi" — normalizeTerminalTitle collapsing our extension's output to a hardcoded literal, discarding the session name and cwd (#16093) "OMP" — driveSyntheticTitleFromHook injecting over it every 80ms Fixed at the source: - normalizeTerminalTitle canonicalizes only the rotating braille frame and keeps the rest, in both spinner positions and through a multiplexer prefix (#8032). Status still round-trips through normalization. - detectAgentStatusFromTitle reads the π state separator, so `π ! <label>` is permission instead of the blanket idle that hid a blocked agent. - normalizeCompatibleAgentTitleForOwner swaps only the brand for the owner's label, so a pane still reads as its launch owner (#6689, #7633, #9077) without losing the session text. - pi/omp set synthesizeWorkingTitle: false — the agent animates its own working title. Terminal states still synthesize; they carry the pane's agent identity downstream. Reverts the ingest-time title rewrite from #16373, whose normalization of runtimePaneTitlesByTabId also changed Windows Shift+Enter bytes (#16376). Co-authored-by: Seongho.Bak <49228032+psh4607@users.noreply.github.com> * fix(terminal): match the state separator only in exact profile casing The separator check runs on every title, so `omp - deploy notes` and `pi - refactor the parser` read as an idle agent. The owner rewrite only ever emits the exact profile labels, so dropping case-insensitivity keeps `OMP - tmp` classifying while ordinary prose stops matching. Co-authored-by: Seongho.Bak <49228032+psh4607@users.noreply.github.com> * test(terminal): pin one real OMP turn to two committed patches Drives 30 working frames as Orca's injected extension emits them plus the idle transition, and asserts what survives the churn gate. Before the fix every frame alternated "⠋ Pi"/"⠋ OMP" and each one committed — ~12 store patches per second on a working tab. Co-authored-by: Seongho.Bak <49228032+psh4607@users.noreply.github.com> * fix(terminal): carry the permission guard inside the separator reader `-` is both a π state separator and the delimiter in the synthetic permission label, so `OMP - action required` read as idle. It resolved correctly only because detectAgentStatusFromTitle happens to check the synthetic label first — and the separator fn is exported, so a direct caller inherited the bug. Also pins the owner rewrite's fixed-point property, which holds only because getAgentLabel does not tokenize omp/pi, and corrects a comment that overstated how tightly the brand swap is scoped. Co-authored-by: Seongho.Bak <49228032+psh4607@users.noreply.github.com> * docs(terminal): name the flag the code actually sets The suite header cited `synthesizeTerminalTitle: false`; the profiles set `synthesizeWorkingTitle: false`. The distinction is the whole reason the narrower flag was chosen — terminal-state frames still carry the pane's agent identity downstream — so the wrong name buried the rationale. Co-authored-by: Seongho.Bak <49228032+psh4607@users.noreply.github.com> --------- Co-authored-by: Seongho.Bak <49228032+psh4607@users.noreply.github.com> |
||
|
|
127fa7fae0 |
refactor(ipc): split repos.ts into focused modules (#16392)
* refactor(ipc): split repos.ts into focused modules
* test: point repo notification mocks at the extracted module
* fix(ipc): repoint the child-process allowlists after the repos split
The type-only `import type { ChildProcess }` moved from repos.ts to
repos/repo-clone-lifecycle.ts, so the import-boundary entry follows it and the
windows-console entry (now stale, and that list only shrinks) is dropped.
Fixture-only; the base file had no runtime child_process use at all.
|
||
|
|
8217e6838f |
refactor runtime contracts and web transports (#16197)
* refactor runtime contracts and transports * test(web): repoint two-phase timeout seam at the transport that now owns call() |
||
|
|
83ffc0df24 |
refactor Electron facilities modules (#16333)
* refactor oversized Electron facilities * fix interactive process timeout and shortcut repeat guard * chore(child-process): drop stale cli-installer allowlist entry cli-installer.ts now routes privileged spawns through runProcess via cli-privileged-processes.ts, so the shrink-only ratchet flags it as stale. * refactor(child-process): extract the bounded output sink runProcess's timeoutMs opt-out (required to preserve the unbounded osascript admin prompt) pushed run-process.ts past the 300-line cap. Move createOutputSink to its own module rather than add a max-lines bypass, which AGENTS.md forbids. Moved verbatim; no behavior change. |
||
|
|
1cf562deea | refactor: split source control AI modules (#16179) | ||
|
|
75103667b0 |
test(cli): ratchet the exec/fork family, not just spawn (#16390)
The pairing ratchet matched spawn|spawnProcess|spawnSync|runProcess only, so
a resolved CLI handed to execFile was the same unpaired launch with none of
the enforcement. codex-trust-grant-host.ts resolves codex and calls
execFileSync, and escaped the ratchet purely through that omission.
Widen to the exec/fork family. The negative lookbehind keeps method calls
such as `RE.exec(` out, which is what made the bare `exec` name safe to
include; a fixture mutation confirms `/x/.exec('x')` does not trip it, and
adding execFileSync(resolvedCli) to a paired file does.
codex-trust-grant-host is allowlisted rather than changed: its only exec is
a wsl.exe identity probe for the binary stamp, and its actual codex launch
is a CodexAppServerInvocation paired centrally in codex-app-server-session.
The entry records what would invalidate it.
|
||
|
|
deae7212d9 |
fix(orchestration): derive inject's agent guidance from the recognized-agent roster (#15874)
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> Co-authored-by: vam <a@a.com> |
||
|
|
48e63c015f |
refactor agent config and auth services (#16195)
* refactor: split agent config and auth services * chore: repoint wsl and global-fetch guards at split module paths * fix: restore merge-base Claude CLI error propagation Drop the secret-redaction rewriting added to Claude CLI error paths in the refactor: spawn errors again reject with the original Error (preserving .code/.errno/.syscall/.stack) and command output/auth-status logs are no longer rewritten. |
||
|
|
2b1b094aa8 |
fix(cli): pair every resolved CLI with its runtime, and ratchet it (#16383)
Follow-up to #16365, which paired 8 spawn sites by hand. Hand-pairing is how the class got introduced, so close it structurally instead. cliPath is now required on CodexAppServerInvocation, `null` only for the guest-side wsl.exe launcher where a host path pairs nothing. Optional let a native builder omit it and silently fall back to pairing against a cmd.exe wrapper with no type error. Every production site already passed it; only test fixtures needed updating, which is the type doing its job. Four more sites now pair. codex-state-db-backfill-recovery spawns the same `codex app-server` subcommand #16365 fixed elsewhere. cli/handlers/account was the worst case: addAgentNodePaths prepends the *newest* version-manager bin, which is not necessarily where the CLI being launched lives, so it actively created the mismatch — pairing now runs last so the CLI's own node wins. commit-message-text-generation and skills/skill-update-run spawn resolved binaries with inherited env. cli/handlers/skills had grown its own buildNpxPath: a weaker local copy that prepended unconditionally, ignored the Windows `Path` key, and special-cased a '.' dirname. Deleted in favor of the shared helper, which checks the sibling node actually exists — the behavior change one test had pinned. The ratchet is the point: any file that resolves a CLI and spawns must reference withCliRuntimeOnPath, with a shrink-only allowlist. It caught skill-update-run, which I had missed. Its first draft required a call paren and so let dependency-injected resolvers (`resolveCommand: resolveCodexCommand`) through — verified by removing a pairing and watching it stay green, then widened until it failed. A second assertion fails on a stale allowlist entry so an exemption cannot outlive its reason. external-editor-launch stays allowlisted: it launches a GUI editor, not a Node CLI whose ABI matters. |
||
|
|
a7505fd911 |
fix(cli): spawn a version-manager CLI with its own node runtime (#16365)
* fix(cli): spawn a version-manager CLI with its own node runtime resolveCliCommand falls back to scanning every version-manager install when PATH misses, so it can hand back ~/.nvm/versions/node/v20.x/bin/codex while PATH still leads with v22. Nothing paired the binary with the runtime it was installed against, so its `#!/usr/bin/env node` shebang loaded a v20-built native module under a v22 ABI and the agent died on first require (#10932). Reproduced with a real addon rather than asserted: a CLI requiring a cpu-features build for NODE_MODULE_VERSION 115, spawned with v24 leading PATH, fails with ERR_DLOPEN_FAILED and exit 1. With the CLI's own bin directory prepended it runs clean. withCliRuntimeOnPath prepends the resolved command's directory when that directory ships a sibling node, and is a no-op otherwise — so a Homebrew or /usr/local CLI is untouched, and the WSL paths pass a bare `codex`/`claude` that is not absolute and so never matches. Host CLI resolution in the Claude login path is now lazy, keeping the WSL branch from resolving a host binary it never spawns. * fix(cli): split PATH on the delimiter we join with, pair app-server too Readiness review findings, all four addressed. withCliRuntimeOnPath chose its join delimiter from the platform option but split with the host's. Passing platform:'win32' from a posix host turned `C:\Windows;C:\Windows\System32` into `C;\Windows;C;\Windows\System32` — every drive letter torn off at its colon. Latent, since no shipped caller passes platform, but the sole win32 test was written against the corrupted value and asserted one split segment, so it green-lit the shredding. That test's other assertion was vacuous: it seeded only `Path`, so the `PATH` key it asserted absent could never exist. Deleting the whole case-dedupe block left the suite green. It now seeds both keys and asserts the full joined string; removing the block fails it. Nothing covered the wiring, and the argument choice is the easy thing to get silently wrong. Note it only diverges on win32 — on posix getSpawnArgsForWindows returns the CLI itself, so pairing the spawn command is indistinguishable there. The new test drives the win32 branch with a .cmd fixture; pairing spawnCmd or dropping the wrapper both fail it now. codex-trust-grant-host and codex-session-index-heal spawn the same `codex app-server` subcommand through runCodexAppServerSession and were left unpaired. Pair centrally there via a new optional cliPath, since invocation.command may be a cmd.exe wrapper. Pairing tests live in their own file: adding them inline pushed codex-fetcher.test.ts past the 800-line ratchet. * fix(cli): read the Windows path key the child will actually use Round-2 review finding. The read was narrower than the delete: the key was picked from exactly two spellings (`Path`, else `PATH`), while the twin dedupe removed every key whose lowercase form is `path`. A block spelling it `path` or `pATh` therefore had its value deleted without ever being read, handing the child a PATH containing only the CLI's own directory — a strictly worse outcome than not pairing at all. Win32 resolves env names case-insensitively and object order preserves block order, so the entry the child reads is the first case-insensitive match. The repo already encodes that rule in resolvePathEnvKey (src/main/pty/windows-path-segment-merge.ts); src/shared cannot import from src/main, so mirror it locally. Verified by execution across six env shapes: lowercase, mixed-case, Path-only, PATH-only, both twins, and a PATHEXT control that must not be touched. All preserve the original PATH; before the fix the first two lost it entirely. Reverting the selector fails the new test and nothing else. |
||
|
|
4a57cfac9a |
fix(terminal): give OMP Pi's Windows Shift+Enter encoding (#16376)
OMP wraps Pi's TUI, so Shift+Enter bytes land in a Pi reader that decodes CSI-u. The omp profile had no `windowsShiftEnterEncoding`, so it fell back to Esc+CR — which submits instead of inserting a newline (#9703). This was latent while an OMP pane's stored title could read either "Pi" or "OMP" depending on which interleaved frame committed first. Pinning the title to the launch owner (#16373) made it deterministically "OMP", so the Windows Shift+Enter fallback now always resolves `omp` and always picks the wrong encoding. `prime-agent` already carries this entry for the identical reason. Co-authored-by: Seongho.Bak <49228032+psh4607@users.noreply.github.com> |
||
|
|
e217fdd10f |
build(orcad): gate orcad's own graph, and prove it loads under plain Node (#16368)
* fix(orcad): close the browser-provider gaps The providers landed without enforced coverage, so a regression in either path would have landed silently. - CI: the external-Chromium integration test was gated on ORCA_BROWSER_EXECUTABLE and nothing ever set it, so it skipped forever. It now runs in its own job against the runner's Chrome and FAILS when Chrome is absent rather than skipping, because an unset variable is exactly how it went uncovered. Timeout raised to 120s: a warm run is ~7s but the first launch against an unseeded profile took 30s and hit Vitest's default, and CI is always that cold case. - Electron provider had no test at all. It is the path anyone with the desktop app hits. - Browser unavailability reported one message for four causes, including telling an operator to set a variable they had already set. Fixes a live defect found while covering it: the runtime advertises browser.tabCreate.known-id.v1 unconditionally, so a web client sends a provisional page id for a page that does not exist yet — and the sidecar's generic requestedPageId branch ran require() on it first and threw. Every known-id create against the Electron provider failed. The adoption logic was already there; only the ordering was wrong. Also updates the workflow-parallelism guard, which correctly caught the new job missing from verify's required-check list, and asserts verify actually reads it. * build(orcad): gate orcad's own graph, and prove it loads under plain Node Two gaps the artifact's own comment asked for. The ratchet measured only orca-runtime + runtime-rpc, but orcad imports ipc/pty directly to install the PTY controller, so its graph is strictly larger. The gate could read zero while the shipped artifact regressed. orcad's entry is now a ratchet entry point, and the baseline stays empty with it included. orcad cannot join plain-node-entry-guard — that is a rollup plugin keyed on electron-vite input names, and orcad is an esbuild artifact. But the half that matters here is the guard's smoke-load: scanning the metafile proves no module NAMES electron, not that the graph resolves under plain Node. A dynamic require, a missing native or a top-level throw all pass the scan and fail at runtime. build-orcad now runs the bundle with a bogus flag and requires the argv rejection that only a fully loaded graph can produce. Verified: a bundle that builds but throws on load fails the gate. |
||
|
|
3b6eb03349 |
fix(terminal): stop OMP tab title flapping between OMP and Pi (#16373)
* fix(terminal): stop OMP tab title flapping between OMP and Pi
OMP wraps Pi, and both share the `pi-compatible` title-identity group. Two
writers publish frames for the same pane under different labels: main's
synthetic spinner injects "<frame> OMP" every 80ms, while the wrapped Pi
harness emits its own "Pi" frames.
`isDecorativeAgentTitleFrameChange` keys on `status:textWithoutSpinner`, so
`working:OMP` and `working:Pi` read as meaningful changes. The alternation
defeated spinner-churn suppression entirely: every 80ms frame committed a
store patch plus a runtime-graph sync, on both the tab-title and
runtime-pane-title paths.
Pin same-group identity frames to the tab's launch owner at both store
choke points, reusing the existing owner-normalization helper already
applied on the sidebar, remote-sync, and mounted-pane paths.
The relabel is scoped to bare identity frames ("⠋ Pi", "Pi ready"); a
semantic session title ("π - <session> - <cwd>") carries text no agent
profile can reproduce and is left untouched, so this does not reintroduce
the generic-label complaint in #16093.
* fix(terminal): scope owner relabel to cross-identity frames
A frame that already names the tab's own agent carries authoritative status
wording, so relabeling it restated bare "Pi" as "Pi ready" and changed a
Pi-owned tab that never flapped. Only relabel when the frame names a
different member of the identity group.
Also fixes the repro suite's types against the project typecheck.
Co-authored-by: Seongho.Bak <49228032+psh4607@users.noreply.github.com>
---------
Co-authored-by: Seongho.Bak <49228032+psh4607@users.noreply.github.com>
|
||
|
|
09048c63d4 |
feat(orcad): add headless browser providers (#16193)
* feat(orcad): add headless browser providers * fix(orcad): merge the duplicate runtime-browser type import |
||
|
|
788575e300 | fix(crash-reporting): stop destroying user crash notes (#15252) | ||
|
|
b516300b8c | refactor agent hook listener modules (#16187) | ||
|
|
50438041a8 | fix(rate-limits): safely surface Codex RPC exit reasons (#16023) |