mirror of
https://github.com/stablyai/orca.git
synced 2026-09-30 16:02:56 +00:00
57ef1fecfb4e8c7cb104720a313e69b3fbf2a866
146
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
20a12a6a46 |
perf(codex): share one launch-prep hook install across a spawn burst (#17669)
* perf(codex): share one launch-prep hook install across a spawn burst Codex launch prep runs a full managed-hook install on every local PTY spawn, and both install lanes serialize globally per Codex home. Opening a multi-pane worktree therefore paid N full installs back to back, and a resumed Codex pane prepares twice. Concurrent spawns for the same runtime home now share one run; the promise is dropped as soon as it settles, so the next launch still re-reads hooks.json and the user's trust state. Also split the `host_env` spawn-timing phase, which spanned the entire Codex preamble and pinned that cost on the env builder that ran last. * refactor(codex): unify the two hook-install single-flight lanes Both the WSL and launch-prep lanes now share one generic in-flight helper instead of duplicating the map bookkeeping. Also routes the WSL launch-prep install through the serialized variant, which closes the same per-spawn serialization gap on WSL that the native lane just got. * refactor: extract the shared in-flight run dedupe The codex hook service and the GitHub conflict-summary cache had grown near-identical private copies of the same single-flight helper. Both now use one module, which also keeps the hook service clear of the 300-line budget. The shared copy keeps the identity check on clear so a late settle cannot evict a newer entry for the same key. |
||
|
|
872bd51d47 |
fix(native-chat): reland large structured command results (#17720)
* fix(native-chat): preserve large structured command results (#17707) * fix(native-chat): preserve large structured command results * chore: place native chat validation artifacts under docs * chore: drop stale root package config * fix(native-chat): enforce rebuilt lifecycle append slots --------- Co-authored-by: Merge Sim <sim@local> * chore: omit native-chat reland planning docs * fix(native-chat): remove journal store import cycle * fix(native-chat): keep journal factory acyclic --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
894ed75abb |
Revert "fix(native-chat): preserve large structured command results (#17707)" (#17719)
This reverts commit
|
||
|
|
5fe37729ea |
fix(native-chat): preserve large structured command results (#17707)
* fix(native-chat): preserve large structured command results * chore: place native chat validation artifacts under docs * chore: drop stale root package config * fix(native-chat): enforce rebuilt lifecycle append slots --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
5ff1aa540e |
fix(codex): re-land WSL direct-home cutover with counsel findings fixed (#16854)
* fix(codex): safely re-land WSL direct homes * fix(codex): finish WSL direct-home cutover * fix(codex): coalesce WSL launch hook installs * perf(codex): avoid duplicate retired WSL session scan * fix(codex): retain canonical WSL retired-home path * fix(codex): fail closed before retiring WSL auth * fix(codex): reopen WSL drain after rollback * fix(codex): preserve WSL source on unknown panes * fix(codex): harden repeated WSL runtime drains * perf(codex): bound pending WSL session scans * fix(codex): recover invalid WSL session watermarks * fix(codex): validate retained WSL scan state * fix(codex): accept durable WSL scan state * test(codex): cover the drain's inode-identity guard against destination replacement Removing the four `target_auth -ef temporary_destination_auth` assertions left all 33 apply-script tests passing, so a regression deleting them would have shipped silently. Reproduced before writing this. A hash check cannot catch the case. The pinned hard link keeps the original inode, so it still hashes correctly after another writer atomically renames a different file over the destination path; only inode identity sees it. Without the guard the script exits 0 and retires the source, leaving the user holding bytes nothing validated. The new case asserts the source survives. The harness is split by responsibility so no file exceeds its max-lines budget: fixtures, the coreutils interference shims, the run types, the apply runner, and the recovery/absent runners. The atomic-rename hook is deliberately separate from the in-place rewrite shim because different guards catch them. * fix(codex): keep the split drain harness inside the child-process boundaries Extracting the harness into non-test modules moved it out of the exemptions the single test file had: three new files import child_process, and two spawned without windowsHide. Adds the three to the import allowlist, and sets windowsHide on the spawns rather than exempting them - the flag is correct for these calls regardless of the ratchet, and they are skipped on win32 anyway. --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
f2e9ba453c |
fix(agent-hooks): route reminted pane keys to canonical identity (STA-3993) (#15714)
* fix(agent-hooks): route reminted pane keys to canonical identity (STA-3993) Spawn was stripping $$<base32>:L$$ ORCA_PANE_KEY values (and the launch token) instead of rewriting them to the metadata-proven tab:leaf key, so OMP hooks never entered last-status.json and sleeping rows stayed working. Alias that exact remint form onto the canonical pane so later posts still route, and keep unmatched tokens from stamping another pane. * fix(agent-hooks): keep reminted pane-key aliases first-pane-wins Remint tokens have no embedded tab identity, so a later spawn that reused the same $$ token with a different tab/leaf was overwriting the alias and routing leftover hook posts onto the new pane. Refuse destination changes for that form while still allowing same-pane pty id updates. * fix(agent-hooks): keep pane alias limit import valid after refactor * fix(agent-hooks): bound pane alias destination keys * fix(ssh): keep pane identity env stripped when hooks disabled --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
585b4086d3 |
test(codex): pin Codex read-repair with a real-binary contract check (#17300)
* test(codex): pin Codex read-repair with a real-binary contract check
Orca's session index-heal depends on a Codex behavior: a `thread/read` of an
unindexed rollout performs a read-repair that inserts the `threads` row. All 55
existing heal tests drive a stub app-server and assert "healed" as "the call did
not error", so if Codex ever dropped the repair they would all stay green while
the subsystem went silently inert.
Adds a real-binary contract check built to the same shape as the Git binary
compatibility contract (src/shared/git-binary-compatibility.test.ts): env-gated
test file, version asserted against the binary, dedicated path-filtered PR job.
Pins only the four arms ablation established Orca relies on:
- a read of an unindexed rollout inserts the state row
- a session with no read inserts nothing (the negative control that makes the
insert causal rather than incidental)
- re-reading an indexed thread inserts nothing
- an archived thread stays archived rather than being resurrected
Written against codex-cli 0.150.1. The job sets ORCA_CODEX_CONTRACT_REQUIRED=1
so a missing or failed CLI install fails red instead of silently skipping.
Existing heal tests are unchanged.
* test(codex): register the contract job in the verify aggregate contract
`pr-workflow-parallelism.test.mjs` pins `verify.needs` exactly, so adding the
job to pr.yml without updating that list failed the shard. Adds the entry, and
adds a workflow contract test mirroring `git-binary-compatibility-workflow.test.mjs`:
- the pinned CODEX_CLI_VERSION is the single source for both the npm install
and the runtime version assertion, so the two cannot drift apart
- the install prefix and the binary path the test is pointed at are the same tree
- ORCA_CODEX_CONTRACT_REQUIRED=1 is set, so a failed install fails red rather
than turning the job into a green no-op
Removing the REQUIRED env from pr.yml reddens the new test, confirming it is live.
* test(codex): make binary version guard exact and bounded
* ci(codex): cover index-heal transport dependencies
* test(ci): pin Codex contract dependency coverage
* test(codex): align contract watchdog with child deadlines
* test(codex): cover three-session contract watchdog
* fix(codex): add sqlite sync-database to index-heal scope
---------
Co-authored-by: Merge Sim <sim@local>
|
||
|
|
1369821bad |
Split Codex hook service responsibilities (#17260)
* Split speech session lifecycle * Split terminal output scheduler pipeline * Split mobile browser pane modules * Prune resolved max-lines suppressions * Split pane tree equalization logic * Extract mobile troubleshoot screen styles * Split external automation manager * Split main window service attachments * Split hosted review creation checks * Split automation dispatch event handling * Split settings navigation metadata * Split daemon initialization lifecycle * Split GitLab item dialog * Split relay dispatcher layers * Split mobile host screen * Retarget mobile view settings source test * Split runtime file client layers * Split ports panel layers * Split runtime environments pane layers * Split local PTY provider responsibilities * Split CDP bridge responsibilities * Split relay Git handler responsibilities * Track moved relay Git fetch audit * Split Linear item drawer responsibilities * Split telemetry event schema responsibilities * Split resource usage status responsibilities * Split remote terminal multiplexer responsibilities * Split Git worktree responsibilities * Split Codex hook service responsibilities * Keep mirrored hook trust type private * Fix F3-speech for #17123 * Fix F1-cycle for #17131 * Fix F4-navtest for #17157 * Fix F2-allowlist for #17161 |
||
|
|
2dfaa676d8 | chore: update oxlint and oxfmt (#17150) | ||
|
|
fd9125ea8c |
feat(native-chat): Codex structured native chat restructure (#16729)
* feat(native-chat): port structured Codex sessions from restructure-recovery Rebuilds the desktop structured native-chat implementation from brennanb2025/native-chat-restructure-recovery (tip 4e31c08db3) on top of current main as a single commit, scoped to the local Codex path. Ported: - Structured agent-session core: durable record store + single-writer lease, canonical journal, agent-session wire host/attach/eviction/subscribers, `agentSession.*` RPC surface (registered via ALL_RPC_METHODS; host-side mobile allowlist included for wire compat), pty write gate, transcript additions, and the Codex app-server adapter/launch resolution. - Renderer: NativeChatStructuredSession view/composer stack, structured launch path with the single-flight guard, local structured session tabs sync, activation gate + structured inventory (read-only `agentSession.handoffStatus` probe), agent-session tabs in the tab strip, AI-vault structured session activation, and the settings pane with the parent Experimental Chat UI toggle plus the nested "Use updated structured native chat" toggle. New sessions require both flags, agent codex, no prompt, and a local non-WSL, non-Windows-host execution host (structured-native-chat-availability). - Fixes 72c013cea6 (verified Codex launch recovery), 8ddbaf5e3d (defer native terminal view switching affordances), and 4e31c08db3 (release the launch gate after a visibility retry) with their regression tests, including the third-launch-after-retry guard case. - Cross-version agent-session wire test + CI lane, packaging entries (proper-lockfile, agent-tooling asar excludes), and the wire-compat doc section. Deliberately not ported: mobile/ changes, the Claude structured runtime (only the claude-transcript-branch-proof and claude-structured-owner-identity leaf modules remain, backing the kept TUI-recovery arms), the terminal↔chat adoption/handoff flow (`agentSession.adoptTerminal`/`requestHandoff`, the handoff request engine, TUI adoption machinery, orca-runtime adoption methods), renderer switching affordances and their dead leftovers, the hook/subagent-status refactor cluster, and unrelated branch changes. The crash-during-acquisition recovery path (restart handoff adjudication, restore/reverse re-acquire, lease schema handoff keys) is kept because every plain direct launch depends on it; a trimmed handoff coordinator exposes only status/restore/close. Branch edits that targeted files main has since split (ipc/pty.ts, worktrees.ts, rpc/methods/terminal.ts, useIpcEvents, pty-connection, store/slices/terminals.ts, runtime-types, web preload) were re-applied to the split modules, preserving main's newer logic (Windows CIM fallback, browser tab close rework, cold-restore resume flow, dispatcher threading). Known seam: the mobile clipboard image-provenance CONSUMER gate ships (agentSession.send refuses unproven mobile image refs with agent_session_image_untrusted) but the producer hunk in rpc/methods/clipboard.ts stays with the unported mobile cluster, so mobile image sends into structured chat fail closed until that side ports. * fix(native-chat): trust only authenticated local image uploads * fix(build): preserve Windows process-tree patch application * test(windows): include process creation time in addon fixture * fix(build): run windows-process-tree node-gyp from the physical package dir gyp expands the node-addon-api dependency by probing node, whose cwd resolves to the package's physical directory in the store, so the emitted target is a store-relative ../../../../node-addon-api@... hop. gyp then resolves that hop against the rebuild cwd; from the node_modules symlink/junction it escapes the store and configure fails with "node_addon_api.gyp not found" (run 32999886072). Rebuild from realpath(package dir) so both bases agree, matching how the package manager itself runs native install scripts. The regression test replays gyp's expansion+resolution against the planned cwd and fails without the fix. * fix(native-chat): keep chat tabs visible through terminal closes and empty-worktree launches Two proven blockers in the native Codex tab contract: closeTerminalTab pre-empted the canonical unified close. With one terminal left it deactivated the worktree on a terminal/editor/browser-only check, blanking a workspace that still held a renderable agent-session tab; with two or more it pre-picked a successor from terminal entities only, re-stamping the group active before closeUnifiedTab's MRU/neighbor repair could land on the chat tab. Successor choice now defers to the unified contract whenever the terminal has a unified row, and deactivation is gated on the unified renderable count (matching leaveWorktreeIfEmpty), with the legacy pre-pick kept only for terminals without a unified row. A structured session created on an empty worktree was published into the host's headless group while preserveLocalLayout froze the local layout, leaving the tab in store but permanently off screen. A preserveLocalLayout owner now always takes client-owned placement — repairing a rendered leaf whose group record is missing, or materializing a rendered group on a truly empty worktree — and applies the client-derived layout repair while still rejecting host-authored layout. Regression tests drive the real store through closeTerminalTab (git worktree and folder workspace) and the real snapshot applier for the empty-worktree adoption states; all fail without the fixes. * fix(native-chat): close stale turns and retry rejected sends * fix(native-chat): retire hosted rows on structured tab activation * fix(native-chat): preserve rpc defaults across main merge * chore: format remote wire compatibility guide * test(native-chat): cover retry after unconfirmed send * fix(native-chat): reload outbox on session switch * docs(settings): disclose structured chat platform limits * fix(native-chat): await Codex launch-home preparation * fix(codex): align child-process allowlist with async trust bridge * test(identity): update inventory for tab surface refactor * fix(windows): preserve process-tree CRLF patch sources * fix(native-chat): anchor an unmatched chat echo where it was sent (#16117) * fix(native-chat): anchor an unmatched chat echo where it was sent The reported symptom was old user messages replaying below every new turn, so the conversation read as scrambled. The cause was not that the echo failed to match a transcript row. Claude consumes a mid-turn send through a `queued_command` attachment and writes no `type:"user"` record for it, so some echoes can never match, and no amount of matching will change that. The cause was WHERE an unmatched echo rendered: buildMobileNativeChatTransientData appended every pending item after the entire transcript, so it re-read below each turn that landed afterwards. Render each echo directly after the transcript row it was sent against, using the baseline the send already captures. An unmatched echo is then at worst a duplicate in the right position rather than a scrambled one, and it stays visible. Echoes sharing an anchor keep send order; a send with no baseline, or one whose anchor folding dropped, still falls back to the tail. Deliberately NOT fixed by deleting the echo. Inferring from send ordering that an echo can never match, then removing it, loses the user's own text for a message the agent did receive, and it cannot fire in the common case anyway - measured drain groups are 1,017 of size 1 against 55 larger. It also escalates an existing gap: the count pass has no baseline-tail guard, unlike the glue pass, while `messages` is a 40-row window that head-trims, resets on reconnect and grows at the front on loadEarlier, so a false landing there would license deleting a DIFFERENT outstanding message. That count-pass gap is real and left for a separate change; anchoring makes its worst case a duplicate in place rather than a scrambled conversation. * fix(native-chat): preserve folded echo anchors * fix(native-chat): preserve forward-folded echo anchors * fix(native-chat): keep leading folded echoes in place * fix(workspace-cleanup): show git status for every row (#16690) * fix(native-chat): refuse structured chat on every Windows execution path canUseStructuredNativeChat only refused win32 when a project runtime resolved, so folder-workspace keys (and other keys with no project runtime) failed open into structured chat on Windows. Fail closed on win32 unconditionally after the host check, matching the settings copy: local macOS/Linux only; Windows/WSL/SSH stay on terminal chat. * fix(native-chat): restore runtime refusals behind the win32 gate |
||
|
|
3558cf943f |
fix(codex): heal WSL hooks before typed launches (#16535)
* fix(codex): heal WSL hooks before typed launches * test(codex): keep launcher fixture type-safe on Windows * fix(build): list codex-home-wsl-env in the CLI typecheck project `managed-home-shell-preflight.ts` is already in the CLI project's include list and now imports `wslCodexRuntimeHomeForGuestHome` from `src/main/pty/codex-home-wsl-env.ts`, which the list did not cover — TS6307, so the CLI typecheck failed on every push. Added the single module rather than a `src/main/pty/**` glob: it is a 31-line leaf with no imports of its own, so it does not widen what the CLI bundle can reach. * fix(codex): converge the two WSL hook install lanes onto one writer Two independent readiness reviews agreed the Orca-terminal boundary holds, but Codex Sol found a P1 the other rated P2: the new just-in-time repair raced the existing relay installer and the two produced DIFFERENT hook and trust representations for the same managed home. Two unserialized writers emitting different formats is worse than the bug this PR fixes, because it fails intermittently rather than cleanly — a pane works or does not depending on which lane won. - Relay Codex installs now delegate to the runtime-home writer, so there is one canonical representation instead of two. Redirected scripts use the runtime path, the readable wrapper, and the prepended group. - `installForRuntimeHomeSerialized` puts every asynchronous WSL caller for a given home on one queue (`wslInstallQueues`), so concurrent panes cannot interleave writes. Also rewrites the stale pin test the new `-x` guard broke. It asserted the defect — "would run the impostor if the preflight carried an unqualified command name", expecting the hijack marker to exist. The guard is a security improvement, so the test now asserts the contract: an unqualified preflight is skipped and the marker is never written. Rewritten to the new behavior, not loosened or deleted. 818 tests pass across the affected suites; typecheck clean. The changed-file quality gate could not run locally — its pnpm engine-warning JSON parser fails under Node 26 — so CI covers it. The boundary both reviews verified is untouched: paired/relay/mobile clients stay hard-blocked from the RPC, params remain shape-locked to the managed home suffix with traversal rejection, nothing is written outside the managed home, and macOS/Linux stay inert. * fix(codex): serialize resolved WSL hook homes * fix(codex): recover managed WSL homes after restart * fix(wsl): translate Codex preflight through WSLENV * fix(cli): cover bounded WSL Codex repair * fix(codex): coalesce duplicate WSL hook repairs * fix(codex): verify reconstructed WSL homes |
||
|
|
cc384c5a3d |
fix(agent-hooks): post posix payloads as json (#11292)
* fix(agent-hooks): post posix payloads as json * fix(agent-hooks): mark header merged envelopes * docs(agent-hooks): describe header merge envelope * fix(agent-hooks): encode posix metadata headers * test(agent-hooks): update WSL JSON hook assertions * fix(agent-hooks): negotiate raw JSON transport * fix(agent-hooks): preserve packed metadata in POSIX shells * test(agent-hooks): include hook envelope in relay boundary inventory --------- Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com> |
||
|
|
928d306b53 |
fix(agent-hooks): stop the Windows hook launcher spelling the AV-denied flag pair (STA-5237) (#16739)
* fix(agent-hooks): stop the Windows hook launcher spelling the AV-denied flag pair (STA-5237) `-WindowStyle Hidden` + `-EncodedCommand` is denied at CreateProcess by Kaspersky on Windows 11, whatever the payload decodes to. Bash reports it as `Permission denied` and every managed hook event fails, so agent status never arrives; the parent shell also briefly cannot spawn anything afterwards, so a denied hook can take the user's next command down with it. Measured on the reporting host (#16003), with a harmless `exit 0` payload: -NoProfile -ExecutionPolicy Bypass -WindowStyle Hidden -EncodedCommand 126 -NoProfile -WindowStyle Hidden -EncodedCommand 126 -WindowStyle Hidden -EncodedCommand 126 -NoProfile -EncodedCommand 0 (5/5) #16576 removed `-ExecutionPolicy Bypass`, which is the one flag of the three NOT in the signature, so hooks kept failing after that fix. The pair that has to stop being spelled is `-WindowStyle Hidden` + `-EncodedCommand`. Because the change is to the shared switch constant, it covers every site that spells the denied pair in one edit: Claude via `wrapWindowsPowerShellEncodedCommand`, gemini/cursor/droid/command-code/copilot via `wrapWindowsHookCommand`, the `runtime-home-hook-command` unsafe-HOME fallback, and the spaced-path fallback for codex/grok/devin/antigravity. Only a flag is removed, so parser and payload compatibility is unchanged for every executor: the string is still a PowerShell command line, still one self-contained token, still base64-shielded. The tradeoff, recorded rather than hidden: `-WindowStyle Hidden` was the shipped fix for #14815 (+#14828, #15117, #15447, #15767), and this removes it. Its suppression was never measured — #14825 confirmed it visually, #16576's author stated it "remains unverified on a real box", and #15506's author argued it cannot help a `.cmd` child with no console to inherit. The console is allocated by the parent chain, not by this command line. A live window measurement is still outstanding and is called out in the PR. Also adds `windows-hook-payload-delivery.test.ts` to the PR CI Windows leg, which had never run it. * test(agent-hooks): keep launcher token out of source grep |
||
|
|
9062494f9b |
fix(ai-vault): stop a whole opencode.db failure reading as one skipped transcript (#16587)
* fix(ai-vault): stop a whole opencode.db failure reading as one skipped transcript #15036 reported "1 transcript skipped / database is locked" with both Agent Session History scopes empty. Two separate defects. The panel counts every unkinded scan issue as a skipped transcript, so a failure that lost an entire *source* was reported as one lost *file*. The whole-database failure is now kinded `scope`, and an unknown `kind` from a newer host degrades to `scope` instead of failing validation and coming back unkinded — a mixed-version remote host previously turned a source-level failure into a phantom skipped transcript. The read also inherited sqlite3's 0 ms busy timeout, so a genuinely contended open failed in ~1 ms. It now opens once with a bounded timeout. No retry loop: sqlite's own busy handler already blocks and retries internally for the whole timeout, and WAL readers do not block on a writer at all (measured: 547/547 cross-process reads at timeout=0 while a writer held open transactions). Measured against a real Ubuntu-24.04 distro, Windows cannot take SQLite's file locks over \\wsl.localhost at all: an idle, never-WAL, nothing-attached database still answers SQLITE_BUSY, a 5 s busy timeout does not change it, and the identical bytes open fine once copied to local disk. So a lock-family error on that share never means "a writer holds it" and no timeout can help. The copy says so rather than sending the user after a write-ahead log that is not the problem. Restoring those sessions needs an in-distro read; that is a follow-up, and this PR no longer pretends a timeout will do it. immutable=1 is deliberately not used as a workaround: over the same share it opens and returns 100 of 150 rows, silently dropping everything still in the uncheckpointed -wal — in a history panel, exactly the newest sessions. * skip the provably futile busy wait on \\wsl.localhost paths |
||
|
|
aaef5e8c9f |
fix(agent-hooks): deliver hook events that fire while Orca is restarting (STA-5329) (#16685)
* fix(agent-hooks): correct durable spool delivery * fix(agent-hooks): spool curl failures after retries * fix(agent-hooks): keep replay out of runtime observations * test(agent-hooks): pin managed hooks inert outside an Orca terminal * fix(agent-hooks): address review findings on the durable spool - claude: pass the literal source; options.agent does not exist (typecheck) - kimi: the windows-local ordering runs its guard pre-stdin and before the function exists, so it no longer spools there (printed command-not-found) - writer: require a readable endpoint file before creating a spool tree - antigravity: carry its out-of-band event name into the record and filter on it - drain: truncate only the bytes consumed, preserving concurrent appends and a torn trailing line * fix(agent-hooks): ignore spool events without pane attribution * fix(agent-hooks): make spool replay and appends robust * test(agent-hooks): type spool replay records * fix(agent-hooks): defer unterminated spool records * fix(agent-hooks): replay spool events through relays * fix(agent-hooks): preserve Codex prompt across child replay * fix(relay): keep startup alive when spool replay fails * fix(relay): simplify spool replay startup guard |
||
|
|
9c01e09ecc |
Revert "fix(codex): launch WSL accounts from direct homes" and "refactor(codex): remove WSL runtime mirror machinery" (#16722)
This reverts commit |
||
|
|
ebcd637db9 |
fix(codex): launch WSL accounts from direct homes (#16504)
* fix(codex): launch WSL accounts from direct homes * fix(codex): coalesce WSL auth drains and validate distro homes * fix(codex): preserve legacy WSL account home metadata * fix(codex): retain marked WSL home compatibility * fix(codex): verify the bytes the WSL drain promotes, not an earlier read The apply script validated the source hash and then re-read it with cp, so a legacy pane rotating in that window put bytes freshness never judged over a valid account home. Codex rewrites auth.json in place, so that read can be torn. Covers it by running the real guest script under sh with a sha256sum shim that rotates the source between the two reads; without the guard it exits 0. * fix(codex): harden WSL auth drain races |
||
|
|
07f2e14c08 |
refactor(codex): make WSL account surfaces direct-home aware (#16499)
* refactor(codex): make WSL account surfaces direct-home aware * fix(wsl): keep Codex relay hooks on managed runtime home * test(wsl): assert relay hooks use managed Codex home |
||
|
|
26721bd632 |
fix(codex): stop blocking the main thread on trust grants (#16441) (#16594)
* fix(codex): stop blocking the main thread on trust grants (#16441) Codex hook trust was granted by blocking the Electron main thread on `spawnSync` of a bundled ELECTRON_RUN_AS_NODE entry for the whole app-server deadline: 15s native, 35s WSL, ~45s on the real-home path (rebase inspect + repair + grant). Cold start and every Codex pane launch showed "Not Responding"; the reported event-loop gap was 15,049 ms. The subprocess only ever existed to donate an event loop to a deliberately blocked parent — `runCodexHookTrustGrantSession` was already the real async implementation. Make the callers async and the fork is unnecessary, so the bridge, the forked entry and its envelope are deleted along with their build/knip/tsconfig registrations. The CLI `agent hooks prepare-codex` handler is already async, so it awaits the in-process session and saves a process spawn per managed-home shell. `resolveCodexTrustGrantHost` is async too; the WSL identity probe moves from `execFileSync` to `runProcess`, dropping that file from the child-process import allowlist. Status reads keep a synchronous native-only stamp path. Two invariants that held only because the lane blocked: - Overlapping capability probes were impossible by construction. `GitCapabilityCache`'s dedupe engine is extracted to a shared `CapabilityProbeCache` and `CodexAppServerCapabilityCache` now inherits it, so concurrent launches against a cold host share one app-server session instead of one each. - Two grants on one `config.toml` could not interleave capture and restore. A reentrant per-file lane now serializes the whole install sequence (managed, WSL runtime, real-home ensure, legacy sweep) and the grant and rebase inside it. Cold-start work moves off the critical path: retained-home reconciliation (N sequential sessions) is fire-and-forget behind the daemon provider, and the startup real-home ensure chains into managed hook reconciliation instead of blocking app init. Every preserved semantic is unchanged: never throws, the ORCA_DISABLE_CODEX_TRUST_RPC kill switch, ledger hits, backfill-pending and cooldown fallbacks, config rollback on every failure path, pre-grant self-computed trust removal, the verify-failure taxonomy, diagnostics and telemetry. * fix(codex): widen the trust-config lane to every config.toml writer Review follow-ups on #16441's async trust grant: - `markCodexProjectTrusted` now runs inside the runtime+system config.toml lanes, so a project-trust write can no longer land inside a hook grant's capture->restore window and be silently reverted. Its callers await it. - `install`/`refreshRuntimeUserHooks`/`remove` hold the system config.toml lane as well as the runtime one — they promote approvals into ~/.codex/config.toml and mirror it back. Lock order is runtime-before-system everywhere. - The real-home ensure chain resumes after a rejection instead of returning the same rejected promise to every later pane launch, and resolving the real home is now inside the module's never-throws boundary. - `buildSpawnEnv` awaits inside a cancelable pending-spawn registration, so shutdown during the (now long) env build stops the PTY from launching. `prepareLocalPtySpawn` generalizes into `awaitCancelableLocalPtySpawn`. - CapabilityProbeCache drops the test-only `nowMs` passthrough; its probe backstop comment now describes what it actually guards. - Preflight is a plain async function; the trust dispatch in orca-runtime collapses into one `markWorkspaceTrustedForAgent`. * test(codex): exercise the trust-config lane under real concurrency The async grant makes two pane launches overlap for the first time. These drive the real modules end to end on real files: a rollback swallowing a sibling's grant, a markCodexProjectTrusted write landing inside a capture -> restore window, shared capability-probe dedupe on a cold host, the host-scoped transient cooldown, and reentrancy from inside an installer. Each was verified to fail against a deliberately broken implementation (lane removed, dedupe disabled, cooldown made global, reentrancy pass- through disabled). * test(codex): stop hook-service suites spawning the developer's real codex The forked grant bundle never existed under vitest, so the RPC lane was unreachable in tests on main. Running it in-process makes these suites spawn a real `codex app-server` when one is installed: 38 spawns and two failures in hook-service-runtime-trust-repair on a machine with codex, green in CI where there is none. Stand in for the missing binary so both environments exercise the same fallback lane. * docs(codex): scope the trust-RPC kill switch comment to what it actually gates The comment read as though the flag forces the fallback lane everywhere. It gates the managed grant only: the real-home rebase still runs its own inspect/repair app-server sessions when Orca's insertion shifts a user's hook positions, and never reads the flag. Verified by exercise, not by reading — with the flag set, both inspect-user-hook-trust and repair-user-hook-trust still ran. Pre-existing: main has no check there either, it just blocked the main thread while doing it. Widening the flag to cover the rebase is a follow-up; this only stops the comment promising something the constant does not do. |
||
|
|
015f904fca |
fix(codex): stop re-scanning all Codex session history on every launch (#16251) (#16593)
* fix(codex): stop re-scanning all Codex session history on every launch (#16251) A launch deleted the backfill completion marker, and a marker could never be written while a Codex pane was open, so every launch re-derived "needs full scan" and walked the entire .codex/sessions tree — on Windows with a large history that read as a hung window. - v4 marker keeps a durable full-history baseline plus a bounded set of pending dates. v3 is read as a baseline, so upgrades pay no full scan. - A launch now marks dates pending instead of deleting the marker, and a full pass certifies the baseline even while a pane is still running; the live pane's own date just stays pending. - Pending dates are persisted, so an abnormal exit or a cross-midnight pane recovers a bounded window instead of a full walk. - A date-limited pass can only extend an existing baseline, never create one, so it can no longer certify history it never looked at. - Marker and index-heal target roots compare through normalizeRuntimePathForComparison, so Windows spellings of one directory stop invalidating each other. - Both append-only ledgers stream instead of readFileSync + whole-file JSON.parse, keeping the main thread responsive on large histories. * fix(codex): keep the backfill marker's full-scan demand durable Review follow-ups on the v4 backfill marker: - markCodexSessionBackfillMarkerPending no longer erases a persisted needsFullScan; the demand survives until a generation-current full walk retires it, and the function now reports it so the launch path folds it into its own in-memory flag (as @rumoii's #16252 does). - A full pass settles the whole pending set instead of subtracting the empty set, so a date a full walk provably covered stops forcing an extra bounded pass on every startup. - isCodexSessionBackfillDate does a real calendar check, so a corrupted marker cannot carry 2026/99/99. No age or future bound: the same guard gates rollout publication and a clock-skewed directory holds real sessions. - 'scans only the current date once a baseline exists' now has a second date directory, so it fails on a full walk instead of passing either way. |
||
|
|
e8005c3325 |
fix(codex): preserve WSL account home trust (#16496)
* fix(codex): preserve WSL account home trust * fix(codex): preserve WSL drive path semantics * fix(codex): preserve mounted-drive WSL config paths * test(codex): preserve WSL path helpers in mock |
||
|
|
c1a7748267 |
fix(agent-hooks): stop the Windows hook launcher spelling the AV-denied flag triple (#16576)
* fix(agent-hooks): stop the hook launcher spelling the AV-denied flag triple
Orca's Windows agent-hook launcher ran
powershell.exe -NoProfile -ExecutionPolicy Bypass -WindowStyle Hidden \
-EncodedCommand <base64>
That exact combination is the textbook "hidden encoded PowerShell" malware
shape, and endpoint security denies it at process creation whatever the
payload decodes to -- even `exit 0`. Every injected hook then failed with
`powershell.exe: Permission denied` (exit 126 from bash's execve, EACCES)
on every turn, for both Claude Code and Codex, with no AV exclusion that
re-enabled it.
Dropping any one of the three flags clears the signature. `-ExecutionPolicy
Bypass` is the one that can move: it sets the Process scope, and so does
`Set-ExecutionPolicy -Scope Process`, which now rides inside the encoded
payload. `-EncodedCommand` is never policy-gated, so the bypass always gets
to run before the managed script does -- which is what keeps Copilot's .ps1
hook working under a Restricted or AllSigned machine policy.
The hidden window and the encoding are unchanged, so nothing regresses for
#14815, #14818 or #6078.
Closes #16003
* fix(agent-hooks): ship the launcher shape #16003 actually measured as allowed
The previous revision of this branch dropped only `-ExecutionPolicy Bypass`
and kept `-WindowStyle Hidden -EncodedCommand`, on the reasoning that
"dropping any one of the three flags clears the signature". That sentence is
not in the bisect. The reporter ran exactly four command lines on the affected
Kaspersky/Windows 11 host:
-NoProfile -WindowStyle Hidden -Command 'exit 0' -> 0
-NoProfile -EncodedCommand <b64> -> 0
-NoProfile -ExecutionPolicy Bypass -Command 'exit 0' -> 0
-NoProfile -ExecutionPolicy Bypass -WindowStyle Hidden -EncodedCommand -> 126
Every passing row drops two flags. No row drops exactly one, so the shape the
branch was about to ship had never been executed on the machine that reports
the bug -- and it is `-WindowStyle Hidden -EncodedCommand`, which is the
"hidden encoded PowerShell" pair the denial is named for in our own comment.
Shipping it would have closed #16003 while leaving every hook on that host
dying at CreateProcess, with no tracking left open.
So emit the measured-passing encoded row instead: `-NoProfile -EncodedCommand`.
Of the two flags there was a choice between, `-EncodedCommand` is the one that
carries correctness -- it is what keeps paths and switches intact across
cmd.exe and MSYS (#6078, #14815). `-WindowStyle Hidden` costs at most a console
flash, and only where the parent has no console to inherit.
Second, the relocated bypass now runs inside try/catch. Under a MachinePolicy
or UserPolicy GPO scope, `Set-ExecutionPolicy -Scope Process` reports that the
process scope did not take. `-ErrorAction SilentlyContinue` covers only the
non-terminating half of that; the command-line switch it replaces was silent
either way. This file already documents that non-stdout PowerShell streams
corrupt consumers merging our output into JSON stdout, so a per-invocation
ErrorRecord on stderr is a regression we should not trade for the switch.
Refs #16003
* fix(agent-hooks): keep the hook console hidden while dropping the AV-denied flag
Round 2 of this PR widened the fix from "stop spelling -ExecutionPolicy Bypass"
to "stop spelling it and -WindowStyle Hidden", on the reasoning that the #16003
reporter never measured a shape that drops exactly one flag, so keeping the
hidden+encoded pair would be extrapolation.
That trades a reproduced regression for an unmeasured one. Window suppression is
the shipped fix for #14815 and its four duplicates (#14828, #15117, #15447,
#15767): a hook launched from a parent with no console gets a fresh console per
event, which takes foreground and eats whatever the user is typing into Orca,
and never closes at all on the stdin-blocking path hook-stdin-contract.ts exists
to guard. That fires on every prompt, tool call and stop of every managed agent.
The AV denial, by contrast, is measured only for the full triple; that the
remaining pair still trips it is a hypothesis. Between a certain regression and
a possible one, keep the certainty.
So the flag that leaves the command line is the policy bypass alone — the only
one of the three with an exact in-payload equivalent, hence the only one that
can move without losing behaviour. If the pair turns out to be denied too, the
answer is a different shape that still hides the window.
* fix(agent-hooks): silence progress before the policy bypass can autoload (#16621)
Hardware-measured on Windows 11 while exercising #16576.
Set-ExecutionPolicy autoloads Microsoft.PowerShell.Security, and that module's
"Preparing modules for first use." progress record is written before any later
assignment can suppress it. Running the bypass first therefore defeated the
silencer that runs immediately after it:
bypass-first stderr = 616 bytes, first merged line '#< CLIXML'
silencer-first stderr = 0 bytes, first merged line '{"decision":"approve"}'
That is precisely the corruption HOOK_PROGRESS_SILENCER's own comment warns
about -- redirected progress becoming CLIXML that can corrupt merged JSON -- so
the PR reintroduced the hazard it documents, one line below documenting it.
Both existing tests asserted the broken order, so they enforced the bug rather
than catching it. Reordered them and added one that pins the ordering itself
rather than the literal string, since the string will drift again.
|
||
|
|
48e63c015f |
refactor agent config and auth services (#16195)
* refactor: split agent config and auth services * chore: repoint wsl and global-fetch guards at split module paths * fix: restore merge-base Claude CLI error propagation Drop the secret-redaction rewriting added to Claude CLI error paths in the refactor: spawn errors again reject with the original Error (preserving .code/.errno/.syscall/.stack) and command output/auth-status logs are no longer rewritten. |
||
|
|
2b1b094aa8 |
fix(cli): pair every resolved CLI with its runtime, and ratchet it (#16383)
Follow-up to #16365, which paired 8 spawn sites by hand. Hand-pairing is how the class got introduced, so close it structurally instead. cliPath is now required on CodexAppServerInvocation, `null` only for the guest-side wsl.exe launcher where a host path pairs nothing. Optional let a native builder omit it and silently fall back to pairing against a cmd.exe wrapper with no type error. Every production site already passed it; only test fixtures needed updating, which is the type doing its job. Four more sites now pair. codex-state-db-backfill-recovery spawns the same `codex app-server` subcommand #16365 fixed elsewhere. cli/handlers/account was the worst case: addAgentNodePaths prepends the *newest* version-manager bin, which is not necessarily where the CLI being launched lives, so it actively created the mismatch — pairing now runs last so the CLI's own node wins. commit-message-text-generation and skills/skill-update-run spawn resolved binaries with inherited env. cli/handlers/skills had grown its own buildNpxPath: a weaker local copy that prepended unconditionally, ignored the Windows `Path` key, and special-cased a '.' dirname. Deleted in favor of the shared helper, which checks the sibling node actually exists — the behavior change one test had pinned. The ratchet is the point: any file that resolves a CLI and spawns must reference withCliRuntimeOnPath, with a shrink-only allowlist. It caught skill-update-run, which I had missed. Its first draft required a call paren and so let dependency-injected resolvers (`resolveCommand: resolveCodexCommand`) through — verified by removing a pairing and watching it stay green, then widened until it failed. A second assertion fails on a stale allowlist entry so an exemption cannot outlive its reason. external-editor-launch stays allowlisted: it launches a GUI editor, not a Node CLI whose ABI matters. |
||
|
|
a7505fd911 |
fix(cli): spawn a version-manager CLI with its own node runtime (#16365)
* fix(cli): spawn a version-manager CLI with its own node runtime resolveCliCommand falls back to scanning every version-manager install when PATH misses, so it can hand back ~/.nvm/versions/node/v20.x/bin/codex while PATH still leads with v22. Nothing paired the binary with the runtime it was installed against, so its `#!/usr/bin/env node` shebang loaded a v20-built native module under a v22 ABI and the agent died on first require (#10932). Reproduced with a real addon rather than asserted: a CLI requiring a cpu-features build for NODE_MODULE_VERSION 115, spawned with v24 leading PATH, fails with ERR_DLOPEN_FAILED and exit 1. With the CLI's own bin directory prepended it runs clean. withCliRuntimeOnPath prepends the resolved command's directory when that directory ships a sibling node, and is a no-op otherwise — so a Homebrew or /usr/local CLI is untouched, and the WSL paths pass a bare `codex`/`claude` that is not absolute and so never matches. Host CLI resolution in the Claude login path is now lazy, keeping the WSL branch from resolving a host binary it never spawns. * fix(cli): split PATH on the delimiter we join with, pair app-server too Readiness review findings, all four addressed. withCliRuntimeOnPath chose its join delimiter from the platform option but split with the host's. Passing platform:'win32' from a posix host turned `C:\Windows;C:\Windows\System32` into `C;\Windows;C;\Windows\System32` — every drive letter torn off at its colon. Latent, since no shipped caller passes platform, but the sole win32 test was written against the corrupted value and asserted one split segment, so it green-lit the shredding. That test's other assertion was vacuous: it seeded only `Path`, so the `PATH` key it asserted absent could never exist. Deleting the whole case-dedupe block left the suite green. It now seeds both keys and asserts the full joined string; removing the block fails it. Nothing covered the wiring, and the argument choice is the easy thing to get silently wrong. Note it only diverges on win32 — on posix getSpawnArgsForWindows returns the CLI itself, so pairing the spawn command is indistinguishable there. The new test drives the win32 branch with a .cmd fixture; pairing spawnCmd or dropping the wrapper both fail it now. codex-trust-grant-host and codex-session-index-heal spawn the same `codex app-server` subcommand through runCodexAppServerSession and were left unpaired. Pair centrally there via a new optional cliPath, since invocation.command may be a cmd.exe wrapper. Pairing tests live in their own file: adding them inline pushed codex-fetcher.test.ts past the 800-line ratchet. * fix(cli): read the Windows path key the child will actually use Round-2 review finding. The read was narrower than the delete: the key was picked from exactly two spellings (`Path`, else `PATH`), while the twin dedupe removed every key whose lowercase form is `path`. A block spelling it `path` or `pATh` therefore had its value deleted without ever being read, handing the child a PATH containing only the CLI's own directory — a strictly worse outcome than not pairing at all. Win32 resolves env names case-insensitively and object order preserves block order, so the entry the child reads is the first case-insensitive match. The repo already encodes that rule in resolvePathEnvKey (src/main/pty/windows-path-segment-merge.ts); src/shared cannot import from src/main, so mirror it locally. Verified by execution across six env shapes: lowercase, mixed-case, Path-only, PATH-only, both twins, and a PATHEXT control that must not be touched. All preserve the original PATH; before the fix the first two lost it entirely. Reverting the selector fails the new test and nothing else. |
||
|
|
8d08d0078d |
fix(codex): stop an unreadable legacy hooks.json from clearing managed trust (#16147)
`cleanupLegacySystemManagedHooks` reads `~/.codex/hooks.json` and, when it finds
no hooks, removes Orca's managed trust entries from the system config.toml and
deletes that home's grant-ledger record.
Since the STA-4823 read classification landed, `readHooksJsonWithRaw` reports a
genuine absence as `{ raw: null, config: {} }` and a failed read as
`{ raw: null, config: null }`. Both still fell into the same branch, so a read
that merely failed discarded hook approvals the user had already given — and the
ledger record that would have let a later pass notice.
Only the definitive-absence answer may reach the removal now.
Found while reviewing #15417 and not covered by it: that PR fixed the classifier
and this is a consumer of it that still collapsed the two answers.
The sibling `assertHooksJsonGeneration` guard in codex-real-home-hook-install.ts
has the same `existsSync ? read : null` shape and is deliberately NOT changed
here. Measured, a permission denial leaves existsSync true and throws from the
read, so that path already fails closed; a guard there could not be made to fail
in a test and would be unprovable code.
|
||
|
|
677718c4a5 |
fix(codex): stop rebuilding shared Codex state from a read that failed (STA-4823) (#15417)
* fix(codex): stop rebuilding shared Codex state from a read that failed (STA-4823) Six shared files were rebuilt, erased or reported healthy after a read that had only failed. Batch A of the STA-4606 split: every one of these is reachable and testable on the host lane, so none of them wait on the WSL work. - `config-toml-trust.ts` upsertHookTrustEntries: `existsSync` reported a locked config.toml as absent, so the base content became '' and the upsert rewrote the file from the trust entries alone — a trust-only stub, with the user's model, provider, MCP servers, approvals and comments gone. It refuses now; every hook-service caller already turns that into "trust entries could not be written. Run /hooks in Codex to approve." - `codex-trust-grant-ledger.ts`: an unreadable ledger degraded to empty and the next write persisted a file holding only the home being written, dropping every other home's grants. The write paths refuse; the read path still degrades, and a corrupt ledger is still rebuilt. - `codex-pane-account-registry.ts`: an unreadable registry erased every pane's attribution AND cached that erasure, so it survived the file recovering. The failure is no longer cached, and both write sites refuse rather than persist a registry derived from an empty stand-in. - `hooks-json-read.ts`: the read arm already separated "no hooks" from "could not read", but the `existsSync` arm in front of it returned a valid empty config for a file that could not be opened. One read now classifies both. - `config-settings-baseline.ts`: absent, unparseable and unreadable all collapsed into `null`, so the snapshot rebuilt a baseline it could not read — recording an in-Codex edit as Orca's own write, after which promotion skips it forever. - `config-sync-stall.ts`: an unreadable runtime config read as absent and the status reported `synced` while the mirror was refusing. It reports `managed-home-unavailable`, the existing reason for exactly this, rather than borrowing a source-side one and blaming the wrong path. Absent and malformed still rebuild throughout — resetting corrupt state is the intent, and conflating it with unreadable would wedge a user on a broken file. * fix(codex): close shared state read-denial gaps * fix(codex): recover oversized settings baselines * fix(codex): name the stalled managed config * fix(codex): preserve hooks after failed source reads * test(codex): correct what the denyExistence rig actually models MEASURED on both platforms: a file-permission denial leaves existsSync TRUE and fails only the content read — chmod 000 gives EACCES on macOS, icacls /deny (R) gives EPERM errno -4048 on Windows, with stat/lstat succeeding in both. The rig's docblock claimed this mode modelled that denial. It does not. What it models is the UNC / \\wsl$ transport, where an unreachable distro reports errno UNKNOWN at every level and existsSync folds it to false. The distinction decides what the D29 guard is worth: under a permission denial the pre-fix code already failed safe, because existsSync was true so it took the read branch and threw. Only a transport that lies about existence reaches the rebuild-from-empty path. No behaviour change; the comment was wrong, not the code. * test(codex): exercise live baseline read denial * fix(codex): retry pane attribution writes * fix(codex): retry reconciliation registry writes * fix(codex): report unreadable sync baselines |
||
|
|
e9e238c883 |
refactor(wsl): delete the environment-policy layer the reviews kept failing on (#16007)
* refactor(wsl): delete the environment-policy layer the reviews kept failing on
A design council (Opus, Grok, GPT-5.6-Sol) reviewed the merged runner after it
took eleven review rounds to land. All three reached the same conclusion: the
invocation half is sound, the environment/probe half is not, and every round had
been debugging the second one.
The finding that settled it, from Opus: `environmentResolved` had **54
references, all in tests and the runner itself. Not one production reader.** The
safety mechanism the strict default existed for was never wired to anything, so
all 19 degrading sites reported absence with full confidence anyway -- #9725
live at every one, under comments claiming it was handled. Two of those comments
say so out loud; I wrote them.
Root cause, in one line: every knob existed only because a failed probe was
fatal. So it no longer is.
- `allowDegradedEnvironment` and `WslGuestEnvironmentUnavailableError` are gone.
A missing login PATH is a fact in the result, not an exception. That deletes
23 opt-outs, six catch-and-remap blocks, the transient/rejected cooldown
split, `probedWithBudget`, and the 1.5x re-probe heuristic -- none of which
had a reason to exist once the case stopped throwing.
- `lane` + `allowDegradedEnvironment` collapse into `loginPath: 'none' |
'preferred'`. 19 of 23 sites passed the opt-out, and two said in comments that
they did not want the login PATH at all: the flag had become the `'none'` the
union was missing.
- The `interactive` lane is deleted. It had zero production callers and kept ~30
lines of fence plumbing alive for tests only.
Net -98 production lines; the runner itself sheds 86 for 38.
Also carries three fixes from the W3 orphan-PR sweep I had not done:
- `WSL_UTF8=1` in the runner. My relay migration deleted the only place setting
it, so wsl.exe's own error text arrived UTF-16LE and read as NUL-riddled.
A regression I introduced. Credit: #9010 (Chang-Jin-Lee).
- `GITLAB_HOST` is now named in WSLENV, so a ported self-hosted host actually
crosses into a distro-routed glab (#12557). Credit: #12558 (makoto-developer).
- The WSL skill-setup command pipes into `sh` instead of `eval "$(...)"`, whose
nested quoting produced `word unexpected (expecting "in")` (#14292). Credit:
#14785 (innocarpe).
* fix(wsl): restore the login PATH for the Codex availability lookup
loginPath:'none' on a PATH lookup reports an nvm-installed codex as absent,
which is #9725. A miss without a resolved environment is now 'could not
check', not 'not installed'.
Also hardens the guards that should have caught it:
- bashism ratchet is per-call, not per-file, and fails closed on lexer desync
- blankStringContents handles regex literals (an apostrophe in /'/g desynced
the lexer, so the scan silently found zero calls)
- windowsHide allowlist 85 -> 80, stale once the lexer parsed those files
Credit: Grok (P0), GPT-Sol (ratchet gaps).
* test(wsl): close the two ratchet gaps that let planted spawns pass
- variable-indirected wsl.exe (`const b = 'wsl.exe'; spawnProcess(b)`) is now
tracked, so the 5 files recorded only in a comment become real allowlist
entries. Three actually spawn that way; the other two never spawned wsl.exe
at all, so the prose record was wrong by three in the hiding direction.
- promisify(renamedAlias) is now resolved, so `const run = promisify(execFile)`
behind an `execFile as x` import can no longer skip windowsHide.
Each verified by planting the violation, watching it fail, restoring, watching
it pass. Credit: GPT-Sol.
* fix(source-scan): stop the regex-literal reader from eating block comments
At index 0 there is no preceding token, so a file opening with a banner
comment had its `/*` read as a pattern and swallowed to the next slash --
110k characters of preload/index.ts, in the direction that hides offenders.
Measured across the tree, old lexer vs new: worst-case over-blanking drops
from -110564 to -1116 characters, and files that desync drop from 51 to 22.
The remaining extra blanking is regex interiors, which is the intent.
Regression tests for both lexer bugs, each verified to fail with its fix
reverted. The first draft of the comment test did not bind -- it asserted on
text after the swallowed span.
* fix(wsl): restore the unverifiable signal on the two remaining probe sites
Round 2. Three call sites used to throw when the login-PATH probe failed;
the redesign rewired one (Codex) and left two reporting confident absence.
- skill-wsl-provider-detection: the script ends in `|| true`, so a lookup
without the login PATH exits 0 with empty stdout -- identical to 'nothing
installed'. Callers skip the ~/.codex and ~/.claude skill roots on an empty
list, losing an nvm-installed provider's skills.
- wsl-cli-installer: the dead catch is replaced by an explicit check. Its
`case ":$PATH:"` probe otherwise answers from the distro default PATH and
Settings states as fact that the CLI is not on PATH. Timeout is checked
first, since a timed-out run also leaves the environment unresolved.
Also narrows the regex-literal prev-token set. '!', '+', '-', '>' and '}' are
value terminators as often as operators, so postfix `n-- / 2` and JSX
`<A size={14} /> : <B` were read as patterns and their spans blanked -- 13
live JSX spans, and one swallowed execFile call that left no desync behind.
False negatives only risk a desync, and desync fails closed.
Plus: WSL_UTF8 on the probe spawn (#9010 reached the runner, not the probe),
and the allowlist header I shuffled by sorting comments along with entries.
Credit: Grok (both P1s), Opus (lexer false positives).
* docs(wsl): drop the lane comments the redesign made false
The interactive lane is gone, so 'both lanes' and the fenced-stdout note
described code that no longer exists. Also states plainly that
environmentResolved is always true under loginPath:'none' -- the field cannot
rescue a PATH lookup that was mislabelled, which is how #9725 came back.
Credit: Grok.
* fix(wsl): stop piping user scripts into the shell's stdin
The W3 migration moved hooks from `wsl.exe --exec bash -c <script>` to a
script piped into `bash -s`. Anything the script runs that reads stdin then
drains the rest of the script, bash hits EOF and exits 0, and the caller logs
success -- an orca.yaml hook of `ssh -T git@github.com || true` followed by
`pnpm install` silently never installs.
Scripts now travel in argv by default, which is what the pre-migration code
did and what --exec makes safe. `scriptDelivery: 'stdin'` stays for the one
caller that needs it: the hook-relay installer embeds a base64 JS bundle far
past any command-line limit, and reads no stdin.
A runner test already described this exact EOF hazard -- for the login shell,
not for the guest command it was itself creating.
Credit: code review.
* fix(skills): make the unverifiable check unconditional, and stop double-probing
Round 3.
- provider detection threw only on an EMPTY result, so a degraded partial hit
slipped through: `claude` visible on the default PATH via Windows interop
plus an nvm-only `codex` returns a plausible ['claude'], and the caller then
skips the ~/.codex skill roots for a provider that is installed. The
installer already got this right with an unconditional throw.
- three sites asked for 'preferred' without needing it. The GROK_HOME probe
runs its own `"$login_shell" -lc`, so the runner's probe was a second login
shell eating up to half an 8s budget; the two skill scans are
find/base64/head/printf/stat over $HOME.
- the indirection binder missed `private readonly x = 'wsl.exe'` (the
modifier was captured as the name), backtick literals, and
`spawnProcess(this.x)`. Commit
|
||
|
|
5651662494 |
fix(wsl): migrate 21 call sites onto the WSL runner (#15923)
* fix(wsl): migrate 21 call sites onto the runner, after five review rounds Rebased onto main now that the runner (#15903) has landed. 21 sites across 15 files move off ad-hoc `execFile('wsl.exe', ...)`. Allowlist 23 -> 16 on the WSL guard; 163 -> 152 on the W1 child_process guard, which moved as a consequence. Five review rounds, each finding real defects -- several introduced by the previous round's fixes: 1. Hooks ran user orca.yaml scripts under dash; probe failure fell back to the login shell, reintroducing the ~/.profile stall the runner exists to remove. 2. An unparseable probe was cached permanently, disabling every WSL feature on the distro; hooks regressed from "runs degraded" to "fails". 3. Exit 127 had no expiry; a starved 5s probe hard-failed the 10s scan behind it; a joiner burned its budget on someone else's probe. 4. The comment stripper blanked live code, so the windowsHide guard walked past a real unguarded spawn and reported the file clean; an ownership-probe timeout silently deselected the user's Claude account. 5. Verification of the guards themselves. The recurring finding -- a call answering "is this installed?" on a degraded PATH -- was eventually fixed structurally rather than per-caller: the runner refuses an unresolved guest PATH unless the caller opts in. Per-site vigilance was demonstrably not holding; 3 of 8 sites had already forgotten the analogous exit-code check. Remaining 16 files need a runner mode that does not exist: a long-lived streaming child (OAuth logins, hook relay), a synchronous caller, or a host-level flag like --status that the guest-command API cannot express. * fix(wsl): close round 5's P1s -- degrade where PATH was never needed Round 5 measured the guards by re-executing their algorithms standalone rather than reading them, and found four things. P1 -- four skill/plugin paths gained a hard dependency on the login-shell probe that they never had. They ran under a plain non-login `sh -c` on main, so a probe failure now breaks WSL skill discovery and install on exactly the distro the runner was built for: one with a slow `~/.profile`. Worse, the throw escapes before each site's own error mapping, so the UI gets a raw internal string. They degrade now, per the rule this branch already wrote down in `wsl-fish-history-cleanup.ts`. P1 -- Codex and Claude were asymmetric. Claude's five credential sites degrade; Codex's were strict, so adding a WSL Codex account failed where adding a Claude one succeeded. Three of the four are byte-equivalent to Claude sites, and their scripts read `$HOME`/`$WSL_DISTRO_NAME`, which wsl.exe supplies without a login shell. `assertWslCodexCliAvailable` stays strict on purpose -- that one really does answer "is this installed?" (#9725). P1 -- the ownership-probe timeout fix did not survive the rebase onto main. A timeout still returned "not owned", which the caller *persists*, clearing the user's account selection. P1 -- `blankStringContents` desynced on a nested template literal (`` `${`x`}` ``), leaving 116 lines of a child_process importer outside the ratchet, with 27 importers structurally at risk. Now tracks template depth. Regenerating against the fixed blanker: 70 -> 68 offenders. Also: the windowsHide vacuity check could not fail while the allowlist alone exceeded its bound -- the exact defect the sibling guard documents avoiding. It now names a file that definitely offends. * fix(wsl): close round 6 -- my blanker fix had traded a false positive for a miss Round 6 re-derived the guard's answer from a TypeScript AST instead of trusting the regex, and caught two things. P1 -- the nested-template fix I shipped in round 5 introduced a worse bug than the one it closed. Switching to "code mode" inside `${...}` without also resetting the quote at a newline meant an apostrophe in a regex literal -- `` `'${value.replace(/'/g, "'\\''")}'` `` , which is exactly the shellQuote shape all over this codebase -- inverted the lexer for the rest of the file. `claude-accounts/service.ts` went blind from line 96, hiding a REAL unguarded `spawn` at :1097: the WSL Claude managed-login path, which opens a console and steals foreground on Windows. Round 5 traded one false positive for one false negative and I did not notice, because the offender count went down. The blanker now resets non-backtick quotes at a newline (the rule stripComments already had) and tracks brace depth per interpolation. The spawn is fixed rather than allowlisted, and the count is 69 -- the number the AST predicted. P1 -- the ownership-timeout guard was dead code: it threw into its own `catch` three lines below, which returned null, which the caller persists as "not owned" and clears the user's account selection. Now a typed sentinel the catch rethrows. P2 -- `WslGuestEnvironmentUnavailableError` reached the UI verbatim from the CLI installer and the Codex availability check. Both mapped. Method note: I had been regenerating the allowlist with a Python transcription of the scanner, and the two drifted -- the same two-implementations problem this workstream keeps finding. The allowlist is now generated by running the shipped test with an empty list and taking what it reports. * fix(guards): stop patching the lexer -- make the scanner fail closed instead Round 7 proved my round-6 fix also did not work, by planting a plainly-named unguarded `spawn` in `claude-accounts/service.ts` and watching the guard pass 3/3. That is three consecutive attempts at an exact lexer, each shipping a desync that hid real calls, and each time the offender count went DOWN, which I read as progress. Round 6's diagnosis was wrong too: the culprit is the `templates` brace-depth stack, which nothing resets, not quote state. So stop trying to be exact. `blankStringContentsDesynced` reports when the lexer lost its bearings, and the guard treats that as an offender. Over-reporting is a nuisance; under-reporting is a false clean, and a false clean is what let a real console-flash spawn out of the ratchet twice. The allowlist goes 69 -> 82: the 13 extra are files whose scan cannot be trusted, now named rather than assumed fine. The planted violation is now caught. Also from round 7: - `SPAWN_CALL` missed promisified and renamed bindings, so `exec('where gemini')` (a real Windows cmd.exe spawn) and a detached `shell: true` in `cli/runtime/launch.ts` were invisible. Added execAsync/execFileAsync/ execFileCb/spawnDetached. - `BASHISM` matched `set -o pipefail` but not `set -euo pipefail`, which is the only spelling this tree uses -- so the check could not have caught the #14292 signature it exists for. Fixed, and it immediately flagged a file; that one turned out to be a comment, so the bashism scan now strips comments too. - The CLI installer error mapping my round-6 commit claimed was "both mapped" was never applied -- only the Codex side had been. Now actually mapped. * fix(guards): close the four holes round 8 found by planting violations Round 8 stopped reasoning about the guard and planted spawns into it. Four holes, none of which reading had found: - `windowsHide: false` **passed**. The check was `args.includes('windowsHide')`, a substring test. Now matches `windowsHide: true`. - A ternary first argument was silently skipped: the method-declaration filter `/^\(\s*\w+\s*[:?]/` also matches `exec(useAlt ? 'a' : 'b', …)`. Now requires a type after the colon. - Renamed bindings were not covered, despite the comment I wrote saying they were -- I had hardcoded three names. Aliases are now resolved from the import. Each is verified closed by planting it and watching the guard fail. `fork` is deliberately still unscanned. Round 8 is right that Node forwards the option, but `ForkOptions` does not declare it, so the two live sites cannot be fixed without a cast. Recorded in the verification doc rather than left as a silent gap, along with two others worth knowing: the allowlist is file-granular, so its ~18 false-positive entries carry a standing pre-approval for real regressions in those files and cannot be retired by fixing code; and `stripComments` has no desync report, so the fail-closed check is only half applied. The doc now also says how to verify a guard change: plant a violation. Every guard fix here that was verified by reading was wrong. * fix(wsl): stop preflight reporting installed CLIs as absent on a slow distro Round 9's merge blocker, and the sharpest finding of the whole workstream: the branch built to close #9725 had reopened it from the other side. `preflight-wsl-command.ts` was one of five sites without `allowDegradedEnvironment`, so a guest-PATH probe failure threw. Every consumer collapses a throw into a verdict: `isCommandAvailable` and `isCommandOnPath` catch to `false` ("not installed"), `isGhAuthenticated` and `isGlabAuthenticated` read an empty payload as "not authenticated". So a slow distro made WSL git, gh and glab read as missing. Two things made it likely rather than theoretical. The probe took two thirds of a 5s budget, leaving the command ~1667ms where main gave it the full 5s inside its own login shell -- a cold WSL VM start routinely lands in that band. And a probe timeout is cached for 30s with a re-probe threshold of 1.5x the failed budget, which a 5s caller can never clear, so every preflight command short-circuited without spawning wsl.exe at all -- and Re-check does not invalidate the cache. Fixes: preflight degrades instead of refusing, and the probe is capped at half the caller's budget and at 4s, so no caller ends up with less time than it had before the runner existed. Also fixes a real console flash found on the way: `preflight-command-exec.ts` spawns git/gh/node through `promisify(execFile)` with no `windowsHide`. Round 9 also confirmed the credential paths are now *safer* than main: all 11 account sites degrade, every destructive guest operation is still marker-gated, and main's `getOwnedManagedAuthPath` could disown an account on a 5s timeout -- which this branch turns into a failed launch instead of a destroyed selection. * fix(wsl): make "Try again" able to succeed, and test the round-9 fix Round 10 returned MERGE with one residual worth closing first. A transient probe failure left the null-resolving promise in `inFlight`, so the only way back was `retryAfter` -- and the 4s probe cap made the 1.5x budget escape unreachable, because no caller can pass more than 4s. For the full 30s window the four non-degrading sites returned their error *without spawning wsl.exe at all*, and each of those errors says "Try again". The advice was guaranteed to fail. The entry is now dropped on a transient outcome and an explicit cooldown gate replaces it, so the window alone decides. The window drops 30s -> 5s: long enough to stop a stampede, short enough that the user's next click reaches a distro that has since warmed up. Round 10 also noted the round-9 fix shipped untested, which was fair. Added: the probe-budget floor for 5s/8s/10s callers, and preflight's degrade opt-in plus its stdout/stderr-carrying rejection, which isGhAuthenticated reads off the caught error as an auth-success fallback. * test(wsl): make the probe-budget guard actually guard Round 11 caught that the regression test I added for the probe cap did not bind: it seeded the guest environment, so the probe resolved in ~0ms and the assertion read the command leg's timeout instead. Reverting the cap to the old 2/3 split left all three cases green. Dropping the seed and asserting on the probe leg fixes it -- verified by reverting the cap and watching all three fail. A regression guard that cannot fail is the shape that has cost the most in this workstream: the windowsHide guard silently passed a real unguarded spawn twice for the same reason. |
||
|
|
8462fa72da | fix(rate-limits): support Codex 0.149 approval policy (#15823) | ||
|
|
a61b39a9a6 |
fix(runtime): stamp a runtime's own project setups as local, and report remote status about the remote (STA-4792) (#15376)
* fix(runtime): stamp a runtime's own project setups as local, and report remote status about the remote (STA-4792) Two independent frame-of-reference bugs, both from code describing one machine while labelled as another. #15366 — projectHostSetup.* persisted the caller's host id verbatim. Those `runtime:<environment-id>` ids are minted by the calling client's own pairing store, so they name a machine only relative to that client. A client sending one is addressing this runtime, and runtimes do not proxy these calls onward, so the host it names is us. Storing the client's spelling made one machine look like a different host to every other client, hid its rows from them, and defeated the (projectId, hostId) duplicate check — two laptops paired to one server each created their own setup for the same checkout. Re-spell it as `local` at the RPC boundary. Rows written earlier keep their old stamp; readers already project `local` back to `runtime:<their-id>`, so the client-visible model is unchanged and no ids are rewritten. STA-4792 defect 4 — `status --environment <name>` hardcoded app.running:false to mean "no desktop on THIS machine" while every other field in the same object described the target, including a desktopWindowStatus echoed straight from it. The result contradicted itself and read as "that run was headless" when the remote GUI was up. `app` now describes the target, keyed off the one window status that requires a live renderer, and the result names its own subject so the frame can't be misread again. The remote pid is not knowable, so it stays null. STA-4792 defect 2 gets a regression test rather than a fix: routing already made the client remote, which is what stops a Windows destination being joined to the local cwd. The test pins the exact reported invocation. * fix(status): share the remote app projection with the SSH host passthrough, and name the version gap on project host setup Two review follow-ups. The SSH host passthrough answered `app.running: true` unconditionally for the Orca host a caller reached over SSH, claiming a desktop app even for a headless `serve`. That is the same defect as the paired-server path, one transport over, so the projection moved to shared and both now answer the question the same way. `--host runtime:<id>` routes project commands to a paired server, which means a client can reach a server that predates project host setup without meaning to. That answered a raw `method_not_found`, which reads as an Orca bug rather than a version gap; the CLI now names it the way the desktop already does. Reverted a third change: making the persistence duplicate check treat `local` and `runtime:*` as one machine. That assumption holds at the RPC boundary, where a `runtime:` host means the runtime being addressed, but not in the store, which also records independent provisioning metadata for machines that are not itself. An existing test covers exactly that, and it was right. The duplicate convergence therefore stays bounded to rows written after the normalization. |
||
|
|
ddaaa4628c |
fix(codex): keep two host-lane records that a failed read used to destroy (STA-4735) (#15289)
`snapshotCodexRuntimeHookTrustProvenance` rebuilt `.orca-hook-trust-provenance.json` from the current `config.toml` on every install and refresh, including when the existing record could not be read. That record is the only thing separating a trust entry Orca wrote from one the user approved inside Codex, so rewriting it after a failed read stamps the approval as Orca-written — and `promoteCodexRuntimeHookApprovalsToSystem`, which runs earlier in the same pass and had already bailed on the same unreadable file, then skips it on every later pass too. One denied read, permanent loss. Only the unreadable case is preserved. A malformed or absent record is still rebuilt, because resetting those IS the intent; conflating the two would wedge a user on a corrupt file forever. `fileContentsEqual` returned `false` from a bare `catch`, so "I could not read this" reached `writeRuntimeAuth` as "these differ" and sent it to the unconditional write below — replacing a refresh token Codex may have rotated a moment earlier with Orca's stale copy. It now reports the difference only when the bytes were actually compared, and the caller refuses instead. `fileContentsMatchExpected`'s `!existsSync` has the same collapse but is left alone deliberately: the write it guards is `writeFileAtomicallyIfUnchanged`, whose rename-and-compare re-checks the real file and refuses on its own, so classifying there would add a guard no test can drive. Noted in a comment. The three `writeRuntimeAuthAtPath` call sites have the same overwrite shape but are all on the WSL lane, which STA-4606 restructures; refusing there without its lane bookkeeping would set a baseline for a write that never happened. |
||
|
|
0b80a773a4 |
fix(codex): stop overwriting and deleting Codex files that were merely unreadable (STA-4737) (#15287)
* fix(codex): stop overwriting and deleting Codex files that were merely unreadable (STA-4737)
Three modules shared by the host and WSL Codex lanes decided a file was absent
from a read that had only failed, and then wrote over it or removed it.
- `codex-config-mirror`: `existsSync` on the RUNTIME config.toml returned false
for a locked file exactly as for an absent one, so the mirror took the
"seed a fresh runtime config" branch and replaced the user's config wholesale.
- `config-settings-promotion`: an unreadable ~/.codex/config.toml counted as
having no promoted settings, and the write path then rebuilt the user's
canonical Codex config from Orca's runtime copy.
- `codex-home-paths`: both delete branches in `linkSystemCodexResource` remove
Orca's mirrored copy because the system resource "is not there". `existsSync`
and `systemResourceIsRegularFile`'s `catch { return false }` both reported
that for a source nobody could read, so one denied read on ~/.codex/AGENTS.md
removed the managed copy on the next launch.
`src/shared/definitive-filesystem-absence.ts` now owns the one errno allowlist —
ENOENT and ENOTDIR, with every other code including unrecognised ones treated as
indeterminate — and `host-codex-managed-home-ownership.ts` drops its private
copy rather than letting the two drift. `codex-path-observation.ts` builds the
three-valued observation on top of it.
The resource sync's two `existsSync`/`statSync` probes collapse into one
resolved stat, which answers reachability and regular-file-ness together and
closes the window between them.
`config-settings-promotion.ts` crossed its max-lines budget, so the write-target
resolution moves to its own module rather than taking a lint exemption.
Deliberately not here: the hook-service trust writes that run after a refused
mirror, and the promotion write target's own classification, which is
unreachable because it always resolves to the same file the read above already
refused. Both are noted in comments rather than half-built.
* fix(codex): preserve resource copies on indeterminate reads
|
||
|
|
3a9f40ed70 |
fix(wsl): read machine output from a fenced login shell (#15290)
* fix(wsl): read machine output from a fenced login shell Orca runs WSL reads through the distro's *interactive* login shell so PATH matches the user's own terminal (nvm, mise and asdf only install into rc files interactive shells read). An interactive shell also runs the distro's rc/motd, and stock Ubuntu 24.04 writes its "run a command as administrator" hint to stdout -- no user customization required. Every caller parsing that stream was reading the banner as data: statPath -> "To run a command as administrator...\n\ndirectory" readPath -> banner prepended to the contents of every file read preflight -> banner prepended to `gh --version` / auth output `.trim()` cannot recover any of these, so a WSL worktree's file explorer sees no valid entry types and file reads return junk. Three call sites had independently grown their own marker to survive this (`__ORCA_AGENT_PATH__`, `ORCA_WSL_GIT_READ_ENV_V1`, and a `>/dev/null` fd dance), which is the tell that it belongs in one place. Fence the payload once, in the shared builder, and hand callers a reader that returns just their bytes. The fence carries a per-call nonce so `cat`-ing a file that happens to quote a marker is not truncated. Exit status is preserved, so the ENOENT mapping still works. wsl-git-read-environment drops its bespoke marker and parsing. * test(wsl): fence the login-shell path-lookup boundary test It asserted a raw interactive login-shell read matched an absolute path, so the distro rc banner made it fail on any stock Ubuntu. It is part of the shell-contracts CI gate, where it skips on Linux and hid the break. * docs(wsl): record the guest command-execution contract Both failure modes are silent - the command runs, exits 0, and returns the wrong bytes - so the rules need to live somewhere a reader will find them before writing the next wsl.exe call site. * fix(codex): fence the WSL Codex identity probe buildWslCodexBinaryStamp reads the login shell's stdout positionally -- path before the first newline, version after -- through an interactive login shell. On a stock Ubuntu the rc banner lands ahead of the payload, so the first newline falls inside the banner and the stamp becomes path="To run a command as administrator..." with the rest as version. Both halves are non-empty, so nothing throws: the stamp is silently wrong, and an unstable stamp reads as "the Codex binary changed" and reissues the trust grant. The identity script ends in `exec`, so it never writes a closing fence; the reader returns everything after the opening one, which is exactly this case. buildWslCodexIdentityArgs becomes buildWslCodexIdentityProbe and returns the reader with the argv so the two cannot drift apart. The other three WSL Codex commands are deliberately left unfenced: availability is exit-code only, and app-server/login hand stdout to a long-running program. * fix(wsl): harden the capture fence after review - readStdout now takes the LAST opening fence, matching the lastIndexOf the wsl-git-read-environment marker used deliberately: a login shell can echo the command text before running it, repeating the fence. - local-worktree-filesystem throws instead of falling back to raw stdout when the fence is missing. The fallback silently reinstated the bug being fixed -- statPath would return the banner as a file type and readPath would return banner+contents, with no signal. Preflight keeps its fallback; its matchers scan the whole blob and tolerate a prefix. - The exit-status test asserted only that the script CONTAINS `exit $?`, which is true for any input and never executed those lines. It now runs a real distro and asserts status 2 reaches the caller, which is what statPath's ENOENT mapping depends on. - Corrected the doc: a sed backreference has no `$`, so `--` never rewrote it. Replaced with the positional and shell-local cases that were measured to differ. * fix(wsl): stop running a login shell for filesystem reads statPath/readPath/rm run coreutils at standard paths and shell builtins. They need nothing from the user's PATH, so there was never a reason to start a login shell -- and starting one is what put the distro's rc/motd on the stdout these callers parse. Fencing that output treated the symptom. Using a plain `sh -c` removes the cause: no profile, no rc, no banner, by construction. The fence and its missing-fence error go away with it. The fence stays where it is actually needed: the three places that must run the user's shell to resolve their PATH (the preflight CLI probe, the WSL git environment probe, and the Codex identity probe). Net -12 lines. --------- Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
6e8da1df8d |
fix(wsl): pass guest argv verbatim through --exec (#15039)
* fix(wsl): pass guest argv verbatim through --exec
`wsl.exe <...> -- <argv>` expands `$name` in every argument against the
guest environment before the guest ever runs. It does this even when no
shell is involved, so `-- /usr/bin/printf %s '$HOME'` prints /home/you.
Every WSL invocation went through that preprocessor, so scripts arrived
already rewritten: `awk '{print $2}'` lost its field reference, and a
POSIX script asking for the literal `$HOME` got the expanded path.
`escapeWslShCommandForWindows` tried to compensate by escaping `$`, but
it skipped any `$` preceded by a backslash, so a script containing `\$`
was still corrupted -- and half the call sites never applied it at all.
Route every invocation through `--exec`, which passes argv through
untouched, and delete the escaper. The direct-git path already used
`--exec`, so this is not a new compatibility dependency.
A guard test fails if the `--` form reappears anywhere in the tree.
Net -50 lines of production code.
* test(wsl): drop remaining escaped-dollar assertions
* fix(wsl): cover the --exec migration's blind spots
An audit of every wsl.exe invocation found sites the first pass missed,
including two it actively broke:
- config/scripts/wsl-git-shell-benchmark.mjs imported
escapeWslShCommandForWindows, which no longer exists, so the script
threw on startup. Its wslShellArgs helper also still used `--`; the
file already had an --exec helper, so route both call sites there.
- classifySubprocessCommand unwrapped `wsl.exe <...> -- <binary>` by
breaking on `--` alone. With every Orca spawn now on --exec it never
found the guest binary and bucketed all WSL subprocesses as plain
"wsl", losing the git/gh/glab breakdown. Break on either separator,
since foreign wsl.exe processes still use `--`.
CliSkillRuntimeSetup builds its setup command as a template literal
rather than an argv array, so no array-shaped search could see it. Its
decoder accepts both separators so commands persisted before this
change still decode.
The guard now scans config/ and tests/ as well as src/, and checks the
command-string spelling alongside the argv one — the two shapes that
have each shipped a regression. It skips comment lines so prose about
the old form stays allowed, and asserts it scanned a plausible file
count so a bad root cannot make it vacuous.
* fix(wsl): restore the guard's multi-line sensitivity
The guard matched line by line, so `'--',\s*'bash'` could not span a
newline -- and every argv array in this repo is formatted one element
per line, which is exactly the shape it exists to catch. Measured
against the pre-migration tree it caught 17 files before and 9 fewer
after. It now strips comment lines and matches the rejoined text, with
a case that pins the multi-line shape so this cannot silently return.
The program list is wider than shells now, which surfaced a false
positive: tmux takes a `--` separator followed by a program too
(`split-window ... -- cat`). Matching is scoped to files that mention
WSL rather than narrowing the list back.
Also:
- Replaced the `sed` regression case, which was vacuous. A backreference
contains no `$`, so it returned `bac` under both separators and would
have passed without the fix. The block claimed every case proved the
bug. Swapped in a positional argument and a shell local, both measured
to differ -- the positional is the shape `wslUncDirectoryExists` uses,
where `--` blanked `$1` so every existing directory probed as missing.
- windows-shell-args.test.ts derived its expected argv from
buildWslExecArgs, the helper under test, so six assertions would still
pass if it regressed to `--`. Spelled the expectation out.
- Dropped two comments citing the removed `--` behavior as rationale.
---------
Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
|
||
|
|
5652fb7469 |
fix(codex): stop resuming a session under the wrong account when a sessions tree is locked (#15093)
* fix(codex): stop resuming a session under the wrong account when a sessions tree is locked Two probes reported "this rollout is not bridged here" for any filesystem error, not just a genuine absence: - codex-session-resume-home.ts used existsSync on each ranked home's sessions directory. existsSync returns false on EBUSY/EPERM, so a briefly locked tree made the scan continue to the next ranked home — and the winning home becomes the resumed pane's CODEX_HOME, so it picks the account. - codex-legacy-session-resume.ts caught every lstat failure for the selected account's candidate rollout and returned null, after which the caller kept the source per-account home. Either way the session resumed under a different account's credentials while the UI still showed the selected one. Only a definitive ENOENT/ENOTDIR now means "not bridged here". Any other error raises the typed temporary-unavailability refusal the ownership gate already uses, which both PTY paths convert into a clean abort before spawn. The refusal is scoped to the SELECTED account's home. An unreadable home that is not the selected account cannot cause a wrong-account resume, so it is still skipped rather than stranding the user. These are pre-existing and independent of the STA-4422 ownership-marker failure: they route to another account today with the gate uninvolved. Fixes STA-4607 * fix(codex): refuse a resume when the selected sessions tree is locked mid-listing Review found the first pass incomplete in two places, both the same category error one layer further down. The preliminary statSync on the selected sessions root was guarded, but the real directory read happens later and listCodexSessionRolloutFilesIncrementally swallows every opendir error. A lock held during enumeration — where nearly all the I/O is, and so the far more likely case — still yielded nothing for the selected account and fell through to another one. The listing now reports directory errors through its existing onDirectoryError hook, and a non-definitive error anywhere under the selected sessions root raises the typed refusal. Non-selected homes and definitive absence still skip. Separately, index.ts wrapped prepareLegacySharedCodexSessionResume in a blanket catch and fell back to the source home, so the typed refusal from the candidate lstat was swallowed and the resume still ran under the peer account's credentials. That catch now rethrows ManagedCodexHomeTemporarilyUnavailableError while ordinary migration failures keep warning and falling back, since a genuine migration failure legitimately should not block a resume. A typed refusal is only as strong as the narrowest catch between the throw and the spawn. The frames between both throw sites and the PTY spawn were audited: findTrustedCodexSessionResume, resolveCodexSessionResumeProvenance and prepareCodexSessionResume have no catches, and the PTY layer already maps the typed error before spawn. Both fixes are mutation-checked. Disabling the listing hook makes the resume resolve to the other account again; the index.ts rethrow is covered only by typecheck, because src/main/index.ts has no unit-test entry point in this repo. * test(codex): pin nested-directory lock coverage; document the resume repin contract Review flagged the listing guard as matching only the exact sessions root, so a nested dated directory would leak. It does not — the guard keys on the root being listed, not the failing directory — but nothing pinned that. Added a test that faults only sessions/2026/07/20 while the root stats fine; it fails under mutation alongside the root case. Also documented why the index.ts rethrow cannot fire today. That launch path pins CODEX_HOME to the account that owns the rollout and deliberately refuses to repin onto whichever account is selected now (#10793), so it does not wire the selected-home resolver. The branch stays as a contract guard so the blanket catch below can never silently swallow a typed refusal if that changes. |
||
|
|
7ae6aedc02 |
fix(codex): stop a transient filesystem error from logging out the active account (#15046)
* fix(codex): stop a transient filesystem error from logging out the active account A single unreadable read of a managed Codex home's ownership marker cleared the user's active account selection, permanently. On Windows any exclusive lock — Defender real-time scanning, a backup agent, a sync client — makes every read of that marker fail with EBUSY, and the background rate-limit poll runs every 15 minutes plus once at every app start. Root cause: the ownership gate answered two very different questions through one channel. "This home is not ours" (a successful observation that failed a trust check) and "we could not read it" both surfaced as a throw, which the caller flattened to null, which three call sites took as proof the home was untrustworthy and wrote activeCodexManagedAccountId: null. Refusing to USE an unverified home is correct. Erasing the user's account selection because a file was briefly locked is not. The gate now returns a tri-state verdict. `untrusted` comes only from a proven trust failure or a definitive ENOENT/ENOTDIR where absence is itself the verdict; every other filesystem exception is `indeterminate`. Only `untrusted` may touch persisted state. Because `null` already meant "fall through to the system default" on both the launch and poll paths, not-clearing on its own would have run a DIFFERENT account behind a UI still showing the selected one. So the refusal needed real channels rather than a sentinel: - the poll returns an explicit skip; returning null would not have skipped at all, since the fetcher maps null to ~/.codex and would have spawned a token-refreshing app-server inside the user's real credential home - pane launch throws a typed temporary-unavailability error that both PTY implementations convert into a clean refusal with a retry message, including the re-resolution after the async auth-readiness wait - automatic session resume resolves the selected home eagerly, so an unreadable account can no longer be silently replaced by another one in the ranking - config-sync status reports a distinct managed-home-unavailable stall instead of "synced", with a bounded renderer retry so it clears on its own Also fixes the ticket's second symptom. The status bar's Sign in button called a re-auth that captured the selection before login and restored it after, so re-authenticating a deselected account restored `null` — a successful login that left the account inactive, with no success toast to distinguish it from failure. It now activates the account it just signed in, but only when the pre-login selection was empty, so it cannot silently switch accounts for multi-account users, and it runs the same restart prompt an explicit switch does. No retry or grace window inside the synchronous gate: it runs on the Electron main process in a loop over accounts, so a sleep there would freeze the UI. Recovery is simply the next readable evaluation. The WSL lane has the same class of defect, including one path that deletes a credential mirror. It is pre-existing, unreachable from these host code paths, and deliberately left for its own change; the host clearing sites cannot reach a WSL account because getSelfContainedManagedHostAccount excludes them. Fixes STA-4422 * test(codex): cover pending reset home ownership |
||
|
|
02ba70a847 |
fix(agent-hooks): make the Windows managed hook survive Claude-hooks-compat consumers (#14825)
* fix(agent-hooks): make the Windows managed hook survive Claude-hooks-compat consumers `~/.claude/settings.json` is not read only by Claude Code. Third-party Claude-hooks-compat layers (cursor-agent, Devin) import the same file and reimplement hook execution, so Orca's entry has to survive consumers that support strictly less than the documented schema. Three separate defects came from assuming otherwise. 1. The entry depended on `args`, which a compat consumer ignores. `args` is valid Claude Code syntax, but cursor-agent spawns `command` alone -- so `conhost.exe` ran bare, which opens an interactive console that never closes. Hook payloads were typed into those stranded shells (#14815). The entry is now one self-contained `command` string that depends on nothing optional. 2. `conhost.exe --headless` never relayed anything. It implements the ConPTY server protocol, not a generic no-window wrapper: it does not wait for the hosted process and relays neither exit code nor stdout. Measured directly -- `conhost --headless cmd /c "echo X& exit /b 42"` yields empty stdout and no exit code, while the replacement returns both and waits. So every hook was fire-and-forget, and whatever it printed was discarded. Replaced with `-WindowStyle Hidden`, which suppresses the window and keeps wait/exit-code/stdout intact. 3. The hook never wrote anything to stdout. Guards exited silently and curl's output went to nul. Claude Code documents empty stdout as "no decision", but cursor-agent treats PreToolUse as a permission gate, fails to parse empty stdout as JSON, and blocks the tool call -- so every shell command in every cursor-agent session on Windows failed (#14818). The script now writes `{}` first, on both the Windows and POSIX branches, which is documented to be identical to writing nothing for real Claude Code. Gemini and Antigravity already did this. Defects 2 and 3 are causally linked: `{}` cannot reach any consumer while conhost is swallowing stdout, so neither fix works without the other. Also fixed while establishing the contract: - The launcher's own missing-script fallback returned empty stdout, reproducing #14818 whenever `~/.orca` was cleaned or an install was half-finished. It now emits `{}` too. - PowerShell serializes progress records to stderr as CLIXML when stderr is redirected; a consumer merging stderr into stdout would see those bytes before the JSON. Every encoded payload now silences progress. - `runtime-home-hook-command.ts` built its own launcher without window suppression -- exactly the drift #14815 asks to prevent. All launcher construction now goes through `windows-powershell-hook-launcher.ts`, so the switch list cannot be present in one installer and missing in another. - Renamed `usesWindowsHeadlessHook` to `usesWindowsPowerShellLauncher`; nothing is headless anymore, and the flag selects a launcher. Testing: the new regression test asserts the effect a consumer observes -- it runs the exact `command` string from settings.json through both cmd.exe and Git Bash, across the guard-exit, reached-curl, and missing-script paths, and parses stdout. Verified it fails when `conhost --headless` is reintroduced. The previous tests all asserted installer intent, which is why they passed through all three defects. * fix(agent-hooks): close hook launcher review gaps --------- Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com> |
||
|
|
9367169888 |
refactor(tests): split every oversized test file off the max-lines suppression list (#14728)
* refactor(tests): split oversized test files off the max-lines suppression list Every `*.test.ts`/`*.spec.ts` that carried an `eslint/oxlint-disable max-lines` directive is now split into focused, behavior-scoped suites that fit the 800-line test budget, with shared setup extracted into co-located `*-test-harness.ts` / `*-test-fixtures.ts` modules (300-line budget). 83 files became ~930; the largest output is 797 effective lines. `orca-runtime.test.ts` is intentionally untouched. Test bodies were moved by scripted line-range slicing rather than retyped, so assertions are byte-identical. The only permitted body edits were mechanical rebinding where a shared value moved into a harness (e.g. `tmpHome` -> `homes.tmpHome`). Registries that enumerate test files were updated in lockstep: - config/max-lines-baseline.txt: pruned 341 -> 258 entries (all 83 removed). - config/reliability-gates.jsonc: 33 gates repointed at the split files, with assertionRefs split per file where a gate's coverage now spans several. - .github/workflows/pr.yml: the real-zsh lane now lists the 4 split files that actually exercise zsh, so they keep running in the dedicated shell lane. Also renamed agent-hooks `server-test-fixtures.ts` to `server.test-fixtures.ts` so the global-fetch call-site audit keeps skipping it, and added `.js` extensions to the CLI suites' dynamic harness imports (node16 resolution) to unbreak `build:cli`. Verification: full suite 52,449 passing vs 52,448 at baseline with zero assertions lost; `pnpm lint`, `pnpm typecheck`, and `pnpm build:cli` all exit 0; the terminal-pane e2e spec runs 31/31 headless. * refactor(tests): split hook-idle arbitration suite that oxfmt pushed over budget The pre-commit oxfmt pass reflowed pty-connection-hook-idle-arbitration.test.ts to 811 effective lines, 11 over the test budget. Split the hook-completion side effect and replacement-agent veto cases into their own suite; both files now sit well under the cap and the 15 tests are unchanged. * test: port upstream test changes into the split files after rebase Rebasing onto main surfaced 27 tests that main had added to files this branch deleted, plus edits to tests that had already moved. Taking the deletion side of those modify/delete conflicts would have dropped that coverage silently, so each upstream change is ported into the split file that now owns the behavior — for example main's six orchestration mailbox tests land across orchestration-runs, -send, and -check. Also repoints `orchestration.notification-mailbox-consistency`, a gate main added after this branch's gate remap, at those same three split files, and re-prunes the max-lines baseline against main's (257 entries). Verified: all 27 upstream test titles present; full suite 52,761 passing with the only diff vs baseline being 12 tests main itself removed and 3 that moved from skipped to passing; lint and typecheck exit 0. * fix(test): flush pending continuations before tearing down terminal test globals CI shard 5/16 failed on both Node 24 and 26 with `ReferenceError: window is not defined` from pty-connection.ts, surfacing through pty-connection-daemon-snapshot-replay.test.ts. The reattach/settle chains `await` a real promise and then touch `window.api`. Under fake timers those continuations cannot run, so they only become schedulable once restoreTerminalTestGlobals() switches back to real timers — which previously happened immediately before `delete globalThis.window`, so a late continuation threw and failed the whole file. Flush async ticks in that window instead. This is latent in the source rather than new: the pre-split 25k-line file kept running other tests after these, which gave the chains time to settle before teardown. Splitting the file moved teardown directly behind them. * fix(test): keep an inert window after terminal test teardown instead of deleting it The async-tick flush was not enough: the reattach/settle chain can resolve after teardown regardless of how long we drain, so CI shard 5/16 still failed with `ReferenceError: window is not defined` from pty-connection.ts. A real renderer never loses `window`, so deleting it was the artificial part. Swap in an inert proxy whose properties resolve to callables and whose calls resolve to undefined, making a late `window.api.pty.*` call a harmless no-op. The next test replaces it wholesale via installTerminalTestGlobals(), and no test asserts that `window` is absent. |
||
|
|
77f23b013f |
refactor(shared): drop the shared/types barrel and import from the real modules (#14447)
#14397 split `shared/types.ts` into 46 per-domain modules but kept the path as a re-export barrel so the import sites did not have to change. This removes the barrel: every consumer now imports from the module that actually declares the type, and `src/shared/types.ts` is deleted. Barrels hide where a type lives, make every consumer look like it depends on the whole domain, and let an unrelated edit invalidate a module that ~2,000 files transitively import. 2,323 import declarations across 2,321 files. Rewritten mechanically: each specifier was resolved to an absolute path via the TypeScript AST and recomputed, rather than string-substituted, so alias forms (`@/../../shared/ types`) and per-specifier `type` modifiers survive. Four cases the mechanical pass had to handle, each found by a gate rather than by reading the diff: - Modules inside `src/shared` import the barrel as `./types`, not `shared/types`. A pre-filter on the latter string skipped 176 of them and left imports dangling at a deleted file, which surfaced as confusing `Property 'x' is optional in type 'Repo' but required in Pick<Repo, ...>` errors rather than "module not found". - The barrel RENAMED one type on the way through (`WorkspaceSource as WorkspaceCreateTelemetrySource`), so the original name in the owning module has to be re-aliased at each consumer. - Three test files put `;(globalThis as ...)` on the line after the import. TypeScript parses that `;` as the import statement's terminator, so replacing through `statement.getEnd()` deletes it and breaks ASI. The rewrite now stops at the module specifier. - A file that already imported directly from a module got a SECOND import from it, because the barrel re-exported those same names — which trips `import/no-duplicates` under `--deny-warnings`. A post-pass merges declarations sharing a specifier and type-only-ness; the `import type` plus `import` pair from one module is left alone, since that form is allowed. Splitting one barrel import into several genuinely adds lines, which pushed `terminal-layout-pty-ownership.ts` to 301 counted lines: its 107-character import must wrap, and neither local type collapses onto one line (101 and 116 characters). Rather than contort a type declaration to fit a line budget, `collectLeafIds` and `pruneLeaves` move to `terminal-pane-layout-tree.ts` — they are pure structural operations on the layout tree and independent of PTY ownership. `visible-worktrees.ts` similarly loses its own mini-barrel re-export of `isDefaultBranchWorkspace`, with the four real consumers repointed at the declaring module. No `max-lines` bypass added. Verified: cold `tsc --noEmit` green on node, cli, and web (buildinfo deleted first — these projects are `composite: true` and reuse stale caches); the full `pnpm lint` green, not just bare oxlint — the narrower local check is what let the duplicate imports reach CI; max-lines ratchet OK at 344. |
||
|
|
537864a248 |
Fix Codex hook trust before manual shell launches (#14326)
* fix codex hook trust before shell launch * fix packaged cli preflight dependency * fix codex shell preflight safety * fix Codex shell preflight settings and startup safety |
||
|
|
5ea7df1a5b |
fix(terminal): make DECSET 2031 subscriptions silent (#13904)
fish arms `CSI ?2031h` before painting each prompt and withdraws it when it hands the tty to a child — a ~1ms window. Orca answered that subscribe with `CSI ?997;Nn` across a 1-3ms renderer hop, so the reply landed after the withdrawal and was read as stdin by the next child, corrupting `brew`/`npx` `[y/N]` prompts. The reply is not stale by Orca's own view when written (measured staleReplies: 0), so no suppress-the-stale-reply scheme can close this — the information needed to suppress does not exist yet. Nothing asked for the reply either. The Contour spec says a terminal "should only send out the DSR when the palette has been updated"; Ghostty (Termio.zig:729 — force=true reachable only from the ?996n DSR), iTerm2 (VT100Terminal.m:995 — flag only) and xterm.js (InputHandler.ts:2035 — flag only) all emit nothing on the DECSET. So stop entering the race: record the subscription, answer nothing. Of 17 real programs measured under a pty, only fish, tmux, claude and opencode subscribe; none block on a reply, and answering produces one redundant palette re-query and zero rendering difference. tmux is the only one that sends `?996n`, which Orca still answers. - Subscribes are record-only at all four emitters (live scan, hidden-gate fact, parked byte watcher, parked responder — the last is deleted, it only replied). - `?996n` answers, the subscription registry, and the theme-flip push are unchanged. `paneLastThemeMode` is still seeded at subscribe so the next appearance re-apply is not read as a flip. - Replay grammar carries `?2031l` alongside `?2031h`, so a late-attaching remote client no longer registers a subscription the TUI already retired. Also closes fish-integration gaps found alongside: `unset` (which fish lacks) becomes `set -e` on paths parsed by the client's login shell, `config.fish` is parsed for agent-home detection, and bracketed-paste startup delivery is made consistent across local/daemon/relay. Regression test drives real fish 4.7.1 under node-pty and asserts on what the child process reads; it fails against pre-fix code with the exact payload from the issue. CI installs fish 4 and fails loudly rather than skipping. Closes #9993 Co-authored-by: Orca <help@stably.ai> |
||
|
|
991a3fe963 |
chore(lint): update oxlint to 1.77 and enable no-op cleanup rules (#13901)
Enable eleven oxlint rules that simplify code without changing behavior, and fix
every existing violation. Each candidate was gated on measured cost rather than
assumption, so rules that regressed runtime performance or type checking were
dropped instead of suppressed.
typescript/no-redundant-type-constituents is the largest addition: 113 sites, no
autofix. Dead constituents are deleted. Where the redundant literal existed to
document intent (`string | 'all'`), it is preserved as `(string & {})`, which
keeps the autocomplete hint the original code was reaching for instead of
flattening it away. The rule also caught a broken import —
remote-shared-control-retirement-probe.ts pulled RuntimeStatus from
src/shared/types, which does not export it, so the type silently degraded to
`any`; no tsconfig covers that file, so tsc never saw it.
oxlint stays at 1.77.0 rather than 1.78.0 because .npmrc sets
minimum-release-age=4320 and 1.78.0 is younger than that window.
Rules evaluated and rejected, with what disqualified each:
- prefer-string-raw: String.raw is a runtime call, not a literal (184x slower)
- prefer-string-replace-all: 26% slower
- text-encoding-identifier-case: ~5% slower, reproducible
- prefer-spread: [...str] is 110% slower than split('') and differs on surrogates
- no-implicit-coercion: `!!x` narrows types and `Boolean(x)` does not (22 tsc errors)
- prefer-arrow-callback: arrows are not constructible, breaking `new` on mocks
- object-shorthand: rewrites source text asserted by a tracked reliability gate
- switch-case-braces: pushes ten files past max-lines, which cannot be suppressed
- no-useless-switch-case: drops `case undefined:` that switch-exhaustiveness-check needs
- arrow-body-style: 115 violations have no fix, and it breaks max-lines
- newline-after-import: false-positives on the leading-semicolon ASI idiom
electron-vite-output-contract asserted on the literal
Object.prototype.hasOwnProperty.call text; retarget it to Object.hasOwn, which
rejects inherited keys identically.
|
||
|
|
a639bd355e |
[perf-remote] perf(codex): skip migration scans for nonshared reattaches (#13654)
* perf(codex): skip migration scans for nonshared reattaches * fix(codex): require provenance to skip migration scans and bound ignored launches Drops the exit-match escape hatch that skipped a scan without route provenance, and ages out ignored reattach launches on the existing 60s/256 lease policy so a lost exit can no longer pin an active launch (and its completion marker) for the process lifetime. Co-authored-by: Orca <help@stably.ai> --------- Co-authored-by: Orca <help@stably.ai> |
||
|
|
7250fde148 |
fix(codex): repair hook trust before restored shells reconnect (#13745)
* fix(codex): reconcile hooks for restored shells * docs: move restored shell validation report * perf(codex): gate retained pane inventory |
||
|
|
2ee43bfc0d |
fix(agent-hooks): refresh existing Orca launchers when agent CLIs are unavailable (#13378)
* fix(agent-hooks): refresh existing shared hook scripts when the CLI is no longer detected A CLI that falls off PATH (moved npm prefix, relocated shim) keeps its user-wide config invoking Orca's launcher script under ~/.orca/agent-hooks, but the presence gate skips install() with no removal — freezing the script at whatever Orca generated last. Anyone in that state kept the pre-#11568 more.com-leaking .cmd forever, because no launcher script is ever deleted and Windows startup deliberately skips shell PATH hydration. Reconcile before gating: every existing shared launcher/statusline script is rewritten to the current template on each install pass. Creating scripts stays behind the presence gate — an existing file is proof of a prior install; a missing one means the gate did its job. Amp and Hermes are deliberately absent: they write provider-native plugin code with its own install lifecycle, not shared launchers. - refreshManagedScriptIfPresent() in installer-utils (no-op unless the file exists) - refreshManagedScripts() on the 11 launcher-writing services (openclaude via the shared Claude class) - reconcile pass in installManagedAgentHooks before presence detection, filtered by the agents option, best-effort per agent - coverage gate: a launcher written to ~/.orca/agent-hooks without a matching refresher entry fails the suite, in both directions * perf(agent-hooks): refresh launchers off the main thread * test(agent-hooks): keep refresh mode assertion POSIX-only |
||
|
|
4b5157b147 |
fix(codex): bound state DB recovery retries (#13109)
Co-authored-by: Jinwoo-H <Jinwoo-H@users.noreply.github.com> |
||
|
|
38275c2aa2 |
fix(codex): publish Windows system-default sessions (#12611)
* fix(codex): publish Windows system-default sessions * fix(codex): close launch-scheduling races in session migration scheduler * fix(codex): bound repeated session migration audits * fix(codex): preserve delayed session publication passes * fix(codex): bound failed session audit events * fix(codex): fence stale session migration markers * chore: preserve main formatting after merge * perf(codex): bound launch session migration scans * perf(codex): preserve coalesced migration scope * fix(codex): preserve scheduled migration recovery * fix(codex): preserve session migration recovery * fix(codex): close session migration launch races * fix(codex): harden session migration completion |
||
|
|
f0443c326a |
fix(codex): recover interrupted state DB backfills (#12617)
* fix(codex): recover interrupted state DB backfills * fix(codex): detect mixed-case backfill timeout * fix(codex): harden backfill recovery review findings * fix(codex): keep process identity retries safe |
||
|
|
4c49989c2e |
refactor(codex): delete the unreachable managed shared-mirror lane (#12614)
PR 9501 shipped real-home routing for the host system default, and the env override that could turn it back off was never a shipped control. The managed-account half of the shared runtime mirror has been unreachable since: every host account routes to its own self-contained CODEX_HOME before that code runs. Delete the flag module and its env plumbing plus the managed branch of syncForCurrentSelection and the six helpers only it called. The three lanes that still use the shared mirror -- Windows, a custom CODEX_HOME, and a hook-lane gate that reports unusable -- are untouched, as are every legacy migration and the WSL read-back helpers. |