Files
orca/src/main/codex/codex-app-server-capability-cache.test.ts
T
Neil 26721bd632 fix(codex): stop blocking the main thread on trust grants (#16441) (#16594)
* fix(codex): stop blocking the main thread on trust grants (#16441)

Codex hook trust was granted by blocking the Electron main thread on
`spawnSync` of a bundled ELECTRON_RUN_AS_NODE entry for the whole
app-server deadline: 15s native, 35s WSL, ~45s on the real-home path
(rebase inspect + repair + grant). Cold start and every Codex pane
launch showed "Not Responding"; the reported event-loop gap was
15,049 ms.

The subprocess only ever existed to donate an event loop to a
deliberately blocked parent — `runCodexHookTrustGrantSession` was
already the real async implementation. Make the callers async and the
fork is unnecessary, so the bridge, the forked entry and its envelope
are deleted along with their build/knip/tsconfig registrations. The CLI
`agent hooks prepare-codex` handler is already async, so it awaits the
in-process session and saves a process spawn per managed-home shell.

`resolveCodexTrustGrantHost` is async too; the WSL identity probe moves
from `execFileSync` to `runProcess`, dropping that file from the
child-process import allowlist. Status reads keep a synchronous
native-only stamp path.

Two invariants that held only because the lane blocked:

- Overlapping capability probes were impossible by construction.
  `GitCapabilityCache`'s dedupe engine is extracted to a shared
  `CapabilityProbeCache` and `CodexAppServerCapabilityCache` now
  inherits it, so concurrent launches against a cold host share one
  app-server session instead of one each.
- Two grants on one `config.toml` could not interleave capture and
  restore. A reentrant per-file lane now serializes the whole install
  sequence (managed, WSL runtime, real-home ensure, legacy sweep) and
  the grant and rebase inside it.

Cold-start work moves off the critical path: retained-home
reconciliation (N sequential sessions) is fire-and-forget behind the
daemon provider, and the startup real-home ensure chains into managed
hook reconciliation instead of blocking app init.

Every preserved semantic is unchanged: never throws, the
ORCA_DISABLE_CODEX_TRUST_RPC kill switch, ledger hits, backfill-pending
and cooldown fallbacks, config rollback on every failure path,
pre-grant self-computed trust removal, the verify-failure taxonomy,
diagnostics and telemetry.

* fix(codex): widen the trust-config lane to every config.toml writer

Review follow-ups on #16441's async trust grant:

- `markCodexProjectTrusted` now runs inside the runtime+system config.toml
  lanes, so a project-trust write can no longer land inside a hook grant's
  capture->restore window and be silently reverted. Its callers await it.
- `install`/`refreshRuntimeUserHooks`/`remove` hold the system config.toml
  lane as well as the runtime one — they promote approvals into
  ~/.codex/config.toml and mirror it back. Lock order is runtime-before-system
  everywhere.
- The real-home ensure chain resumes after a rejection instead of returning
  the same rejected promise to every later pane launch, and resolving the real
  home is now inside the module's never-throws boundary.
- `buildSpawnEnv` awaits inside a cancelable pending-spawn registration, so
  shutdown during the (now long) env build stops the PTY from launching.
  `prepareLocalPtySpawn` generalizes into `awaitCancelableLocalPtySpawn`.
- CapabilityProbeCache drops the test-only `nowMs` passthrough; its probe
  backstop comment now describes what it actually guards.
- Preflight is a plain async function; the trust dispatch in orca-runtime
  collapses into one `markWorkspaceTrustedForAgent`.

* test(codex): exercise the trust-config lane under real concurrency

The async grant makes two pane launches overlap for the first time. These
drive the real modules end to end on real files: a rollback swallowing a
sibling's grant, a markCodexProjectTrusted write landing inside a capture
-> restore window, shared capability-probe dedupe on a cold host, the
host-scoped transient cooldown, and reentrancy from inside an installer.

Each was verified to fail against a deliberately broken implementation
(lane removed, dedupe disabled, cooldown made global, reentrancy pass-
through disabled).

* test(codex): stop hook-service suites spawning the developer's real codex

The forked grant bundle never existed under vitest, so the RPC lane was
unreachable in tests on main. Running it in-process makes these suites
spawn a real `codex app-server` when one is installed: 38 spawns and two
failures in hook-service-runtime-trust-repair on a machine with codex,
green in CI where there is none. Stand in for the missing binary so both
environments exercise the same fallback lane.

* docs(codex): scope the trust-RPC kill switch comment to what it actually gates

The comment read as though the flag forces the fallback lane everywhere. It
gates the managed grant only: the real-home rebase still runs its own
inspect/repair app-server sessions when Orca's insertion shifts a user's hook
positions, and never reads the flag.

Verified by exercise, not by reading — with the flag set, both
inspect-user-hook-trust and repair-user-hook-trust still ran. Pre-existing:
main has no check there either, it just blocked the main thread while doing it.

Widening the flag to cover the rebase is a follow-up; this only stops the
comment promising something the constant does not do.
2026-08-26 16:44:55 -07:00

213 lines
6.8 KiB
TypeScript

import { describe, expect, it, vi } from 'vitest'
import {
CODEX_APP_SERVER_CAPABILITY_RETRY_INTERVAL_MS,
CodexAppServerCapabilityCache,
getCodexAppServerHostKey
} from './codex-app-server-capability-cache'
const unsupportedError = new Error('unsupported')
const isUnsupported = (error: unknown): boolean => error === unsupportedError
describe('CodexAppServerCapabilityCache', () => {
it('retries a host after the compatibility interval', () => {
const cache = new CodexAppServerCapabilityCache()
cache.rememberUnsupported('native', 1_000)
expect(
cache.shouldTry('native', 1_000 + CODEX_APP_SERVER_CAPABILITY_RETRY_INTERVAL_MS - 1)
).toBe(false)
expect(cache.shouldTry('native', 1_000 + CODEX_APP_SERVER_CAPABILITY_RETRY_INTERVAL_MS)).toBe(
true
)
})
it('falls back on the first unsupported probe and skips the probe on later calls', async () => {
const cache = new CodexAppServerCapabilityCache()
const firstPreferred = vi.fn(() => Promise.reject(unsupportedError))
await expect(
cache.runWithFallback(
'native',
firstPreferred,
() => Promise.resolve('first-fallback'),
isUnsupported
)
).resolves.toBe('first-fallback')
expect(firstPreferred).toHaveBeenCalledTimes(1)
const laterPreferred = vi.fn(() => Promise.resolve('unexpected-preferred'))
await expect(
cache.runWithFallback(
'native',
laterPreferred,
() => Promise.resolve('cached-fallback'),
isUnsupported
)
).resolves.toBe('cached-fallback')
await expect(
cache.runWithFallback(
'native',
laterPreferred,
() => Promise.resolve('cached-fallback'),
isUnsupported
)
).resolves.toBe('cached-fallback')
expect(laterPreferred).not.toHaveBeenCalled()
})
it('isolates capability state per execution host', async () => {
const cache = new CodexAppServerCapabilityCache()
cache.rememberUnsupported('wsl:Ubuntu', 1_000)
expect(cache.shouldTry('wsl:Ubuntu', 1_001)).toBe(false)
expect(cache.shouldTry('native', 1_001)).toBe(true)
expect(cache.shouldTry('wsl:Debian', 1_001)).toBe(true)
const nativePreferred = vi.fn(() => Promise.resolve('native-result'))
await expect(
cache.runWithFallback(
'native',
nativePreferred,
() => Promise.resolve('unexpected'),
isUnsupported
)
).resolves.toBe('native-result')
expect(nativePreferred).toHaveBeenCalledTimes(1)
})
it('drops known support when a later call reports the capability unsupported', async () => {
const cache = new CodexAppServerCapabilityCache()
await expect(
cache.runWithFallback(
'native',
() => Promise.resolve('supported'),
() => Promise.resolve('unexpected'),
isUnsupported
)
).resolves.toBe('supported')
expect(cache.isKnownSupported('native')).toBe(true)
await expect(
cache.runWithFallback(
'native',
() => Promise.reject(unsupportedError),
() => Promise.resolve('fallback'),
isUnsupported
)
).resolves.toBe('fallback')
expect(cache.isKnownSupported('native')).toBe(false)
const laterPreferred = vi.fn(() => Promise.resolve('unexpected-preferred'))
await expect(
cache.runWithFallback(
'native',
laterPreferred,
() => Promise.resolve('cached-fallback'),
isUnsupported
)
).resolves.toBe('cached-fallback')
expect(laterPreferred).not.toHaveBeenCalled()
})
it('rethrows transient errors without marking the host unsupported', async () => {
const cache = new CodexAppServerCapabilityCache()
const transient = new Error('spawn ETIMEDOUT')
await expect(
cache.runWithFallback(
'native',
() => Promise.reject(transient),
() => Promise.resolve('unexpected-fallback'),
isUnsupported
)
).rejects.toBe(transient)
expect(cache.shouldTry('native', 2)).toBe(true)
})
// Why (#16441): grants no longer block the main thread, so two pane launches
// can reach a cold host at once. Without dedupe each one pays its own
// app-server session against a codex that has no such RPC surface.
it('dedupes concurrent probes on one host to a single app-server session', async () => {
const cache = new CodexAppServerCapabilityCache()
let releaseProbe!: (error: unknown) => void
const preferred = vi.fn(
() =>
new Promise<string>((_resolve, reject) => {
releaseProbe = reject
})
)
const first = cache.runWithFallback(
'native',
preferred,
() => Promise.resolve('fallback'),
isUnsupported
)
const second = cache.runWithFallback(
'native',
preferred,
() => Promise.resolve('fallback'),
isUnsupported
)
await Promise.resolve()
releaseProbe(unsupportedError)
await expect(first).resolves.toBe('fallback')
await expect(second).resolves.toBe('fallback')
expect(preferred).toHaveBeenCalledTimes(1)
})
it('lets a waiter run its own work once the in-flight probe reports support', async () => {
const cache = new CodexAppServerCapabilityCache()
let releaseProbe!: (value: string) => void
const firstPreferred = vi.fn(
() =>
new Promise<string>((resolve) => {
releaseProbe = resolve
})
)
const secondPreferred = vi.fn(() => Promise.resolve('second'))
const first = cache.runWithFallback(
'native',
firstPreferred,
() => Promise.resolve('fallback'),
isUnsupported
)
const second = cache.runWithFallback(
'native',
secondPreferred,
() => Promise.resolve('fallback'),
isUnsupported
)
await Promise.resolve()
releaseProbe('first')
await expect(first).resolves.toBe('first')
await expect(second).resolves.toBe('second')
expect(secondPreferred).toHaveBeenCalledTimes(1)
})
it('isolates in-flight probes per host so a cold WSL distro never waits on native', async () => {
const cache = new CodexAppServerCapabilityCache()
const nativePreferred = vi.fn(() => new Promise<string>(() => {}))
void cache.runWithFallback(
'native',
nativePreferred,
() => Promise.resolve('fallback'),
isUnsupported
)
const wslPreferred = vi.fn(() => Promise.resolve('wsl-result'))
await expect(
cache.runWithFallback(
'wsl:Ubuntu',
wslPreferred,
() => Promise.resolve('fallback'),
isUnsupported
)
).resolves.toBe('wsl-result')
})
it('builds host keys that keep WSL distros apart', () => {
expect(getCodexAppServerHostKey({ kind: 'native' })).toBe('native')
expect(getCodexAppServerHostKey({ kind: 'wsl', distro: 'Ubuntu' })).toBe('wsl:Ubuntu')
expect(getCodexAppServerHostKey({ kind: 'wsl', distro: 'Debian' })).toBe('wsl:Debian')
})
})