mirror of
https://github.com/stablyai/orca.git
synced 2026-09-28 08:02:43 +00:00
* fix(agent-trust): bound the remaining Codex trust writes a launch waits on #23148 capped the trust write on the agentTrust:markTrusted IPC handler and the worktree-remote startup path, but three call sites still awaited markCodexProjectTrusted with no bound: the main-process worktree startup used by orchestration worker-start, Codex quick-launch preparation, and Codex session resume preparation. All three share the per-config.toml lane with hook installs and app-server trust grants, which has no cap on queue depth, and the SSH writer adds a resolveHome round trip plus unbounded SFTP over a possibly half-open link. A wedged lane left the user clicking "start agent" with nothing happening and no error. Each site now routes through the existing awaitAgentTrustWriteWithinDeadline helper with unchanged semantics for its two current callers. Abandoning at the deadline never cancels the write, writes nothing else, and never substitutes a local fallback, so the workspace stays untrusted and Codex raises its own trust prompt. The await still sits ahead of the runtime-home resolution and hook repair that the PTY spawn waits on, so the ordering #23148 established is kept. Reliability only; there is no speed gain. The win is that a launch cannot hang. Verified with 2,755 agent-trust/Codex/startup/runtime tests, the 53-test reliability-gate command, the gate manifest check, typecheck:node, changed-code quality, oxlint and oxfmt. Red-green: with the three sites reverted to a bare await the new bounding tests hang to the 30s vitest timeout (3 fail/14 pass); restored, 17 pass. Co-Authored-By: Claude <noreply@anthropic.com> * docs(reliability-gates): stop the trust-preflight gate claiming a bounded launch The gate's performance budget said a stuck predecessor "degrades to the agent's own trust prompt instead of an indefinite caller wait." That is false at two of this branch's three new sites. In startup/codex-launch-preparation.ts the next awaited call after the abandoned write is codexHookService.prepareRuntimeHomeForLaunch; in startup/codex-session-resume-launch.ts it is installForLaunchPrep or refreshRuntimeUserHooksForLaunchPrep. Both reach runExclusivelyForRuntimeAndSystemTrustConfig, which takes the same runExclusivelyForCodexTrustConfig queue - FIFO per config.toml, no depth cap and no timeout - on the runtime home and on ~/.codex/config.toml that markCodexProjectTrusted itself takes. A wedged lane still stalls those two launches one step later, with the abandoned write holding its queue slot ahead of the hook step. Only markLocalWorktreeTrusted has nothing after its Codex branch and so is bounded end to end. The budget, invariant, oracle and coverage notes now say what is true: all five markCodexProjectTrusted call sites are bounded, so no trust write hangs a caller indefinitely, but the bound is on the write and not on the launch. A new knownGaps entry records the residual stall and that the new launch-prep tests mock codexHookService wholesale, so nothing in these suites can catch it. The already-complete-write case stays labelled a timer control, not a red. Wording only; no production code, test or count changed. The cited command still reports 6 files / 53 tests, and the gate manifest check passes for 138 gates. Co-Authored-By: Claude <noreply@anthropic.com> * docs(reliability-gates): name the first unbounded re-entry, not the hook-service call The trust-preflight gate cited codexHookService.prepareRuntimeHomeForLaunch as the immediate next step after the abandoned write in startup/codex-launch-preparation.ts. Tracing the module, the first awaited step is ensureRealHomeHooksIfSelected. It is conditional: only when the target is not WSL and runtimeHome.isHostSystemDefaultRealHomeSelected(launchEnv) is true does it call ensureRealHomeCodexHookState, which chains behind that module's own serial ensure promise and then acquires runExclusivelyForCodexTrustConfig directly on ~/.codex/config.toml - the same lane markCodexProjectTrusted's inner acquire takes. When the real home is not selected that call returns without awaiting the lane, runtimeHome.prepareForCodexLaunchAsync is synchronous on the non-WSL path, and only then is prepareRuntimeHomeForLaunch the first re-entry. So the gate could miss a launch stall one step earlier than the one it named. startup/codex-session-resume-launch.ts had the same omission: its hook branch takes ensureRealHomeCodexHookState when the resume home is ~/.codex, and only otherwise installForLaunchPrep or refreshRuntimeUserHooksForLaunchPrep. The coverage notes also understated the mocking - codex-launch-trust-write-deadline mocks ensureRealHomeCodexHookState as well as codexHookService and defaults the real-home selection to false, so neither lane is driven by any suite here. The performance budget, coverage notes and the residual-stall knownGap now name the first re-entry on each path and say when it applies, rather than only the later hook-service call. Wording only; no production code, test or count changed. The cited command still reports 6 files / 53 tests, and the gate manifest check passes for 138 gates. Co-Authored-By: Claude <noreply@anthropic.com> * docs(reliability-gates): scope the trust-preflight invariant to the local writes The invariant claimed every Codex trust write a launch waits on is bounded or abandoned at a deadline. The entry's own knownGaps contradicted that: the remote write reached through markRemoteWorktreeTrusted in runtime/runtime-worktree-agent-startup.ts is awaited with no timer, and it is a Codex write, not an adjacent one - it reads TUI_AGENT_CONFIG[agent].preflightTrust, which is 'codex' for the codex agent, and markRemoteAgentWorkspaceTrusted then branches into markRemoteCodexProjectTrusted. Its one production caller, markRemoteWorkspaceTrustedForAgent, is reached with a connectionId from seven launch entry points; five await it directly before the createTerminal that spawns the agent, and two gate the launch command they hand back for their caller to spawn. It chains a session.resolveHome round trip plus SFTP realpath/read/mkdir/ write, none with a timeout, over a possibly half-open link - the same hang this branch bounds locally. The invariant now claims only the five local markCodexProjectTrusted sites and names the remote exception inline instead of leaving it to knownGaps. Audited the rest of the entry against the code in the same pass: performanceBudget said "no trust write can hold its caller indefinitely" (now scoped to those five); the "all five awaited Codex trust writes" tally is now "all five awaited local" and points at the unbounded sixth; the remote gap no longer calls that path merely "preset-agnostic and out of Codex scope"; oracle and coverageNotes now record that nothing here drives markRemoteWorktreeTrusted, since the remote-preset suite only checks what markRemoteAgentWorkspaceTrusted writes, never how long it may take; and surfaces gains the remote writer that suite actually covers. Re-verified the rest rather than assuming it: the launch-prep re-entry chain, the FIFO no-cap no-timeout trust-config queue, the IPC handler capping both its remote and local writes with no cross-branch fallback, and that upsertProjectTrustLevel and upsertProjectTrustLevelInContent remain the only producers of a project trust_level, so no further writer needs covering. The already-complete-write case stays labelled a timer control, not a red. Wording only; no production code, test or count changed. The cited command still reports 6 files / 53 tests, and the gate manifest check passes for 140 gates. Co-Authored-By: Claude <noreply@anthropic.com> --------- Co-authored-by: Claude <noreply@anthropic.com>