Commit Graph
12 Commits
Author SHA1 Message Date
Neil 0a81211994 fix(cursor): read desktop login outside the main thread (#24572)
* fix(cursor): read desktop login outside the main thread

* fix(cursor): await native worker retirement before respawning
2026-10-02 03:48:33 -07:00
Brennan Benson cfa43e7eab fix(codex): opening a terminal no longer strips Codex hooks from the real ~/.codex (#23552)
* fix(codex): a real-home restore leaves a file alone once someone else changed it

Orca writes ~/.codex/hooks.json (and a trust rebase writes config.toml), then
runs a Codex trust session for up to 10 s, then restores the original bytes if
the session fails. The restore wrote unconditionally, so a save that landed
during the session, from the user or another Orca, was silently reverted.

Each restore now compares first: it writes the original back only while the
file still holds the generation Orca's mutation left, and otherwise logs and
leaves it alone. This covers the real-home install and opt-out sweep
(restoreRealHomeHooksJson), the legacy sweep's hooks restore, and config.toml
rollback (restoreCodexTrustConfig).

For hooks.json the generation is the exact bytes Orca wrote. For a config.toml
that a trust rebase changed it is the file as the rebase left it. When Codex
itself wrote config.toml inside the session that just failed, Orca never knew
those bytes, so that rollback compares against the file as the session settled.
The next commit keeps other Orca instances out of that window; a user edit made
during such a session can still be rolled back.

* fix(codex): serialize real-home Codex writes across Orca instances

Every Orca on one HOME (a dev and a packaged app, or an offline CLI) writes the
same ~/.codex/hooks.json, config.toml and ~/.orca/agent-hooks/codex-hook.sh.
The per-file lane that orders capture, mutate and restore was in-process only,
so another instance could write inside this one's restore window, or undo it.

The lane for the user's real config.toml now also holds the existing
crash-safe managed-hook install lock (~/.orca/managed-hook-install.lock, the
one relay installers take for the same home). It is taken only by the
outermost acquire, because the lock file is not reentrant and grants and trust
rebases nest inside an install. Managed-home installs, the real-home install
and opt-out sweep, and the legacy sweep all enter through it. Compare-and-swap
on restore stays as the backstop.

A lock that cannot be taken within its 10 s wait fails that install, which is
already best effort: launch prep logs it, and the real-home lane falls back to
the managed lane until its retry.

* fix(codex): opening a terminal no longer strips the shared Codex entry from ~/.codex

Every Orca instance on one HOME writes the same status-hook entry into the
user's ~/.codex/hooks.json, with its trust in config.toml. Launch prep runs on
every pane spawn, and under a managed Codex account it ran the legacy system
sweep. That sweep matched Orca entries by script file name, so it removed the
current shared entry and the trust blocks the grant ledger recorded. On a live
laptop hooks.json went 4139 -> 18 bytes about 150 ms before a new pane opened.
With hooks off, the real-home lane's launch prep swept the same way.

Now nothing automatic removes the current entry or its trust:
- The legacy sweep removes only an enumerated list of retired command forms
  that no build writes any more (#1019's double-quoted form, #1536's
  exec-guarded form, and Windows' per-userData bare path), plus their trust.
- ensureRealHomeCodexHookState with hooks off writes nothing; that covers
  launch prep, session resume and startup.
- Only the user's explicit opt-out (codexHookService.remove()) strips the entry
  and its ledger-recorded trust from the real home.
- The sweep-suppression gate existed only to stop the sweep from deleting the
  current entry, so it is deleted with its main-process wiring.

Startup with hooks off already skipped the real-home install; with this change
the first pane's launch prep with hooks off also leaves ~/.codex untouched.

* fix(codex): a pane's prepare-codex only repairs a home its own HOME's app installed

On macOS a pane starts through login(1), so it gets the user's real HOME even
when its Orca app runs with another one. The pane's `codex()` preflight
installed hooks in the CLI process with that real HOME: it rewrote
~/.orca/agent-hooks/codex-hook.sh, promoted trust into the real config.toml,
and wrote the real HOME's script path into the app's managed home.

The preflight now acts only when the managed home's hooks already run this
process's own shared script, which proves the app that installed them shares
its HOME. Otherwise it writes nothing; the app installed the home at spawn.

Why not a no-op: the preflight was added (#14326) because trust can go stale
between opening a pane and typing `codex`, for example in a pane that survives
an app update, and Codex then stops in hook review. For a same-HOME pane it
still repairs that. Why keep promotion: the install drops runtime trust the
system config does not back, so skipping promotion would delete approvals the
user gave inside Orca-launched Codex.

* test(agent-hooks): await every installer in the refresher coverage test

The test fired each managed installer without awaiting it and read
~/.orca/agent-hooks straight after. Codex's install now takes the
cross-process real-home lock before it writes its script, so the script
landed after the read. Await the installers, and stub Codex's trust sessions
so the awaited install cannot start a real `codex app-server`.

* fix(codex): retire the two real-home command forms the list missed

The real-home lane wrote two Codex hook forms into ~/.codex that no build
writes any more and that the enumerated retired list did not name:
- POSIX, #9501 until #10885: the file-guarded form draining with a bare `cat`.
- Windows, #9501 until #10221 took Windows off the real-home lane: the encoded
  PowerShell launcher for a non-cmd-safe script path.

The file-name sweep removed both before; the enumerated sweep left them in
place, trusted, still passing the script's exit status to Codex. Both now
match as frozen literals.

Also corrects the startup ordering comment: the real-home install runs first
so its in-slot upgrade lands before the managed install's sweep retires the
prior command; nothing re-arms a legacy sweep any more.

* fix(codex): take the real-home lock only when a write is needed

The previous commit made every entry to the real-home config lane take the
cross-process lock. That lane runs on every pane spawn and every typed
`codex` preflight, so the steady state paid an owner probe (a `ps` spawn on
macOS) and could wait up to 10 s behind another instance's trust session,
even though it wrote nothing.

Each real-home writer now compares the desired state with the files on disk
first, without the lock. Only when a write is needed does it take the lock,
re-read and recheck, then write:
- real-home install: the planned hooks.json, the shared script and the
  ledger-recorded grant are compared; the locked path re-plans from disk.
- legacy sweep: locks only when a retired entry is present; the sweep re-reads.
- approval promotion: locks only when there is something to promote; the
  promotions are recomputed under the lock.
- the shared ~/.orca/agent-hooks script: locks only when its bytes differ.
The explicit opt-out always takes the lock. The lock is reentrant through
async context, since grants and rebases nest inside an install, so the
config-lane option the previous commit added is removed.

* fix(codex): a shared script without its exec bit is not the steady state

The compare-first check matched the shared ~/.orca/agent-hooks script on bytes
alone. writeManagedScript also restores 0755 on every call, and the POSIX hook
guard skips a script that is not executable, so a script whose mode was lost
(a dotfiles restore, a plain copy) now stayed that way: every Codex hook
drained stdin and reported nothing until an app restart refreshed the script.

The check now also requires the mode the writer sets, so that case takes the
lock and the write path repairs it.

* test(codex): the retired encoded launcher never matches today's shared one

The shared encoded Windows launcher is still current for other agents, so the
comment claiming today's launcher is never encoded was wrong. What keeps the
retired matcher off it is the exact payload: since #14825 the shared launcher
prefixes its payload and drops -ExecutionPolicy Bypass. Pin that with a case.

* fix(codex): the pane step recognises its own script under a home path with an apostrophe

The same-HOME check looked for the script path wrapped in bare single quotes,
but both hook writers escape an apostrophe inside the quotes. A home such as
C:\Users\O'Brien never matched, so the pane-step repair never ran there.

* fix(codex): the trust-RPC escape hatch still keeps the real home off its lane

The no-write check reported a recorded grant as current, so with
ORCA_DISABLE_CODEX_TRUST_RPC set the real-home lane stayed in use. The grant
itself refuses before reading its ledger; the check now does the same.

* fix(codex): the shared script write no longer waits on the real-home lock

The write is atomic and skips identical bytes; waiting behind another
instance's trust session could only fail a pane's managed-home install.

* fix(codex): an in-Orca approval survives a launch that cannot get the real-home lock

The install drops runtime trust the system config does not back, so a
promotion skipped for want of the lock lost the approval for good. It
now writes unlocked, as it did before the lock existed.

* refactor(codex): take the cross-process real-home lock back out

The lock fixed no observed failure. The three that were observed each have
their own fix in this series: the legacy sweep matches only frozen retired
command forms, hooks-off launch prep writes nothing, and a pane's
prepare-codex repairs only a home its own HOME's app installed. The lock
instead brought its own defects: a steady-state spawn waiting behind another
instance's trust session, a compare-first split to avoid that, a script
write and an approval promotion that could fail for want of the lock.

Removed, with their tests: the real-home write lock and its async-context
reentrancy, the plan/compare split that kept it off steady-state spawns, the
compare-first legacy sweep, the locked approval promotion and its unlocked
fallback, the compare-first shared script write (writeManagedScript already
skips identical bytes and restores the exec bit), and the CLI tsconfig
entries the lock pulled in.

Kept: the retired-forms matcher, the hooks-off no-op, removal only on an
explicit opt-out, the pane own-script check, and the compare-and-swap
rollbacks. Every instance now writes identical bytes idempotently.

* fix(codex): an opt-out that cannot read hooks.json keeps Orca's trust and ledger

The opt-out swept the real-home entry, then dropped Orca's ledger-proven trust
whenever a ledger existed, even when the sweep could not read hooks.json. The
entry could still be there, now untrusted, and the ledger that proves ownership
was gone for the retry. Drop that trust only after a sweep that read the file.

* refactor(agent-hooks): one predicate for whether an agent's status hooks are on

"Global switch on and this agent not turned off" was spelled out separately
in the startup controls, the settings reconcile, the retained-home
reconcile, the WSL preflight RPC, the CLI preflight and the OpenCode plugin
selection. They now share one function, in a module light enough for the
CLI's per-launch Codex preflight to load. The PTY spawn env derives the
Codex flag from the switch and opt-out list it already carries, the same way
it does for OpenCode and Pi, instead of receiving a second copy.

* fix(codex): launch and resume prep honour Codex's per-agent hook opt-out

Turning Codex off in the per-agent hook settings removes Orca's Codex hook
entry, but launch prep and session resume read only the global hooks switch,
so the next Codex launch or resume wrote the entry straight back into the
real ~/.codex or the account's home. Both now read the per-agent predicate,
which the PTY spawn env and startup already honoured.

* fix(codex): turning Codex off per agent clears the real ~/.codex entry

While the real-home lane owns ~/.codex/hooks.json, the legacy system-home
sweep stands down. That gate read only the global switch, so turning Codex
off per agent ran remove() with the sweep still suppressed and left Orca's
entry in the real ~/.codex. The gate now reads the per-agent predicate, the
same as turning every hook off.

* test(codex): cover the system ~/.codex sweep gate for Codex turned off

The gate that lets the legacy system-home sweep run was an inline closure in
startup, so reverting it to the global switch left CI green. It is now a
pure function beside the gate it feeds, with a table test and a remove()
test on a seeded ~/.codex: turning Codex off strips Orca's entry and keeps
user hooks; with Codex on the entry stays.

* fix(cli): keep the agent-status hooks predicate loadable by the packaged CLI

The CLI's prepare-codex handler imported the predicate from src/main, but
the Electron build rebuilds out/main from its declared entries only, so the
packaged `orca agent hooks` commands could not load it (package jobs and the
CLI bundle-parity test were red). The predicate reads only settings, so it
now lives in src/shared, which the CLI compiles itself.

* feat(codex): every Orca build writes one frozen Codex hook command

The Codex hook command was built from this build's wrapper, so two builds
on one HOME disagreed about the bytes of the shared ~/.codex entry and kept
rewriting it, with a Codex trust session each time.

The command is now fixed per form and carries its form number:
- POSIX: one command with no path in it. It runs the shared script only in
  an Orca pane with hooks on (pane key and hook port set), drains stdin
  everywhere else, and always exits 0. A branch for a per-build script root
  is written now and stays dormant until Orca sets ORCA_AGENT_HOOK_ROOT, so
  that change will not move these bytes.
- Windows: the bare forward-slash path to the shared .cmd, which runs under
  PowerShell 7 and 5.1, Codex's hook hosts. A profile path that is not one
  PowerShell token gets a plain PowerShell form with the same branches.

The literals live in the form module, so a change to the shared hook
constants cannot move them; goldens pin the bytes. Every form keeps
`agent-hooks/codex-hook.*` in plain text, so older builds still recognize it.

* fix(codex): one main-process owner adds the real-home entry; nothing restores files

Each Orca writer of ~/.codex decided what Orca's entry must be from its own
build and instance, then removed or reverted whatever differed: launch prep
rewrote any Orca-shaped entry to this build's command and stripped Orca
entries from events this build does not use, and a failed trust session
restored hooks.json and config.toml from snapshots. With several instances
and builds on one HOME, every disagreement became a deletion or a revert.

The main process is now the one writer, and its writes are add-only:
- A launch or resume adds Orca's frozen entry to an event that has none and
  leaves every Orca entry it finds, so a running older build is never fought.
- App start also converts an older Orca form to the frozen command, once, in
  its own slot: one hooks.json write (one .bak) and one trust grant per home.
- A newer form is never rewritten or appended beside, and Orca entries in
  events this build does not use are kept.
- After a failed trust grant, only an entry this call wrote that is still
  untrusted is withdrawn, putting back the handler it replaced. Both files
  are re-read, so a concurrent edit, or the identical entry another Orca
  trusted meanwhile, survives.

Deleted: the compare-and-swap hooks.json restore, the config.toml snapshot
restore after a grant session and after a user-trust re-key, and the
rollback module. A grant session writes trust only at Orca's own keys, and
every caller settles those keys itself. A failed re-key of moved user hooks
now keeps the write and reports it; Codex lists those hooks for review.

* fix(codex): the pane CLI asks the app to prepare its Codex home

`orca agent hooks prepare-codex` ran Codex's install inside the pane. That
process can have the real HOME (login(1)) and runs outside the app's
in-process queues, so it was a second writer of ~/.codex and ~/.orca beside
the app. A check that the home ran "its own script" guarded it.

The pane step now only asks the app, over the same kind of local RPC the WSL
pane step already uses (agentHooks.prepareCodexForPane). The app checks that
the pane's CODEX_HOME is one its own userData owns, reads its own hooks
setting, and installs on its own queue. An app that is not running, or is
too old to know the method, makes the step a no-op, as it is on WSL. The
own-script check and the CLI's settings read are gone, and the preflight
module leaves the CLI bundle.

* fix(codex): delete the pane step on native hosts

The previous commit had `orca agent hooks prepare-codex` ask the app to
prepare the pane's Codex home. The case it existed for (#14326, a pane that
survives an app update with stale hook trust) did not reproduce, and no other
desktop agent host writes agent config from a terminal or launch wrapper.

- Deleted: the agentHooks.prepareCodexForPane RPC method, its params and
  catalog entry, and prepareManagedCodexHomeBeforeShellLaunch with its module,
  tests and CLI build entry.
- `agent hooks prepare-codex` is a no-op on native hosts. It stays for one
  release so shell wrappers from older builds, which still call it, exit 0.
- WSL panes are unchanged: they still ask the app over
  agentHooks.prepareCodexForWslPane.

The shell wrappers and ORCA_CODEX_LAUNCH_PREFLIGHT stay, because WSL panes
use the same wrappers and variable (forwarded through WSLENV). A native pane
still starts the CLI once per `codex` it runs; skipping that is a follow-up.

* test(codex): a failed trust session keeps concurrent edits to both files

QA case 9 at host level, on a real file system in a temp HOME: Codex's trust
session fails after another writer saved hooks.json and config.toml.

- Both saves survive, and no Orca entry is left that Codex would list for
  review: this call's entry is withdrawn.
- A failed one-time conversion puts the older Orca entry back in its slot and
  keeps both saves.

Both tests fail on the previous head, which restored config.toml from a
snapshot and left the untrusted entries in hooks.json. Removing the
withdrawal turns both red.

* feat(codex): read whether an Orca entry's stored trust is still current

A Codex release that changes how it hashes a hook leaves Orca's stored trust
stale: the entry is present, but Codex lists it as modified. Checking only
whether the entry is missing cannot see that.

readOrcaEntryTrust sorts a present entry into four states:
- trusted: the stored hash is the current one;
- untrusted: there is no stored hash;
- stale: the stored hash is not the current one;
- disabled: the user turned the entry off.

The caller can pass Codex's current hash, for example one a grant recorded.
The failed-grant withdrawal now uses it, and also keeps an entry the user
turned off. Nothing re-grants on 'stale' yet.

* fix(codex): a slow Codex start retries on the next launch, never for minutes

On a loaded Mac a cold `codex app-server` took over 10 s (QA case 4). The
grant timed out, the entry was withdrawn, and a 5-minute cooldown in both the
grant and the real-home install then refused every retry.

- The native session deadline is 30 s, the same as WSL's.
- A timeout starts no cooldown in the grant or in the real-home install. The
  next launch retries. Other failures keep their cooldown.
- Launches that queue behind a slow session share one follow-up run, so a
  launch waits for at most two sessions, not one per earlier launch.

Tests: a 15 s cold start still grants and keeps the entry; after a timeout,
the next launch runs a session at once; four queued launches run two
sessions. Each is red on the previous head, and each mechanism was removed in
turn to confirm its test turns red.

* fix(codex): Orca's automatic writes never move a user hook

Codex keys a hook's trust by its position in hooks.json. App start's collapse
of Orca duplicates removed every Orca entry and appended one at the end. That
moved any user hook that followed a removed entry, so the write waited on a
session to re-key the moved hook's trust.

App start now:
- converts the first Orca entry that sits in a plain slot to the frozen
  command, in place;
- drops any other Orca entry only when that moves no user hook;
- keeps a duplicate that a user hook follows, and trusts every frozen copy,
  so none is listed for review;
- appends only when no frozen entry is left.

Tests check user positions and user trust blocks byte-for-byte for each
automatic write: add-missing (append), the one-time conversion (in place),
a trailing duplicate, a duplicate before a user hook, and older duplicates
normalized to one entry. The three collapse cases fail on the previous head.
Removing the position check, or the in-place conversion, turns its tests red.
Only the explicit opt-out still removes an entry that user hooks follow.

* fix(codex): removing an Orca entry never waits on a Codex session

Removing an Orca entry from ~/.codex/hooks.json moves every user hook behind
it up a slot, and Codex keys trust by slot. The retired-form sweep, the
opt-out and a failed-grant withdrawal all asked a `codex app-server` session
to list the old trust before writing, and to re-key it afterwards. A timeout
there threw before the write and latched a 5-minute cooldown, so a slow cold
start blocked the retired-form sweep at boot (QA case 4).

Each moved hook's [hooks.state] block now moves to its new key, body bytes
unchanged, straight after the hooks.json write. Codex hashes a hook's content,
not its position or its file path, so the moved block stays exactly as valid
as it was: a trusted hook stays trusted, an untrusted one stays untrusted, and
one the user turned off stays off. No removal waits on or depends on a
session. A failed config.toml write keeps the hooks write and logs.

Deleted: the inspect and repair sessions, their client, and their cooldown.
The generation guards on the hooks.json writes stay, for other processes.

Tests: the retired sweep removes the retired entry and carries the trust of
the user hook behind it while every Codex session times out (red on the
previous head); the opt-out carries an appended user hook's trust; the move
carries trusted, disabled and untrusted states byte for byte. Removing the
move turns all of them red.

* fix(codex): a Codex launch never waits on Codex's approval of Orca's entry

A launch on the real-home lane awaited Codex's trust grant for the entry it
had just added. A cold `codex app-server` on a loaded Mac took over 10 s, so
the launch could wait that long, and a failure then latched a 5-minute
cooldown.

- Codex's approval runs in the background, with a 30 s cold-start budget.
- A launch uses the real home only when the ledger shows trust is already
  current. Otherwise it goes to the managed home at once, and the next launch
  picks up the finished grant.
- A launch that arrives while a grant runs does no work and does not queue
  behind it.
- A resume into the real home has no managed home to fall back to. It waits
  for the grant, but no longer than the 10 s a launch always could.
- A background grant that times out starts no cooldown; the next launch
  retries. Any other failure backs off for 10 s instead of 5 minutes.
  Success is what the ledger remembers.
- A failed grant still withdraws only what that install added and is still
  unapproved. The log now says how many entries it took back and when the
  next try comes.

Managed-home grants keep their 10 s deadline and stay on launch prep, as
before; they fall back to Orca-computed trust.

Tests:
- A 15 s start: the launch returns in under a second on the managed home, a
  second launch starts no session, the grant lands in the background, and the
  next launch uses the real home.
- A timeout sets no cooldown, withdraws its adds and logs it.
- Another failure retries after 10 s, not before.
- A resume waits only as long as allowed.
- Case 9 checks the log line and the retry.

Making the launch await the grant, a 10 s budget, either timeout cooldown, and
a 5-minute backoff were each tried, and each turns its test red.

* fix(codex): move a hook's trust only when every stored key has the known shape

Orca now edits Codex's trust store directly when a removal moves a user
hook. Three safeguards keep that honest:

- Fail safe. If any [hooks.state] key in config.toml does not have the
  shape `<path>:<event>:<group>:<handler>`, nothing moves and Codex asks the
  user to review. That shape was checked unchanged from Codex 0.141 to 0.158.
- Targeted. The file is read immediately before the atomic rename, and only
  the moved keys' blocks change. Every other byte stays, and no snapshot is
  restored.
- Verbatim. Each block's body moves as Codex wrote it, including fields
  Orca does not know. No hash is ever computed, and a hook with no block
  gets none.

Tests:
- An unknown key shape stops every move.
- Everything except the moved block survives byte for byte, and the moved
  body keeps an unknown field.
- In case 9, a hook the user approved during the failed session keeps its
  approval when the withdrawal moves it, beside the concurrent project edit.

Removing the shape check, or writing a computed block instead of the stored
body, turns these tests red.

* refactor(codex): keep only the trust read the failed-grant withdrawal uses

A capture across Codex 0.141, 0.150 and 0.158, switching in all six
directions, showed Orca's entry keeps the same hash and stays trusted. A
Codex upgrade does not make its trust stale, so nothing needs to re-grant
on staleness.

readOrcaEntryTrust keeps the four states the withdrawal needs, but loses
the parameter that let a caller pass a different current hash, and the test
for a Codex that hashes differently.

* fix(codex): native panes no longer start the Orca CLI before each codex

The pane step is a no-op on native hosts, but native panes still carried
ORCA_CODEX_LAUNCH_PREFLIGHT, so every `codex` typed in a pane started the
Orca CLI for nothing. Only a packaged Windows build's WSL pane now gets the
variable; the app prepares every native Codex home itself.

The resolver loses the dev-launcher path and its userDataPath option, which
only native panes used.

Tests: a native macOS, Linux and Windows pane gets no preflight, packaged or
not, even with the bundled CLI present; a WSL pane still gets the verified
absolute launcher. Letting native panes through again turns them red.

* chore(cli): say when the native prepare-codex no-op can go

Native pane wrappers from builds up to v1.4.216 still call it. It can be
deleted once no supported build's wrapper does.

* test(codex): check the WSL launcher path instead of asserting it

* fix(codex): a launch no longer waits behind the background real-home approval

The background grant ran its whole codex app-server session inside the shared
~/.codex/config.toml lane, and on a cold host its session was also the shared
capability probe. A launch sent to the managed home then waited on both: the
managed install and the project-trust write queue on that lane, and the
managed install's own grant waited for the probe. On a cold app-server that
was up to 30 s per launch.

The lane was held across the session only to protect the retired
capture-and-restore. Codex writes its own records, so the lane is now taken
only around Orca's own pre-grant write. The background grant runs its session
without publishing it as the shared probe, and the whole grant is bounded by
its deadline, so a hang outside the session cannot leave the lane 'granting'.

* fix(codex): a failed re-grant no longer strips Codex's own approval of Orca's entries

Before each trust session, the grant deleted every Orca record whose hash
matched the one Orca computes. That exists because a managed home's fallback
writes Orca-computed trust under both Windows path-separator spellings, and
Codex rewrites only its own spelling, so the other copy would linger. On
failure the managed and WSL fallbacks write that trust back, and before this
fold a snapshot restore covered it.

The real ~/.codex has neither: Orca never writes computed trust there (the
real-home lane does not run on Windows at all), so a matching record there is
Codex's own approval. After a ledger miss (another Orca profile, a Codex
update, a lost ledger) and a failed session, nothing put it back, and every
Orca entry showed "Hooks need review".

The clear now runs only for homes whose fallback writes that trust.

* fix(codex): a real-home resume spawns only once Orca's entry is approved or withdrawn

A resume that must run in ~/.codex waited at most 10 s for the background
approval, then spawned anyway. On a cold app-server that left Codex beside an
unapproved Orca entry, so the resumed pane showed hook review.

The resume now waits for the grant to settle. Settled means Codex approved the
entry, or the grant failed and withdrew its own unapproved write; the grant's
deadline bounds the wait (30 s, the cold-start budget), and a failed approval
never fails the resume.

Why this over the alternatives:
- Spawning at 10 s keeps the review prompt this fold exists to remove.
- Withdrawing at 10 s from the resume races the still-running session: Codex
  can write the frozen entry's hash after the withdrawal, and for a converted
  entry that marks the older command Orca put back as modified.
- A resume cannot use the managed home: the session lives in ~/.codex.
So the only states that cannot race Codex are the grant's own settle. The cost
is a longer worst case on a cold app-server (up to the 30 s deadline, plus any
managed-home install that holds the config.toml lane); a warm approval takes
seconds, and an approved entry costs no wait.

* fix(codex): keep the 5-minute trust cooldown for launch-path grants

The fold shortened the host's trust-grant cooldown from 5 minutes to 10
seconds for every grant. That was meant for the background ~/.codex approval,
which blocks no launch. The managed-home and WSL grants run inline on the
launch path, so with a hung app-server every launch more than 10 s after the
last failure paid the full inline timeout again (10 s native, 30 s WSL).

Cooldowns are now kept per lane: inline grants keep 5 minutes, the background
grant retries after 10 s, and neither lane's failure cools the other down. A
success, or a proven-missing surface, still clears both. The real-home
install's own retries (an unreadable hooks.json, unknown keys) are back on the
5-minute interval they had before the fold.

The cooldown moves to its own module so the grant stays within the file limit.

* fix(codex): a failed grant withdraws the exact copy it wrote

The withdrawal re-found "this call's" entry by command, taking the first
frozen handler in the event. When app start converted a later slot while an
earlier frozen copy sat in a matcher group (which conversion skips), a failed
grant acted on that earlier copy: it put the older command into it, or skipped
it, and left the converted, unapproved copy in place.

Each write now records where its handler landed, after any duplicate drops,
and the withdrawal acts only on that slot. A copy that has since moved is left
alone; the next launch's grant retries it.

* fix(codex): the failed-grant withdrawal checks hooks.json is unchanged before writing

The install and the retired-form sweep both refuse to replace ~/.codex/hooks.json
if it changed since they read it. The withdrawal did not: a save landing
between its read and its atomic replace was lost. The window is small, since
the withdrawal is synchronous, but it now carries the same guard.

* refactor(codex): drop rationale left over from the snapshot restore; name the trust-move module for what it does

Comments on the config.toml lanes still justified them by a grant's
capture-and-restore window, which the fold deleted, and the trust-write
deadline still counted a grant session holding the lane. They now give the
reason that remains: Orca's own multi-step reads and writes, and managed-home
installs that hold the lane across their inline grant.

codex-user-hook-trust-rebase no longer rebases through Codex; it moves stored
trust records, so it is now codex-user-hook-trust-moves.

The grant test that pinned two sessions on one config.toml to run one at a
time is removed: its reason was an interleaved capture and restore. Callers
that write config.toml around a grant hold their own lane, which the nested
installer test still covers.

* build(cli): list the trust-grant cooldown module in the CLI program

The CLI's agent-hooks handler loads the hook controls, which reach the Codex
trust grant; the CLI project is composite, so every module in that graph must
be listed.

* docs(codex): say which Windows hosts each hook command form runs under

Codex runs a hook under the turn's shell (PowerShell 7 or 5.1 in every
captured session) and, with no single local turn shell, under %COMSPEC% /C.
The bare forward-slash path ran under all three in the Windows host census.
The PowerShell form used for a profile path with a space does not parse under
cmd.exe; no form valid in all three hosts has been run for such a path, so the
form stays and the gap is stated here and in the PR.

* test(codex): type the withdrawal seam without an assertion

* fix(codex): a real-home resume starts at once, trusting Orca's entries for that process

A resume that must run in ~/.codex waited for Codex's background approval of
Orca's newly written hook entry: up to 30-40 s on a cold app-server. That made
the user's resume wait on bookkeeping, and the alternatives (start at 10 s with
Codex's hook review showing, or withdraw the entry and race Codex's own write)
were worse.

Codex reads hook trust from its session-flag config layer as well as the user's
config.toml, merged per key, and has since hook trust shipped. So the resume no
longer waits. When Orca's own frozen entries in ~/.codex are untrusted (or hold
a stale hash), the resume command carries
`-c hooks.state={'<key>'={trusted_hash='<hash>'},...}` for exactly those entries:
the key under both the logical and the real path of ~/.codex (Codex keys an
explicit CODEX_HOME by its real path), and the hash of that entry's content, so
it can trust nothing else at that slot. The user's hooks are never included,
nothing is written, and the background approval still runs for later plain
`codex` launches. An approved entry adds nothing; a Codex known to lack hook
trust gets nothing.

One inline table, because Codex splits a `-c` key on every `.` and the key holds
`.codex/hooks.json`. TOML literal strings keep `"` out of Windows native-argument
quoting. The flag goes before `resume <id>`, quoted for the pane's shell (portable
Unix, PowerShell or cmd), in the launch command and in the setup-sequenced copy of
it; a cmd line whose path cmd would expand, or a key with an apostrophe, is left
unchanged. SSH and WSL resumes get no preparation, so no local path reaches them.

* Revert "fix(codex): a real-home resume starts at once, trusting Orca's entries for that process"

This reverts commit 1bd30651d6.

* fix(codex): a real-home resume starts at once, without waiting for approval

A resume into the real ~/.codex waited until the background approval settled,
up to its 30 s deadline on a cold app-server: bookkeeping for later launches
gating the resume the user asked for. It now starts at once. If the approval is
still running, that first resume can show Codex's hook review once; the
approval then lands and later resumes and plain codex launches are trusted.

Trusting Orca's entries per process was the alternative, but the resume command
is typed into the pane's shell, and hook settings stay out of typed commands.

* test(codex): read real-home hook groups with the installer's own type

* fix(codex): a background approval is bounded only by its session's own deadline

Review loop 2, L3. grantWithinDeadline raced a second 30 s timer against
the background approval. Loop 1 added it so that a hang upstream of the
session could not leave the lane 'granting' forever.

That hang cannot happen. The only caller is the native real-home grant
(its plan is always host 'native'; the real-home lane is off on Windows,
so WSL never reaches it). Everything before the session is synchronous
there: command resolution and binary stamp, the ledger read, the
state-db backfill check, the capability and cooldown checks, and
runUnshared awaits no shared probe. A synchronous hang would freeze the
main thread, which no timer can rescue. The session itself starts a kill
timer right after spawn (runCodexAppServerSession), with the same 30 s,
and it kills the app-server tree when it fires.

So the outer timer was a second copy of that bound. Because it started
first, it won by the spawn time. It then settled the lane and cleared
backgroundGrant while the app-server was still alive, and the next
launch could start a second concurrent session. It abandoned the
session rather than cancelling it. Deleted, not moved: the session's
own timer is the one bound, and it cancels.

Test: codex-real-home-slow-app-server.test.ts "runs one session at a
time, ended by its own deadline". The fake session starts its timer
after a simulated spawn, as the real one does. A launch at 30 s finds
the session still running and starts none; the lane settles when the
session times out. It replaces the "settles a grant that never answers"
test, whose never-answering session could not time out at all.

* fix(codex): a background approval's retry has one schedule, the real-home lane's

Review loop 2, L4. A non-timeout background failure set two 10 s
schedules for one failure: the real-home lane's installRetryAfterMs,
which gates ensure, and a `<host>#background` cooldown in the grant
module. ensure's gate always tripped first, so the second one was
consulted only after something reset the first (turning hooks off).
Then it answered 'retry-cached', which wrote the entry into
~/.codex/hooks.json only to withdraw it again: churn, not protection.

Background plans now neither start nor consult a grant-module cooldown.
The real-home lane (installRetryAfterMs) is the one source of truth for
when a background approval runs again, and its 10 s interval moves into
codex-real-home-background-grant.ts, the module that sets it. The
cooldown module is back to one host-keyed map for launch-path grants,
with the same 5-minute interval as main. A success or a proven-missing
surface from either lane still clears the host's cooldown.

Tests:
- codex-hook-trust-grant.test.ts "neither starts nor waits on a
  cooldown for a background grant": two failing background grants each
  run a session and leave no cooldown; an inline failure still cools
  down inline grants and not the background one.
- codex-real-home-slow-app-server.test.ts "has one retry schedule:
  turning hooks off and on after a failure retries at once": after a
  failed approval, hooks off then on runs a session and installs,
  instead of a retry-cached write-and-withdraw.

* fix(codex): hooks turned off and on during an approval re-add Orca's entry

Review loop 2, L1. ensure returned at once whenever a background
approval was running, whatever the lane. Turning hooks off during an
approval sets the lane to 'removed' (usable), so turning them back on
returned 'removed' without re-adding the entry. Launches in that window
spawned in ~/.codex with no Orca hook and got no status for their
lifetime, for up to 30 s, until the approval settled and a later launch
re-added it.

ensure now returns early only while the lane is 'granting', which is
what the early return exists for: a launch never waits on Codex's
approval and uses the managed home until it lands. Any other lane runs
the normal add-missing install.

That install can start a second approval while the first is still
running. Approvals are now chained, so Codex still runs one session at
a time, and a finished approval clears the handle only if it is still
the latest one (before, an older approval's finally could clear a newer
one's handle). The older approval's result is already dropped by the
lane generation check.

Test: codex-real-home-slow-app-server.test.ts "re-adds the entry when
hooks go off and on during an approval, one session at a time". While
the approval hangs: opt-out removes the entry; re-enable re-adds every
entry, keeps launches on the managed home, and starts no second
session; once Codex answers, the lane is installed and every entry is
approved.

* test(codex): a launch during the real-home approval shows what it waits on

Review loop 2, M2. The launch test's fake Codex failed every
managed-home session at once with ENOENT, so the managed home's own
approval was an instant "unsupported" fallback, and the test could not
show that a launch sent to the managed home still waits on that home's
inline approval when its ledger misses (first use, a Codex update, a
lost ledger), up to 10 s, as on main.

Now the managed-home session behaves like a real one:
- "settles on the managed home with its hooks and the project trust
  written": the managed app-server answers; two launches settle in
  under 2 s while the real-home approval hangs, and the second launch
  finds the managed approval in its ledger (one managed session).
- new "waits up to the managed home's own 10 s approval when that home
  is cold too": the managed session fails at its own deadline, as the
  real one does. The first launch is still pending at 9.999 s and
  settles on the managed home at 10 s; the request asked for 10 s. The
  next launch settles at once, because the failed inline approval cools
  down for 5 minutes.

No product change.

* refactor(codex): the managed and WSL installs own their pre-approval trust clear

Review loop 2, L7. Before a Codex approval session, a managed or WSL home
clears the approvals Orca itself computed, because on Windows its
fallback writes them under both path spellings and Codex's canonical key
may not overwrite the other one. The fallback writes them back if the
session fails. ~/.codex has no such fallback, so there the clear would
only delete Codex's own records (loop-1 H2). The grant module carried
this as a plan flag, fallbackWritesSelfComputedTrust, and took the
config.toml lane around the clear itself.

The reviewer proposed moving the clear into the two callers. A literal
move, clearing before the grant call, is NOT behaviour-neutral, so this
does not do that:
- The grant first checks its ledger, which compares the stored hash
  with the one Codex recorded. Codex's hash equals Orca's computed one
  (the premise of readOrcaEntryTrust), so a clear before that check
  deletes exactly the record the ledger proves. Every managed launch
  would then miss the ledger and run an inline session (up to 10 s).
- Checked, not inferred: with the clear moved before the call in the
  managed install, codex-launch-during-real-home-grant.test.ts "settles
  on the managed home..." fails (2 managed sessions instead of 1).
  Log: ~/orca-qa/codex-real-home-leak/fb6/l7-literal-move.log

What this does instead: each caller passes its clear as the grant's
`beforeSession` step, which the grant runs only when a session will
actually run (after a ledger miss, and not on a cooldown or cached
fallback), exactly where the flag ran it. So:
- the flag and its "never set for the real home" rule are gone; the
  real-home grant passes no step, so the grant module has no path left
  that deletes a trust record in ~/.codex;
- the grant module's own lane acquisition around the clear is gone. It
  was always a pass-through: both callers already hold that file's
  lane (the managed install holds the runtime and system lanes, the
  WSL install holds its config.toml lane) across the whole grant.

No behaviour change. The loop-1 probes still pass as fixed: trust-strip
prints every entry trusted after a failed re-grant, and lane-hold
prints managedInstall=settled projectTrust=settled.

Tests (codex-hook-trust-grant.test.ts):
- "removes equivalent Windows fallback keys before the RPC writes
  canonical trust" now passes the managed caller's step;
- new "runs the caller's pre-session step only when a session runs":
  the step runs once for a session and not on the ledger hit after it.

* chore(codex): comments stop describing a lock held across the session, or a rollback

Review loop 2, L6 comment sweep (comments and one test name only):
- codex-trust-config-concurrent-launch.test.ts: the test named "does not
  let a failing launch roll back a concurrent launch" said the per-file
  lane was the only thing left and that the doomed run's rollback must
  not resurrect the file. There is no lane across a session and no
  rollback now. Retargeted to what it covers: "leaves a concurrent
  grant's records in place when a sibling grant fails" (a restore would
  still turn it red).
- codex-trust-grant-ledger.ts: "a grant session blocks launch prep" is
  true only of inline grants; the background one still costs an
  app-server start. The drift clause no longer says "before the pane
  launches", which is false for the real home.
- agent-trust-write-deadline.ts: a stray hard wrap.
The install.ts:105 comment was fixed with L1. A sweep of src/main/codex,
src/main/startup, src/main/agent-hooks, the trust presets and the CLI
handlers for rollback, restore, rebase, capture/restore, and a lane held
across a grant or session found nothing else stale; the remaining "no
restore" comments state the current rule.

* fix(codex): a real-home resume waits for the one running approval, up to its 30 s limit

Review loop 2, M1; coordinator ruling. A resume into ~/.codex has no
managed home to fall back to. 5a737261d8 let it start at once beside an
Orca entry still awaiting Codex's approval. Codex's TUI then shows a
full-screen hook-review picker before the session and waits for keys:
"Trust all and continue" also trusts the user's own unreviewed hooks,
and "Continue without trusting" leaves that session with no Orca status
for its whole life, because Codex does not reload hooks when Orca's
approval lands later. Panes restored at app start after an update hit
it too, since the start-time conversion leaves every entry awaiting
approval.

The resume now waits, but only while Orca's entry in ~/.codex is
written and a grant is approving it (lane 'granting'). Every resume
waits on that same in-flight grant: ensure never starts a second one
while the lane is 'granting', so panes restored together share one
session. The bound is the grant's own session limit (30 s). The grant
settles only after Codex approved the entry, or after it withdrew its
own unapproved adds, so the resumed session starts either trusted or
with no Orca entry: never beside an unapproved one, and no picker. On
a withdrawal that session has no Orca status, as on main after its
10 s wait. A failed approval never fails the resume.

Tests (codex-launch-during-real-home-grant.test.ts):
- "waits for a warm approval, and spawns with the entries approved";
- "spawns at the approval session limit with Orca entries withdrawn"
  (fake timers: pending at 29.999 s, spawns at 30 s with no Orca entry);
- "makes panes restored together wait on one approval session" (three
  resumes, one session, all settle once it lands).
codex-launch-per-agent-hook-opt-out.test.ts: a resume into ~/.codex
awaits the approval; a resume into a managed account home does not.

* fix(codex): repeated background approval timeouts back off, growing to 5 minutes

Review loop 2, M3; coordinator ruling. A timeout of the ~/.codex
approval starts no cooldown, so the next launch retries at once. On a
host where codex app-server never starts within 30 s, every launch then
wrote Orca's entry into ~/.codex/hooks.json, withdrew it again, and
started another 30 s session, for the rest of the process: an unbounded
retry with no exit.

After 3 timeouts in a row the retry now waits 10 s, then 1 minute, then
5 minutes for every later one. The first two timeouts still retry on the
next launch, so a slow cold start is not punished. Any other outcome
ends the streak (a success, or any other failure, which keeps its own
10 s wait). The streak lives only in memory, so every app start begins
at zero and a slow boot can never latch.

Tests (codex-real-home-slow-app-server.test.ts):
- "backs off after three timeouts in a row, growing to 5 minutes, and a
  success resets it": the first two timeouts retry at once, then 10 s,
  1 min, 5 min, 5 min; after a success, a fresh approval gets two
  immediate retries again and a 10 s backoff after the third;
- "keeps trying after timeouts during a slow first start, once the app
  server answers": three timeouts, then the next attempt at 10 s
  installs.

* refactor(codex): one approval at a time, decided under the config.toml lane

The real-home check kept a lane label, a generation stamp, a promise chain of
ensures and a chain of approvals, and decided from the label at call time.
Concurrent resumes from any state other than 'granting' each started their own
approval (N x 30 s), a chained approval ran a plan an earlier failure had
withdrawn, a hooks-off check during an approval released a waiting resume beside
unapproved entries, and an app-start conversion during an approval was dropped.

Now each check is one step under the real config.toml lane: an approval in
flight answers 'approving' (unusable), hooks off answers 'removed', an open
retry window answers 'unavailable', and otherwise the unchanged install runs and
starts at most one approval. The approval settles under the lane: it withdraws
its own unapproved adds on failure, sets the retry, and derives the verdict from
the settings and the outcome, then runs an owed conversion. A resume waits only
while an approval runs and an unapproved Orca entry is on disk. The opt-out
sweep moves verbatim into its own module.

* fix(codex): only a success or app start resets the approval timeout streak

The ruling is that three timeouts in a row back off, and the count resets on
success and at app start. A non-timeout failure or an unexpected error also
reset it, so a host alternating those with timeouts never backed off.

* fix(codex): a Windows profile path the shells cannot carry bare runs through cmd.exe

The Windows hook command was the bare forward-slash script path, or, for a
profile path that is not one PowerShell word, a PowerShell script. That script
cannot parse under cmd.exe, which Codex uses when a session has no single local
turn shell, so such a profile got no status there.

A path of only letters, digits and _ . : / ~ - stays bare. Any other path,
including one with a space, & ^ $ ` ' ! ( ) or a non-ASCII character, is written
as cmd --% /d /c @"<path>", which ran under PowerShell 7, Windows PowerShell 5.1
and cmd.exe for each of those characters with a real Codex 0.158.0. The choice
depends only on the path, so every build on a machine writes the same bytes. A
machine holding the earlier PowerShell spelling converts it once at app start.

* build(cli): list the real-home hook sweep module in the CLI program

* fix(codex): the Windows cmd spelling names the system cmd.exe and turns off delayed expansion

A profile path the shells cannot carry bare was written as
cmd --% /d /c @"<path>". Under Codex's cmd.exe host the outer cmd.exe resolves
a bare `cmd` from the hook's working directory first, so a repo holding
cmd.bat (or .cmd, .com, .exe) at the session cwd would run on every hook event.
And with delayed expansion turned on in the registry, a `!` in the path was
dropped.

The spelling is now <SystemRoot>/System32/cmd.exe --% /d /v:off /c @"<path>",
unquoted (PowerShell reads a quoted first token as an expression) and with
forward slashes. The Windows directory comes from %SystemRoot% when written,
else from the directory above %ComSpec%'s System32, so both give the same bytes;
if neither is a drive-absolute path it can spell unquoted, it is C:/Windows,
which is still absolute. The bytes stay a pure function of the profile path and
that directory, so every build on a machine writes the same command. Safe
profile paths keep the bare path. Older Orca forms, including the bare-cmd
spelling, convert once; the new spelling is never swept as retired.

* refactor(codex): an approval's settle runs no deferred conversion

An app-start conversion that arrived while an approval ran was remembered and
run by that approval's settle. The settle then rewrote an older entry in place,
unapproved, and started a second approval inside the same wait that releases
every resume, so a resume could start beside an entry Codex would put up for
review.

That path could not happen: the only conversion caller is app start, and it is
the process's first check, so no approval can be running when it arrives. The
deferral and the settle's second check are deleted. A conversion that met an
approval would now be skipped until the next start, and the test for this case
pins that the settle writes nothing new and runs one session.

* fix(codex): an approval's settle keeps a failed opt-out's verdict and ends only its own flight

With hooks read off, an approval's settle always concluded 'removed', which the
routing check treats as usable. If an opt-out during that approval could not
read hooks.json, it had concluded 'unavailable' because the entry may still be
there, and the settle overwrote that. The settle now keeps 'unavailable' when
hooks are off; the next hooks-off check or opt-out re-derives it as before.

The settle's fallback when it cannot run now clears the running approval only
if it is still its own, and the routing check's comment states its rule: never
usable while an approval runs.

* fix(codex): spell the system cmd.exe with backslashes

Under Codex's cmd.exe host the outer cmd.exe hands the typed program text to
the child verbatim, and cmd.exe scans its whole command line for switches, so a
forward-slash C:/Windows/System32/cmd.exe is read as switches: the hook never
runs ("The syntax of the command is incorrect.") and /d is lost. Measured live
on Windows; both PowerShell hosts rewrite argv0 and were unaffected. The script
path after @" keeps forward slashes.

* chore(codex): say why the cmd.exe path is absolute, as measured on Windows

* test(codex): Windows managed-install tests expect the frozen command

They still asserted main's PowerShell text and a backslash bare path; they only
run on Windows, so nothing here caught it. Also correct the /v:off comment: a
lone ! is never dropped, only a !NAME! pair expands.

* ci: run the Codex managed-install tests in the Windows job

Its Windows-only cases skip everywhere else, so nothing ran them; three of them
still asserted a command this branch no longer writes.

* ci: a change to the Codex managed-install tests starts the Windows job

Also say what the missing-script case asserts: a non-zero exit, which
PowerShell reports as 1.

* chore(codex): name the hook trust key pattern for what it matches

* test(codex): the managed-install tests remove folders with the retrying helper

Now that they run in the Windows lane, a raw recursive rm there can throw EPERM
after the assertions pass.

* refactor(codex): one Codex hook-trust key pattern for the trust move and #23958's carry

* test(codex): the trust move carries a block in Codex's quoted spelling and leaves no second table
2026-09-30 14:26:23 -07:00
28d617d0d0 fix: close abandoned PDF.js development asset streams (#23054)
* fix: close abandoned PDF.js development asset streams

* test: keep casting safety comments attached after formatting

* fix(i18n): restore diff note draft catalog entries

* fix: guard closed PDF streams and preserve foreign listeners

* fix(i18n): make AI the recipient of diff notes

* test(browser): report the phase of WebRTC probe timeouts

* test: pin the historical hourly package version fixture

* test: provide the historical package version through build options

---------

Co-authored-by: OrcaWin <alpha-eng@stably.ai>
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 15:08:43 -07:00
OrcaWinandm4air 8416e8de10 refactor(persistence): retire ordinary JSON profile writes (#23202)
* refactor(persistence): retire ordinary JSON profile writes

Require SQLite for writable profiles and keep import, compatibility export, and recovery in a documented legacy-json boundary.

* fix(cli): preserve dynamic profile imports in release output

* test(persistence): exercise SQL races and verify packaged CLI imports

* test(persistence): consolidate shared fixture imports

* test(persistence): close SQLite fixtures before cleanup and await launcher output

* test(automations): use SQLite fixtures for dispatch fencing and skip coalescing

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 12:56:32 -07:00
OrcaWin 82412dab8b Persist profile state in SQLite with background writes (#22612)
Migrate profile state to SQLite and move writes and backups into a background worker. Acknowledge terminal, SSH and automation changes only after durable saves. Preserve JSON import, recovery, rollback and compatibility exports.

Validate migration, worker failures, maintenance, cross-profile moves and terminal lifetime races with unit, integration and end-to-end coverage.
2026-09-25 22:47:33 -07:00
OrcaWin 54a19c2ba3 fix(pdf): render CJK text with pdf.js resources
Merged after PR-specific checks passed. CI failures are unrelated baseline findings in ClientHostedBrowserPagePane.markup.test.tsx and pane-title-update-global-scan-budget.test.tsx.
2026-09-19 01:28:09 -07:00
Jinwoo Hong 631b51f508 perf(usage): run the Claude/Codex/OpenCode usage scans on a worker thread (#21114)
* perf(codex-usage): resume rollout scans at the last parsed byte

Codex rollout files are append-only and grow all day, but any append
changed both mtime and size, so `canReuse` discarded the cached entry and
the scanner re-read the whole file from byte 0 on the Electron main
process. On one real corpus that was 6.59 GB re-read per cycle across
26.63 GB / 21,110 files.

Each parsed file now persists a resume point: the offset just past the
last newline-terminated line, the parse context at that offset (session
id, cwd, model, running totals), a sha256 of the 4 KiB before it, and the
file's dev:ino. A grown file resumes there and merges the appended
rollup into the cached one; anything unproven falls back to a full
reparse — truncation, an in-place rewrite, rotation, a counted tail with
no trailing newline, a legacy copied-session suffix offset, or a file
that must reclaim deferred fork claims. Resume never depends on mtime
equality, so a coarse-mtime filesystem cannot hide an append.

Fixture: a 75,737-byte rollout with a 758-byte append re-read 76,495
bytes before and 8,950 after (the append plus two bounded 4 KiB boundary
windows).

Also bounds the automation-attribution force predicate for both Codex and
Claude: it keyed on `lastScanError`, so a persistently failing scan forced
a fresh full rescan on every single lookup. It now keys on the most recent
scan attempt, which is one forced scan per run regardless of outcome.

* perf(usage): run the Claude/Codex/OpenCode usage scans on a worker thread

The three first-party usage scans walk whole rollout and transcript corpora
and read OpenCode's SQLite synchronously, all on the Electron main process.
They rarely produce a long stall — the JSONL reader streams, so it yields to
the loop between chunks — but they pin the main-process event loop at ~95%
utilization for the scan's whole duration, which is what every IPC message,
timer and window event then queues behind.

Move that work to one lazily-spawned, unref'd worker thread shared by all
three providers, following the OpenCode SQLite scanner precedent (#8864).
Measured on a synthetic 4,000-rollout corpus (25.8 MB cache): a cold scan
drops from 2,147 ms of main-thread time to 31 ms, and a steady-state
incremental scan from 165 ms to 64 ms.

The worker is stateless and the cache crosses the boundary both ways. That
costs ~64 ms of structured clone at this corpus size, against 2,147 ms saved
on the cold path, and it keeps the persisted cache the single source of
truth — a worker-owned copy would need an invalidation protocol and a second
resident copy of the same multi-MB array.

Failure is closed, never a silent empty result: a worker that cannot spawn,
times out, or crash-loops rejects, and the store records the scan error and
keeps the previous projection.

Two clients already carried the same FIFO/timeout/crash-cap machinery, so
extract it once as WorkerThreadRequestQueue (with the packaged entry-path
resolver as worker-thread-entry-path) and move all three onto it, rather
than adding a third copy. Their existing tests pass unchanged.

The oracle is event-loop utilization on the calling thread, not a stopwatch:
usage-scan-worker-event-loop.test.ts runs the same scan both ways and asserts
the worker leg leaves the caller idle while the main-thread leg does not, so
CI load moves both legs together (#18788).

* test(usage): compare the two scan arms instead of two fixed thresholds

The event-loop oracle claimed to be self-calibrating — its header said "the
ratio is self-calibrating, so CI load moves both legs together (#18788)
instead of tipping a fixed millisecond threshold." It computed no ratio. Two
separate `it()` blocks each asserted an absolute threshold against its own
arm, run separately, so load moved them independently. The comment described
a test nobody wrote, and the flake it promised was impossible is the one that
landed: `activeRatio > 0.8` on the calling-thread arm measured 0.764 on an
ubuntu runner.

Fixing the comment is not enough, because the fraction is the wrong quantity.
CPU contention drags the calling-thread arm's active/wall fraction *down*
toward the worker's, since the loop parks waiting on a contended libuv pool.
A 4-vCPU Linux container measured that arm at 0.175-0.756 across twenty runs,
idle and loaded — never once above 0.8. Active *milliseconds* move the other
way: contention stretches the caller's JS time far more than it stretches the
worker arm's fixed post-and-deserialize cost, so the gap widens under load.

Merge the two arms into one case over one corpus and assert the worker arm
costs the caller under a fifth of the inline arm's active milliseconds. Same
twenty Linux runs: 10.9x-83.6x, passing throughout. Keep the presence
preconditions on both arms — an arm that silently scanned nothing satisfies
the comparison trivially — and extend them to the calling-thread arm, which
previously checked only file and session counts.

* fix(ports): name the dropped command when the probe queue is full

The shared-queue extraction turned `Port scan command queue is full; dropped
${command}.` into a constant string, because `describeFull` was given no way
to see the request. Pile-up is per-probe, so the name is the only thing in
that log that identifies which of lsof/ps/netstat was shed.

Pass the rejected request to `describeFull` and restore the name. The request
is built before the cap check so it exists to be named; the id it burns is a
correlation token, so a gap costs nothing.

The existing overflow test asserted only the error class, which is why the
regression escaped a 29-test suite. It now dispatches the overflow under a
different command than the accepted ones and asserts the message text, so a
message that names the wrong request fails too.

Also add a direct WorkerThreadRequestQueue test. Three subsystems share the
queue and each client test only sees the parts its own protocol exercises,
with `queueCap` reachable from port-scan alone. Covers one-at-a-time FIFO
dispatch, the deadline starting at dispatch rather than enqueue, the
consecutive-death cap, and both points where that count clears.

And record the child-process hazard at the usage worker entry. `terminate()`
reaps nothing the thread spawned, and OpenCode discovery reaches a fork
today: `wslGated*` forks the WSL transcript sidecar for a `\\wsl$\...` path,
which a Windows `OPENCODE_DB` or `XDG_DATA_HOME` can be. One scan through
that entry with a UNC `OPENCODE_DB` forked a sidecar that outlived
`terminate()`.

* test(ai-vault): assert the OpenCode worker messages exactly, not by fragment

Checked every message string in the two clients the shared-queue extraction
rewrote against origin/main. Only the port-scan queue-full one regressed
(fixed in the previous commit); the OpenCode SQLite client's four messages
render identically, the remaining source diffs being renames — `error.message`
to `lastError`, `call.timeoutMs` and `CALL_DEADLINE_MS` to `timeoutMs`.
`session-scanner-worker-client.ts` was not touched by the extraction.

But its suite could not have caught it either. `/timed out/`, `/exited with
code/` and a bare `rejects.toThrow()` all still match a message that has lost
its interpolated value, which is the same blind spot that let the port-scan
regression through. Assert the rendered text instead: the timeout names its
deadline, the exit names its code, and the crash-loop drain still carries the
text of the fault that killed the run.

* fix(usage): correct the worker entry's child-process note

The previous note said `worker.terminate()` leaves a forked sidecar orphaned.
It does not, and the reproduction that appeared to show it used a stub sidecar
missing the `process.on('disconnect', () => process.exit(0))` the real entry
has. With a faithful one: the sidecar lives exactly as long as the thread and
is gone within 2s of `terminate()`, because tearing the thread down closes the
IPC channel it owned. Two worker lifecycles forked two sidecars and leaked
neither, and the pre-worker main-thread path reaps its sidecar the same way,
on host exit.

What is true and worth recording: a fork is reachable from this bundle at all,
which is easy to miss; it survives only as long as the channel does; and the
sidecar is now re-forked per worker lifecycle instead of pooled for the app's
life. State those, and warn that a future child which does not exit on channel
close would not get the same free cleanup.

* fix(usage): kill a wedged scan worker on no progress, not on wall clock

`USAGE_SCAN_TIMEOUT_MS` was a 10-minute deadline on the whole scan. A cold
scan of a real history is legitimately minutes — 637 s measured on a 30 GB
corpus with 300 worktrees before the per-cwd memo, ~51 s after — so a
larger corpus or a slower disk crosses it. Crossing it killed the worker,
recorded a scan error and left the cache unadvanced, so the next refresh
started cold and died at the same point, forever.

The deadline is now a no-progress window. The worker posts a file counter
as it walks the corpus (`UsageScanWorkerProgress`, rate-limited to one
message a second), and `WorkerThreadRequestQueue` re-arms the active
call's timer on each one via the new optional `isProgress`. Clients that
do not pass it keep the plain wall-clock deadline. `MAX_CONSECUTIVE_DEATHS`
and idle teardown are unchanged.

* refactor(usage): report scan progress as a file count, not one call per file

Claude's scanner walks batches, so a per-file callback made it loop just
to bump a counter.
2026-09-16 23:03:37 -04:00
Neil 26721bd632 fix(codex): stop blocking the main thread on trust grants (#16441) (#16594)
* fix(codex): stop blocking the main thread on trust grants (#16441)

Codex hook trust was granted by blocking the Electron main thread on
`spawnSync` of a bundled ELECTRON_RUN_AS_NODE entry for the whole
app-server deadline: 15s native, 35s WSL, ~45s on the real-home path
(rebase inspect + repair + grant). Cold start and every Codex pane
launch showed "Not Responding"; the reported event-loop gap was
15,049 ms.

The subprocess only ever existed to donate an event loop to a
deliberately blocked parent — `runCodexHookTrustGrantSession` was
already the real async implementation. Make the callers async and the
fork is unnecessary, so the bridge, the forked entry and its envelope
are deleted along with their build/knip/tsconfig registrations. The CLI
`agent hooks prepare-codex` handler is already async, so it awaits the
in-process session and saves a process spawn per managed-home shell.

`resolveCodexTrustGrantHost` is async too; the WSL identity probe moves
from `execFileSync` to `runProcess`, dropping that file from the
child-process import allowlist. Status reads keep a synchronous
native-only stamp path.

Two invariants that held only because the lane blocked:

- Overlapping capability probes were impossible by construction.
  `GitCapabilityCache`'s dedupe engine is extracted to a shared
  `CapabilityProbeCache` and `CodexAppServerCapabilityCache` now
  inherits it, so concurrent launches against a cold host share one
  app-server session instead of one each.
- Two grants on one `config.toml` could not interleave capture and
  restore. A reentrant per-file lane now serializes the whole install
  sequence (managed, WSL runtime, real-home ensure, legacy sweep) and
  the grant and rebase inside it.

Cold-start work moves off the critical path: retained-home
reconciliation (N sequential sessions) is fire-and-forget behind the
daemon provider, and the startup real-home ensure chains into managed
hook reconciliation instead of blocking app init.

Every preserved semantic is unchanged: never throws, the
ORCA_DISABLE_CODEX_TRUST_RPC kill switch, ledger hits, backfill-pending
and cooldown fallbacks, config rollback on every failure path,
pre-grant self-computed trust removal, the verify-failure taxonomy,
diagnostics and telemetry.

* fix(codex): widen the trust-config lane to every config.toml writer

Review follow-ups on #16441's async trust grant:

- `markCodexProjectTrusted` now runs inside the runtime+system config.toml
  lanes, so a project-trust write can no longer land inside a hook grant's
  capture->restore window and be silently reverted. Its callers await it.
- `install`/`refreshRuntimeUserHooks`/`remove` hold the system config.toml
  lane as well as the runtime one — they promote approvals into
  ~/.codex/config.toml and mirror it back. Lock order is runtime-before-system
  everywhere.
- The real-home ensure chain resumes after a rejection instead of returning
  the same rejected promise to every later pane launch, and resolving the real
  home is now inside the module's never-throws boundary.
- `buildSpawnEnv` awaits inside a cancelable pending-spawn registration, so
  shutdown during the (now long) env build stops the PTY from launching.
  `prepareLocalPtySpawn` generalizes into `awaitCancelableLocalPtySpawn`.
- CapabilityProbeCache drops the test-only `nowMs` passthrough; its probe
  backstop comment now describes what it actually guards.
- Preflight is a plain async function; the trust dispatch in orca-runtime
  collapses into one `markWorkspaceTrustedForAgent`.

* test(codex): exercise the trust-config lane under real concurrency

The async grant makes two pane launches overlap for the first time. These
drive the real modules end to end on real files: a rollback swallowing a
sibling's grant, a markCodexProjectTrusted write landing inside a capture
-> restore window, shared capability-probe dedupe on a cold host, the
host-scoped transient cooldown, and reentrancy from inside an installer.

Each was verified to fail against a deliberately broken implementation
(lane removed, dedupe disabled, cooldown made global, reentrancy pass-
through disabled).

* test(codex): stop hook-service suites spawning the developer's real codex

The forked grant bundle never existed under vitest, so the RPC lane was
unreachable in tests on main. Running it in-process makes these suites
spawn a real `codex app-server` when one is installed: 38 spawns and two
failures in hook-service-runtime-trust-repair on a machine with codex,
green in CI where there is none. Stand in for the missing binary so both
environments exercise the same fallback lane.

* docs(codex): scope the trust-RPC kill switch comment to what it actually gates

The comment read as though the flag forces the fallback lane everywhere. It
gates the managed grant only: the real-home rebase still runs its own
inspect/repair app-server sessions when Orca's insertion shifts a user's hook
positions, and never reads the flag.

Verified by exercise, not by reading — with the flag set, both
inspect-user-hook-trust and repair-user-hook-trust still ran. Pre-existing:
main has no check there either, it just blocked the main thread while doing it.

Widening the flag to cover the rebase is a follow-up; this only stops the
comment promising something the constant does not do.
2026-08-26 16:44:55 -07:00
OrcaWinandOrcaWin 4ec6bbf588 Kill hung WSL transcript filesystem operations via child process with route quarantine (#15381)
* fix(native-chat): kill hung WSL operations via child process

Stalled UNC file operations hold libuv permits even after the gate
timeout expires, blocking Chat tab recovery. Two stalled operations
fill both permits and freeze all WSL access until restart.

Fork file I/O for UNC paths into a separate child process. On deadline
expiry, kill the process to force the hung syscall to exit. This frees
the permit for the affected tab's next read. Temporarily quarantine the
stalled route to avoid retry storms.

* chore: drop internal review artifact from the repo root

* fix(native-chat): harden the WSL transcript fs sidecar

Review follow-ups on the sidecar isolation change:

- Only the deadline may abort running gate work. The sole waiter's
  same-duration timeout fired first, killed healthy children on caller
  abandonment, and settled the task before the deadline could quarantine
  a stalled route - leaving the back-off dead for every dedupe:false op.
- Resolve the fork entry from out/main/chunks too: the resolver compiles
  into a shared chunk, and the scanner service child has no
  process.resourcesPath, so packaged WSL vault scans threw entry-not-found
  (masked as an empty tree).
- Allowlist the fork env instead of spreading process.env; ambient
  NODE_OPTIONS would halt or --require code into every child.
- Wrap transport faults (spawn failure, child death) in
  WslTranscriptFsError('unavailable') so discovery reports them as scan
  issues instead of misreading them as missing paths or empty trees.
- Gate the vitest in-process fallback on the vitest worker global so a
  leaked VITEST=true cannot revert production to in-process UNC syscalls.
- Reap idle sidecar processes after 60s instead of holding them for the
  app session.
- Split 'open' into its own protocol union member so the reusable-call
  Exclude actually strips it from the pooled-process API.
- Guard kill('SIGKILL') against the teardown race where an exiting child
  emits an unlistened 'error', and dispatch reads by handle kind before
  path spelling.

* fix(native-chat): probe stalled WSL routes instead of a fixed quarantine

Remaining review follow-ups:

- Escalating route quarantine: first strike lifts after 5s so a distro
  that was cold-booting when its op hit the deadline recovers on the
  next poll (~35s total instead of ~90s); repeat stalls double the
  back-off toward the prior 2x-timeout cap, and any settle the deadline
  did not force clears the strikes. Queued same-route tasks fail fast
  at quarantine instead of stranding one waiter deadline per file in
  sequential scans.
- Single request implementation: the vitest in-process fallback now runs
  the child's own dispatcher (WslTranscriptFsProcessOperations + decode),
  so unit suites exercise exactly what the forked process executes and
  the per-call-site fallback closures are gone. Dirent fixtures gained
  the full kind-flag set the serializer reads.
- Dropped the production-dead per-route close queue; UNC FileHandles
  (test fallback only) mirror the process-handle close contract.
- Error class, messages, and factories move to wsl-transcript-fs-error
  (re-exported from the gate) to keep the gate under the lines budget.

* fix(native-chat): harden WSL transcript fs with route quarantine strike

Extract quarantine logic into a dedicated module with strike decay: stalls older
than 5 minutes restart from base back-off, and concurrent-lane timeouts count as
one incident. Allow joining live in-flight tasks on quarantined routes (they cost
no new I/O). Preserve quarantine across transport faults (child death). Handle
file shrinking during tail reads by detecting short reads and returning empty.
Defer file closes that arrive mid-read instead of refusing, preventing slot
leaks. Separate process slot and boundary-finding concerns into focused modules.

* fix(native-chat): enforce route quarantine windows and isolate lanes per

A late result arriving after the deadline was incorrectly lifting the route
quarantine, allowing subsequent work to start before the back-off period
expired. Now late results are correctly recognized as stale and never cut
the quarantine short.

Process work is now isolated per (route, priority) lane so a scan stall
cannot block exact reads on the same distro. Each lane gets its own client
and process pool; late results and handle faults stay scoped to their lane.

Tests now fake performance.now() alongside timers (the quarantine clock
depends on it) and wait for the full back-off window to expire rather than
advancing by 0. Gate state is reset between test cases since late releases
never lift the quarantine.

---------

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-08-18 16:57:35 -07:00
Jinjing 4c69e45552 Strengthen plain-node-entry-guard with entry name validation (#12761)
* Strengthen plain-node-entry-guard with entry name validation

- Add buildStart hook to validate guarded entry names exist in rollup
  inputs, preventing stale names from silently stopping guards
- Extend electron require detection to subpaths (electron/main, etc)
- Improve smoke test signal/exit handling and use constants
- Add comprehensive tests for entry validation and new behaviors

* Add SIGKILL escalation to plain-node-entry-guard timeout

Switch from spawnSync to async spawn to properly handle daemons that trap
SIGTERM. spawnSync's timeout only sends the signal and waits, so a daemon
that ignores SIGTERM causes the build to hang. The new runDaemonEntry
function escalates to SIGKILL after a grace period to enforce the deadline.
Configurable timeouts and grace periods via SmokeTimings type; closeBundle
hook becomes async to support the change.
2026-08-11 12:20:20 -07:00
Brennan Benson 5df2ddbc9c perf(ai-vault): isolate tab title resolution (#13377)
* perf(ai-vault): isolate tab title resolution

* fix(ai-vault): preserve background scan caches

* fix(ai-vault): resolve nested worker from chunks
2026-08-09 16:45:52 -07:00
Jinjing fde816e4ee move folders (#12758) 2026-08-05 12:09:24 -07:00