mirror of
https://github.com/stablyai/orca.git
synced 2026-10-04 00:02:21 +00:00
cd8d03bc06398b856f6db734aef11df4e2afe15a
12
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0a81211994 |
fix(cursor): read desktop login outside the main thread (#24572)
* fix(cursor): read desktop login outside the main thread * fix(cursor): await native worker retirement before respawning |
||
|
|
cfa43e7eab |
fix(codex): opening a terminal no longer strips Codex hooks from the real ~/.codex (#23552)
* fix(codex): a real-home restore leaves a file alone once someone else changed it Orca writes ~/.codex/hooks.json (and a trust rebase writes config.toml), then runs a Codex trust session for up to 10 s, then restores the original bytes if the session fails. The restore wrote unconditionally, so a save that landed during the session, from the user or another Orca, was silently reverted. Each restore now compares first: it writes the original back only while the file still holds the generation Orca's mutation left, and otherwise logs and leaves it alone. This covers the real-home install and opt-out sweep (restoreRealHomeHooksJson), the legacy sweep's hooks restore, and config.toml rollback (restoreCodexTrustConfig). For hooks.json the generation is the exact bytes Orca wrote. For a config.toml that a trust rebase changed it is the file as the rebase left it. When Codex itself wrote config.toml inside the session that just failed, Orca never knew those bytes, so that rollback compares against the file as the session settled. The next commit keeps other Orca instances out of that window; a user edit made during such a session can still be rolled back. * fix(codex): serialize real-home Codex writes across Orca instances Every Orca on one HOME (a dev and a packaged app, or an offline CLI) writes the same ~/.codex/hooks.json, config.toml and ~/.orca/agent-hooks/codex-hook.sh. The per-file lane that orders capture, mutate and restore was in-process only, so another instance could write inside this one's restore window, or undo it. The lane for the user's real config.toml now also holds the existing crash-safe managed-hook install lock (~/.orca/managed-hook-install.lock, the one relay installers take for the same home). It is taken only by the outermost acquire, because the lock file is not reentrant and grants and trust rebases nest inside an install. Managed-home installs, the real-home install and opt-out sweep, and the legacy sweep all enter through it. Compare-and-swap on restore stays as the backstop. A lock that cannot be taken within its 10 s wait fails that install, which is already best effort: launch prep logs it, and the real-home lane falls back to the managed lane until its retry. * fix(codex): opening a terminal no longer strips the shared Codex entry from ~/.codex Every Orca instance on one HOME writes the same status-hook entry into the user's ~/.codex/hooks.json, with its trust in config.toml. Launch prep runs on every pane spawn, and under a managed Codex account it ran the legacy system sweep. That sweep matched Orca entries by script file name, so it removed the current shared entry and the trust blocks the grant ledger recorded. On a live laptop hooks.json went 4139 -> 18 bytes about 150 ms before a new pane opened. With hooks off, the real-home lane's launch prep swept the same way. Now nothing automatic removes the current entry or its trust: - The legacy sweep removes only an enumerated list of retired command forms that no build writes any more (#1019's double-quoted form, #1536's exec-guarded form, and Windows' per-userData bare path), plus their trust. - ensureRealHomeCodexHookState with hooks off writes nothing; that covers launch prep, session resume and startup. - Only the user's explicit opt-out (codexHookService.remove()) strips the entry and its ledger-recorded trust from the real home. - The sweep-suppression gate existed only to stop the sweep from deleting the current entry, so it is deleted with its main-process wiring. Startup with hooks off already skipped the real-home install; with this change the first pane's launch prep with hooks off also leaves ~/.codex untouched. * fix(codex): a pane's prepare-codex only repairs a home its own HOME's app installed On macOS a pane starts through login(1), so it gets the user's real HOME even when its Orca app runs with another one. The pane's `codex()` preflight installed hooks in the CLI process with that real HOME: it rewrote ~/.orca/agent-hooks/codex-hook.sh, promoted trust into the real config.toml, and wrote the real HOME's script path into the app's managed home. The preflight now acts only when the managed home's hooks already run this process's own shared script, which proves the app that installed them shares its HOME. Otherwise it writes nothing; the app installed the home at spawn. Why not a no-op: the preflight was added (#14326) because trust can go stale between opening a pane and typing `codex`, for example in a pane that survives an app update, and Codex then stops in hook review. For a same-HOME pane it still repairs that. Why keep promotion: the install drops runtime trust the system config does not back, so skipping promotion would delete approvals the user gave inside Orca-launched Codex. * test(agent-hooks): await every installer in the refresher coverage test The test fired each managed installer without awaiting it and read ~/.orca/agent-hooks straight after. Codex's install now takes the cross-process real-home lock before it writes its script, so the script landed after the read. Await the installers, and stub Codex's trust sessions so the awaited install cannot start a real `codex app-server`. * fix(codex): retire the two real-home command forms the list missed The real-home lane wrote two Codex hook forms into ~/.codex that no build writes any more and that the enumerated retired list did not name: - POSIX, #9501 until #10885: the file-guarded form draining with a bare `cat`. - Windows, #9501 until #10221 took Windows off the real-home lane: the encoded PowerShell launcher for a non-cmd-safe script path. The file-name sweep removed both before; the enumerated sweep left them in place, trusted, still passing the script's exit status to Codex. Both now match as frozen literals. Also corrects the startup ordering comment: the real-home install runs first so its in-slot upgrade lands before the managed install's sweep retires the prior command; nothing re-arms a legacy sweep any more. * fix(codex): take the real-home lock only when a write is needed The previous commit made every entry to the real-home config lane take the cross-process lock. That lane runs on every pane spawn and every typed `codex` preflight, so the steady state paid an owner probe (a `ps` spawn on macOS) and could wait up to 10 s behind another instance's trust session, even though it wrote nothing. Each real-home writer now compares the desired state with the files on disk first, without the lock. Only when a write is needed does it take the lock, re-read and recheck, then write: - real-home install: the planned hooks.json, the shared script and the ledger-recorded grant are compared; the locked path re-plans from disk. - legacy sweep: locks only when a retired entry is present; the sweep re-reads. - approval promotion: locks only when there is something to promote; the promotions are recomputed under the lock. - the shared ~/.orca/agent-hooks script: locks only when its bytes differ. The explicit opt-out always takes the lock. The lock is reentrant through async context, since grants and rebases nest inside an install, so the config-lane option the previous commit added is removed. * fix(codex): a shared script without its exec bit is not the steady state The compare-first check matched the shared ~/.orca/agent-hooks script on bytes alone. writeManagedScript also restores 0755 on every call, and the POSIX hook guard skips a script that is not executable, so a script whose mode was lost (a dotfiles restore, a plain copy) now stayed that way: every Codex hook drained stdin and reported nothing until an app restart refreshed the script. The check now also requires the mode the writer sets, so that case takes the lock and the write path repairs it. * test(codex): the retired encoded launcher never matches today's shared one The shared encoded Windows launcher is still current for other agents, so the comment claiming today's launcher is never encoded was wrong. What keeps the retired matcher off it is the exact payload: since #14825 the shared launcher prefixes its payload and drops -ExecutionPolicy Bypass. Pin that with a case. * fix(codex): the pane step recognises its own script under a home path with an apostrophe The same-HOME check looked for the script path wrapped in bare single quotes, but both hook writers escape an apostrophe inside the quotes. A home such as C:\Users\O'Brien never matched, so the pane-step repair never ran there. * fix(codex): the trust-RPC escape hatch still keeps the real home off its lane The no-write check reported a recorded grant as current, so with ORCA_DISABLE_CODEX_TRUST_RPC set the real-home lane stayed in use. The grant itself refuses before reading its ledger; the check now does the same. * fix(codex): the shared script write no longer waits on the real-home lock The write is atomic and skips identical bytes; waiting behind another instance's trust session could only fail a pane's managed-home install. * fix(codex): an in-Orca approval survives a launch that cannot get the real-home lock The install drops runtime trust the system config does not back, so a promotion skipped for want of the lock lost the approval for good. It now writes unlocked, as it did before the lock existed. * refactor(codex): take the cross-process real-home lock back out The lock fixed no observed failure. The three that were observed each have their own fix in this series: the legacy sweep matches only frozen retired command forms, hooks-off launch prep writes nothing, and a pane's prepare-codex repairs only a home its own HOME's app installed. The lock instead brought its own defects: a steady-state spawn waiting behind another instance's trust session, a compare-first split to avoid that, a script write and an approval promotion that could fail for want of the lock. Removed, with their tests: the real-home write lock and its async-context reentrancy, the plan/compare split that kept it off steady-state spawns, the compare-first legacy sweep, the locked approval promotion and its unlocked fallback, the compare-first shared script write (writeManagedScript already skips identical bytes and restores the exec bit), and the CLI tsconfig entries the lock pulled in. Kept: the retired-forms matcher, the hooks-off no-op, removal only on an explicit opt-out, the pane own-script check, and the compare-and-swap rollbacks. Every instance now writes identical bytes idempotently. * fix(codex): an opt-out that cannot read hooks.json keeps Orca's trust and ledger The opt-out swept the real-home entry, then dropped Orca's ledger-proven trust whenever a ledger existed, even when the sweep could not read hooks.json. The entry could still be there, now untrusted, and the ledger that proves ownership was gone for the retry. Drop that trust only after a sweep that read the file. * refactor(agent-hooks): one predicate for whether an agent's status hooks are on "Global switch on and this agent not turned off" was spelled out separately in the startup controls, the settings reconcile, the retained-home reconcile, the WSL preflight RPC, the CLI preflight and the OpenCode plugin selection. They now share one function, in a module light enough for the CLI's per-launch Codex preflight to load. The PTY spawn env derives the Codex flag from the switch and opt-out list it already carries, the same way it does for OpenCode and Pi, instead of receiving a second copy. * fix(codex): launch and resume prep honour Codex's per-agent hook opt-out Turning Codex off in the per-agent hook settings removes Orca's Codex hook entry, but launch prep and session resume read only the global hooks switch, so the next Codex launch or resume wrote the entry straight back into the real ~/.codex or the account's home. Both now read the per-agent predicate, which the PTY spawn env and startup already honoured. * fix(codex): turning Codex off per agent clears the real ~/.codex entry While the real-home lane owns ~/.codex/hooks.json, the legacy system-home sweep stands down. That gate read only the global switch, so turning Codex off per agent ran remove() with the sweep still suppressed and left Orca's entry in the real ~/.codex. The gate now reads the per-agent predicate, the same as turning every hook off. * test(codex): cover the system ~/.codex sweep gate for Codex turned off The gate that lets the legacy system-home sweep run was an inline closure in startup, so reverting it to the global switch left CI green. It is now a pure function beside the gate it feeds, with a table test and a remove() test on a seeded ~/.codex: turning Codex off strips Orca's entry and keeps user hooks; with Codex on the entry stays. * fix(cli): keep the agent-status hooks predicate loadable by the packaged CLI The CLI's prepare-codex handler imported the predicate from src/main, but the Electron build rebuilds out/main from its declared entries only, so the packaged `orca agent hooks` commands could not load it (package jobs and the CLI bundle-parity test were red). The predicate reads only settings, so it now lives in src/shared, which the CLI compiles itself. * feat(codex): every Orca build writes one frozen Codex hook command The Codex hook command was built from this build's wrapper, so two builds on one HOME disagreed about the bytes of the shared ~/.codex entry and kept rewriting it, with a Codex trust session each time. The command is now fixed per form and carries its form number: - POSIX: one command with no path in it. It runs the shared script only in an Orca pane with hooks on (pane key and hook port set), drains stdin everywhere else, and always exits 0. A branch for a per-build script root is written now and stays dormant until Orca sets ORCA_AGENT_HOOK_ROOT, so that change will not move these bytes. - Windows: the bare forward-slash path to the shared .cmd, which runs under PowerShell 7 and 5.1, Codex's hook hosts. A profile path that is not one PowerShell token gets a plain PowerShell form with the same branches. The literals live in the form module, so a change to the shared hook constants cannot move them; goldens pin the bytes. Every form keeps `agent-hooks/codex-hook.*` in plain text, so older builds still recognize it. * fix(codex): one main-process owner adds the real-home entry; nothing restores files Each Orca writer of ~/.codex decided what Orca's entry must be from its own build and instance, then removed or reverted whatever differed: launch prep rewrote any Orca-shaped entry to this build's command and stripped Orca entries from events this build does not use, and a failed trust session restored hooks.json and config.toml from snapshots. With several instances and builds on one HOME, every disagreement became a deletion or a revert. The main process is now the one writer, and its writes are add-only: - A launch or resume adds Orca's frozen entry to an event that has none and leaves every Orca entry it finds, so a running older build is never fought. - App start also converts an older Orca form to the frozen command, once, in its own slot: one hooks.json write (one .bak) and one trust grant per home. - A newer form is never rewritten or appended beside, and Orca entries in events this build does not use are kept. - After a failed trust grant, only an entry this call wrote that is still untrusted is withdrawn, putting back the handler it replaced. Both files are re-read, so a concurrent edit, or the identical entry another Orca trusted meanwhile, survives. Deleted: the compare-and-swap hooks.json restore, the config.toml snapshot restore after a grant session and after a user-trust re-key, and the rollback module. A grant session writes trust only at Orca's own keys, and every caller settles those keys itself. A failed re-key of moved user hooks now keeps the write and reports it; Codex lists those hooks for review. * fix(codex): the pane CLI asks the app to prepare its Codex home `orca agent hooks prepare-codex` ran Codex's install inside the pane. That process can have the real HOME (login(1)) and runs outside the app's in-process queues, so it was a second writer of ~/.codex and ~/.orca beside the app. A check that the home ran "its own script" guarded it. The pane step now only asks the app, over the same kind of local RPC the WSL pane step already uses (agentHooks.prepareCodexForPane). The app checks that the pane's CODEX_HOME is one its own userData owns, reads its own hooks setting, and installs on its own queue. An app that is not running, or is too old to know the method, makes the step a no-op, as it is on WSL. The own-script check and the CLI's settings read are gone, and the preflight module leaves the CLI bundle. * fix(codex): delete the pane step on native hosts The previous commit had `orca agent hooks prepare-codex` ask the app to prepare the pane's Codex home. The case it existed for (#14326, a pane that survives an app update with stale hook trust) did not reproduce, and no other desktop agent host writes agent config from a terminal or launch wrapper. - Deleted: the agentHooks.prepareCodexForPane RPC method, its params and catalog entry, and prepareManagedCodexHomeBeforeShellLaunch with its module, tests and CLI build entry. - `agent hooks prepare-codex` is a no-op on native hosts. It stays for one release so shell wrappers from older builds, which still call it, exit 0. - WSL panes are unchanged: they still ask the app over agentHooks.prepareCodexForWslPane. The shell wrappers and ORCA_CODEX_LAUNCH_PREFLIGHT stay, because WSL panes use the same wrappers and variable (forwarded through WSLENV). A native pane still starts the CLI once per `codex` it runs; skipping that is a follow-up. * test(codex): a failed trust session keeps concurrent edits to both files QA case 9 at host level, on a real file system in a temp HOME: Codex's trust session fails after another writer saved hooks.json and config.toml. - Both saves survive, and no Orca entry is left that Codex would list for review: this call's entry is withdrawn. - A failed one-time conversion puts the older Orca entry back in its slot and keeps both saves. Both tests fail on the previous head, which restored config.toml from a snapshot and left the untrusted entries in hooks.json. Removing the withdrawal turns both red. * feat(codex): read whether an Orca entry's stored trust is still current A Codex release that changes how it hashes a hook leaves Orca's stored trust stale: the entry is present, but Codex lists it as modified. Checking only whether the entry is missing cannot see that. readOrcaEntryTrust sorts a present entry into four states: - trusted: the stored hash is the current one; - untrusted: there is no stored hash; - stale: the stored hash is not the current one; - disabled: the user turned the entry off. The caller can pass Codex's current hash, for example one a grant recorded. The failed-grant withdrawal now uses it, and also keeps an entry the user turned off. Nothing re-grants on 'stale' yet. * fix(codex): a slow Codex start retries on the next launch, never for minutes On a loaded Mac a cold `codex app-server` took over 10 s (QA case 4). The grant timed out, the entry was withdrawn, and a 5-minute cooldown in both the grant and the real-home install then refused every retry. - The native session deadline is 30 s, the same as WSL's. - A timeout starts no cooldown in the grant or in the real-home install. The next launch retries. Other failures keep their cooldown. - Launches that queue behind a slow session share one follow-up run, so a launch waits for at most two sessions, not one per earlier launch. Tests: a 15 s cold start still grants and keeps the entry; after a timeout, the next launch runs a session at once; four queued launches run two sessions. Each is red on the previous head, and each mechanism was removed in turn to confirm its test turns red. * fix(codex): Orca's automatic writes never move a user hook Codex keys a hook's trust by its position in hooks.json. App start's collapse of Orca duplicates removed every Orca entry and appended one at the end. That moved any user hook that followed a removed entry, so the write waited on a session to re-key the moved hook's trust. App start now: - converts the first Orca entry that sits in a plain slot to the frozen command, in place; - drops any other Orca entry only when that moves no user hook; - keeps a duplicate that a user hook follows, and trusts every frozen copy, so none is listed for review; - appends only when no frozen entry is left. Tests check user positions and user trust blocks byte-for-byte for each automatic write: add-missing (append), the one-time conversion (in place), a trailing duplicate, a duplicate before a user hook, and older duplicates normalized to one entry. The three collapse cases fail on the previous head. Removing the position check, or the in-place conversion, turns its tests red. Only the explicit opt-out still removes an entry that user hooks follow. * fix(codex): removing an Orca entry never waits on a Codex session Removing an Orca entry from ~/.codex/hooks.json moves every user hook behind it up a slot, and Codex keys trust by slot. The retired-form sweep, the opt-out and a failed-grant withdrawal all asked a `codex app-server` session to list the old trust before writing, and to re-key it afterwards. A timeout there threw before the write and latched a 5-minute cooldown, so a slow cold start blocked the retired-form sweep at boot (QA case 4). Each moved hook's [hooks.state] block now moves to its new key, body bytes unchanged, straight after the hooks.json write. Codex hashes a hook's content, not its position or its file path, so the moved block stays exactly as valid as it was: a trusted hook stays trusted, an untrusted one stays untrusted, and one the user turned off stays off. No removal waits on or depends on a session. A failed config.toml write keeps the hooks write and logs. Deleted: the inspect and repair sessions, their client, and their cooldown. The generation guards on the hooks.json writes stay, for other processes. Tests: the retired sweep removes the retired entry and carries the trust of the user hook behind it while every Codex session times out (red on the previous head); the opt-out carries an appended user hook's trust; the move carries trusted, disabled and untrusted states byte for byte. Removing the move turns all of them red. * fix(codex): a Codex launch never waits on Codex's approval of Orca's entry A launch on the real-home lane awaited Codex's trust grant for the entry it had just added. A cold `codex app-server` on a loaded Mac took over 10 s, so the launch could wait that long, and a failure then latched a 5-minute cooldown. - Codex's approval runs in the background, with a 30 s cold-start budget. - A launch uses the real home only when the ledger shows trust is already current. Otherwise it goes to the managed home at once, and the next launch picks up the finished grant. - A launch that arrives while a grant runs does no work and does not queue behind it. - A resume into the real home has no managed home to fall back to. It waits for the grant, but no longer than the 10 s a launch always could. - A background grant that times out starts no cooldown; the next launch retries. Any other failure backs off for 10 s instead of 5 minutes. Success is what the ledger remembers. - A failed grant still withdraws only what that install added and is still unapproved. The log now says how many entries it took back and when the next try comes. Managed-home grants keep their 10 s deadline and stay on launch prep, as before; they fall back to Orca-computed trust. Tests: - A 15 s start: the launch returns in under a second on the managed home, a second launch starts no session, the grant lands in the background, and the next launch uses the real home. - A timeout sets no cooldown, withdraws its adds and logs it. - Another failure retries after 10 s, not before. - A resume waits only as long as allowed. - Case 9 checks the log line and the retry. Making the launch await the grant, a 10 s budget, either timeout cooldown, and a 5-minute backoff were each tried, and each turns its test red. * fix(codex): move a hook's trust only when every stored key has the known shape Orca now edits Codex's trust store directly when a removal moves a user hook. Three safeguards keep that honest: - Fail safe. If any [hooks.state] key in config.toml does not have the shape `<path>:<event>:<group>:<handler>`, nothing moves and Codex asks the user to review. That shape was checked unchanged from Codex 0.141 to 0.158. - Targeted. The file is read immediately before the atomic rename, and only the moved keys' blocks change. Every other byte stays, and no snapshot is restored. - Verbatim. Each block's body moves as Codex wrote it, including fields Orca does not know. No hash is ever computed, and a hook with no block gets none. Tests: - An unknown key shape stops every move. - Everything except the moved block survives byte for byte, and the moved body keeps an unknown field. - In case 9, a hook the user approved during the failed session keeps its approval when the withdrawal moves it, beside the concurrent project edit. Removing the shape check, or writing a computed block instead of the stored body, turns these tests red. * refactor(codex): keep only the trust read the failed-grant withdrawal uses A capture across Codex 0.141, 0.150 and 0.158, switching in all six directions, showed Orca's entry keeps the same hash and stays trusted. A Codex upgrade does not make its trust stale, so nothing needs to re-grant on staleness. readOrcaEntryTrust keeps the four states the withdrawal needs, but loses the parameter that let a caller pass a different current hash, and the test for a Codex that hashes differently. * fix(codex): native panes no longer start the Orca CLI before each codex The pane step is a no-op on native hosts, but native panes still carried ORCA_CODEX_LAUNCH_PREFLIGHT, so every `codex` typed in a pane started the Orca CLI for nothing. Only a packaged Windows build's WSL pane now gets the variable; the app prepares every native Codex home itself. The resolver loses the dev-launcher path and its userDataPath option, which only native panes used. Tests: a native macOS, Linux and Windows pane gets no preflight, packaged or not, even with the bundled CLI present; a WSL pane still gets the verified absolute launcher. Letting native panes through again turns them red. * chore(cli): say when the native prepare-codex no-op can go Native pane wrappers from builds up to v1.4.216 still call it. It can be deleted once no supported build's wrapper does. * test(codex): check the WSL launcher path instead of asserting it * fix(codex): a launch no longer waits behind the background real-home approval The background grant ran its whole codex app-server session inside the shared ~/.codex/config.toml lane, and on a cold host its session was also the shared capability probe. A launch sent to the managed home then waited on both: the managed install and the project-trust write queue on that lane, and the managed install's own grant waited for the probe. On a cold app-server that was up to 30 s per launch. The lane was held across the session only to protect the retired capture-and-restore. Codex writes its own records, so the lane is now taken only around Orca's own pre-grant write. The background grant runs its session without publishing it as the shared probe, and the whole grant is bounded by its deadline, so a hang outside the session cannot leave the lane 'granting'. * fix(codex): a failed re-grant no longer strips Codex's own approval of Orca's entries Before each trust session, the grant deleted every Orca record whose hash matched the one Orca computes. That exists because a managed home's fallback writes Orca-computed trust under both Windows path-separator spellings, and Codex rewrites only its own spelling, so the other copy would linger. On failure the managed and WSL fallbacks write that trust back, and before this fold a snapshot restore covered it. The real ~/.codex has neither: Orca never writes computed trust there (the real-home lane does not run on Windows at all), so a matching record there is Codex's own approval. After a ledger miss (another Orca profile, a Codex update, a lost ledger) and a failed session, nothing put it back, and every Orca entry showed "Hooks need review". The clear now runs only for homes whose fallback writes that trust. * fix(codex): a real-home resume spawns only once Orca's entry is approved or withdrawn A resume that must run in ~/.codex waited at most 10 s for the background approval, then spawned anyway. On a cold app-server that left Codex beside an unapproved Orca entry, so the resumed pane showed hook review. The resume now waits for the grant to settle. Settled means Codex approved the entry, or the grant failed and withdrew its own unapproved write; the grant's deadline bounds the wait (30 s, the cold-start budget), and a failed approval never fails the resume. Why this over the alternatives: - Spawning at 10 s keeps the review prompt this fold exists to remove. - Withdrawing at 10 s from the resume races the still-running session: Codex can write the frozen entry's hash after the withdrawal, and for a converted entry that marks the older command Orca put back as modified. - A resume cannot use the managed home: the session lives in ~/.codex. So the only states that cannot race Codex are the grant's own settle. The cost is a longer worst case on a cold app-server (up to the 30 s deadline, plus any managed-home install that holds the config.toml lane); a warm approval takes seconds, and an approved entry costs no wait. * fix(codex): keep the 5-minute trust cooldown for launch-path grants The fold shortened the host's trust-grant cooldown from 5 minutes to 10 seconds for every grant. That was meant for the background ~/.codex approval, which blocks no launch. The managed-home and WSL grants run inline on the launch path, so with a hung app-server every launch more than 10 s after the last failure paid the full inline timeout again (10 s native, 30 s WSL). Cooldowns are now kept per lane: inline grants keep 5 minutes, the background grant retries after 10 s, and neither lane's failure cools the other down. A success, or a proven-missing surface, still clears both. The real-home install's own retries (an unreadable hooks.json, unknown keys) are back on the 5-minute interval they had before the fold. The cooldown moves to its own module so the grant stays within the file limit. * fix(codex): a failed grant withdraws the exact copy it wrote The withdrawal re-found "this call's" entry by command, taking the first frozen handler in the event. When app start converted a later slot while an earlier frozen copy sat in a matcher group (which conversion skips), a failed grant acted on that earlier copy: it put the older command into it, or skipped it, and left the converted, unapproved copy in place. Each write now records where its handler landed, after any duplicate drops, and the withdrawal acts only on that slot. A copy that has since moved is left alone; the next launch's grant retries it. * fix(codex): the failed-grant withdrawal checks hooks.json is unchanged before writing The install and the retired-form sweep both refuse to replace ~/.codex/hooks.json if it changed since they read it. The withdrawal did not: a save landing between its read and its atomic replace was lost. The window is small, since the withdrawal is synchronous, but it now carries the same guard. * refactor(codex): drop rationale left over from the snapshot restore; name the trust-move module for what it does Comments on the config.toml lanes still justified them by a grant's capture-and-restore window, which the fold deleted, and the trust-write deadline still counted a grant session holding the lane. They now give the reason that remains: Orca's own multi-step reads and writes, and managed-home installs that hold the lane across their inline grant. codex-user-hook-trust-rebase no longer rebases through Codex; it moves stored trust records, so it is now codex-user-hook-trust-moves. The grant test that pinned two sessions on one config.toml to run one at a time is removed: its reason was an interleaved capture and restore. Callers that write config.toml around a grant hold their own lane, which the nested installer test still covers. * build(cli): list the trust-grant cooldown module in the CLI program The CLI's agent-hooks handler loads the hook controls, which reach the Codex trust grant; the CLI project is composite, so every module in that graph must be listed. * docs(codex): say which Windows hosts each hook command form runs under Codex runs a hook under the turn's shell (PowerShell 7 or 5.1 in every captured session) and, with no single local turn shell, under %COMSPEC% /C. The bare forward-slash path ran under all three in the Windows host census. The PowerShell form used for a profile path with a space does not parse under cmd.exe; no form valid in all three hosts has been run for such a path, so the form stays and the gap is stated here and in the PR. * test(codex): type the withdrawal seam without an assertion * fix(codex): a real-home resume starts at once, trusting Orca's entries for that process A resume that must run in ~/.codex waited for Codex's background approval of Orca's newly written hook entry: up to 30-40 s on a cold app-server. That made the user's resume wait on bookkeeping, and the alternatives (start at 10 s with Codex's hook review showing, or withdraw the entry and race Codex's own write) were worse. Codex reads hook trust from its session-flag config layer as well as the user's config.toml, merged per key, and has since hook trust shipped. So the resume no longer waits. When Orca's own frozen entries in ~/.codex are untrusted (or hold a stale hash), the resume command carries `-c hooks.state={'<key>'={trusted_hash='<hash>'},...}` for exactly those entries: the key under both the logical and the real path of ~/.codex (Codex keys an explicit CODEX_HOME by its real path), and the hash of that entry's content, so it can trust nothing else at that slot. The user's hooks are never included, nothing is written, and the background approval still runs for later plain `codex` launches. An approved entry adds nothing; a Codex known to lack hook trust gets nothing. One inline table, because Codex splits a `-c` key on every `.` and the key holds `.codex/hooks.json`. TOML literal strings keep `"` out of Windows native-argument quoting. The flag goes before `resume <id>`, quoted for the pane's shell (portable Unix, PowerShell or cmd), in the launch command and in the setup-sequenced copy of it; a cmd line whose path cmd would expand, or a key with an apostrophe, is left unchanged. SSH and WSL resumes get no preparation, so no local path reaches them. * Revert "fix(codex): a real-home resume starts at once, trusting Orca's entries for that process" This reverts commit |
||
|
|
28d617d0d0 |
fix: close abandoned PDF.js development asset streams (#23054)
* fix: close abandoned PDF.js development asset streams * test: keep casting safety comments attached after formatting * fix(i18n): restore diff note draft catalog entries * fix: guard closed PDF streams and preserve foreign listeners * fix(i18n): make AI the recipient of diff notes * test(browser): report the phase of WebRTC probe timeouts * test: pin the historical hourly package version fixture * test: provide the historical package version through build options --------- Co-authored-by: OrcaWin <alpha-eng@stably.ai> Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
8416e8de10 |
refactor(persistence): retire ordinary JSON profile writes (#23202)
* refactor(persistence): retire ordinary JSON profile writes Require SQLite for writable profiles and keep import, compatibility export, and recovery in a documented legacy-json boundary. * fix(cli): preserve dynamic profile imports in release output * test(persistence): exercise SQL races and verify packaged CLI imports * test(persistence): consolidate shared fixture imports * test(persistence): close SQLite fixtures before cleanup and await launcher output * test(automations): use SQLite fixtures for dispatch fencing and skip coalescing --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
82412dab8b |
Persist profile state in SQLite with background writes (#22612)
Migrate profile state to SQLite and move writes and backups into a background worker. Acknowledge terminal, SSH and automation changes only after durable saves. Preserve JSON import, recovery, rollback and compatibility exports. Validate migration, worker failures, maintenance, cross-profile moves and terminal lifetime races with unit, integration and end-to-end coverage. |
||
|
|
54a19c2ba3 |
fix(pdf): render CJK text with pdf.js resources
Merged after PR-specific checks passed. CI failures are unrelated baseline findings in ClientHostedBrowserPagePane.markup.test.tsx and pane-title-update-global-scan-budget.test.tsx. |
||
|
|
631b51f508 |
perf(usage): run the Claude/Codex/OpenCode usage scans on a worker thread (#21114)
* perf(codex-usage): resume rollout scans at the last parsed byte Codex rollout files are append-only and grow all day, but any append changed both mtime and size, so `canReuse` discarded the cached entry and the scanner re-read the whole file from byte 0 on the Electron main process. On one real corpus that was 6.59 GB re-read per cycle across 26.63 GB / 21,110 files. Each parsed file now persists a resume point: the offset just past the last newline-terminated line, the parse context at that offset (session id, cwd, model, running totals), a sha256 of the 4 KiB before it, and the file's dev:ino. A grown file resumes there and merges the appended rollup into the cached one; anything unproven falls back to a full reparse — truncation, an in-place rewrite, rotation, a counted tail with no trailing newline, a legacy copied-session suffix offset, or a file that must reclaim deferred fork claims. Resume never depends on mtime equality, so a coarse-mtime filesystem cannot hide an append. Fixture: a 75,737-byte rollout with a 758-byte append re-read 76,495 bytes before and 8,950 after (the append plus two bounded 4 KiB boundary windows). Also bounds the automation-attribution force predicate for both Codex and Claude: it keyed on `lastScanError`, so a persistently failing scan forced a fresh full rescan on every single lookup. It now keys on the most recent scan attempt, which is one forced scan per run regardless of outcome. * perf(usage): run the Claude/Codex/OpenCode usage scans on a worker thread The three first-party usage scans walk whole rollout and transcript corpora and read OpenCode's SQLite synchronously, all on the Electron main process. They rarely produce a long stall — the JSONL reader streams, so it yields to the loop between chunks — but they pin the main-process event loop at ~95% utilization for the scan's whole duration, which is what every IPC message, timer and window event then queues behind. Move that work to one lazily-spawned, unref'd worker thread shared by all three providers, following the OpenCode SQLite scanner precedent (#8864). Measured on a synthetic 4,000-rollout corpus (25.8 MB cache): a cold scan drops from 2,147 ms of main-thread time to 31 ms, and a steady-state incremental scan from 165 ms to 64 ms. The worker is stateless and the cache crosses the boundary both ways. That costs ~64 ms of structured clone at this corpus size, against 2,147 ms saved on the cold path, and it keeps the persisted cache the single source of truth — a worker-owned copy would need an invalidation protocol and a second resident copy of the same multi-MB array. Failure is closed, never a silent empty result: a worker that cannot spawn, times out, or crash-loops rejects, and the store records the scan error and keeps the previous projection. Two clients already carried the same FIFO/timeout/crash-cap machinery, so extract it once as WorkerThreadRequestQueue (with the packaged entry-path resolver as worker-thread-entry-path) and move all three onto it, rather than adding a third copy. Their existing tests pass unchanged. The oracle is event-loop utilization on the calling thread, not a stopwatch: usage-scan-worker-event-loop.test.ts runs the same scan both ways and asserts the worker leg leaves the caller idle while the main-thread leg does not, so CI load moves both legs together (#18788). * test(usage): compare the two scan arms instead of two fixed thresholds The event-loop oracle claimed to be self-calibrating — its header said "the ratio is self-calibrating, so CI load moves both legs together (#18788) instead of tipping a fixed millisecond threshold." It computed no ratio. Two separate `it()` blocks each asserted an absolute threshold against its own arm, run separately, so load moved them independently. The comment described a test nobody wrote, and the flake it promised was impossible is the one that landed: `activeRatio > 0.8` on the calling-thread arm measured 0.764 on an ubuntu runner. Fixing the comment is not enough, because the fraction is the wrong quantity. CPU contention drags the calling-thread arm's active/wall fraction *down* toward the worker's, since the loop parks waiting on a contended libuv pool. A 4-vCPU Linux container measured that arm at 0.175-0.756 across twenty runs, idle and loaded — never once above 0.8. Active *milliseconds* move the other way: contention stretches the caller's JS time far more than it stretches the worker arm's fixed post-and-deserialize cost, so the gap widens under load. Merge the two arms into one case over one corpus and assert the worker arm costs the caller under a fifth of the inline arm's active milliseconds. Same twenty Linux runs: 10.9x-83.6x, passing throughout. Keep the presence preconditions on both arms — an arm that silently scanned nothing satisfies the comparison trivially — and extend them to the calling-thread arm, which previously checked only file and session counts. * fix(ports): name the dropped command when the probe queue is full The shared-queue extraction turned `Port scan command queue is full; dropped ${command}.` into a constant string, because `describeFull` was given no way to see the request. Pile-up is per-probe, so the name is the only thing in that log that identifies which of lsof/ps/netstat was shed. Pass the rejected request to `describeFull` and restore the name. The request is built before the cap check so it exists to be named; the id it burns is a correlation token, so a gap costs nothing. The existing overflow test asserted only the error class, which is why the regression escaped a 29-test suite. It now dispatches the overflow under a different command than the accepted ones and asserts the message text, so a message that names the wrong request fails too. Also add a direct WorkerThreadRequestQueue test. Three subsystems share the queue and each client test only sees the parts its own protocol exercises, with `queueCap` reachable from port-scan alone. Covers one-at-a-time FIFO dispatch, the deadline starting at dispatch rather than enqueue, the consecutive-death cap, and both points where that count clears. And record the child-process hazard at the usage worker entry. `terminate()` reaps nothing the thread spawned, and OpenCode discovery reaches a fork today: `wslGated*` forks the WSL transcript sidecar for a `\\wsl$\...` path, which a Windows `OPENCODE_DB` or `XDG_DATA_HOME` can be. One scan through that entry with a UNC `OPENCODE_DB` forked a sidecar that outlived `terminate()`. * test(ai-vault): assert the OpenCode worker messages exactly, not by fragment Checked every message string in the two clients the shared-queue extraction rewrote against origin/main. Only the port-scan queue-full one regressed (fixed in the previous commit); the OpenCode SQLite client's four messages render identically, the remaining source diffs being renames — `error.message` to `lastError`, `call.timeoutMs` and `CALL_DEADLINE_MS` to `timeoutMs`. `session-scanner-worker-client.ts` was not touched by the extraction. But its suite could not have caught it either. `/timed out/`, `/exited with code/` and a bare `rejects.toThrow()` all still match a message that has lost its interpolated value, which is the same blind spot that let the port-scan regression through. Assert the rendered text instead: the timeout names its deadline, the exit names its code, and the crash-loop drain still carries the text of the fault that killed the run. * fix(usage): correct the worker entry's child-process note The previous note said `worker.terminate()` leaves a forked sidecar orphaned. It does not, and the reproduction that appeared to show it used a stub sidecar missing the `process.on('disconnect', () => process.exit(0))` the real entry has. With a faithful one: the sidecar lives exactly as long as the thread and is gone within 2s of `terminate()`, because tearing the thread down closes the IPC channel it owned. Two worker lifecycles forked two sidecars and leaked neither, and the pre-worker main-thread path reaps its sidecar the same way, on host exit. What is true and worth recording: a fork is reachable from this bundle at all, which is easy to miss; it survives only as long as the channel does; and the sidecar is now re-forked per worker lifecycle instead of pooled for the app's life. State those, and warn that a future child which does not exit on channel close would not get the same free cleanup. * fix(usage): kill a wedged scan worker on no progress, not on wall clock `USAGE_SCAN_TIMEOUT_MS` was a 10-minute deadline on the whole scan. A cold scan of a real history is legitimately minutes — 637 s measured on a 30 GB corpus with 300 worktrees before the per-cwd memo, ~51 s after — so a larger corpus or a slower disk crosses it. Crossing it killed the worker, recorded a scan error and left the cache unadvanced, so the next refresh started cold and died at the same point, forever. The deadline is now a no-progress window. The worker posts a file counter as it walks the corpus (`UsageScanWorkerProgress`, rate-limited to one message a second), and `WorkerThreadRequestQueue` re-arms the active call's timer on each one via the new optional `isProgress`. Clients that do not pass it keep the plain wall-clock deadline. `MAX_CONSECUTIVE_DEATHS` and idle teardown are unchanged. * refactor(usage): report scan progress as a file count, not one call per file Claude's scanner walks batches, so a per-file callback made it loop just to bump a counter. |
||
|
|
26721bd632 |
fix(codex): stop blocking the main thread on trust grants (#16441) (#16594)
* fix(codex): stop blocking the main thread on trust grants (#16441) Codex hook trust was granted by blocking the Electron main thread on `spawnSync` of a bundled ELECTRON_RUN_AS_NODE entry for the whole app-server deadline: 15s native, 35s WSL, ~45s on the real-home path (rebase inspect + repair + grant). Cold start and every Codex pane launch showed "Not Responding"; the reported event-loop gap was 15,049 ms. The subprocess only ever existed to donate an event loop to a deliberately blocked parent — `runCodexHookTrustGrantSession` was already the real async implementation. Make the callers async and the fork is unnecessary, so the bridge, the forked entry and its envelope are deleted along with their build/knip/tsconfig registrations. The CLI `agent hooks prepare-codex` handler is already async, so it awaits the in-process session and saves a process spawn per managed-home shell. `resolveCodexTrustGrantHost` is async too; the WSL identity probe moves from `execFileSync` to `runProcess`, dropping that file from the child-process import allowlist. Status reads keep a synchronous native-only stamp path. Two invariants that held only because the lane blocked: - Overlapping capability probes were impossible by construction. `GitCapabilityCache`'s dedupe engine is extracted to a shared `CapabilityProbeCache` and `CodexAppServerCapabilityCache` now inherits it, so concurrent launches against a cold host share one app-server session instead of one each. - Two grants on one `config.toml` could not interleave capture and restore. A reentrant per-file lane now serializes the whole install sequence (managed, WSL runtime, real-home ensure, legacy sweep) and the grant and rebase inside it. Cold-start work moves off the critical path: retained-home reconciliation (N sequential sessions) is fire-and-forget behind the daemon provider, and the startup real-home ensure chains into managed hook reconciliation instead of blocking app init. Every preserved semantic is unchanged: never throws, the ORCA_DISABLE_CODEX_TRUST_RPC kill switch, ledger hits, backfill-pending and cooldown fallbacks, config rollback on every failure path, pre-grant self-computed trust removal, the verify-failure taxonomy, diagnostics and telemetry. * fix(codex): widen the trust-config lane to every config.toml writer Review follow-ups on #16441's async trust grant: - `markCodexProjectTrusted` now runs inside the runtime+system config.toml lanes, so a project-trust write can no longer land inside a hook grant's capture->restore window and be silently reverted. Its callers await it. - `install`/`refreshRuntimeUserHooks`/`remove` hold the system config.toml lane as well as the runtime one — they promote approvals into ~/.codex/config.toml and mirror it back. Lock order is runtime-before-system everywhere. - The real-home ensure chain resumes after a rejection instead of returning the same rejected promise to every later pane launch, and resolving the real home is now inside the module's never-throws boundary. - `buildSpawnEnv` awaits inside a cancelable pending-spawn registration, so shutdown during the (now long) env build stops the PTY from launching. `prepareLocalPtySpawn` generalizes into `awaitCancelableLocalPtySpawn`. - CapabilityProbeCache drops the test-only `nowMs` passthrough; its probe backstop comment now describes what it actually guards. - Preflight is a plain async function; the trust dispatch in orca-runtime collapses into one `markWorkspaceTrustedForAgent`. * test(codex): exercise the trust-config lane under real concurrency The async grant makes two pane launches overlap for the first time. These drive the real modules end to end on real files: a rollback swallowing a sibling's grant, a markCodexProjectTrusted write landing inside a capture -> restore window, shared capability-probe dedupe on a cold host, the host-scoped transient cooldown, and reentrancy from inside an installer. Each was verified to fail against a deliberately broken implementation (lane removed, dedupe disabled, cooldown made global, reentrancy pass- through disabled). * test(codex): stop hook-service suites spawning the developer's real codex The forked grant bundle never existed under vitest, so the RPC lane was unreachable in tests on main. Running it in-process makes these suites spawn a real `codex app-server` when one is installed: 38 spawns and two failures in hook-service-runtime-trust-repair on a machine with codex, green in CI where there is none. Stand in for the missing binary so both environments exercise the same fallback lane. * docs(codex): scope the trust-RPC kill switch comment to what it actually gates The comment read as though the flag forces the fallback lane everywhere. It gates the managed grant only: the real-home rebase still runs its own inspect/repair app-server sessions when Orca's insertion shifts a user's hook positions, and never reads the flag. Verified by exercise, not by reading — with the flag set, both inspect-user-hook-trust and repair-user-hook-trust still ran. Pre-existing: main has no check there either, it just blocked the main thread while doing it. Widening the flag to cover the rebase is a follow-up; this only stops the comment promising something the constant does not do. |
||
|
|
4ec6bbf588 |
Kill hung WSL transcript filesystem operations via child process with route quarantine (#15381)
* fix(native-chat): kill hung WSL operations via child process
Stalled UNC file operations hold libuv permits even after the gate
timeout expires, blocking Chat tab recovery. Two stalled operations
fill both permits and freeze all WSL access until restart.
Fork file I/O for UNC paths into a separate child process. On deadline
expiry, kill the process to force the hung syscall to exit. This frees
the permit for the affected tab's next read. Temporarily quarantine the
stalled route to avoid retry storms.
* chore: drop internal review artifact from the repo root
* fix(native-chat): harden the WSL transcript fs sidecar
Review follow-ups on the sidecar isolation change:
- Only the deadline may abort running gate work. The sole waiter's
same-duration timeout fired first, killed healthy children on caller
abandonment, and settled the task before the deadline could quarantine
a stalled route - leaving the back-off dead for every dedupe:false op.
- Resolve the fork entry from out/main/chunks too: the resolver compiles
into a shared chunk, and the scanner service child has no
process.resourcesPath, so packaged WSL vault scans threw entry-not-found
(masked as an empty tree).
- Allowlist the fork env instead of spreading process.env; ambient
NODE_OPTIONS would halt or --require code into every child.
- Wrap transport faults (spawn failure, child death) in
WslTranscriptFsError('unavailable') so discovery reports them as scan
issues instead of misreading them as missing paths or empty trees.
- Gate the vitest in-process fallback on the vitest worker global so a
leaked VITEST=true cannot revert production to in-process UNC syscalls.
- Reap idle sidecar processes after 60s instead of holding them for the
app session.
- Split 'open' into its own protocol union member so the reusable-call
Exclude actually strips it from the pooled-process API.
- Guard kill('SIGKILL') against the teardown race where an exiting child
emits an unlistened 'error', and dispatch reads by handle kind before
path spelling.
* fix(native-chat): probe stalled WSL routes instead of a fixed quarantine
Remaining review follow-ups:
- Escalating route quarantine: first strike lifts after 5s so a distro
that was cold-booting when its op hit the deadline recovers on the
next poll (~35s total instead of ~90s); repeat stalls double the
back-off toward the prior 2x-timeout cap, and any settle the deadline
did not force clears the strikes. Queued same-route tasks fail fast
at quarantine instead of stranding one waiter deadline per file in
sequential scans.
- Single request implementation: the vitest in-process fallback now runs
the child's own dispatcher (WslTranscriptFsProcessOperations + decode),
so unit suites exercise exactly what the forked process executes and
the per-call-site fallback closures are gone. Dirent fixtures gained
the full kind-flag set the serializer reads.
- Dropped the production-dead per-route close queue; UNC FileHandles
(test fallback only) mirror the process-handle close contract.
- Error class, messages, and factories move to wsl-transcript-fs-error
(re-exported from the gate) to keep the gate under the lines budget.
* fix(native-chat): harden WSL transcript fs with route quarantine strike
Extract quarantine logic into a dedicated module with strike decay: stalls older
than 5 minutes restart from base back-off, and concurrent-lane timeouts count as
one incident. Allow joining live in-flight tasks on quarantined routes (they cost
no new I/O). Preserve quarantine across transport faults (child death). Handle
file shrinking during tail reads by detecting short reads and returning empty.
Defer file closes that arrive mid-read instead of refusing, preventing slot
leaks. Separate process slot and boundary-finding concerns into focused modules.
* fix(native-chat): enforce route quarantine windows and isolate lanes per
A late result arriving after the deadline was incorrectly lifting the route
quarantine, allowing subsequent work to start before the back-off period
expired. Now late results are correctly recognized as stale and never cut
the quarantine short.
Process work is now isolated per (route, priority) lane so a scan stall
cannot block exact reads on the same distro. Each lane gets its own client
and process pool; late results and handle faults stay scoped to their lane.
Tests now fake performance.now() alongside timers (the quarantine clock
depends on it) and wait for the full back-off window to expire rather than
advancing by 0. Gate state is reset between test cases since late releases
never lift the quarantine.
---------
Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
|
||
|
|
4c69e45552 |
Strengthen plain-node-entry-guard with entry name validation (#12761)
* Strengthen plain-node-entry-guard with entry name validation - Add buildStart hook to validate guarded entry names exist in rollup inputs, preventing stale names from silently stopping guards - Extend electron require detection to subpaths (electron/main, etc) - Improve smoke test signal/exit handling and use constants - Add comprehensive tests for entry validation and new behaviors * Add SIGKILL escalation to plain-node-entry-guard timeout Switch from spawnSync to async spawn to properly handle daemons that trap SIGTERM. spawnSync's timeout only sends the signal and waits, so a daemon that ignores SIGTERM causes the build to hang. The new runDaemonEntry function escalates to SIGKILL after a grace period to enforce the deadline. Configurable timeouts and grace periods via SmokeTimings type; closeBundle hook becomes async to support the change. |
||
|
|
5df2ddbc9c |
perf(ai-vault): isolate tab title resolution (#13377)
* perf(ai-vault): isolate tab title resolution * fix(ai-vault): preserve background scan caches * fix(ai-vault): resolve nested worker from chunks |
||
|
|
fde816e4ee | move folders (#12758) |