Commit Graph
518 Commits
Author SHA1 Message Date
Neil 13ecf051c3 Reuse prepared Windows native builds in SSH CI (#24555)
* ci: reuse qualified Windows server slots for SSH host tests

* ci: reuse prepared relay addons after an exact native cache hit
2026-10-02 03:31:57 -07:00
Neil efbf651c7b Reduce CI setup costs and fixture failures (#24537)
* Let scheduled CI warmers wait and measure WebRTC startup

* Measure a smaller daemon shutdown fixture image

* Counterbalance WebRTC startup and verify retained fixture files

* Record CI fixture measurements and remove temporary pilots

* Clarify fixture build dependency cleanup evidence

* Make coalesced snapshot fixture delivery deterministic

* test: type the PTY write delay observer
2026-10-02 02:46:02 -07:00
Neil 6153fbcfe4 Reduce redundant headless server CI work (#24527)
* ci: avoid unrelated headless server qualification

* ci: skip headless detection for ineligible draft PRs

* ci: preserve cross-host qualification and skip supplied prerequisites

* ci: include Windows server cache validation in change detection
2026-10-02 01:42:41 -07:00
Neil 5a56636f66 Bound E2E package setup and retain cancellation traces (#24617)
* test: align source-control fixtures with current store contracts

* Bound E2E package setup and retain cancelled-job traces
2026-10-02 01:03:48 -07:00
Neil 8ff6296bc7 Speed up serializer checks and keep native caches stable (#24476)
* Reuse serializer oracle cells and isolate native cache policy

* Preserve native cache post-save paths and record hosted oracle gain

* Record native cache reuse and separate cancel-test startup budget
2026-10-01 21:43:56 -07:00
Brennan Benson 444f1952c7 ci: run every cross-version wire test, picked up by folder so new ones can't be skipped (#24499)
* ci(cross-version-wire): run the whole directory so no compatibility test is left out

Three cross-version tests ran in no CI job because the job named its files by hand.
Run the directory instead, ratchet that every file kept out of the unit shards
runs in some PR job, and re-run the job when the modules the newly running
tests guard change.

* test(cross-version): give the orchestration downgrade test its siblings' 120 s budget

* ci(unit-exclusion): count only merge-gating jobs, and require each excluded file's job to fire on it

The coverage check counted any pr.yml job, including e2e, terminal IME and Windows WSL, which are
left out of verify.needs and so cannot block a merge. It now reads verify.needs and the reusable
workflows those jobs call.

It also only proved that some step names each excluded file, not that the job runs when the file
changes. The structured-session zsh login-shell test runs only in shell_contracts, whose path
trigger matched neither it, its harness nor its subject, so a PR touching only those ran it
nowhere. The check now asserts a change to each excluded file fires a gating job that names it,
and the shell trigger gains those three paths.

* ci(cross-version-wire): trigger on the turn-outcome vocabulary and the schema version-skew resolver

A change confined to src/shared/agent-turn-outcome (the arms a newer host publishes) or to
orchestration-schema-version-skew (how current code reopens a downgraded database) skipped the
job whose tests guard exactly those contracts. Also corrects the publish/read direction in the
turn-end comment.

* test(cross-version): state why the orchestration downgrade test needs 120 s

* test(ci): glob the unit tree once for the unit-exclusion coverage checks
2026-10-01 20:53:51 -07:00
Neil f69052e113 Reuse qualified Windows server builds and dependency verification records (#24448) 2026-10-01 16:19:40 -07:00
Brennan Benson 976dc00337 fix(native-chat): Stop's pause is worked out from the chat's history, so a steered message is never re-sent (#24072)
* test(native-chat): a Stop over a card sent now into the running turn keeps it paused

Red on main: Codex's turn end withdraws the steered hand-off and the queue
sends the card again as a new host turn, with no pause recorded.

* fix(native-chat): a Stop's queue pause holds a card whose hand-off is still unanswered

A card sent now into the running turn was still pending when Stop judged the
pause, so nothing was recorded; the interrupt then withdrew the hand-off, the
card went back to waiting unpaused, and the queue sent it again as a new turn.
The pause now counts a hand-off that may still return to waiting, judged with
the appended row applied, so a withdrawal lands under the pause and an
acceptance retires it in that same write. Codex and Claude both hit it.

* test(native-chat): the Claude re-send case fails on its diff, inside the test's budget

* test(native-chat): a pause held by an unanswered hand-off ends on every path that ends it

The provider's answer, the provider dying, the chat closing, a restart, a
withdrawal still owed at open, and a /clear (refused until the hand-off ends,
then carrying every waiting card paused 'cleared'); each ends with the queue
sending again.

* fix(native-chat): narrow the pause's settled hand-off, and assert the queued receipt's card

* refactor(native-chat): derive the queue's pause from Stop and Resume journal rows

Stop now appends one journal row where it takes effect, before the interrupt,
whatever the queue holds; Resume appends its own. The pause is a pure function
of the fold: the latest Stop with no later Resume and no later accepted turn a
person asked for. A /clear's carried cards name their source, which is the
replacement's 'cleared' pause. Host-origin turns never lift either.

One predicate decides which cards a pause holds; by default every waiting card
without a hold of its own, including one queued after the Stop. The drain's
consume re-judges it inside its own transaction.

The rows are tombstones of an id no item takes, carrying the mark: a released
host reads an unknown row kind as corruption and truncates the journal there.

Deletes the stored pause (recordPause, the retire hook on every appended row,
the settle-before-record step, mayReturnToWaiting and its row overlay) and the
tests that only proved it retires. The queued_message_pauses table stays in the
schema, unread and unwritten, for downgrade safety.

* fix(native-chat): a card queued after a Stop sends normally, never ahead of held ones

A Stop's pause now holds only the cards queued before its row, plus a steer it
withdrew, which returns to its own place. Each card records the journal
position it was queued at, and the one hold rule compares that with the Stop
row. A card queued after the Stop is a new instruction: it sends as usual, but
the drain still stops at the first held card, so it never overtakes them.
/clear's pause holds the cards it carried. Holding every card again is a
one-line switch in that rule.

* fix(native-chat): the queue's own send re-checks the no-overtake rule in its transaction

The drain's pick and its consume now read one function, nextSendableQueuedCard,
so a Stop row that lands between them holds a newer card behind an older held
one exactly as the pick would. Notes why Stop and Resume ride a tombstone row.

* fix(native-chat): stop creating the unused queue pause table

The queue's pause is derived from journal rows, so nothing reads or writes
queued_message_pauses. It was still created on every open "for downgrade
safety", but an older build creates it itself when it opens the database, so
the table only sat empty in every new database. The tests now pin that no
pause table exists.

* fix(native-chat): a Stop's pause never hides the restart pause

A Stop holds only the cards queued before it. The pause derivation still
returned the Stop alone whenever it was in force, so the restart pause was
never considered: a card queued after the Stop, written by a host process
that has since exited, sent by itself after Orca restarted, with no pause
header and no Resume. A /clear pause that held nothing could hide it the
same way.

Every pause in force is now derived. A card is held if any of them holds
it, and it names the first that does. The drain's pick, the consume
transaction's re-check and the published header all read that one rule;
the header names the pause holding the first card Resume would send.

* test(native-chat): pin the Stop's no-resend, lift and held-card rules

- The Claude and Codex Stop-withdraws-a-steer tests checked "not sent
  again" at one instant, before a queue ignoring the pause re-sends. They
  now wait for the stopped turn to end and re-check after a quiet window.
- The deleted-card test read a card queued after the Stop, which sends
  whether or not a person's turn lifts it; it now reads the Stop's pause
  before and after that turn.
- Unit cases pin that a Stop holds a card with no recorded position and one
  queued before a rewind.

* refactor(native-chat): a Stop writes one Stop event with its reason, turn and caller

The Stop row that paused the queue becomes the general Stop event
{ reason, turnId?, at, caller? }, whose reason is the host's existing stop
cause. It still rides a tombstone of a host-only id (a released host deletes
the journal from the first unknown row kind), and Resume keeps its own marker
on its own id. Only a person's Stop (reason user-stop) pauses the queue.

* test(native-chat): a rewind keeps a lifted /clear pause lifted and restates the same Stop event

* test(native-chat): pin that Stop and Resume rows never reach apps or count as history

* test(native-chat): only a person's Stop event pauses the queue

* test(native-chat): pin that a Stop's event precedes the interrupt and the at-start stop

Through the real host: the event names the turn and who asked and is in the
journal when the interrupt reaches the agent; at an agent still starting it is
there before the start is ended and holds a card queued before it; an idle Stop
writes one only when it withdrew a send; and the queue's claim re-judges a
pause that landed after its pick.

* test(native-chat): a card held at a starting agent is checked before the Stop's timing

Also says precisely what the claim's in-transaction pause check defends
against: the Stop and the drain share one serialized lane.

* test(native-chat): a released build keeps and folds a journal holding Stop events

Replays this build's rows from the released build's own journal database: every
row is kept, the history after the Stop still folds, and an older client is sent
only removed ids no item uses.

* style(native-chat): format the Stop event changes

* test(native-chat): type the released build's exports through one checked helper

* fix(native-chat): the Stop/Resume row guard narrows to those tombstones only

* test(native-chat): run the Stop-event downgrade test in CI, and cover a writable downgrade

The Stop-event downgrade test ran in no CI lane: unit shards exclude the
cross-version folder, and the cross-version lane runs a fixed file list that
did not name it. It is now on that list.

Its only case replayed the rows into a release's own fresh database, because
that release cannot open the current host database. A second case opens the
journal this build wrote with a main build that shares the database: it opens
writable, keeps every row, appends, and this build then reopens it with the
person's Stop still pausing the queue.

* fix(native-chat): a Stop that stops nothing new writes no Stop event

A Stop reaching a running agent wrote a Stop event on every press. Two
presses before the first interrupt landed wrote two events, so a card
queued between them counted as before the latest Stop and was held,
though a card queued after a Stop should send normally. A Stop naming a
turn that had already ended, as a phone sends late, also wrote an event
for a turn it never stopped.

It now writes one only when it withdrew a queued send, or stops something
no event records yet: not a turn the journal no longer runs, and not the
live turn a Stop still in force already names, unless a card was handed
over into it since, which this Stop's interrupt sends back and must hold.
The interrupt and the "already finished" note are unchanged. A Stop at a
starting agent still always writes.

* test(native-chat): pin that a later host, eviction or close Stop never lifts a person's Stop

* chore(native-chat): put each Stop-row doc on its own declaration, and say only user-stop is journaled

* fix(native-chat): any later Stop event ends a person's Stop pause

A person's Stop paused the queue until their next accepted turn or Resume,
and a later Stop of another reason (the host stopping the agent, an
eviction, a close) was ignored. Now the pause is the latest Stop event's:
a later Stop of any reason ends a person's pause, and only a person's Stop
pauses. The fold keeps the latest Stop event whatever its reason.

An eviction of a resting chat writes no Stop event (a Stop that stops
nothing writes nothing), so it cannot release held cards; a test pins that
no event means no lift.

* fix(native-chat): a second Stop press is a repeat even when the first came before the turn showed

A Stop pressed before the agent's turn shows in the journal (before
Claude's echo, or before Codex opens the turn) records no turn. A second
press once the turn showed compared that missing turn with the live one,
wrote a second Stop event, and held a card queued between the presses.

A repeat is now judged by what was sent since the Stop in force: with
nothing sent after it (a refused send aside), a Stop that named no turn,
or named the live one, is repeated and writes nothing. Anything sent since
and not refused, including a send whose fate is unknown, makes the new
press write, since its interrupt may send that card back to waiting.

Tests: the two-press case across the turn showing; a steer between the
presses settled unknown; and a Stop naming a turn that ended while the next
card is sent but shows no turn yet, which writes and holds that card. The
fold test that claimed an eviction path is renamed.

* fix(native-chat): the queue's pause ignores a Stop or Resume row holding a value no build writes

A Stop or Resume row's value is read from disk with no shape check, and
the pause fold stored whatever it found. A stored `stopEvent: null` would
then throw on every pause check for that chat: the queue's pick, its
send, and every queue update to clients. No build writes such a row, so
this is hardening.

The fold now reads a Stop only when it is an object with a string reason
and a finite time, and a Resume only when it is `true`. Anything else is
ignored: it pauses nothing and ends nothing. The row is still not treated
as malformed, which could cut the history short.

* refactor(native-chat): one reading of a Stop's turn for its event and its note

A Stop's event and its note each worked out the same two facts on their own:
which turn the Stop is about (the one it named, else the one running), and
whether a named turn is the one the journal shows running. The event decides
before the interrupt; the note and whether the session ends decide after the
provider's answer, so those decisions stay separate, but the facts they read
are now one helper each in structured-agent-session-turn-stop-notes.ts:
structuredAgentSessionStoppedTurnId and
structuredAgentSessionStopNamesTurnNotLive. The event's turn, the note's key,
the session-ending condition, the running-command check and the repeat check
all read them. No behavior change.

Tests: a Stop naming no turn records the running turn on its event, and
rewrites that turn's note as a Stop naming it does.

* refactor(native-chat): a failed-interrupt Stop reads its turn through the shared helper

The new branch that ends a Codex child after a failed interrupt asked
whether the Stop's turn still runs with `turnId ?? liveTurnId`, a third
copy of "the turn a Stop is about". It now reads
structuredAgentSessionStoppedTurnId, the value the note key already uses,
read at the same point before the cancel. No behavior change.

Test: a Codex Stop whose interrupt failed ends the child, holds the card
queued before it with the queue paused, and writes its Stop event before
the turn's end.
2026-10-01 13:21:58 -07:00
Jinwoo HongandClaude Opus 5.5 ae41eb414a fix(terminal): give plain fish tabs Orca's codex function without changing fish's startup (#24284)
* fix(terminal): give plain fish tabs Orca's codex function without changing fish's startup

A `codex` typed into a plain fish tab ran without --no-daemon because only
wrapped fish tabs (startup command / ready marker) got Orca's codex function.

Plain fish spawns now prepend an Orca data dir to XDG_DATA_DIRS and record the
exact prefix in ORCA_FISH_XDG_DATA_DIRS_PREFIX. Fish sources the dir's
fish/vendor_conf.d snippet, which first restores XDG_DATA_DIRS (unset again if it
was unset), erases the marker, drops its dir from fish's derived vendor/function/
completion paths, then defines the shared fish codex function at the first prompt
so the user's config.fish still wins. fish argv is unchanged; wrapped tabs keep
their existing -C path. A local fallback to another shell restores the user's
XDG_DATA_DIRS instead of deleting it.

Bumps the terminal daemon protocol to v39 so new tabs move to a daemon that
injects the env; v38 owners stay attachable.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(fish): skip the XDG handoff for -N/--no-config and empty XDG_DATA_DIRS

fish never reads vendor_conf.d under -N/--no-config (also abbreviated or
clustered), so the snippet could not undo the prefix; and the restore cannot
tell an empty XDG_DATA_DIRS from an unset one. Both now launch untouched.
Run the real-fish handoff tests in the shell contracts job, where fish is
required, so they no longer skip in CI.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(fish): compare the unset-restore case against a fish without Orca

Ubuntu runners ship snapd's fish vendor snippet, which sets XDG_DATA_DIRS on
every fish start, so "unset" was never the right oracle there.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(fish): treat an empty XDG_DATA_DIRS like unset so the tab still gets the codex hook

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(pty): put back the user's own launch env on a shell fallback

The primary shell's launch config now records the pre-launch value of each
key it writes. A fallback shell restores those values (unsetting keys that
had none) instead of deleting the keys, which hands back an inherited
XDG_DATA_DIRS after a fish fallback and an inherited ZDOTDIR after a
zsh->bash fallback, with no per-shell special case.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor(fish): drop the Node restore twin and simplify the vendor snippet

- Remove restoreFishXdgDataDirs; the generic fallback restore covers it.
- Snippet: read ":$XDG_DATA_DIRS:" directly and filter Orca's vendor dirs
  with one string match per variable.
- Require inheritedXdgDataDirs in both getShellLaunchConfig option shapes.
- Drop the test-only FISH_XDG_DATA_DIRS_HANDOFF_DAEMON_PROTOCOL_VERSION.
- Fix stale fish comments and trim redundant -N launch cases.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor(fish): stop scrubbing fish's lookup paths after the handoff

Only XDG_DATA_DIRS is restored, by exact prefix; Orca's dir holds nothing but this snippet, so leaving it on fish's derived paths loads nothing else and drops the glob match.

* docs(fish): drop the comment for the removed vendor-dir cleanup

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 15:38:53 -04:00
Jinwoo Hong 07dad6739a refactor(relay): sample fleet health inside the same-cap roll instead of a separate monitor run (#24443)
* refactor(relay): sample fleet health inside the same-cap roll instead of a separate monitor run

A same-cap wave no longer consumes a 15-minute monitor dry-run and its sealed,
single-use, five-minute-fresh evidence. Each apply wave now samples fleet health
itself right before isolation, with the monitor's evaluator, thresholds, and
tolerances, for a window sized to the cell's host count (3/5/8 min), plus three
lookback rules: no cell container exit in 10 min, no minute over 500 director
503s in 10 min, and director concurrency p99 within the monitor bar over 4 min.

Removes the monitor-run inputs, the gate's consume/authorize steps, the
break-glass override, and the same-cap-only authorization shapes in
relay-monitor-evidence.mjs. The monitor workflow and the rehome enable path are
unchanged.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): bound the pre-drain sample overrun and keep the drain token fresh

Review follow-ups: alternating tolerated readings could hold the sample open
until its step timeout, so cap the overrun at three samples past the window;
record why a read failed; mint a fresh admin ID token for the drain after the
sample; raise the job timeout to 90 min so a long sample cannot cancel the
job past the failsafe.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* feat(relay): exempt the rolled cell and existing-only cells from the pre-drain crash rule

The exit rule counted every relay container exit fleet-wide, so a cell that
crashes every few hours (c25, 12 a week) blocked the very roll that fixes it,
and existing-only legacy cells (c5, 15 a week) blocked rolls they take no part
in. Exits are now grouped by instance, each instance is named by its own newest
runtime-metrics log line, and only exits on general or migration-only cells
other than the target count. An exit no configured cell can be named for trips
the rule; a failed lookup is a failed read. relay-observability.tf joins the
evidence-code set because the rule depends on its filter.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* test(relay): cover re-asking for an unnamed exiting instance; note the boot-exit risk

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-10-01 15:36:17 -04:00
Jinwoo Hong 051ca4d34e feat(relay): declare US cells c32 and c33 at the 3,000-host shape (#24444)
* feat(relay): declare US cells c32 and c33 at the 3,000-host shape

Declares two us-central1 cells at the Asia shape (cap 3000, 6000 request
units, e2-standard-4) with the US default pool of 10, and generalises the
Asia topology and admission ladder to derive each wave's region from its
reviewed zone, leaving every Asia wave's behaviour unchanged.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* docs(relay): note the US canary tie-break and leave the fleet pool list to promotion

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): plan C32 and C33 as one topology wave

The live-image overlay refuses a declared non-target cell with no template,
so a lone C32 plan would fail on C33. Registration and promotion stay
one cell at a time.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-10-01 15:34:57 -04:00
Neil 197ea3a3b3 Free PR CI capacity by avoiding repeated setup and real-time test waits (#24355)
* Reduce repeated PR setup and transcript timing waits; add hosted comparisons

* Align parallelism contract with Node-only external rebuild toolchain

* Record hosted coverage and launch package, store, and cancellation comparisons

* Apply hosted Windows setup savings and remove measured test waits

* Keep measured PR package gains and remove completed comparison jobs

* Report measured test counts with precise units
2026-10-01 11:51:43 -07:00
Jinwoo Hong 56c7642aae test(orcad): skip the live-terminal runtime hand-over across a protocol bump (#24429)
* test(orcad): skip the Bun-to-Node live-terminal hand-over across a protocol bump

The last Bun orcad's daemon reports protocol 38 forever, so asserting the
adopted daemon matches this checkout's PROTOCOL_VERSION failed every bump.
Ask the Bun slot's daemon for its protocol once, run the hand-over when it
matches, and skip with the two versions named when it does not: a daemon
at another protocol is never adopted across an update.

* test(orcad): clean up the Bun protocol probe even when its launch fails

The probe's cleanup ran only after a successful launch, so a launch that
timed out or threw left its orcad and daemon running. One finally now
stops the orcad, kills what it launched, and kills any daemon named by a
pid file in the probe's data root.
2026-10-01 14:28:05 -04:00
OrcaWinandm4air 6d1a97ef98 fix(ssh): launch the Windows relay outside sshd's job so standard users work (#24224)
* fix(ssh): launch the Windows relay outside sshd's job without WMI

Win32-OpenSSH kills a session's job on close but allows breakaway. relay.js
gains a one-shot launcher mode that starts the detached relay with
CREATE_BREAKAWAY_FROM_JOB through the staged process-tree addon, so a standard
user no longer needs a WMI Remote Enable grant. WMI stays as the fallback for a
relay without the addon, and a refusal there is named. The Windows SSH-host
lanes drop their WMI grant and assert the breakaway route and adoption.

* fix(ssh): find runtime holds without WMI on a standard-user Windows host

The store GC read held runtimes through Get-CimInstance Win32_Process, which
WMI refuses to a standard user's SSH logon, so the pass kept every runtime.
On a refusal it now reads this account's own process image paths through
Get-Process.

* build(relay): ship the Windows relay launcher addon in every desktop package

macOS and Linux packages carried Windows relays without windows-process-tree.node,
so a legacy-runtime relay they uploaded to a Windows SSH host could not launch
outside sshd's job and fell back to WMI, which a standard user is refused.

A reusable Windows job now compiles the x64 and arm64 addons once and uploads
them; release-cut, release-mac-build, and the hourly/daily/adhoc mac builds
download them before build:release and require both arches. Staging now rejects
a binary with the wrong PE machine, the ReadProcessMemory import, or no
spawnOutsideJob export, so a stale pre-launcher build cannot ship.

* ci(ssh): run the Windows SSH-host lanes when the relay process-tree build scripts change

The staging and gyp-rebuild scripts decide which windows-process-tree addon the
relay ships, so a change to either must re-prove the Windows host cells.

* test(ci): find the mac orcad-template download by artifact name

The release mac job now also downloads the relay Windows process-tree addons, so
the first download-artifact step is no longer the template's.

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 05:32:09 -07:00
8afa1db50c feat(ssh): rung B glibc 2.17 compat runtime; gate remote vault on host node:sqlite (#24148)
* feat(ssh): wire rung B to the glibc 2.17 compat runtime; gate rung C vault on full node:sqlite

- COMPAT_RELAY_RUNTIMES lists linux-x64-glibc217; rung B plans the compat slot and compat
  pinned Node when glibc is below 2.28 or rung A refused with libc_floor/missing_lib.
- The relay version folds the compat runtime's executable hash; refusals are cached per runtime.
- The orcad template stages an optional linux-x64-glibc217 target (base package + compat
  node-pty slot + compat runtime marker); the verifier and materializer accept it.
- node-pty slot loader falls back to the compat slot when the default slot is missing or
  needs a newer glibc.
- Runtime store GC keeps the compat pin beside the default one on every relay connect.
- hasNodeSqliteReaderApi (DatabaseSync + backup) gates relay session search and the relay
  OpenCode reader, which now names the host Node version in its unavailable reason; the SSH
  vault reader installs the compat Node on old-glibc hosts and uploads nothing when no
  pinned Node can run.
- Rung D: a remembered noexec reports home_noexec and never advises installing Node.

* fix(ssh): re-prove a replayed noexec after rung D so allowing exec recovers the host

* fix(ssh): keep the rung B compat runtime pinned in the relay-connect store GC

* test(ssh): mock deployment-target facts in the Windows OpenCode runtime tests

* ci(ssh): build the glibc 2.17 compat slot for the hostile-host matrix; CentOS 7 lands on rung B

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>
2026-10-01 05:32:05 -07:00
Jinwoo Hong 744e7722c2 fix(relay): accept MIG version-name reconciliation and recreate stranded cells without rewriting the MIG (#24373)
* fix(relay): accept MIG version-name reconciliation and recreate stranded cells without rewriting the MIG

The stranded-rollback recovery ran a gcloud rolling action, which renames the
MIG version outside Terraform. Every later plan for that cell then reverted the
label, and the capacity-plan validator refused the revert as an unreviewed MIG
change, so the cell could be neither rolled nor rolled back.

The validator now accepts a MIG field moving back to what relay-gce-cells.tf
declares (version name and update policy), in every mode, and a test pins those
values to the Terraform file. The stranded branch recreates the cell's single
instance with recreate-instances, which leaves the MIG untouched.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): let a label-only MIG plan through and recreate on it in a stranded rollback

A stranded rollback whose template is already in place plans only the version
name revert. The validator still required the MIG template to move, so that
plan was refused, and the recreate gate (changes == 0) would have skipped a
plan of one change and left the drain flag set. Require the template move only
when no declared field reconciles, and recreate whenever the template was not
replaced (changes < 2).

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-10-01 08:12:49 -04:00
OrcaWinandm4air f9940d5354 ci(ssh): macOS SSH-host lane for the pinned relay; fix uploads under a symlinked root (#24179)
* test(ssh): upload a root reached through a symlinked parent

The upload-root realpath fix landed with #24180; this keeps macoshost's case
where the root is passed explicitly beneath a symlinked parent.

* ci(ssh): macOS hostile-host lane on a loopback user-level sshd

Adds local-sshd cells for darwin-arm64 (macos-14) and darwin-x64
(macos-15-intel): a non-root sshd on 127.0.0.1 logs in as the runner user
with SetEnv PATH=<shims>:/usr/bin:/bin:/usr/sbin:/sbin and an empty HOME, so
no rc file restores Homebrew. The driver asserts rung A, terminal echo,
cached runtime reuse, GC keeping the in-use runtime, no toolchain or xattr
calls, and that the SFTP-uploaded Node carries no quarantine and runs as
uploaded. Docker cells are unchanged; each machine runs only cells it can host.

* test(ssh): fail a hostile-host run that would skip every named or hostable cell

A cell named for the wrong OS or arch was silently skipped, so a macOS job on a
mismatched runner went green having deployed nothing. Named cells must now be
hostable here, and a gated run must select at least one cell.

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 04:43:46 -07:00
OrcaWinandm4air 0ad77ea2f7 ci(ssh): Windows SSH-host lanes (inbox + preview OpenSSH) for the pinned relay (#24180)
* ci(ssh): import the private Windows OpenSSH provisioning harness

Copied unchanged from origin/OrcaWin/np-windows-ssh-provider-diagnostic
(config/ci/windows-ssh-provider/preview-ssh/ at 1242f3c4c8, commits 78b3857a0d,
c24adccff0, 069aa7b38b): a private LocalSystem sshd service on 127.0.0.1 for a
dedicated standard user, either the Microsoft-signed Win32-OpenSSH
10.0.0.0p2-Preview ZIP (archive and every binary pinned by sha256) or the inbox
OpenSSH.Server capability binaries. The following commits extend it for the
pinned-Node relay host lanes.

* test(ssh): run hostile-host cells through a host-agnostic driver

The Docker matrix drove the relay deploy and inspected the container with
inline docker exec calls, so no other host could reuse it. Split it into:

- ssh-hostile-host-test-harness.ts: the deploy, ladder observation, terminal
  echo (per-shell probe), runtime reuse and GC-keeps-in-use assertions, now
  also capturing every command the deploy sent the host.
- ssh-hostile-host-observer.ts: how a driver inspects the host outside SSH;
  docker exec for containers, the local filesystem for a loopback host.
- a legacy_opt_out outcome: the ladder never runs and nothing enters the
  pinned store, whatever the host-Node path does.

Launched cells now also check the runtime's sha256 on the host and that every
slot file (the Windows bundled ConPTY pair included) landed in the relay dir.

* ci(ssh): Windows SSH-host lanes for the pinned-Node relay

Phase 2 exit gate, Windows half: the real deployAndLaunchRelay through a real
SshConnection against Win32-OpenSSH on 127.0.0.1, on windows-2022 (x64) and
windows-11-arm (arm64), for both the inbox OpenSSH.Server capability and the
Microsoft-signed 10.0.0.0p2-Preview release (ZIP; archive and each binary
pinned by sha256 and Authenticode, as in the imported harness).

Builds on the provisioning harness from
origin/OrcaWin/np-windows-ssh-provider-diagnostic (previous import commit):
- one private standard account per cell, so every cell starts from an empty
  runtime store;
- -HiddenTools: the private sshd service's own Environment carries a PATH
  without any machine PATH entry holding node/npm/compilers, led by logging
  .cmd shims; a session probe fails the job if node.exe still resolves;
- DefaultShell set per cell by invoke-pinned-relay-cells.ps1 and restored at
  cleanup (dispatch proven per cell via %COMSPEC%).

Cells (src/main/ssh/ssh-windows-host-cells.ts): pinned-cmd (stock sshd),
pinned-powershell (DefaultShell = Windows PowerShell) expect rung A on the
pinned node.exe with the relay self-test passing, terminal echo, runtime
reuse, GC keeping the in-use runtime, stage identity through node.exe and no
Add-Type in any decoded session command; legacy-opt-out expects the ladder
never to run and an untouched pinned store.

* fix(ssh-ci): tolerate absent-drive PATH entries and retry Windows userData teardown

Join-Path throws on a machine PATH entry naming a drive the runner lacks,
which would abort provisioning before any cell ran; the toolchain split now
probes with [IO.File]::Exists over [IO.Path]::Combine, and the self-test
covers an absent drive. The hostile-host harness removes its throwaway
userData with removeTreeSync so a transient Windows lock cannot fail the
lane's afterAll.

* fix(ssh-ci): stop the account list rebinding the typed -Accounts param

PowerShell variable names are case-insensitive, so $accounts=[List[hashtable]] assigned into
the [int]$Accounts parameter and every Windows host job died before provisioning. Rename the
list and make the provisioning self-test reject script-scope assignments that shadow a param.

* fix(ssh-ci): hide the host toolchain by ACL, since sessions ignore the service PATH

Win32-OpenSSH builds a session's PATH from the machine and user registry values, so the private
service's Environment never reached SSH sessions and host node.exe stayed visible. Deny the private
accounts the toolchain PATH directories, put the logging shims on each account's own PATH, and
record failing sshd and client log lines so a refused login is diagnosable from the receipt.

* fix(ssh): resolve the upload root before checking entries stay inside it

uploadDirectory compared each entry's realpath against the root as given, so a root reached
through a symlink, junction or Windows 8.3 short name (C:\Users\RUNNER~1 in TEMP) rejected every
entry as escaped and the pinned runtime upload never started.

* fix(ssh-ci): fail cells on a vitest failure and give each account its own keys file

The cells script read $LASTEXITCODE under the workflow's GetNewClosure callback, which sees a
stale captured copy, so failed cells reported exit 0 and the job passed. Read the global value.
Inbox sshd 8.1 checks authorized_keys with read_ok=0, refusing a file other accounts can read;
use one keys file per account via %u.

* fix(ssh-ci): keep the account name in inbox mode and surface the WMI launch gap

The inbox binary-verification loop reused $name, so later SSH and SFTP probes logged in as
'sftp-server.exe'. Before the cells run, probe whether a standard SSH user can call WMI
Win32_Process.Create (the Windows relay launch path); when refused, warn and grant the cell
accounts Remote Enable on root\cimv2 for the run so the remaining assertions execute.

* test(ssh): keep the first terminal session answering keepalives through GC

The hostile-host driver disposed the first session's multiplexer before the GC and reconnect
steps, so a slow Windows GC let the relay reap the silent owner as 'local' and the reconnect then
waited out the full owner grace. Keep the session live until the connection closes, as the app
does, and resend the terminal probe until the shell evaluates it: ConPTY PowerShell can drop
typeahead sent before its first prompt.

* ci(ssh): keep each cell's relay logs in the receipts

* fix(relay): detach an ended socket client as peer-closed before destroying it

The listener destroyed a socket on 'end' but detached its client only on 'close'. A relay write
in that window failed with 'Relay socket is closed', and the dispatcher closed the client as
'local', so its PTY owner kept the full 30s grace instead of the peer-closed floor and a quick
reconnect was refused. The Windows host lanes logged this race on the named-pipe endpoint.

* test(relay): drive the peer-end listener test with a real dispatcher instead of a cast stub

The stub was an unchecked 'as unknown as RelayDispatcher' that failed the changed-code casting gate.

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 04:03:07 -07:00
OrcaWinandm4air 554f7f4ce5 feat(packaging): ship the orcad server template in desktop builds (#24155)
* build(orcad): merge per-runner prebuild slot trees into one matrix

Each node-server lane builds only its own node-pty slot. Release CI needs
their union before `build:orcad-prebuilds --require-slots` and the
template build can run; merge-orcad-prebuilds.mjs verifies every lane's
files against its own manifest, refuses duplicate slots and mismatched
node-pty/N-API/Node-header builds, then writes one merged manifest.

* build(orcad): keep agent-browser out of the desktop deployment template

The template rides inside every desktop build (design D2). Seven ~10 MB
agent-browser binaries would be ~76 MB, more than the rest of the template;
design D2's package contents never listed it, and a slot without one
already reports no headless browser. ORCAD_OMIT_AGENT_BROWSER=1 skips the
copy; standalone build:orcad still includes it.

* feat(packaging): ship the orcad deployment template in desktop builds

Design D2: the server JS and every target's addons ship inside the app,
as out/relay does; the ~120 MB Node runtimes stay excluded and are
downloaded on demand. electron-builder copies out/orcad-template to
Resources/orcad-template on every desktop OS, which is the first path
materializeOrcadArtifact tries (process.resourcesPath).

Platform signing rewrites native bytes the template manifest hashes:
- macOS: the tree is signIgnored (codesign rejects its ELF/PE payloads);
  afterPack signs the darwin targets' Mach-O files with the app identity,
  as notarization requires, then reseals only those manifest entries.
- Windows: SignPath signs after packaging, so release CI reseals from the
  inner-signing list (packaged-orcad-template.cjs --reseal-signed).
Every other file must still match the build's hashes; afterPack verifies.

ORCA_REQUIRE_ORCAD_TEMPLATE=1 makes a missing template fail beforePack and
afterPack; without it a build ships none and SSH relays keep the legacy
path. verify-packaged-orcad-template.test.mjs's "unused, excluded"
contract is reversed on purpose.

* ci(release): build the orcad template from qualified lanes and package it

node-server-tests.yml becomes callable with a ref and build_template.
With build_template, each lane that owns a release slot (macOS, Windows,
the glibc 2.28 and Alpine lanes, and the glibc 2.17 compat lane) uploads
its qualified out/orcad-prebuilds, the Windows lane also uploads both
process-table addons, and desktop_template merges them, gates the full
matrix plus the compat slot with --require-slots, runs
build:orcad-template and uploads the orcad-template artifact.

release-cut calls it at the release tag beside the other gates. The
build and build-mac jobs wait for it, download it into out/orcad-template
(the mac workflow from the parent run), and require it via
ORCA_REQUIRE_ORCAD_TEMPLATE. The Windows signing staging skips the
template's Linux/macOS payloads, and a reseal step records SignPath's
bytes before the installer rebuild. A template-scoped concurrency group
keeps a release call and main's push runs from cancelling each other.

* test(orcad): keep the packaged-lookup imports clear of the compat-slot import edits

* ci(orcad): let a rerun lane replace its template artifacts

upload-artifact v4 refuses a second upload under an existing name in the same
run, so rerunning a flaky node-server lane during a release would fail at the
upload instead of re-qualifying the slot.

* ci(node-server): build the template's Windows addons before the lane switches to Node 18

The addon build script imports TypeScript, which Node 18 cannot load, so every
build_template run (release-cut included) failed on windows-2022.

* fix(build): ship the orcad template's shared node_modules

electron-builder's extraResources filter always drops the root node_modules of
a source directory, so packaged apps lost orcad-template/node_modules and the
afterPack verify failed. Copy it through its own resource entry.

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 04:01:26 -07:00
Jinwoo Hong c422936a71 fix(relay): anchor same-cap monitor evidence freshness to the run's authorisation, not job startup (#24349)
* fix(relay): anchor same-cap monitor evidence freshness to the run's authorisation, not job startup

The same-cap gate now verifies the dry-run on its own clock and records the
authorisation instant in the single-use consumed marker; each cell job checks
the evidence age at that instant and bounds its own start after it.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): refuse a re-run same-cap gate before it consumes evidence; tighten order tests

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-10-01 06:52:31 -04:00
14d4bb2e2a fix(ssh): Windows hosts without Add-Type staging; runtime-store GC on Windows (#24149)
* fix(ssh): collect the pinned-Node runtime store on Windows hosts

Windows SSH hosts now run runtime-store GC instead of skipping it: one
PowerShell inventory reads .runtime-ref-node-<sha> and .runtime-node refs from
every version dir, and one Get-CimInstance Win32_Process query filtered on an
image path under runtimes\ adds process holds (never by image name; a failed
query keeps everything). Stale upload stages are swept with the same rule as
POSIX. Promotion and the post-upload hold check now take the store lock on
Windows too, and the lock's own commands run unwrapped there.

Windows relay version-dir liveness now honours .relay-pid (design D5): a live
PID answers ALIVE before any pipe is touched, a dead one (ESRCH) plus refusing
pipes is exited, anything else is unverifiable. The runtime probe adopts a
pinned node.exe an earlier vault reader left without a .verified marker after
running it.

* fix(ssh): Windows stage fencing and vault runtime go through the verified node.exe

Upload-stage file identity on Windows no longer compiles an Add-Type P/Invoke
helper when the relay runs on Orca's verified pinned node.exe: the stage
commands run a fixed fs.lstatSync(..., {bigint:true}) script through it. It
prints the legacy helper's vol:high:low lowercase hex, and identity files are
compared after normalising hex spelling, so old and new clients recover each
other's stages. Host-Node relays keep the legacy helper; the choice is
documented in windows-edr-posture.md.

The Windows OpenCode vault reader now installs the pinned runtime through
ensureRemoteOrcadNodeRuntime (official zip, host-side extraction, .verified,
store lock) instead of uploading a client-extracted node.exe, and the relay dir
gains a .runtime-ref-node-<sha> so store GC keeps the runtime the vault uses.

* test(ssh): run the Windows stage-identity and store-GC tests on the Windows lane

The legacy/node.exe identity compatibility test and the Win32_Process hold path
were gated to win32 but no CI lane ran them. Add both files to the Windows
package lane and a real running-node.exe hold test.

* test(ssh): tear down Windows-lane temp trees through removeTreeSync

* test(ssh): grant the store lock to the Windows OpenCode runtime setup test

The Windows promote now runs under runtimes/.store-lock, so the mocked host
must answer the lock's CreateNew step.

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>
2026-10-01 03:25:42 -07:00
OrcaWinandm4air 6aed05471c ci(ssh): hostile-host matrix for the relay runtime ladder (#24146)
* fix(ssh): classify a musl host missing libstdc++ as missing_lib, not wrong_libc

musl's loader follows each missing-library line with one 'Error relocating ... symbol
not found' per unresolved symbol, and the relocation pattern was checked first. Check
missing libraries before relocation errors; the ld-linux/ld-musl interpreter case stays
wrong_libc.

* build(orcad): allow a partial deployment template for CI

build-orcad-template --targets a,b builds and verifies only the named slots, so a CI job
that can fill just the x64 Linux prebuild slots can still materialize rung A/C addons.
Without the flag every target is still built and verified.

* ci(ssh): hostile-host matrix for the relay runtime ladder

Drives the real client-side relay deploy against Docker sshd targets and asserts the
design D6 rung each lands on: Debian 10 and AlmaLinux 8 (glibc 2.28) and Alpine (musl)
on rung A; Alpine without libstdc++ refused missing_lib down to D; Ubuntu 22.04 with a
host Node 20 and a noexec home straight to D (home_noexec); CentOS 7 (glibc 2.17)
refused libc_floor at A and C, falling to a host-npm path with no Node; and a
no-egress Debian 10 still on rung A. Launched cells also prove the terminal echoes,
no npm or compiler ran, a second connect reuses the uploaded runtime, and runtime GC
keeps the in-use runtime while collecting an idle one.

New workflow ssh-hostile-hosts.yml runs on dispatch and on path-filtered PRs.

* test(ci): pin the hostile-host workflow to the headless-server builder images

The matrix builds its runtime slots in copies of the node-server lanes' Alpine
and manylinux images; this contract fails when NODE_RUNTIME_PIN or either
builder digest moves in one workflow and not the other.

* test(ssh): reconnect as the same client and retry a grace-held PTY owner in the hostile-host matrix

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 03:24:17 -07:00
Neil bd90da7a5b ci: share PR planning setup and reuse the static native cache (#24329) 2026-10-01 02:46:18 -07:00
Jinwoo Hong 9ed7b39b3c fix(relay): run the same-cap headroom gate in the modes the job actually receives (#24343)
The parent workflow collapses canary-apply and batch-apply into the job mode
apply, so the headroom step's canary-apply/batch-apply condition never held and
the gate was skipped on every real roll. Run it wherever the drain runs (apply,
rollback before its restart) and in read-only verify; skip only a resumed
rollback, which drains nothing. A new workflow-shape test fails on any job
step comparing against a mode the parent cannot pass, and on a drain that can
run without the headroom check.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-10-01 05:18:42 -04:00
OrcaWinandm4air ddd4927a0b build(orcad): server node-pty slots at glibc 2.28, plus a glibc 2.17 compat slot (#24134)
* build(orcad): build server glibc slots on glibc 2.28 and add the glibc 2.17 compat slot

Design D6: the default linux-{x64,arm64}-glibc node-pty slots now build in
manylinux_2_28 (digest-pinned) and are gated at glibc 2.28 / GLIBCXX_3.4.25
through a floor profile on verify-linux-glibc-floor.cjs; the desktop keeps
its Ubuntu 20.04 (2.31) default.

Adds the opt-in linux-x64-glibc217 compat target: NODE_RUNTIME_COMPAT_ASSETS
pins the unofficial glibc-217 Node (update/check pin scripts cover it, outside
SERVER_TARGETS), and a new CI lane builds the compat slot in manylinux2014
with static libstdc++, gates it at glibc 2.17 with no shared C++ runtime in
DT_NEEDED, and smokes it under the glibc-217 Node.

* refactor(node-runtime-pin): route compat lookups through isCompatServerTarget; keep the glibc doc's slot-name paragraph intact

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 01:10:57 -07:00
OrcaWinandm4air 6593d7d194 feat(orcad): run orcad on the pinned Node instead of Bun (#24110)
* ci(daemon): gate PRs on daemon protocol crossing from the newest release

Lands daemon-protocol-facts.mjs from the Windows update diagnostic branch with a
stricter parser, and adds check-daemon-protocol-crossing.mjs (rule R1): the working
tree must attach the newest release tag's daemon. Rollback crossing is reported only.
Runs in the cross-version-wire job, which already has full tags; tag selection moves
to config/scripts/stable-release-tags.mjs so both use one rule.

* feat(persistence): run profile backups in the worker whenever its entry is bundled

* refactor(orcad): make profile and native preflight runtime-neutral

The profile preflight parser now takes the expected runtime identity from the
caller (shipped callers pass the pinned Bun identity), and the native
preflight is renamed to orcad-runtime-native-preflight with neutral wording.

* feat(runtime): pin the Node 24.21.0 server runtime with an offline CI check

Add src/shared/node-runtime-pin.ts (NODE_RUNTIME_PIN, SERVER_TARGETS,
NODE_RUNTIME_ASSETS for all 8 server targets plus the headers tarball),
generated by config/scripts/update-node-runtime-pin.mjs from the nodejs.org
and unofficial-builds SHASUMS. check-node-runtime-pin.mjs verifies, with no
network, that the pin tracks the locked Electron, matches engines.node's
major, and covers exactly SERVER_TARGETS; it runs in the static analysis job.

ORCAD_BUN_TARGETS consumers now read SERVER_TARGETS so there is one target
list; orcad's Bun runtime and build output are unchanged.

* test(persistence): skip plain-Node backup selection tests in the Bun profile suite

* fix(runtime): reject a pinned archive that belongs to another target

* ci(daemon): fail PRs that swap a runtime launcher and bump the daemon protocol

D7.1 R3: hosting orcad or the daemon on another runtime is not a protocol change,
so one PR must not do both. The launcher file list lives in the check script; the
allow-runtime-launcher-protocol-bump label overrides it.

* feat(orcad): select pinned-Node slots by a .runtime-node marker

D7.1 R5: a Node slot names its shared runtimes/node-<sha256>/node through
.runtime-node instead of .build-target, so Bun-era clients read it as a legacy
slot rather than exiting 78 on a missing bun-runtime. Nothing builds the marker yet.

* fix(runtime): load the Node pin without the typeless-module warning

check-node-runtime-pin.mjs now requires the pin and takes nodeDistArchiveName from
its own module, so it no longer loads the update script's build graph.

* fix(orcad): resolve Node slots to the design's runtimes/node-<sha>/bin/node layout

* feat(orcad): 8-slot node-pty prebuilds against the pinned Node headers at N-API 8

- build-orcad-prebuilds.mjs adds win32-x64/arm64 (conpty.node, the vendored
  conpty.dll/OpenConsole.exe, upstream's N-API conpty_console_list.node), compiles
  in a scratch copy against the hash-verified pinned headers (node.lib pinned per
  Windows arch) with NAPI_VERSION=8, rejects post-8 node_api_* imports, and writes
  a schema 2 manifest with per-file sha256, N-API level and the glibc need.
- --require-slots [slots] verifies files against hashes; --smoke loads the slot
  under the pinned Node and spawns a PTY; --print-slot names the host slot.
- The slot installer gates on N-API, libc, arch, glibc and file hashes instead of
  the exact NODE_MODULE_VERSION, and installs nested files (conpty/).
- bun-profile-tests.yml builds, verifies and smokes each runner's slot.

* fix(orcad): scope node-pty's glibc .symver pins to glibc on musl prebuild slots

musl's unversioned libc cannot satisfy openpty@GLIBC_* references at link
time, so the Alpine slot compile would fail. Pin the staged pty.cc guard to
__GLIBC__ and assert both musl transforms against the installed patch.

* feat(orcad): run orcad on the pinned Node instead of Bun

A packaged orcad slot now references the pinned Node 24.21.0 by its
executableSha256 (`.runtime-node`, `.server-target`) instead of carrying
bun-runtime, and ships node-pty from the slot's prebuild, only its own
ripgrep, and no Windows Bun PTY gate. The runtime lives beside the slots
at runtimes/node-<sha>/bin/node (node.exe on Windows, upstream name).

- build:orcad (build-orcad-node.mjs) builds the host slot's prebuild when
  missing and places the pinned runtime; the template is schema 3 with
  per-target files.
- handoffToBundledOrcad() resolves the slot's runtime reference and checks
  process.versions.node against the pin; a host Node >= 18 still hands off.
  Startup preflight keys on running as that runtime; callers expect 'node'.
- orcad and its daemon use node-pty (ConPTY + windows-pty-job on Windows);
  the Bun PTY sources, gate entry and canUseBunPty branches are removed.
- SSH deploy uploads the official archive once per pin, extracts and
  hash-checks it on the host, and self-tests it before publishing. Bun
  slots stay launchable for rollback; Node slots never use host Node.
- The runtime materializer is generic over pinned assets; the Bun wrapper
  remains only for the OpenCode vault reader (design Phase 2).
- Cross-runtime test: a profile DB written by Bun 1.4.2 (WAL left by
  SIGKILL) opens and backs up under the pinned Node, and the reverse.

No daemon PROTOCOL_VERSION change (design D7.1 R3).

* docs(ci): name the headless lanes after the pinned Node, drop Bun shard timings

Design D10: ci-demand-rollout.md and ci-runner-efficiency.md follow the
bun-profile-tests.yml -> node-server-tests.yml rename; shard timings drop the
deleted Bun PTY tests and follow the renamed ones.

* chore(ci): count the runtime archive download as a runtime launcher path

* fix(orcad): pin the macOS C++ standard for node-pty prebuilds

The official Node headers' config.gypi sets clang: 0, so common.gypi skips its
gnu++20 xcode_settings and Apple clang 15 (macos-14 runners) compiles
node-addon-api as C++98.

* fix(orcad): resolve the preflight's slot through realpath, as the handoff does

A symlinked orcad.js handed off to its real slot's pinned Node, but the
startup and profile preflights read the symlink's directory, found no
runtime marker there, and silently skipped the readiness check.

* refactor(ssh): drop materializeCachedNodeRuntime, which nothing calls

Deploys upload the verified official archive (design D5); no client path
needs an extracted Node executable cached by digest.

* test(orcad): gate the Bun-to-Node upgrade and Node-to-Bun rollback with live terminals

Design D7.1 R1/R3/R4 and D7.2. The last Bun orcad and this checkout's Node
slot are installed side by side under ~/.orca-remote, launched and stopped
with the client's own deploy commands, and share one data root. Each
direction proves the incoming orcad adopts the outgoing runtime's daemon
(same PID, same shell, output continues), opens its profile database and
backs it up with its own shipped worker, and that GC keeps the slot the
live daemon was forked from.

The node-server Linux lanes provide Bun 1.4.2 and build that Bun orcad
from main, and run with --cross-runtime. --artifact and --cross-runtime
now make their tests fail on a missing input instead of skipping.

* ci(node-server): pin node:24.21.0-alpine by its multi-arch index digest

* test(ssh): name the runtime archive fixture after its role

* test(node-server): load node-pty from the packaged slot in artifact runs

The node-server lane installs dependencies without building node-pty, and
Linux has no upstream prebuild, so the real-PTY failed-I/O teardown test
(picked up by the pty-subprocess selector) could not load pty.node. In
--artifact runs, alias node-pty to out/orcad's shipped slot so the test
exercises the addon orcad actually runs under the pinned Node.

* fix(orcad): let the Windows profile preflight exit after its PTY probe

On Windows, node-pty keeps the conout worker thread and pseudoconsole alive
until kill(), even after the shell exits. The PTY health probe never killed a
cleanly exited probe, so the packaged preflight printed its readiness line
and then hung until the build's 30s timeout, reported with an empty stderr.

- The probe kills its PTY on Windows after exit and uses the bundled ConPTY
  the daemon spawns with.
- The preflight exits once stdout is flushed; its owner reads to EOF.
- Preflight failures now report code, signal, timeout, stdout and stderr.

* test(node-server): load the slot's node-pty in the real-PTY test, not by alias

A vite alias redirected only ESM imports of node-pty; windows-pty-job and
local-pty-utils resolve it through require, so Windows loaded two conpty.node
copies and the Git Bash job-membership proof read an empty job. The failed-I/O
teardown test now loads node-pty through a fixture that picks the packaged slot
in artifact lanes.

The pty-subprocess selector was a prefix that also pulled in its POSIX-host
sibling unit tests, which pr.yml runs and which were never qualified on
Windows. Select the directory plus the two sibling files that belong here.

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 00:39:00 -07:00
Jinwoo Hong 3e5c8d9f8e feat(relay): declare Asia cell c31 at the c30 shape (#24310)
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-10-01 03:11:32 -04:00
OrcaWinandm4air d2dfc79764 ci(daemon): runtime-launcher protocol ratchet and Node slot marker (#24108)
* ci(daemon): gate PRs on daemon protocol crossing from the newest release

Lands daemon-protocol-facts.mjs from the Windows update diagnostic branch with a
stricter parser, and adds check-daemon-protocol-crossing.mjs (rule R1): the working
tree must attach the newest release tag's daemon. Rollback crossing is reported only.
Runs in the cross-version-wire job, which already has full tags; tag selection moves
to config/scripts/stable-release-tags.mjs so both use one rule.

* feat(persistence): run profile backups in the worker whenever its entry is bundled

* refactor(orcad): make profile and native preflight runtime-neutral

The profile preflight parser now takes the expected runtime identity from the
caller (shipped callers pass the pinned Bun identity), and the native
preflight is renamed to orcad-runtime-native-preflight with neutral wording.

* feat(runtime): pin the Node 24.21.0 server runtime with an offline CI check

Add src/shared/node-runtime-pin.ts (NODE_RUNTIME_PIN, SERVER_TARGETS,
NODE_RUNTIME_ASSETS for all 8 server targets plus the headers tarball),
generated by config/scripts/update-node-runtime-pin.mjs from the nodejs.org
and unofficial-builds SHASUMS. check-node-runtime-pin.mjs verifies, with no
network, that the pin tracks the locked Electron, matches engines.node's
major, and covers exactly SERVER_TARGETS; it runs in the static analysis job.

ORCAD_BUN_TARGETS consumers now read SERVER_TARGETS so there is one target
list; orcad's Bun runtime and build output are unchanged.

* test(persistence): skip plain-Node backup selection tests in the Bun profile suite

* fix(runtime): reject a pinned archive that belongs to another target

* ci(daemon): fail PRs that swap a runtime launcher and bump the daemon protocol

D7.1 R3: hosting orcad or the daemon on another runtime is not a protocol change,
so one PR must not do both. The launcher file list lives in the check script; the
allow-runtime-launcher-protocol-bump label overrides it.

* feat(orcad): select pinned-Node slots by a .runtime-node marker

D7.1 R5: a Node slot names its shared runtimes/node-<sha256>/node through
.runtime-node instead of .build-target, so Bun-era clients read it as a legacy
slot rather than exiting 78 on a missing bun-runtime. Nothing builds the marker yet.

* fix(runtime): load the Node pin without the typeless-module warning

check-node-runtime-pin.mjs now requires the pin and takes nodeDistArchiveName from
its own module, so it no longer loads the update script's build graph.

* fix(orcad): resolve Node slots to the design's runtimes/node-<sha>/bin/node layout

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 00:10:57 -07:00
OrcaWinandm4air 2a83c9536f ci(daemon): gate PRs on daemon protocol crossing from the newest release (#24089)
Lands daemon-protocol-facts.mjs from the Windows update diagnostic branch with a
stricter parser, and adds check-daemon-protocol-crossing.mjs (rule R1): the working
tree must attach the newest release tag's daemon. Rollback crossing is reported only.
Runs in the cross-version-wire job, which already has full tags; tag selection moves
to config/scripts/stable-release-tags.mjs so both use one rule.

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-30 23:23:28 -07:00
OrcaWinandm4air 3135fbbf49 feat(runtime): pin the Node 24.21.0 server runtime with an offline CI check (#24087)
* feat(runtime): pin the Node 24.21.0 server runtime with an offline CI check

Add src/shared/node-runtime-pin.ts (NODE_RUNTIME_PIN, SERVER_TARGETS,
NODE_RUNTIME_ASSETS for all 8 server targets plus the headers tarball),
generated by config/scripts/update-node-runtime-pin.mjs from the nodejs.org
and unofficial-builds SHASUMS. check-node-runtime-pin.mjs verifies, with no
network, that the pin tracks the locked Electron, matches engines.node's
major, and covers exactly SERVER_TARGETS; it runs in the static analysis job.

ORCAD_BUN_TARGETS consumers now read SERVER_TARGETS so there is one target
list; orcad's Bun runtime and build output are unchanged.

* fix(runtime): reject a pinned archive that belongs to another target

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-30 22:57:10 -07:00
Jinwoo Hong 7c119465b0 fix(relay): define restart-safe by the cell runtime, and refuse waves without headroom (#24259)
* fix(relay): let a same-cap drain finish when only unplaceable hosts remain

The c28 canary on 2026-10-01 drained the cell to zero live connections, but four
hosts with no free slot anywhere kept redialling and held director leases on it,
so the restart-safe wait timed out and left the cell isolated and empty.

The drain wait now also passes once the runtime has carried nothing for a
sustained quiet window while a small, capped number of leases remain, and logs
the escape. Apply modes also refuse a cell whose hosts exceed 80% of the free
slots on the other general cells, so a wave cannot strand hosts in the first place.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): define restart-safe by the cell runtime, not director leases

Replaces the opt-in stranded-host escape with a corrected definition. A
restart is safe when the cell runtime carries nothing live and no migration
is open, sustained for the drain pace window. Director activity leases lag
hosts that already left or cannot be placed, so they are reported in a
progress line and the verified result instead of blocking the restart.

The same-cap drain passes its existing pace window. The headroom script is
added to the trusted evidence code paths with the other production scripts.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): print stranded director counts on every restart-safe sample

Each restart-safe poll now prints its sample count and the director's
restart-blocking leases, request units, reserved remainder, and migrations
under `stranded`; the verified line carries the same object. Open migrations
still block because each is pinned to the cell incarnation a restart replaces.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): require the pace window for every live restart-safe wait

Pre-auth and total connections no longer reset the restart-safe window:
on drained c28 they flickered with unauthenticated redials in a third of
samples, which a restart does not lose. They stay in the progress output.

Every live restart-safe call must now pass --pace-window-ms. The capacity
job and staging proof drain unpaced, so they pass the production 300000 ms
window, and the calls that relied on the 180000 ms default get 480000 ms.

Headroom free slots now follow the director's placement rule: the admission
pause minus the larger of observed and enforced units, minus outstanding
control reservations.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-09-30 22:25:36 -04:00
Brennan Benson 5cda0f4508 refactor(native-chat): keep agent-session records in the chat journal database (#24006)
* fix(native-chat): report a failed startup chat reconcile instead of failing app startup

At startup the chat host re-checks every saved chat's lease and writes the
result to agent-sessions.json. If that write failed (the file lock gave up,
the file could not be written, or the file was written by a newer Orca and
is read-only here), reconcileRestartLeases rejected, the startup IPC call
rejected, and the renderer fell into its degraded "Session restore failed.
Changes won't be saved until restart" mode.

The reconcile is bookkeeping: a lease left unreconciled grants no writer,
and every attach, send and read of a chat reconciles its own lease again.
So the startup reconcile now reports its failure through a new optional
host dependency, onStartupReconcileFailure, and resolves. The runtime
routes it to its onError sink under the scope
structured-agent-session-startup-reconcile, or logs it when no sink is
installed (the desktop installs none).

* fix(native-chat): read restored chats without waiting on lease bookkeeping

With native chat on and a chat tab open at quit, the renderer's startup
also awaits the chat tab restore (session.tabs.listAll). That restore
re-ran the lease reconcile before reading each chat and rethrew its store
failure, then recorded each restored tab as visible through a store
transaction that throws on a held lock or a read-only store. Either one
failed the restore, so startup still fell into "Session restore failed".

Reading a chat grants no writer, so the reconcile startup and the restore
run is now a reader's: createReaderReconcile never throws, answers whether
every lease is settled (recovery is resolved only then; the journal opens
either way), and reports each distinct failure once until a reconcile
settles. Attach and agent start keep the strict reconcile. The restore's
tab republish logs a failed visibility write and still publishes the tab,
since a client drops every unpublished chat tab; user-driven publishes
still refuse.

The host dependency is renamed onLeaseReconcileFailure (scope
structured-agent-session-lease-reconcile), since it now also reports for
reads.

* fix(native-chat): keep every record-store write off the startup chat read path

Round-2 review found two more writes on the startup chat restore that
could still fail it and put the app into "Session restore failed":
republishing a /clear replacement recorded its tab visibility strictly,
and resolving a chat's recovery rethrew its store error. The restore
also paid one lock wait per tab and per batch of chats while the lock
stayed held.

The restore now derives tabs from state it already holds:
- publishStructuredAgentSessionTab splits into the strict write and
  projectStructuredAgentSessionTab, which only updates the runtime's
  snapshot. The restore and /clear replacements only project: a saved
  tab index already lists every restored chat, and a /clear moves the
  tab in the same write that commits it. visibilityWriteMayFail is gone.
- Chats a legacy profile restores that the index does not list are
  recorded in one best-effort transaction (store.showSessionTabs), so a
  failure leaves the index absent to seed again rather than partial.
- The read restore's recovery resolution is caught and reported through
  onLeaseReconcileFailure, deduplicated with the reconcile's reports.
- Once lease bookkeeping fails in a restore pass, the rest of that pass
  skips it, so a held lock costs one wait for the startup reconcile and
  one for the restore, however many chats are open.

User actions (create, reveal, attach, send, the /clear commit) keep
their strict writes.

* test: open, seed and read the agent-session record store through one harness

Tests that open the durable agent-session record store, seed it, or read
back what it persisted now go through agent-session-record-store-test-harness.ts
instead of calling AgentSessionRecordStore.open or touching agent-sessions.json
themselves. A later change that moves the store into the chat database then
changes the harness instead of every test. No production code changes.

Tests whose subject is the JSON file itself (its .bak recovery, salvage,
schema versions, permissions, and what older builds read back) keep reading
and writing the file directly; the storage move rewrites or deletes them.

* fix(native-chat): start each restore pass from one lease check and stop its bookkeeping at the first failure

The restore now runs one reader lease check for the pass and lets each chat
re-check and resolve recovery only while the pass is still settled. The
first refusal or failed write clears it for the rest of the pass, and every
chat is still opened for reading. With another process holding the lock,
startup waits on it once in prepare and once in the restore, however many
chats are open; a legacy profile waits once more for its tab-index seed.

* docs(native-chat): correct restore comments and a test name to match the final design

* test: address the record-store harness by the host's state directory

The harness took the store's own folder, so each caller picked one
(join(root, 'store'), or 'agent-sessions' where a test read the store the
runtime owns). A later change that moves the store into the state
directory's journal database could not tell those apart, and would have
had to edit every caller again.

Every harness function now takes the state directory, the one the test's
journal database and recovery capsule already live in, and keeps the
store in the same subfolder the runtime uses. Callers pass that directory;
store-only tests pass their temp directory unchanged. Format tests that
share a directory with harness calls take the file path from
testAgentSessionStoreFilePath.

The folder name moves from a private constant in the runtime to
AGENT_SESSION_STORE_DIR_NAME beside the store's file name, so the harness
shares it without importing the runtime. Its value and every path built
from it are unchanged.

* refactor(native-chat): keep agent-session records in the chat journal database

The record store's records, operation ledger, retired claim keys and chat tab
index become tables in agent-session-journal.db (user_version 4). The version-4
migration copies agent-sessions.json in its own transaction and never writes,
renames or deletes that file or its .bak. Each store write is one journal
transaction over exactly the rows it changed, checked with the load rules; the
file lock, the external-change refresh and its hash, the .bak rotation, salvage
and the hot-path recovery fence are gone from the store.

* wip: importer tests

* test(native-chat): cover the records migration, the import, row writes and read-only records

* docs(native-chat): retire comments that describe the records file as the live store

* test(native-chat): drop the record-store harness's leftover file path and type the import fixture

* test(native-chat): let the host harness cleanup wait out a recovery-offer read's lock

* fix(native-chat): let Stop reach the agent when its ledger row cannot be written

Stop's operation-ledger row now shares the database with the chat history, so
damage, a full disk or a stranded transaction on that write refused the Stop
before the interrupt. A cancel plan now takes its decision from the committed
ledger in memory, runs without settling, and warns that the row was skipped.
Other mutations answer proven damage with the typed "Unable to load this chat."
refusal instead of the raw SQLite error.

* fix(native-chat): answer whether a profile holds chats from the database's rows

Every host install creates agent-session-journal.db, chats or not, and the
version probe created it too, so its mere existence made every profile that
ever installed the host wait on host install and reconcile at startup. The
check now opens the database read-only and looks for a record or tab row,
lets the records file answer while its import is still owed, and counts an
unreadable database as present. The version probe no longer creates the file.

* fix(native-chat): open a chat from history when its tab index cannot be written

Over records a newer Orca wrote, every write is refused, so opening a closed
chat from Agent Session History failed on the tab-visibility write and the chat
read as unreachable. Like closing a tab, opening one now reports a failed
restore-index write and still publishes the tab.

* fix(native-chat): keep the records import owed when the backup read fails transiently

A torn records file whose .bak could not be read (EACCES, EIO) was reported as
unusable, so the migration completed with nothing copied and never retried.
A non-ENOENT read failure of either copy now carries its cause, which the
importer classifies as a read that can clear.

* test(native-chat): pin that an unreadable records file never falls back to its backup

* fix(native-chat): restore imported chats' tabs when the records file had no tab index

A chat created while the import was owed recorded a tab index holding only
itself. When the file it later imported had no index, that index still read
as recorded, so the imported chats' tabs never came back. The import now
clears the recorded marker in that case, and restore falls back to the
profile's tabs.

* refactor(native-chat): drop the unused in-transaction store write

Nothing called it, and it bypassed the write queue and the read-only refusal.

* docs(native-chat): say that an unusable records file is left untouched but never re-imported

* refactor(native-chat): keep the provider handle chain check as main has it

The chain-validation refactor has no measured need in this change.

* docs(native-chat): retire lease-renewer comments that describe the records file as the live store

* fix(native-chat): keep a throwing failure sink from failing the startup chat read

The lease bookkeeping failure reporter called the host's failure sink
directly, so a sink that threw turned a reported, recoverable store failure
back into a rejected startup reconcile or read restore. The reporter now
catches a sink throw and logs both the original failure and the sink error
with console.warn.

* test(native-chat): wait for a replaced host's restart-offer writes before cleanup

A restart test replaces the host without tearing the old one down, so the old
host's fire-and-forget restart-offer withdrawal could still hold the recovery
capsule's lock directory when cleanup removed the test directory (ENOTEMPTY).
The harness now hands hosts a capsule that tracks running operations and waits
for them before removing the directory, replacing the rm retries.

* docs(native-chat): retire the abandon helper's note that the store re-creates its directory

* fix(native-chat): restore a chat opened while the import was owed beside the profile's chats

When the imported records file had no tab index, restore fell back to the
profile's saved tabs, which never list a Claude chat, and the seed then
rewrote the tab table without the chat opened while the import was owed.
The tab rows that chat left are now loaded as unrecorded, restore takes them
together with the profile's chats, and the seed keeps their tab ids.

* test(native-chat): pin that a create whose tab index write fails still opens the chat

* docs(native-chat): say why restore puts chats opened while the import was owed first

* test(native-chat): replace a ledger row rather than change it in place in the Send-now rerun test

The record store freezes published rows in tests, so setting a row's outcome
in place threw; the test now swaps in a changed copy, as its sibling cases do.
2026-09-30 14:45:32 -07:00
Brennan Benson cfa43e7eab fix(codex): opening a terminal no longer strips Codex hooks from the real ~/.codex (#23552)
* fix(codex): a real-home restore leaves a file alone once someone else changed it

Orca writes ~/.codex/hooks.json (and a trust rebase writes config.toml), then
runs a Codex trust session for up to 10 s, then restores the original bytes if
the session fails. The restore wrote unconditionally, so a save that landed
during the session, from the user or another Orca, was silently reverted.

Each restore now compares first: it writes the original back only while the
file still holds the generation Orca's mutation left, and otherwise logs and
leaves it alone. This covers the real-home install and opt-out sweep
(restoreRealHomeHooksJson), the legacy sweep's hooks restore, and config.toml
rollback (restoreCodexTrustConfig).

For hooks.json the generation is the exact bytes Orca wrote. For a config.toml
that a trust rebase changed it is the file as the rebase left it. When Codex
itself wrote config.toml inside the session that just failed, Orca never knew
those bytes, so that rollback compares against the file as the session settled.
The next commit keeps other Orca instances out of that window; a user edit made
during such a session can still be rolled back.

* fix(codex): serialize real-home Codex writes across Orca instances

Every Orca on one HOME (a dev and a packaged app, or an offline CLI) writes the
same ~/.codex/hooks.json, config.toml and ~/.orca/agent-hooks/codex-hook.sh.
The per-file lane that orders capture, mutate and restore was in-process only,
so another instance could write inside this one's restore window, or undo it.

The lane for the user's real config.toml now also holds the existing
crash-safe managed-hook install lock (~/.orca/managed-hook-install.lock, the
one relay installers take for the same home). It is taken only by the
outermost acquire, because the lock file is not reentrant and grants and trust
rebases nest inside an install. Managed-home installs, the real-home install
and opt-out sweep, and the legacy sweep all enter through it. Compare-and-swap
on restore stays as the backstop.

A lock that cannot be taken within its 10 s wait fails that install, which is
already best effort: launch prep logs it, and the real-home lane falls back to
the managed lane until its retry.

* fix(codex): opening a terminal no longer strips the shared Codex entry from ~/.codex

Every Orca instance on one HOME writes the same status-hook entry into the
user's ~/.codex/hooks.json, with its trust in config.toml. Launch prep runs on
every pane spawn, and under a managed Codex account it ran the legacy system
sweep. That sweep matched Orca entries by script file name, so it removed the
current shared entry and the trust blocks the grant ledger recorded. On a live
laptop hooks.json went 4139 -> 18 bytes about 150 ms before a new pane opened.
With hooks off, the real-home lane's launch prep swept the same way.

Now nothing automatic removes the current entry or its trust:
- The legacy sweep removes only an enumerated list of retired command forms
  that no build writes any more (#1019's double-quoted form, #1536's
  exec-guarded form, and Windows' per-userData bare path), plus their trust.
- ensureRealHomeCodexHookState with hooks off writes nothing; that covers
  launch prep, session resume and startup.
- Only the user's explicit opt-out (codexHookService.remove()) strips the entry
  and its ledger-recorded trust from the real home.
- The sweep-suppression gate existed only to stop the sweep from deleting the
  current entry, so it is deleted with its main-process wiring.

Startup with hooks off already skipped the real-home install; with this change
the first pane's launch prep with hooks off also leaves ~/.codex untouched.

* fix(codex): a pane's prepare-codex only repairs a home its own HOME's app installed

On macOS a pane starts through login(1), so it gets the user's real HOME even
when its Orca app runs with another one. The pane's `codex()` preflight
installed hooks in the CLI process with that real HOME: it rewrote
~/.orca/agent-hooks/codex-hook.sh, promoted trust into the real config.toml,
and wrote the real HOME's script path into the app's managed home.

The preflight now acts only when the managed home's hooks already run this
process's own shared script, which proves the app that installed them shares
its HOME. Otherwise it writes nothing; the app installed the home at spawn.

Why not a no-op: the preflight was added (#14326) because trust can go stale
between opening a pane and typing `codex`, for example in a pane that survives
an app update, and Codex then stops in hook review. For a same-HOME pane it
still repairs that. Why keep promotion: the install drops runtime trust the
system config does not back, so skipping promotion would delete approvals the
user gave inside Orca-launched Codex.

* test(agent-hooks): await every installer in the refresher coverage test

The test fired each managed installer without awaiting it and read
~/.orca/agent-hooks straight after. Codex's install now takes the
cross-process real-home lock before it writes its script, so the script
landed after the read. Await the installers, and stub Codex's trust sessions
so the awaited install cannot start a real `codex app-server`.

* fix(codex): retire the two real-home command forms the list missed

The real-home lane wrote two Codex hook forms into ~/.codex that no build
writes any more and that the enumerated retired list did not name:
- POSIX, #9501 until #10885: the file-guarded form draining with a bare `cat`.
- Windows, #9501 until #10221 took Windows off the real-home lane: the encoded
  PowerShell launcher for a non-cmd-safe script path.

The file-name sweep removed both before; the enumerated sweep left them in
place, trusted, still passing the script's exit status to Codex. Both now
match as frozen literals.

Also corrects the startup ordering comment: the real-home install runs first
so its in-slot upgrade lands before the managed install's sweep retires the
prior command; nothing re-arms a legacy sweep any more.

* fix(codex): take the real-home lock only when a write is needed

The previous commit made every entry to the real-home config lane take the
cross-process lock. That lane runs on every pane spawn and every typed
`codex` preflight, so the steady state paid an owner probe (a `ps` spawn on
macOS) and could wait up to 10 s behind another instance's trust session,
even though it wrote nothing.

Each real-home writer now compares the desired state with the files on disk
first, without the lock. Only when a write is needed does it take the lock,
re-read and recheck, then write:
- real-home install: the planned hooks.json, the shared script and the
  ledger-recorded grant are compared; the locked path re-plans from disk.
- legacy sweep: locks only when a retired entry is present; the sweep re-reads.
- approval promotion: locks only when there is something to promote; the
  promotions are recomputed under the lock.
- the shared ~/.orca/agent-hooks script: locks only when its bytes differ.
The explicit opt-out always takes the lock. The lock is reentrant through
async context, since grants and rebases nest inside an install, so the
config-lane option the previous commit added is removed.

* fix(codex): a shared script without its exec bit is not the steady state

The compare-first check matched the shared ~/.orca/agent-hooks script on bytes
alone. writeManagedScript also restores 0755 on every call, and the POSIX hook
guard skips a script that is not executable, so a script whose mode was lost
(a dotfiles restore, a plain copy) now stayed that way: every Codex hook
drained stdin and reported nothing until an app restart refreshed the script.

The check now also requires the mode the writer sets, so that case takes the
lock and the write path repairs it.

* test(codex): the retired encoded launcher never matches today's shared one

The shared encoded Windows launcher is still current for other agents, so the
comment claiming today's launcher is never encoded was wrong. What keeps the
retired matcher off it is the exact payload: since #14825 the shared launcher
prefixes its payload and drops -ExecutionPolicy Bypass. Pin that with a case.

* fix(codex): the pane step recognises its own script under a home path with an apostrophe

The same-HOME check looked for the script path wrapped in bare single quotes,
but both hook writers escape an apostrophe inside the quotes. A home such as
C:\Users\O'Brien never matched, so the pane-step repair never ran there.

* fix(codex): the trust-RPC escape hatch still keeps the real home off its lane

The no-write check reported a recorded grant as current, so with
ORCA_DISABLE_CODEX_TRUST_RPC set the real-home lane stayed in use. The grant
itself refuses before reading its ledger; the check now does the same.

* fix(codex): the shared script write no longer waits on the real-home lock

The write is atomic and skips identical bytes; waiting behind another
instance's trust session could only fail a pane's managed-home install.

* fix(codex): an in-Orca approval survives a launch that cannot get the real-home lock

The install drops runtime trust the system config does not back, so a
promotion skipped for want of the lock lost the approval for good. It
now writes unlocked, as it did before the lock existed.

* refactor(codex): take the cross-process real-home lock back out

The lock fixed no observed failure. The three that were observed each have
their own fix in this series: the legacy sweep matches only frozen retired
command forms, hooks-off launch prep writes nothing, and a pane's
prepare-codex repairs only a home its own HOME's app installed. The lock
instead brought its own defects: a steady-state spawn waiting behind another
instance's trust session, a compare-first split to avoid that, a script
write and an approval promotion that could fail for want of the lock.

Removed, with their tests: the real-home write lock and its async-context
reentrancy, the plan/compare split that kept it off steady-state spawns, the
compare-first legacy sweep, the locked approval promotion and its unlocked
fallback, the compare-first shared script write (writeManagedScript already
skips identical bytes and restores the exec bit), and the CLI tsconfig
entries the lock pulled in.

Kept: the retired-forms matcher, the hooks-off no-op, removal only on an
explicit opt-out, the pane own-script check, and the compare-and-swap
rollbacks. Every instance now writes identical bytes idempotently.

* fix(codex): an opt-out that cannot read hooks.json keeps Orca's trust and ledger

The opt-out swept the real-home entry, then dropped Orca's ledger-proven trust
whenever a ledger existed, even when the sweep could not read hooks.json. The
entry could still be there, now untrusted, and the ledger that proves ownership
was gone for the retry. Drop that trust only after a sweep that read the file.

* refactor(agent-hooks): one predicate for whether an agent's status hooks are on

"Global switch on and this agent not turned off" was spelled out separately
in the startup controls, the settings reconcile, the retained-home
reconcile, the WSL preflight RPC, the CLI preflight and the OpenCode plugin
selection. They now share one function, in a module light enough for the
CLI's per-launch Codex preflight to load. The PTY spawn env derives the
Codex flag from the switch and opt-out list it already carries, the same way
it does for OpenCode and Pi, instead of receiving a second copy.

* fix(codex): launch and resume prep honour Codex's per-agent hook opt-out

Turning Codex off in the per-agent hook settings removes Orca's Codex hook
entry, but launch prep and session resume read only the global hooks switch,
so the next Codex launch or resume wrote the entry straight back into the
real ~/.codex or the account's home. Both now read the per-agent predicate,
which the PTY spawn env and startup already honoured.

* fix(codex): turning Codex off per agent clears the real ~/.codex entry

While the real-home lane owns ~/.codex/hooks.json, the legacy system-home
sweep stands down. That gate read only the global switch, so turning Codex
off per agent ran remove() with the sweep still suppressed and left Orca's
entry in the real ~/.codex. The gate now reads the per-agent predicate, the
same as turning every hook off.

* test(codex): cover the system ~/.codex sweep gate for Codex turned off

The gate that lets the legacy system-home sweep run was an inline closure in
startup, so reverting it to the global switch left CI green. It is now a
pure function beside the gate it feeds, with a table test and a remove()
test on a seeded ~/.codex: turning Codex off strips Orca's entry and keeps
user hooks; with Codex on the entry stays.

* fix(cli): keep the agent-status hooks predicate loadable by the packaged CLI

The CLI's prepare-codex handler imported the predicate from src/main, but
the Electron build rebuilds out/main from its declared entries only, so the
packaged `orca agent hooks` commands could not load it (package jobs and the
CLI bundle-parity test were red). The predicate reads only settings, so it
now lives in src/shared, which the CLI compiles itself.

* feat(codex): every Orca build writes one frozen Codex hook command

The Codex hook command was built from this build's wrapper, so two builds
on one HOME disagreed about the bytes of the shared ~/.codex entry and kept
rewriting it, with a Codex trust session each time.

The command is now fixed per form and carries its form number:
- POSIX: one command with no path in it. It runs the shared script only in
  an Orca pane with hooks on (pane key and hook port set), drains stdin
  everywhere else, and always exits 0. A branch for a per-build script root
  is written now and stays dormant until Orca sets ORCA_AGENT_HOOK_ROOT, so
  that change will not move these bytes.
- Windows: the bare forward-slash path to the shared .cmd, which runs under
  PowerShell 7 and 5.1, Codex's hook hosts. A profile path that is not one
  PowerShell token gets a plain PowerShell form with the same branches.

The literals live in the form module, so a change to the shared hook
constants cannot move them; goldens pin the bytes. Every form keeps
`agent-hooks/codex-hook.*` in plain text, so older builds still recognize it.

* fix(codex): one main-process owner adds the real-home entry; nothing restores files

Each Orca writer of ~/.codex decided what Orca's entry must be from its own
build and instance, then removed or reverted whatever differed: launch prep
rewrote any Orca-shaped entry to this build's command and stripped Orca
entries from events this build does not use, and a failed trust session
restored hooks.json and config.toml from snapshots. With several instances
and builds on one HOME, every disagreement became a deletion or a revert.

The main process is now the one writer, and its writes are add-only:
- A launch or resume adds Orca's frozen entry to an event that has none and
  leaves every Orca entry it finds, so a running older build is never fought.
- App start also converts an older Orca form to the frozen command, once, in
  its own slot: one hooks.json write (one .bak) and one trust grant per home.
- A newer form is never rewritten or appended beside, and Orca entries in
  events this build does not use are kept.
- After a failed trust grant, only an entry this call wrote that is still
  untrusted is withdrawn, putting back the handler it replaced. Both files
  are re-read, so a concurrent edit, or the identical entry another Orca
  trusted meanwhile, survives.

Deleted: the compare-and-swap hooks.json restore, the config.toml snapshot
restore after a grant session and after a user-trust re-key, and the
rollback module. A grant session writes trust only at Orca's own keys, and
every caller settles those keys itself. A failed re-key of moved user hooks
now keeps the write and reports it; Codex lists those hooks for review.

* fix(codex): the pane CLI asks the app to prepare its Codex home

`orca agent hooks prepare-codex` ran Codex's install inside the pane. That
process can have the real HOME (login(1)) and runs outside the app's
in-process queues, so it was a second writer of ~/.codex and ~/.orca beside
the app. A check that the home ran "its own script" guarded it.

The pane step now only asks the app, over the same kind of local RPC the WSL
pane step already uses (agentHooks.prepareCodexForPane). The app checks that
the pane's CODEX_HOME is one its own userData owns, reads its own hooks
setting, and installs on its own queue. An app that is not running, or is
too old to know the method, makes the step a no-op, as it is on WSL. The
own-script check and the CLI's settings read are gone, and the preflight
module leaves the CLI bundle.

* fix(codex): delete the pane step on native hosts

The previous commit had `orca agent hooks prepare-codex` ask the app to
prepare the pane's Codex home. The case it existed for (#14326, a pane that
survives an app update with stale hook trust) did not reproduce, and no other
desktop agent host writes agent config from a terminal or launch wrapper.

- Deleted: the agentHooks.prepareCodexForPane RPC method, its params and
  catalog entry, and prepareManagedCodexHomeBeforeShellLaunch with its module,
  tests and CLI build entry.
- `agent hooks prepare-codex` is a no-op on native hosts. It stays for one
  release so shell wrappers from older builds, which still call it, exit 0.
- WSL panes are unchanged: they still ask the app over
  agentHooks.prepareCodexForWslPane.

The shell wrappers and ORCA_CODEX_LAUNCH_PREFLIGHT stay, because WSL panes
use the same wrappers and variable (forwarded through WSLENV). A native pane
still starts the CLI once per `codex` it runs; skipping that is a follow-up.

* test(codex): a failed trust session keeps concurrent edits to both files

QA case 9 at host level, on a real file system in a temp HOME: Codex's trust
session fails after another writer saved hooks.json and config.toml.

- Both saves survive, and no Orca entry is left that Codex would list for
  review: this call's entry is withdrawn.
- A failed one-time conversion puts the older Orca entry back in its slot and
  keeps both saves.

Both tests fail on the previous head, which restored config.toml from a
snapshot and left the untrusted entries in hooks.json. Removing the
withdrawal turns both red.

* feat(codex): read whether an Orca entry's stored trust is still current

A Codex release that changes how it hashes a hook leaves Orca's stored trust
stale: the entry is present, but Codex lists it as modified. Checking only
whether the entry is missing cannot see that.

readOrcaEntryTrust sorts a present entry into four states:
- trusted: the stored hash is the current one;
- untrusted: there is no stored hash;
- stale: the stored hash is not the current one;
- disabled: the user turned the entry off.

The caller can pass Codex's current hash, for example one a grant recorded.
The failed-grant withdrawal now uses it, and also keeps an entry the user
turned off. Nothing re-grants on 'stale' yet.

* fix(codex): a slow Codex start retries on the next launch, never for minutes

On a loaded Mac a cold `codex app-server` took over 10 s (QA case 4). The
grant timed out, the entry was withdrawn, and a 5-minute cooldown in both the
grant and the real-home install then refused every retry.

- The native session deadline is 30 s, the same as WSL's.
- A timeout starts no cooldown in the grant or in the real-home install. The
  next launch retries. Other failures keep their cooldown.
- Launches that queue behind a slow session share one follow-up run, so a
  launch waits for at most two sessions, not one per earlier launch.

Tests: a 15 s cold start still grants and keeps the entry; after a timeout,
the next launch runs a session at once; four queued launches run two
sessions. Each is red on the previous head, and each mechanism was removed in
turn to confirm its test turns red.

* fix(codex): Orca's automatic writes never move a user hook

Codex keys a hook's trust by its position in hooks.json. App start's collapse
of Orca duplicates removed every Orca entry and appended one at the end. That
moved any user hook that followed a removed entry, so the write waited on a
session to re-key the moved hook's trust.

App start now:
- converts the first Orca entry that sits in a plain slot to the frozen
  command, in place;
- drops any other Orca entry only when that moves no user hook;
- keeps a duplicate that a user hook follows, and trusts every frozen copy,
  so none is listed for review;
- appends only when no frozen entry is left.

Tests check user positions and user trust blocks byte-for-byte for each
automatic write: add-missing (append), the one-time conversion (in place),
a trailing duplicate, a duplicate before a user hook, and older duplicates
normalized to one entry. The three collapse cases fail on the previous head.
Removing the position check, or the in-place conversion, turns its tests red.
Only the explicit opt-out still removes an entry that user hooks follow.

* fix(codex): removing an Orca entry never waits on a Codex session

Removing an Orca entry from ~/.codex/hooks.json moves every user hook behind
it up a slot, and Codex keys trust by slot. The retired-form sweep, the
opt-out and a failed-grant withdrawal all asked a `codex app-server` session
to list the old trust before writing, and to re-key it afterwards. A timeout
there threw before the write and latched a 5-minute cooldown, so a slow cold
start blocked the retired-form sweep at boot (QA case 4).

Each moved hook's [hooks.state] block now moves to its new key, body bytes
unchanged, straight after the hooks.json write. Codex hashes a hook's content,
not its position or its file path, so the moved block stays exactly as valid
as it was: a trusted hook stays trusted, an untrusted one stays untrusted, and
one the user turned off stays off. No removal waits on or depends on a
session. A failed config.toml write keeps the hooks write and logs.

Deleted: the inspect and repair sessions, their client, and their cooldown.
The generation guards on the hooks.json writes stay, for other processes.

Tests: the retired sweep removes the retired entry and carries the trust of
the user hook behind it while every Codex session times out (red on the
previous head); the opt-out carries an appended user hook's trust; the move
carries trusted, disabled and untrusted states byte for byte. Removing the
move turns all of them red.

* fix(codex): a Codex launch never waits on Codex's approval of Orca's entry

A launch on the real-home lane awaited Codex's trust grant for the entry it
had just added. A cold `codex app-server` on a loaded Mac took over 10 s, so
the launch could wait that long, and a failure then latched a 5-minute
cooldown.

- Codex's approval runs in the background, with a 30 s cold-start budget.
- A launch uses the real home only when the ledger shows trust is already
  current. Otherwise it goes to the managed home at once, and the next launch
  picks up the finished grant.
- A launch that arrives while a grant runs does no work and does not queue
  behind it.
- A resume into the real home has no managed home to fall back to. It waits
  for the grant, but no longer than the 10 s a launch always could.
- A background grant that times out starts no cooldown; the next launch
  retries. Any other failure backs off for 10 s instead of 5 minutes.
  Success is what the ledger remembers.
- A failed grant still withdraws only what that install added and is still
  unapproved. The log now says how many entries it took back and when the
  next try comes.

Managed-home grants keep their 10 s deadline and stay on launch prep, as
before; they fall back to Orca-computed trust.

Tests:
- A 15 s start: the launch returns in under a second on the managed home, a
  second launch starts no session, the grant lands in the background, and the
  next launch uses the real home.
- A timeout sets no cooldown, withdraws its adds and logs it.
- Another failure retries after 10 s, not before.
- A resume waits only as long as allowed.
- Case 9 checks the log line and the retry.

Making the launch await the grant, a 10 s budget, either timeout cooldown, and
a 5-minute backoff were each tried, and each turns its test red.

* fix(codex): move a hook's trust only when every stored key has the known shape

Orca now edits Codex's trust store directly when a removal moves a user
hook. Three safeguards keep that honest:

- Fail safe. If any [hooks.state] key in config.toml does not have the
  shape `<path>:<event>:<group>:<handler>`, nothing moves and Codex asks the
  user to review. That shape was checked unchanged from Codex 0.141 to 0.158.
- Targeted. The file is read immediately before the atomic rename, and only
  the moved keys' blocks change. Every other byte stays, and no snapshot is
  restored.
- Verbatim. Each block's body moves as Codex wrote it, including fields
  Orca does not know. No hash is ever computed, and a hook with no block
  gets none.

Tests:
- An unknown key shape stops every move.
- Everything except the moved block survives byte for byte, and the moved
  body keeps an unknown field.
- In case 9, a hook the user approved during the failed session keeps its
  approval when the withdrawal moves it, beside the concurrent project edit.

Removing the shape check, or writing a computed block instead of the stored
body, turns these tests red.

* refactor(codex): keep only the trust read the failed-grant withdrawal uses

A capture across Codex 0.141, 0.150 and 0.158, switching in all six
directions, showed Orca's entry keeps the same hash and stays trusted. A
Codex upgrade does not make its trust stale, so nothing needs to re-grant
on staleness.

readOrcaEntryTrust keeps the four states the withdrawal needs, but loses
the parameter that let a caller pass a different current hash, and the test
for a Codex that hashes differently.

* fix(codex): native panes no longer start the Orca CLI before each codex

The pane step is a no-op on native hosts, but native panes still carried
ORCA_CODEX_LAUNCH_PREFLIGHT, so every `codex` typed in a pane started the
Orca CLI for nothing. Only a packaged Windows build's WSL pane now gets the
variable; the app prepares every native Codex home itself.

The resolver loses the dev-launcher path and its userDataPath option, which
only native panes used.

Tests: a native macOS, Linux and Windows pane gets no preflight, packaged or
not, even with the bundled CLI present; a WSL pane still gets the verified
absolute launcher. Letting native panes through again turns them red.

* chore(cli): say when the native prepare-codex no-op can go

Native pane wrappers from builds up to v1.4.216 still call it. It can be
deleted once no supported build's wrapper does.

* test(codex): check the WSL launcher path instead of asserting it

* fix(codex): a launch no longer waits behind the background real-home approval

The background grant ran its whole codex app-server session inside the shared
~/.codex/config.toml lane, and on a cold host its session was also the shared
capability probe. A launch sent to the managed home then waited on both: the
managed install and the project-trust write queue on that lane, and the
managed install's own grant waited for the probe. On a cold app-server that
was up to 30 s per launch.

The lane was held across the session only to protect the retired
capture-and-restore. Codex writes its own records, so the lane is now taken
only around Orca's own pre-grant write. The background grant runs its session
without publishing it as the shared probe, and the whole grant is bounded by
its deadline, so a hang outside the session cannot leave the lane 'granting'.

* fix(codex): a failed re-grant no longer strips Codex's own approval of Orca's entries

Before each trust session, the grant deleted every Orca record whose hash
matched the one Orca computes. That exists because a managed home's fallback
writes Orca-computed trust under both Windows path-separator spellings, and
Codex rewrites only its own spelling, so the other copy would linger. On
failure the managed and WSL fallbacks write that trust back, and before this
fold a snapshot restore covered it.

The real ~/.codex has neither: Orca never writes computed trust there (the
real-home lane does not run on Windows at all), so a matching record there is
Codex's own approval. After a ledger miss (another Orca profile, a Codex
update, a lost ledger) and a failed session, nothing put it back, and every
Orca entry showed "Hooks need review".

The clear now runs only for homes whose fallback writes that trust.

* fix(codex): a real-home resume spawns only once Orca's entry is approved or withdrawn

A resume that must run in ~/.codex waited at most 10 s for the background
approval, then spawned anyway. On a cold app-server that left Codex beside an
unapproved Orca entry, so the resumed pane showed hook review.

The resume now waits for the grant to settle. Settled means Codex approved the
entry, or the grant failed and withdrew its own unapproved write; the grant's
deadline bounds the wait (30 s, the cold-start budget), and a failed approval
never fails the resume.

Why this over the alternatives:
- Spawning at 10 s keeps the review prompt this fold exists to remove.
- Withdrawing at 10 s from the resume races the still-running session: Codex
  can write the frozen entry's hash after the withdrawal, and for a converted
  entry that marks the older command Orca put back as modified.
- A resume cannot use the managed home: the session lives in ~/.codex.
So the only states that cannot race Codex are the grant's own settle. The cost
is a longer worst case on a cold app-server (up to the 30 s deadline, plus any
managed-home install that holds the config.toml lane); a warm approval takes
seconds, and an approved entry costs no wait.

* fix(codex): keep the 5-minute trust cooldown for launch-path grants

The fold shortened the host's trust-grant cooldown from 5 minutes to 10
seconds for every grant. That was meant for the background ~/.codex approval,
which blocks no launch. The managed-home and WSL grants run inline on the
launch path, so with a hung app-server every launch more than 10 s after the
last failure paid the full inline timeout again (10 s native, 30 s WSL).

Cooldowns are now kept per lane: inline grants keep 5 minutes, the background
grant retries after 10 s, and neither lane's failure cools the other down. A
success, or a proven-missing surface, still clears both. The real-home
install's own retries (an unreadable hooks.json, unknown keys) are back on the
5-minute interval they had before the fold.

The cooldown moves to its own module so the grant stays within the file limit.

* fix(codex): a failed grant withdraws the exact copy it wrote

The withdrawal re-found "this call's" entry by command, taking the first
frozen handler in the event. When app start converted a later slot while an
earlier frozen copy sat in a matcher group (which conversion skips), a failed
grant acted on that earlier copy: it put the older command into it, or skipped
it, and left the converted, unapproved copy in place.

Each write now records where its handler landed, after any duplicate drops,
and the withdrawal acts only on that slot. A copy that has since moved is left
alone; the next launch's grant retries it.

* fix(codex): the failed-grant withdrawal checks hooks.json is unchanged before writing

The install and the retired-form sweep both refuse to replace ~/.codex/hooks.json
if it changed since they read it. The withdrawal did not: a save landing
between its read and its atomic replace was lost. The window is small, since
the withdrawal is synchronous, but it now carries the same guard.

* refactor(codex): drop rationale left over from the snapshot restore; name the trust-move module for what it does

Comments on the config.toml lanes still justified them by a grant's
capture-and-restore window, which the fold deleted, and the trust-write
deadline still counted a grant session holding the lane. They now give the
reason that remains: Orca's own multi-step reads and writes, and managed-home
installs that hold the lane across their inline grant.

codex-user-hook-trust-rebase no longer rebases through Codex; it moves stored
trust records, so it is now codex-user-hook-trust-moves.

The grant test that pinned two sessions on one config.toml to run one at a
time is removed: its reason was an interleaved capture and restore. Callers
that write config.toml around a grant hold their own lane, which the nested
installer test still covers.

* build(cli): list the trust-grant cooldown module in the CLI program

The CLI's agent-hooks handler loads the hook controls, which reach the Codex
trust grant; the CLI project is composite, so every module in that graph must
be listed.

* docs(codex): say which Windows hosts each hook command form runs under

Codex runs a hook under the turn's shell (PowerShell 7 or 5.1 in every
captured session) and, with no single local turn shell, under %COMSPEC% /C.
The bare forward-slash path ran under all three in the Windows host census.
The PowerShell form used for a profile path with a space does not parse under
cmd.exe; no form valid in all three hosts has been run for such a path, so the
form stays and the gap is stated here and in the PR.

* test(codex): type the withdrawal seam without an assertion

* fix(codex): a real-home resume starts at once, trusting Orca's entries for that process

A resume that must run in ~/.codex waited for Codex's background approval of
Orca's newly written hook entry: up to 30-40 s on a cold app-server. That made
the user's resume wait on bookkeeping, and the alternatives (start at 10 s with
Codex's hook review showing, or withdraw the entry and race Codex's own write)
were worse.

Codex reads hook trust from its session-flag config layer as well as the user's
config.toml, merged per key, and has since hook trust shipped. So the resume no
longer waits. When Orca's own frozen entries in ~/.codex are untrusted (or hold
a stale hash), the resume command carries
`-c hooks.state={'<key>'={trusted_hash='<hash>'},...}` for exactly those entries:
the key under both the logical and the real path of ~/.codex (Codex keys an
explicit CODEX_HOME by its real path), and the hash of that entry's content, so
it can trust nothing else at that slot. The user's hooks are never included,
nothing is written, and the background approval still runs for later plain
`codex` launches. An approved entry adds nothing; a Codex known to lack hook
trust gets nothing.

One inline table, because Codex splits a `-c` key on every `.` and the key holds
`.codex/hooks.json`. TOML literal strings keep `"` out of Windows native-argument
quoting. The flag goes before `resume <id>`, quoted for the pane's shell (portable
Unix, PowerShell or cmd), in the launch command and in the setup-sequenced copy of
it; a cmd line whose path cmd would expand, or a key with an apostrophe, is left
unchanged. SSH and WSL resumes get no preparation, so no local path reaches them.

* Revert "fix(codex): a real-home resume starts at once, trusting Orca's entries for that process"

This reverts commit 1bd30651d6.

* fix(codex): a real-home resume starts at once, without waiting for approval

A resume into the real ~/.codex waited until the background approval settled,
up to its 30 s deadline on a cold app-server: bookkeeping for later launches
gating the resume the user asked for. It now starts at once. If the approval is
still running, that first resume can show Codex's hook review once; the
approval then lands and later resumes and plain codex launches are trusted.

Trusting Orca's entries per process was the alternative, but the resume command
is typed into the pane's shell, and hook settings stay out of typed commands.

* test(codex): read real-home hook groups with the installer's own type

* fix(codex): a background approval is bounded only by its session's own deadline

Review loop 2, L3. grantWithinDeadline raced a second 30 s timer against
the background approval. Loop 1 added it so that a hang upstream of the
session could not leave the lane 'granting' forever.

That hang cannot happen. The only caller is the native real-home grant
(its plan is always host 'native'; the real-home lane is off on Windows,
so WSL never reaches it). Everything before the session is synchronous
there: command resolution and binary stamp, the ledger read, the
state-db backfill check, the capability and cooldown checks, and
runUnshared awaits no shared probe. A synchronous hang would freeze the
main thread, which no timer can rescue. The session itself starts a kill
timer right after spawn (runCodexAppServerSession), with the same 30 s,
and it kills the app-server tree when it fires.

So the outer timer was a second copy of that bound. Because it started
first, it won by the spawn time. It then settled the lane and cleared
backgroundGrant while the app-server was still alive, and the next
launch could start a second concurrent session. It abandoned the
session rather than cancelling it. Deleted, not moved: the session's
own timer is the one bound, and it cancels.

Test: codex-real-home-slow-app-server.test.ts "runs one session at a
time, ended by its own deadline". The fake session starts its timer
after a simulated spawn, as the real one does. A launch at 30 s finds
the session still running and starts none; the lane settles when the
session times out. It replaces the "settles a grant that never answers"
test, whose never-answering session could not time out at all.

* fix(codex): a background approval's retry has one schedule, the real-home lane's

Review loop 2, L4. A non-timeout background failure set two 10 s
schedules for one failure: the real-home lane's installRetryAfterMs,
which gates ensure, and a `<host>#background` cooldown in the grant
module. ensure's gate always tripped first, so the second one was
consulted only after something reset the first (turning hooks off).
Then it answered 'retry-cached', which wrote the entry into
~/.codex/hooks.json only to withdraw it again: churn, not protection.

Background plans now neither start nor consult a grant-module cooldown.
The real-home lane (installRetryAfterMs) is the one source of truth for
when a background approval runs again, and its 10 s interval moves into
codex-real-home-background-grant.ts, the module that sets it. The
cooldown module is back to one host-keyed map for launch-path grants,
with the same 5-minute interval as main. A success or a proven-missing
surface from either lane still clears the host's cooldown.

Tests:
- codex-hook-trust-grant.test.ts "neither starts nor waits on a
  cooldown for a background grant": two failing background grants each
  run a session and leave no cooldown; an inline failure still cools
  down inline grants and not the background one.
- codex-real-home-slow-app-server.test.ts "has one retry schedule:
  turning hooks off and on after a failure retries at once": after a
  failed approval, hooks off then on runs a session and installs,
  instead of a retry-cached write-and-withdraw.

* fix(codex): hooks turned off and on during an approval re-add Orca's entry

Review loop 2, L1. ensure returned at once whenever a background
approval was running, whatever the lane. Turning hooks off during an
approval sets the lane to 'removed' (usable), so turning them back on
returned 'removed' without re-adding the entry. Launches in that window
spawned in ~/.codex with no Orca hook and got no status for their
lifetime, for up to 30 s, until the approval settled and a later launch
re-added it.

ensure now returns early only while the lane is 'granting', which is
what the early return exists for: a launch never waits on Codex's
approval and uses the managed home until it lands. Any other lane runs
the normal add-missing install.

That install can start a second approval while the first is still
running. Approvals are now chained, so Codex still runs one session at
a time, and a finished approval clears the handle only if it is still
the latest one (before, an older approval's finally could clear a newer
one's handle). The older approval's result is already dropped by the
lane generation check.

Test: codex-real-home-slow-app-server.test.ts "re-adds the entry when
hooks go off and on during an approval, one session at a time". While
the approval hangs: opt-out removes the entry; re-enable re-adds every
entry, keeps launches on the managed home, and starts no second
session; once Codex answers, the lane is installed and every entry is
approved.

* test(codex): a launch during the real-home approval shows what it waits on

Review loop 2, M2. The launch test's fake Codex failed every
managed-home session at once with ENOENT, so the managed home's own
approval was an instant "unsupported" fallback, and the test could not
show that a launch sent to the managed home still waits on that home's
inline approval when its ledger misses (first use, a Codex update, a
lost ledger), up to 10 s, as on main.

Now the managed-home session behaves like a real one:
- "settles on the managed home with its hooks and the project trust
  written": the managed app-server answers; two launches settle in
  under 2 s while the real-home approval hangs, and the second launch
  finds the managed approval in its ledger (one managed session).
- new "waits up to the managed home's own 10 s approval when that home
  is cold too": the managed session fails at its own deadline, as the
  real one does. The first launch is still pending at 9.999 s and
  settles on the managed home at 10 s; the request asked for 10 s. The
  next launch settles at once, because the failed inline approval cools
  down for 5 minutes.

No product change.

* refactor(codex): the managed and WSL installs own their pre-approval trust clear

Review loop 2, L7. Before a Codex approval session, a managed or WSL home
clears the approvals Orca itself computed, because on Windows its
fallback writes them under both path spellings and Codex's canonical key
may not overwrite the other one. The fallback writes them back if the
session fails. ~/.codex has no such fallback, so there the clear would
only delete Codex's own records (loop-1 H2). The grant module carried
this as a plan flag, fallbackWritesSelfComputedTrust, and took the
config.toml lane around the clear itself.

The reviewer proposed moving the clear into the two callers. A literal
move, clearing before the grant call, is NOT behaviour-neutral, so this
does not do that:
- The grant first checks its ledger, which compares the stored hash
  with the one Codex recorded. Codex's hash equals Orca's computed one
  (the premise of readOrcaEntryTrust), so a clear before that check
  deletes exactly the record the ledger proves. Every managed launch
  would then miss the ledger and run an inline session (up to 10 s).
- Checked, not inferred: with the clear moved before the call in the
  managed install, codex-launch-during-real-home-grant.test.ts "settles
  on the managed home..." fails (2 managed sessions instead of 1).
  Log: ~/orca-qa/codex-real-home-leak/fb6/l7-literal-move.log

What this does instead: each caller passes its clear as the grant's
`beforeSession` step, which the grant runs only when a session will
actually run (after a ledger miss, and not on a cooldown or cached
fallback), exactly where the flag ran it. So:
- the flag and its "never set for the real home" rule are gone; the
  real-home grant passes no step, so the grant module has no path left
  that deletes a trust record in ~/.codex;
- the grant module's own lane acquisition around the clear is gone. It
  was always a pass-through: both callers already hold that file's
  lane (the managed install holds the runtime and system lanes, the
  WSL install holds its config.toml lane) across the whole grant.

No behaviour change. The loop-1 probes still pass as fixed: trust-strip
prints every entry trusted after a failed re-grant, and lane-hold
prints managedInstall=settled projectTrust=settled.

Tests (codex-hook-trust-grant.test.ts):
- "removes equivalent Windows fallback keys before the RPC writes
  canonical trust" now passes the managed caller's step;
- new "runs the caller's pre-session step only when a session runs":
  the step runs once for a session and not on the ledger hit after it.

* chore(codex): comments stop describing a lock held across the session, or a rollback

Review loop 2, L6 comment sweep (comments and one test name only):
- codex-trust-config-concurrent-launch.test.ts: the test named "does not
  let a failing launch roll back a concurrent launch" said the per-file
  lane was the only thing left and that the doomed run's rollback must
  not resurrect the file. There is no lane across a session and no
  rollback now. Retargeted to what it covers: "leaves a concurrent
  grant's records in place when a sibling grant fails" (a restore would
  still turn it red).
- codex-trust-grant-ledger.ts: "a grant session blocks launch prep" is
  true only of inline grants; the background one still costs an
  app-server start. The drift clause no longer says "before the pane
  launches", which is false for the real home.
- agent-trust-write-deadline.ts: a stray hard wrap.
The install.ts:105 comment was fixed with L1. A sweep of src/main/codex,
src/main/startup, src/main/agent-hooks, the trust presets and the CLI
handlers for rollback, restore, rebase, capture/restore, and a lane held
across a grant or session found nothing else stale; the remaining "no
restore" comments state the current rule.

* fix(codex): a real-home resume waits for the one running approval, up to its 30 s limit

Review loop 2, M1; coordinator ruling. A resume into ~/.codex has no
managed home to fall back to. 5a737261d8 let it start at once beside an
Orca entry still awaiting Codex's approval. Codex's TUI then shows a
full-screen hook-review picker before the session and waits for keys:
"Trust all and continue" also trusts the user's own unreviewed hooks,
and "Continue without trusting" leaves that session with no Orca status
for its whole life, because Codex does not reload hooks when Orca's
approval lands later. Panes restored at app start after an update hit
it too, since the start-time conversion leaves every entry awaiting
approval.

The resume now waits, but only while Orca's entry in ~/.codex is
written and a grant is approving it (lane 'granting'). Every resume
waits on that same in-flight grant: ensure never starts a second one
while the lane is 'granting', so panes restored together share one
session. The bound is the grant's own session limit (30 s). The grant
settles only after Codex approved the entry, or after it withdrew its
own unapproved adds, so the resumed session starts either trusted or
with no Orca entry: never beside an unapproved one, and no picker. On
a withdrawal that session has no Orca status, as on main after its
10 s wait. A failed approval never fails the resume.

Tests (codex-launch-during-real-home-grant.test.ts):
- "waits for a warm approval, and spawns with the entries approved";
- "spawns at the approval session limit with Orca entries withdrawn"
  (fake timers: pending at 29.999 s, spawns at 30 s with no Orca entry);
- "makes panes restored together wait on one approval session" (three
  resumes, one session, all settle once it lands).
codex-launch-per-agent-hook-opt-out.test.ts: a resume into ~/.codex
awaits the approval; a resume into a managed account home does not.

* fix(codex): repeated background approval timeouts back off, growing to 5 minutes

Review loop 2, M3; coordinator ruling. A timeout of the ~/.codex
approval starts no cooldown, so the next launch retries at once. On a
host where codex app-server never starts within 30 s, every launch then
wrote Orca's entry into ~/.codex/hooks.json, withdrew it again, and
started another 30 s session, for the rest of the process: an unbounded
retry with no exit.

After 3 timeouts in a row the retry now waits 10 s, then 1 minute, then
5 minutes for every later one. The first two timeouts still retry on the
next launch, so a slow cold start is not punished. Any other outcome
ends the streak (a success, or any other failure, which keeps its own
10 s wait). The streak lives only in memory, so every app start begins
at zero and a slow boot can never latch.

Tests (codex-real-home-slow-app-server.test.ts):
- "backs off after three timeouts in a row, growing to 5 minutes, and a
  success resets it": the first two timeouts retry at once, then 10 s,
  1 min, 5 min, 5 min; after a success, a fresh approval gets two
  immediate retries again and a 10 s backoff after the third;
- "keeps trying after timeouts during a slow first start, once the app
  server answers": three timeouts, then the next attempt at 10 s
  installs.

* refactor(codex): one approval at a time, decided under the config.toml lane

The real-home check kept a lane label, a generation stamp, a promise chain of
ensures and a chain of approvals, and decided from the label at call time.
Concurrent resumes from any state other than 'granting' each started their own
approval (N x 30 s), a chained approval ran a plan an earlier failure had
withdrawn, a hooks-off check during an approval released a waiting resume beside
unapproved entries, and an app-start conversion during an approval was dropped.

Now each check is one step under the real config.toml lane: an approval in
flight answers 'approving' (unusable), hooks off answers 'removed', an open
retry window answers 'unavailable', and otherwise the unchanged install runs and
starts at most one approval. The approval settles under the lane: it withdraws
its own unapproved adds on failure, sets the retry, and derives the verdict from
the settings and the outcome, then runs an owed conversion. A resume waits only
while an approval runs and an unapproved Orca entry is on disk. The opt-out
sweep moves verbatim into its own module.

* fix(codex): only a success or app start resets the approval timeout streak

The ruling is that three timeouts in a row back off, and the count resets on
success and at app start. A non-timeout failure or an unexpected error also
reset it, so a host alternating those with timeouts never backed off.

* fix(codex): a Windows profile path the shells cannot carry bare runs through cmd.exe

The Windows hook command was the bare forward-slash script path, or, for a
profile path that is not one PowerShell word, a PowerShell script. That script
cannot parse under cmd.exe, which Codex uses when a session has no single local
turn shell, so such a profile got no status there.

A path of only letters, digits and _ . : / ~ - stays bare. Any other path,
including one with a space, & ^ $ ` ' ! ( ) or a non-ASCII character, is written
as cmd --% /d /c @"<path>", which ran under PowerShell 7, Windows PowerShell 5.1
and cmd.exe for each of those characters with a real Codex 0.158.0. The choice
depends only on the path, so every build on a machine writes the same bytes. A
machine holding the earlier PowerShell spelling converts it once at app start.

* build(cli): list the real-home hook sweep module in the CLI program

* fix(codex): the Windows cmd spelling names the system cmd.exe and turns off delayed expansion

A profile path the shells cannot carry bare was written as
cmd --% /d /c @"<path>". Under Codex's cmd.exe host the outer cmd.exe resolves
a bare `cmd` from the hook's working directory first, so a repo holding
cmd.bat (or .cmd, .com, .exe) at the session cwd would run on every hook event.
And with delayed expansion turned on in the registry, a `!` in the path was
dropped.

The spelling is now <SystemRoot>/System32/cmd.exe --% /d /v:off /c @"<path>",
unquoted (PowerShell reads a quoted first token as an expression) and with
forward slashes. The Windows directory comes from %SystemRoot% when written,
else from the directory above %ComSpec%'s System32, so both give the same bytes;
if neither is a drive-absolute path it can spell unquoted, it is C:/Windows,
which is still absolute. The bytes stay a pure function of the profile path and
that directory, so every build on a machine writes the same command. Safe
profile paths keep the bare path. Older Orca forms, including the bare-cmd
spelling, convert once; the new spelling is never swept as retired.

* refactor(codex): an approval's settle runs no deferred conversion

An app-start conversion that arrived while an approval ran was remembered and
run by that approval's settle. The settle then rewrote an older entry in place,
unapproved, and started a second approval inside the same wait that releases
every resume, so a resume could start beside an entry Codex would put up for
review.

That path could not happen: the only conversion caller is app start, and it is
the process's first check, so no approval can be running when it arrives. The
deferral and the settle's second check are deleted. A conversion that met an
approval would now be skipped until the next start, and the test for this case
pins that the settle writes nothing new and runs one session.

* fix(codex): an approval's settle keeps a failed opt-out's verdict and ends only its own flight

With hooks read off, an approval's settle always concluded 'removed', which the
routing check treats as usable. If an opt-out during that approval could not
read hooks.json, it had concluded 'unavailable' because the entry may still be
there, and the settle overwrote that. The settle now keeps 'unavailable' when
hooks are off; the next hooks-off check or opt-out re-derives it as before.

The settle's fallback when it cannot run now clears the running approval only
if it is still its own, and the routing check's comment states its rule: never
usable while an approval runs.

* fix(codex): spell the system cmd.exe with backslashes

Under Codex's cmd.exe host the outer cmd.exe hands the typed program text to
the child verbatim, and cmd.exe scans its whole command line for switches, so a
forward-slash C:/Windows/System32/cmd.exe is read as switches: the hook never
runs ("The syntax of the command is incorrect.") and /d is lost. Measured live
on Windows; both PowerShell hosts rewrite argv0 and were unaffected. The script
path after @" keeps forward slashes.

* chore(codex): say why the cmd.exe path is absolute, as measured on Windows

* test(codex): Windows managed-install tests expect the frozen command

They still asserted main's PowerShell text and a backslash bare path; they only
run on Windows, so nothing here caught it. Also correct the /v:off comment: a
lone ! is never dropped, only a !NAME! pair expands.

* ci: run the Codex managed-install tests in the Windows job

Its Windows-only cases skip everywhere else, so nothing ran them; three of them
still asserted a command this branch no longer writes.

* ci: a change to the Codex managed-install tests starts the Windows job

Also say what the missing-script case asserts: a non-zero exit, which
PowerShell reports as 1.

* chore(codex): name the hook trust key pattern for what it matches

* test(codex): the managed-install tests remove folders with the retrying helper

Now that they run in the Windows lane, a raw recursive rm there can throw EPERM
after the assertions pass.

* refactor(codex): one Codex hook-trust key pattern for the trust move and #23958's carry

* test(codex): the trust move carries a block in Codex's quoted spelling and leaves no second table
2026-09-30 14:26:23 -07:00
Neil cfe4c633eb fix(mobile): publish the Android APK's size and checksum with the release (#24037)
An APK that fails to install with a missing certificate or a package-parse error
is usually a download that died near the end: the signature block sits in the
last ~100 KB of a 133 MB file, so a truncated APK looks complete and carries no
signature at all. The release published neither a size nor a digest, so there
was no way to tell that apart from a bad build without deriving both from the
asset by hand.

The release now uploads app-release.apk.sha256 next to the APK in
`sha256sum -c` format (binary marker, so Git Bash cannot translate line endings
while hashing) and puts the exact byte size and digest in the release body,
naming `shasum -a 256 -c` for readers on macOS.

The upload path rewrites the body too: --clobber replaces the APK, so a digest
left over from the previous build would describe a file nobody can download,
and a reader comparing against it would reject a good APK. Both paths reserve
the section's own length out of the release-body cap before truncating, so the
section always survives and the body always fits; MAX_RELEASE_BODY_LENGTH is
exported from the desktop release script rather than restated.

Refs #24011, #12248, #11444.
2026-09-30 12:14:13 -07:00
Neil d2dbe2c385 fix(windows): replace the managed CLI launcher with a native one (#24094)
* docs(security): add the antivirus clearance path for future releases

Every AV false positive here has been handled one vendor and one shipped
version at a time. Document the programs that clear future releases instead --
signer and product enrollment rather than per-build sample submission -- and add
a script that reports an RC's current detection state by hash, so a verdict is
found before users meet it in an issue report.

Hash lookup only by default; --upload transmits the artifact and stays manual.

* fix(windows): replace the managed CLI launcher with a native one

resources\bin\orca.exe was a csc-compiled MSIL assembly: a small, freshly
compiled .NET image in a user-writable directory that mutates environment
variables and proxies a child process. That is the shape .NET dropper
heuristics are trained on, and every verdict against it named the family --
MSILHeracles from two vendors, Wacatac!ml from a third. Signing the file does
not change its shape, so signing never cleared it.

Rebuild it in Rust. Same resolution, same environment contract, same argv
passthrough that keeps newline-bearing orchestration bodies intact (#8374), and
the child still inherits our environment block rather than an explicit map, so
a block carrying both PATH and Path survives (#12046). The PE now carries
publisher, version, icon and an asInvoker manifest from build.rs.

Refs #23383

* ci(windows): install the Rust toolchain before building the CLI launcher

The hosted runners happen to ship cargo, but a real Windows dev box does not --
verified on our own Windows QA host, where cargo and rustc were both absent.
Relying on the image means a future image change fails deep inside
electron-builder's native hook instead of at an obvious step.
2026-09-30 02:49:03 -07:00
Brennan Benson e594cb06af test(mobile): record RPC goldens without a pinned commit, and check recorded requests against the desktop's params rules (#23732)
* test(mobile): add rpc:diff to decode what a recording change moved

The RPC recording goldens are content-addressed JSON, so their raw git diff is
pool hashes. `pnpm --dir mobile rpc:diff [<base>]` decodes both sides and prints,
per golden, the checkpoint, field and JSON path that moved with both values,
grouped across checkpoints, plus added and removed goldens. `--summary <file>`
appends a Markdown report capped for GitHub's step-summary limit.

It reads any pooled format, so it can prove the next commit's format change
moves no recorded value. Checkpoints are matched by occurrence because an id can
repeat within one golden.

This commit adds files under the recorder directory, which moves the header
digest every golden pins; the next commit removes that header.

* test(mobile): record RPC goldens without a pinned commit or input digests

Every golden carried a pinned `baseline` commit plus digests of the recorder,
its mount adapter and its scenario, and the record script refused to run unless
the product tree matched the pin. So every behaviour change repinned to its own
branch commit and rewrote all ~790 files, the squash made that commit
unreachable, and main's pin job stayed red until a hand-made repin pull request
landed (22 of them in 12 days). The digests could only fail when an input moved
and the recording did not, which is exactly the change that carries no
information; every run already re-derives each golden from the current tree and
compares it.

Format 6 keeps the format version, operation, family, named deltas, the value
pool and the recording. Removed: the pin and fence, the three digest modules and
their test, the pin guard and its CI job, and the dead scenario `version` field
(the manifest reader now refuses `baseline` and `version` with a message).

- `pnpm --dir mobile rpc:record [<golden-id>...] [--prune]` records all or some
  goldens; orphans are listed, and deleted only with `--prune`. Every derived
  test title now starts with its golden id so an id selects it.
- `compareGolden` reports every difference in one failure (identity fields by
  name, the checkpoint list, each checkpoint/field/path grouped), keeps the
  final byte compare, and ends with the command to re-record that golden.
- `unhandled-recording.test.ts` now drives a detached rejection through
  `runRecording` into a checkpoint and the cleanup checkpoint; no golden carries
  one, and disconnecting the capture passed every suite before.
- Seam rules that existed only to keep a digest honest are gone; the
  mutant-reachability, register-completeness and one-exposure rules stay.
- CI: `Mobile tests on main` runs the whole mobile suite on every merge that
  touches mobile/, src/shared/, the root lockfile or the host RPC paths, since
  `verify` never runs on main. A new `Mobile RPC Recording Replay` workflow
  replays the recordings on pull requests that touch src/shared/ or the root
  lockfile without touching mobile/. `verify` writes the `rpc:diff` report to
  the job summary.

Proof: `rpc:diff` against the parent reports no recorded behaviour moved; each
golden only loses its ten header lines.

* test(mobile): check every recorded request against the host's params contract

The goldens script the host's replies, so a scenario could record a success
for a request the real host would refuse, and a desktop change that tightens a
params schema moved no golden at all.

`recorded-request-params.test.ts` parses every distinct request the corpus puts
on the wire with the host dispatcher's own `parseRpcRequestParams` and the
schema `rpc-params-catalog.generated.ts` binds to that method. It fails on a
method the host lacks, params it refuses, params sent to a method that takes
none (the dispatcher never reads them), and keys the schema silently strips
unless an inventory entry gives the reason; a stale entry fails too. Each rule
is also shown firing on a made-up request, since the corpus has no instance of
three of them. It imports the desktop dispatcher, so it sits beside the other
Node-side tests outside the RN test program, and the params-contract boundary
now exempts test files, which are never bundled.

It found twelve requests the host would refuse, all from invented fixture
values, not product code, fixed at their source:
- git.branchDiff sent `base-oid`/`head-oid`/`merge-base` where the host needs
  full object ids (diff-review and source-control adapters, and the branch
  compare replies in the manifest that feed them);
- an iOS push registration without `apnsEnvironment`, which a real iOS token
  always carries (`push-token.ts`); the adapter now defaults to `production`;
- `settings.update` given Linear's `assigned` filter as a GitHub preset, which
  the product type forbids; the scenario now picks `my-issues`;
- GitLab `projectRef` as a string where the host and the product type take
  `{ host, path }` (7 methods, 5 adapters and the manifest).

46 goldens move, and a decoded comparison of every one of them shows no change
other than those substitutions; `rpc:diff` lists them.

* ci(mobile): detect a mobile change without a SIGPIPE-prone grep pipe

Under the runner's pipefail, grep -q exiting on its first match SIGPIPEs git
diff on a long file list, so a large pull request touching mobile/ read as
uncovered and replayed the recordings a second time.

* test(mobile): drop comments that still describe the golden header and digests

Eleven adapters justified an import rule by the header a golden no longer
carries, and that rule's test is gone. The census failure now names the
rpc:record and --prune commands.

* test(mobile): refuse a golden that keeps a key no recording writes

Decoding dropped unknown top-level keys, so an old header left behind by a
hand-resolved merge conflict passed every compare unseen.

* ci(mobile): summarize RPC recording changes after a failed test step too

* test(mobile): stream rpc:record output instead of capturing it

A captured run stayed silent for its whole duration and clipped its tail,
where the failure summary sits, past 8 MB.

* test(ci): let the Ruby-gate contract skip the always-run RPC summary step

fef088d8f4 gave the summary step an `if: ${{ !cancelled() }}`, and this test lists every gated step in `verify` and expects each to be gated on the Ruby scope.

* test(mobile): replay only a golden file that is exactly what rpc:record writes

Replay compared two re-encodings of decoded values, so anything decoding drops (a leftover header
key, a hand edit) sat in the committed file uncompared; a key allow-list covered one case of that.
Replay now passes only if the file text equals the formatted golden for the run, sharing one
formatter with writeGolden, and keeps the field-level report as the failure message. The allow-list
goes; the value-based compareGolden stays for the bridged run, which has no file.

* test(mobile): end a corrupt or hand-edited golden's failure with the re-record command

A hand edit to a pooled value failed in decode with only "Golden value <hash> does not hash to its
pool key": no golden id and no command to fix it. readGolden now prefixes parse and decode failures
with the golden id and ends them with the rpc:record command. The rpc:diff header also said it
always exits 0; it exits non-zero when git or a golden cannot be read, and now says so.

* ci(mobile): run Mobile Checks on every src/shared and root lockfile change

Replaces the replay-only workflow: mobile imports hundreds of shared modules, so a shared edit can
move a golden or break mobile's typecheck, and the full job catches both before merge. A root
lockfile-only change skips the Ruby release checks, which read no root Node dependency.

* test(mobile): list or prune orphaned goldens even when the recording run fails

Orphans come from the manifest, not the run, so a failed or timed-out rpc:record still reports
them; the exit code stays non-zero. README: say what a failed replay reports (first differing
path per field, capped groups) and what rpc:diff compares with and without a base.

* test(mobile): pin the RPC recording goldens to LF so a CRLF checkout still replays

Replay now requires the committed golden text to equal exactly what rpc:record writes, which is
LF. A Windows checkout with core.autocrlf=true converted every golden to CRLF and failed all 790
with "holds the same recording but is not the file rpc:record writes for it".
2026-09-29 23:26:55 -07:00
Jinwoo Hong 9f4311598f fix(codex): trust the worktree Codex starts in, not a guessed repo root (#23937)
* fix(codex): trust the path Codex checks for bare-repo worktrees

Codex keys a linked worktree's trust on the main checkout only when that
checkout's .git leads back to the common git dir; otherwise (bare repo,
--separate-git-dir) it keys on the worktree itself. Orca always wrote the
main-checkout key, so Codex showed its trust prompt and worker-start
failed at agent_readiness.

Mirror trust.rs exactly, and pin it with a real-binary contract that
runs in the existing Codex contract CI job.

Fixes #23847

* fix(codex): trust the worktree path itself instead of mirroring trust.rs

Codex looks up the cwd's own [projects] entry before any repo root
(config_toml.rs get_active_project, loader decision_for_dir), so trusting
the workspace realpath satisfies every git layout. Drops the
resolve_root_git_project_for_trust mirror: simpler, cannot drift from
Codex, and never widens trust past the folder Orca launched in. Cost is
one config entry per worktree; entries older Orca wrote on main
checkouts stay valid.

The six real-git layout tests now assert the workspace key and, under
the contract, that real Codex starts workspaceWrite for each. The
contract probes the binary version once and fails at load when required
but missing; its CI path filter now includes config-toml-trust.
2026-09-29 20:27:02 -04:00
Jinwoo Hong 9420d49bcb fix(terminal): run Codex in Orca terminals without the shared background server (#23900) 2026-09-29 10:22:27 -07:00
Neil f8f656ca19 perf(ci): spend fewer concurrency slots per pull request (#23810)
A concurrency slot is charged per job, not per core, and the account's cap is
the scarce resource: standard runner minutes are free and unlimited on a public
repository. Two paths spent slots that bought nothing.

The unit matrix ran eight fixed shards averaging 6.5 minutes each, 3384
job-slots a day and 68% of all slot demand, while the arm pool queued 10.5
minutes at p95 — the queue was the oversharding. Five shards run the same work
in ~10.5 minutes each for three fewer slots per run.

Bun profile persistence escalated to all six platforms on `config/`,
`resources/` and `.github/` wholesale, which took 36.5% of the last 1100
commits through the full matrix where a platform-flavoured predicate takes 19%.
A pull request now qualifies one platform unless the change is platform-
flavoured, and the push to main re-qualifies all six, so an unescalated miss
surfaces minutes after merge rather than at the next cron. Missing changed-file
evidence and an unavailable dependency graph still fail closed to all six.
2026-09-29 00:13:33 -07:00
Neil f6324f242a chore(ci): stop auto-filing community PRs onto the project board (#23796)
Removes the Track Community PRs workflow. Community pull requests will no
longer be added to project stablyai/13 automatically.
2026-09-28 22:15:36 -07:00
Neil 2ea3fb1d46 perf(ci): take advisory unit-selection evidence off the gate (#23776)
selection_evidence is continue-on-error on both the job and its comparison step,
so it can never fail a PR -- it downloads the shard reports, compares selection
against the full results and uploads a review artifact. But a caller's
`needs: test` waits for every job in the called workflow, so living inside
unit-tests.yml it held verify for ~36s after the last shard finished.

It moves to its own reusable workflow called as a sibling, so it still runs on
every PR and still uploads its artifact, but verify no longer waits for it. It is
deliberately absent from verify's needs, and a contract test pins both that and
its advisory status so it cannot drift back onto the critical path.

Measured on a recent run: the shards finished, then selection_evidence ran 36s,
then verify 3s. Only the last of those gates anything.
2026-09-28 22:13:35 -07:00
Brennan Benson c5fc0c6f26 fix(ci): keep a squash-merged RPC recording pin reachable through its pull request (#23720)
* fix(ci): keep a squash-merged RPC recording pin reachable through its pull request

Main's "RPC recording pin" check has been red since #22762: that branch pinned
the recording corpus to its own commit 03995ae, and the squash-merge left that
commit out of main's history. Every behaviour-change squash did the same, and
each needed a hand-made repin PR to clear it (#23565, #23535, #23046 and more).

The guard now accepts a pin that is either in this history or in the head of the
pull request whose squash wrote it into the manifest. It finds that pull request
from the `(#n)` subject of the commit that added the pin and fetches
`refs/pull/<n>/head`, which GitHub keeps after the branch is deleted. The
reproduce step uses the same lookup, so it can still check the pinned tree out.

* fix(ci): give the recording pin lookup room to walk a blobless clone

In CI's blobless clone, `git log -S` fetches the manifest's blobs one commit at a
time, a few seconds each. Under the 30 s process default the walk was killed after
a handful of manifest commits, which main's history already exceeds (up to 7
manifest commits between a pin landing and the next pin change), and the guard
then failed with an empty "Could not find the commit that pinned" error. The
lookup and the pull request fetch now carry explicit budgets and say when they
timed out.

The not-an-ancestor instruction now names the pull request whose head was
checked, or says the commit that pinned it names none.

Adds the two merge-preview shapes the guard runs on: a branch opened after a
squash resolves main's pin through the squash's pull request, and a branch whose
rebase dropped its own pinned commit fails on its pull request instead of on main.

* fix(mobile): tell a missing recording pin apart from product drift

After a squash the pinned commit can live only in its pull request's head, so a
clone that never fetched it makes `git diff --quiet <baseline>` exit 128. The
recorder reported that as "Product sources or lockfile differ from the pinned
main baseline", which sends the developer to repin a tree that may match. It now
prints git's error and the command that fetches the pin.

* fix(ci): ask GitHub which pull request holds a squash-dropped recording pin

The recording pin guard found the pull request that keeps a squash-dropped
pin by walking main's first-parent history for the commit that wrote the pin
into the manifest and reading "(#n)" off its subject. A merger who edits the
squash title loses the number, and the push to main turns red anyway. That
already happened on main: of the 22 squashes that left a pin outside main's
history, #21674's title had no "(#n)".

The guard now asks GitHub for the pull requests associated with the pinned
commit (GET /repos/{owner}/{repo}/commits/{sha}/pulls) and, for each in turn,
fetches refs/pull/<n>/head and accepts only when git proves the pin is an
ancestor of that head. GitHub only nominates candidates, so a wrong answer can
fail the guard but never pass it. The endpoint named the right pull request
for all 22 historical cases, #21674 included, and names none for commits a
force-push orphaned.

This removes the pickaxe walk, its 600 s budget and its lazy blob fetches in
a blobless clone, the first-parent subtlety, and the subject regex. A revert
that restores an older pull-request-only pin now resolves too, because the
lookup is by the pin itself rather than by the commit that last wrote it.

CI passes the job token to both guard steps and grants the job
pull-requests: read. Local runs work without a token on this public repo and
send GITHUB_TOKEN or GH_TOKEN when set. A failed lookup throws with the HTTP
status, and names the rate limit when an unauthenticated call is refused.
2026-09-28 21:17:07 -07:00
Neil ec9f35e2ee perf(ci): plan the unit shards before the static-analysis gate instead of behind it (#23743)
A caller's `needs` gate the whole called workflow, so while the plan job lived in
unit-tests.yml it could not start until static analysis and typecheck had both
finished and passed -- and the shard matrix then waited on it. The two hops were
serial when they did not need to be: planning reads the checkout, a git diff
against HEAD^1, the import graph and the checked-in timing baseline in
config/scripts/ci-shard-timings.json, and consumes nothing that static analysis,
typecheck or the native-cache primer produce.

Planning moves to its own reusable workflow so pr.yml can run it against
code_paths alone, overlapping it with the gate. Measured across 99 runs, the
shard matrix is created a median 93s earlier (p25 47s, p90 241s, never later).
Planning stays a required predecessor of the shards, so an empty assignment
cannot expand the matrix.

The gate itself is deliberately left in place. It fires on 22% of runs, and the
shard queue wait knees hard above ~9 concurrent ARM jobs -- 4s median below that
against 218s at 15-19 -- so admitting 8 doomed shards per failed run would cost
more in queue pressure than it returns in latency.

Cost is one 37s ubuntu-latest job, which does not touch the ARM pool the shards
contend for.

A planning failure still fails the PR: the shards are skipped, and verify's
check_job requires success whenever the classifier says tests should run, so it
reports `test: expected success, got skipped`.
2026-09-28 19:22:10 -07:00
Brennan Benson 2ca4ecbc61 feat(orchestration): let a structured chat run orchestration as itself (#22568)
* feat(orchestration): inject the Orca session id into structured children and let the CLI act as it

Every structured session's child (native Claude, native Codex, and the terminal
view) carries ORCA_AGENT_SESSION_ID and reaches the Orca CLI. The CLI sends the id
in the orchestration envelope; when present it is the caller, and a caller flag
naming anyone else is refused before any request. The id is stripped from
inherited PTY env and from the SSH host-CLI passthrough, and crosses into WSL so
the host can refuse the cross-host claim.

* test(orchestration): pin session id injection for native Claude, native Codex, the terminal view, WSL, PTY inheritance and SSH

* test(orchestration): pin one caller precedence rule across every CLI verb that names its caller

Adds the per-verb table (flagless acts as the session; a conflicting --from or
--terminal is refused before any request; the session's own spellings are
accepted), the enumerated guess population with its positive control, the
structured worker's own handle, the identity-less refusal for an older child,
the unchanged terminal agent, and the envelope. dispatch-show's --from only fills
preview text, so it passes through unfenced and a session's flagless preview
names the address the real dispatch writes.

* refactor(orchestration): keep the identity-less marker reader to the marker; the id is checked first

* test(orchestration): pin that a host refusal of the session surfaces verbatim from the CLI

* fix(orchestration): keep the identity-less marker beside the id for CLIs that predate it

A CLI older than the id, reached through a global install when a shell rc resets
PATH, would otherwise guess a sibling's terminal in a chat that no longer carries
the marker. It refuses on the marker instead; a current CLI checks the id first,
so the marker never makes a session with an id identity-less.

* fix(orchestration): refuse a conflicting --from on gate-list and task-list scoped by --run

A --run listing needs no caller, so both handlers skipped the resolver and a
--from naming another actor was dropped silently under a session. The conflict
check now runs on that branch too; terminal callers are unchanged.

* fix(orchestration): name this app's CLI by absolute path for a structured session's login shells

A provider can run each command in a login shell: Codex runs zsh -lc, and the
profile rebuilds PATH, putting a global install (possibly an older Orca) ahead of
the directory Orca prepended. ORCA_CLI_COMMAND, which an agent resolves the CLI
from first, is now the absolute launcher in that directory (the native launcher
on Windows), so no shell's startup files can swap it. The PATH prepend stays for
shells that read no profile. Found by the live coordinator run of the next PR.

* test(orchestration): pin a structured worker's CLI command as this app's absolute launcher

* test(orchestration): run the zsh login-shell arm in the real-shell lane that installs zsh

The ordinary Linux unit lane has no /bin/zsh, so the zsh arm failed there with
ENOENT. It moves to a live-shell file registered in the shell-contracts lane; the
bash arm keeps running in every lane. The lane guard's detector now also sees a
zsh spawned through the ProcessSpec program field, which is how this test
escaped it.

* fix(orchestration): omit a structured child's CLI command when no launcher resolves, and pin its instance

A bare `orca` fallback named GNOME's screen reader on packaged Linux, and an inherited value named
another app's CLI. The builder now deletes any inherited value, sets the absolute launcher only when
one resolved, and pins ORCA_USER_DATA_PATH so a current CLI dials the instance that minted the id.
Renames the marker reader to hasStructuredSessionMarker and records why the terminal view carries
the id without the marker.

* fix(terminal): name this app's CLI launcher by absolute path in every local terminal

ORCA_CLI_COMMAND meant three things by lane: an absolute launcher for a structured session, a bare
name for WSL, and nothing for any other terminal, so a structured session's terminal view lost it.
Local terminals now get the same absolute launcher the structured lane gets; WSL keeps its guest
command name, and a terminal whose launcher does not resolve still gets none.

* feat(cli): hand a command to the session's own CLI when another Orca CLI was invoked

A login shell can reorder PATH behind a global install, and an agent or its helper script can run
bare `orca`, so the binary that answered depended on the agent following instructions. Orca's
packaged launchers and bare-orca shims now export ORCA_CLI_SELF (outermost wins). At the CLI entry,
when it names a different launcher than ORCA_CLI_COMMAND, the command re-runs once through the named
launcher with ORCA_CLI_REEXEC=1 and exits with its status; both variables are consumed so no child
inherits them. Dev launchers export no self on purpose, WSL and SSH names never qualify, and a
launcher that cannot start leaves the command to run here. The Windows launcher no longer rewrites
ORCA_CLI_COMMAND; the legacy ask protocol normalizes its resume command itself.

* refactor(orchestration): declare which flag names the caller on each spec and refuse at the CLI entry

Each handler hand-classified its --from/--terminal as the caller or a target, and the refusal of a
conflicting caller flag ran inside the caller resolver plus two standalone calls for --run listings,
so a new verb that read its flag raw would pass a sibling's handle to a pre-session host. Specs now
declare identityFlagRoles, the CLI entry refuses a conflicting caller flag once from the spec, the
resolver only applies the id-wins rule, and a test fails any orchestration verb that accepts --from
or --terminal without classifying it.

* perf(cli): keep the session caller check off the actor codec's module graph

The check runs at the CLI entry for every command, and the actor codec pulls zod through the session
record. Compare the session's own spellings as plain strings instead.

* refactor(cli): spell a session's address from the one prefix constant, off the codec's module graph

The Orca session address prefix moves to a leaf module with no imports, re-exported by
the address codec, so the CLI entry check derives `session:<id>` from that constant
instead of re-typing it and still stays off the codec's zod graph. Prose and test names
say caller or Orca session id, not actor.

* refactor(orchestration): drop the session id's terminal-view spawn now that the handoff is gone

The terminal handoff was removed, so no terminal is ever a structured session:
- delete the terminal-view identity env and its WSL passthrough, and their tests;
- strip the session caller keys from every terminal's env unconditionally;
- the CLI's own-address spelling moves beside the injected id in src/shared, with
  a test pinning it to the address the host's party resolver gives that session.

* fix(terminal): run the Codex launch preflight through the CLI the terminal names

Packaged Linux names the userData shim in ORCA_CLI_COMMAND, while the preflight
ran the bundled launcher behind it. The CLI saw a different launcher and handed
the preflight off to the shim, booting Electron twice before every codex launch.

* revert(terminal): keep terminals on main's ORCA_CLI_COMMAND and Codex preflight

Only a structured session needs an absolute ORCA_CLI_COMMAND; local terminals go back to
naming none (WSL keeps its guest command), and the Codex launch preflight goes back to the
bundled launcher. The CLI handoff is scoped to sessions, so a terminal's preflight can no
longer be handed off and start Electron twice.

This reverts commit d2cefb6c03 and commit dd2853a5a9.

* fix(cli): hand off to the session's CLI only inside a structured session

The handoff ran whenever an Orca launcher's ORCA_CLI_SELF differed from an absolute
ORCA_CLI_COMMAND, so any process with both - a terminal, a script - ran another install's CLI
instead of the one invoked: a beta's --version lied, and an AppImage command from a terminal
that outlived its Orca failed. It now requires the injected session id, the identity it exists
to deliver. The launcher variables are still consumed in every process.

* fix(cli): name the packaged Windows command after the handoff decision

The launcher stopped writing orca/orca-ide over ORCA_CLI_COMMAND so the handoff could see a
session's absolute launcher, which also changed what every Windows terminal's CLI read. The CLI
entry now applies the launcher's rule itself once the handoff is decided, so terminals and the
legacy ask resume command see exactly what they saw before, and the resume-command reader
goes back to its original form.

* refactor(cli): decide the session handoff from the CLI's own entry, not a launcher export

Every packaged launcher, shim and dispatcher exported ORCA_CLI_SELF so the CLI could tell which
launcher ran it, and compared that with the session's ORCA_CLI_COMMAND. Two launchers of the same
app are different files, so a session that reached its own app through a global orca-ide on Linux
still handed off and started Electron twice, and the export rode artifacts every terminal uses.

A structured session now also names the JS entry its launcher runs (ORCA_SESSION_CLI_ENTRY), and
the CLI compares its own argv entry with it: any launcher of the same app stays, another install
hands off. The launcher scripts, Linux shim and dispatcher go back to main; the Windows launcher
keeps only leaving ORCA_CLI_COMMAND for the CLI to name after the handoff decision.

* refactor(cli): drop the session CLI handoff; the pinned instance and injected id already bind any current CLI

Every current Orca CLI dials the instance ORCA_USER_DATA_PATH names and sends the injected
session id in the orchestration envelope, so a bare `orca` that reaches another install's
current CLI already acts as the session. An older CLI has no handoff code and refuses on the
marker. The handoff only lined up versions between two current CLIs, and comparing two
separately derived paths kept misfiring (an AppImage's mount against its registered
extraction started the CLI twice on every call).

Removes the re-exec, ORCA_SESSION_CLI_ENTRY and ORCA_CLI_REEXEC, and the CLI-side Windows
command naming; the packaged Windows launcher rewrites ORCA_CLI_COMMAND again, as on main,
inside its own process only. resolveHostCliEntryPath goes back to the SSH passthrough.

* test(orchestration): say why the registered worker case pins the handle, now that every session's env is populated
2026-09-28 15:19:44 -07:00
Neil 8c61a5df1f fix(windows): require signed release binaries and identify CLI launcher (#23680) 2026-09-28 14:34:08 -07:00
Neil 0f52bb8be5 perf(ci): use four ARM test workers and remove repeated compilation (#23685)
* ci: benchmark per-job Node compile caching on full unit shards

* ci: measure unit shards with three and four workers

* ci: benchmark localization extraction CLI patch

* perf(build): reuse identical relay bundles across platforms

* ci: compare Vitest 4 and 5 on complete ARM shards

* perf(ci): upgrade localization extraction to skip irrelevant syntax walks

* perf(ci): use all four ARM cores and remove benchmark workflows

* ci: preserve failures while capturing unit source revision

* fix(ci): preserve commented and escaped localization calls

* ci: remove corrected localization benchmark harness
2026-09-28 14:05:30 -07:00
Jinjing 95a16e3f67 fix(release): stop the release policy from deleting pipeline-cut releases (#23669)
* fix(release): stop the release policy from deleting pipeline-cut releases

The policy judged a release by who created the release object. Cut Release
reuses an existing draft, so a CI-built v1.4.216 whose draft a person had
created was deleted (tag included) when its notes were edited, and Latest
fell back to v1.4.214 because v1.4.215 was also published by a person.

- Authorize a desktop release when its annotated tag was created by the
  release pipeline and points at its `release: vX` commit, not only by author.
- Only delete on `published`; an edit never deletes a release or tag.
- Pick Latest from the highest authorized stable using the same check.
- Move the policy into config/scripts/release-policy.mjs with tests.

* fix(release): load the policy module from the tagged commit

Release events run the workflow file from the tag's commit, so checking out
the default branch could pair an old workflow with a newer module.
2026-09-28 12:03:58 -07:00
Neil cf20423ff3 ci: skip unrelated installs and share xterm build dependencies (#23607)
* ci: pilot shared xterm installed dependencies

* ci: bound xterm cache production to verified main entries

* ci: benchmark xterm reuse on the production ARM runner

* ci: avoid installing Orca dependencies for standalone xterm checks

* ci: use Node-only setup in the production xterm job

* ci: remove completed xterm benchmark workflow
2026-09-28 03:46:22 -07:00
Neil 0fe8974354 perf(ci): move six more jobs to the free ARM runner (#23594)
* perf(ci): move six more jobs to the free ARM runner

Follows the static-analysis move, which measured 172s to 128s. Each of these was
checked for an x86 requirement rather than assumed portable.

pr.yml:
  cross-version-wire        source-only, tagged checkout plus in-process vitest
  managed_hook_node18       Node 18 publishes linux-arm64; the per-platform
                            runtime files are read as data, so host arch is moot
  codex_index_heal_contract @openai/codex ships @openai/codex-linux-arm64
  shell_contracts           fish 4.x is published for noble/arm64, zsh is in the
                            arm64 archive, so the fatal fish-4 gate still holds

mobile.yml:
  verify                    209s of its 298s is Vitest; no Android SDK, emulator,
                            gradle, Hermes or Watchman, no docker, no artifacts.
                            Gemfile.lock lists the generic `ruby` platform, so
                            frozen bundler resolves without an aarch64-linux entry
  recording-pin             pure Node plus git; the golden comparison masks
                            `platform`

Left on x86 deliberately:
  package                   builds --x64 targets, its docker gates are
                            --platform linux/amd64, and it is where the glibc
                            floor check runs. node-pty's .symver pin is
                            arch-specific, so flipping would validate the arm64
                            pin and stop validating the shipped x64 one
  orcad_browser             Google ships no Linux arm64 Chrome
  mobile_web_app            same Chrome wall; its render check fails closed
  git_compatibility         its cache key carries runner.arch and the warmer is
                            x86, so flipping alone means a cold `make git` every
                            run
  relay_integration         no technical blocker, but x86 relay coverage is a
                            documented placement and the reusable workflow has no
                            per-job runner input
  e2e and the ssh lanes     the ssh jobs would silently retarget the tested
                            remote from linux-x64 to linux-arm64

* perf(ci): move xterm_patch_sync to ARM too

The patch check rebuilds 4 packages x 2 builds and byte-compares against the
checked-in bundles. Ran it on darwin-arm64: exit 0, in sync at 1362743 bytes,
with the full fetch-and-rebuild path exercised rather than a short-circuit. The
bundles were generated on Linux x64 and reproduce byte-for-byte on a different
arch and a different OS, so the output is host-independent.
2026-09-28 02:42:41 -07:00
Neil 2f8f4f576d perf(ci): run static analysis on the free ARM runner (#23576)
Measured 128s against 172s on ubuntu-latest, with every compute-bound step
faster: type-aware 24s to 15s, anti-slop 28s to 19s, localization extraction
67s to 46s, and the orcad terminal smoke 39s to 14s. Checkout and the install
were unchanged at 12s and 16s.

The toolchain resolves on arm64: both lint engines ship linux-arm64 bindings
(@oxlint/binding-linux-arm64-gnu, @oxlint-tsgolint/linux-arm64), and
build-orcad-bun.mjs derives its target from process.arch. The orcad smoke
booting and round-tripping a real PTY is the evidence that node-pty compiled
and that Bun, the bundled ripgrep and @parcel/watcher all resolved.

The runner is free for public repositories, the same one the typecheck job
already uses.
2026-09-28 01:59:34 -07:00