mirror of
https://github.com/stablyai/orca.git
synced 2026-10-02 16:02:15 +00:00
8ff6296bc7844fc049357e4e6494aa3b42df8eba
514
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
8ff6296bc7 |
Speed up serializer checks and keep native caches stable (#24476)
* Reuse serializer oracle cells and isolate native cache policy * Preserve native cache post-save paths and record hosted oracle gain * Record native cache reuse and separate cancel-test startup budget |
||
|
|
444f1952c7 |
ci: run every cross-version wire test, picked up by folder so new ones can't be skipped (#24499)
* ci(cross-version-wire): run the whole directory so no compatibility test is left out Three cross-version tests ran in no CI job because the job named its files by hand. Run the directory instead, ratchet that every file kept out of the unit shards runs in some PR job, and re-run the job when the modules the newly running tests guard change. * test(cross-version): give the orchestration downgrade test its siblings' 120 s budget * ci(unit-exclusion): count only merge-gating jobs, and require each excluded file's job to fire on it The coverage check counted any pr.yml job, including e2e, terminal IME and Windows WSL, which are left out of verify.needs and so cannot block a merge. It now reads verify.needs and the reusable workflows those jobs call. It also only proved that some step names each excluded file, not that the job runs when the file changes. The structured-session zsh login-shell test runs only in shell_contracts, whose path trigger matched neither it, its harness nor its subject, so a PR touching only those ran it nowhere. The check now asserts a change to each excluded file fires a gating job that names it, and the shell trigger gains those three paths. * ci(cross-version-wire): trigger on the turn-outcome vocabulary and the schema version-skew resolver A change confined to src/shared/agent-turn-outcome (the arms a newer host publishes) or to orchestration-schema-version-skew (how current code reopens a downgraded database) skipped the job whose tests guard exactly those contracts. Also corrects the publish/read direction in the turn-end comment. * test(cross-version): state why the orchestration downgrade test needs 120 s * test(ci): glob the unit tree once for the unit-exclusion coverage checks |
||
|
|
f69052e113 | Reuse qualified Windows server builds and dependency verification records (#24448) | ||
|
|
976dc00337 |
fix(native-chat): Stop's pause is worked out from the chat's history, so a steered message is never re-sent (#24072)
* test(native-chat): a Stop over a card sent now into the running turn keeps it paused
Red on main: Codex's turn end withdraws the steered hand-off and the queue
sends the card again as a new host turn, with no pause recorded.
* fix(native-chat): a Stop's queue pause holds a card whose hand-off is still unanswered
A card sent now into the running turn was still pending when Stop judged the
pause, so nothing was recorded; the interrupt then withdrew the hand-off, the
card went back to waiting unpaused, and the queue sent it again as a new turn.
The pause now counts a hand-off that may still return to waiting, judged with
the appended row applied, so a withdrawal lands under the pause and an
acceptance retires it in that same write. Codex and Claude both hit it.
* test(native-chat): the Claude re-send case fails on its diff, inside the test's budget
* test(native-chat): a pause held by an unanswered hand-off ends on every path that ends it
The provider's answer, the provider dying, the chat closing, a restart, a
withdrawal still owed at open, and a /clear (refused until the hand-off ends,
then carrying every waiting card paused 'cleared'); each ends with the queue
sending again.
* fix(native-chat): narrow the pause's settled hand-off, and assert the queued receipt's card
* refactor(native-chat): derive the queue's pause from Stop and Resume journal rows
Stop now appends one journal row where it takes effect, before the interrupt,
whatever the queue holds; Resume appends its own. The pause is a pure function
of the fold: the latest Stop with no later Resume and no later accepted turn a
person asked for. A /clear's carried cards name their source, which is the
replacement's 'cleared' pause. Host-origin turns never lift either.
One predicate decides which cards a pause holds; by default every waiting card
without a hold of its own, including one queued after the Stop. The drain's
consume re-judges it inside its own transaction.
The rows are tombstones of an id no item takes, carrying the mark: a released
host reads an unknown row kind as corruption and truncates the journal there.
Deletes the stored pause (recordPause, the retire hook on every appended row,
the settle-before-record step, mayReturnToWaiting and its row overlay) and the
tests that only proved it retires. The queued_message_pauses table stays in the
schema, unread and unwritten, for downgrade safety.
* fix(native-chat): a card queued after a Stop sends normally, never ahead of held ones
A Stop's pause now holds only the cards queued before its row, plus a steer it
withdrew, which returns to its own place. Each card records the journal
position it was queued at, and the one hold rule compares that with the Stop
row. A card queued after the Stop is a new instruction: it sends as usual, but
the drain still stops at the first held card, so it never overtakes them.
/clear's pause holds the cards it carried. Holding every card again is a
one-line switch in that rule.
* fix(native-chat): the queue's own send re-checks the no-overtake rule in its transaction
The drain's pick and its consume now read one function, nextSendableQueuedCard,
so a Stop row that lands between them holds a newer card behind an older held
one exactly as the pick would. Notes why Stop and Resume ride a tombstone row.
* fix(native-chat): stop creating the unused queue pause table
The queue's pause is derived from journal rows, so nothing reads or writes
queued_message_pauses. It was still created on every open "for downgrade
safety", but an older build creates it itself when it opens the database, so
the table only sat empty in every new database. The tests now pin that no
pause table exists.
* fix(native-chat): a Stop's pause never hides the restart pause
A Stop holds only the cards queued before it. The pause derivation still
returned the Stop alone whenever it was in force, so the restart pause was
never considered: a card queued after the Stop, written by a host process
that has since exited, sent by itself after Orca restarted, with no pause
header and no Resume. A /clear pause that held nothing could hide it the
same way.
Every pause in force is now derived. A card is held if any of them holds
it, and it names the first that does. The drain's pick, the consume
transaction's re-check and the published header all read that one rule;
the header names the pause holding the first card Resume would send.
* test(native-chat): pin the Stop's no-resend, lift and held-card rules
- The Claude and Codex Stop-withdraws-a-steer tests checked "not sent
again" at one instant, before a queue ignoring the pause re-sends. They
now wait for the stopped turn to end and re-check after a quiet window.
- The deleted-card test read a card queued after the Stop, which sends
whether or not a person's turn lifts it; it now reads the Stop's pause
before and after that turn.
- Unit cases pin that a Stop holds a card with no recorded position and one
queued before a rewind.
* refactor(native-chat): a Stop writes one Stop event with its reason, turn and caller
The Stop row that paused the queue becomes the general Stop event
{ reason, turnId?, at, caller? }, whose reason is the host's existing stop
cause. It still rides a tombstone of a host-only id (a released host deletes
the journal from the first unknown row kind), and Resume keeps its own marker
on its own id. Only a person's Stop (reason user-stop) pauses the queue.
* test(native-chat): a rewind keeps a lifted /clear pause lifted and restates the same Stop event
* test(native-chat): pin that Stop and Resume rows never reach apps or count as history
* test(native-chat): only a person's Stop event pauses the queue
* test(native-chat): pin that a Stop's event precedes the interrupt and the at-start stop
Through the real host: the event names the turn and who asked and is in the
journal when the interrupt reaches the agent; at an agent still starting it is
there before the start is ended and holds a card queued before it; an idle Stop
writes one only when it withdrew a send; and the queue's claim re-judges a
pause that landed after its pick.
* test(native-chat): a card held at a starting agent is checked before the Stop's timing
Also says precisely what the claim's in-transaction pause check defends
against: the Stop and the drain share one serialized lane.
* test(native-chat): a released build keeps and folds a journal holding Stop events
Replays this build's rows from the released build's own journal database: every
row is kept, the history after the Stop still folds, and an older client is sent
only removed ids no item uses.
* style(native-chat): format the Stop event changes
* test(native-chat): type the released build's exports through one checked helper
* fix(native-chat): the Stop/Resume row guard narrows to those tombstones only
* test(native-chat): run the Stop-event downgrade test in CI, and cover a writable downgrade
The Stop-event downgrade test ran in no CI lane: unit shards exclude the
cross-version folder, and the cross-version lane runs a fixed file list that
did not name it. It is now on that list.
Its only case replayed the rows into a release's own fresh database, because
that release cannot open the current host database. A second case opens the
journal this build wrote with a main build that shares the database: it opens
writable, keeps every row, appends, and this build then reopens it with the
person's Stop still pausing the queue.
* fix(native-chat): a Stop that stops nothing new writes no Stop event
A Stop reaching a running agent wrote a Stop event on every press. Two
presses before the first interrupt landed wrote two events, so a card
queued between them counted as before the latest Stop and was held,
though a card queued after a Stop should send normally. A Stop naming a
turn that had already ended, as a phone sends late, also wrote an event
for a turn it never stopped.
It now writes one only when it withdrew a queued send, or stops something
no event records yet: not a turn the journal no longer runs, and not the
live turn a Stop still in force already names, unless a card was handed
over into it since, which this Stop's interrupt sends back and must hold.
The interrupt and the "already finished" note are unchanged. A Stop at a
starting agent still always writes.
* test(native-chat): pin that a later host, eviction or close Stop never lifts a person's Stop
* chore(native-chat): put each Stop-row doc on its own declaration, and say only user-stop is journaled
* fix(native-chat): any later Stop event ends a person's Stop pause
A person's Stop paused the queue until their next accepted turn or Resume,
and a later Stop of another reason (the host stopping the agent, an
eviction, a close) was ignored. Now the pause is the latest Stop event's:
a later Stop of any reason ends a person's pause, and only a person's Stop
pauses. The fold keeps the latest Stop event whatever its reason.
An eviction of a resting chat writes no Stop event (a Stop that stops
nothing writes nothing), so it cannot release held cards; a test pins that
no event means no lift.
* fix(native-chat): a second Stop press is a repeat even when the first came before the turn showed
A Stop pressed before the agent's turn shows in the journal (before
Claude's echo, or before Codex opens the turn) records no turn. A second
press once the turn showed compared that missing turn with the live one,
wrote a second Stop event, and held a card queued between the presses.
A repeat is now judged by what was sent since the Stop in force: with
nothing sent after it (a refused send aside), a Stop that named no turn,
or named the live one, is repeated and writes nothing. Anything sent since
and not refused, including a send whose fate is unknown, makes the new
press write, since its interrupt may send that card back to waiting.
Tests: the two-press case across the turn showing; a steer between the
presses settled unknown; and a Stop naming a turn that ended while the next
card is sent but shows no turn yet, which writes and holds that card. The
fold test that claimed an eviction path is renamed.
* fix(native-chat): the queue's pause ignores a Stop or Resume row holding a value no build writes
A Stop or Resume row's value is read from disk with no shape check, and
the pause fold stored whatever it found. A stored `stopEvent: null` would
then throw on every pause check for that chat: the queue's pick, its
send, and every queue update to clients. No build writes such a row, so
this is hardening.
The fold now reads a Stop only when it is an object with a string reason
and a finite time, and a Resume only when it is `true`. Anything else is
ignored: it pauses nothing and ends nothing. The row is still not treated
as malformed, which could cut the history short.
* refactor(native-chat): one reading of a Stop's turn for its event and its note
A Stop's event and its note each worked out the same two facts on their own:
which turn the Stop is about (the one it named, else the one running), and
whether a named turn is the one the journal shows running. The event decides
before the interrupt; the note and whether the session ends decide after the
provider's answer, so those decisions stay separate, but the facts they read
are now one helper each in structured-agent-session-turn-stop-notes.ts:
structuredAgentSessionStoppedTurnId and
structuredAgentSessionStopNamesTurnNotLive. The event's turn, the note's key,
the session-ending condition, the running-command check and the repeat check
all read them. No behavior change.
Tests: a Stop naming no turn records the running turn on its event, and
rewrites that turn's note as a Stop naming it does.
* refactor(native-chat): a failed-interrupt Stop reads its turn through the shared helper
The new branch that ends a Codex child after a failed interrupt asked
whether the Stop's turn still runs with `turnId ?? liveTurnId`, a third
copy of "the turn a Stop is about". It now reads
structuredAgentSessionStoppedTurnId, the value the note key already uses,
read at the same point before the cancel. No behavior change.
Test: a Codex Stop whose interrupt failed ends the child, holds the card
queued before it with the queue paused, and writes its Stop event before
the turn's end.
|
||
|
|
ae41eb414a |
fix(terminal): give plain fish tabs Orca's codex function without changing fish's startup (#24284)
* fix(terminal): give plain fish tabs Orca's codex function without changing fish's startup A `codex` typed into a plain fish tab ran without --no-daemon because only wrapped fish tabs (startup command / ready marker) got Orca's codex function. Plain fish spawns now prepend an Orca data dir to XDG_DATA_DIRS and record the exact prefix in ORCA_FISH_XDG_DATA_DIRS_PREFIX. Fish sources the dir's fish/vendor_conf.d snippet, which first restores XDG_DATA_DIRS (unset again if it was unset), erases the marker, drops its dir from fish's derived vendor/function/ completion paths, then defines the shared fish codex function at the first prompt so the user's config.fish still wins. fish argv is unchanged; wrapped tabs keep their existing -C path. A local fallback to another shell restores the user's XDG_DATA_DIRS instead of deleting it. Bumps the terminal daemon protocol to v39 so new tabs move to a daemon that injects the env; v38 owners stay attachable. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(fish): skip the XDG handoff for -N/--no-config and empty XDG_DATA_DIRS fish never reads vendor_conf.d under -N/--no-config (also abbreviated or clustered), so the snippet could not undo the prefix; and the restore cannot tell an empty XDG_DATA_DIRS from an unset one. Both now launch untouched. Run the real-fish handoff tests in the shell contracts job, where fish is required, so they no longer skip in CI. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test(fish): compare the unset-restore case against a fish without Orca Ubuntu runners ship snapd's fish vendor snippet, which sets XDG_DATA_DIRS on every fish start, so "unset" was never the right oracle there. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(fish): treat an empty XDG_DATA_DIRS like unset so the tab still gets the codex hook Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(pty): put back the user's own launch env on a shell fallback The primary shell's launch config now records the pre-launch value of each key it writes. A fallback shell restores those values (unsetting keys that had none) instead of deleting the keys, which hands back an inherited XDG_DATA_DIRS after a fish fallback and an inherited ZDOTDIR after a zsh->bash fallback, with no per-shell special case. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * refactor(fish): drop the Node restore twin and simplify the vendor snippet - Remove restoreFishXdgDataDirs; the generic fallback restore covers it. - Snippet: read ":$XDG_DATA_DIRS:" directly and filter Orca's vendor dirs with one string match per variable. - Require inheritedXdgDataDirs in both getShellLaunchConfig option shapes. - Drop the test-only FISH_XDG_DATA_DIRS_HANDOFF_DAEMON_PROTOCOL_VERSION. - Fix stale fish comments and trim redundant -N launch cases. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * refactor(fish): stop scrubbing fish's lookup paths after the handoff Only XDG_DATA_DIRS is restored, by exact prefix; Orca's dir holds nothing but this snippet, so leaving it on fish's derived paths loads nothing else and drops the glob match. * docs(fish): drop the comment for the removed vendor-dir cleanup --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
07dad6739a |
refactor(relay): sample fleet health inside the same-cap roll instead of a separate monitor run (#24443)
* refactor(relay): sample fleet health inside the same-cap roll instead of a separate monitor run A same-cap wave no longer consumes a 15-minute monitor dry-run and its sealed, single-use, five-minute-fresh evidence. Each apply wave now samples fleet health itself right before isolation, with the monitor's evaluator, thresholds, and tolerances, for a window sized to the cell's host count (3/5/8 min), plus three lookback rules: no cell container exit in 10 min, no minute over 500 director 503s in 10 min, and director concurrency p99 within the monitor bar over 4 min. Removes the monitor-run inputs, the gate's consume/authorize steps, the break-glass override, and the same-cap-only authorization shapes in relay-monitor-evidence.mjs. The monitor workflow and the rehome enable path are unchanged. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): bound the pre-drain sample overrun and keep the drain token fresh Review follow-ups: alternating tolerated readings could hold the sample open until its step timeout, so cap the overrun at three samples past the window; record why a read failed; mint a fresh admin ID token for the drain after the sample; raise the job timeout to 90 min so a long sample cannot cancel the job past the failsafe. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * feat(relay): exempt the rolled cell and existing-only cells from the pre-drain crash rule The exit rule counted every relay container exit fleet-wide, so a cell that crashes every few hours (c25, 12 a week) blocked the very roll that fixes it, and existing-only legacy cells (c5, 15 a week) blocked rolls they take no part in. Exits are now grouped by instance, each instance is named by its own newest runtime-metrics log line, and only exits on general or migration-only cells other than the target count. An exit no configured cell can be named for trips the rule; a failed lookup is a failed read. relay-observability.tf joins the evidence-code set because the rule depends on its filter. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * test(relay): cover re-asking for an unnamed exiting instance; note the boot-exit risk Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
051ca4d34e |
feat(relay): declare US cells c32 and c33 at the 3,000-host shape (#24444)
* feat(relay): declare US cells c32 and c33 at the 3,000-host shape Declares two us-central1 cells at the Asia shape (cap 3000, 6000 request units, e2-standard-4) with the US default pool of 10, and generalises the Asia topology and admission ladder to derive each wave's region from its reviewed zone, leaving every Asia wave's behaviour unchanged. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * docs(relay): note the US canary tie-break and leave the fleet pool list to promotion Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): plan C32 and C33 as one topology wave The live-image overlay refuses a declared non-target cell with no template, so a lone C32 plan would fail on C33. Registration and promotion stay one cell at a time. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
197ea3a3b3 |
Free PR CI capacity by avoiding repeated setup and real-time test waits (#24355)
* Reduce repeated PR setup and transcript timing waits; add hosted comparisons * Align parallelism contract with Node-only external rebuild toolchain * Record hosted coverage and launch package, store, and cancellation comparisons * Apply hosted Windows setup savings and remove measured test waits * Keep measured PR package gains and remove completed comparison jobs * Report measured test counts with precise units |
||
|
|
56c7642aae |
test(orcad): skip the live-terminal runtime hand-over across a protocol bump (#24429)
* test(orcad): skip the Bun-to-Node live-terminal hand-over across a protocol bump The last Bun orcad's daemon reports protocol 38 forever, so asserting the adopted daemon matches this checkout's PROTOCOL_VERSION failed every bump. Ask the Bun slot's daemon for its protocol once, run the hand-over when it matches, and skip with the two versions named when it does not: a daemon at another protocol is never adopted across an update. * test(orcad): clean up the Bun protocol probe even when its launch fails The probe's cleanup ran only after a successful launch, so a launch that timed out or threw left its orcad and daemon running. One finally now stops the orcad, kills what it launched, and kills any daemon named by a pid file in the probe's data root. |
||
|
|
6d1a97ef98 |
fix(ssh): launch the Windows relay outside sshd's job so standard users work (#24224)
* fix(ssh): launch the Windows relay outside sshd's job without WMI Win32-OpenSSH kills a session's job on close but allows breakaway. relay.js gains a one-shot launcher mode that starts the detached relay with CREATE_BREAKAWAY_FROM_JOB through the staged process-tree addon, so a standard user no longer needs a WMI Remote Enable grant. WMI stays as the fallback for a relay without the addon, and a refusal there is named. The Windows SSH-host lanes drop their WMI grant and assert the breakaway route and adoption. * fix(ssh): find runtime holds without WMI on a standard-user Windows host The store GC read held runtimes through Get-CimInstance Win32_Process, which WMI refuses to a standard user's SSH logon, so the pass kept every runtime. On a refusal it now reads this account's own process image paths through Get-Process. * build(relay): ship the Windows relay launcher addon in every desktop package macOS and Linux packages carried Windows relays without windows-process-tree.node, so a legacy-runtime relay they uploaded to a Windows SSH host could not launch outside sshd's job and fell back to WMI, which a standard user is refused. A reusable Windows job now compiles the x64 and arm64 addons once and uploads them; release-cut, release-mac-build, and the hourly/daily/adhoc mac builds download them before build:release and require both arches. Staging now rejects a binary with the wrong PE machine, the ReadProcessMemory import, or no spawnOutsideJob export, so a stale pre-launcher build cannot ship. * ci(ssh): run the Windows SSH-host lanes when the relay process-tree build scripts change The staging and gyp-rebuild scripts decide which windows-process-tree addon the relay ships, so a change to either must re-prove the Windows host cells. * test(ci): find the mac orcad-template download by artifact name The release mac job now also downloads the relay Windows process-tree addons, so the first download-artifact step is no longer the template's. --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
8afa1db50c |
feat(ssh): rung B glibc 2.17 compat runtime; gate remote vault on host node:sqlite (#24148)
* feat(ssh): wire rung B to the glibc 2.17 compat runtime; gate rung C vault on full node:sqlite - COMPAT_RELAY_RUNTIMES lists linux-x64-glibc217; rung B plans the compat slot and compat pinned Node when glibc is below 2.28 or rung A refused with libc_floor/missing_lib. - The relay version folds the compat runtime's executable hash; refusals are cached per runtime. - The orcad template stages an optional linux-x64-glibc217 target (base package + compat node-pty slot + compat runtime marker); the verifier and materializer accept it. - node-pty slot loader falls back to the compat slot when the default slot is missing or needs a newer glibc. - Runtime store GC keeps the compat pin beside the default one on every relay connect. - hasNodeSqliteReaderApi (DatabaseSync + backup) gates relay session search and the relay OpenCode reader, which now names the host Node version in its unavailable reason; the SSH vault reader installs the compat Node on old-glibc hosts and uploads nothing when no pinned Node can run. - Rung D: a remembered noexec reports home_noexec and never advises installing Node. * fix(ssh): re-prove a replayed noexec after rung D so allowing exec recovers the host * fix(ssh): keep the rung B compat runtime pinned in the relay-connect store GC * test(ssh): mock deployment-target facts in the Windows OpenCode runtime tests * ci(ssh): build the glibc 2.17 compat slot for the hostile-host matrix; CentOS 7 lands on rung B --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> Co-authored-by: m4air <m4air@Mac.localdomain> |
||
|
|
744e7722c2 |
fix(relay): accept MIG version-name reconciliation and recreate stranded cells without rewriting the MIG (#24373)
* fix(relay): accept MIG version-name reconciliation and recreate stranded cells without rewriting the MIG The stranded-rollback recovery ran a gcloud rolling action, which renames the MIG version outside Terraform. Every later plan for that cell then reverted the label, and the capacity-plan validator refused the revert as an unreviewed MIG change, so the cell could be neither rolled nor rolled back. The validator now accepts a MIG field moving back to what relay-gce-cells.tf declares (version name and update policy), in every mode, and a test pins those values to the Terraform file. The stranded branch recreates the cell's single instance with recreate-instances, which leaves the MIG untouched. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): let a label-only MIG plan through and recreate on it in a stranded rollback A stranded rollback whose template is already in place plans only the version name revert. The validator still required the MIG template to move, so that plan was refused, and the recreate gate (changes == 0) would have skipped a plan of one change and left the drain flag set. Require the template move only when no declared field reconciles, and recreate whenever the template was not replaced (changes < 2). Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
f9940d5354 |
ci(ssh): macOS SSH-host lane for the pinned relay; fix uploads under a symlinked root (#24179)
* test(ssh): upload a root reached through a symlinked parent The upload-root realpath fix landed with #24180; this keeps macoshost's case where the root is passed explicitly beneath a symlinked parent. * ci(ssh): macOS hostile-host lane on a loopback user-level sshd Adds local-sshd cells for darwin-arm64 (macos-14) and darwin-x64 (macos-15-intel): a non-root sshd on 127.0.0.1 logs in as the runner user with SetEnv PATH=<shims>:/usr/bin:/bin:/usr/sbin:/sbin and an empty HOME, so no rc file restores Homebrew. The driver asserts rung A, terminal echo, cached runtime reuse, GC keeping the in-use runtime, no toolchain or xattr calls, and that the SFTP-uploaded Node carries no quarantine and runs as uploaded. Docker cells are unchanged; each machine runs only cells it can host. * test(ssh): fail a hostile-host run that would skip every named or hostable cell A cell named for the wrong OS or arch was silently skipped, so a macOS job on a mismatched runner went green having deployed nothing. Named cells must now be hostable here, and a gated run must select at least one cell. --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
0ad77ea2f7 |
ci(ssh): Windows SSH-host lanes (inbox + preview OpenSSH) for the pinned relay (#24180)
* ci(ssh): import the private Windows OpenSSH provisioning harness Copied unchanged from origin/OrcaWin/np-windows-ssh-provider-diagnostic (config/ci/windows-ssh-provider/preview-ssh/ at |
||
|
|
554f7f4ce5 |
feat(packaging): ship the orcad server template in desktop builds (#24155)
* build(orcad): merge per-runner prebuild slot trees into one matrix Each node-server lane builds only its own node-pty slot. Release CI needs their union before `build:orcad-prebuilds --require-slots` and the template build can run; merge-orcad-prebuilds.mjs verifies every lane's files against its own manifest, refuses duplicate slots and mismatched node-pty/N-API/Node-header builds, then writes one merged manifest. * build(orcad): keep agent-browser out of the desktop deployment template The template rides inside every desktop build (design D2). Seven ~10 MB agent-browser binaries would be ~76 MB, more than the rest of the template; design D2's package contents never listed it, and a slot without one already reports no headless browser. ORCAD_OMIT_AGENT_BROWSER=1 skips the copy; standalone build:orcad still includes it. * feat(packaging): ship the orcad deployment template in desktop builds Design D2: the server JS and every target's addons ship inside the app, as out/relay does; the ~120 MB Node runtimes stay excluded and are downloaded on demand. electron-builder copies out/orcad-template to Resources/orcad-template on every desktop OS, which is the first path materializeOrcadArtifact tries (process.resourcesPath). Platform signing rewrites native bytes the template manifest hashes: - macOS: the tree is signIgnored (codesign rejects its ELF/PE payloads); afterPack signs the darwin targets' Mach-O files with the app identity, as notarization requires, then reseals only those manifest entries. - Windows: SignPath signs after packaging, so release CI reseals from the inner-signing list (packaged-orcad-template.cjs --reseal-signed). Every other file must still match the build's hashes; afterPack verifies. ORCA_REQUIRE_ORCAD_TEMPLATE=1 makes a missing template fail beforePack and afterPack; without it a build ships none and SSH relays keep the legacy path. verify-packaged-orcad-template.test.mjs's "unused, excluded" contract is reversed on purpose. * ci(release): build the orcad template from qualified lanes and package it node-server-tests.yml becomes callable with a ref and build_template. With build_template, each lane that owns a release slot (macOS, Windows, the glibc 2.28 and Alpine lanes, and the glibc 2.17 compat lane) uploads its qualified out/orcad-prebuilds, the Windows lane also uploads both process-table addons, and desktop_template merges them, gates the full matrix plus the compat slot with --require-slots, runs build:orcad-template and uploads the orcad-template artifact. release-cut calls it at the release tag beside the other gates. The build and build-mac jobs wait for it, download it into out/orcad-template (the mac workflow from the parent run), and require it via ORCA_REQUIRE_ORCAD_TEMPLATE. The Windows signing staging skips the template's Linux/macOS payloads, and a reseal step records SignPath's bytes before the installer rebuild. A template-scoped concurrency group keeps a release call and main's push runs from cancelling each other. * test(orcad): keep the packaged-lookup imports clear of the compat-slot import edits * ci(orcad): let a rerun lane replace its template artifacts upload-artifact v4 refuses a second upload under an existing name in the same run, so rerunning a flaky node-server lane during a release would fail at the upload instead of re-qualifying the slot. * ci(node-server): build the template's Windows addons before the lane switches to Node 18 The addon build script imports TypeScript, which Node 18 cannot load, so every build_template run (release-cut included) failed on windows-2022. * fix(build): ship the orcad template's shared node_modules electron-builder's extraResources filter always drops the root node_modules of a source directory, so packaged apps lost orcad-template/node_modules and the afterPack verify failed. Copy it through its own resource entry. --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
c422936a71 |
fix(relay): anchor same-cap monitor evidence freshness to the run's authorisation, not job startup (#24349)
* fix(relay): anchor same-cap monitor evidence freshness to the run's authorisation, not job startup The same-cap gate now verifies the dry-run on its own clock and records the authorisation instant in the single-use consumed marker; each cell job checks the evidence age at that instant and bounds its own start after it. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): refuse a re-run same-cap gate before it consumes evidence; tighten order tests Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
14d4bb2e2a |
fix(ssh): Windows hosts without Add-Type staging; runtime-store GC on Windows (#24149)
* fix(ssh): collect the pinned-Node runtime store on Windows hosts
Windows SSH hosts now run runtime-store GC instead of skipping it: one
PowerShell inventory reads .runtime-ref-node-<sha> and .runtime-node refs from
every version dir, and one Get-CimInstance Win32_Process query filtered on an
image path under runtimes\ adds process holds (never by image name; a failed
query keeps everything). Stale upload stages are swept with the same rule as
POSIX. Promotion and the post-upload hold check now take the store lock on
Windows too, and the lock's own commands run unwrapped there.
Windows relay version-dir liveness now honours .relay-pid (design D5): a live
PID answers ALIVE before any pipe is touched, a dead one (ESRCH) plus refusing
pipes is exited, anything else is unverifiable. The runtime probe adopts a
pinned node.exe an earlier vault reader left without a .verified marker after
running it.
* fix(ssh): Windows stage fencing and vault runtime go through the verified node.exe
Upload-stage file identity on Windows no longer compiles an Add-Type P/Invoke
helper when the relay runs on Orca's verified pinned node.exe: the stage
commands run a fixed fs.lstatSync(..., {bigint:true}) script through it. It
prints the legacy helper's vol:high:low lowercase hex, and identity files are
compared after normalising hex spelling, so old and new clients recover each
other's stages. Host-Node relays keep the legacy helper; the choice is
documented in windows-edr-posture.md.
The Windows OpenCode vault reader now installs the pinned runtime through
ensureRemoteOrcadNodeRuntime (official zip, host-side extraction, .verified,
store lock) instead of uploading a client-extracted node.exe, and the relay dir
gains a .runtime-ref-node-<sha> so store GC keeps the runtime the vault uses.
* test(ssh): run the Windows stage-identity and store-GC tests on the Windows lane
The legacy/node.exe identity compatibility test and the Win32_Process hold path
were gated to win32 but no CI lane ran them. Add both files to the Windows
package lane and a real running-node.exe hold test.
* test(ssh): tear down Windows-lane temp trees through removeTreeSync
* test(ssh): grant the store lock to the Windows OpenCode runtime setup test
The Windows promote now runs under runtimes/.store-lock, so the mocked host
must answer the lock's CreateNew step.
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>
|
||
|
|
6aed05471c |
ci(ssh): hostile-host matrix for the relay runtime ladder (#24146)
* fix(ssh): classify a musl host missing libstdc++ as missing_lib, not wrong_libc musl's loader follows each missing-library line with one 'Error relocating ... symbol not found' per unresolved symbol, and the relocation pattern was checked first. Check missing libraries before relocation errors; the ld-linux/ld-musl interpreter case stays wrong_libc. * build(orcad): allow a partial deployment template for CI build-orcad-template --targets a,b builds and verifies only the named slots, so a CI job that can fill just the x64 Linux prebuild slots can still materialize rung A/C addons. Without the flag every target is still built and verified. * ci(ssh): hostile-host matrix for the relay runtime ladder Drives the real client-side relay deploy against Docker sshd targets and asserts the design D6 rung each lands on: Debian 10 and AlmaLinux 8 (glibc 2.28) and Alpine (musl) on rung A; Alpine without libstdc++ refused missing_lib down to D; Ubuntu 22.04 with a host Node 20 and a noexec home straight to D (home_noexec); CentOS 7 (glibc 2.17) refused libc_floor at A and C, falling to a host-npm path with no Node; and a no-egress Debian 10 still on rung A. Launched cells also prove the terminal echoes, no npm or compiler ran, a second connect reuses the uploaded runtime, and runtime GC keeps the in-use runtime while collecting an idle one. New workflow ssh-hostile-hosts.yml runs on dispatch and on path-filtered PRs. * test(ci): pin the hostile-host workflow to the headless-server builder images The matrix builds its runtime slots in copies of the node-server lanes' Alpine and manylinux images; this contract fails when NODE_RUNTIME_PIN or either builder digest moves in one workflow and not the other. * test(ssh): reconnect as the same client and retry a grace-held PTY owner in the hostile-host matrix --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
bd90da7a5b | ci: share PR planning setup and reuse the static native cache (#24329) | ||
|
|
9ed7b39b3c |
fix(relay): run the same-cap headroom gate in the modes the job actually receives (#24343)
The parent workflow collapses canary-apply and batch-apply into the job mode apply, so the headroom step's canary-apply/batch-apply condition never held and the gate was skipped on every real roll. Run it wherever the drain runs (apply, rollback before its restart) and in read-only verify; skip only a resumed rollback, which drains nothing. A new workflow-shape test fails on any job step comparing against a mode the parent cannot pass, and on a drain that can run without the headroom check. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
ddd4927a0b |
build(orcad): server node-pty slots at glibc 2.28, plus a glibc 2.17 compat slot (#24134)
* build(orcad): build server glibc slots on glibc 2.28 and add the glibc 2.17 compat slot
Design D6: the default linux-{x64,arm64}-glibc node-pty slots now build in
manylinux_2_28 (digest-pinned) and are gated at glibc 2.28 / GLIBCXX_3.4.25
through a floor profile on verify-linux-glibc-floor.cjs; the desktop keeps
its Ubuntu 20.04 (2.31) default.
Adds the opt-in linux-x64-glibc217 compat target: NODE_RUNTIME_COMPAT_ASSETS
pins the unofficial glibc-217 Node (update/check pin scripts cover it, outside
SERVER_TARGETS), and a new CI lane builds the compat slot in manylinux2014
with static libstdc++, gates it at glibc 2.17 with no shared C++ runtime in
DT_NEEDED, and smokes it under the glibc-217 Node.
* refactor(node-runtime-pin): route compat lookups through isCompatServerTarget; keep the glibc doc's slot-name paragraph intact
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
|
||
|
|
6593d7d194 |
feat(orcad): run orcad on the pinned Node instead of Bun (#24110)
* ci(daemon): gate PRs on daemon protocol crossing from the newest release Lands daemon-protocol-facts.mjs from the Windows update diagnostic branch with a stricter parser, and adds check-daemon-protocol-crossing.mjs (rule R1): the working tree must attach the newest release tag's daemon. Rollback crossing is reported only. Runs in the cross-version-wire job, which already has full tags; tag selection moves to config/scripts/stable-release-tags.mjs so both use one rule. * feat(persistence): run profile backups in the worker whenever its entry is bundled * refactor(orcad): make profile and native preflight runtime-neutral The profile preflight parser now takes the expected runtime identity from the caller (shipped callers pass the pinned Bun identity), and the native preflight is renamed to orcad-runtime-native-preflight with neutral wording. * feat(runtime): pin the Node 24.21.0 server runtime with an offline CI check Add src/shared/node-runtime-pin.ts (NODE_RUNTIME_PIN, SERVER_TARGETS, NODE_RUNTIME_ASSETS for all 8 server targets plus the headers tarball), generated by config/scripts/update-node-runtime-pin.mjs from the nodejs.org and unofficial-builds SHASUMS. check-node-runtime-pin.mjs verifies, with no network, that the pin tracks the locked Electron, matches engines.node's major, and covers exactly SERVER_TARGETS; it runs in the static analysis job. ORCAD_BUN_TARGETS consumers now read SERVER_TARGETS so there is one target list; orcad's Bun runtime and build output are unchanged. * test(persistence): skip plain-Node backup selection tests in the Bun profile suite * fix(runtime): reject a pinned archive that belongs to another target * ci(daemon): fail PRs that swap a runtime launcher and bump the daemon protocol D7.1 R3: hosting orcad or the daemon on another runtime is not a protocol change, so one PR must not do both. The launcher file list lives in the check script; the allow-runtime-launcher-protocol-bump label overrides it. * feat(orcad): select pinned-Node slots by a .runtime-node marker D7.1 R5: a Node slot names its shared runtimes/node-<sha256>/node through .runtime-node instead of .build-target, so Bun-era clients read it as a legacy slot rather than exiting 78 on a missing bun-runtime. Nothing builds the marker yet. * fix(runtime): load the Node pin without the typeless-module warning check-node-runtime-pin.mjs now requires the pin and takes nodeDistArchiveName from its own module, so it no longer loads the update script's build graph. * fix(orcad): resolve Node slots to the design's runtimes/node-<sha>/bin/node layout * feat(orcad): 8-slot node-pty prebuilds against the pinned Node headers at N-API 8 - build-orcad-prebuilds.mjs adds win32-x64/arm64 (conpty.node, the vendored conpty.dll/OpenConsole.exe, upstream's N-API conpty_console_list.node), compiles in a scratch copy against the hash-verified pinned headers (node.lib pinned per Windows arch) with NAPI_VERSION=8, rejects post-8 node_api_* imports, and writes a schema 2 manifest with per-file sha256, N-API level and the glibc need. - --require-slots [slots] verifies files against hashes; --smoke loads the slot under the pinned Node and spawns a PTY; --print-slot names the host slot. - The slot installer gates on N-API, libc, arch, glibc and file hashes instead of the exact NODE_MODULE_VERSION, and installs nested files (conpty/). - bun-profile-tests.yml builds, verifies and smokes each runner's slot. * fix(orcad): scope node-pty's glibc .symver pins to glibc on musl prebuild slots musl's unversioned libc cannot satisfy openpty@GLIBC_* references at link time, so the Alpine slot compile would fail. Pin the staged pty.cc guard to __GLIBC__ and assert both musl transforms against the installed patch. * feat(orcad): run orcad on the pinned Node instead of Bun A packaged orcad slot now references the pinned Node 24.21.0 by its executableSha256 (`.runtime-node`, `.server-target`) instead of carrying bun-runtime, and ships node-pty from the slot's prebuild, only its own ripgrep, and no Windows Bun PTY gate. The runtime lives beside the slots at runtimes/node-<sha>/bin/node (node.exe on Windows, upstream name). - build:orcad (build-orcad-node.mjs) builds the host slot's prebuild when missing and places the pinned runtime; the template is schema 3 with per-target files. - handoffToBundledOrcad() resolves the slot's runtime reference and checks process.versions.node against the pin; a host Node >= 18 still hands off. Startup preflight keys on running as that runtime; callers expect 'node'. - orcad and its daemon use node-pty (ConPTY + windows-pty-job on Windows); the Bun PTY sources, gate entry and canUseBunPty branches are removed. - SSH deploy uploads the official archive once per pin, extracts and hash-checks it on the host, and self-tests it before publishing. Bun slots stay launchable for rollback; Node slots never use host Node. - The runtime materializer is generic over pinned assets; the Bun wrapper remains only for the OpenCode vault reader (design Phase 2). - Cross-runtime test: a profile DB written by Bun 1.4.2 (WAL left by SIGKILL) opens and backs up under the pinned Node, and the reverse. No daemon PROTOCOL_VERSION change (design D7.1 R3). * docs(ci): name the headless lanes after the pinned Node, drop Bun shard timings Design D10: ci-demand-rollout.md and ci-runner-efficiency.md follow the bun-profile-tests.yml -> node-server-tests.yml rename; shard timings drop the deleted Bun PTY tests and follow the renamed ones. * chore(ci): count the runtime archive download as a runtime launcher path * fix(orcad): pin the macOS C++ standard for node-pty prebuilds The official Node headers' config.gypi sets clang: 0, so common.gypi skips its gnu++20 xcode_settings and Apple clang 15 (macos-14 runners) compiles node-addon-api as C++98. * fix(orcad): resolve the preflight's slot through realpath, as the handoff does A symlinked orcad.js handed off to its real slot's pinned Node, but the startup and profile preflights read the symlink's directory, found no runtime marker there, and silently skipped the readiness check. * refactor(ssh): drop materializeCachedNodeRuntime, which nothing calls Deploys upload the verified official archive (design D5); no client path needs an extracted Node executable cached by digest. * test(orcad): gate the Bun-to-Node upgrade and Node-to-Bun rollback with live terminals Design D7.1 R1/R3/R4 and D7.2. The last Bun orcad and this checkout's Node slot are installed side by side under ~/.orca-remote, launched and stopped with the client's own deploy commands, and share one data root. Each direction proves the incoming orcad adopts the outgoing runtime's daemon (same PID, same shell, output continues), opens its profile database and backs it up with its own shipped worker, and that GC keeps the slot the live daemon was forked from. The node-server Linux lanes provide Bun 1.4.2 and build that Bun orcad from main, and run with --cross-runtime. --artifact and --cross-runtime now make their tests fail on a missing input instead of skipping. * ci(node-server): pin node:24.21.0-alpine by its multi-arch index digest * test(ssh): name the runtime archive fixture after its role * test(node-server): load node-pty from the packaged slot in artifact runs The node-server lane installs dependencies without building node-pty, and Linux has no upstream prebuild, so the real-PTY failed-I/O teardown test (picked up by the pty-subprocess selector) could not load pty.node. In --artifact runs, alias node-pty to out/orcad's shipped slot so the test exercises the addon orcad actually runs under the pinned Node. * fix(orcad): let the Windows profile preflight exit after its PTY probe On Windows, node-pty keeps the conout worker thread and pseudoconsole alive until kill(), even after the shell exits. The PTY health probe never killed a cleanly exited probe, so the packaged preflight printed its readiness line and then hung until the build's 30s timeout, reported with an empty stderr. - The probe kills its PTY on Windows after exit and uses the bundled ConPTY the daemon spawns with. - The preflight exits once stdout is flushed; its owner reads to EOF. - Preflight failures now report code, signal, timeout, stdout and stderr. * test(node-server): load the slot's node-pty in the real-PTY test, not by alias A vite alias redirected only ESM imports of node-pty; windows-pty-job and local-pty-utils resolve it through require, so Windows loaded two conpty.node copies and the Git Bash job-membership proof read an empty job. The failed-I/O teardown test now loads node-pty through a fixture that picks the packaged slot in artifact lanes. The pty-subprocess selector was a prefix that also pulled in its POSIX-host sibling unit tests, which pr.yml runs and which were never qualified on Windows. Select the directory plus the two sibling files that belong here. --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
3e5c8d9f8e |
feat(relay): declare Asia cell c31 at the c30 shape (#24310)
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
d2dfc79764 |
ci(daemon): runtime-launcher protocol ratchet and Node slot marker (#24108)
* ci(daemon): gate PRs on daemon protocol crossing from the newest release Lands daemon-protocol-facts.mjs from the Windows update diagnostic branch with a stricter parser, and adds check-daemon-protocol-crossing.mjs (rule R1): the working tree must attach the newest release tag's daemon. Rollback crossing is reported only. Runs in the cross-version-wire job, which already has full tags; tag selection moves to config/scripts/stable-release-tags.mjs so both use one rule. * feat(persistence): run profile backups in the worker whenever its entry is bundled * refactor(orcad): make profile and native preflight runtime-neutral The profile preflight parser now takes the expected runtime identity from the caller (shipped callers pass the pinned Bun identity), and the native preflight is renamed to orcad-runtime-native-preflight with neutral wording. * feat(runtime): pin the Node 24.21.0 server runtime with an offline CI check Add src/shared/node-runtime-pin.ts (NODE_RUNTIME_PIN, SERVER_TARGETS, NODE_RUNTIME_ASSETS for all 8 server targets plus the headers tarball), generated by config/scripts/update-node-runtime-pin.mjs from the nodejs.org and unofficial-builds SHASUMS. check-node-runtime-pin.mjs verifies, with no network, that the pin tracks the locked Electron, matches engines.node's major, and covers exactly SERVER_TARGETS; it runs in the static analysis job. ORCAD_BUN_TARGETS consumers now read SERVER_TARGETS so there is one target list; orcad's Bun runtime and build output are unchanged. * test(persistence): skip plain-Node backup selection tests in the Bun profile suite * fix(runtime): reject a pinned archive that belongs to another target * ci(daemon): fail PRs that swap a runtime launcher and bump the daemon protocol D7.1 R3: hosting orcad or the daemon on another runtime is not a protocol change, so one PR must not do both. The launcher file list lives in the check script; the allow-runtime-launcher-protocol-bump label overrides it. * feat(orcad): select pinned-Node slots by a .runtime-node marker D7.1 R5: a Node slot names its shared runtimes/node-<sha256>/node through .runtime-node instead of .build-target, so Bun-era clients read it as a legacy slot rather than exiting 78 on a missing bun-runtime. Nothing builds the marker yet. * fix(runtime): load the Node pin without the typeless-module warning check-node-runtime-pin.mjs now requires the pin and takes nodeDistArchiveName from its own module, so it no longer loads the update script's build graph. * fix(orcad): resolve Node slots to the design's runtimes/node-<sha>/bin/node layout --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
2a83c9536f |
ci(daemon): gate PRs on daemon protocol crossing from the newest release (#24089)
Lands daemon-protocol-facts.mjs from the Windows update diagnostic branch with a stricter parser, and adds check-daemon-protocol-crossing.mjs (rule R1): the working tree must attach the newest release tag's daemon. Rollback crossing is reported only. Runs in the cross-version-wire job, which already has full tags; tag selection moves to config/scripts/stable-release-tags.mjs so both use one rule. Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
3135fbbf49 |
feat(runtime): pin the Node 24.21.0 server runtime with an offline CI check (#24087)
* feat(runtime): pin the Node 24.21.0 server runtime with an offline CI check Add src/shared/node-runtime-pin.ts (NODE_RUNTIME_PIN, SERVER_TARGETS, NODE_RUNTIME_ASSETS for all 8 server targets plus the headers tarball), generated by config/scripts/update-node-runtime-pin.mjs from the nodejs.org and unofficial-builds SHASUMS. check-node-runtime-pin.mjs verifies, with no network, that the pin tracks the locked Electron, matches engines.node's major, and covers exactly SERVER_TARGETS; it runs in the static analysis job. ORCAD_BUN_TARGETS consumers now read SERVER_TARGETS so there is one target list; orcad's Bun runtime and build output are unchanged. * fix(runtime): reject a pinned archive that belongs to another target --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
7c119465b0 |
fix(relay): define restart-safe by the cell runtime, and refuse waves without headroom (#24259)
* fix(relay): let a same-cap drain finish when only unplaceable hosts remain The c28 canary on 2026-10-01 drained the cell to zero live connections, but four hosts with no free slot anywhere kept redialling and held director leases on it, so the restart-safe wait timed out and left the cell isolated and empty. The drain wait now also passes once the runtime has carried nothing for a sustained quiet window while a small, capped number of leases remain, and logs the escape. Apply modes also refuse a cell whose hosts exceed 80% of the free slots on the other general cells, so a wave cannot strand hosts in the first place. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): define restart-safe by the cell runtime, not director leases Replaces the opt-in stranded-host escape with a corrected definition. A restart is safe when the cell runtime carries nothing live and no migration is open, sustained for the drain pace window. Director activity leases lag hosts that already left or cannot be placed, so they are reported in a progress line and the verified result instead of blocking the restart. The same-cap drain passes its existing pace window. The headroom script is added to the trusted evidence code paths with the other production scripts. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): print stranded director counts on every restart-safe sample Each restart-safe poll now prints its sample count and the director's restart-blocking leases, request units, reserved remainder, and migrations under `stranded`; the verified line carries the same object. Open migrations still block because each is pinned to the cell incarnation a restart replaces. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): require the pace window for every live restart-safe wait Pre-auth and total connections no longer reset the restart-safe window: on drained c28 they flickered with unauthenticated redials in a third of samples, which a restart does not lose. They stay in the progress output. Every live restart-safe call must now pass --pace-window-ms. The capacity job and staging proof drain unpaced, so they pass the production 300000 ms window, and the calls that relied on the 180000 ms default get 480000 ms. Headroom free slots now follow the director's placement rule: the admission pause minus the larger of observed and enforced units, minus outstanding control reservations. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
5cda0f4508 |
refactor(native-chat): keep agent-session records in the chat journal database (#24006)
* fix(native-chat): report a failed startup chat reconcile instead of failing app startup At startup the chat host re-checks every saved chat's lease and writes the result to agent-sessions.json. If that write failed (the file lock gave up, the file could not be written, or the file was written by a newer Orca and is read-only here), reconcileRestartLeases rejected, the startup IPC call rejected, and the renderer fell into its degraded "Session restore failed. Changes won't be saved until restart" mode. The reconcile is bookkeeping: a lease left unreconciled grants no writer, and every attach, send and read of a chat reconciles its own lease again. So the startup reconcile now reports its failure through a new optional host dependency, onStartupReconcileFailure, and resolves. The runtime routes it to its onError sink under the scope structured-agent-session-startup-reconcile, or logs it when no sink is installed (the desktop installs none). * fix(native-chat): read restored chats without waiting on lease bookkeeping With native chat on and a chat tab open at quit, the renderer's startup also awaits the chat tab restore (session.tabs.listAll). That restore re-ran the lease reconcile before reading each chat and rethrew its store failure, then recorded each restored tab as visible through a store transaction that throws on a held lock or a read-only store. Either one failed the restore, so startup still fell into "Session restore failed". Reading a chat grants no writer, so the reconcile startup and the restore run is now a reader's: createReaderReconcile never throws, answers whether every lease is settled (recovery is resolved only then; the journal opens either way), and reports each distinct failure once until a reconcile settles. Attach and agent start keep the strict reconcile. The restore's tab republish logs a failed visibility write and still publishes the tab, since a client drops every unpublished chat tab; user-driven publishes still refuse. The host dependency is renamed onLeaseReconcileFailure (scope structured-agent-session-lease-reconcile), since it now also reports for reads. * fix(native-chat): keep every record-store write off the startup chat read path Round-2 review found two more writes on the startup chat restore that could still fail it and put the app into "Session restore failed": republishing a /clear replacement recorded its tab visibility strictly, and resolving a chat's recovery rethrew its store error. The restore also paid one lock wait per tab and per batch of chats while the lock stayed held. The restore now derives tabs from state it already holds: - publishStructuredAgentSessionTab splits into the strict write and projectStructuredAgentSessionTab, which only updates the runtime's snapshot. The restore and /clear replacements only project: a saved tab index already lists every restored chat, and a /clear moves the tab in the same write that commits it. visibilityWriteMayFail is gone. - Chats a legacy profile restores that the index does not list are recorded in one best-effort transaction (store.showSessionTabs), so a failure leaves the index absent to seed again rather than partial. - The read restore's recovery resolution is caught and reported through onLeaseReconcileFailure, deduplicated with the reconcile's reports. - Once lease bookkeeping fails in a restore pass, the rest of that pass skips it, so a held lock costs one wait for the startup reconcile and one for the restore, however many chats are open. User actions (create, reveal, attach, send, the /clear commit) keep their strict writes. * test: open, seed and read the agent-session record store through one harness Tests that open the durable agent-session record store, seed it, or read back what it persisted now go through agent-session-record-store-test-harness.ts instead of calling AgentSessionRecordStore.open or touching agent-sessions.json themselves. A later change that moves the store into the chat database then changes the harness instead of every test. No production code changes. Tests whose subject is the JSON file itself (its .bak recovery, salvage, schema versions, permissions, and what older builds read back) keep reading and writing the file directly; the storage move rewrites or deletes them. * fix(native-chat): start each restore pass from one lease check and stop its bookkeeping at the first failure The restore now runs one reader lease check for the pass and lets each chat re-check and resolve recovery only while the pass is still settled. The first refusal or failed write clears it for the rest of the pass, and every chat is still opened for reading. With another process holding the lock, startup waits on it once in prepare and once in the restore, however many chats are open; a legacy profile waits once more for its tab-index seed. * docs(native-chat): correct restore comments and a test name to match the final design * test: address the record-store harness by the host's state directory The harness took the store's own folder, so each caller picked one (join(root, 'store'), or 'agent-sessions' where a test read the store the runtime owns). A later change that moves the store into the state directory's journal database could not tell those apart, and would have had to edit every caller again. Every harness function now takes the state directory, the one the test's journal database and recovery capsule already live in, and keeps the store in the same subfolder the runtime uses. Callers pass that directory; store-only tests pass their temp directory unchanged. Format tests that share a directory with harness calls take the file path from testAgentSessionStoreFilePath. The folder name moves from a private constant in the runtime to AGENT_SESSION_STORE_DIR_NAME beside the store's file name, so the harness shares it without importing the runtime. Its value and every path built from it are unchanged. * refactor(native-chat): keep agent-session records in the chat journal database The record store's records, operation ledger, retired claim keys and chat tab index become tables in agent-session-journal.db (user_version 4). The version-4 migration copies agent-sessions.json in its own transaction and never writes, renames or deletes that file or its .bak. Each store write is one journal transaction over exactly the rows it changed, checked with the load rules; the file lock, the external-change refresh and its hash, the .bak rotation, salvage and the hot-path recovery fence are gone from the store. * wip: importer tests * test(native-chat): cover the records migration, the import, row writes and read-only records * docs(native-chat): retire comments that describe the records file as the live store * test(native-chat): drop the record-store harness's leftover file path and type the import fixture * test(native-chat): let the host harness cleanup wait out a recovery-offer read's lock * fix(native-chat): let Stop reach the agent when its ledger row cannot be written Stop's operation-ledger row now shares the database with the chat history, so damage, a full disk or a stranded transaction on that write refused the Stop before the interrupt. A cancel plan now takes its decision from the committed ledger in memory, runs without settling, and warns that the row was skipped. Other mutations answer proven damage with the typed "Unable to load this chat." refusal instead of the raw SQLite error. * fix(native-chat): answer whether a profile holds chats from the database's rows Every host install creates agent-session-journal.db, chats or not, and the version probe created it too, so its mere existence made every profile that ever installed the host wait on host install and reconcile at startup. The check now opens the database read-only and looks for a record or tab row, lets the records file answer while its import is still owed, and counts an unreadable database as present. The version probe no longer creates the file. * fix(native-chat): open a chat from history when its tab index cannot be written Over records a newer Orca wrote, every write is refused, so opening a closed chat from Agent Session History failed on the tab-visibility write and the chat read as unreachable. Like closing a tab, opening one now reports a failed restore-index write and still publishes the tab. * fix(native-chat): keep the records import owed when the backup read fails transiently A torn records file whose .bak could not be read (EACCES, EIO) was reported as unusable, so the migration completed with nothing copied and never retried. A non-ENOENT read failure of either copy now carries its cause, which the importer classifies as a read that can clear. * test(native-chat): pin that an unreadable records file never falls back to its backup * fix(native-chat): restore imported chats' tabs when the records file had no tab index A chat created while the import was owed recorded a tab index holding only itself. When the file it later imported had no index, that index still read as recorded, so the imported chats' tabs never came back. The import now clears the recorded marker in that case, and restore falls back to the profile's tabs. * refactor(native-chat): drop the unused in-transaction store write Nothing called it, and it bypassed the write queue and the read-only refusal. * docs(native-chat): say that an unusable records file is left untouched but never re-imported * refactor(native-chat): keep the provider handle chain check as main has it The chain-validation refactor has no measured need in this change. * docs(native-chat): retire lease-renewer comments that describe the records file as the live store * fix(native-chat): keep a throwing failure sink from failing the startup chat read The lease bookkeeping failure reporter called the host's failure sink directly, so a sink that threw turned a reported, recoverable store failure back into a rejected startup reconcile or read restore. The reporter now catches a sink throw and logs both the original failure and the sink error with console.warn. * test(native-chat): wait for a replaced host's restart-offer writes before cleanup A restart test replaces the host without tearing the old one down, so the old host's fire-and-forget restart-offer withdrawal could still hold the recovery capsule's lock directory when cleanup removed the test directory (ENOTEMPTY). The harness now hands hosts a capsule that tracks running operations and waits for them before removing the directory, replacing the rm retries. * docs(native-chat): retire the abandon helper's note that the store re-creates its directory * fix(native-chat): restore a chat opened while the import was owed beside the profile's chats When the imported records file had no tab index, restore fell back to the profile's saved tabs, which never list a Claude chat, and the seed then rewrote the tab table without the chat opened while the import was owed. The tab rows that chat left are now loaded as unrecorded, restore takes them together with the profile's chats, and the seed keeps their tab ids. * test(native-chat): pin that a create whose tab index write fails still opens the chat * docs(native-chat): say why restore puts chats opened while the import was owed first * test(native-chat): replace a ledger row rather than change it in place in the Send-now rerun test The record store freezes published rows in tests, so setting a row's outcome in place threw; the test now swaps in a changed copy, as its sibling cases do. |
||
|
|
cfa43e7eab |
fix(codex): opening a terminal no longer strips Codex hooks from the real ~/.codex (#23552)
* fix(codex): a real-home restore leaves a file alone once someone else changed it Orca writes ~/.codex/hooks.json (and a trust rebase writes config.toml), then runs a Codex trust session for up to 10 s, then restores the original bytes if the session fails. The restore wrote unconditionally, so a save that landed during the session, from the user or another Orca, was silently reverted. Each restore now compares first: it writes the original back only while the file still holds the generation Orca's mutation left, and otherwise logs and leaves it alone. This covers the real-home install and opt-out sweep (restoreRealHomeHooksJson), the legacy sweep's hooks restore, and config.toml rollback (restoreCodexTrustConfig). For hooks.json the generation is the exact bytes Orca wrote. For a config.toml that a trust rebase changed it is the file as the rebase left it. When Codex itself wrote config.toml inside the session that just failed, Orca never knew those bytes, so that rollback compares against the file as the session settled. The next commit keeps other Orca instances out of that window; a user edit made during such a session can still be rolled back. * fix(codex): serialize real-home Codex writes across Orca instances Every Orca on one HOME (a dev and a packaged app, or an offline CLI) writes the same ~/.codex/hooks.json, config.toml and ~/.orca/agent-hooks/codex-hook.sh. The per-file lane that orders capture, mutate and restore was in-process only, so another instance could write inside this one's restore window, or undo it. The lane for the user's real config.toml now also holds the existing crash-safe managed-hook install lock (~/.orca/managed-hook-install.lock, the one relay installers take for the same home). It is taken only by the outermost acquire, because the lock file is not reentrant and grants and trust rebases nest inside an install. Managed-home installs, the real-home install and opt-out sweep, and the legacy sweep all enter through it. Compare-and-swap on restore stays as the backstop. A lock that cannot be taken within its 10 s wait fails that install, which is already best effort: launch prep logs it, and the real-home lane falls back to the managed lane until its retry. * fix(codex): opening a terminal no longer strips the shared Codex entry from ~/.codex Every Orca instance on one HOME writes the same status-hook entry into the user's ~/.codex/hooks.json, with its trust in config.toml. Launch prep runs on every pane spawn, and under a managed Codex account it ran the legacy system sweep. That sweep matched Orca entries by script file name, so it removed the current shared entry and the trust blocks the grant ledger recorded. On a live laptop hooks.json went 4139 -> 18 bytes about 150 ms before a new pane opened. With hooks off, the real-home lane's launch prep swept the same way. Now nothing automatic removes the current entry or its trust: - The legacy sweep removes only an enumerated list of retired command forms that no build writes any more (#1019's double-quoted form, #1536's exec-guarded form, and Windows' per-userData bare path), plus their trust. - ensureRealHomeCodexHookState with hooks off writes nothing; that covers launch prep, session resume and startup. - Only the user's explicit opt-out (codexHookService.remove()) strips the entry and its ledger-recorded trust from the real home. - The sweep-suppression gate existed only to stop the sweep from deleting the current entry, so it is deleted with its main-process wiring. Startup with hooks off already skipped the real-home install; with this change the first pane's launch prep with hooks off also leaves ~/.codex untouched. * fix(codex): a pane's prepare-codex only repairs a home its own HOME's app installed On macOS a pane starts through login(1), so it gets the user's real HOME even when its Orca app runs with another one. The pane's `codex()` preflight installed hooks in the CLI process with that real HOME: it rewrote ~/.orca/agent-hooks/codex-hook.sh, promoted trust into the real config.toml, and wrote the real HOME's script path into the app's managed home. The preflight now acts only when the managed home's hooks already run this process's own shared script, which proves the app that installed them shares its HOME. Otherwise it writes nothing; the app installed the home at spawn. Why not a no-op: the preflight was added (#14326) because trust can go stale between opening a pane and typing `codex`, for example in a pane that survives an app update, and Codex then stops in hook review. For a same-HOME pane it still repairs that. Why keep promotion: the install drops runtime trust the system config does not back, so skipping promotion would delete approvals the user gave inside Orca-launched Codex. * test(agent-hooks): await every installer in the refresher coverage test The test fired each managed installer without awaiting it and read ~/.orca/agent-hooks straight after. Codex's install now takes the cross-process real-home lock before it writes its script, so the script landed after the read. Await the installers, and stub Codex's trust sessions so the awaited install cannot start a real `codex app-server`. * fix(codex): retire the two real-home command forms the list missed The real-home lane wrote two Codex hook forms into ~/.codex that no build writes any more and that the enumerated retired list did not name: - POSIX, #9501 until #10885: the file-guarded form draining with a bare `cat`. - Windows, #9501 until #10221 took Windows off the real-home lane: the encoded PowerShell launcher for a non-cmd-safe script path. The file-name sweep removed both before; the enumerated sweep left them in place, trusted, still passing the script's exit status to Codex. Both now match as frozen literals. Also corrects the startup ordering comment: the real-home install runs first so its in-slot upgrade lands before the managed install's sweep retires the prior command; nothing re-arms a legacy sweep any more. * fix(codex): take the real-home lock only when a write is needed The previous commit made every entry to the real-home config lane take the cross-process lock. That lane runs on every pane spawn and every typed `codex` preflight, so the steady state paid an owner probe (a `ps` spawn on macOS) and could wait up to 10 s behind another instance's trust session, even though it wrote nothing. Each real-home writer now compares the desired state with the files on disk first, without the lock. Only when a write is needed does it take the lock, re-read and recheck, then write: - real-home install: the planned hooks.json, the shared script and the ledger-recorded grant are compared; the locked path re-plans from disk. - legacy sweep: locks only when a retired entry is present; the sweep re-reads. - approval promotion: locks only when there is something to promote; the promotions are recomputed under the lock. - the shared ~/.orca/agent-hooks script: locks only when its bytes differ. The explicit opt-out always takes the lock. The lock is reentrant through async context, since grants and rebases nest inside an install, so the config-lane option the previous commit added is removed. * fix(codex): a shared script without its exec bit is not the steady state The compare-first check matched the shared ~/.orca/agent-hooks script on bytes alone. writeManagedScript also restores 0755 on every call, and the POSIX hook guard skips a script that is not executable, so a script whose mode was lost (a dotfiles restore, a plain copy) now stayed that way: every Codex hook drained stdin and reported nothing until an app restart refreshed the script. The check now also requires the mode the writer sets, so that case takes the lock and the write path repairs it. * test(codex): the retired encoded launcher never matches today's shared one The shared encoded Windows launcher is still current for other agents, so the comment claiming today's launcher is never encoded was wrong. What keeps the retired matcher off it is the exact payload: since #14825 the shared launcher prefixes its payload and drops -ExecutionPolicy Bypass. Pin that with a case. * fix(codex): the pane step recognises its own script under a home path with an apostrophe The same-HOME check looked for the script path wrapped in bare single quotes, but both hook writers escape an apostrophe inside the quotes. A home such as C:\Users\O'Brien never matched, so the pane-step repair never ran there. * fix(codex): the trust-RPC escape hatch still keeps the real home off its lane The no-write check reported a recorded grant as current, so with ORCA_DISABLE_CODEX_TRUST_RPC set the real-home lane stayed in use. The grant itself refuses before reading its ledger; the check now does the same. * fix(codex): the shared script write no longer waits on the real-home lock The write is atomic and skips identical bytes; waiting behind another instance's trust session could only fail a pane's managed-home install. * fix(codex): an in-Orca approval survives a launch that cannot get the real-home lock The install drops runtime trust the system config does not back, so a promotion skipped for want of the lock lost the approval for good. It now writes unlocked, as it did before the lock existed. * refactor(codex): take the cross-process real-home lock back out The lock fixed no observed failure. The three that were observed each have their own fix in this series: the legacy sweep matches only frozen retired command forms, hooks-off launch prep writes nothing, and a pane's prepare-codex repairs only a home its own HOME's app installed. The lock instead brought its own defects: a steady-state spawn waiting behind another instance's trust session, a compare-first split to avoid that, a script write and an approval promotion that could fail for want of the lock. Removed, with their tests: the real-home write lock and its async-context reentrancy, the plan/compare split that kept it off steady-state spawns, the compare-first legacy sweep, the locked approval promotion and its unlocked fallback, the compare-first shared script write (writeManagedScript already skips identical bytes and restores the exec bit), and the CLI tsconfig entries the lock pulled in. Kept: the retired-forms matcher, the hooks-off no-op, removal only on an explicit opt-out, the pane own-script check, and the compare-and-swap rollbacks. Every instance now writes identical bytes idempotently. * fix(codex): an opt-out that cannot read hooks.json keeps Orca's trust and ledger The opt-out swept the real-home entry, then dropped Orca's ledger-proven trust whenever a ledger existed, even when the sweep could not read hooks.json. The entry could still be there, now untrusted, and the ledger that proves ownership was gone for the retry. Drop that trust only after a sweep that read the file. * refactor(agent-hooks): one predicate for whether an agent's status hooks are on "Global switch on and this agent not turned off" was spelled out separately in the startup controls, the settings reconcile, the retained-home reconcile, the WSL preflight RPC, the CLI preflight and the OpenCode plugin selection. They now share one function, in a module light enough for the CLI's per-launch Codex preflight to load. The PTY spawn env derives the Codex flag from the switch and opt-out list it already carries, the same way it does for OpenCode and Pi, instead of receiving a second copy. * fix(codex): launch and resume prep honour Codex's per-agent hook opt-out Turning Codex off in the per-agent hook settings removes Orca's Codex hook entry, but launch prep and session resume read only the global hooks switch, so the next Codex launch or resume wrote the entry straight back into the real ~/.codex or the account's home. Both now read the per-agent predicate, which the PTY spawn env and startup already honoured. * fix(codex): turning Codex off per agent clears the real ~/.codex entry While the real-home lane owns ~/.codex/hooks.json, the legacy system-home sweep stands down. That gate read only the global switch, so turning Codex off per agent ran remove() with the sweep still suppressed and left Orca's entry in the real ~/.codex. The gate now reads the per-agent predicate, the same as turning every hook off. * test(codex): cover the system ~/.codex sweep gate for Codex turned off The gate that lets the legacy system-home sweep run was an inline closure in startup, so reverting it to the global switch left CI green. It is now a pure function beside the gate it feeds, with a table test and a remove() test on a seeded ~/.codex: turning Codex off strips Orca's entry and keeps user hooks; with Codex on the entry stays. * fix(cli): keep the agent-status hooks predicate loadable by the packaged CLI The CLI's prepare-codex handler imported the predicate from src/main, but the Electron build rebuilds out/main from its declared entries only, so the packaged `orca agent hooks` commands could not load it (package jobs and the CLI bundle-parity test were red). The predicate reads only settings, so it now lives in src/shared, which the CLI compiles itself. * feat(codex): every Orca build writes one frozen Codex hook command The Codex hook command was built from this build's wrapper, so two builds on one HOME disagreed about the bytes of the shared ~/.codex entry and kept rewriting it, with a Codex trust session each time. The command is now fixed per form and carries its form number: - POSIX: one command with no path in it. It runs the shared script only in an Orca pane with hooks on (pane key and hook port set), drains stdin everywhere else, and always exits 0. A branch for a per-build script root is written now and stays dormant until Orca sets ORCA_AGENT_HOOK_ROOT, so that change will not move these bytes. - Windows: the bare forward-slash path to the shared .cmd, which runs under PowerShell 7 and 5.1, Codex's hook hosts. A profile path that is not one PowerShell token gets a plain PowerShell form with the same branches. The literals live in the form module, so a change to the shared hook constants cannot move them; goldens pin the bytes. Every form keeps `agent-hooks/codex-hook.*` in plain text, so older builds still recognize it. * fix(codex): one main-process owner adds the real-home entry; nothing restores files Each Orca writer of ~/.codex decided what Orca's entry must be from its own build and instance, then removed or reverted whatever differed: launch prep rewrote any Orca-shaped entry to this build's command and stripped Orca entries from events this build does not use, and a failed trust session restored hooks.json and config.toml from snapshots. With several instances and builds on one HOME, every disagreement became a deletion or a revert. The main process is now the one writer, and its writes are add-only: - A launch or resume adds Orca's frozen entry to an event that has none and leaves every Orca entry it finds, so a running older build is never fought. - App start also converts an older Orca form to the frozen command, once, in its own slot: one hooks.json write (one .bak) and one trust grant per home. - A newer form is never rewritten or appended beside, and Orca entries in events this build does not use are kept. - After a failed trust grant, only an entry this call wrote that is still untrusted is withdrawn, putting back the handler it replaced. Both files are re-read, so a concurrent edit, or the identical entry another Orca trusted meanwhile, survives. Deleted: the compare-and-swap hooks.json restore, the config.toml snapshot restore after a grant session and after a user-trust re-key, and the rollback module. A grant session writes trust only at Orca's own keys, and every caller settles those keys itself. A failed re-key of moved user hooks now keeps the write and reports it; Codex lists those hooks for review. * fix(codex): the pane CLI asks the app to prepare its Codex home `orca agent hooks prepare-codex` ran Codex's install inside the pane. That process can have the real HOME (login(1)) and runs outside the app's in-process queues, so it was a second writer of ~/.codex and ~/.orca beside the app. A check that the home ran "its own script" guarded it. The pane step now only asks the app, over the same kind of local RPC the WSL pane step already uses (agentHooks.prepareCodexForPane). The app checks that the pane's CODEX_HOME is one its own userData owns, reads its own hooks setting, and installs on its own queue. An app that is not running, or is too old to know the method, makes the step a no-op, as it is on WSL. The own-script check and the CLI's settings read are gone, and the preflight module leaves the CLI bundle. * fix(codex): delete the pane step on native hosts The previous commit had `orca agent hooks prepare-codex` ask the app to prepare the pane's Codex home. The case it existed for (#14326, a pane that survives an app update with stale hook trust) did not reproduce, and no other desktop agent host writes agent config from a terminal or launch wrapper. - Deleted: the agentHooks.prepareCodexForPane RPC method, its params and catalog entry, and prepareManagedCodexHomeBeforeShellLaunch with its module, tests and CLI build entry. - `agent hooks prepare-codex` is a no-op on native hosts. It stays for one release so shell wrappers from older builds, which still call it, exit 0. - WSL panes are unchanged: they still ask the app over agentHooks.prepareCodexForWslPane. The shell wrappers and ORCA_CODEX_LAUNCH_PREFLIGHT stay, because WSL panes use the same wrappers and variable (forwarded through WSLENV). A native pane still starts the CLI once per `codex` it runs; skipping that is a follow-up. * test(codex): a failed trust session keeps concurrent edits to both files QA case 9 at host level, on a real file system in a temp HOME: Codex's trust session fails after another writer saved hooks.json and config.toml. - Both saves survive, and no Orca entry is left that Codex would list for review: this call's entry is withdrawn. - A failed one-time conversion puts the older Orca entry back in its slot and keeps both saves. Both tests fail on the previous head, which restored config.toml from a snapshot and left the untrusted entries in hooks.json. Removing the withdrawal turns both red. * feat(codex): read whether an Orca entry's stored trust is still current A Codex release that changes how it hashes a hook leaves Orca's stored trust stale: the entry is present, but Codex lists it as modified. Checking only whether the entry is missing cannot see that. readOrcaEntryTrust sorts a present entry into four states: - trusted: the stored hash is the current one; - untrusted: there is no stored hash; - stale: the stored hash is not the current one; - disabled: the user turned the entry off. The caller can pass Codex's current hash, for example one a grant recorded. The failed-grant withdrawal now uses it, and also keeps an entry the user turned off. Nothing re-grants on 'stale' yet. * fix(codex): a slow Codex start retries on the next launch, never for minutes On a loaded Mac a cold `codex app-server` took over 10 s (QA case 4). The grant timed out, the entry was withdrawn, and a 5-minute cooldown in both the grant and the real-home install then refused every retry. - The native session deadline is 30 s, the same as WSL's. - A timeout starts no cooldown in the grant or in the real-home install. The next launch retries. Other failures keep their cooldown. - Launches that queue behind a slow session share one follow-up run, so a launch waits for at most two sessions, not one per earlier launch. Tests: a 15 s cold start still grants and keeps the entry; after a timeout, the next launch runs a session at once; four queued launches run two sessions. Each is red on the previous head, and each mechanism was removed in turn to confirm its test turns red. * fix(codex): Orca's automatic writes never move a user hook Codex keys a hook's trust by its position in hooks.json. App start's collapse of Orca duplicates removed every Orca entry and appended one at the end. That moved any user hook that followed a removed entry, so the write waited on a session to re-key the moved hook's trust. App start now: - converts the first Orca entry that sits in a plain slot to the frozen command, in place; - drops any other Orca entry only when that moves no user hook; - keeps a duplicate that a user hook follows, and trusts every frozen copy, so none is listed for review; - appends only when no frozen entry is left. Tests check user positions and user trust blocks byte-for-byte for each automatic write: add-missing (append), the one-time conversion (in place), a trailing duplicate, a duplicate before a user hook, and older duplicates normalized to one entry. The three collapse cases fail on the previous head. Removing the position check, or the in-place conversion, turns its tests red. Only the explicit opt-out still removes an entry that user hooks follow. * fix(codex): removing an Orca entry never waits on a Codex session Removing an Orca entry from ~/.codex/hooks.json moves every user hook behind it up a slot, and Codex keys trust by slot. The retired-form sweep, the opt-out and a failed-grant withdrawal all asked a `codex app-server` session to list the old trust before writing, and to re-key it afterwards. A timeout there threw before the write and latched a 5-minute cooldown, so a slow cold start blocked the retired-form sweep at boot (QA case 4). Each moved hook's [hooks.state] block now moves to its new key, body bytes unchanged, straight after the hooks.json write. Codex hashes a hook's content, not its position or its file path, so the moved block stays exactly as valid as it was: a trusted hook stays trusted, an untrusted one stays untrusted, and one the user turned off stays off. No removal waits on or depends on a session. A failed config.toml write keeps the hooks write and logs. Deleted: the inspect and repair sessions, their client, and their cooldown. The generation guards on the hooks.json writes stay, for other processes. Tests: the retired sweep removes the retired entry and carries the trust of the user hook behind it while every Codex session times out (red on the previous head); the opt-out carries an appended user hook's trust; the move carries trusted, disabled and untrusted states byte for byte. Removing the move turns all of them red. * fix(codex): a Codex launch never waits on Codex's approval of Orca's entry A launch on the real-home lane awaited Codex's trust grant for the entry it had just added. A cold `codex app-server` on a loaded Mac took over 10 s, so the launch could wait that long, and a failure then latched a 5-minute cooldown. - Codex's approval runs in the background, with a 30 s cold-start budget. - A launch uses the real home only when the ledger shows trust is already current. Otherwise it goes to the managed home at once, and the next launch picks up the finished grant. - A launch that arrives while a grant runs does no work and does not queue behind it. - A resume into the real home has no managed home to fall back to. It waits for the grant, but no longer than the 10 s a launch always could. - A background grant that times out starts no cooldown; the next launch retries. Any other failure backs off for 10 s instead of 5 minutes. Success is what the ledger remembers. - A failed grant still withdraws only what that install added and is still unapproved. The log now says how many entries it took back and when the next try comes. Managed-home grants keep their 10 s deadline and stay on launch prep, as before; they fall back to Orca-computed trust. Tests: - A 15 s start: the launch returns in under a second on the managed home, a second launch starts no session, the grant lands in the background, and the next launch uses the real home. - A timeout sets no cooldown, withdraws its adds and logs it. - Another failure retries after 10 s, not before. - A resume waits only as long as allowed. - Case 9 checks the log line and the retry. Making the launch await the grant, a 10 s budget, either timeout cooldown, and a 5-minute backoff were each tried, and each turns its test red. * fix(codex): move a hook's trust only when every stored key has the known shape Orca now edits Codex's trust store directly when a removal moves a user hook. Three safeguards keep that honest: - Fail safe. If any [hooks.state] key in config.toml does not have the shape `<path>:<event>:<group>:<handler>`, nothing moves and Codex asks the user to review. That shape was checked unchanged from Codex 0.141 to 0.158. - Targeted. The file is read immediately before the atomic rename, and only the moved keys' blocks change. Every other byte stays, and no snapshot is restored. - Verbatim. Each block's body moves as Codex wrote it, including fields Orca does not know. No hash is ever computed, and a hook with no block gets none. Tests: - An unknown key shape stops every move. - Everything except the moved block survives byte for byte, and the moved body keeps an unknown field. - In case 9, a hook the user approved during the failed session keeps its approval when the withdrawal moves it, beside the concurrent project edit. Removing the shape check, or writing a computed block instead of the stored body, turns these tests red. * refactor(codex): keep only the trust read the failed-grant withdrawal uses A capture across Codex 0.141, 0.150 and 0.158, switching in all six directions, showed Orca's entry keeps the same hash and stays trusted. A Codex upgrade does not make its trust stale, so nothing needs to re-grant on staleness. readOrcaEntryTrust keeps the four states the withdrawal needs, but loses the parameter that let a caller pass a different current hash, and the test for a Codex that hashes differently. * fix(codex): native panes no longer start the Orca CLI before each codex The pane step is a no-op on native hosts, but native panes still carried ORCA_CODEX_LAUNCH_PREFLIGHT, so every `codex` typed in a pane started the Orca CLI for nothing. Only a packaged Windows build's WSL pane now gets the variable; the app prepares every native Codex home itself. The resolver loses the dev-launcher path and its userDataPath option, which only native panes used. Tests: a native macOS, Linux and Windows pane gets no preflight, packaged or not, even with the bundled CLI present; a WSL pane still gets the verified absolute launcher. Letting native panes through again turns them red. * chore(cli): say when the native prepare-codex no-op can go Native pane wrappers from builds up to v1.4.216 still call it. It can be deleted once no supported build's wrapper does. * test(codex): check the WSL launcher path instead of asserting it * fix(codex): a launch no longer waits behind the background real-home approval The background grant ran its whole codex app-server session inside the shared ~/.codex/config.toml lane, and on a cold host its session was also the shared capability probe. A launch sent to the managed home then waited on both: the managed install and the project-trust write queue on that lane, and the managed install's own grant waited for the probe. On a cold app-server that was up to 30 s per launch. The lane was held across the session only to protect the retired capture-and-restore. Codex writes its own records, so the lane is now taken only around Orca's own pre-grant write. The background grant runs its session without publishing it as the shared probe, and the whole grant is bounded by its deadline, so a hang outside the session cannot leave the lane 'granting'. * fix(codex): a failed re-grant no longer strips Codex's own approval of Orca's entries Before each trust session, the grant deleted every Orca record whose hash matched the one Orca computes. That exists because a managed home's fallback writes Orca-computed trust under both Windows path-separator spellings, and Codex rewrites only its own spelling, so the other copy would linger. On failure the managed and WSL fallbacks write that trust back, and before this fold a snapshot restore covered it. The real ~/.codex has neither: Orca never writes computed trust there (the real-home lane does not run on Windows at all), so a matching record there is Codex's own approval. After a ledger miss (another Orca profile, a Codex update, a lost ledger) and a failed session, nothing put it back, and every Orca entry showed "Hooks need review". The clear now runs only for homes whose fallback writes that trust. * fix(codex): a real-home resume spawns only once Orca's entry is approved or withdrawn A resume that must run in ~/.codex waited at most 10 s for the background approval, then spawned anyway. On a cold app-server that left Codex beside an unapproved Orca entry, so the resumed pane showed hook review. The resume now waits for the grant to settle. Settled means Codex approved the entry, or the grant failed and withdrew its own unapproved write; the grant's deadline bounds the wait (30 s, the cold-start budget), and a failed approval never fails the resume. Why this over the alternatives: - Spawning at 10 s keeps the review prompt this fold exists to remove. - Withdrawing at 10 s from the resume races the still-running session: Codex can write the frozen entry's hash after the withdrawal, and for a converted entry that marks the older command Orca put back as modified. - A resume cannot use the managed home: the session lives in ~/.codex. So the only states that cannot race Codex are the grant's own settle. The cost is a longer worst case on a cold app-server (up to the 30 s deadline, plus any managed-home install that holds the config.toml lane); a warm approval takes seconds, and an approved entry costs no wait. * fix(codex): keep the 5-minute trust cooldown for launch-path grants The fold shortened the host's trust-grant cooldown from 5 minutes to 10 seconds for every grant. That was meant for the background ~/.codex approval, which blocks no launch. The managed-home and WSL grants run inline on the launch path, so with a hung app-server every launch more than 10 s after the last failure paid the full inline timeout again (10 s native, 30 s WSL). Cooldowns are now kept per lane: inline grants keep 5 minutes, the background grant retries after 10 s, and neither lane's failure cools the other down. A success, or a proven-missing surface, still clears both. The real-home install's own retries (an unreadable hooks.json, unknown keys) are back on the 5-minute interval they had before the fold. The cooldown moves to its own module so the grant stays within the file limit. * fix(codex): a failed grant withdraws the exact copy it wrote The withdrawal re-found "this call's" entry by command, taking the first frozen handler in the event. When app start converted a later slot while an earlier frozen copy sat in a matcher group (which conversion skips), a failed grant acted on that earlier copy: it put the older command into it, or skipped it, and left the converted, unapproved copy in place. Each write now records where its handler landed, after any duplicate drops, and the withdrawal acts only on that slot. A copy that has since moved is left alone; the next launch's grant retries it. * fix(codex): the failed-grant withdrawal checks hooks.json is unchanged before writing The install and the retired-form sweep both refuse to replace ~/.codex/hooks.json if it changed since they read it. The withdrawal did not: a save landing between its read and its atomic replace was lost. The window is small, since the withdrawal is synchronous, but it now carries the same guard. * refactor(codex): drop rationale left over from the snapshot restore; name the trust-move module for what it does Comments on the config.toml lanes still justified them by a grant's capture-and-restore window, which the fold deleted, and the trust-write deadline still counted a grant session holding the lane. They now give the reason that remains: Orca's own multi-step reads and writes, and managed-home installs that hold the lane across their inline grant. codex-user-hook-trust-rebase no longer rebases through Codex; it moves stored trust records, so it is now codex-user-hook-trust-moves. The grant test that pinned two sessions on one config.toml to run one at a time is removed: its reason was an interleaved capture and restore. Callers that write config.toml around a grant hold their own lane, which the nested installer test still covers. * build(cli): list the trust-grant cooldown module in the CLI program The CLI's agent-hooks handler loads the hook controls, which reach the Codex trust grant; the CLI project is composite, so every module in that graph must be listed. * docs(codex): say which Windows hosts each hook command form runs under Codex runs a hook under the turn's shell (PowerShell 7 or 5.1 in every captured session) and, with no single local turn shell, under %COMSPEC% /C. The bare forward-slash path ran under all three in the Windows host census. The PowerShell form used for a profile path with a space does not parse under cmd.exe; no form valid in all three hosts has been run for such a path, so the form stays and the gap is stated here and in the PR. * test(codex): type the withdrawal seam without an assertion * fix(codex): a real-home resume starts at once, trusting Orca's entries for that process A resume that must run in ~/.codex waited for Codex's background approval of Orca's newly written hook entry: up to 30-40 s on a cold app-server. That made the user's resume wait on bookkeeping, and the alternatives (start at 10 s with Codex's hook review showing, or withdraw the entry and race Codex's own write) were worse. Codex reads hook trust from its session-flag config layer as well as the user's config.toml, merged per key, and has since hook trust shipped. So the resume no longer waits. When Orca's own frozen entries in ~/.codex are untrusted (or hold a stale hash), the resume command carries `-c hooks.state={'<key>'={trusted_hash='<hash>'},...}` for exactly those entries: the key under both the logical and the real path of ~/.codex (Codex keys an explicit CODEX_HOME by its real path), and the hash of that entry's content, so it can trust nothing else at that slot. The user's hooks are never included, nothing is written, and the background approval still runs for later plain `codex` launches. An approved entry adds nothing; a Codex known to lack hook trust gets nothing. One inline table, because Codex splits a `-c` key on every `.` and the key holds `.codex/hooks.json`. TOML literal strings keep `"` out of Windows native-argument quoting. The flag goes before `resume <id>`, quoted for the pane's shell (portable Unix, PowerShell or cmd), in the launch command and in the setup-sequenced copy of it; a cmd line whose path cmd would expand, or a key with an apostrophe, is left unchanged. SSH and WSL resumes get no preparation, so no local path reaches them. * Revert "fix(codex): a real-home resume starts at once, trusting Orca's entries for that process" This reverts commit |
||
|
|
cfe4c633eb |
fix(mobile): publish the Android APK's size and checksum with the release (#24037)
An APK that fails to install with a missing certificate or a package-parse error is usually a download that died near the end: the signature block sits in the last ~100 KB of a 133 MB file, so a truncated APK looks complete and carries no signature at all. The release published neither a size nor a digest, so there was no way to tell that apart from a bad build without deriving both from the asset by hand. The release now uploads app-release.apk.sha256 next to the APK in `sha256sum -c` format (binary marker, so Git Bash cannot translate line endings while hashing) and puts the exact byte size and digest in the release body, naming `shasum -a 256 -c` for readers on macOS. The upload path rewrites the body too: --clobber replaces the APK, so a digest left over from the previous build would describe a file nobody can download, and a reader comparing against it would reject a good APK. Both paths reserve the section's own length out of the release-body cap before truncating, so the section always survives and the body always fits; MAX_RELEASE_BODY_LENGTH is exported from the desktop release script rather than restated. Refs #24011, #12248, #11444. |
||
|
|
d2dbe2c385 |
fix(windows): replace the managed CLI launcher with a native one (#24094)
* docs(security): add the antivirus clearance path for future releases Every AV false positive here has been handled one vendor and one shipped version at a time. Document the programs that clear future releases instead -- signer and product enrollment rather than per-build sample submission -- and add a script that reports an RC's current detection state by hash, so a verdict is found before users meet it in an issue report. Hash lookup only by default; --upload transmits the artifact and stays manual. * fix(windows): replace the managed CLI launcher with a native one resources\bin\orca.exe was a csc-compiled MSIL assembly: a small, freshly compiled .NET image in a user-writable directory that mutates environment variables and proxies a child process. That is the shape .NET dropper heuristics are trained on, and every verdict against it named the family -- MSILHeracles from two vendors, Wacatac!ml from a third. Signing the file does not change its shape, so signing never cleared it. Rebuild it in Rust. Same resolution, same environment contract, same argv passthrough that keeps newline-bearing orchestration bodies intact (#8374), and the child still inherits our environment block rather than an explicit map, so a block carrying both PATH and Path survives (#12046). The PE now carries publisher, version, icon and an asInvoker manifest from build.rs. Refs #23383 * ci(windows): install the Rust toolchain before building the CLI launcher The hosted runners happen to ship cargo, but a real Windows dev box does not -- verified on our own Windows QA host, where cargo and rustc were both absent. Relying on the image means a future image change fails deep inside electron-builder's native hook instead of at an obvious step. |
||
|
|
e594cb06af |
test(mobile): record RPC goldens without a pinned commit, and check recorded requests against the desktop's params rules (#23732)
* test(mobile): add rpc:diff to decode what a recording change moved
The RPC recording goldens are content-addressed JSON, so their raw git diff is
pool hashes. `pnpm --dir mobile rpc:diff [<base>]` decodes both sides and prints,
per golden, the checkpoint, field and JSON path that moved with both values,
grouped across checkpoints, plus added and removed goldens. `--summary <file>`
appends a Markdown report capped for GitHub's step-summary limit.
It reads any pooled format, so it can prove the next commit's format change
moves no recorded value. Checkpoints are matched by occurrence because an id can
repeat within one golden.
This commit adds files under the recorder directory, which moves the header
digest every golden pins; the next commit removes that header.
* test(mobile): record RPC goldens without a pinned commit or input digests
Every golden carried a pinned `baseline` commit plus digests of the recorder,
its mount adapter and its scenario, and the record script refused to run unless
the product tree matched the pin. So every behaviour change repinned to its own
branch commit and rewrote all ~790 files, the squash made that commit
unreachable, and main's pin job stayed red until a hand-made repin pull request
landed (22 of them in 12 days). The digests could only fail when an input moved
and the recording did not, which is exactly the change that carries no
information; every run already re-derives each golden from the current tree and
compares it.
Format 6 keeps the format version, operation, family, named deltas, the value
pool and the recording. Removed: the pin and fence, the three digest modules and
their test, the pin guard and its CI job, and the dead scenario `version` field
(the manifest reader now refuses `baseline` and `version` with a message).
- `pnpm --dir mobile rpc:record [<golden-id>...] [--prune]` records all or some
goldens; orphans are listed, and deleted only with `--prune`. Every derived
test title now starts with its golden id so an id selects it.
- `compareGolden` reports every difference in one failure (identity fields by
name, the checkpoint list, each checkpoint/field/path grouped), keeps the
final byte compare, and ends with the command to re-record that golden.
- `unhandled-recording.test.ts` now drives a detached rejection through
`runRecording` into a checkpoint and the cleanup checkpoint; no golden carries
one, and disconnecting the capture passed every suite before.
- Seam rules that existed only to keep a digest honest are gone; the
mutant-reachability, register-completeness and one-exposure rules stay.
- CI: `Mobile tests on main` runs the whole mobile suite on every merge that
touches mobile/, src/shared/, the root lockfile or the host RPC paths, since
`verify` never runs on main. A new `Mobile RPC Recording Replay` workflow
replays the recordings on pull requests that touch src/shared/ or the root
lockfile without touching mobile/. `verify` writes the `rpc:diff` report to
the job summary.
Proof: `rpc:diff` against the parent reports no recorded behaviour moved; each
golden only loses its ten header lines.
* test(mobile): check every recorded request against the host's params contract
The goldens script the host's replies, so a scenario could record a success
for a request the real host would refuse, and a desktop change that tightens a
params schema moved no golden at all.
`recorded-request-params.test.ts` parses every distinct request the corpus puts
on the wire with the host dispatcher's own `parseRpcRequestParams` and the
schema `rpc-params-catalog.generated.ts` binds to that method. It fails on a
method the host lacks, params it refuses, params sent to a method that takes
none (the dispatcher never reads them), and keys the schema silently strips
unless an inventory entry gives the reason; a stale entry fails too. Each rule
is also shown firing on a made-up request, since the corpus has no instance of
three of them. It imports the desktop dispatcher, so it sits beside the other
Node-side tests outside the RN test program, and the params-contract boundary
now exempts test files, which are never bundled.
It found twelve requests the host would refuse, all from invented fixture
values, not product code, fixed at their source:
- git.branchDiff sent `base-oid`/`head-oid`/`merge-base` where the host needs
full object ids (diff-review and source-control adapters, and the branch
compare replies in the manifest that feed them);
- an iOS push registration without `apnsEnvironment`, which a real iOS token
always carries (`push-token.ts`); the adapter now defaults to `production`;
- `settings.update` given Linear's `assigned` filter as a GitHub preset, which
the product type forbids; the scenario now picks `my-issues`;
- GitLab `projectRef` as a string where the host and the product type take
`{ host, path }` (7 methods, 5 adapters and the manifest).
46 goldens move, and a decoded comparison of every one of them shows no change
other than those substitutions; `rpc:diff` lists them.
* ci(mobile): detect a mobile change without a SIGPIPE-prone grep pipe
Under the runner's pipefail, grep -q exiting on its first match SIGPIPEs git
diff on a long file list, so a large pull request touching mobile/ read as
uncovered and replayed the recordings a second time.
* test(mobile): drop comments that still describe the golden header and digests
Eleven adapters justified an import rule by the header a golden no longer
carries, and that rule's test is gone. The census failure now names the
rpc:record and --prune commands.
* test(mobile): refuse a golden that keeps a key no recording writes
Decoding dropped unknown top-level keys, so an old header left behind by a
hand-resolved merge conflict passed every compare unseen.
* ci(mobile): summarize RPC recording changes after a failed test step too
* test(mobile): stream rpc:record output instead of capturing it
A captured run stayed silent for its whole duration and clipped its tail,
where the failure summary sits, past 8 MB.
* test(ci): let the Ruby-gate contract skip the always-run RPC summary step
|
||
|
|
9f4311598f |
fix(codex): trust the worktree Codex starts in, not a guessed repo root (#23937)
* fix(codex): trust the path Codex checks for bare-repo worktrees Codex keys a linked worktree's trust on the main checkout only when that checkout's .git leads back to the common git dir; otherwise (bare repo, --separate-git-dir) it keys on the worktree itself. Orca always wrote the main-checkout key, so Codex showed its trust prompt and worker-start failed at agent_readiness. Mirror trust.rs exactly, and pin it with a real-binary contract that runs in the existing Codex contract CI job. Fixes #23847 * fix(codex): trust the worktree path itself instead of mirroring trust.rs Codex looks up the cwd's own [projects] entry before any repo root (config_toml.rs get_active_project, loader decision_for_dir), so trusting the workspace realpath satisfies every git layout. Drops the resolve_root_git_project_for_trust mirror: simpler, cannot drift from Codex, and never widens trust past the folder Orca launched in. Cost is one config entry per worktree; entries older Orca wrote on main checkouts stay valid. The six real-git layout tests now assert the workspace key and, under the contract, that real Codex starts workspaceWrite for each. The contract probes the binary version once and fails at load when required but missing; its CI path filter now includes config-toml-trust. |
||
|
|
9420d49bcb | fix(terminal): run Codex in Orca terminals without the shared background server (#23900) | ||
|
|
f8f656ca19 |
perf(ci): spend fewer concurrency slots per pull request (#23810)
A concurrency slot is charged per job, not per core, and the account's cap is the scarce resource: standard runner minutes are free and unlimited on a public repository. Two paths spent slots that bought nothing. The unit matrix ran eight fixed shards averaging 6.5 minutes each, 3384 job-slots a day and 68% of all slot demand, while the arm pool queued 10.5 minutes at p95 — the queue was the oversharding. Five shards run the same work in ~10.5 minutes each for three fewer slots per run. Bun profile persistence escalated to all six platforms on `config/`, `resources/` and `.github/` wholesale, which took 36.5% of the last 1100 commits through the full matrix where a platform-flavoured predicate takes 19%. A pull request now qualifies one platform unless the change is platform- flavoured, and the push to main re-qualifies all six, so an unescalated miss surfaces minutes after merge rather than at the next cron. Missing changed-file evidence and an unavailable dependency graph still fail closed to all six. |
||
|
|
f6324f242a |
chore(ci): stop auto-filing community PRs onto the project board (#23796)
Removes the Track Community PRs workflow. Community pull requests will no longer be added to project stablyai/13 automatically. |
||
|
|
2ea3fb1d46 |
perf(ci): take advisory unit-selection evidence off the gate (#23776)
selection_evidence is continue-on-error on both the job and its comparison step, so it can never fail a PR -- it downloads the shard reports, compares selection against the full results and uploads a review artifact. But a caller's `needs: test` waits for every job in the called workflow, so living inside unit-tests.yml it held verify for ~36s after the last shard finished. It moves to its own reusable workflow called as a sibling, so it still runs on every PR and still uploads its artifact, but verify no longer waits for it. It is deliberately absent from verify's needs, and a contract test pins both that and its advisory status so it cannot drift back onto the critical path. Measured on a recent run: the shards finished, then selection_evidence ran 36s, then verify 3s. Only the last of those gates anything. |
||
|
|
c5fc0c6f26 |
fix(ci): keep a squash-merged RPC recording pin reachable through its pull request (#23720)
* fix(ci): keep a squash-merged RPC recording pin reachable through its pull request
Main's "RPC recording pin" check has been red since #22762: that branch pinned
the recording corpus to its own commit
|
||
|
|
ec9f35e2ee |
perf(ci): plan the unit shards before the static-analysis gate instead of behind it (#23743)
A caller's `needs` gate the whole called workflow, so while the plan job lived in unit-tests.yml it could not start until static analysis and typecheck had both finished and passed -- and the shard matrix then waited on it. The two hops were serial when they did not need to be: planning reads the checkout, a git diff against HEAD^1, the import graph and the checked-in timing baseline in config/scripts/ci-shard-timings.json, and consumes nothing that static analysis, typecheck or the native-cache primer produce. Planning moves to its own reusable workflow so pr.yml can run it against code_paths alone, overlapping it with the gate. Measured across 99 runs, the shard matrix is created a median 93s earlier (p25 47s, p90 241s, never later). Planning stays a required predecessor of the shards, so an empty assignment cannot expand the matrix. The gate itself is deliberately left in place. It fires on 22% of runs, and the shard queue wait knees hard above ~9 concurrent ARM jobs -- 4s median below that against 218s at 15-19 -- so admitting 8 doomed shards per failed run would cost more in queue pressure than it returns in latency. Cost is one 37s ubuntu-latest job, which does not touch the ARM pool the shards contend for. A planning failure still fails the PR: the shards are skipped, and verify's check_job requires success whenever the classifier says tests should run, so it reports `test: expected success, got skipped`. |
||
|
|
2ca4ecbc61 |
feat(orchestration): let a structured chat run orchestration as itself (#22568)
* feat(orchestration): inject the Orca session id into structured children and let the CLI act as it Every structured session's child (native Claude, native Codex, and the terminal view) carries ORCA_AGENT_SESSION_ID and reaches the Orca CLI. The CLI sends the id in the orchestration envelope; when present it is the caller, and a caller flag naming anyone else is refused before any request. The id is stripped from inherited PTY env and from the SSH host-CLI passthrough, and crosses into WSL so the host can refuse the cross-host claim. * test(orchestration): pin session id injection for native Claude, native Codex, the terminal view, WSL, PTY inheritance and SSH * test(orchestration): pin one caller precedence rule across every CLI verb that names its caller Adds the per-verb table (flagless acts as the session; a conflicting --from or --terminal is refused before any request; the session's own spellings are accepted), the enumerated guess population with its positive control, the structured worker's own handle, the identity-less refusal for an older child, the unchanged terminal agent, and the envelope. dispatch-show's --from only fills preview text, so it passes through unfenced and a session's flagless preview names the address the real dispatch writes. * refactor(orchestration): keep the identity-less marker reader to the marker; the id is checked first * test(orchestration): pin that a host refusal of the session surfaces verbatim from the CLI * fix(orchestration): keep the identity-less marker beside the id for CLIs that predate it A CLI older than the id, reached through a global install when a shell rc resets PATH, would otherwise guess a sibling's terminal in a chat that no longer carries the marker. It refuses on the marker instead; a current CLI checks the id first, so the marker never makes a session with an id identity-less. * fix(orchestration): refuse a conflicting --from on gate-list and task-list scoped by --run A --run listing needs no caller, so both handlers skipped the resolver and a --from naming another actor was dropped silently under a session. The conflict check now runs on that branch too; terminal callers are unchanged. * fix(orchestration): name this app's CLI by absolute path for a structured session's login shells A provider can run each command in a login shell: Codex runs zsh -lc, and the profile rebuilds PATH, putting a global install (possibly an older Orca) ahead of the directory Orca prepended. ORCA_CLI_COMMAND, which an agent resolves the CLI from first, is now the absolute launcher in that directory (the native launcher on Windows), so no shell's startup files can swap it. The PATH prepend stays for shells that read no profile. Found by the live coordinator run of the next PR. * test(orchestration): pin a structured worker's CLI command as this app's absolute launcher * test(orchestration): run the zsh login-shell arm in the real-shell lane that installs zsh The ordinary Linux unit lane has no /bin/zsh, so the zsh arm failed there with ENOENT. It moves to a live-shell file registered in the shell-contracts lane; the bash arm keeps running in every lane. The lane guard's detector now also sees a zsh spawned through the ProcessSpec program field, which is how this test escaped it. * fix(orchestration): omit a structured child's CLI command when no launcher resolves, and pin its instance A bare `orca` fallback named GNOME's screen reader on packaged Linux, and an inherited value named another app's CLI. The builder now deletes any inherited value, sets the absolute launcher only when one resolved, and pins ORCA_USER_DATA_PATH so a current CLI dials the instance that minted the id. Renames the marker reader to hasStructuredSessionMarker and records why the terminal view carries the id without the marker. * fix(terminal): name this app's CLI launcher by absolute path in every local terminal ORCA_CLI_COMMAND meant three things by lane: an absolute launcher for a structured session, a bare name for WSL, and nothing for any other terminal, so a structured session's terminal view lost it. Local terminals now get the same absolute launcher the structured lane gets; WSL keeps its guest command name, and a terminal whose launcher does not resolve still gets none. * feat(cli): hand a command to the session's own CLI when another Orca CLI was invoked A login shell can reorder PATH behind a global install, and an agent or its helper script can run bare `orca`, so the binary that answered depended on the agent following instructions. Orca's packaged launchers and bare-orca shims now export ORCA_CLI_SELF (outermost wins). At the CLI entry, when it names a different launcher than ORCA_CLI_COMMAND, the command re-runs once through the named launcher with ORCA_CLI_REEXEC=1 and exits with its status; both variables are consumed so no child inherits them. Dev launchers export no self on purpose, WSL and SSH names never qualify, and a launcher that cannot start leaves the command to run here. The Windows launcher no longer rewrites ORCA_CLI_COMMAND; the legacy ask protocol normalizes its resume command itself. * refactor(orchestration): declare which flag names the caller on each spec and refuse at the CLI entry Each handler hand-classified its --from/--terminal as the caller or a target, and the refusal of a conflicting caller flag ran inside the caller resolver plus two standalone calls for --run listings, so a new verb that read its flag raw would pass a sibling's handle to a pre-session host. Specs now declare identityFlagRoles, the CLI entry refuses a conflicting caller flag once from the spec, the resolver only applies the id-wins rule, and a test fails any orchestration verb that accepts --from or --terminal without classifying it. * perf(cli): keep the session caller check off the actor codec's module graph The check runs at the CLI entry for every command, and the actor codec pulls zod through the session record. Compare the session's own spellings as plain strings instead. * refactor(cli): spell a session's address from the one prefix constant, off the codec's module graph The Orca session address prefix moves to a leaf module with no imports, re-exported by the address codec, so the CLI entry check derives `session:<id>` from that constant instead of re-typing it and still stays off the codec's zod graph. Prose and test names say caller or Orca session id, not actor. * refactor(orchestration): drop the session id's terminal-view spawn now that the handoff is gone The terminal handoff was removed, so no terminal is ever a structured session: - delete the terminal-view identity env and its WSL passthrough, and their tests; - strip the session caller keys from every terminal's env unconditionally; - the CLI's own-address spelling moves beside the injected id in src/shared, with a test pinning it to the address the host's party resolver gives that session. * fix(terminal): run the Codex launch preflight through the CLI the terminal names Packaged Linux names the userData shim in ORCA_CLI_COMMAND, while the preflight ran the bundled launcher behind it. The CLI saw a different launcher and handed the preflight off to the shim, booting Electron twice before every codex launch. * revert(terminal): keep terminals on main's ORCA_CLI_COMMAND and Codex preflight Only a structured session needs an absolute ORCA_CLI_COMMAND; local terminals go back to naming none (WSL keeps its guest command), and the Codex launch preflight goes back to the bundled launcher. The CLI handoff is scoped to sessions, so a terminal's preflight can no longer be handed off and start Electron twice. This reverts commit |
||
|
|
8c61a5df1f | fix(windows): require signed release binaries and identify CLI launcher (#23680) | ||
|
|
0f52bb8be5 |
perf(ci): use four ARM test workers and remove repeated compilation (#23685)
* ci: benchmark per-job Node compile caching on full unit shards * ci: measure unit shards with three and four workers * ci: benchmark localization extraction CLI patch * perf(build): reuse identical relay bundles across platforms * ci: compare Vitest 4 and 5 on complete ARM shards * perf(ci): upgrade localization extraction to skip irrelevant syntax walks * perf(ci): use all four ARM cores and remove benchmark workflows * ci: preserve failures while capturing unit source revision * fix(ci): preserve commented and escaped localization calls * ci: remove corrected localization benchmark harness |
||
|
|
95a16e3f67 |
fix(release): stop the release policy from deleting pipeline-cut releases (#23669)
* fix(release): stop the release policy from deleting pipeline-cut releases The policy judged a release by who created the release object. Cut Release reuses an existing draft, so a CI-built v1.4.216 whose draft a person had created was deleted (tag included) when its notes were edited, and Latest fell back to v1.4.214 because v1.4.215 was also published by a person. - Authorize a desktop release when its annotated tag was created by the release pipeline and points at its `release: vX` commit, not only by author. - Only delete on `published`; an edit never deletes a release or tag. - Pick Latest from the highest authorized stable using the same check. - Move the policy into config/scripts/release-policy.mjs with tests. * fix(release): load the policy module from the tagged commit Release events run the workflow file from the tag's commit, so checking out the default branch could pair an old workflow with a newer module. |
||
|
|
cf20423ff3 |
ci: skip unrelated installs and share xterm build dependencies (#23607)
* ci: pilot shared xterm installed dependencies * ci: bound xterm cache production to verified main entries * ci: benchmark xterm reuse on the production ARM runner * ci: avoid installing Orca dependencies for standalone xterm checks * ci: use Node-only setup in the production xterm job * ci: remove completed xterm benchmark workflow |
||
|
|
0fe8974354 |
perf(ci): move six more jobs to the free ARM runner (#23594)
* perf(ci): move six more jobs to the free ARM runner
Follows the static-analysis move, which measured 172s to 128s. Each of these was
checked for an x86 requirement rather than assumed portable.
pr.yml:
cross-version-wire source-only, tagged checkout plus in-process vitest
managed_hook_node18 Node 18 publishes linux-arm64; the per-platform
runtime files are read as data, so host arch is moot
codex_index_heal_contract @openai/codex ships @openai/codex-linux-arm64
shell_contracts fish 4.x is published for noble/arm64, zsh is in the
arm64 archive, so the fatal fish-4 gate still holds
mobile.yml:
verify 209s of its 298s is Vitest; no Android SDK, emulator,
gradle, Hermes or Watchman, no docker, no artifacts.
Gemfile.lock lists the generic `ruby` platform, so
frozen bundler resolves without an aarch64-linux entry
recording-pin pure Node plus git; the golden comparison masks
`platform`
Left on x86 deliberately:
package builds --x64 targets, its docker gates are
--platform linux/amd64, and it is where the glibc
floor check runs. node-pty's .symver pin is
arch-specific, so flipping would validate the arm64
pin and stop validating the shipped x64 one
orcad_browser Google ships no Linux arm64 Chrome
mobile_web_app same Chrome wall; its render check fails closed
git_compatibility its cache key carries runner.arch and the warmer is
x86, so flipping alone means a cold `make git` every
run
relay_integration no technical blocker, but x86 relay coverage is a
documented placement and the reusable workflow has no
per-job runner input
e2e and the ssh lanes the ssh jobs would silently retarget the tested
remote from linux-x64 to linux-arm64
* perf(ci): move xterm_patch_sync to ARM too
The patch check rebuilds 4 packages x 2 builds and byte-compares against the
checked-in bundles. Ran it on darwin-arm64: exit 0, in sync at 1362743 bytes,
with the full fetch-and-rebuild path exercised rather than a short-circuit. The
bundles were generated on Linux x64 and reproduce byte-for-byte on a different
arch and a different OS, so the output is host-independent.
|
||
|
|
2f8f4f576d |
perf(ci): run static analysis on the free ARM runner (#23576)
Measured 128s against 172s on ubuntu-latest, with every compute-bound step faster: type-aware 24s to 15s, anti-slop 28s to 19s, localization extraction 67s to 46s, and the orcad terminal smoke 39s to 14s. Checkout and the install were unchanged at 12s and 16s. The toolchain resolves on arm64: both lint engines ship linux-arm64 bindings (@oxlint/binding-linux-arm64-gnu, @oxlint-tsgolint/linux-arm64), and build-orcad-bun.mjs derives its target from process.arch. The orcad smoke booting and round-tripping a real PTY is the evidence that node-pty compiled and that Bun, the bundled ripgrep and @parcel/watcher all resolved. The runner is free for public repositories, the same one the typecheck job already uses. |
||
|
|
d05080175b |
perf(ci): stop duplicating shared Linux download caches (#23578)
* perf(ci): pilot pnpm verification record caching on Linux * test(ci): review pnpm verification record in mobile cache audit * perf(ci): share Electron downloads and clean closed PR caches * test(ci): retain cache ordering checks for restore-only consumers * perf(ci): limit archive sharing rollout to primed Linux hosts |
||
|
|
e6fbbdf684 |
perf(ci): cache pnpm verification records on Linux (#23568)
* perf(ci): pilot pnpm verification record caching on Linux * test(ci): review pnpm verification record in mobile cache audit |
||
|
|
080c562898 |
perf(ci): diff against the merge commit's first parent so PR checkouts can be shallow (#23562)
Every changed-path gate asked git for `--merge-base "$BASE_SHA" "$HEAD_SHA"`, which needs the event payload's base SHA to be in the local graph. That is the only reason two jobs cloned all 8127 refs' history. On a pull_request checkout HEAD is already the merge commit, so its first parent is the base side and no merge base has to be computed. config/scripts/git-pull-request-diff-base.mjs resolved that for the two Node gates; the workflow's inline gates now use the same helper through a small CLI rather than open-coding it. code_paths gates all 22 jobs, so its checkout is charged to the start of every one of them: measured 20.7s to 1.6s, keeping blob:none because its sparse tree is ~7 files and leaves no blobs to refetch. Static analysis drops the filter instead, since populating all 30,226 files makes blob:none force a second promisor fetch: 23s to ~11s. Verified on a real merge ref. At depth 50 the old and new forms produce identical changed-file sets. At depth 2 the new form still works and the old one fails with `fatal: bad object`, which is the failure a stale base would have caused once the checkout stopped being complete. Also drops the dead resolveBase + merge-base prelude in the changed-code gate, whose result resolvePullRequestDiffBase already discarded on every PR. |
||
|
|
7a24d3d335 |
fix(native-chat): the conversation outlives its agent (#22835)
* refactor(native-chat): remove the unused terminal handoff No client ever called agentSession.requestHandoff or mounted the handoff chrome. Delete the handoff coordinator, the terminal-owner runtime, the proof write path and the unmounted UI. Keep agentSession.handoffStatus, which released desktop clients read for worktree activation, and let records an older build left mid handoff reconcile through the ordinary restart and recovery paths. * fix(native-chat): never let the pre-stop snapshot hold a chat's stop Eviction now drains delivered events before quit's resume-offer snapshot. An unbounded wait there sits ahead of the provider stop, so a sink whose journal write stalls kept the child running until the step deadline aborted the eviction. The offer is advisory: bound the drain and stop the child regardless. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(native-chat): drop helpers only the terminal handoff called `claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and `queryWindowsProcessRowsFresh` lost their last caller with the handoff. The fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`, the teardown path that still depends on that contract. Co-Authored-By: Claude <noreply@anthropic.com> * docs(native-chat): stop citing the removed handoff in lifecycle comments Six comments still named the handoff coordinator, a handoff suspend, or a terminal-owned session as live participants in the flows they describe. Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): type the stalled snapshot drain without a cast Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): pin that a start dead before proving owes no settlement The removed restart handoff test pinned this branch; nothing else did. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): keep the owner-status read behind an in-flight attach The handoff removal dropped the per-session queue from `handoffStatus`, so a read landing mid-start reported the reservation (no owner) instead of the settled chat owner, and shipped desktop clients blocked worktree activation on it. The read is queued again, as it was before the removal. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(terminal): remove the agent-session PTY write gate The gate only refused a write when a PTY had been bound to a chat session, and the only code that ever bound one was the terminal handoff this branch removes. With it gone, every admit/readmit returned "admitted" unconditionally, so the checks on the renderer write path, the runtime controller backstop, terminal.send, agent prompts, preview input and orchestration pointers, the refusal fields on terminal.send and worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane orchestration routing could no longer run. Ordinary writes take the same path in the same order as before. Co-Authored-By: Claude <noreply@anthropic.com> * refactor(native-chat): drop the transcript helpers only the handoff called appendLegacyTranscriptMessages fed the terminal transcript catch-up and proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost their last caller with the handoff. Their tests now go through the live entry points instead: the roster bounds through the legacy import, the pinned-read and growth tests through the ancestry replay the history window uses, and the marker rules through the string proof in their own file rather than the session-file resolver's. Co-Authored-By: Claude <noreply@anthropic.com> * fix(native-chat): stop calling a starting chat "mid-handoff" A send refused because the chat's owner is not settled showed "The session is mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that reach it are a chat that is still starting, or one whose previous agent process has not yet been confirmed stopped. The message now says which of the two it is. The refusal code is unchanged. Co-Authored-By: Claude <noreply@anthropic.com> * test(native-chat): type the stand-in roster decoder without a cast Co-Authored-By: Claude <noreply@anthropic.com> * refactor(codex): name the pinned rollout lookup for what it does With the terminal handoff gone, the module named codex-tui-rollout-proof holds only the pinned rollout lookup that structured Codex launches use to resume a thread, so the name described code that no longer exists. Rename the module and its options type. Also drop a mobile allowlist assertion that pinned the removed agentSession.requestHandoff method, which no longer exists to allow. * refactor(native-chat): type the owner-status reply as the host sends it The handoffStatus reply type still listed the terminal handoff's fields and states (terminal placement, host label, proof retry, queued and waiting phases, the to-terminal direction). No host writes them any more and the only client reader parses the reply as unknown, so they described nothing. The reply on the wire is unchanged. * refactor(native-chat): normalize terminal-handoff lease values once at decode Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types still admitted them, so readers across the host kept branches for values no path produces and the compiler could not point at them. The store now validates the on-disk shape, which still accepts those values so an older record is not quarantined, and maps them once while parsing: - `preparing` and `old-owner-stopped` become `recovering` - a `tui` lease becomes `native`; when it records a process it also becomes `conflicted`, the claim every build probes but never stops. A plain native owner would be stopped by restart recovery, here and in older builds. Revisions are taken over the normalized state on both sides of every compare, and the mapped record reaches disk with the store's first transaction, the same way the tab-id backfill does. The in-memory types narrow to what this build writes, and the branches that existed only for the removed values go. Structured-worker identity keeps its verdict for a former terminal owner by refusing a conflicted claim rather than a non-native kind. * refactor(native-chat): stop threading the owner kind through a reservation A reservation only ever names a native owner now, so the request no longer carries a kind and the reserved lease records `native` directly. The attach params keep `runtimeKind`: agentSession.ensure and create accept it, and the operation fingerprint stored in the ledger covers it. * test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else Hiding a tab also committed the visibility index, so the no-op transaction wrote the file even when its open-time revision was wrong. Committing the index first leaves the pending rewrite as the only reason to write. * fix(native-chat): name a chat write by its target, not the owner generation A write carried the fence of the last frame the pane read, and the host refused it unless that fence was still current. An idle release and the restart after it each move the fence, and the release publishes nothing, so a send after a release was refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a cold start was refused as stale. Every write already names what it acts on: a send its conversation, a cancel its turn, a prompt answer its item revision, a rewind its epoch; an option is last-writer-wins. So admission stops comparing the client's fence, and the rebase that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it. The writer-lease check stays, and so does the attach's compare-and-swap. Frames now stamp the fence read when each frame is sent instead of a copy each subscriber kept, which went stale on the same release. * fix(native-chat): every journal append reaches the chats that are open A journal write and its delivery to open readers were two calls, and some writers made only the first. A failed start whose lease could not be handed back, a provider revision with no frame behind it, and eviction's settlement were all journaled without reaching an open chat. A journal handle now reports every durable change, and the host's session map binds that report to the session's readers when the handle is set. Writers no longer publish what they append; the per-writer publish calls are deleted. * test(native-chat): an epoch replacement reaches the open chat * test(native-chat): each row reaches an open chat once, and a live handle enters only through the map * test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite The seeded record had no surface tab id, so the next open backfilled one and that rewrite alone made the no-op transaction write. The test passed with the legacy-lease rewrite signal removed. * test(worktree-activation): restore the OMP surfaced-agent resume test The handoff removal deleted it alongside the terminal-owner tests, but it covers the surfaced-PTY block that still guards resume, including an agent whose ownership is unknown. * perf(native-chat): a publish behind a delivered commit reads nothing Each commit now delivers itself, so the publish a provider frame still sends afterwards found every reader caught up but still read rows and rebuilt the timeline for each one. A caught-up reader now skips the read. * test(native-chat): state why the teardown test's fake journal is safe to cast * docs(native-chat): say mutation admission checks only the writer lease * docs(native-chat): drop the send rebase from comments that still described it * fix(native-chat): a message is accepted, then delivered A send to a chat with no running agent restarted the agent inside the send call, before the message was recorded, so the client waited for the whole start and a failed restart refused the message. Claude held prompts sent during startup, and those could settle as "unconfirmed". A send is now accepted inside the session's serialized queue: one ledger row and one submission row marked handoverRecorded, published, answered pending. A per-session delivery loop exists while a message is queued. It starts the agent through the same serialized attach a hold uses, waits outside the queue for a Claude child to prove its start, and hands the oldest queued message over as its own serialized step, writing dispatch{pending} before the adapter call. A start it needed and did not get writes one error-tone row and rejects every queued message with the same words; a start Stop cancelled writes none. Settlement follows from the rows. A queued message is provably unwritten, so a close, an eviction or an exit rejects it. A handed-over message stays in doubt. A queued row at or below the sequence a handle found when it opened was left by an earlier process and is rejected at open, with no latch. Stop withdraws queued messages with no writer lease and no fence. An attach failure keeps the conversation open, and the attach adopts its journal. Owed work counts the loop and queued rows. A compaction or rewind found prepared when a conversation opens was started under a child this process no longer has, so the open settles it rather than leaving it to refuse every send until a view attaches. The open cursor is scoped to its epoch, because sequences restart when an epoch is replaced. Deleted: restart-before-admission, recordFailedRestart, the fence rebase, Claude's startup gate, the attach's forget on failure and its own crash boundary. Clients without agent-session.accepted-send.v1 get their reply held until the handover; the desktop and paired desktop lists advertise it. * fix(native-chat): settle queued messages only for the child that ended A child that proved its start and then exited before its message was handed over left the message queued: the exit settlement returned early when nothing else was in flight. Delivery then started another child for it, and a child that died the same way started another, without end and without a row. A retried settlement for an earlier generation, run by the attach that delivery started, did the opposite: with that generation's turn unfinished it rejected the message queued for the child being attached. The settlement now takes the rejection for queued messages from its caller. The unexpected exit and the eviction pass one, and it applies even with no other work in flight; the retry for an earlier generation passes none. * fix(native-chat): an adoption that fails to import keeps the conversation open The attach now writes into the conversation's own open journal, but a failed transcript import still closed it as if it were the attach's provisional one. The conversation stayed indexed with a closed journal, so every later send answered "could not be recorded" and every attach failed again until the app restarted. The import now closes only a journal the attach opened for itself. * perf(native-chat): the recovering open reads the journal once Every conversation open now goes through the recovering open, including the read restore of every chat at startup, which used to replay its journal once. The recovering open replayed it twice: once to probe it and again inside the open. The probe is now handed to the open as its load. * fix(native-chat): an attach that fails after indexing its child leaves no child behind A failed attach now keeps the conversation open, but a failure after `onAttached` indexed the child (the rewind or compaction recovery, or the attach's own success record) left that entry claiming a child the failure path had already released. The next send found the phantom, skipped the start, and wrote at a fence the journal had moved past, so the message stayed queued for good. The entry now drops the released child and its event sink, and follows the record's fence, as a failure before indexing already did. * fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer The error strip for a message the host accepted and then did not deliver matched the entry before the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not send your message" with nothing to retry. It now reads the reconciled entry. A rejection the journal records before the send's own pending answer lands is final as well: that answer no longer puts the entry back to dispatching with no Retry. * fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down The preamble waits for its submission to be delivered while the worker's agent starts. When that wait ran out it threw operation_unknown, and the failed-start teardown then closed the session, which rejected the very preamble the host was about to deliver. It now reports a turn start nobody observed yet: the worker is start-unknown with its session kept, the host delivers the preamble when the agent starts, and the worker's report settles the dispatch as for any unobserved start. The receipt no longer suggests reading a screen a structured worker lacks. * fix(native-chat): a message rejected while its chat was closed reads as not sent A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked every later message behind a Retry and no reason, and the delivery probe, seeing the journal already answered, never ran. The reconcile now settles it as rejected like a dispatching one. * test(orchestration): name why the readiness settlement fakes are cast * fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent * docs(native-chat): drop the fence from the admission the send effects run behind * docs(native-chat): give the fence move on release the reason that still holds * docs(native-chat): stop citing a write fence check in launch and mailbox comments Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease. * refactor(native-chat): the provider child is its own record A conversation now outlives any number of provider children, so the child is one record on the conversation's entry instead of five loose fields beside its journal. It is written in one place: indexed only once an attach has fully succeeded, and ended through one function that an exit, a failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence. - A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence patch after it are gone. - Conversation writes read the record's fence, the way mutation admission already does; a child's own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of the conversation's fence, are gone. - The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer dropped when an attach replaced the whole entry. - Stop on a child still proving its start stops only the child: its lease goes back and the chat is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus the conversation's close. - The settlement retry uses the conversation's own journal, opened through the host's one open. * fix(native-chat): the delivery loop alone settles a message its start or child failed A queued message was settled by whichever path happened to end the child first: the loop, the unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that rejected every pending row. That gave two failure rows with different tones for one start, a loop that could hand over to a different child than the one it waited on, and a Claude start that died while starting reading unlike every other failed start. - The loop remembers the child it waited on. At handover, if that child is gone or replaced, it reads how it ended: a Stop continues; anything else writes one failure row and rejects every queued message with the same words, then stops. A child still starting whose start the adapter says did not land fails the same way. The exit, eviction and the settlement retry only settle the handed-over and legacy rows of the child that ended. - One failure row, always an error, keyed by the start. A start a view began that dies with nothing queued writes the same row through the same builder, so a second report revises it. - The open no longer rejects leftovers; the loop's first step does, and the open wakes it. - `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the failure before the exit is processed. - Quit closes every conversation the way closing a chat does: what is still queued is rejected as closed, with or without a child, and a start the loop already has in flight is waited for so the child it produces is stopped rather than left behind. * refactor(native-chat): a stopped child ends on the one reading of its stop The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that verdict to the child's ending, so the host never forms a second view of whether the root is gone. Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition, and a failed re-attach passes what its release saw. The end-of-child record can therefore also carry a stop whose root was not seen to go, which nothing ends on yet. * feat(native-chat): the host says it accepts a send before any agent has it The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same string capable clients already send. A client can then tell a host that answers a send at acceptance, and admits a Stop with no writer before a turn starts, from an older one that still restarts the agent inside the send. Additive: an older client ignores a capability it does not know. * refactor(native-chat): an attach never opens a journal of its own The attach adopts the conversation's open journal, which outlives it, so it no longer opens one for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag that told the two cases apart is gone. Tests that attach without a host open the conversation the way a host does. * fix(native-chat): a moved fence resends nothing on a host that accepts first The outbox treated any fence change as a new owner: it dropped the answer of a send in flight, queued that send to go out again under the same id, and unblocked a refused head. On an older host that is how a send the restart refused, unrecorded, gets another try. On a host that records every send before it starts an agent, a fence moves because that start ran, so the same rule resent into every failed start. With a fence stamped on every frame, that became a loop. The outbox now reacts to a fence change only when the host has not advertised that it accepts a send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed start reaches the client as a rejected message it keeps with its Retry. Against an older host, or before one has answered, the outbox behaves as it did. Desktop and paired web share this hook. * refactor(native-chat): a child's end says whether the user or the host stopped it The end-of-child record's cause now tells a user's Stop from the host stopping the child for a cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a user's Stop, as before, and fails the start it was waiting on after a host stop, with the one error row and every queued message rejected, in the stop's reason when it gave one. The reason stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet. * fix(native-chat): a chat whose only work is a queued message is not offered for resume A message accepted while the agent was starting counts as working in the chat, and quit rejects it as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a chat whose agent never had the message. The snapshot now reads only what was handed over. * fix(native-chat): the conversation outlives its agent Opening a chat no longer starts its agent. A conversation is reached through one host accessor that opens its journal at rest, and a send is what starts the agent, through the delivery loop. One idle sweep, every five minutes, stops an agent that has been quiet for thirty minutes and owes no work, then drops an open journal handle that is only a cache. Its record, tab, status row and readers stay. - hold and release are no-ops; hold still builds the host for shipped mobile builds. - The holders, the holds, the release clock and the exit respawn are deleted. - Options, the model list, the goal and the context meter answer at rest; a model pick at rest is recorded as intent for the next start. - Compact, rewind, clear and goal changes start the agent first. A send does too when a rewind is still in doubt after the conversation opens. - Orchestration routes mail and group addresses on ownership (the record plus the chat tab), not on whether the process runs. An open dispatch keeps its worker running. - The restart continuation is a send; Resume all holds each slot until the message is handed over or rejected. - A read error never replaces a loaded transcript, and shows the host's own words. * test(native-chat): type the queued-message fixtures in the resume-offer tests * fix(native-chat): a start that dies while a message waits on it is that message's failed start Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When that start died, its exit wrote the start's error row and left the message queued, so the delivery loop started a second agent into the same failure and wrote a second row. A child's end now records where the conversation's journal stood, and the loop settles a message accepted before a failed start ended with that start: one row, under its key, and no second start. A message sent after the failure still gets a fresh start. * docs(native-chat): say what an attach's open conversation and unconfirmed ids are now * test(native-chat): pin what a failed start settles, and what a resume offer names A view's child that dies while a sent message waits settles that message only when it died starting and no child has taken its place: a proven child's crash, or a second start since, gets the message delivered. The resume offer names the handed-over message, never a newer one still queued. * test(native-chat): the failed-start pins fail on what the message became, not on a timeout * fix(native-chat): a restart offer ends when the chat's agent starts again The offer used to end only when the chat's newest user message changed, because opening a chat started its agent and that start could not be told apart from real activity. Opening a chat starts nothing now, so the host reads the fact it already publishes: a chat's status row goes from not host-owned to host-owned exactly when its agent is started. At that edge the offer and any failure record for the chat are withdrawn, unless the start is a resume action's own (its continuation is the oldest undelivered message). A continuation and a message racing to be first are decided at acceptance: the continuation is refused, quietly and with nothing filed, when any other message was accepted since the restart. A failed continuation start leaves the offer retryable, and each resume action sends its own message id. Deleted: the newest-user-message comparison, its journal reader, the continuation filter, and the failure ledger's own "answered by the chat" check. The marker still carries its message id for one release, so the previous build can read it. * fix(runtime): end a transcript stream when its client unsubscribes Desktop: the IPC subscription controller was dropped as soon as the streaming handler returned, which for most streams is right after it binds. A later runtime:unsubscribe then found nothing to abort, so the host kept the subscriber and derived and sent every publish to a channel no one listened to. The controller now lives until the renderer unsubscribes, resubscribes the same id, or goes away. Mobile: disposing an agentSession.subscribe stream now sends agentSession.unsubscribe with the stream's frame id, so the host ends that subscriber and leaves a sibling stream on the same socket running. The direct path now passes the frame id the relay path already passed. * test(orchestration): the preamble's host stub is typed, not cast The preamble send now takes only what it reads of the host, the send, the settlement wait and the record's fence, so its test builds that host with real types instead of `as never`. * fix(native-chat): one fact ends a restart offer: the chat moved on since the restart The offer is live while no other message has been accepted in the chat since the restart and its agent has not proved a start since. The offer list, the resume's reservation check and the continuation's acceptance check all read that one fact, so a message whose start then failed withdraws the offer too, and a stale click finds nothing to act on. The fact is read off the conversation's open handle, which the restart closed, so it is retired durably whenever it may have changed: a message accepted, a start proven. A close and reopen within the same run therefore cannot bring the offer back. A continuation rejected before it reached the agent does not count, so a retry after a failed start still runs. Deleted: the quit-time gate on withdrawal, which changed nothing because the withdrawal and the quit's own offer write share one queue; the per-action "withdrawn" flag and the separate acceptance check it paired with. * test(native-chat): an older build reads the restart offer this build records The offer lives in a file the previous release reads after a downgrade. Pin that against the pinned release's own capsule, and run the lane when the marker or the capsule changes. * fix(native-chat): read a restart offer against where the journal stood when it was taken "Since the restart" was read off the conversation's open handle, which the idle sweep closes: after a reopen, a message the user had already sent looked older than the handle and the withdrawn offer came back. The offer now records the journal position (epoch and sequence) at the moment it is taken, and a message accepted after that position, or a journal on another epoch, means the chat moved on. That is derived from the journal, so it holds across any number of closes and reopens. An older build's offer has no position; only a start withdraws it. Because the message half is now durable, the offer is no longer rewritten in the recovery file on every accepted message; a proven start still writes it, since only the host that saw the start knows of it. * test(native-chat): wait for the listing's retire write before reading the recovery file * fix(native-chat): keep the terminal-backed chat's read error over its local echoes Messages winning over a read error is right for the structured chat, whose read retries and whose messages came from the transcript. The terminal-backed view assembles its list from local echoes too (a launch prompt, a pending send), so a failed read there showed only those bubbles and no error. Only the structured pane now keeps messages over an error. * fix(native-chat): a start retries the exit settlement a failed journal write left owed An agent exit whose journal settlement write failed releases the lease latched until a retry lands. Reopening the chat used to be that retry; with reveal now only opening the journal, nothing retried it before the next app launch, and every send was refused. The start the send needs now runs the retry first, where the attach would. * perf(native-chat): answer the owner check without opening the chat Worktree activation calls agentSession.handoffStatus for every chat tab in the worktree, and the answer comes from the session record alone. Reaching it through the accessor opened each resting chat's journal (a full read, the crash-boundary write and a restored status publish), then kept it open for the idle window. It now checks the record and the adapter's support, as before this series, and opens nothing. * fix(native-chat): a read waiting on the session lock opens nothing once quit began The accessor checked for quit before queueing the open, so a read queued behind a session task ran its open after teardown had begun and indexed a journal no teardown step would close. The check now runs at the open itself. * fix(native-chat): read a failed resume's chat before calling it retryable Whether a failed resume is retryable is the offer's own rule: the chat has not moved on since the restart, read from its journal. The failure list read it only for a chat already open, so once the idle sweep closed a chat the user had moved on in, its failure showed Retry again, and the click did nothing. The list now opens the failed chats first, as the offer list does. * test(native-chat): type the provider event sink the settlement test reaches for * fix(native-chat): say the structured read keeps trying only where it does The structured pane's "Orca keeps trying to load it" line never showed: the view state filled in an untranslated fallback whenever the read error had no text, and the empty state prefers any message. The view state now leaves the message out, so the structured pane shows that line and the terminal-backed pane its own translated one. Mobile's structured lane does not resubscribe after an error frame, so it no longer makes the claim. * test(native-chat): await the send's settlement instead of polling for the start The at-rest send tests polled for the provider start with vi.waitFor's one-second default, which a loaded machine outran. They now await the host's own settlement of the message. * fix(native-chat): a restart offer resumes any time after the quit, and knows its own continuations The continuation's message id was dated by the quit, and the ledger refuses a new id dated more than a day back, so Resume or Retry a day after quitting was always refused (on main too). It is now dated by the resume action. Telling a rejected continuation from the user's own message read the operation ledger, whose rows expire after about a day; after that a failed resume stopped being retryable. The offer now records the continuation each action sends on its own capsule entry, bounded to the newest 16, so the ids end with the offer. The ledger read is deleted. * fix(orchestration): route no mail to a structured worker its orchestration released A structured worker is routed on ownership, and a resting worker's lease is released, so ownership held while its chat tab stayed listed. A worker the coordinator abandoned and then released, found at rest by the release, therefore still took peer mail and @worktree: broadcasts, and each one restarted its agent. Routing now also reads the orchestration's own resource row: once it is released, direct mail, group addressing and worker-show's addressable answer drop the worker, as they would a terminal worker whose terminal closed. The chat tab stays, and nothing new is stored. * fix(native-chat): a failed retry names the user's prompt, not Orca's continuation A resume's continuation is written to the chat before its start, so after a failed attempt the chat's newest user message is that rejected continuation. A second failure then showed Orca's own restart text as the chat's prompt. A retry now keeps the prompt its first failure named. * fix(orchestration): read the released row optionally, as the authority does worker-show's observation called the row lookup directly, which a runtime double without it threw on and failed the structured tab-retirement release. * fix(native-chat): the status bar drops a restart offer the chat moved on from The renderer re-read the host's restart offer only when a failed chat showed activity, so after a message withdrew a pending offer the host answered no chats while the status bar kept counting one, and clicking it opened nothing. The same watch now covers pending offers: a status change in an offered chat asks the host again, once. * test(native-chat): a roster of idle or finished children does not keep an agent awake The sweep reads owed background work through the shared child-work liveness that upstream's release clock adopted; a child that went idle or finished is not work the agent still owes. * fix(orchestration): a task dispatched into a resting structured worker keeps it running The sweep's open-dispatch check read only the worker-start dispatch that owns the worker's terminal resource, so a task later dispatched to the same worker (orchestration dispatch --to, which writes a dispatch with no worker row) did not count: after thirty quiet minutes the worker was stopped while that task was open, and its coordinator read exited. Any unsettled dispatch addressed to the worker's process incarnation now counts, derived from the existing rows. * docs(native-chat): comments stop describing the hold this PR removed Eight comments still justified orderings and teardown choices by a viewer or dispatch hold that pinned the provider child. Nothing holds any more; the orderings stand for the binding's redrive subscription and parked mail, and a chat's agent runs from a send until the idle sweep rests it. Comment-only. * fix(native-chat): a restart offer keeps the start its own continuation made Whose start ended an offer was decided at read time, from whether the offer's continuation was still the queued message. Once the provider refused that continuation, the child it had started read as someone else's start, so the offer ended and its failure showed no Retry. The delivery loop now records which queued message a start is for on the in-memory child, and the child's end carries it; the offer counts a start as its own when that message is one of its continuations. * fix(native-chat): an agent gets a full idle window after its owed work ends The sweep measured quiet only from the last journal row, so once a subagent, command, monitor or dispatch that had outlived the window ended, the agent was stopped at the next tick. A child can read done before the lead's wake-up turn writes anything, and stopping in that gap loses the wake-up. The sweep now counts owed work it observes as activity, which gives the agent the full window afterwards, as the release clock it replaced did. * test(claude): the options-read fixture runs a live child The fixture marked its conversation running with a hasProviderChild field the session type does not have, so the read took the at-rest path and refused a session with no record. It now carries a child, which is what the read checks. * test(native-chat): host tests reach its collaborators through a typed seam The rest-test rig and three test files read the host's private members with Reflect.get and cast the result. The host now exposes one test-only accessor, collaboratorsForTests(), and the subscribers class a subscriberCountForTests() beside its existing retainedActivityCountForTests(), so the tests are checked against the real types and the casts are gone. * refactor(orchestration): one owner answers a structured worker's custody Routing, group addressing, worker-show and the idle sweep each composed their own reading of whether orchestration still holds a structured worker, so each new obligation or retirement state had to be added to every reader. structured-worker-custody now derives both answers from the worker-terminal list state coordinators see in worker-list: addressable is owned and not released, and owed work is an active custody or an unsettled task dispatched to the same incarnation. The owner's state is read through the remote dispatch attachment too, as the terminal transfer lookup already does. Behaviour is unchanged; a settled worker awaiting its coordinator still rests. * refactor(orchestration): owed work is an open dispatch on the worker's incarnation A supervised worker's own dispatch context stays open exactly while the worker is active, so the separate active-custody branch only repeated it. Owed work is now one fact, which also states the policy that a worker awaiting its coordinator's decision may rest, and both custody decisions are written once at the top of the module. * fix(native-chat): a restart offer knows its continuations by a tag in their id The offer recorded each continuation id in a list on its capsule entry, capped at 16, and a running action's id in memory. Both could disagree with the journal: past the cap an old rejected continuation read as the chat moving on, and a crash during a retry restored the failure's older entry, which lacked the retry's id. Each continuation id now carries a tag derived from the offer (its teardown and chat), then the action's own part, so any continuation of this offer, queued or rejected, is recognised from the journal row and the marker alone. The persisted list, its cap and the in-memory action map are deleted; the agent-start withdrawal keeps an offer whose own continuation the start was for, read against the stored marker. * test(runtime): the legacy-worker reveal test judges its stale snapshot inside the wait The tui-idle probe reads through readTerminal, which now awaits the structured worker check before the PTY read, so the probe's snapshot request starts a microtask later. vi.waitFor missed it on its first check and polled again at 50 ms, the same moment the wait's own 50 ms timeout fired. The stale snapshot then resolved after the wait had already timed out, so the test passed without judging it, and the rejection landed before any handler was attached. Vitest reported that as an unhandled error and failed the shard. Polling every 1 ms sees the request within a few ms, so the snapshot is judged while the wait is still pending. * fix(native-chat): a view never restarts a chat whose last start failed A Claude chat whose CLI exits during startup left one red row per start, and every time a view bound to it (the chat opening right after its create died, or the user switching back to it) the hold started the CLI again, so the same launch-failure row repeated. Only a send retries a failed start now, the same rule provider-exit recovery already applied; the rule lives in one predicate the hold, exit recovery and the delivery loop share. * fix(native-chat): the idle sweep reads owed work every tick Owed work counted as activity, but the sweep read it only once the idle window had elapsed, so it refreshed the clock at most once a window. Work that ended just before the next read left the agent to be stopped at that read, moments after the work ended, which is the gap the refresh was meant to cover. The sweep now reads owed work on every tick for a started agent, so the window always runs from the last tick that saw work owed. * fix(native-chat): a continuation handed to the agent stays sent The offer read its own continuation as not reaching the agent while its dispatch was pending, which also covered one already handed over and still unanswered. When the wait for that answer ended first, the failure it filed read as retryable, and a retry sent a second continuation to an agent that may have acted on the first. Only a continuation still queued, or rejected, is now read as unsent. * test(native-chat): start the child the loop waits on with an attach, not a second view A view no longer starts a child whose last start failed, so the R2 case that waits on a child started since the failure now gets that child from a client attach, the one non-send starter left. * fix(native-chat): settle a gone generation's turn wherever a conversation opens A send that opens a chat this process had not read yet (after a crash, from a phone or the CLI) went through the delivery open, which never settled what the dead generation left running; only the read restore and a successful acquire did. When the send's start then failed, the turn stayed running for every reader. The settlement now runs in the one journal open, at the crash boundary, for every opener except an acquisition, which settles from the evidence it read before its reserve; the read restore's separate step is gone. * test(native-chat): prove the next child's start settles the turn an earlier child left The R1 case lost its only settlement assertion when the latch it checked was deleted. It now seeds the running turn the earlier child left and asserts it ends at the exit's receipt, with the exit's row, before the message is handed to the new child. * test(native-chat): count a failed start's rows by row, not by text Comparing the set of texts passed when two different rows carried the same words, which is the duplicate the test exists to catch. * test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget * test(native-chat): the interrupted create's own retry continues again The merge of main's lease-latch fix replaced that test's retry of the interrupted create, under its own operation id, with a fresh start whose result nothing read. That fresh start passes with the released-reservation continuation deleted, so the case the fix exists for went untested. The retry and its assertion are main's again. * docs(native-chat): three comments that still had views starting agents A start with nothing queued now comes from a command, goal change or rewind; an interrupted compaction left alone would refuse every send, so no agent would ever start to finish it; and a current host raises the unattached read refusal only once quit began, with the attach window belonging to an older host. * test(native-chat): pin the open's and the send's start and row counts, however the view binds Opening a fresh chat whose starts fail makes one start and one row, with two views bound before or after the create's child died; one send makes one more of each. * test(native-chat): a reader's open settles the turn a failed exit settlement left running An exit whose settlement write failed leaves its turn running in the open journal. PR 1's open now settles it, and this pins the two reads that reach it here: a reader reopening a chat the idle sweep closed, and a read that opens the chat before the restart restore reaches it. * test(native-chat): the view-start test's starting window outlasts two subscriptions on a loaded runner A subscription reads the conversation before it returns, so under load the two views took longer than the create child's 300 ms start, which then exited before the test checked that it had not. The child now takes a second to fail. * fix(native-chat): settle a gone generation's turn at every open but an acquisition's The journal open skipped the settlement whenever the lease read reserved or live, to leave an acquisition's own open to the acquisition. But a lease a crashed process left in recovery also reads live, until the next acquire resolves it. A send that opened such a chat, from a phone or the CLI after a crash on a host that could not prove the old owner gone, skipped the settlement; when its start then failed, the dead turn stayed running for every reader. The acquisition now says it is the opener, and every other open settles, whatever the lease still claims. * test(native-chat): hold the create's start open until the views bind The "view binds while the create is still starting" case gave the create a 300 ms head start and asserted the views bound before it died. On a loaded runner the holds took longer, the create's exit landed first, and the case failed its own precondition. The create's initialize now waits on a gate the test releases once the views are bound. * test(native-chat): a read that reaches a crashed chat before the startup reconcile settles its turn On desktop the chat on screen at relaunch reads before startup reconciles the leases, while the dead process's lease still reads live. The open settles the turn it left running anyway, and the restore that follows finds it settled. * refactor(native-chat): drop the composer's second error formatter After the merge with main, every chat write in the composer path reports its failure as a typed outcome worded by the refusal-notice table, so the send's catch sees only a local throw. The {code, message} formatter this branch added for it has no payload left to format, and its claim to be the one way a chat words a failure is no longer true. The composer send is main's again. * test(native-chat): pin the reason on a message rejected while its chat was closed The reopen test checked only that the message reads as not sent; it now also checks the Retry row carries the host's reason. * docs(native-chat): drop the removed dispatch hold from six comments A worker's session no longer takes a dispatch hold, and no release clock rests a chat by visibility; the agent-launch comments, the abandon test, the teardown test and the refusal census still said so. * test(native-chat): rest the owner-status chat through the idle sweep, not a hold The activation-gate test from #22808 put its chat at rest by holding and releasing it, and passed the release-clock grace. This branch deleted both, so the case threw before it reached its assertions. It now moves the host's clock past the idle window and lets the sweep stop the agent and close the conversation, then asserts the same owner answer and activation gate. * fix(native-chat): show the structured pane's retrying line when a read fails The read transport always hands the pane the host's words, so the error state's "Orca keeps trying to load it" line, which showed only when there were none, was never seen: the pane showed the host's text twice, as its subtitle and on the status line under it. The structured pane now always says its read keeps retrying, and the host's text stays on the status line. The terminal-backed chat is unchanged. * test(native-chat): wait for a send's background start before the refusal oracle removes its store An accepted send wakes the delivery loop, which starts the agent in the background. The oracle's teardown disposed the loop but did not wait for that start, so its lease write could create a temp file in the store directory while the directory was being removed, failing the test with ENOTEMPTY about one run in four. The teardown now drains tracked starts before it closes the journals. * fix(native-chat): a start a message waited on gets one failure row, the delivery loop's When a queued message's start failed, two writers could report it under the same row: the delivery loop, when the adapter settled the start without proving it, and the exit settlement, when the child's exit landed. The last one won, so the chat's row could name a different cause than the one the message was rejected with, or be written twice. The exit settlement now writes the start's row only when no message is queued and the loop has not already recorded that start. A start for a command, goal change or rewind, with nothing queued, still gets its row from the exit. --------- Co-authored-by: Claude <noreply@anthropic.com> |