* fix(ci): run mobile typechecks without concurrent dependency refresh
* test(ci): check effective Linux E2E package list
* test(ci): preserve the mobile production compiler barrier
---------
Co-authored-by: Orca Integration Recovery <orca-validation@invalid.example>
* Let measured cache producers keep stores without downloading them
* Check that restore-only callers do not publish a producer path
* Enable the measured producer mode and record hosted comparisons
* ci(release): publish after a skipped orcad template
#24872 skips orcad-template for tags that predate it, but a skipped ancestor
skips every job that keeps the implicit success(), so publish-release and the
post-release jobs never ran for v1.4.219.
* test: brace-free filter in the orcad downstream contract
* ci: reuse pnpm verification records in Alpine builders
* ci: qualify consumers of the verification restore action
* ci: match Linux verification cache archive paths
* ci: defer headless dependency installation until graph analysis is needed
* docs: align headless CI rollout with platform and cache policy
* test: isolate headless detector output from the parent CI step
Orca downloads a newer agent-state-rules.json from a fixed GitHub release (stable or next channel), validates it like the bundled rules, and applies it without a restart; a local override wins over the download, which wins over the bundled rules. A hand-started workflow from main is the only publisher; merging publishes nothing.
* ci: bound unit jobs to one hour of execution
* docs: keep CI budget notes clear of the headless follow-up
* docs: keep CI deadline evidence in the pull request
* Let scheduled CI warmers wait and measure WebRTC startup
* Measure a smaller daemon shutdown fixture image
* Counterbalance WebRTC startup and verify retained fixture files
* Record CI fixture measurements and remove temporary pilots
* Clarify fixture build dependency cleanup evidence
* Make coalesced snapshot fixture delivery deterministic
* test: type the PTY write delay observer
* ci: avoid unrelated headless server qualification
* ci: skip headless detection for ineligible draft PRs
* ci: preserve cross-host qualification and skip supplied prerequisites
* ci: include Windows server cache validation in change detection
* Reuse serializer oracle cells and isolate native cache policy
* Preserve native cache post-save paths and record hosted oracle gain
* Record native cache reuse and separate cancel-test startup budget
* ci(cross-version-wire): run the whole directory so no compatibility test is left out
Three cross-version tests ran in no CI job because the job named its files by hand.
Run the directory instead, ratchet that every file kept out of the unit shards
runs in some PR job, and re-run the job when the modules the newly running
tests guard change.
* test(cross-version): give the orchestration downgrade test its siblings' 120 s budget
* ci(unit-exclusion): count only merge-gating jobs, and require each excluded file's job to fire on it
The coverage check counted any pr.yml job, including e2e, terminal IME and Windows WSL, which are
left out of verify.needs and so cannot block a merge. It now reads verify.needs and the reusable
workflows those jobs call.
It also only proved that some step names each excluded file, not that the job runs when the file
changes. The structured-session zsh login-shell test runs only in shell_contracts, whose path
trigger matched neither it, its harness nor its subject, so a PR touching only those ran it
nowhere. The check now asserts a change to each excluded file fires a gating job that names it,
and the shell trigger gains those three paths.
* ci(cross-version-wire): trigger on the turn-outcome vocabulary and the schema version-skew resolver
A change confined to src/shared/agent-turn-outcome (the arms a newer host publishes) or to
orchestration-schema-version-skew (how current code reopens a downgraded database) skipped the
job whose tests guard exactly those contracts. Also corrects the publish/read direction in the
turn-end comment.
* test(cross-version): state why the orchestration downgrade test needs 120 s
* test(ci): glob the unit tree once for the unit-exclusion coverage checks
* test(native-chat): a Stop over a card sent now into the running turn keeps it paused
Red on main: Codex's turn end withdraws the steered hand-off and the queue
sends the card again as a new host turn, with no pause recorded.
* fix(native-chat): a Stop's queue pause holds a card whose hand-off is still unanswered
A card sent now into the running turn was still pending when Stop judged the
pause, so nothing was recorded; the interrupt then withdrew the hand-off, the
card went back to waiting unpaused, and the queue sent it again as a new turn.
The pause now counts a hand-off that may still return to waiting, judged with
the appended row applied, so a withdrawal lands under the pause and an
acceptance retires it in that same write. Codex and Claude both hit it.
* test(native-chat): the Claude re-send case fails on its diff, inside the test's budget
* test(native-chat): a pause held by an unanswered hand-off ends on every path that ends it
The provider's answer, the provider dying, the chat closing, a restart, a
withdrawal still owed at open, and a /clear (refused until the hand-off ends,
then carrying every waiting card paused 'cleared'); each ends with the queue
sending again.
* fix(native-chat): narrow the pause's settled hand-off, and assert the queued receipt's card
* refactor(native-chat): derive the queue's pause from Stop and Resume journal rows
Stop now appends one journal row where it takes effect, before the interrupt,
whatever the queue holds; Resume appends its own. The pause is a pure function
of the fold: the latest Stop with no later Resume and no later accepted turn a
person asked for. A /clear's carried cards name their source, which is the
replacement's 'cleared' pause. Host-origin turns never lift either.
One predicate decides which cards a pause holds; by default every waiting card
without a hold of its own, including one queued after the Stop. The drain's
consume re-judges it inside its own transaction.
The rows are tombstones of an id no item takes, carrying the mark: a released
host reads an unknown row kind as corruption and truncates the journal there.
Deletes the stored pause (recordPause, the retire hook on every appended row,
the settle-before-record step, mayReturnToWaiting and its row overlay) and the
tests that only proved it retires. The queued_message_pauses table stays in the
schema, unread and unwritten, for downgrade safety.
* fix(native-chat): a card queued after a Stop sends normally, never ahead of held ones
A Stop's pause now holds only the cards queued before its row, plus a steer it
withdrew, which returns to its own place. Each card records the journal
position it was queued at, and the one hold rule compares that with the Stop
row. A card queued after the Stop is a new instruction: it sends as usual, but
the drain still stops at the first held card, so it never overtakes them.
/clear's pause holds the cards it carried. Holding every card again is a
one-line switch in that rule.
* fix(native-chat): the queue's own send re-checks the no-overtake rule in its transaction
The drain's pick and its consume now read one function, nextSendableQueuedCard,
so a Stop row that lands between them holds a newer card behind an older held
one exactly as the pick would. Notes why Stop and Resume ride a tombstone row.
* fix(native-chat): stop creating the unused queue pause table
The queue's pause is derived from journal rows, so nothing reads or writes
queued_message_pauses. It was still created on every open "for downgrade
safety", but an older build creates it itself when it opens the database, so
the table only sat empty in every new database. The tests now pin that no
pause table exists.
* fix(native-chat): a Stop's pause never hides the restart pause
A Stop holds only the cards queued before it. The pause derivation still
returned the Stop alone whenever it was in force, so the restart pause was
never considered: a card queued after the Stop, written by a host process
that has since exited, sent by itself after Orca restarted, with no pause
header and no Resume. A /clear pause that held nothing could hide it the
same way.
Every pause in force is now derived. A card is held if any of them holds
it, and it names the first that does. The drain's pick, the consume
transaction's re-check and the published header all read that one rule;
the header names the pause holding the first card Resume would send.
* test(native-chat): pin the Stop's no-resend, lift and held-card rules
- The Claude and Codex Stop-withdraws-a-steer tests checked "not sent
again" at one instant, before a queue ignoring the pause re-sends. They
now wait for the stopped turn to end and re-check after a quiet window.
- The deleted-card test read a card queued after the Stop, which sends
whether or not a person's turn lifts it; it now reads the Stop's pause
before and after that turn.
- Unit cases pin that a Stop holds a card with no recorded position and one
queued before a rewind.
* refactor(native-chat): a Stop writes one Stop event with its reason, turn and caller
The Stop row that paused the queue becomes the general Stop event
{ reason, turnId?, at, caller? }, whose reason is the host's existing stop
cause. It still rides a tombstone of a host-only id (a released host deletes
the journal from the first unknown row kind), and Resume keeps its own marker
on its own id. Only a person's Stop (reason user-stop) pauses the queue.
* test(native-chat): a rewind keeps a lifted /clear pause lifted and restates the same Stop event
* test(native-chat): pin that Stop and Resume rows never reach apps or count as history
* test(native-chat): only a person's Stop event pauses the queue
* test(native-chat): pin that a Stop's event precedes the interrupt and the at-start stop
Through the real host: the event names the turn and who asked and is in the
journal when the interrupt reaches the agent; at an agent still starting it is
there before the start is ended and holds a card queued before it; an idle Stop
writes one only when it withdrew a send; and the queue's claim re-judges a
pause that landed after its pick.
* test(native-chat): a card held at a starting agent is checked before the Stop's timing
Also says precisely what the claim's in-transaction pause check defends
against: the Stop and the drain share one serialized lane.
* test(native-chat): a released build keeps and folds a journal holding Stop events
Replays this build's rows from the released build's own journal database: every
row is kept, the history after the Stop still folds, and an older client is sent
only removed ids no item uses.
* style(native-chat): format the Stop event changes
* test(native-chat): type the released build's exports through one checked helper
* fix(native-chat): the Stop/Resume row guard narrows to those tombstones only
* test(native-chat): run the Stop-event downgrade test in CI, and cover a writable downgrade
The Stop-event downgrade test ran in no CI lane: unit shards exclude the
cross-version folder, and the cross-version lane runs a fixed file list that
did not name it. It is now on that list.
Its only case replayed the rows into a release's own fresh database, because
that release cannot open the current host database. A second case opens the
journal this build wrote with a main build that shares the database: it opens
writable, keeps every row, appends, and this build then reopens it with the
person's Stop still pausing the queue.
* fix(native-chat): a Stop that stops nothing new writes no Stop event
A Stop reaching a running agent wrote a Stop event on every press. Two
presses before the first interrupt landed wrote two events, so a card
queued between them counted as before the latest Stop and was held,
though a card queued after a Stop should send normally. A Stop naming a
turn that had already ended, as a phone sends late, also wrote an event
for a turn it never stopped.
It now writes one only when it withdrew a queued send, or stops something
no event records yet: not a turn the journal no longer runs, and not the
live turn a Stop still in force already names, unless a card was handed
over into it since, which this Stop's interrupt sends back and must hold.
The interrupt and the "already finished" note are unchanged. A Stop at a
starting agent still always writes.
* test(native-chat): pin that a later host, eviction or close Stop never lifts a person's Stop
* chore(native-chat): put each Stop-row doc on its own declaration, and say only user-stop is journaled
* fix(native-chat): any later Stop event ends a person's Stop pause
A person's Stop paused the queue until their next accepted turn or Resume,
and a later Stop of another reason (the host stopping the agent, an
eviction, a close) was ignored. Now the pause is the latest Stop event's:
a later Stop of any reason ends a person's pause, and only a person's Stop
pauses. The fold keeps the latest Stop event whatever its reason.
An eviction of a resting chat writes no Stop event (a Stop that stops
nothing writes nothing), so it cannot release held cards; a test pins that
no event means no lift.
* fix(native-chat): a second Stop press is a repeat even when the first came before the turn showed
A Stop pressed before the agent's turn shows in the journal (before
Claude's echo, or before Codex opens the turn) records no turn. A second
press once the turn showed compared that missing turn with the live one,
wrote a second Stop event, and held a card queued between the presses.
A repeat is now judged by what was sent since the Stop in force: with
nothing sent after it (a refused send aside), a Stop that named no turn,
or named the live one, is repeated and writes nothing. Anything sent since
and not refused, including a send whose fate is unknown, makes the new
press write, since its interrupt may send that card back to waiting.
Tests: the two-press case across the turn showing; a steer between the
presses settled unknown; and a Stop naming a turn that ended while the next
card is sent but shows no turn yet, which writes and holds that card. The
fold test that claimed an eviction path is renamed.
* fix(native-chat): the queue's pause ignores a Stop or Resume row holding a value no build writes
A Stop or Resume row's value is read from disk with no shape check, and
the pause fold stored whatever it found. A stored `stopEvent: null` would
then throw on every pause check for that chat: the queue's pick, its
send, and every queue update to clients. No build writes such a row, so
this is hardening.
The fold now reads a Stop only when it is an object with a string reason
and a finite time, and a Resume only when it is `true`. Anything else is
ignored: it pauses nothing and ends nothing. The row is still not treated
as malformed, which could cut the history short.
* refactor(native-chat): one reading of a Stop's turn for its event and its note
A Stop's event and its note each worked out the same two facts on their own:
which turn the Stop is about (the one it named, else the one running), and
whether a named turn is the one the journal shows running. The event decides
before the interrupt; the note and whether the session ends decide after the
provider's answer, so those decisions stay separate, but the facts they read
are now one helper each in structured-agent-session-turn-stop-notes.ts:
structuredAgentSessionStoppedTurnId and
structuredAgentSessionStopNamesTurnNotLive. The event's turn, the note's key,
the session-ending condition, the running-command check and the repeat check
all read them. No behavior change.
Tests: a Stop naming no turn records the running turn on its event, and
rewrites that turn's note as a Stop naming it does.
* refactor(native-chat): a failed-interrupt Stop reads its turn through the shared helper
The new branch that ends a Codex child after a failed interrupt asked
whether the Stop's turn still runs with `turnId ?? liveTurnId`, a third
copy of "the turn a Stop is about". It now reads
structuredAgentSessionStoppedTurnId, the value the note key already uses,
read at the same point before the cancel. No behavior change.
Test: a Codex Stop whose interrupt failed ends the child, holds the card
queued before it with the queue paused, and writes its Stop event before
the turn's end.
* fix(terminal): give plain fish tabs Orca's codex function without changing fish's startup
A `codex` typed into a plain fish tab ran without --no-daemon because only
wrapped fish tabs (startup command / ready marker) got Orca's codex function.
Plain fish spawns now prepend an Orca data dir to XDG_DATA_DIRS and record the
exact prefix in ORCA_FISH_XDG_DATA_DIRS_PREFIX. Fish sources the dir's
fish/vendor_conf.d snippet, which first restores XDG_DATA_DIRS (unset again if it
was unset), erases the marker, drops its dir from fish's derived vendor/function/
completion paths, then defines the shared fish codex function at the first prompt
so the user's config.fish still wins. fish argv is unchanged; wrapped tabs keep
their existing -C path. A local fallback to another shell restores the user's
XDG_DATA_DIRS instead of deleting it.
Bumps the terminal daemon protocol to v39 so new tabs move to a daemon that
injects the env; v38 owners stay attachable.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(fish): skip the XDG handoff for -N/--no-config and empty XDG_DATA_DIRS
fish never reads vendor_conf.d under -N/--no-config (also abbreviated or
clustered), so the snippet could not undo the prefix; and the restore cannot
tell an empty XDG_DATA_DIRS from an unset one. Both now launch untouched.
Run the real-fish handoff tests in the shell contracts job, where fish is
required, so they no longer skip in CI.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test(fish): compare the unset-restore case against a fish without Orca
Ubuntu runners ship snapd's fish vendor snippet, which sets XDG_DATA_DIRS on
every fish start, so "unset" was never the right oracle there.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(fish): treat an empty XDG_DATA_DIRS like unset so the tab still gets the codex hook
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(pty): put back the user's own launch env on a shell fallback
The primary shell's launch config now records the pre-launch value of each
key it writes. A fallback shell restores those values (unsetting keys that
had none) instead of deleting the keys, which hands back an inherited
XDG_DATA_DIRS after a fish fallback and an inherited ZDOTDIR after a
zsh->bash fallback, with no per-shell special case.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* refactor(fish): drop the Node restore twin and simplify the vendor snippet
- Remove restoreFishXdgDataDirs; the generic fallback restore covers it.
- Snippet: read ":$XDG_DATA_DIRS:" directly and filter Orca's vendor dirs
with one string match per variable.
- Require inheritedXdgDataDirs in both getShellLaunchConfig option shapes.
- Drop the test-only FISH_XDG_DATA_DIRS_HANDOFF_DAEMON_PROTOCOL_VERSION.
- Fix stale fish comments and trim redundant -N launch cases.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* refactor(fish): stop scrubbing fish's lookup paths after the handoff
Only XDG_DATA_DIRS is restored, by exact prefix; Orca's dir holds nothing but this snippet, so leaving it on fish's derived paths loads nothing else and drops the glob match.
* docs(fish): drop the comment for the removed vendor-dir cleanup
---------
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* refactor(relay): sample fleet health inside the same-cap roll instead of a separate monitor run
A same-cap wave no longer consumes a 15-minute monitor dry-run and its sealed,
single-use, five-minute-fresh evidence. Each apply wave now samples fleet health
itself right before isolation, with the monitor's evaluator, thresholds, and
tolerances, for a window sized to the cell's host count (3/5/8 min), plus three
lookback rules: no cell container exit in 10 min, no minute over 500 director
503s in 10 min, and director concurrency p99 within the monitor bar over 4 min.
Removes the monitor-run inputs, the gate's consume/authorize steps, the
break-glass override, and the same-cap-only authorization shapes in
relay-monitor-evidence.mjs. The monitor workflow and the rehome enable path are
unchanged.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): bound the pre-drain sample overrun and keep the drain token fresh
Review follow-ups: alternating tolerated readings could hold the sample open
until its step timeout, so cap the overrun at three samples past the window;
record why a read failed; mint a fresh admin ID token for the drain after the
sample; raise the job timeout to 90 min so a long sample cannot cancel the
job past the failsafe.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* feat(relay): exempt the rolled cell and existing-only cells from the pre-drain crash rule
The exit rule counted every relay container exit fleet-wide, so a cell that
crashes every few hours (c25, 12 a week) blocked the very roll that fixes it,
and existing-only legacy cells (c5, 15 a week) blocked rolls they take no part
in. Exits are now grouped by instance, each instance is named by its own newest
runtime-metrics log line, and only exits on general or migration-only cells
other than the target count. An exit no configured cell can be named for trips
the rule; a failed lookup is a failed read. relay-observability.tf joins the
evidence-code set because the rule depends on its filter.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* test(relay): cover re-asking for an unnamed exiting instance; note the boot-exit risk
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* feat(relay): declare US cells c32 and c33 at the 3,000-host shape
Declares two us-central1 cells at the Asia shape (cap 3000, 6000 request
units, e2-standard-4) with the US default pool of 10, and generalises the
Asia topology and admission ladder to derive each wave's region from its
reviewed zone, leaving every Asia wave's behaviour unchanged.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* docs(relay): note the US canary tie-break and leave the fleet pool list to promotion
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): plan C32 and C33 as one topology wave
The live-image overlay refuses a declared non-target cell with no template,
so a lone C32 plan would fail on C33. Registration and promotion stay
one cell at a time.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* Reduce repeated PR setup and transcript timing waits; add hosted comparisons
* Align parallelism contract with Node-only external rebuild toolchain
* Record hosted coverage and launch package, store, and cancellation comparisons
* Apply hosted Windows setup savings and remove measured test waits
* Keep measured PR package gains and remove completed comparison jobs
* Report measured test counts with precise units
* test(orcad): skip the Bun-to-Node live-terminal hand-over across a protocol bump
The last Bun orcad's daemon reports protocol 38 forever, so asserting the
adopted daemon matches this checkout's PROTOCOL_VERSION failed every bump.
Ask the Bun slot's daemon for its protocol once, run the hand-over when it
matches, and skip with the two versions named when it does not: a daemon
at another protocol is never adopted across an update.
* test(orcad): clean up the Bun protocol probe even when its launch fails
The probe's cleanup ran only after a successful launch, so a launch that
timed out or threw left its orcad and daemon running. One finally now
stops the orcad, kills what it launched, and kills any daemon named by a
pid file in the probe's data root.
* fix(ssh): launch the Windows relay outside sshd's job without WMI
Win32-OpenSSH kills a session's job on close but allows breakaway. relay.js
gains a one-shot launcher mode that starts the detached relay with
CREATE_BREAKAWAY_FROM_JOB through the staged process-tree addon, so a standard
user no longer needs a WMI Remote Enable grant. WMI stays as the fallback for a
relay without the addon, and a refusal there is named. The Windows SSH-host
lanes drop their WMI grant and assert the breakaway route and adoption.
* fix(ssh): find runtime holds without WMI on a standard-user Windows host
The store GC read held runtimes through Get-CimInstance Win32_Process, which
WMI refuses to a standard user's SSH logon, so the pass kept every runtime.
On a refusal it now reads this account's own process image paths through
Get-Process.
* build(relay): ship the Windows relay launcher addon in every desktop package
macOS and Linux packages carried Windows relays without windows-process-tree.node,
so a legacy-runtime relay they uploaded to a Windows SSH host could not launch
outside sshd's job and fell back to WMI, which a standard user is refused.
A reusable Windows job now compiles the x64 and arm64 addons once and uploads
them; release-cut, release-mac-build, and the hourly/daily/adhoc mac builds
download them before build:release and require both arches. Staging now rejects
a binary with the wrong PE machine, the ReadProcessMemory import, or no
spawnOutsideJob export, so a stale pre-launcher build cannot ship.
* ci(ssh): run the Windows SSH-host lanes when the relay process-tree build scripts change
The staging and gyp-rebuild scripts decide which windows-process-tree addon the
relay ships, so a change to either must re-prove the Windows host cells.
* test(ci): find the mac orcad-template download by artifact name
The release mac job now also downloads the relay Windows process-tree addons, so
the first download-artifact step is no longer the template's.
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
* feat(ssh): wire rung B to the glibc 2.17 compat runtime; gate rung C vault on full node:sqlite
- COMPAT_RELAY_RUNTIMES lists linux-x64-glibc217; rung B plans the compat slot and compat
pinned Node when glibc is below 2.28 or rung A refused with libc_floor/missing_lib.
- The relay version folds the compat runtime's executable hash; refusals are cached per runtime.
- The orcad template stages an optional linux-x64-glibc217 target (base package + compat
node-pty slot + compat runtime marker); the verifier and materializer accept it.
- node-pty slot loader falls back to the compat slot when the default slot is missing or
needs a newer glibc.
- Runtime store GC keeps the compat pin beside the default one on every relay connect.
- hasNodeSqliteReaderApi (DatabaseSync + backup) gates relay session search and the relay
OpenCode reader, which now names the host Node version in its unavailable reason; the SSH
vault reader installs the compat Node on old-glibc hosts and uploads nothing when no
pinned Node can run.
- Rung D: a remembered noexec reports home_noexec and never advises installing Node.
* fix(ssh): re-prove a replayed noexec after rung D so allowing exec recovers the host
* fix(ssh): keep the rung B compat runtime pinned in the relay-connect store GC
* test(ssh): mock deployment-target facts in the Windows OpenCode runtime tests
* ci(ssh): build the glibc 2.17 compat slot for the hostile-host matrix; CentOS 7 lands on rung B
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>
* fix(relay): accept MIG version-name reconciliation and recreate stranded cells without rewriting the MIG
The stranded-rollback recovery ran a gcloud rolling action, which renames the
MIG version outside Terraform. Every later plan for that cell then reverted the
label, and the capacity-plan validator refused the revert as an unreviewed MIG
change, so the cell could be neither rolled nor rolled back.
The validator now accepts a MIG field moving back to what relay-gce-cells.tf
declares (version name and update policy), in every mode, and a test pins those
values to the Terraform file. The stranded branch recreates the cell's single
instance with recreate-instances, which leaves the MIG untouched.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): let a label-only MIG plan through and recreate on it in a stranded rollback
A stranded rollback whose template is already in place plans only the version
name revert. The validator still required the MIG template to move, so that
plan was refused, and the recreate gate (changes == 0) would have skipped a
plan of one change and left the drain flag set. Require the template move only
when no declared field reconciles, and recreate whenever the template was not
replaced (changes < 2).
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* test(ssh): upload a root reached through a symlinked parent
The upload-root realpath fix landed with #24180; this keeps macoshost's case
where the root is passed explicitly beneath a symlinked parent.
* ci(ssh): macOS hostile-host lane on a loopback user-level sshd
Adds local-sshd cells for darwin-arm64 (macos-14) and darwin-x64
(macos-15-intel): a non-root sshd on 127.0.0.1 logs in as the runner user
with SetEnv PATH=<shims>:/usr/bin:/bin:/usr/sbin:/sbin and an empty HOME, so
no rc file restores Homebrew. The driver asserts rung A, terminal echo,
cached runtime reuse, GC keeping the in-use runtime, no toolchain or xattr
calls, and that the SFTP-uploaded Node carries no quarantine and runs as
uploaded. Docker cells are unchanged; each machine runs only cells it can host.
* test(ssh): fail a hostile-host run that would skip every named or hostable cell
A cell named for the wrong OS or arch was silently skipped, so a macOS job on a
mismatched runner went green having deployed nothing. Named cells must now be
hostable here, and a gated run must select at least one cell.
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
* ci(ssh): import the private Windows OpenSSH provisioning harness
Copied unchanged from origin/OrcaWin/np-windows-ssh-provider-diagnostic
(config/ci/windows-ssh-provider/preview-ssh/ at 1242f3c4c8, commits 78b3857a0d,
c24adccff0, 069aa7b38b): a private LocalSystem sshd service on 127.0.0.1 for a
dedicated standard user, either the Microsoft-signed Win32-OpenSSH
10.0.0.0p2-Preview ZIP (archive and every binary pinned by sha256) or the inbox
OpenSSH.Server capability binaries. The following commits extend it for the
pinned-Node relay host lanes.
* test(ssh): run hostile-host cells through a host-agnostic driver
The Docker matrix drove the relay deploy and inspected the container with
inline docker exec calls, so no other host could reuse it. Split it into:
- ssh-hostile-host-test-harness.ts: the deploy, ladder observation, terminal
echo (per-shell probe), runtime reuse and GC-keeps-in-use assertions, now
also capturing every command the deploy sent the host.
- ssh-hostile-host-observer.ts: how a driver inspects the host outside SSH;
docker exec for containers, the local filesystem for a loopback host.
- a legacy_opt_out outcome: the ladder never runs and nothing enters the
pinned store, whatever the host-Node path does.
Launched cells now also check the runtime's sha256 on the host and that every
slot file (the Windows bundled ConPTY pair included) landed in the relay dir.
* ci(ssh): Windows SSH-host lanes for the pinned-Node relay
Phase 2 exit gate, Windows half: the real deployAndLaunchRelay through a real
SshConnection against Win32-OpenSSH on 127.0.0.1, on windows-2022 (x64) and
windows-11-arm (arm64), for both the inbox OpenSSH.Server capability and the
Microsoft-signed 10.0.0.0p2-Preview release (ZIP; archive and each binary
pinned by sha256 and Authenticode, as in the imported harness).
Builds on the provisioning harness from
origin/OrcaWin/np-windows-ssh-provider-diagnostic (previous import commit):
- one private standard account per cell, so every cell starts from an empty
runtime store;
- -HiddenTools: the private sshd service's own Environment carries a PATH
without any machine PATH entry holding node/npm/compilers, led by logging
.cmd shims; a session probe fails the job if node.exe still resolves;
- DefaultShell set per cell by invoke-pinned-relay-cells.ps1 and restored at
cleanup (dispatch proven per cell via %COMSPEC%).
Cells (src/main/ssh/ssh-windows-host-cells.ts): pinned-cmd (stock sshd),
pinned-powershell (DefaultShell = Windows PowerShell) expect rung A on the
pinned node.exe with the relay self-test passing, terminal echo, runtime
reuse, GC keeping the in-use runtime, stage identity through node.exe and no
Add-Type in any decoded session command; legacy-opt-out expects the ladder
never to run and an untouched pinned store.
* fix(ssh-ci): tolerate absent-drive PATH entries and retry Windows userData teardown
Join-Path throws on a machine PATH entry naming a drive the runner lacks,
which would abort provisioning before any cell ran; the toolchain split now
probes with [IO.File]::Exists over [IO.Path]::Combine, and the self-test
covers an absent drive. The hostile-host harness removes its throwaway
userData with removeTreeSync so a transient Windows lock cannot fail the
lane's afterAll.
* fix(ssh-ci): stop the account list rebinding the typed -Accounts param
PowerShell variable names are case-insensitive, so $accounts=[List[hashtable]] assigned into
the [int]$Accounts parameter and every Windows host job died before provisioning. Rename the
list and make the provisioning self-test reject script-scope assignments that shadow a param.
* fix(ssh-ci): hide the host toolchain by ACL, since sessions ignore the service PATH
Win32-OpenSSH builds a session's PATH from the machine and user registry values, so the private
service's Environment never reached SSH sessions and host node.exe stayed visible. Deny the private
accounts the toolchain PATH directories, put the logging shims on each account's own PATH, and
record failing sshd and client log lines so a refused login is diagnosable from the receipt.
* fix(ssh): resolve the upload root before checking entries stay inside it
uploadDirectory compared each entry's realpath against the root as given, so a root reached
through a symlink, junction or Windows 8.3 short name (C:\Users\RUNNER~1 in TEMP) rejected every
entry as escaped and the pinned runtime upload never started.
* fix(ssh-ci): fail cells on a vitest failure and give each account its own keys file
The cells script read $LASTEXITCODE under the workflow's GetNewClosure callback, which sees a
stale captured copy, so failed cells reported exit 0 and the job passed. Read the global value.
Inbox sshd 8.1 checks authorized_keys with read_ok=0, refusing a file other accounts can read;
use one keys file per account via %u.
* fix(ssh-ci): keep the account name in inbox mode and surface the WMI launch gap
The inbox binary-verification loop reused $name, so later SSH and SFTP probes logged in as
'sftp-server.exe'. Before the cells run, probe whether a standard SSH user can call WMI
Win32_Process.Create (the Windows relay launch path); when refused, warn and grant the cell
accounts Remote Enable on root\cimv2 for the run so the remaining assertions execute.
* test(ssh): keep the first terminal session answering keepalives through GC
The hostile-host driver disposed the first session's multiplexer before the GC and reconnect
steps, so a slow Windows GC let the relay reap the silent owner as 'local' and the reconnect then
waited out the full owner grace. Keep the session live until the connection closes, as the app
does, and resend the terminal probe until the shell evaluates it: ConPTY PowerShell can drop
typeahead sent before its first prompt.
* ci(ssh): keep each cell's relay logs in the receipts
* fix(relay): detach an ended socket client as peer-closed before destroying it
The listener destroyed a socket on 'end' but detached its client only on 'close'. A relay write
in that window failed with 'Relay socket is closed', and the dispatcher closed the client as
'local', so its PTY owner kept the full 30s grace instead of the peer-closed floor and a quick
reconnect was refused. The Windows host lanes logged this race on the named-pipe endpoint.
* test(relay): drive the peer-end listener test with a real dispatcher instead of a cast stub
The stub was an unchecked 'as unknown as RelayDispatcher' that failed the changed-code casting gate.
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
* build(orcad): merge per-runner prebuild slot trees into one matrix
Each node-server lane builds only its own node-pty slot. Release CI needs
their union before `build:orcad-prebuilds --require-slots` and the
template build can run; merge-orcad-prebuilds.mjs verifies every lane's
files against its own manifest, refuses duplicate slots and mismatched
node-pty/N-API/Node-header builds, then writes one merged manifest.
* build(orcad): keep agent-browser out of the desktop deployment template
The template rides inside every desktop build (design D2). Seven ~10 MB
agent-browser binaries would be ~76 MB, more than the rest of the template;
design D2's package contents never listed it, and a slot without one
already reports no headless browser. ORCAD_OMIT_AGENT_BROWSER=1 skips the
copy; standalone build:orcad still includes it.
* feat(packaging): ship the orcad deployment template in desktop builds
Design D2: the server JS and every target's addons ship inside the app,
as out/relay does; the ~120 MB Node runtimes stay excluded and are
downloaded on demand. electron-builder copies out/orcad-template to
Resources/orcad-template on every desktop OS, which is the first path
materializeOrcadArtifact tries (process.resourcesPath).
Platform signing rewrites native bytes the template manifest hashes:
- macOS: the tree is signIgnored (codesign rejects its ELF/PE payloads);
afterPack signs the darwin targets' Mach-O files with the app identity,
as notarization requires, then reseals only those manifest entries.
- Windows: SignPath signs after packaging, so release CI reseals from the
inner-signing list (packaged-orcad-template.cjs --reseal-signed).
Every other file must still match the build's hashes; afterPack verifies.
ORCA_REQUIRE_ORCAD_TEMPLATE=1 makes a missing template fail beforePack and
afterPack; without it a build ships none and SSH relays keep the legacy
path. verify-packaged-orcad-template.test.mjs's "unused, excluded"
contract is reversed on purpose.
* ci(release): build the orcad template from qualified lanes and package it
node-server-tests.yml becomes callable with a ref and build_template.
With build_template, each lane that owns a release slot (macOS, Windows,
the glibc 2.28 and Alpine lanes, and the glibc 2.17 compat lane) uploads
its qualified out/orcad-prebuilds, the Windows lane also uploads both
process-table addons, and desktop_template merges them, gates the full
matrix plus the compat slot with --require-slots, runs
build:orcad-template and uploads the orcad-template artifact.
release-cut calls it at the release tag beside the other gates. The
build and build-mac jobs wait for it, download it into out/orcad-template
(the mac workflow from the parent run), and require it via
ORCA_REQUIRE_ORCAD_TEMPLATE. The Windows signing staging skips the
template's Linux/macOS payloads, and a reseal step records SignPath's
bytes before the installer rebuild. A template-scoped concurrency group
keeps a release call and main's push runs from cancelling each other.
* test(orcad): keep the packaged-lookup imports clear of the compat-slot import edits
* ci(orcad): let a rerun lane replace its template artifacts
upload-artifact v4 refuses a second upload under an existing name in the same
run, so rerunning a flaky node-server lane during a release would fail at the
upload instead of re-qualifying the slot.
* ci(node-server): build the template's Windows addons before the lane switches to Node 18
The addon build script imports TypeScript, which Node 18 cannot load, so every
build_template run (release-cut included) failed on windows-2022.
* fix(build): ship the orcad template's shared node_modules
electron-builder's extraResources filter always drops the root node_modules of
a source directory, so packaged apps lost orcad-template/node_modules and the
afterPack verify failed. Copy it through its own resource entry.
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
* fix(relay): anchor same-cap monitor evidence freshness to the run's authorisation, not job startup
The same-cap gate now verifies the dry-run on its own clock and records the
authorisation instant in the single-use consumed marker; each cell job checks
the evidence age at that instant and bounds its own start after it.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(relay): refuse a re-run same-cap gate before it consumes evidence; tighten order tests
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* fix(ssh): collect the pinned-Node runtime store on Windows hosts
Windows SSH hosts now run runtime-store GC instead of skipping it: one
PowerShell inventory reads .runtime-ref-node-<sha> and .runtime-node refs from
every version dir, and one Get-CimInstance Win32_Process query filtered on an
image path under runtimes\ adds process holds (never by image name; a failed
query keeps everything). Stale upload stages are swept with the same rule as
POSIX. Promotion and the post-upload hold check now take the store lock on
Windows too, and the lock's own commands run unwrapped there.
Windows relay version-dir liveness now honours .relay-pid (design D5): a live
PID answers ALIVE before any pipe is touched, a dead one (ESRCH) plus refusing
pipes is exited, anything else is unverifiable. The runtime probe adopts a
pinned node.exe an earlier vault reader left without a .verified marker after
running it.
* fix(ssh): Windows stage fencing and vault runtime go through the verified node.exe
Upload-stage file identity on Windows no longer compiles an Add-Type P/Invoke
helper when the relay runs on Orca's verified pinned node.exe: the stage
commands run a fixed fs.lstatSync(..., {bigint:true}) script through it. It
prints the legacy helper's vol:high:low lowercase hex, and identity files are
compared after normalising hex spelling, so old and new clients recover each
other's stages. Host-Node relays keep the legacy helper; the choice is
documented in windows-edr-posture.md.
The Windows OpenCode vault reader now installs the pinned runtime through
ensureRemoteOrcadNodeRuntime (official zip, host-side extraction, .verified,
store lock) instead of uploading a client-extracted node.exe, and the relay dir
gains a .runtime-ref-node-<sha> so store GC keeps the runtime the vault uses.
* test(ssh): run the Windows stage-identity and store-GC tests on the Windows lane
The legacy/node.exe identity compatibility test and the Win32_Process hold path
were gated to win32 but no CI lane ran them. Add both files to the Windows
package lane and a real running-node.exe hold test.
* test(ssh): tear down Windows-lane temp trees through removeTreeSync
* test(ssh): grant the store lock to the Windows OpenCode runtime setup test
The Windows promote now runs under runtimes/.store-lock, so the mocked host
must answer the lock's CreateNew step.
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>
* fix(ssh): classify a musl host missing libstdc++ as missing_lib, not wrong_libc
musl's loader follows each missing-library line with one 'Error relocating ... symbol
not found' per unresolved symbol, and the relocation pattern was checked first. Check
missing libraries before relocation errors; the ld-linux/ld-musl interpreter case stays
wrong_libc.
* build(orcad): allow a partial deployment template for CI
build-orcad-template --targets a,b builds and verifies only the named slots, so a CI job
that can fill just the x64 Linux prebuild slots can still materialize rung A/C addons.
Without the flag every target is still built and verified.
* ci(ssh): hostile-host matrix for the relay runtime ladder
Drives the real client-side relay deploy against Docker sshd targets and asserts the
design D6 rung each lands on: Debian 10 and AlmaLinux 8 (glibc 2.28) and Alpine (musl)
on rung A; Alpine without libstdc++ refused missing_lib down to D; Ubuntu 22.04 with a
host Node 20 and a noexec home straight to D (home_noexec); CentOS 7 (glibc 2.17)
refused libc_floor at A and C, falling to a host-npm path with no Node; and a
no-egress Debian 10 still on rung A. Launched cells also prove the terminal echoes,
no npm or compiler ran, a second connect reuses the uploaded runtime, and runtime GC
keeps the in-use runtime while collecting an idle one.
New workflow ssh-hostile-hosts.yml runs on dispatch and on path-filtered PRs.
* test(ci): pin the hostile-host workflow to the headless-server builder images
The matrix builds its runtime slots in copies of the node-server lanes' Alpine
and manylinux images; this contract fails when NODE_RUNTIME_PIN or either
builder digest moves in one workflow and not the other.
* test(ssh): reconnect as the same client and retry a grace-held PTY owner in the hostile-host matrix
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
The parent workflow collapses canary-apply and batch-apply into the job mode
apply, so the headroom step's canary-apply/batch-apply condition never held and
the gate was skipped on every real roll. Run it wherever the drain runs (apply,
rollback before its restart) and in read-only verify; skip only a resumed
rollback, which drains nothing. A new workflow-shape test fails on any job
step comparing against a mode the parent cannot pass, and on a drain that can
run without the headroom check.
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
* build(orcad): build server glibc slots on glibc 2.28 and add the glibc 2.17 compat slot
Design D6: the default linux-{x64,arm64}-glibc node-pty slots now build in
manylinux_2_28 (digest-pinned) and are gated at glibc 2.28 / GLIBCXX_3.4.25
through a floor profile on verify-linux-glibc-floor.cjs; the desktop keeps
its Ubuntu 20.04 (2.31) default.
Adds the opt-in linux-x64-glibc217 compat target: NODE_RUNTIME_COMPAT_ASSETS
pins the unofficial glibc-217 Node (update/check pin scripts cover it, outside
SERVER_TARGETS), and a new CI lane builds the compat slot in manylinux2014
with static libstdc++, gates it at glibc 2.17 with no shared C++ runtime in
DT_NEEDED, and smokes it under the glibc-217 Node.
* refactor(node-runtime-pin): route compat lookups through isCompatServerTarget; keep the glibc doc's slot-name paragraph intact
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
* ci(daemon): gate PRs on daemon protocol crossing from the newest release
Lands daemon-protocol-facts.mjs from the Windows update diagnostic branch with a
stricter parser, and adds check-daemon-protocol-crossing.mjs (rule R1): the working
tree must attach the newest release tag's daemon. Rollback crossing is reported only.
Runs in the cross-version-wire job, which already has full tags; tag selection moves
to config/scripts/stable-release-tags.mjs so both use one rule.
* feat(persistence): run profile backups in the worker whenever its entry is bundled
* refactor(orcad): make profile and native preflight runtime-neutral
The profile preflight parser now takes the expected runtime identity from the
caller (shipped callers pass the pinned Bun identity), and the native
preflight is renamed to orcad-runtime-native-preflight with neutral wording.
* feat(runtime): pin the Node 24.21.0 server runtime with an offline CI check
Add src/shared/node-runtime-pin.ts (NODE_RUNTIME_PIN, SERVER_TARGETS,
NODE_RUNTIME_ASSETS for all 8 server targets plus the headers tarball),
generated by config/scripts/update-node-runtime-pin.mjs from the nodejs.org
and unofficial-builds SHASUMS. check-node-runtime-pin.mjs verifies, with no
network, that the pin tracks the locked Electron, matches engines.node's
major, and covers exactly SERVER_TARGETS; it runs in the static analysis job.
ORCAD_BUN_TARGETS consumers now read SERVER_TARGETS so there is one target
list; orcad's Bun runtime and build output are unchanged.
* test(persistence): skip plain-Node backup selection tests in the Bun profile suite
* fix(runtime): reject a pinned archive that belongs to another target
* ci(daemon): fail PRs that swap a runtime launcher and bump the daemon protocol
D7.1 R3: hosting orcad or the daemon on another runtime is not a protocol change,
so one PR must not do both. The launcher file list lives in the check script; the
allow-runtime-launcher-protocol-bump label overrides it.
* feat(orcad): select pinned-Node slots by a .runtime-node marker
D7.1 R5: a Node slot names its shared runtimes/node-<sha256>/node through
.runtime-node instead of .build-target, so Bun-era clients read it as a legacy
slot rather than exiting 78 on a missing bun-runtime. Nothing builds the marker yet.
* fix(runtime): load the Node pin without the typeless-module warning
check-node-runtime-pin.mjs now requires the pin and takes nodeDistArchiveName from
its own module, so it no longer loads the update script's build graph.
* fix(orcad): resolve Node slots to the design's runtimes/node-<sha>/bin/node layout
* feat(orcad): 8-slot node-pty prebuilds against the pinned Node headers at N-API 8
- build-orcad-prebuilds.mjs adds win32-x64/arm64 (conpty.node, the vendored
conpty.dll/OpenConsole.exe, upstream's N-API conpty_console_list.node), compiles
in a scratch copy against the hash-verified pinned headers (node.lib pinned per
Windows arch) with NAPI_VERSION=8, rejects post-8 node_api_* imports, and writes
a schema 2 manifest with per-file sha256, N-API level and the glibc need.
- --require-slots [slots] verifies files against hashes; --smoke loads the slot
under the pinned Node and spawns a PTY; --print-slot names the host slot.
- The slot installer gates on N-API, libc, arch, glibc and file hashes instead of
the exact NODE_MODULE_VERSION, and installs nested files (conpty/).
- bun-profile-tests.yml builds, verifies and smokes each runner's slot.
* fix(orcad): scope node-pty's glibc .symver pins to glibc on musl prebuild slots
musl's unversioned libc cannot satisfy openpty@GLIBC_* references at link
time, so the Alpine slot compile would fail. Pin the staged pty.cc guard to
__GLIBC__ and assert both musl transforms against the installed patch.
* feat(orcad): run orcad on the pinned Node instead of Bun
A packaged orcad slot now references the pinned Node 24.21.0 by its
executableSha256 (`.runtime-node`, `.server-target`) instead of carrying
bun-runtime, and ships node-pty from the slot's prebuild, only its own
ripgrep, and no Windows Bun PTY gate. The runtime lives beside the slots
at runtimes/node-<sha>/bin/node (node.exe on Windows, upstream name).
- build:orcad (build-orcad-node.mjs) builds the host slot's prebuild when
missing and places the pinned runtime; the template is schema 3 with
per-target files.
- handoffToBundledOrcad() resolves the slot's runtime reference and checks
process.versions.node against the pin; a host Node >= 18 still hands off.
Startup preflight keys on running as that runtime; callers expect 'node'.
- orcad and its daemon use node-pty (ConPTY + windows-pty-job on Windows);
the Bun PTY sources, gate entry and canUseBunPty branches are removed.
- SSH deploy uploads the official archive once per pin, extracts and
hash-checks it on the host, and self-tests it before publishing. Bun
slots stay launchable for rollback; Node slots never use host Node.
- The runtime materializer is generic over pinned assets; the Bun wrapper
remains only for the OpenCode vault reader (design Phase 2).
- Cross-runtime test: a profile DB written by Bun 1.4.2 (WAL left by
SIGKILL) opens and backs up under the pinned Node, and the reverse.
No daemon PROTOCOL_VERSION change (design D7.1 R3).
* docs(ci): name the headless lanes after the pinned Node, drop Bun shard timings
Design D10: ci-demand-rollout.md and ci-runner-efficiency.md follow the
bun-profile-tests.yml -> node-server-tests.yml rename; shard timings drop the
deleted Bun PTY tests and follow the renamed ones.
* chore(ci): count the runtime archive download as a runtime launcher path
* fix(orcad): pin the macOS C++ standard for node-pty prebuilds
The official Node headers' config.gypi sets clang: 0, so common.gypi skips its
gnu++20 xcode_settings and Apple clang 15 (macos-14 runners) compiles
node-addon-api as C++98.
* fix(orcad): resolve the preflight's slot through realpath, as the handoff does
A symlinked orcad.js handed off to its real slot's pinned Node, but the
startup and profile preflights read the symlink's directory, found no
runtime marker there, and silently skipped the readiness check.
* refactor(ssh): drop materializeCachedNodeRuntime, which nothing calls
Deploys upload the verified official archive (design D5); no client path
needs an extracted Node executable cached by digest.
* test(orcad): gate the Bun-to-Node upgrade and Node-to-Bun rollback with live terminals
Design D7.1 R1/R3/R4 and D7.2. The last Bun orcad and this checkout's Node
slot are installed side by side under ~/.orca-remote, launched and stopped
with the client's own deploy commands, and share one data root. Each
direction proves the incoming orcad adopts the outgoing runtime's daemon
(same PID, same shell, output continues), opens its profile database and
backs it up with its own shipped worker, and that GC keeps the slot the
live daemon was forked from.
The node-server Linux lanes provide Bun 1.4.2 and build that Bun orcad
from main, and run with --cross-runtime. --artifact and --cross-runtime
now make their tests fail on a missing input instead of skipping.
* ci(node-server): pin node:24.21.0-alpine by its multi-arch index digest
* test(ssh): name the runtime archive fixture after its role
* test(node-server): load node-pty from the packaged slot in artifact runs
The node-server lane installs dependencies without building node-pty, and
Linux has no upstream prebuild, so the real-PTY failed-I/O teardown test
(picked up by the pty-subprocess selector) could not load pty.node. In
--artifact runs, alias node-pty to out/orcad's shipped slot so the test
exercises the addon orcad actually runs under the pinned Node.
* fix(orcad): let the Windows profile preflight exit after its PTY probe
On Windows, node-pty keeps the conout worker thread and pseudoconsole alive
until kill(), even after the shell exits. The PTY health probe never killed a
cleanly exited probe, so the packaged preflight printed its readiness line
and then hung until the build's 30s timeout, reported with an empty stderr.
- The probe kills its PTY on Windows after exit and uses the bundled ConPTY
the daemon spawns with.
- The preflight exits once stdout is flushed; its owner reads to EOF.
- Preflight failures now report code, signal, timeout, stdout and stderr.
* test(node-server): load the slot's node-pty in the real-PTY test, not by alias
A vite alias redirected only ESM imports of node-pty; windows-pty-job and
local-pty-utils resolve it through require, so Windows loaded two conpty.node
copies and the Git Bash job-membership proof read an empty job. The failed-I/O
teardown test now loads node-pty through a fixture that picks the packaged slot
in artifact lanes.
The pty-subprocess selector was a prefix that also pulled in its POSIX-host
sibling unit tests, which pr.yml runs and which were never qualified on
Windows. Select the directory plus the two sibling files that belong here.
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
* ci(daemon): gate PRs on daemon protocol crossing from the newest release
Lands daemon-protocol-facts.mjs from the Windows update diagnostic branch with a
stricter parser, and adds check-daemon-protocol-crossing.mjs (rule R1): the working
tree must attach the newest release tag's daemon. Rollback crossing is reported only.
Runs in the cross-version-wire job, which already has full tags; tag selection moves
to config/scripts/stable-release-tags.mjs so both use one rule.
* feat(persistence): run profile backups in the worker whenever its entry is bundled
* refactor(orcad): make profile and native preflight runtime-neutral
The profile preflight parser now takes the expected runtime identity from the
caller (shipped callers pass the pinned Bun identity), and the native
preflight is renamed to orcad-runtime-native-preflight with neutral wording.
* feat(runtime): pin the Node 24.21.0 server runtime with an offline CI check
Add src/shared/node-runtime-pin.ts (NODE_RUNTIME_PIN, SERVER_TARGETS,
NODE_RUNTIME_ASSETS for all 8 server targets plus the headers tarball),
generated by config/scripts/update-node-runtime-pin.mjs from the nodejs.org
and unofficial-builds SHASUMS. check-node-runtime-pin.mjs verifies, with no
network, that the pin tracks the locked Electron, matches engines.node's
major, and covers exactly SERVER_TARGETS; it runs in the static analysis job.
ORCAD_BUN_TARGETS consumers now read SERVER_TARGETS so there is one target
list; orcad's Bun runtime and build output are unchanged.
* test(persistence): skip plain-Node backup selection tests in the Bun profile suite
* fix(runtime): reject a pinned archive that belongs to another target
* ci(daemon): fail PRs that swap a runtime launcher and bump the daemon protocol
D7.1 R3: hosting orcad or the daemon on another runtime is not a protocol change,
so one PR must not do both. The launcher file list lives in the check script; the
allow-runtime-launcher-protocol-bump label overrides it.
* feat(orcad): select pinned-Node slots by a .runtime-node marker
D7.1 R5: a Node slot names its shared runtimes/node-<sha256>/node through
.runtime-node instead of .build-target, so Bun-era clients read it as a legacy
slot rather than exiting 78 on a missing bun-runtime. Nothing builds the marker yet.
* fix(runtime): load the Node pin without the typeless-module warning
check-node-runtime-pin.mjs now requires the pin and takes nodeDistArchiveName from
its own module, so it no longer loads the update script's build graph.
* fix(orcad): resolve Node slots to the design's runtimes/node-<sha>/bin/node layout
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>