mirror of
https://github.com/stablyai/orca.git
synced 2026-10-07 16:02:29 +00:00
eb44511992a11ef1ed787173a0c19df40e27a97f
562
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
328caa2160 |
fix(git): reduce queries and preserve data across execution hosts (#24602)
* fix(git): reduce queries and preserve data across execution hosts * fix(ci): exercise pinned Git and serialize mobile dependency entrypoints * fix(relay): preserve fresh diff retries after hung shared reads * test(git): wait for fetch barrier before canceling preparation * fix(i18n): describe index-preserving discard in every locale * fix(git): retain clone diagnostics and allow WSL policy startup * test(git): refresh default-base and branch-safety fixtures |
||
|
|
75f8f34ce3 | Overlap independent Linux headless runtime builds (#24910) | ||
|
|
1241ce1b49 |
Skip slower root package-store restores in macOS PR jobs (#24908)
* Skip slower root package-store restores in macOS PR jobs * Update the companion cache-policy contract |
||
|
|
a824fb74ab |
Reuse the headless detector compiler without installing full dependencies (#24895)
* Reuse the headless detector compiler without full dependency setup * Keep optional compiler-cache saves from failing cache warming |
||
|
|
58cf72d48e | Skip slower root package-store restores in Linux PR jobs (#24896) | ||
|
|
add1c55590 |
Skip slower Windows root package-store restores in CI (#24885)
* Skip slower Windows root package-store restores in CI * Update reviewed mobile dependency-store cache expression |
||
|
|
75be95fd8c | ci: move ARM Mac qualification to macOS 15 (#24760) | ||
|
|
61836f6026 |
Reduce scheduled CI cache warming to every six hours (#24881)
* Reduce scheduled CI cache warming to every six hours * Document cache warmer recovery interval and measured tradeoff |
||
|
|
d3a406fcbe |
ci(release): publish after a skipped orcad template (#24882)
* ci(release): publish after a skipped orcad template #24872 skips orcad-template for tags that predate it, but a skipped ancestor skips every job that keeps the implicit success(), so publish-release and the post-release jobs never ran for v1.4.219. * test: brace-free filter in the orcad downstream contract |
||
|
|
cc73c8e1a7 | ci: overlap ARM SSH setup and independent observation waits (#24714) | ||
|
|
b94c75c4bd |
Reuse pnpm verification records in Alpine CI (#24817)
* ci: reuse pnpm verification records in Alpine builders * ci: qualify consumers of the verification restore action * ci: match Linux verification cache archive paths |
||
|
|
ac28e8c85e |
Skip dependency installation for known headless build inputs (#24716)
* ci: defer headless dependency installation until graph analysis is needed * docs: align headless CI rollout with platform and cache policy * test: isolate headless detector output from the parent CI step |
||
|
|
d3e592365e |
ci(release): skip the orcad template for tags that predate it (#24872)
A patch cut from a base older than #24155 has no orcad template source, so the template job could never pass and every desktop build waited on it. |
||
|
|
564f4d021a |
feat: live updates for agent state rules (#24387)
Orca downloads a newer agent-state-rules.json from a fixed GitHub release (stable or next channel), validates it like the bundled rules, and applies it without a restart; a local override wins over the download, which wins over the bundled rules. A hand-started workflow from main is the only publisher; merging publishes nothing. |
||
|
|
e2f2b707d2 | ci: skip installed glibc tools in SSH host qualification (#24733) | ||
|
|
76b1a90ff6 |
chore(deps): update reviewed dependencies across Orca (#24561)
* chore(deps): update reviewed desktop dependencies and tooling * chore(deps): update compatible mobile packages and Fastlane * chore(deps): update cloud transports and enforce release age * chore(deps): patch documentation dependencies and record review * chore: remove dependency review reports * test(linear): smoke-load resolved SDK through CommonJS loader * fix(deps): keep native rebuilds from reinstalling addon dependencies * fix(native): invoke installed node-gyp directly for Node rebuilds * test(cloud): exclude observer probes from row-lock timing budget * test(mobile): preserve the CSS writer receiver in viewport spy * test(native): remove obsolete batch-shim fixture exception * Stream native rebuild output through the process wrapper |
||
|
|
ff452661e7 |
Stop stalled unit jobs after an hour (#24583)
* ci: bound unit jobs to one hour of execution * docs: keep CI budget notes clear of the headless follow-up * docs: keep CI deadline evidence in the pull request |
||
|
|
302526411d | ci: share bounded apt setup with the E2E native cache job (#24758) | ||
|
|
13ecf051c3 |
Reuse prepared Windows native builds in SSH CI (#24555)
* ci: reuse qualified Windows server slots for SSH host tests * ci: reuse prepared relay addons after an exact native cache hit |
||
|
|
efbf651c7b |
Reduce CI setup costs and fixture failures (#24537)
* Let scheduled CI warmers wait and measure WebRTC startup * Measure a smaller daemon shutdown fixture image * Counterbalance WebRTC startup and verify retained fixture files * Record CI fixture measurements and remove temporary pilots * Clarify fixture build dependency cleanup evidence * Make coalesced snapshot fixture delivery deterministic * test: type the PTY write delay observer |
||
|
|
6153fbcfe4 |
Reduce redundant headless server CI work (#24527)
* ci: avoid unrelated headless server qualification * ci: skip headless detection for ineligible draft PRs * ci: preserve cross-host qualification and skip supplied prerequisites * ci: include Windows server cache validation in change detection |
||
|
|
5a56636f66 |
Bound E2E package setup and retain cancellation traces (#24617)
* test: align source-control fixtures with current store contracts * Bound E2E package setup and retain cancelled-job traces |
||
|
|
8ff6296bc7 |
Speed up serializer checks and keep native caches stable (#24476)
* Reuse serializer oracle cells and isolate native cache policy * Preserve native cache post-save paths and record hosted oracle gain * Record native cache reuse and separate cancel-test startup budget |
||
|
|
444f1952c7 |
ci: run every cross-version wire test, picked up by folder so new ones can't be skipped (#24499)
* ci(cross-version-wire): run the whole directory so no compatibility test is left out Three cross-version tests ran in no CI job because the job named its files by hand. Run the directory instead, ratchet that every file kept out of the unit shards runs in some PR job, and re-run the job when the modules the newly running tests guard change. * test(cross-version): give the orchestration downgrade test its siblings' 120 s budget * ci(unit-exclusion): count only merge-gating jobs, and require each excluded file's job to fire on it The coverage check counted any pr.yml job, including e2e, terminal IME and Windows WSL, which are left out of verify.needs and so cannot block a merge. It now reads verify.needs and the reusable workflows those jobs call. It also only proved that some step names each excluded file, not that the job runs when the file changes. The structured-session zsh login-shell test runs only in shell_contracts, whose path trigger matched neither it, its harness nor its subject, so a PR touching only those ran it nowhere. The check now asserts a change to each excluded file fires a gating job that names it, and the shell trigger gains those three paths. * ci(cross-version-wire): trigger on the turn-outcome vocabulary and the schema version-skew resolver A change confined to src/shared/agent-turn-outcome (the arms a newer host publishes) or to orchestration-schema-version-skew (how current code reopens a downgraded database) skipped the job whose tests guard exactly those contracts. Also corrects the publish/read direction in the turn-end comment. * test(cross-version): state why the orchestration downgrade test needs 120 s * test(ci): glob the unit tree once for the unit-exclusion coverage checks |
||
|
|
f69052e113 | Reuse qualified Windows server builds and dependency verification records (#24448) | ||
|
|
976dc00337 |
fix(native-chat): Stop's pause is worked out from the chat's history, so a steered message is never re-sent (#24072)
* test(native-chat): a Stop over a card sent now into the running turn keeps it paused
Red on main: Codex's turn end withdraws the steered hand-off and the queue
sends the card again as a new host turn, with no pause recorded.
* fix(native-chat): a Stop's queue pause holds a card whose hand-off is still unanswered
A card sent now into the running turn was still pending when Stop judged the
pause, so nothing was recorded; the interrupt then withdrew the hand-off, the
card went back to waiting unpaused, and the queue sent it again as a new turn.
The pause now counts a hand-off that may still return to waiting, judged with
the appended row applied, so a withdrawal lands under the pause and an
acceptance retires it in that same write. Codex and Claude both hit it.
* test(native-chat): the Claude re-send case fails on its diff, inside the test's budget
* test(native-chat): a pause held by an unanswered hand-off ends on every path that ends it
The provider's answer, the provider dying, the chat closing, a restart, a
withdrawal still owed at open, and a /clear (refused until the hand-off ends,
then carrying every waiting card paused 'cleared'); each ends with the queue
sending again.
* fix(native-chat): narrow the pause's settled hand-off, and assert the queued receipt's card
* refactor(native-chat): derive the queue's pause from Stop and Resume journal rows
Stop now appends one journal row where it takes effect, before the interrupt,
whatever the queue holds; Resume appends its own. The pause is a pure function
of the fold: the latest Stop with no later Resume and no later accepted turn a
person asked for. A /clear's carried cards name their source, which is the
replacement's 'cleared' pause. Host-origin turns never lift either.
One predicate decides which cards a pause holds; by default every waiting card
without a hold of its own, including one queued after the Stop. The drain's
consume re-judges it inside its own transaction.
The rows are tombstones of an id no item takes, carrying the mark: a released
host reads an unknown row kind as corruption and truncates the journal there.
Deletes the stored pause (recordPause, the retire hook on every appended row,
the settle-before-record step, mayReturnToWaiting and its row overlay) and the
tests that only proved it retires. The queued_message_pauses table stays in the
schema, unread and unwritten, for downgrade safety.
* fix(native-chat): a card queued after a Stop sends normally, never ahead of held ones
A Stop's pause now holds only the cards queued before its row, plus a steer it
withdrew, which returns to its own place. Each card records the journal
position it was queued at, and the one hold rule compares that with the Stop
row. A card queued after the Stop is a new instruction: it sends as usual, but
the drain still stops at the first held card, so it never overtakes them.
/clear's pause holds the cards it carried. Holding every card again is a
one-line switch in that rule.
* fix(native-chat): the queue's own send re-checks the no-overtake rule in its transaction
The drain's pick and its consume now read one function, nextSendableQueuedCard,
so a Stop row that lands between them holds a newer card behind an older held
one exactly as the pick would. Notes why Stop and Resume ride a tombstone row.
* fix(native-chat): stop creating the unused queue pause table
The queue's pause is derived from journal rows, so nothing reads or writes
queued_message_pauses. It was still created on every open "for downgrade
safety", but an older build creates it itself when it opens the database, so
the table only sat empty in every new database. The tests now pin that no
pause table exists.
* fix(native-chat): a Stop's pause never hides the restart pause
A Stop holds only the cards queued before it. The pause derivation still
returned the Stop alone whenever it was in force, so the restart pause was
never considered: a card queued after the Stop, written by a host process
that has since exited, sent by itself after Orca restarted, with no pause
header and no Resume. A /clear pause that held nothing could hide it the
same way.
Every pause in force is now derived. A card is held if any of them holds
it, and it names the first that does. The drain's pick, the consume
transaction's re-check and the published header all read that one rule;
the header names the pause holding the first card Resume would send.
* test(native-chat): pin the Stop's no-resend, lift and held-card rules
- The Claude and Codex Stop-withdraws-a-steer tests checked "not sent
again" at one instant, before a queue ignoring the pause re-sends. They
now wait for the stopped turn to end and re-check after a quiet window.
- The deleted-card test read a card queued after the Stop, which sends
whether or not a person's turn lifts it; it now reads the Stop's pause
before and after that turn.
- Unit cases pin that a Stop holds a card with no recorded position and one
queued before a rewind.
* refactor(native-chat): a Stop writes one Stop event with its reason, turn and caller
The Stop row that paused the queue becomes the general Stop event
{ reason, turnId?, at, caller? }, whose reason is the host's existing stop
cause. It still rides a tombstone of a host-only id (a released host deletes
the journal from the first unknown row kind), and Resume keeps its own marker
on its own id. Only a person's Stop (reason user-stop) pauses the queue.
* test(native-chat): a rewind keeps a lifted /clear pause lifted and restates the same Stop event
* test(native-chat): pin that Stop and Resume rows never reach apps or count as history
* test(native-chat): only a person's Stop event pauses the queue
* test(native-chat): pin that a Stop's event precedes the interrupt and the at-start stop
Through the real host: the event names the turn and who asked and is in the
journal when the interrupt reaches the agent; at an agent still starting it is
there before the start is ended and holds a card queued before it; an idle Stop
writes one only when it withdrew a send; and the queue's claim re-judges a
pause that landed after its pick.
* test(native-chat): a card held at a starting agent is checked before the Stop's timing
Also says precisely what the claim's in-transaction pause check defends
against: the Stop and the drain share one serialized lane.
* test(native-chat): a released build keeps and folds a journal holding Stop events
Replays this build's rows from the released build's own journal database: every
row is kept, the history after the Stop still folds, and an older client is sent
only removed ids no item uses.
* style(native-chat): format the Stop event changes
* test(native-chat): type the released build's exports through one checked helper
* fix(native-chat): the Stop/Resume row guard narrows to those tombstones only
* test(native-chat): run the Stop-event downgrade test in CI, and cover a writable downgrade
The Stop-event downgrade test ran in no CI lane: unit shards exclude the
cross-version folder, and the cross-version lane runs a fixed file list that
did not name it. It is now on that list.
Its only case replayed the rows into a release's own fresh database, because
that release cannot open the current host database. A second case opens the
journal this build wrote with a main build that shares the database: it opens
writable, keeps every row, appends, and this build then reopens it with the
person's Stop still pausing the queue.
* fix(native-chat): a Stop that stops nothing new writes no Stop event
A Stop reaching a running agent wrote a Stop event on every press. Two
presses before the first interrupt landed wrote two events, so a card
queued between them counted as before the latest Stop and was held,
though a card queued after a Stop should send normally. A Stop naming a
turn that had already ended, as a phone sends late, also wrote an event
for a turn it never stopped.
It now writes one only when it withdrew a queued send, or stops something
no event records yet: not a turn the journal no longer runs, and not the
live turn a Stop still in force already names, unless a card was handed
over into it since, which this Stop's interrupt sends back and must hold.
The interrupt and the "already finished" note are unchanged. A Stop at a
starting agent still always writes.
* test(native-chat): pin that a later host, eviction or close Stop never lifts a person's Stop
* chore(native-chat): put each Stop-row doc on its own declaration, and say only user-stop is journaled
* fix(native-chat): any later Stop event ends a person's Stop pause
A person's Stop paused the queue until their next accepted turn or Resume,
and a later Stop of another reason (the host stopping the agent, an
eviction, a close) was ignored. Now the pause is the latest Stop event's:
a later Stop of any reason ends a person's pause, and only a person's Stop
pauses. The fold keeps the latest Stop event whatever its reason.
An eviction of a resting chat writes no Stop event (a Stop that stops
nothing writes nothing), so it cannot release held cards; a test pins that
no event means no lift.
* fix(native-chat): a second Stop press is a repeat even when the first came before the turn showed
A Stop pressed before the agent's turn shows in the journal (before
Claude's echo, or before Codex opens the turn) records no turn. A second
press once the turn showed compared that missing turn with the live one,
wrote a second Stop event, and held a card queued between the presses.
A repeat is now judged by what was sent since the Stop in force: with
nothing sent after it (a refused send aside), a Stop that named no turn,
or named the live one, is repeated and writes nothing. Anything sent since
and not refused, including a send whose fate is unknown, makes the new
press write, since its interrupt may send that card back to waiting.
Tests: the two-press case across the turn showing; a steer between the
presses settled unknown; and a Stop naming a turn that ended while the next
card is sent but shows no turn yet, which writes and holds that card. The
fold test that claimed an eviction path is renamed.
* fix(native-chat): the queue's pause ignores a Stop or Resume row holding a value no build writes
A Stop or Resume row's value is read from disk with no shape check, and
the pause fold stored whatever it found. A stored `stopEvent: null` would
then throw on every pause check for that chat: the queue's pick, its
send, and every queue update to clients. No build writes such a row, so
this is hardening.
The fold now reads a Stop only when it is an object with a string reason
and a finite time, and a Resume only when it is `true`. Anything else is
ignored: it pauses nothing and ends nothing. The row is still not treated
as malformed, which could cut the history short.
* refactor(native-chat): one reading of a Stop's turn for its event and its note
A Stop's event and its note each worked out the same two facts on their own:
which turn the Stop is about (the one it named, else the one running), and
whether a named turn is the one the journal shows running. The event decides
before the interrupt; the note and whether the session ends decide after the
provider's answer, so those decisions stay separate, but the facts they read
are now one helper each in structured-agent-session-turn-stop-notes.ts:
structuredAgentSessionStoppedTurnId and
structuredAgentSessionStopNamesTurnNotLive. The event's turn, the note's key,
the session-ending condition, the running-command check and the repeat check
all read them. No behavior change.
Tests: a Stop naming no turn records the running turn on its event, and
rewrites that turn's note as a Stop naming it does.
* refactor(native-chat): a failed-interrupt Stop reads its turn through the shared helper
The new branch that ends a Codex child after a failed interrupt asked
whether the Stop's turn still runs with `turnId ?? liveTurnId`, a third
copy of "the turn a Stop is about". It now reads
structuredAgentSessionStoppedTurnId, the value the note key already uses,
read at the same point before the cancel. No behavior change.
Test: a Codex Stop whose interrupt failed ends the child, holds the card
queued before it with the queue paused, and writes its Stop event before
the turn's end.
|
||
|
|
ae41eb414a |
fix(terminal): give plain fish tabs Orca's codex function without changing fish's startup (#24284)
* fix(terminal): give plain fish tabs Orca's codex function without changing fish's startup A `codex` typed into a plain fish tab ran without --no-daemon because only wrapped fish tabs (startup command / ready marker) got Orca's codex function. Plain fish spawns now prepend an Orca data dir to XDG_DATA_DIRS and record the exact prefix in ORCA_FISH_XDG_DATA_DIRS_PREFIX. Fish sources the dir's fish/vendor_conf.d snippet, which first restores XDG_DATA_DIRS (unset again if it was unset), erases the marker, drops its dir from fish's derived vendor/function/ completion paths, then defines the shared fish codex function at the first prompt so the user's config.fish still wins. fish argv is unchanged; wrapped tabs keep their existing -C path. A local fallback to another shell restores the user's XDG_DATA_DIRS instead of deleting it. Bumps the terminal daemon protocol to v39 so new tabs move to a daemon that injects the env; v38 owners stay attachable. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(fish): skip the XDG handoff for -N/--no-config and empty XDG_DATA_DIRS fish never reads vendor_conf.d under -N/--no-config (also abbreviated or clustered), so the snippet could not undo the prefix; and the restore cannot tell an empty XDG_DATA_DIRS from an unset one. Both now launch untouched. Run the real-fish handoff tests in the shell contracts job, where fish is required, so they no longer skip in CI. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test(fish): compare the unset-restore case against a fish without Orca Ubuntu runners ship snapd's fish vendor snippet, which sets XDG_DATA_DIRS on every fish start, so "unset" was never the right oracle there. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(fish): treat an empty XDG_DATA_DIRS like unset so the tab still gets the codex hook Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(pty): put back the user's own launch env on a shell fallback The primary shell's launch config now records the pre-launch value of each key it writes. A fallback shell restores those values (unsetting keys that had none) instead of deleting the keys, which hands back an inherited XDG_DATA_DIRS after a fish fallback and an inherited ZDOTDIR after a zsh->bash fallback, with no per-shell special case. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * refactor(fish): drop the Node restore twin and simplify the vendor snippet - Remove restoreFishXdgDataDirs; the generic fallback restore covers it. - Snippet: read ":$XDG_DATA_DIRS:" directly and filter Orca's vendor dirs with one string match per variable. - Require inheritedXdgDataDirs in both getShellLaunchConfig option shapes. - Drop the test-only FISH_XDG_DATA_DIRS_HANDOFF_DAEMON_PROTOCOL_VERSION. - Fix stale fish comments and trim redundant -N launch cases. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * refactor(fish): stop scrubbing fish's lookup paths after the handoff Only XDG_DATA_DIRS is restored, by exact prefix; Orca's dir holds nothing but this snippet, so leaving it on fish's derived paths loads nothing else and drops the glob match. * docs(fish): drop the comment for the removed vendor-dir cleanup --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
07dad6739a |
refactor(relay): sample fleet health inside the same-cap roll instead of a separate monitor run (#24443)
* refactor(relay): sample fleet health inside the same-cap roll instead of a separate monitor run A same-cap wave no longer consumes a 15-minute monitor dry-run and its sealed, single-use, five-minute-fresh evidence. Each apply wave now samples fleet health itself right before isolation, with the monitor's evaluator, thresholds, and tolerances, for a window sized to the cell's host count (3/5/8 min), plus three lookback rules: no cell container exit in 10 min, no minute over 500 director 503s in 10 min, and director concurrency p99 within the monitor bar over 4 min. Removes the monitor-run inputs, the gate's consume/authorize steps, the break-glass override, and the same-cap-only authorization shapes in relay-monitor-evidence.mjs. The monitor workflow and the rehome enable path are unchanged. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): bound the pre-drain sample overrun and keep the drain token fresh Review follow-ups: alternating tolerated readings could hold the sample open until its step timeout, so cap the overrun at three samples past the window; record why a read failed; mint a fresh admin ID token for the drain after the sample; raise the job timeout to 90 min so a long sample cannot cancel the job past the failsafe. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * feat(relay): exempt the rolled cell and existing-only cells from the pre-drain crash rule The exit rule counted every relay container exit fleet-wide, so a cell that crashes every few hours (c25, 12 a week) blocked the very roll that fixes it, and existing-only legacy cells (c5, 15 a week) blocked rolls they take no part in. Exits are now grouped by instance, each instance is named by its own newest runtime-metrics log line, and only exits on general or migration-only cells other than the target count. An exit no configured cell can be named for trips the rule; a failed lookup is a failed read. relay-observability.tf joins the evidence-code set because the rule depends on its filter. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * test(relay): cover re-asking for an unnamed exiting instance; note the boot-exit risk Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
051ca4d34e |
feat(relay): declare US cells c32 and c33 at the 3,000-host shape (#24444)
* feat(relay): declare US cells c32 and c33 at the 3,000-host shape Declares two us-central1 cells at the Asia shape (cap 3000, 6000 request units, e2-standard-4) with the US default pool of 10, and generalises the Asia topology and admission ladder to derive each wave's region from its reviewed zone, leaving every Asia wave's behaviour unchanged. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * docs(relay): note the US canary tie-break and leave the fleet pool list to promotion Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): plan C32 and C33 as one topology wave The live-image overlay refuses a declared non-target cell with no template, so a lone C32 plan would fail on C33. Registration and promotion stay one cell at a time. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
197ea3a3b3 |
Free PR CI capacity by avoiding repeated setup and real-time test waits (#24355)
* Reduce repeated PR setup and transcript timing waits; add hosted comparisons * Align parallelism contract with Node-only external rebuild toolchain * Record hosted coverage and launch package, store, and cancellation comparisons * Apply hosted Windows setup savings and remove measured test waits * Keep measured PR package gains and remove completed comparison jobs * Report measured test counts with precise units |
||
|
|
56c7642aae |
test(orcad): skip the live-terminal runtime hand-over across a protocol bump (#24429)
* test(orcad): skip the Bun-to-Node live-terminal hand-over across a protocol bump The last Bun orcad's daemon reports protocol 38 forever, so asserting the adopted daemon matches this checkout's PROTOCOL_VERSION failed every bump. Ask the Bun slot's daemon for its protocol once, run the hand-over when it matches, and skip with the two versions named when it does not: a daemon at another protocol is never adopted across an update. * test(orcad): clean up the Bun protocol probe even when its launch fails The probe's cleanup ran only after a successful launch, so a launch that timed out or threw left its orcad and daemon running. One finally now stops the orcad, kills what it launched, and kills any daemon named by a pid file in the probe's data root. |
||
|
|
6d1a97ef98 |
fix(ssh): launch the Windows relay outside sshd's job so standard users work (#24224)
* fix(ssh): launch the Windows relay outside sshd's job without WMI Win32-OpenSSH kills a session's job on close but allows breakaway. relay.js gains a one-shot launcher mode that starts the detached relay with CREATE_BREAKAWAY_FROM_JOB through the staged process-tree addon, so a standard user no longer needs a WMI Remote Enable grant. WMI stays as the fallback for a relay without the addon, and a refusal there is named. The Windows SSH-host lanes drop their WMI grant and assert the breakaway route and adoption. * fix(ssh): find runtime holds without WMI on a standard-user Windows host The store GC read held runtimes through Get-CimInstance Win32_Process, which WMI refuses to a standard user's SSH logon, so the pass kept every runtime. On a refusal it now reads this account's own process image paths through Get-Process. * build(relay): ship the Windows relay launcher addon in every desktop package macOS and Linux packages carried Windows relays without windows-process-tree.node, so a legacy-runtime relay they uploaded to a Windows SSH host could not launch outside sshd's job and fell back to WMI, which a standard user is refused. A reusable Windows job now compiles the x64 and arm64 addons once and uploads them; release-cut, release-mac-build, and the hourly/daily/adhoc mac builds download them before build:release and require both arches. Staging now rejects a binary with the wrong PE machine, the ReadProcessMemory import, or no spawnOutsideJob export, so a stale pre-launcher build cannot ship. * ci(ssh): run the Windows SSH-host lanes when the relay process-tree build scripts change The staging and gyp-rebuild scripts decide which windows-process-tree addon the relay ships, so a change to either must re-prove the Windows host cells. * test(ci): find the mac orcad-template download by artifact name The release mac job now also downloads the relay Windows process-tree addons, so the first download-artifact step is no longer the template's. --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
8afa1db50c |
feat(ssh): rung B glibc 2.17 compat runtime; gate remote vault on host node:sqlite (#24148)
* feat(ssh): wire rung B to the glibc 2.17 compat runtime; gate rung C vault on full node:sqlite - COMPAT_RELAY_RUNTIMES lists linux-x64-glibc217; rung B plans the compat slot and compat pinned Node when glibc is below 2.28 or rung A refused with libc_floor/missing_lib. - The relay version folds the compat runtime's executable hash; refusals are cached per runtime. - The orcad template stages an optional linux-x64-glibc217 target (base package + compat node-pty slot + compat runtime marker); the verifier and materializer accept it. - node-pty slot loader falls back to the compat slot when the default slot is missing or needs a newer glibc. - Runtime store GC keeps the compat pin beside the default one on every relay connect. - hasNodeSqliteReaderApi (DatabaseSync + backup) gates relay session search and the relay OpenCode reader, which now names the host Node version in its unavailable reason; the SSH vault reader installs the compat Node on old-glibc hosts and uploads nothing when no pinned Node can run. - Rung D: a remembered noexec reports home_noexec and never advises installing Node. * fix(ssh): re-prove a replayed noexec after rung D so allowing exec recovers the host * fix(ssh): keep the rung B compat runtime pinned in the relay-connect store GC * test(ssh): mock deployment-target facts in the Windows OpenCode runtime tests * ci(ssh): build the glibc 2.17 compat slot for the hostile-host matrix; CentOS 7 lands on rung B --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> Co-authored-by: m4air <m4air@Mac.localdomain> |
||
|
|
744e7722c2 |
fix(relay): accept MIG version-name reconciliation and recreate stranded cells without rewriting the MIG (#24373)
* fix(relay): accept MIG version-name reconciliation and recreate stranded cells without rewriting the MIG The stranded-rollback recovery ran a gcloud rolling action, which renames the MIG version outside Terraform. Every later plan for that cell then reverted the label, and the capacity-plan validator refused the revert as an unreviewed MIG change, so the cell could be neither rolled nor rolled back. The validator now accepts a MIG field moving back to what relay-gce-cells.tf declares (version name and update policy), in every mode, and a test pins those values to the Terraform file. The stranded branch recreates the cell's single instance with recreate-instances, which leaves the MIG untouched. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): let a label-only MIG plan through and recreate on it in a stranded rollback A stranded rollback whose template is already in place plans only the version name revert. The validator still required the MIG template to move, so that plan was refused, and the recreate gate (changes == 0) would have skipped a plan of one change and left the drain flag set. Require the template move only when no declared field reconciles, and recreate whenever the template was not replaced (changes < 2). Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
f9940d5354 |
ci(ssh): macOS SSH-host lane for the pinned relay; fix uploads under a symlinked root (#24179)
* test(ssh): upload a root reached through a symlinked parent The upload-root realpath fix landed with #24180; this keeps macoshost's case where the root is passed explicitly beneath a symlinked parent. * ci(ssh): macOS hostile-host lane on a loopback user-level sshd Adds local-sshd cells for darwin-arm64 (macos-14) and darwin-x64 (macos-15-intel): a non-root sshd on 127.0.0.1 logs in as the runner user with SetEnv PATH=<shims>:/usr/bin:/bin:/usr/sbin:/sbin and an empty HOME, so no rc file restores Homebrew. The driver asserts rung A, terminal echo, cached runtime reuse, GC keeping the in-use runtime, no toolchain or xattr calls, and that the SFTP-uploaded Node carries no quarantine and runs as uploaded. Docker cells are unchanged; each machine runs only cells it can host. * test(ssh): fail a hostile-host run that would skip every named or hostable cell A cell named for the wrong OS or arch was silently skipped, so a macOS job on a mismatched runner went green having deployed nothing. Named cells must now be hostable here, and a gated run must select at least one cell. --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
0ad77ea2f7 |
ci(ssh): Windows SSH-host lanes (inbox + preview OpenSSH) for the pinned relay (#24180)
* ci(ssh): import the private Windows OpenSSH provisioning harness Copied unchanged from origin/OrcaWin/np-windows-ssh-provider-diagnostic (config/ci/windows-ssh-provider/preview-ssh/ at |
||
|
|
554f7f4ce5 |
feat(packaging): ship the orcad server template in desktop builds (#24155)
* build(orcad): merge per-runner prebuild slot trees into one matrix Each node-server lane builds only its own node-pty slot. Release CI needs their union before `build:orcad-prebuilds --require-slots` and the template build can run; merge-orcad-prebuilds.mjs verifies every lane's files against its own manifest, refuses duplicate slots and mismatched node-pty/N-API/Node-header builds, then writes one merged manifest. * build(orcad): keep agent-browser out of the desktop deployment template The template rides inside every desktop build (design D2). Seven ~10 MB agent-browser binaries would be ~76 MB, more than the rest of the template; design D2's package contents never listed it, and a slot without one already reports no headless browser. ORCAD_OMIT_AGENT_BROWSER=1 skips the copy; standalone build:orcad still includes it. * feat(packaging): ship the orcad deployment template in desktop builds Design D2: the server JS and every target's addons ship inside the app, as out/relay does; the ~120 MB Node runtimes stay excluded and are downloaded on demand. electron-builder copies out/orcad-template to Resources/orcad-template on every desktop OS, which is the first path materializeOrcadArtifact tries (process.resourcesPath). Platform signing rewrites native bytes the template manifest hashes: - macOS: the tree is signIgnored (codesign rejects its ELF/PE payloads); afterPack signs the darwin targets' Mach-O files with the app identity, as notarization requires, then reseals only those manifest entries. - Windows: SignPath signs after packaging, so release CI reseals from the inner-signing list (packaged-orcad-template.cjs --reseal-signed). Every other file must still match the build's hashes; afterPack verifies. ORCA_REQUIRE_ORCAD_TEMPLATE=1 makes a missing template fail beforePack and afterPack; without it a build ships none and SSH relays keep the legacy path. verify-packaged-orcad-template.test.mjs's "unused, excluded" contract is reversed on purpose. * ci(release): build the orcad template from qualified lanes and package it node-server-tests.yml becomes callable with a ref and build_template. With build_template, each lane that owns a release slot (macOS, Windows, the glibc 2.28 and Alpine lanes, and the glibc 2.17 compat lane) uploads its qualified out/orcad-prebuilds, the Windows lane also uploads both process-table addons, and desktop_template merges them, gates the full matrix plus the compat slot with --require-slots, runs build:orcad-template and uploads the orcad-template artifact. release-cut calls it at the release tag beside the other gates. The build and build-mac jobs wait for it, download it into out/orcad-template (the mac workflow from the parent run), and require it via ORCA_REQUIRE_ORCAD_TEMPLATE. The Windows signing staging skips the template's Linux/macOS payloads, and a reseal step records SignPath's bytes before the installer rebuild. A template-scoped concurrency group keeps a release call and main's push runs from cancelling each other. * test(orcad): keep the packaged-lookup imports clear of the compat-slot import edits * ci(orcad): let a rerun lane replace its template artifacts upload-artifact v4 refuses a second upload under an existing name in the same run, so rerunning a flaky node-server lane during a release would fail at the upload instead of re-qualifying the slot. * ci(node-server): build the template's Windows addons before the lane switches to Node 18 The addon build script imports TypeScript, which Node 18 cannot load, so every build_template run (release-cut included) failed on windows-2022. * fix(build): ship the orcad template's shared node_modules electron-builder's extraResources filter always drops the root node_modules of a source directory, so packaged apps lost orcad-template/node_modules and the afterPack verify failed. Copy it through its own resource entry. --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
c422936a71 |
fix(relay): anchor same-cap monitor evidence freshness to the run's authorisation, not job startup (#24349)
* fix(relay): anchor same-cap monitor evidence freshness to the run's authorisation, not job startup The same-cap gate now verifies the dry-run on its own clock and records the authorisation instant in the single-use consumed marker; each cell job checks the evidence age at that instant and bounds its own start after it. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): refuse a re-run same-cap gate before it consumes evidence; tighten order tests Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
14d4bb2e2a |
fix(ssh): Windows hosts without Add-Type staging; runtime-store GC on Windows (#24149)
* fix(ssh): collect the pinned-Node runtime store on Windows hosts
Windows SSH hosts now run runtime-store GC instead of skipping it: one
PowerShell inventory reads .runtime-ref-node-<sha> and .runtime-node refs from
every version dir, and one Get-CimInstance Win32_Process query filtered on an
image path under runtimes\ adds process holds (never by image name; a failed
query keeps everything). Stale upload stages are swept with the same rule as
POSIX. Promotion and the post-upload hold check now take the store lock on
Windows too, and the lock's own commands run unwrapped there.
Windows relay version-dir liveness now honours .relay-pid (design D5): a live
PID answers ALIVE before any pipe is touched, a dead one (ESRCH) plus refusing
pipes is exited, anything else is unverifiable. The runtime probe adopts a
pinned node.exe an earlier vault reader left without a .verified marker after
running it.
* fix(ssh): Windows stage fencing and vault runtime go through the verified node.exe
Upload-stage file identity on Windows no longer compiles an Add-Type P/Invoke
helper when the relay runs on Orca's verified pinned node.exe: the stage
commands run a fixed fs.lstatSync(..., {bigint:true}) script through it. It
prints the legacy helper's vol:high:low lowercase hex, and identity files are
compared after normalising hex spelling, so old and new clients recover each
other's stages. Host-Node relays keep the legacy helper; the choice is
documented in windows-edr-posture.md.
The Windows OpenCode vault reader now installs the pinned runtime through
ensureRemoteOrcadNodeRuntime (official zip, host-side extraction, .verified,
store lock) instead of uploading a client-extracted node.exe, and the relay dir
gains a .runtime-ref-node-<sha> so store GC keeps the runtime the vault uses.
* test(ssh): run the Windows stage-identity and store-GC tests on the Windows lane
The legacy/node.exe identity compatibility test and the Win32_Process hold path
were gated to win32 but no CI lane ran them. Add both files to the Windows
package lane and a real running-node.exe hold test.
* test(ssh): tear down Windows-lane temp trees through removeTreeSync
* test(ssh): grant the store lock to the Windows OpenCode runtime setup test
The Windows promote now runs under runtimes/.store-lock, so the mocked host
must answer the lock's CreateNew step.
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>
|
||
|
|
6aed05471c |
ci(ssh): hostile-host matrix for the relay runtime ladder (#24146)
* fix(ssh): classify a musl host missing libstdc++ as missing_lib, not wrong_libc musl's loader follows each missing-library line with one 'Error relocating ... symbol not found' per unresolved symbol, and the relocation pattern was checked first. Check missing libraries before relocation errors; the ld-linux/ld-musl interpreter case stays wrong_libc. * build(orcad): allow a partial deployment template for CI build-orcad-template --targets a,b builds and verifies only the named slots, so a CI job that can fill just the x64 Linux prebuild slots can still materialize rung A/C addons. Without the flag every target is still built and verified. * ci(ssh): hostile-host matrix for the relay runtime ladder Drives the real client-side relay deploy against Docker sshd targets and asserts the design D6 rung each lands on: Debian 10 and AlmaLinux 8 (glibc 2.28) and Alpine (musl) on rung A; Alpine without libstdc++ refused missing_lib down to D; Ubuntu 22.04 with a host Node 20 and a noexec home straight to D (home_noexec); CentOS 7 (glibc 2.17) refused libc_floor at A and C, falling to a host-npm path with no Node; and a no-egress Debian 10 still on rung A. Launched cells also prove the terminal echoes, no npm or compiler ran, a second connect reuses the uploaded runtime, and runtime GC keeps the in-use runtime while collecting an idle one. New workflow ssh-hostile-hosts.yml runs on dispatch and on path-filtered PRs. * test(ci): pin the hostile-host workflow to the headless-server builder images The matrix builds its runtime slots in copies of the node-server lanes' Alpine and manylinux images; this contract fails when NODE_RUNTIME_PIN or either builder digest moves in one workflow and not the other. * test(ssh): reconnect as the same client and retry a grace-held PTY owner in the hostile-host matrix --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
bd90da7a5b | ci: share PR planning setup and reuse the static native cache (#24329) | ||
|
|
9ed7b39b3c |
fix(relay): run the same-cap headroom gate in the modes the job actually receives (#24343)
The parent workflow collapses canary-apply and batch-apply into the job mode apply, so the headroom step's canary-apply/batch-apply condition never held and the gate was skipped on every real roll. Run it wherever the drain runs (apply, rollback before its restart) and in read-only verify; skip only a resumed rollback, which drains nothing. A new workflow-shape test fails on any job step comparing against a mode the parent cannot pass, and on a drain that can run without the headroom check. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
ddd4927a0b |
build(orcad): server node-pty slots at glibc 2.28, plus a glibc 2.17 compat slot (#24134)
* build(orcad): build server glibc slots on glibc 2.28 and add the glibc 2.17 compat slot
Design D6: the default linux-{x64,arm64}-glibc node-pty slots now build in
manylinux_2_28 (digest-pinned) and are gated at glibc 2.28 / GLIBCXX_3.4.25
through a floor profile on verify-linux-glibc-floor.cjs; the desktop keeps
its Ubuntu 20.04 (2.31) default.
Adds the opt-in linux-x64-glibc217 compat target: NODE_RUNTIME_COMPAT_ASSETS
pins the unofficial glibc-217 Node (update/check pin scripts cover it, outside
SERVER_TARGETS), and a new CI lane builds the compat slot in manylinux2014
with static libstdc++, gates it at glibc 2.17 with no shared C++ runtime in
DT_NEEDED, and smokes it under the glibc-217 Node.
* refactor(node-runtime-pin): route compat lookups through isCompatServerTarget; keep the glibc doc's slot-name paragraph intact
---------
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
|
||
|
|
6593d7d194 |
feat(orcad): run orcad on the pinned Node instead of Bun (#24110)
* ci(daemon): gate PRs on daemon protocol crossing from the newest release Lands daemon-protocol-facts.mjs from the Windows update diagnostic branch with a stricter parser, and adds check-daemon-protocol-crossing.mjs (rule R1): the working tree must attach the newest release tag's daemon. Rollback crossing is reported only. Runs in the cross-version-wire job, which already has full tags; tag selection moves to config/scripts/stable-release-tags.mjs so both use one rule. * feat(persistence): run profile backups in the worker whenever its entry is bundled * refactor(orcad): make profile and native preflight runtime-neutral The profile preflight parser now takes the expected runtime identity from the caller (shipped callers pass the pinned Bun identity), and the native preflight is renamed to orcad-runtime-native-preflight with neutral wording. * feat(runtime): pin the Node 24.21.0 server runtime with an offline CI check Add src/shared/node-runtime-pin.ts (NODE_RUNTIME_PIN, SERVER_TARGETS, NODE_RUNTIME_ASSETS for all 8 server targets plus the headers tarball), generated by config/scripts/update-node-runtime-pin.mjs from the nodejs.org and unofficial-builds SHASUMS. check-node-runtime-pin.mjs verifies, with no network, that the pin tracks the locked Electron, matches engines.node's major, and covers exactly SERVER_TARGETS; it runs in the static analysis job. ORCAD_BUN_TARGETS consumers now read SERVER_TARGETS so there is one target list; orcad's Bun runtime and build output are unchanged. * test(persistence): skip plain-Node backup selection tests in the Bun profile suite * fix(runtime): reject a pinned archive that belongs to another target * ci(daemon): fail PRs that swap a runtime launcher and bump the daemon protocol D7.1 R3: hosting orcad or the daemon on another runtime is not a protocol change, so one PR must not do both. The launcher file list lives in the check script; the allow-runtime-launcher-protocol-bump label overrides it. * feat(orcad): select pinned-Node slots by a .runtime-node marker D7.1 R5: a Node slot names its shared runtimes/node-<sha256>/node through .runtime-node instead of .build-target, so Bun-era clients read it as a legacy slot rather than exiting 78 on a missing bun-runtime. Nothing builds the marker yet. * fix(runtime): load the Node pin without the typeless-module warning check-node-runtime-pin.mjs now requires the pin and takes nodeDistArchiveName from its own module, so it no longer loads the update script's build graph. * fix(orcad): resolve Node slots to the design's runtimes/node-<sha>/bin/node layout * feat(orcad): 8-slot node-pty prebuilds against the pinned Node headers at N-API 8 - build-orcad-prebuilds.mjs adds win32-x64/arm64 (conpty.node, the vendored conpty.dll/OpenConsole.exe, upstream's N-API conpty_console_list.node), compiles in a scratch copy against the hash-verified pinned headers (node.lib pinned per Windows arch) with NAPI_VERSION=8, rejects post-8 node_api_* imports, and writes a schema 2 manifest with per-file sha256, N-API level and the glibc need. - --require-slots [slots] verifies files against hashes; --smoke loads the slot under the pinned Node and spawns a PTY; --print-slot names the host slot. - The slot installer gates on N-API, libc, arch, glibc and file hashes instead of the exact NODE_MODULE_VERSION, and installs nested files (conpty/). - bun-profile-tests.yml builds, verifies and smokes each runner's slot. * fix(orcad): scope node-pty's glibc .symver pins to glibc on musl prebuild slots musl's unversioned libc cannot satisfy openpty@GLIBC_* references at link time, so the Alpine slot compile would fail. Pin the staged pty.cc guard to __GLIBC__ and assert both musl transforms against the installed patch. * feat(orcad): run orcad on the pinned Node instead of Bun A packaged orcad slot now references the pinned Node 24.21.0 by its executableSha256 (`.runtime-node`, `.server-target`) instead of carrying bun-runtime, and ships node-pty from the slot's prebuild, only its own ripgrep, and no Windows Bun PTY gate. The runtime lives beside the slots at runtimes/node-<sha>/bin/node (node.exe on Windows, upstream name). - build:orcad (build-orcad-node.mjs) builds the host slot's prebuild when missing and places the pinned runtime; the template is schema 3 with per-target files. - handoffToBundledOrcad() resolves the slot's runtime reference and checks process.versions.node against the pin; a host Node >= 18 still hands off. Startup preflight keys on running as that runtime; callers expect 'node'. - orcad and its daemon use node-pty (ConPTY + windows-pty-job on Windows); the Bun PTY sources, gate entry and canUseBunPty branches are removed. - SSH deploy uploads the official archive once per pin, extracts and hash-checks it on the host, and self-tests it before publishing. Bun slots stay launchable for rollback; Node slots never use host Node. - The runtime materializer is generic over pinned assets; the Bun wrapper remains only for the OpenCode vault reader (design Phase 2). - Cross-runtime test: a profile DB written by Bun 1.4.2 (WAL left by SIGKILL) opens and backs up under the pinned Node, and the reverse. No daemon PROTOCOL_VERSION change (design D7.1 R3). * docs(ci): name the headless lanes after the pinned Node, drop Bun shard timings Design D10: ci-demand-rollout.md and ci-runner-efficiency.md follow the bun-profile-tests.yml -> node-server-tests.yml rename; shard timings drop the deleted Bun PTY tests and follow the renamed ones. * chore(ci): count the runtime archive download as a runtime launcher path * fix(orcad): pin the macOS C++ standard for node-pty prebuilds The official Node headers' config.gypi sets clang: 0, so common.gypi skips its gnu++20 xcode_settings and Apple clang 15 (macos-14 runners) compiles node-addon-api as C++98. * fix(orcad): resolve the preflight's slot through realpath, as the handoff does A symlinked orcad.js handed off to its real slot's pinned Node, but the startup and profile preflights read the symlink's directory, found no runtime marker there, and silently skipped the readiness check. * refactor(ssh): drop materializeCachedNodeRuntime, which nothing calls Deploys upload the verified official archive (design D5); no client path needs an extracted Node executable cached by digest. * test(orcad): gate the Bun-to-Node upgrade and Node-to-Bun rollback with live terminals Design D7.1 R1/R3/R4 and D7.2. The last Bun orcad and this checkout's Node slot are installed side by side under ~/.orca-remote, launched and stopped with the client's own deploy commands, and share one data root. Each direction proves the incoming orcad adopts the outgoing runtime's daemon (same PID, same shell, output continues), opens its profile database and backs it up with its own shipped worker, and that GC keeps the slot the live daemon was forked from. The node-server Linux lanes provide Bun 1.4.2 and build that Bun orcad from main, and run with --cross-runtime. --artifact and --cross-runtime now make their tests fail on a missing input instead of skipping. * ci(node-server): pin node:24.21.0-alpine by its multi-arch index digest * test(ssh): name the runtime archive fixture after its role * test(node-server): load node-pty from the packaged slot in artifact runs The node-server lane installs dependencies without building node-pty, and Linux has no upstream prebuild, so the real-PTY failed-I/O teardown test (picked up by the pty-subprocess selector) could not load pty.node. In --artifact runs, alias node-pty to out/orcad's shipped slot so the test exercises the addon orcad actually runs under the pinned Node. * fix(orcad): let the Windows profile preflight exit after its PTY probe On Windows, node-pty keeps the conout worker thread and pseudoconsole alive until kill(), even after the shell exits. The PTY health probe never killed a cleanly exited probe, so the packaged preflight printed its readiness line and then hung until the build's 30s timeout, reported with an empty stderr. - The probe kills its PTY on Windows after exit and uses the bundled ConPTY the daemon spawns with. - The preflight exits once stdout is flushed; its owner reads to EOF. - Preflight failures now report code, signal, timeout, stdout and stderr. * test(node-server): load the slot's node-pty in the real-PTY test, not by alias A vite alias redirected only ESM imports of node-pty; windows-pty-job and local-pty-utils resolve it through require, so Windows loaded two conpty.node copies and the Git Bash job-membership proof read an empty job. The failed-I/O teardown test now loads node-pty through a fixture that picks the packaged slot in artifact lanes. The pty-subprocess selector was a prefix that also pulled in its POSIX-host sibling unit tests, which pr.yml runs and which were never qualified on Windows. Select the directory plus the two sibling files that belong here. --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
3e5c8d9f8e |
feat(relay): declare Asia cell c31 at the c30 shape (#24310)
Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
d2dfc79764 |
ci(daemon): runtime-launcher protocol ratchet and Node slot marker (#24108)
* ci(daemon): gate PRs on daemon protocol crossing from the newest release Lands daemon-protocol-facts.mjs from the Windows update diagnostic branch with a stricter parser, and adds check-daemon-protocol-crossing.mjs (rule R1): the working tree must attach the newest release tag's daemon. Rollback crossing is reported only. Runs in the cross-version-wire job, which already has full tags; tag selection moves to config/scripts/stable-release-tags.mjs so both use one rule. * feat(persistence): run profile backups in the worker whenever its entry is bundled * refactor(orcad): make profile and native preflight runtime-neutral The profile preflight parser now takes the expected runtime identity from the caller (shipped callers pass the pinned Bun identity), and the native preflight is renamed to orcad-runtime-native-preflight with neutral wording. * feat(runtime): pin the Node 24.21.0 server runtime with an offline CI check Add src/shared/node-runtime-pin.ts (NODE_RUNTIME_PIN, SERVER_TARGETS, NODE_RUNTIME_ASSETS for all 8 server targets plus the headers tarball), generated by config/scripts/update-node-runtime-pin.mjs from the nodejs.org and unofficial-builds SHASUMS. check-node-runtime-pin.mjs verifies, with no network, that the pin tracks the locked Electron, matches engines.node's major, and covers exactly SERVER_TARGETS; it runs in the static analysis job. ORCAD_BUN_TARGETS consumers now read SERVER_TARGETS so there is one target list; orcad's Bun runtime and build output are unchanged. * test(persistence): skip plain-Node backup selection tests in the Bun profile suite * fix(runtime): reject a pinned archive that belongs to another target * ci(daemon): fail PRs that swap a runtime launcher and bump the daemon protocol D7.1 R3: hosting orcad or the daemon on another runtime is not a protocol change, so one PR must not do both. The launcher file list lives in the check script; the allow-runtime-launcher-protocol-bump label overrides it. * feat(orcad): select pinned-Node slots by a .runtime-node marker D7.1 R5: a Node slot names its shared runtimes/node-<sha256>/node through .runtime-node instead of .build-target, so Bun-era clients read it as a legacy slot rather than exiting 78 on a missing bun-runtime. Nothing builds the marker yet. * fix(runtime): load the Node pin without the typeless-module warning check-node-runtime-pin.mjs now requires the pin and takes nodeDistArchiveName from its own module, so it no longer loads the update script's build graph. * fix(orcad): resolve Node slots to the design's runtimes/node-<sha>/bin/node layout --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
2a83c9536f |
ci(daemon): gate PRs on daemon protocol crossing from the newest release (#24089)
Lands daemon-protocol-facts.mjs from the Windows update diagnostic branch with a stricter parser, and adds check-daemon-protocol-crossing.mjs (rule R1): the working tree must attach the newest release tag's daemon. Rollback crossing is reported only. Runs in the cross-version-wire job, which already has full tags; tag selection moves to config/scripts/stable-release-tags.mjs so both use one rule. Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
3135fbbf49 |
feat(runtime): pin the Node 24.21.0 server runtime with an offline CI check (#24087)
* feat(runtime): pin the Node 24.21.0 server runtime with an offline CI check Add src/shared/node-runtime-pin.ts (NODE_RUNTIME_PIN, SERVER_TARGETS, NODE_RUNTIME_ASSETS for all 8 server targets plus the headers tarball), generated by config/scripts/update-node-runtime-pin.mjs from the nodejs.org and unofficial-builds SHASUMS. check-node-runtime-pin.mjs verifies, with no network, that the pin tracks the locked Electron, matches engines.node's major, and covers exactly SERVER_TARGETS; it runs in the static analysis job. ORCAD_BUN_TARGETS consumers now read SERVER_TARGETS so there is one target list; orcad's Bun runtime and build output are unchanged. * fix(runtime): reject a pinned archive that belongs to another target --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
7c119465b0 |
fix(relay): define restart-safe by the cell runtime, and refuse waves without headroom (#24259)
* fix(relay): let a same-cap drain finish when only unplaceable hosts remain The c28 canary on 2026-10-01 drained the cell to zero live connections, but four hosts with no free slot anywhere kept redialling and held director leases on it, so the restart-safe wait timed out and left the cell isolated and empty. The drain wait now also passes once the runtime has carried nothing for a sustained quiet window while a small, capped number of leases remain, and logs the escape. Apply modes also refuse a cell whose hosts exceed 80% of the free slots on the other general cells, so a wave cannot strand hosts in the first place. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): define restart-safe by the cell runtime, not director leases Replaces the opt-in stranded-host escape with a corrected definition. A restart is safe when the cell runtime carries nothing live and no migration is open, sustained for the drain pace window. Director activity leases lag hosts that already left or cannot be placed, so they are reported in a progress line and the verified result instead of blocking the restart. The same-cap drain passes its existing pace window. The headroom script is added to the trusted evidence code paths with the other production scripts. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): print stranded director counts on every restart-safe sample Each restart-safe poll now prints its sample count and the director's restart-blocking leases, request units, reserved remainder, and migrations under `stranded`; the verified line carries the same object. Open migrations still block because each is pinned to the cell incarnation a restart replaces. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 * fix(relay): require the pace window for every live restart-safe wait Pre-auth and total connections no longer reset the restart-safe window: on drained c28 they flickered with unauthenticated redials in a third of samples, which a restart does not lose. They stay in the progress output. Every live restart-safe call must now pass --pace-window-ms. The capacity job and staging proof drain unpaced, so they pass the production 300000 ms window, and the calls that relied on the 180000 ms default get 480000 ms. Headroom free slots now follow the director's placement rule: the admission pause minus the larger of observed and enforced units, minus outstanding control reservations. Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010 |
||
|
|
5cda0f4508 |
refactor(native-chat): keep agent-session records in the chat journal database (#24006)
* fix(native-chat): report a failed startup chat reconcile instead of failing app startup At startup the chat host re-checks every saved chat's lease and writes the result to agent-sessions.json. If that write failed (the file lock gave up, the file could not be written, or the file was written by a newer Orca and is read-only here), reconcileRestartLeases rejected, the startup IPC call rejected, and the renderer fell into its degraded "Session restore failed. Changes won't be saved until restart" mode. The reconcile is bookkeeping: a lease left unreconciled grants no writer, and every attach, send and read of a chat reconciles its own lease again. So the startup reconcile now reports its failure through a new optional host dependency, onStartupReconcileFailure, and resolves. The runtime routes it to its onError sink under the scope structured-agent-session-startup-reconcile, or logs it when no sink is installed (the desktop installs none). * fix(native-chat): read restored chats without waiting on lease bookkeeping With native chat on and a chat tab open at quit, the renderer's startup also awaits the chat tab restore (session.tabs.listAll). That restore re-ran the lease reconcile before reading each chat and rethrew its store failure, then recorded each restored tab as visible through a store transaction that throws on a held lock or a read-only store. Either one failed the restore, so startup still fell into "Session restore failed". Reading a chat grants no writer, so the reconcile startup and the restore run is now a reader's: createReaderReconcile never throws, answers whether every lease is settled (recovery is resolved only then; the journal opens either way), and reports each distinct failure once until a reconcile settles. Attach and agent start keep the strict reconcile. The restore's tab republish logs a failed visibility write and still publishes the tab, since a client drops every unpublished chat tab; user-driven publishes still refuse. The host dependency is renamed onLeaseReconcileFailure (scope structured-agent-session-lease-reconcile), since it now also reports for reads. * fix(native-chat): keep every record-store write off the startup chat read path Round-2 review found two more writes on the startup chat restore that could still fail it and put the app into "Session restore failed": republishing a /clear replacement recorded its tab visibility strictly, and resolving a chat's recovery rethrew its store error. The restore also paid one lock wait per tab and per batch of chats while the lock stayed held. The restore now derives tabs from state it already holds: - publishStructuredAgentSessionTab splits into the strict write and projectStructuredAgentSessionTab, which only updates the runtime's snapshot. The restore and /clear replacements only project: a saved tab index already lists every restored chat, and a /clear moves the tab in the same write that commits it. visibilityWriteMayFail is gone. - Chats a legacy profile restores that the index does not list are recorded in one best-effort transaction (store.showSessionTabs), so a failure leaves the index absent to seed again rather than partial. - The read restore's recovery resolution is caught and reported through onLeaseReconcileFailure, deduplicated with the reconcile's reports. - Once lease bookkeeping fails in a restore pass, the rest of that pass skips it, so a held lock costs one wait for the startup reconcile and one for the restore, however many chats are open. User actions (create, reveal, attach, send, the /clear commit) keep their strict writes. * test: open, seed and read the agent-session record store through one harness Tests that open the durable agent-session record store, seed it, or read back what it persisted now go through agent-session-record-store-test-harness.ts instead of calling AgentSessionRecordStore.open or touching agent-sessions.json themselves. A later change that moves the store into the chat database then changes the harness instead of every test. No production code changes. Tests whose subject is the JSON file itself (its .bak recovery, salvage, schema versions, permissions, and what older builds read back) keep reading and writing the file directly; the storage move rewrites or deletes them. * fix(native-chat): start each restore pass from one lease check and stop its bookkeeping at the first failure The restore now runs one reader lease check for the pass and lets each chat re-check and resolve recovery only while the pass is still settled. The first refusal or failed write clears it for the rest of the pass, and every chat is still opened for reading. With another process holding the lock, startup waits on it once in prepare and once in the restore, however many chats are open; a legacy profile waits once more for its tab-index seed. * docs(native-chat): correct restore comments and a test name to match the final design * test: address the record-store harness by the host's state directory The harness took the store's own folder, so each caller picked one (join(root, 'store'), or 'agent-sessions' where a test read the store the runtime owns). A later change that moves the store into the state directory's journal database could not tell those apart, and would have had to edit every caller again. Every harness function now takes the state directory, the one the test's journal database and recovery capsule already live in, and keeps the store in the same subfolder the runtime uses. Callers pass that directory; store-only tests pass their temp directory unchanged. Format tests that share a directory with harness calls take the file path from testAgentSessionStoreFilePath. The folder name moves from a private constant in the runtime to AGENT_SESSION_STORE_DIR_NAME beside the store's file name, so the harness shares it without importing the runtime. Its value and every path built from it are unchanged. * refactor(native-chat): keep agent-session records in the chat journal database The record store's records, operation ledger, retired claim keys and chat tab index become tables in agent-session-journal.db (user_version 4). The version-4 migration copies agent-sessions.json in its own transaction and never writes, renames or deletes that file or its .bak. Each store write is one journal transaction over exactly the rows it changed, checked with the load rules; the file lock, the external-change refresh and its hash, the .bak rotation, salvage and the hot-path recovery fence are gone from the store. * wip: importer tests * test(native-chat): cover the records migration, the import, row writes and read-only records * docs(native-chat): retire comments that describe the records file as the live store * test(native-chat): drop the record-store harness's leftover file path and type the import fixture * test(native-chat): let the host harness cleanup wait out a recovery-offer read's lock * fix(native-chat): let Stop reach the agent when its ledger row cannot be written Stop's operation-ledger row now shares the database with the chat history, so damage, a full disk or a stranded transaction on that write refused the Stop before the interrupt. A cancel plan now takes its decision from the committed ledger in memory, runs without settling, and warns that the row was skipped. Other mutations answer proven damage with the typed "Unable to load this chat." refusal instead of the raw SQLite error. * fix(native-chat): answer whether a profile holds chats from the database's rows Every host install creates agent-session-journal.db, chats or not, and the version probe created it too, so its mere existence made every profile that ever installed the host wait on host install and reconcile at startup. The check now opens the database read-only and looks for a record or tab row, lets the records file answer while its import is still owed, and counts an unreadable database as present. The version probe no longer creates the file. * fix(native-chat): open a chat from history when its tab index cannot be written Over records a newer Orca wrote, every write is refused, so opening a closed chat from Agent Session History failed on the tab-visibility write and the chat read as unreachable. Like closing a tab, opening one now reports a failed restore-index write and still publishes the tab. * fix(native-chat): keep the records import owed when the backup read fails transiently A torn records file whose .bak could not be read (EACCES, EIO) was reported as unusable, so the migration completed with nothing copied and never retried. A non-ENOENT read failure of either copy now carries its cause, which the importer classifies as a read that can clear. * test(native-chat): pin that an unreadable records file never falls back to its backup * fix(native-chat): restore imported chats' tabs when the records file had no tab index A chat created while the import was owed recorded a tab index holding only itself. When the file it later imported had no index, that index still read as recorded, so the imported chats' tabs never came back. The import now clears the recorded marker in that case, and restore falls back to the profile's tabs. * refactor(native-chat): drop the unused in-transaction store write Nothing called it, and it bypassed the write queue and the read-only refusal. * docs(native-chat): say that an unusable records file is left untouched but never re-imported * refactor(native-chat): keep the provider handle chain check as main has it The chain-validation refactor has no measured need in this change. * docs(native-chat): retire lease-renewer comments that describe the records file as the live store * fix(native-chat): keep a throwing failure sink from failing the startup chat read The lease bookkeeping failure reporter called the host's failure sink directly, so a sink that threw turned a reported, recoverable store failure back into a rejected startup reconcile or read restore. The reporter now catches a sink throw and logs both the original failure and the sink error with console.warn. * test(native-chat): wait for a replaced host's restart-offer writes before cleanup A restart test replaces the host without tearing the old one down, so the old host's fire-and-forget restart-offer withdrawal could still hold the recovery capsule's lock directory when cleanup removed the test directory (ENOTEMPTY). The harness now hands hosts a capsule that tracks running operations and waits for them before removing the directory, replacing the rm retries. * docs(native-chat): retire the abandon helper's note that the store re-creates its directory * fix(native-chat): restore a chat opened while the import was owed beside the profile's chats When the imported records file had no tab index, restore fell back to the profile's saved tabs, which never list a Claude chat, and the seed then rewrote the tab table without the chat opened while the import was owed. The tab rows that chat left are now loaded as unrecorded, restore takes them together with the profile's chats, and the seed keeps their tab ids. * test(native-chat): pin that a create whose tab index write fails still opens the chat * docs(native-chat): say why restore puts chats opened while the import was owed first * test(native-chat): replace a ledger row rather than change it in place in the Send-now rerun test The record store freezes published rows in tests, so setting a row's outcome in place threw; the test now swaps in a changed copy, as its sibling cases do. |