Commit Graph
968 Commits
Author SHA1 Message Date
Jinwoo Hong 40e60b7385 refactor(ai-vault): delete the unused session-scanner worker thread (#24607)
* refactor(ai-vault): delete the unused session-scanner worker thread

Production always scans through the forked session-scanner service process;
the worker thread was reachable only under NODE_ENV=test or the
undocumented ORCA_AI_VAULT_SERVICE_PROCESS=0 switch, and nothing fell back
to it on service failure. Remove the thread (spawn, client, protocol, entry,
tests), its build entry, knip and plain-node-guard listings, and the
backend switch, so session-scanner-background always routes to the service.

- Move the scan options type to the service protocol as
  AiVaultServiceScanOptions.
- Tests now mock session-scanner-service-spawn, the seam production calls.
- Repoint the hot-path listing reliability gate from the worker-client test
  to the service-client test, which covers the same bounded-queue,
  cancellation, and fault-restart properties for the real executor.

STA-9122

* test(ai-vault): cover the service's Claude-vs-OMP subagent lister choice

Runs the real service entry and subagent reader, replacing only the two
per-agent listers, so a swapped lister choice fails.

STA-9122
2026-10-02 19:58:33 -04:00
Jinwoo Hong b06f40f3ba fix(opencode): read the binder's session store off the main thread; ship the reader worker in orcad (#24638)
* fix(opencode): read the binder's session store on the foreign SQLite reader worker (STA-9122)

Before: the OpenCode session binder listed new sessions from opencode.db with
node:sqlite on the main thread every 60 s (and on SessionStart kicks), so a
large or contended store could stall the app the same way Cursor's did.

After: the read is a pure openCodeBinderSessions reader in
foreign-sqlite-readers/readers/, run only on the worker. The binder's
correlation, pane snapshot and process sweep stay where they were.

- The binder round awaits listSessions and re-checks its generation right
  after, so a stop() during the read discards the round before it touches the
  unbound map or the watermark.
- The client's in-flight dedupe key now includes the cursor, so a stale round
  from before a restart cannot hand its rows to the restarted round.
- Idle teardown is per reader. The binder lane keeps its thread for 120 s,
  longer than its 60 s poll, so the thread is not respawned every round.
- A timeout, crash, malformed reply or unstartable worker resolves to [] (no
  sessions), the value the old read already returned on failure.
- An absent store still reads as [] without a log line, and a permission or
  corrupt-file failure still logs (kept from #24577, now in the reader: it
  stats the path and throws anything but ENOENT/ENOTDIR to the client's log).
- The binder lane inherits #24572's limits from the shared lane: no respawn
  until a timed-out worker has exited, 2 consecutive deaths, a queue cap of
  8. Its timeout stays 60 s, matching its poll.
- dispatch switches on the destructured kind, so a new kind without a case
  still fails to compile.

orcad: the hook server runs there too, so orcad now ships
foreign-sqlite-reader-entry.js beside orcad.js (ORCAD_ARTIFACTS, built as an
orcad child). build-orcad runs a smoke check that starts the built worker
under the build's Node and under the pinned runtime, and does a real binder
read on a fixture DB, a Cursor read of a missing file and an OpenCode history
list. The OpenCode history scanner uses the same entry and was bundled into
orcad without it, so on orcad it always failed closed; it can now run.

Tests: reader (cursor, same-ms ids, OpenCode 2 rows, missing then created,
corrupt, inaccessible directory), retirement gate for the binder lane, dispatch
routing, client lane (rows, failure -> [], dedupe per cursor, own thread, idle
teardown default and override), binder loop with an async listSessions
(failure -> [], stop during the read), orcad path resolution through orcad's
host adapters, artifact list, and the smoke check against good, missing and
non-reading entries.

* test(opencode): cover the binder read deadline with fake timers and name the failure test accurately (STA-9122)
2026-10-02 19:58:29 -04:00
Neil 533446dde6 Stop mocked renderer imports from qualifying headless CI (#24902)
* Decouple headless running-work tests from the renderer

* Keep the shared running-work probe contract documented
2026-10-02 16:56:43 -07:00
Neil a824fb74ab Reuse the headless detector compiler without installing full dependencies (#24895)
* Reuse the headless detector compiler without full dependency setup

* Keep optional compiler-cache saves from failing cache warming
2026-10-02 16:56:08 -07:00
Neil 58cf72d48e Skip slower root package-store restores in Linux PR jobs (#24896) 2026-10-02 16:09:59 -07:00
Jinwoo Hong f8a31edd6f refactor(cursor): move the desktop-login read onto the shared foreign SQLite reader worker (#24603)
* refactor(sqlite): rename the OpenCode SQLite worker entry to foreign-sqlite-reader (STA-9122)

The worker thread that reads OpenCode's database off the main thread is about
to read other apps' databases too, so its entry is renamed to what it is:
src/main/foreign-sqlite-readers/foreign-sqlite-reader-entry.ts, built as
out/main/foreign-sqlite-reader-entry.js.

Why now: #24572 fixed the Cursor focus freeze with a second, dedicated
worker. Rather than grow one worker per foreign app, the next commit moves
Cursor onto this entry and deletes that worker. This commit is the rename
only; #24572's cursor-desktop-profile-worker-entry lines stay until then.

It moves out of ai-vault/ into a new foreign-sqlite-readers/ module because
it will no longer be session-scanner code; the module will own the readers,
their dispatch, protocol and main-process client.

The entry still routes only OpenCode kinds in this commit. The OpenCode
dispatch, protocol and process entry stay in ai-vault/ and stay OpenCode-only,
because the SSH/WSL relay reader bundles them (build-relay.mjs).

Every reference is updated: electron.vite.config.ts input key, knip entry,
the plain-node entry guard and its test, the asarUnpack list (the scanner
service still spawns this entry under ELECTRON_RUN_AS_NODE), and the
electron-builder test that reads the filename. The filename and the
beside-or-one-up (Rollup chunks) lookup now live in
foreign-sqlite-reader-entry-path.ts, which the OpenCode spawn reuses, plus an
Electron-main resolver that uses the packaged app.asar path.

* fix(cursor): move the desktop-login read from its dedicated worker onto the foreign SQLite reader (STA-9122)

#24572 fixed the Cursor focus freeze (#24360) with a dedicated worker
(rate-limits/cursor-desktop-profile-worker*.ts). Orca already runs OpenCode's
database reads on a worker, and more foreign-app SQLite reads are coming, so
keeping one worker per app means one entry, build input, asarUnpack line, knip
entry and guard line each. This keeps one pattern instead: Cursor's
state.vscdb read runs on the shared foreign SQLite reader entry, and the
dedicated worker, its entry and its config lines are deleted.

What moves:
- The read itself is a pure cursorProfile reader in
  foreign-sqlite-readers/readers/ (was rate-limits/cursor-desktop-state-db.ts),
  run only on the worker. A separate dispatch owns the new kinds and refuses
  an unknown kind. The entry routes OpenCode kinds to the untouched OpenCode
  dispatch, so the relay's OpenCode reader stays byte-identical.
- ForeignSqliteReaderClient gives each reader its own WorkerThreadRequestQueue
  lane (own lazily started, idle-torn-down thread; one shared factory) with
  in-flight dedupe per database path. Any failure resolves to the reader's
  existing failure value and never falls back to the main thread.

Kept from #24572, so every reader gets them:
- Await worker retirement before respawning. Worker.terminate() cannot
  interrupt a native SQLite call (e.g. a WAL-index rebuild), so the old
  thread lives on until that call returns; respawning at once stacked a new
  thread on the same work for every timed-out read (#24572 measured three
  live workers). This belongs in the shared host, which fire-and-forgot
  terminate(): LazyWorkerThreadHost now takes awaitRetirement and refuses to
  spawn until the terminated worker settles, and the queue fails calls closed
  meanwhile. Opt-in, because pure-JS clients (session scanner abort, port
  scan) respawn right after an abort. A rejected terminate() also ends
  retirement, so it cannot latch the reader off (raised in #24572's review).
- 10 s Cursor timeout, 2 consecutive deaths, a queue cap of 8.
- #24572's worker tests, rewritten against the shared client: responsive
  caller plus coalesced probes, unavailable worker without path leaks,
  stalled-worker recovery, no respawn before retirement, dispose settles.

Tests: reader, dispatch, client (timeout, 10 s default, crash, malformed,
unavailable without a main-thread read, dedupe, queue cap, own thread per
reader), queue retirement (stalled and rejected terminate), import boundary,
and an event-loop test reading a ~50 MB WAL with no -shm on a real worker.

* test(sqlite): walk the reader import boundary with the shared source-tree scan (STA-9122)

* fix(sqlite): key reader dedupe on a caller-supplied key, not the path alone (STA-9122)
2026-10-02 19:00:41 -04:00
Neil add1c55590 Skip slower Windows root package-store restores in CI (#24885)
* Skip slower Windows root package-store restores in CI

* Update reviewed mobile dependency-store cache expression
2026-10-02 15:24:30 -07:00
Neil 1aa0860f7e Keep large Markdown previews responsive (#24880)
* Keep large Markdown previews responsive

* Fix large preview review navigation and Find budgets

* Initialize preview scroll caches once and check viewport visibility

* Restore large previews after loaded rows are measured

* Refresh loaded Markdown rows after viewport changes

* Keep Markdown revisions visible and reuse bounded search text
2026-10-02 15:22:03 -07:00
Neil 75be95fd8c ci: move ARM Mac qualification to macOS 15 (#24760) 2026-10-02 14:37:48 -07:00
Neil 61836f6026 Reduce scheduled CI cache warming to every six hours (#24881)
* Reduce scheduled CI cache warming to every six hours

* Document cache warmer recovery interval and measured tradeoff
2026-10-02 14:35:06 -07:00
Jinwoo Hong d3a406fcbe ci(release): publish after a skipped orcad template (#24882)
* ci(release): publish after a skipped orcad template

#24872 skips orcad-template for tags that predate it, but a skipped ancestor
skips every job that keeps the implicit success(), so publish-release and the
post-release jobs never ran for v1.4.219.

* test: brace-free filter in the orcad downstream contract
2026-10-02 17:32:33 -04:00
Neil cc73c8e1a7 ci: overlap ARM SSH setup and independent observation waits (#24714) 2026-10-02 14:00:49 -07:00
Neil b94c75c4bd Reuse pnpm verification records in Alpine CI (#24817)
* ci: reuse pnpm verification records in Alpine builders

* ci: qualify consumers of the verification restore action

* ci: match Linux verification cache archive paths
2026-10-02 13:54:04 -07:00
Neil ac28e8c85e Skip dependency installation for known headless build inputs (#24716)
* ci: defer headless dependency installation until graph analysis is needed

* docs: align headless CI rollout with platform and cache policy

* test: isolate headless detector output from the parent CI step
2026-10-02 13:53:56 -07:00
Jinwoo Hong d3e592365e ci(release): skip the orcad template for tags that predate it (#24872)
A patch cut from a base older than #24155 has no orcad template source, so
the template job could never pass and every desktop build waited on it.
2026-10-02 16:05:00 -04:00
Jinwoo Hong bdab91e531 fix(windows): link the CLI launcher's C runtime statically so orca.exe runs without the VC++ Redistributable (#24484) 2026-10-02 15:25:30 -04:00
Jinwoo Hong b3f39ee178 fix(release): accept the build identity as a later declarator in the minified telemetry check (#24219) 2026-10-02 15:25:25 -04:00
Jinwoo Hong 564f4d021a feat: live updates for agent state rules (#24387)
Orca downloads a newer agent-state-rules.json from a fixed GitHub release (stable or next channel), validates it like the bundled rules, and applies it without a restart; a local override wins over the download, which wins over the bundled rules. A hand-started workflow from main is the only publisher; merging publishes nothing.
2026-10-02 14:59:24 -04:00
Neil 5706d144b7 Check E2E packages where the installer reads them (#24785)
* test: check plugin fixture worktree cleanup

* test: check E2E installer package environment
2026-10-02 05:49:30 -07:00
Neil 76b1a90ff6 chore(deps): update reviewed dependencies across Orca (#24561)
* chore(deps): update reviewed desktop dependencies and tooling

* chore(deps): update compatible mobile packages and Fastlane

* chore(deps): update cloud transports and enforce release age

* chore(deps): patch documentation dependencies and record review

* chore: remove dependency review reports

* test(linear): smoke-load resolved SDK through CommonJS loader

* fix(deps): keep native rebuilds from reinstalling addon dependencies

* fix(native): invoke installed node-gyp directly for Node rebuilds

* test(cloud): exclude observer probes from row-lock timing budget

* test(mobile): preserve the CSS writer receiver in viewport spy

* test(native): remove obsolete batch-shim fixture exception

* Stream native rebuild output through the process wrapper
2026-10-02 05:05:43 -07:00
Neil 13ecf051c3 Reuse prepared Windows native builds in SSH CI (#24555)
* ci: reuse qualified Windows server slots for SSH host tests

* ci: reuse prepared relay addons after an exact native cache hit
2026-10-02 03:31:57 -07:00
Neil efbf651c7b Reduce CI setup costs and fixture failures (#24537)
* Let scheduled CI warmers wait and measure WebRTC startup

* Measure a smaller daemon shutdown fixture image

* Counterbalance WebRTC startup and verify retained fixture files

* Record CI fixture measurements and remove temporary pilots

* Clarify fixture build dependency cleanup evidence

* Make coalesced snapshot fixture delivery deterministic

* test: type the PTY write delay observer
2026-10-02 02:46:02 -07:00
Neil 6153fbcfe4 Reduce redundant headless server CI work (#24527)
* ci: avoid unrelated headless server qualification

* ci: skip headless detection for ineligible draft PRs

* ci: preserve cross-host qualification and skip supplied prerequisites

* ci: include Windows server cache validation in change detection
2026-10-02 01:42:41 -07:00
OrcaWinandm4air 5f308bfa9c revert: take the 26 Phase 3 (#16741 port) PRs back out of main (#24559)
* Revert "feat(orcad): source-side dormant export of a relay-hosted SSH target (#16741 T6-8) (#24519)"

This reverts commit 783101b304.

* Revert "feat(ssh): update, roll back, recover and stop a managed orcad server (#16741 T6-5 follow-up) (#24463)"

This reverts commit 38c2d1dcb9.

* Revert "feat(ssh): deploy and pair an empty managed orcad server over SSH (#16741 T6-5) (#24453)"

This reverts commit 8b76683b40.

* Revert "fix(ssh): orcad GC honors the activation journal; readiness requires proven daemon coverage (#16741 T6 follow-up) (#24451)"

This reverts commit d3f8c5063b.

* Revert "feat(ssh): remote orcad stop by request file and journaled decommission (#16741 T6-4) (#24449)"

This reverts commit 43d9b43d3f.

* Revert "feat(orcad): supervisable server: stop requests, managed stop receipts and a lifetime that keeps its lock on failed teardown (#16741 T6-3) (#24433)"

This reverts commit b093d3ab20.

* Revert "feat(ssh): crash-safe orcad activation, rollback and recovery (#16741 T6-2) (#24423)"

This reverts commit 1a9ac0e955.

* Revert "feat(runtime): SSH access links for paired servers in a downgrade-safe sidecar (#16741 T5-1+T5-2) (#24420)"

This reverts commit 99db2bfae4.

* Revert "feat(relay): capability-gated owner reset with a durable preparation journal (#16741 T3 R1) (#24418)"

This reverts commit 34a582bd39.

* Revert "feat(ssh): track connection-manager drains, test probes and provider continuations (#16741 T2 P3+P8a) (#24407)"

This reverts commit d53063d2b1.

* Revert "feat(daemon): idle retirement, session census and recovery-only provider (#16741 T2 P4b) (#24409)"

This reverts commit ff212dbbef.

* Revert "feat(ssh): add pty.resumeClient and split SSH PTY process listing (#16741 T2 P5+P6) (#24414)"

This reverts commit 92cb71765e.

* Revert "feat(relay): await owned watcher and agent children on shutdown (#16741 T2 P1) (#24400)"

This reverts commit 6b36e4f85b.

* Revert "feat(session): retry failed renderer session writes and verify local folder PTYs (#16741 T2 P9) (#24406)"

This reverts commit d23ecef301.

* Revert "feat(ssh): remote orcad primitives on the pinned Node runtime (#16741 T6-1) (#24419)"

This reverts commit dd87ae578d.

* Revert "fix(runtime): fence runtime-environment subscriptions and status probes by identity (#16741 T5-3) (#24421)"

This reverts commit ece9e4d2e3.

* Revert "feat(orcad): migration manifest and dormant-state contracts (#16741 T6-7) (#24422)"

This reverts commit 3fbdaba262.

* Revert "feat(ssh): wire SshConnection through the work and transport close ledgers (#16741 T2 P2) (#24401)"

This reverts commit 4e8edc8872.

* Revert "feat(profiles): carry markdown frontmatter visibility in project transfers (#16741 T2 P7) (#24405)"

This reverts commit 60c93263cc.

* Revert "fix(runtime): project the PTY incarnation onto mobile session tabs (#24413)"

This reverts commit 99e0303572.

* Revert "feat(daemon): tag daemon stream data with the PTY incarnation id (#16741 T2 P4a) (#24402)"

This reverts commit 817af768b0.

* Revert "feat(ssh): port the SSH connection work ledger and transport close ledger (#16741 T2) (#24210)"

This reverts commit c9918931c8.

* Revert "feat(relay): fence and drain file and git response streams on shutdown (#24185)"

This reverts commit dc08ffeba9.

* Revert "refactor(runtime-rpc): extract the Node WebSocket lifecycle; opt-in pinned port (#24186)"

This reverts commit a789233bbb.

* Revert "feat(relay): route relay handlers through work admission; producer publication drain (#24181)"

This reverts commit 0b812bd698.

* Revert "feat(relay): land the #16741 T1 seam (work drain, publication drain, release gate) (#24156)"

This reverts commit 3aa2d3af7c.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
2026-10-02 00:52:32 -07:00
Neil 8ff6296bc7 Speed up serializer checks and keep native caches stable (#24476)
* Reuse serializer oracle cells and isolate native cache policy

* Preserve native cache post-save paths and record hosted oracle gain

* Record native cache reuse and separate cancel-test startup budget
2026-10-01 21:43:56 -07:00
Brennan Benson 444f1952c7 ci: run every cross-version wire test, picked up by folder so new ones can't be skipped (#24499)
* ci(cross-version-wire): run the whole directory so no compatibility test is left out

Three cross-version tests ran in no CI job because the job named its files by hand.
Run the directory instead, ratchet that every file kept out of the unit shards
runs in some PR job, and re-run the job when the modules the newly running
tests guard change.

* test(cross-version): give the orchestration downgrade test its siblings' 120 s budget

* ci(unit-exclusion): count only merge-gating jobs, and require each excluded file's job to fire on it

The coverage check counted any pr.yml job, including e2e, terminal IME and Windows WSL, which are
left out of verify.needs and so cannot block a merge. It now reads verify.needs and the reusable
workflows those jobs call.

It also only proved that some step names each excluded file, not that the job runs when the file
changes. The structured-session zsh login-shell test runs only in shell_contracts, whose path
trigger matched neither it, its harness nor its subject, so a PR touching only those ran it
nowhere. The check now asserts a change to each excluded file fires a gating job that names it,
and the shell trigger gains those three paths.

* ci(cross-version-wire): trigger on the turn-outcome vocabulary and the schema version-skew resolver

A change confined to src/shared/agent-turn-outcome (the arms a newer host publishes) or to
orchestration-schema-version-skew (how current code reopens a downgraded database) skipped the
job whose tests guard exactly those contracts. Also corrects the publish/read direction in the
turn-end comment.

* test(cross-version): state why the orchestration downgrade test needs 120 s

* test(ci): glob the unit tree once for the unit-exclusion coverage checks
2026-10-01 20:53:51 -07:00
Neil f69052e113 Reuse qualified Windows server builds and dependency verification records (#24448) 2026-10-01 16:19:40 -07:00
OrcaWinandm4air 1a9ac0e955 feat(ssh): crash-safe orcad activation, rollback and recovery (#16741 T6-2) (#24423)
Journal every orcad activation and rollback under a host fence so an interrupted one recovers to exactly the slot the activation record names. D7: planOrcadUpdate and assessOrcadRollback refuse a restart whose incoming build cannot attach the live terminal daemon's protocol. POSIX-only and inert: no production caller.

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 12:34:11 -07:00
Neil 197ea3a3b3 Free PR CI capacity by avoiding repeated setup and real-time test waits (#24355)
* Reduce repeated PR setup and transcript timing waits; add hosted comparisons

* Align parallelism contract with Node-only external rebuild toolchain

* Record hosted coverage and launch package, store, and cancellation comparisons

* Apply hosted Windows setup savings and remove measured test waits

* Keep measured PR package gains and remove completed comparison jobs

* Report measured test counts with precise units
2026-10-01 11:51:43 -07:00
Zun 35c8887a76 fix(i18n): correct Korean working and shell labels (#24341) 2026-10-01 14:22:05 -04:00
OrcaWinandm4air 6d1a97ef98 fix(ssh): launch the Windows relay outside sshd's job so standard users work (#24224)
* fix(ssh): launch the Windows relay outside sshd's job without WMI

Win32-OpenSSH kills a session's job on close but allows breakaway. relay.js
gains a one-shot launcher mode that starts the detached relay with
CREATE_BREAKAWAY_FROM_JOB through the staged process-tree addon, so a standard
user no longer needs a WMI Remote Enable grant. WMI stays as the fallback for a
relay without the addon, and a refusal there is named. The Windows SSH-host
lanes drop their WMI grant and assert the breakaway route and adoption.

* fix(ssh): find runtime holds without WMI on a standard-user Windows host

The store GC read held runtimes through Get-CimInstance Win32_Process, which
WMI refuses to a standard user's SSH logon, so the pass kept every runtime.
On a refusal it now reads this account's own process image paths through
Get-Process.

* build(relay): ship the Windows relay launcher addon in every desktop package

macOS and Linux packages carried Windows relays without windows-process-tree.node,
so a legacy-runtime relay they uploaded to a Windows SSH host could not launch
outside sshd's job and fell back to WMI, which a standard user is refused.

A reusable Windows job now compiles the x64 and arm64 addons once and uploads
them; release-cut, release-mac-build, and the hourly/daily/adhoc mac builds
download them before build:release and require both arches. Staging now rejects
a binary with the wrong PE machine, the ReadProcessMemory import, or no
spawnOutsideJob export, so a stale pre-launcher build cannot ship.

* ci(ssh): run the Windows SSH-host lanes when the relay process-tree build scripts change

The staging and gyp-rebuild scripts decide which windows-process-tree addon the
relay ships, so a change to either must re-prove the Windows host cells.

* test(ci): find the mac orcad-template download by artifact name

The release mac job now also downloads the relay Windows process-tree addons, so
the first download-artifact step is no longer the template's.

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 05:32:09 -07:00
8afa1db50c feat(ssh): rung B glibc 2.17 compat runtime; gate remote vault on host node:sqlite (#24148)
* feat(ssh): wire rung B to the glibc 2.17 compat runtime; gate rung C vault on full node:sqlite

- COMPAT_RELAY_RUNTIMES lists linux-x64-glibc217; rung B plans the compat slot and compat
  pinned Node when glibc is below 2.28 or rung A refused with libc_floor/missing_lib.
- The relay version folds the compat runtime's executable hash; refusals are cached per runtime.
- The orcad template stages an optional linux-x64-glibc217 target (base package + compat
  node-pty slot + compat runtime marker); the verifier and materializer accept it.
- node-pty slot loader falls back to the compat slot when the default slot is missing or
  needs a newer glibc.
- Runtime store GC keeps the compat pin beside the default one on every relay connect.
- hasNodeSqliteReaderApi (DatabaseSync + backup) gates relay session search and the relay
  OpenCode reader, which now names the host Node version in its unavailable reason; the SSH
  vault reader installs the compat Node on old-glibc hosts and uploads nothing when no
  pinned Node can run.
- Rung D: a remembered noexec reports home_noexec and never advises installing Node.

* fix(ssh): re-prove a replayed noexec after rung D so allowing exec recovers the host

* fix(ssh): keep the rung B compat runtime pinned in the relay-connect store GC

* test(ssh): mock deployment-target facts in the Windows OpenCode runtime tests

* ci(ssh): build the glibc 2.17 compat slot for the hostile-host matrix; CentOS 7 lands on rung B

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>
2026-10-01 05:32:05 -07:00
OrcaWinandm4air f9940d5354 ci(ssh): macOS SSH-host lane for the pinned relay; fix uploads under a symlinked root (#24179)
* test(ssh): upload a root reached through a symlinked parent

The upload-root realpath fix landed with #24180; this keeps macoshost's case
where the root is passed explicitly beneath a symlinked parent.

* ci(ssh): macOS hostile-host lane on a loopback user-level sshd

Adds local-sshd cells for darwin-arm64 (macos-14) and darwin-x64
(macos-15-intel): a non-root sshd on 127.0.0.1 logs in as the runner user
with SetEnv PATH=<shims>:/usr/bin:/bin:/usr/sbin:/sbin and an empty HOME, so
no rc file restores Homebrew. The driver asserts rung A, terminal echo,
cached runtime reuse, GC keeping the in-use runtime, no toolchain or xattr
calls, and that the SFTP-uploaded Node carries no quarantine and runs as
uploaded. Docker cells are unchanged; each machine runs only cells it can host.

* test(ssh): fail a hostile-host run that would skip every named or hostable cell

A cell named for the wrong OS or arch was silently skipped, so a macOS job on a
mismatched runner went green having deployed nothing. Named cells must now be
hostable here, and a gated run must select at least one cell.

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 04:43:46 -07:00
OrcaWinandm4air 0ad77ea2f7 ci(ssh): Windows SSH-host lanes (inbox + preview OpenSSH) for the pinned relay (#24180)
* ci(ssh): import the private Windows OpenSSH provisioning harness

Copied unchanged from origin/OrcaWin/np-windows-ssh-provider-diagnostic
(config/ci/windows-ssh-provider/preview-ssh/ at 1242f3c4c8, commits 78b3857a0d,
c24adccff0, 069aa7b38b): a private LocalSystem sshd service on 127.0.0.1 for a
dedicated standard user, either the Microsoft-signed Win32-OpenSSH
10.0.0.0p2-Preview ZIP (archive and every binary pinned by sha256) or the inbox
OpenSSH.Server capability binaries. The following commits extend it for the
pinned-Node relay host lanes.

* test(ssh): run hostile-host cells through a host-agnostic driver

The Docker matrix drove the relay deploy and inspected the container with
inline docker exec calls, so no other host could reuse it. Split it into:

- ssh-hostile-host-test-harness.ts: the deploy, ladder observation, terminal
  echo (per-shell probe), runtime reuse and GC-keeps-in-use assertions, now
  also capturing every command the deploy sent the host.
- ssh-hostile-host-observer.ts: how a driver inspects the host outside SSH;
  docker exec for containers, the local filesystem for a loopback host.
- a legacy_opt_out outcome: the ladder never runs and nothing enters the
  pinned store, whatever the host-Node path does.

Launched cells now also check the runtime's sha256 on the host and that every
slot file (the Windows bundled ConPTY pair included) landed in the relay dir.

* ci(ssh): Windows SSH-host lanes for the pinned-Node relay

Phase 2 exit gate, Windows half: the real deployAndLaunchRelay through a real
SshConnection against Win32-OpenSSH on 127.0.0.1, on windows-2022 (x64) and
windows-11-arm (arm64), for both the inbox OpenSSH.Server capability and the
Microsoft-signed 10.0.0.0p2-Preview release (ZIP; archive and each binary
pinned by sha256 and Authenticode, as in the imported harness).

Builds on the provisioning harness from
origin/OrcaWin/np-windows-ssh-provider-diagnostic (previous import commit):
- one private standard account per cell, so every cell starts from an empty
  runtime store;
- -HiddenTools: the private sshd service's own Environment carries a PATH
  without any machine PATH entry holding node/npm/compilers, led by logging
  .cmd shims; a session probe fails the job if node.exe still resolves;
- DefaultShell set per cell by invoke-pinned-relay-cells.ps1 and restored at
  cleanup (dispatch proven per cell via %COMSPEC%).

Cells (src/main/ssh/ssh-windows-host-cells.ts): pinned-cmd (stock sshd),
pinned-powershell (DefaultShell = Windows PowerShell) expect rung A on the
pinned node.exe with the relay self-test passing, terminal echo, runtime
reuse, GC keeping the in-use runtime, stage identity through node.exe and no
Add-Type in any decoded session command; legacy-opt-out expects the ladder
never to run and an untouched pinned store.

* fix(ssh-ci): tolerate absent-drive PATH entries and retry Windows userData teardown

Join-Path throws on a machine PATH entry naming a drive the runner lacks,
which would abort provisioning before any cell ran; the toolchain split now
probes with [IO.File]::Exists over [IO.Path]::Combine, and the self-test
covers an absent drive. The hostile-host harness removes its throwaway
userData with removeTreeSync so a transient Windows lock cannot fail the
lane's afterAll.

* fix(ssh-ci): stop the account list rebinding the typed -Accounts param

PowerShell variable names are case-insensitive, so $accounts=[List[hashtable]] assigned into
the [int]$Accounts parameter and every Windows host job died before provisioning. Rename the
list and make the provisioning self-test reject script-scope assignments that shadow a param.

* fix(ssh-ci): hide the host toolchain by ACL, since sessions ignore the service PATH

Win32-OpenSSH builds a session's PATH from the machine and user registry values, so the private
service's Environment never reached SSH sessions and host node.exe stayed visible. Deny the private
accounts the toolchain PATH directories, put the logging shims on each account's own PATH, and
record failing sshd and client log lines so a refused login is diagnosable from the receipt.

* fix(ssh): resolve the upload root before checking entries stay inside it

uploadDirectory compared each entry's realpath against the root as given, so a root reached
through a symlink, junction or Windows 8.3 short name (C:\Users\RUNNER~1 in TEMP) rejected every
entry as escaped and the pinned runtime upload never started.

* fix(ssh-ci): fail cells on a vitest failure and give each account its own keys file

The cells script read $LASTEXITCODE under the workflow's GetNewClosure callback, which sees a
stale captured copy, so failed cells reported exit 0 and the job passed. Read the global value.
Inbox sshd 8.1 checks authorized_keys with read_ok=0, refusing a file other accounts can read;
use one keys file per account via %u.

* fix(ssh-ci): keep the account name in inbox mode and surface the WMI launch gap

The inbox binary-verification loop reused $name, so later SSH and SFTP probes logged in as
'sftp-server.exe'. Before the cells run, probe whether a standard SSH user can call WMI
Win32_Process.Create (the Windows relay launch path); when refused, warn and grant the cell
accounts Remote Enable on root\cimv2 for the run so the remaining assertions execute.

* test(ssh): keep the first terminal session answering keepalives through GC

The hostile-host driver disposed the first session's multiplexer before the GC and reconnect
steps, so a slow Windows GC let the relay reap the silent owner as 'local' and the reconnect then
waited out the full owner grace. Keep the session live until the connection closes, as the app
does, and resend the terminal probe until the shell evaluates it: ConPTY PowerShell can drop
typeahead sent before its first prompt.

* ci(ssh): keep each cell's relay logs in the receipts

* fix(relay): detach an ended socket client as peer-closed before destroying it

The listener destroyed a socket on 'end' but detached its client only on 'close'. A relay write
in that window failed with 'Relay socket is closed', and the dispatcher closed the client as
'local', so its PTY owner kept the full 30s grace instead of the peer-closed floor and a quick
reconnect was refused. The Windows host lanes logged this race on the named-pipe endpoint.

* test(relay): drive the peer-end listener test with a real dispatcher instead of a cast stub

The stub was an unchecked 'as unknown as RelayDispatcher' that failed the changed-code casting gate.

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 04:03:07 -07:00
OrcaWinandm4air 554f7f4ce5 feat(packaging): ship the orcad server template in desktop builds (#24155)
* build(orcad): merge per-runner prebuild slot trees into one matrix

Each node-server lane builds only its own node-pty slot. Release CI needs
their union before `build:orcad-prebuilds --require-slots` and the
template build can run; merge-orcad-prebuilds.mjs verifies every lane's
files against its own manifest, refuses duplicate slots and mismatched
node-pty/N-API/Node-header builds, then writes one merged manifest.

* build(orcad): keep agent-browser out of the desktop deployment template

The template rides inside every desktop build (design D2). Seven ~10 MB
agent-browser binaries would be ~76 MB, more than the rest of the template;
design D2's package contents never listed it, and a slot without one
already reports no headless browser. ORCAD_OMIT_AGENT_BROWSER=1 skips the
copy; standalone build:orcad still includes it.

* feat(packaging): ship the orcad deployment template in desktop builds

Design D2: the server JS and every target's addons ship inside the app,
as out/relay does; the ~120 MB Node runtimes stay excluded and are
downloaded on demand. electron-builder copies out/orcad-template to
Resources/orcad-template on every desktop OS, which is the first path
materializeOrcadArtifact tries (process.resourcesPath).

Platform signing rewrites native bytes the template manifest hashes:
- macOS: the tree is signIgnored (codesign rejects its ELF/PE payloads);
  afterPack signs the darwin targets' Mach-O files with the app identity,
  as notarization requires, then reseals only those manifest entries.
- Windows: SignPath signs after packaging, so release CI reseals from the
  inner-signing list (packaged-orcad-template.cjs --reseal-signed).
Every other file must still match the build's hashes; afterPack verifies.

ORCA_REQUIRE_ORCAD_TEMPLATE=1 makes a missing template fail beforePack and
afterPack; without it a build ships none and SSH relays keep the legacy
path. verify-packaged-orcad-template.test.mjs's "unused, excluded"
contract is reversed on purpose.

* ci(release): build the orcad template from qualified lanes and package it

node-server-tests.yml becomes callable with a ref and build_template.
With build_template, each lane that owns a release slot (macOS, Windows,
the glibc 2.28 and Alpine lanes, and the glibc 2.17 compat lane) uploads
its qualified out/orcad-prebuilds, the Windows lane also uploads both
process-table addons, and desktop_template merges them, gates the full
matrix plus the compat slot with --require-slots, runs
build:orcad-template and uploads the orcad-template artifact.

release-cut calls it at the release tag beside the other gates. The
build and build-mac jobs wait for it, download it into out/orcad-template
(the mac workflow from the parent run), and require it via
ORCA_REQUIRE_ORCAD_TEMPLATE. The Windows signing staging skips the
template's Linux/macOS payloads, and a reseal step records SignPath's
bytes before the installer rebuild. A template-scoped concurrency group
keeps a release call and main's push runs from cancelling each other.

* test(orcad): keep the packaged-lookup imports clear of the compat-slot import edits

* ci(orcad): let a rerun lane replace its template artifacts

upload-artifact v4 refuses a second upload under an existing name in the same
run, so rerunning a flaky node-server lane during a release would fail at the
upload instead of re-qualifying the slot.

* ci(node-server): build the template's Windows addons before the lane switches to Node 18

The addon build script imports TypeScript, which Node 18 cannot load, so every
build_template run (release-cut included) failed on windows-2022.

* fix(build): ship the orcad template's shared node_modules

electron-builder's extraResources filter always drops the root node_modules of
a source directory, so packaged apps lost orcad-template/node_modules and the
afterPack verify failed. Copy it through its own resource entry.

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 04:01:26 -07:00
14d4bb2e2a fix(ssh): Windows hosts without Add-Type staging; runtime-store GC on Windows (#24149)
* fix(ssh): collect the pinned-Node runtime store on Windows hosts

Windows SSH hosts now run runtime-store GC instead of skipping it: one
PowerShell inventory reads .runtime-ref-node-<sha> and .runtime-node refs from
every version dir, and one Get-CimInstance Win32_Process query filtered on an
image path under runtimes\ adds process holds (never by image name; a failed
query keeps everything). Stale upload stages are swept with the same rule as
POSIX. Promotion and the post-upload hold check now take the store lock on
Windows too, and the lock's own commands run unwrapped there.

Windows relay version-dir liveness now honours .relay-pid (design D5): a live
PID answers ALIVE before any pipe is touched, a dead one (ESRCH) plus refusing
pipes is exited, anything else is unverifiable. The runtime probe adopts a
pinned node.exe an earlier vault reader left without a .verified marker after
running it.

* fix(ssh): Windows stage fencing and vault runtime go through the verified node.exe

Upload-stage file identity on Windows no longer compiles an Add-Type P/Invoke
helper when the relay runs on Orca's verified pinned node.exe: the stage
commands run a fixed fs.lstatSync(..., {bigint:true}) script through it. It
prints the legacy helper's vol:high:low lowercase hex, and identity files are
compared after normalising hex spelling, so old and new clients recover each
other's stages. Host-Node relays keep the legacy helper; the choice is
documented in windows-edr-posture.md.

The Windows OpenCode vault reader now installs the pinned runtime through
ensureRemoteOrcadNodeRuntime (official zip, host-side extraction, .verified,
store lock) instead of uploading a client-extracted node.exe, and the relay dir
gains a .runtime-ref-node-<sha> so store GC keeps the runtime the vault uses.

* test(ssh): run the Windows stage-identity and store-GC tests on the Windows lane

The legacy/node.exe identity compatibility test and the Win32_Process hold path
were gated to win32 but no CI lane ran them. Add both files to the Windows
package lane and a real running-node.exe hold test.

* test(ssh): tear down Windows-lane temp trees through removeTreeSync

* test(ssh): grant the store lock to the Windows OpenCode runtime setup test

The Windows promote now runs under runtimes/.store-lock, so the mocked host
must answer the lock's CreateNew step.

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>
2026-10-01 03:25:42 -07:00
OrcaWinandm4air 6aed05471c ci(ssh): hostile-host matrix for the relay runtime ladder (#24146)
* fix(ssh): classify a musl host missing libstdc++ as missing_lib, not wrong_libc

musl's loader follows each missing-library line with one 'Error relocating ... symbol
not found' per unresolved symbol, and the relocation pattern was checked first. Check
missing libraries before relocation errors; the ld-linux/ld-musl interpreter case stays
wrong_libc.

* build(orcad): allow a partial deployment template for CI

build-orcad-template --targets a,b builds and verifies only the named slots, so a CI job
that can fill just the x64 Linux prebuild slots can still materialize rung A/C addons.
Without the flag every target is still built and verified.

* ci(ssh): hostile-host matrix for the relay runtime ladder

Drives the real client-side relay deploy against Docker sshd targets and asserts the
design D6 rung each lands on: Debian 10 and AlmaLinux 8 (glibc 2.28) and Alpine (musl)
on rung A; Alpine without libstdc++ refused missing_lib down to D; Ubuntu 22.04 with a
host Node 20 and a noexec home straight to D (home_noexec); CentOS 7 (glibc 2.17)
refused libc_floor at A and C, falling to a host-npm path with no Node; and a
no-egress Debian 10 still on rung A. Launched cells also prove the terminal echoes,
no npm or compiler ran, a second connect reuses the uploaded runtime, and runtime GC
keeps the in-use runtime while collecting an idle one.

New workflow ssh-hostile-hosts.yml runs on dispatch and on path-filtered PRs.

* test(ci): pin the hostile-host workflow to the headless-server builder images

The matrix builds its runtime slots in copies of the node-server lanes' Alpine
and manylinux images; this contract fails when NODE_RUNTIME_PIN or either
builder digest moves in one workflow and not the other.

* test(ssh): reconnect as the same client and retry a grace-held PTY owner in the hostile-host matrix

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 03:24:17 -07:00
Neil bd90da7a5b ci: share PR planning setup and reuse the static native cache (#24329) 2026-10-01 02:46:18 -07:00
OrcaWinandm4air 53fd2dea0b feat(ssh): relay runtime fallback ladder, telemetry and host runtime setting (#24133)
* feat(ssh): complete the relay runtime fallback ladder (D6 rungs B slot, C, D)

Rung C runs the relay on the host's Node >= 18 with Orca's prebuilt N-API
addons and no npm (addon-only probe mode). Rung B is a data-driven slot chosen
only when a compat runtime is listed. Rung D fails the connect with a
classified reason carried as a TerminalUnavailableCause. The ladder steps
down only on classified refusals; unanswered probes throw. The rung decision
is persisted per host keyed by (glibc, runtime hash, Orca major), and
ssh_remote_runtime_resolved reports it once per host per session.

* feat(settings): SSH host runtime choice (Auto | Orca-managed Node | Host Node)

* docs(telemetry): describe ssh_remote_runtime_resolved

* fix(ssh): let a passing rung C disprove a remembered noexec; allow glibc-less compat runtimes

A remembered rung A noexec was re-persisted even after rung C self-tested addons from the same
~/.orca-remote tree, so rung A stayed skipped until the key changed. Rung B's evaluator also could
never match a musl compat runtime.

* test(ssh): import node:fs once in the host-node addon test

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 02:08:26 -07:00
OrcaWinandm4air ddd4927a0b build(orcad): server node-pty slots at glibc 2.28, plus a glibc 2.17 compat slot (#24134)
* build(orcad): build server glibc slots on glibc 2.28 and add the glibc 2.17 compat slot

Design D6: the default linux-{x64,arm64}-glibc node-pty slots now build in
manylinux_2_28 (digest-pinned) and are gated at glibc 2.28 / GLIBCXX_3.4.25
through a floor profile on verify-linux-glibc-floor.cjs; the desktop keeps
its Ubuntu 20.04 (2.31) default.

Adds the opt-in linux-x64-glibc217 compat target: NODE_RUNTIME_COMPAT_ASSETS
pins the unofficial glibc-217 Node (update/check pin scripts cover it, outside
SERVER_TARGETS), and a new CI lane builds the compat slot in manylinux2014
with static libstdc++, gates it at glibc 2.17 with no shared C++ runtime in
DT_NEEDED, and smokes it under the glibc-217 Node.

* refactor(node-runtime-pin): route compat lookups through isCompatServerTarget; keep the glibc doc's slot-name paragraph intact

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 01:10:57 -07:00
OrcaWinandm4air 8c2cd7d331 feat(ai-vault): read remote OpenCode history with the pinned Node; remove Bun (#24128)
* feat(ai-vault): read OpenCode history with the pinned Node instead of Bun

SSH hosts whose Node lacks node:sqlite (or its backup(), which 22.13-22.15
omit) now get the pinned Node in the shared ~/.orca-remote/runtimes/node-<sha>
store orcad uses: POSIX hosts receive the official archive and extract and
hash-verify it on the host; Windows hosts receive the verified node.exe the
client extracted, promoted by host Node with the same hash check. WSL distros
use the same layout and checks under ~/.cache/orca/runtimes/.

The Bun release pin table and its materializer are deleted. Old relays keep
reading their vault-sqlite/<sha>/bun references; nothing deletes those files.
An unconfirmed runtime upload now keeps its stage instead of removing it.

* refactor(sqlite): drop the Bun SQLite adapter; node:sqlite is the only backend

Nothing outside Electron runs on Bun any more (design D4), so SyncDatabase
loses its Bun branch, and bun-sqlite-database, bun-sqlite-statement and
bun-readonly-wal go, with the relay's bun:sqlite external. The profile-state
backup worker admits Electron or an entry that exists, and startup errors
name the pinned Node. The D7 cross-runtime gate still runs Bun 1.4.2, now
reaching Bun's SQLite through its node:sqlite.

* test(native-chat): drop the Bun SQLite driver case now that node:sqlite is the only backend

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 01:10:31 -07:00
OrcaWinandm4air 6593d7d194 feat(orcad): run orcad on the pinned Node instead of Bun (#24110)
* ci(daemon): gate PRs on daemon protocol crossing from the newest release

Lands daemon-protocol-facts.mjs from the Windows update diagnostic branch with a
stricter parser, and adds check-daemon-protocol-crossing.mjs (rule R1): the working
tree must attach the newest release tag's daemon. Rollback crossing is reported only.
Runs in the cross-version-wire job, which already has full tags; tag selection moves
to config/scripts/stable-release-tags.mjs so both use one rule.

* feat(persistence): run profile backups in the worker whenever its entry is bundled

* refactor(orcad): make profile and native preflight runtime-neutral

The profile preflight parser now takes the expected runtime identity from the
caller (shipped callers pass the pinned Bun identity), and the native
preflight is renamed to orcad-runtime-native-preflight with neutral wording.

* feat(runtime): pin the Node 24.21.0 server runtime with an offline CI check

Add src/shared/node-runtime-pin.ts (NODE_RUNTIME_PIN, SERVER_TARGETS,
NODE_RUNTIME_ASSETS for all 8 server targets plus the headers tarball),
generated by config/scripts/update-node-runtime-pin.mjs from the nodejs.org
and unofficial-builds SHASUMS. check-node-runtime-pin.mjs verifies, with no
network, that the pin tracks the locked Electron, matches engines.node's
major, and covers exactly SERVER_TARGETS; it runs in the static analysis job.

ORCAD_BUN_TARGETS consumers now read SERVER_TARGETS so there is one target
list; orcad's Bun runtime and build output are unchanged.

* test(persistence): skip plain-Node backup selection tests in the Bun profile suite

* fix(runtime): reject a pinned archive that belongs to another target

* ci(daemon): fail PRs that swap a runtime launcher and bump the daemon protocol

D7.1 R3: hosting orcad or the daemon on another runtime is not a protocol change,
so one PR must not do both. The launcher file list lives in the check script; the
allow-runtime-launcher-protocol-bump label overrides it.

* feat(orcad): select pinned-Node slots by a .runtime-node marker

D7.1 R5: a Node slot names its shared runtimes/node-<sha256>/node through
.runtime-node instead of .build-target, so Bun-era clients read it as a legacy
slot rather than exiting 78 on a missing bun-runtime. Nothing builds the marker yet.

* fix(runtime): load the Node pin without the typeless-module warning

check-node-runtime-pin.mjs now requires the pin and takes nodeDistArchiveName from
its own module, so it no longer loads the update script's build graph.

* fix(orcad): resolve Node slots to the design's runtimes/node-<sha>/bin/node layout

* feat(orcad): 8-slot node-pty prebuilds against the pinned Node headers at N-API 8

- build-orcad-prebuilds.mjs adds win32-x64/arm64 (conpty.node, the vendored
  conpty.dll/OpenConsole.exe, upstream's N-API conpty_console_list.node), compiles
  in a scratch copy against the hash-verified pinned headers (node.lib pinned per
  Windows arch) with NAPI_VERSION=8, rejects post-8 node_api_* imports, and writes
  a schema 2 manifest with per-file sha256, N-API level and the glibc need.
- --require-slots [slots] verifies files against hashes; --smoke loads the slot
  under the pinned Node and spawns a PTY; --print-slot names the host slot.
- The slot installer gates on N-API, libc, arch, glibc and file hashes instead of
  the exact NODE_MODULE_VERSION, and installs nested files (conpty/).
- bun-profile-tests.yml builds, verifies and smokes each runner's slot.

* fix(orcad): scope node-pty's glibc .symver pins to glibc on musl prebuild slots

musl's unversioned libc cannot satisfy openpty@GLIBC_* references at link
time, so the Alpine slot compile would fail. Pin the staged pty.cc guard to
__GLIBC__ and assert both musl transforms against the installed patch.

* feat(orcad): run orcad on the pinned Node instead of Bun

A packaged orcad slot now references the pinned Node 24.21.0 by its
executableSha256 (`.runtime-node`, `.server-target`) instead of carrying
bun-runtime, and ships node-pty from the slot's prebuild, only its own
ripgrep, and no Windows Bun PTY gate. The runtime lives beside the slots
at runtimes/node-<sha>/bin/node (node.exe on Windows, upstream name).

- build:orcad (build-orcad-node.mjs) builds the host slot's prebuild when
  missing and places the pinned runtime; the template is schema 3 with
  per-target files.
- handoffToBundledOrcad() resolves the slot's runtime reference and checks
  process.versions.node against the pin; a host Node >= 18 still hands off.
  Startup preflight keys on running as that runtime; callers expect 'node'.
- orcad and its daemon use node-pty (ConPTY + windows-pty-job on Windows);
  the Bun PTY sources, gate entry and canUseBunPty branches are removed.
- SSH deploy uploads the official archive once per pin, extracts and
  hash-checks it on the host, and self-tests it before publishing. Bun
  slots stay launchable for rollback; Node slots never use host Node.
- The runtime materializer is generic over pinned assets; the Bun wrapper
  remains only for the OpenCode vault reader (design Phase 2).
- Cross-runtime test: a profile DB written by Bun 1.4.2 (WAL left by
  SIGKILL) opens and backs up under the pinned Node, and the reverse.

No daemon PROTOCOL_VERSION change (design D7.1 R3).

* docs(ci): name the headless lanes after the pinned Node, drop Bun shard timings

Design D10: ci-demand-rollout.md and ci-runner-efficiency.md follow the
bun-profile-tests.yml -> node-server-tests.yml rename; shard timings drop the
deleted Bun PTY tests and follow the renamed ones.

* chore(ci): count the runtime archive download as a runtime launcher path

* fix(orcad): pin the macOS C++ standard for node-pty prebuilds

The official Node headers' config.gypi sets clang: 0, so common.gypi skips its
gnu++20 xcode_settings and Apple clang 15 (macos-14 runners) compiles
node-addon-api as C++98.

* fix(orcad): resolve the preflight's slot through realpath, as the handoff does

A symlinked orcad.js handed off to its real slot's pinned Node, but the
startup and profile preflights read the symlink's directory, found no
runtime marker there, and silently skipped the readiness check.

* refactor(ssh): drop materializeCachedNodeRuntime, which nothing calls

Deploys upload the verified official archive (design D5); no client path
needs an extracted Node executable cached by digest.

* test(orcad): gate the Bun-to-Node upgrade and Node-to-Bun rollback with live terminals

Design D7.1 R1/R3/R4 and D7.2. The last Bun orcad and this checkout's Node
slot are installed side by side under ~/.orca-remote, launched and stopped
with the client's own deploy commands, and share one data root. Each
direction proves the incoming orcad adopts the outgoing runtime's daemon
(same PID, same shell, output continues), opens its profile database and
backs it up with its own shipped worker, and that GC keeps the slot the
live daemon was forked from.

The node-server Linux lanes provide Bun 1.4.2 and build that Bun orcad
from main, and run with --cross-runtime. --artifact and --cross-runtime
now make their tests fail on a missing input instead of skipping.

* ci(node-server): pin node:24.21.0-alpine by its multi-arch index digest

* test(ssh): name the runtime archive fixture after its role

* test(node-server): load node-pty from the packaged slot in artifact runs

The node-server lane installs dependencies without building node-pty, and
Linux has no upstream prebuild, so the real-PTY failed-I/O teardown test
(picked up by the pty-subprocess selector) could not load pty.node. In
--artifact runs, alias node-pty to out/orcad's shipped slot so the test
exercises the addon orcad actually runs under the pinned Node.

* fix(orcad): let the Windows profile preflight exit after its PTY probe

On Windows, node-pty keeps the conout worker thread and pseudoconsole alive
until kill(), even after the shell exits. The PTY health probe never killed a
cleanly exited probe, so the packaged preflight printed its readiness line
and then hung until the build's 30s timeout, reported with an empty stderr.

- The probe kills its PTY on Windows after exit and uses the bundled ConPTY
  the daemon spawns with.
- The preflight exits once stdout is flushed; its owner reads to EOF.
- Preflight failures now report code, signal, timeout, stdout and stderr.

* test(node-server): load the slot's node-pty in the real-PTY test, not by alias

A vite alias redirected only ESM imports of node-pty; windows-pty-job and
local-pty-utils resolve it through require, so Windows loaded two conpty.node
copies and the Git Bash job-membership proof read an empty job. The failed-I/O
teardown test now loads node-pty through a fixture that picks the packaged slot
in artifact lanes.

The pty-subprocess selector was a prefix that also pulled in its POSIX-host
sibling unit tests, which pr.yml runs and which were never qualified on
Windows. Select the directory plus the two sibling files that belong here.

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 00:39:00 -07:00
OrcaWinandm4air d2dfc79764 ci(daemon): runtime-launcher protocol ratchet and Node slot marker (#24108)
* ci(daemon): gate PRs on daemon protocol crossing from the newest release

Lands daemon-protocol-facts.mjs from the Windows update diagnostic branch with a
stricter parser, and adds check-daemon-protocol-crossing.mjs (rule R1): the working
tree must attach the newest release tag's daemon. Rollback crossing is reported only.
Runs in the cross-version-wire job, which already has full tags; tag selection moves
to config/scripts/stable-release-tags.mjs so both use one rule.

* feat(persistence): run profile backups in the worker whenever its entry is bundled

* refactor(orcad): make profile and native preflight runtime-neutral

The profile preflight parser now takes the expected runtime identity from the
caller (shipped callers pass the pinned Bun identity), and the native
preflight is renamed to orcad-runtime-native-preflight with neutral wording.

* feat(runtime): pin the Node 24.21.0 server runtime with an offline CI check

Add src/shared/node-runtime-pin.ts (NODE_RUNTIME_PIN, SERVER_TARGETS,
NODE_RUNTIME_ASSETS for all 8 server targets plus the headers tarball),
generated by config/scripts/update-node-runtime-pin.mjs from the nodejs.org
and unofficial-builds SHASUMS. check-node-runtime-pin.mjs verifies, with no
network, that the pin tracks the locked Electron, matches engines.node's
major, and covers exactly SERVER_TARGETS; it runs in the static analysis job.

ORCAD_BUN_TARGETS consumers now read SERVER_TARGETS so there is one target
list; orcad's Bun runtime and build output are unchanged.

* test(persistence): skip plain-Node backup selection tests in the Bun profile suite

* fix(runtime): reject a pinned archive that belongs to another target

* ci(daemon): fail PRs that swap a runtime launcher and bump the daemon protocol

D7.1 R3: hosting orcad or the daemon on another runtime is not a protocol change,
so one PR must not do both. The launcher file list lives in the check script; the
allow-runtime-launcher-protocol-bump label overrides it.

* feat(orcad): select pinned-Node slots by a .runtime-node marker

D7.1 R5: a Node slot names its shared runtimes/node-<sha256>/node through
.runtime-node instead of .build-target, so Bun-era clients read it as a legacy
slot rather than exiting 78 on a missing bun-runtime. Nothing builds the marker yet.

* fix(runtime): load the Node pin without the typeless-module warning

check-node-runtime-pin.mjs now requires the pin and takes nodeDistArchiveName from
its own module, so it no longer loads the update script's build graph.

* fix(orcad): resolve Node slots to the design's runtimes/node-<sha>/bin/node layout

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 00:10:57 -07:00
OrcaWinandm4air 2a83c9536f ci(daemon): gate PRs on daemon protocol crossing from the newest release (#24089)
Lands daemon-protocol-facts.mjs from the Windows update diagnostic branch with a
stricter parser, and adds check-daemon-protocol-crossing.mjs (rule R1): the working
tree must attach the newest release tag's daemon. Rollback crossing is reported only.
Runs in the cross-version-wire job, which already has full tags; tag selection moves
to config/scripts/stable-release-tags.mjs so both use one rule.

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-30 23:23:28 -07:00
OrcaWinandm4air 49a83deaef refactor(orcad): make profile backup and preflight runtime-neutral (#24088)
* feat(persistence): run profile backups in the worker whenever its entry is bundled

* refactor(orcad): make profile and native preflight runtime-neutral

The profile preflight parser now takes the expected runtime identity from the
caller (shipped callers pass the pinned Bun identity), and the native
preflight is renamed to orcad-runtime-native-preflight with neutral wording.

* test(persistence): skip plain-Node backup selection tests in the Bun profile suite

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-30 23:23:23 -07:00
OrcaWinandm4air 3135fbbf49 feat(runtime): pin the Node 24.21.0 server runtime with an offline CI check (#24087)
* feat(runtime): pin the Node 24.21.0 server runtime with an offline CI check

Add src/shared/node-runtime-pin.ts (NODE_RUNTIME_PIN, SERVER_TARGETS,
NODE_RUNTIME_ASSETS for all 8 server targets plus the headers tarball),
generated by config/scripts/update-node-runtime-pin.mjs from the nodejs.org
and unofficial-builds SHASUMS. check-node-runtime-pin.mjs verifies, with no
network, that the pin tracks the locked Electron, matches engines.node's
major, and covers exactly SERVER_TARGETS; it runs in the static analysis job.

ORCAD_BUN_TARGETS consumers now read SERVER_TARGETS so there is one target
list; orcad's Bun runtime and build output are unchanged.

* fix(runtime): reject a pinned archive that belongs to another target

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-30 22:57:10 -07:00
Jinwoo Hong 1762a138f7 feat(mobile): slide the page's host stack on push and Back (#24268)
expo-router's Stack on web renders native-stack's web view, which flips display and ignores animation. The page's host stack now keeps expo-router's StackRouter under its public Navigator and draws the slide with the Web Animations API; a popped screen stays mounted until it has slid out. Native is a pure move.
2026-10-01 01:19:35 -04:00
Jinwoo Hong e9ec63168f fix(mobile): hold-to-dictate, repeat keys and the browser long-press survive the page's long-press (#24277)
On the OTA page a held press died ~500 ms in: the WebView's long-press selected nearby text and that selection's selectionchange/touchcancel ended the press. Page text is now unselectable unless it opts in (as native), hold surfaces declare onLongPress, the browser pane refuses contextmenu termination, and the chat mic's swapped icons no longer steal the touch target.
2026-10-01 01:18:24 -04:00
Brennan Benson 12b8ef8c0b fix(worktree): update local main safely, once per branch, alongside the checkout (#23698)
* fix(worktree): retry local main refresh through git lock contention and skip false alarms

* fix(worktree): overlap the local main refresh with the checkout and run one refresh per repo at a time

* fix(worktree): skip the local base refresh when the create makes that branch itself

Creating a workspace named feature-x from origin/feature-x runs `worktree add -b feature-x`,
which now overlaps the refresh. The refresh's drift probe could see refs/heads/feature-x
missing and its presence probe then see it (the add just wrote it), which reported
"not fast-forward" and showed a sticky "Local feature-x was not refreshed" warning.
`-b` refuses an existing branch, so there is nothing to refresh in that case: skip it on
the local, prepared-checkout and SSH create paths.

The SSH overlap tests move to their own file so the existing suite stays under the line limit.

* fix(worktree): say plainly what happens after the local base refresh queue wait expires

* test(worktree): prove SSH local base refreshes of one repo run one at a time

* test(worktree): drop type assertions from the SSH refresh overlap test mocks

* fix(worktree): fast-forward local main with one host-owned merge --ff-only per branch

Moves the whole local base refresh into one shared routine that runs on the
execution host (main process for local and WSL repos, the relay for SSH), so
the app no longer keeps a second copy of the checks, queue and retry.

A checked-out branch now moves with merge --ff-only (hooks, auto-gc and
autostash off) instead of status-then-reset --hard, which silently overwrote an
untracked file the new commit adds and could discard an edit or a commit made
after the check. A free branch moves with a compare-and-swap update-ref that
writes a reflog message. Status reads no longer take index.lock.

Creates of one branch share one run plus at most one trailing run; a create
waits at most 30 s and never starts a competing mutation. The failure toast is
keyed by repo and branch because every create that joined a run reports the
same fact.

* fix(worktree): fast-forward local main even when the repo requires signed merges

With merge.verifySignatures=true, the owner-checkout fast-forward refused an
unsigned origin/main tip, so every create warned "Local main was not
refreshed" where the old reset moved main. The new workspace is already
created from that same unsigned commit, and a branch that is not checked out
moves without a signature check, so the refusal protected nothing. Turn the
setting off for this one merge, like the hooks, gc and autostash overrides.

* fix(worktree): clear git read caches when a shared local main update lands late

The update of local main can finish after a create stopped waiting for it, so
the shared run now invalidates git read caches itself. The index.lock real-git
test also no longer reads the developer's global git config.

* fix(worktree): never overwrite an ignored file when fast-forwarding local main

A plain `git merge --ff-only` silently replaces an ignored file (for example a
local `.env`) at a path the new commit starts tracking. Pass
`--no-overwrite-ignore` so git refuses instead and the create reports the
checkout as having local changes. Supported on the fast-forward path since
well before Git 2.25.

Also make the relay test for one-refresh-per-branch hold the first merge until
the second request has reached the relay, so it fails without the coalescing.

* fix(worktree): keep the local main update a plain fast-forward whatever the user's merge settings say

A per-branch mergeOptions such as '-s ours' or '--squash', or pull.twohead=ours,
made the update create a merge commit that dropped upstream, or stage upstream
without moving main, while reporting success. The command now clears the
branch's mergeOptions and passes the strategy and signature choice on the
command line, which beats any config. After the move Orca confirms local main
is exactly the target before reporting it updated. The exact command also runs
in the Git 2.25 compatibility suite.

* fix(worktree): make the Git 2.25 fast-forward contract pass in CI and rerun on every change to it

The new real-Git contract for the local main fast-forward wrote a post-merge hook into .git/hooks, which does not exist when the repo is created by the uninstalled Git 2.25.5 build CI uses (no templates), so the Git compatibility check failed. Create the directory first.

The Git compatibility check also did not run when only the fast-forward module changed, so a later edit to its merge arguments (for example a flag Git 2.25 lacks) would skip the one check that tests them. Add the module to the check's paths.

* fix(worktree): answer every create from a local main update toward its own base

Creates from different remotes' main (origin/main and upstream/main) shared one queued update
per repo and branch, which ran only the latest caller's target: a create could get no result
for its own base, or a false "not refreshed" warning computed for another remote's main.

The per-branch runner now queues one run per distinct target, still one at a time per branch,
and only callers toward the same target share a queued run. Applied in the app and the relay.

* test(worktree): record the third create's result in the mixed-remote burst tests and update the toast id rationale

* chore(worktree): correct the toast id rationale
2026-09-30 18:30:22 -07:00
Brennan Benson 6f2a7d05c9 fix(worktrees): let git delete removed checkouts so chat sends never wait behind them (#23837)
* fix(worktrees): delete removed checkouts in git, not in Orca's file pool

Local worktree removal renamed the checkout into a sibling trash root and
deleted it in the background with a recursive fs.rm in the main process.
That queued one request per entry on libuv's shared 4-thread file pool, so
for minutes every other async fs call in the main process (the agent-session
store behind chat sends, file explorer reads) waited behind the delete.

`git worktree remove` now deletes the checkout inline in git's own process
again, so the card stays in its Deleting state for the length of the delete
while Orca's file pool stays free. No timeout applies to the call, so a
large delete is never killed halfway.

If git reports success but the path still exists (Git for Windows leaves
junctions and their parent directories in place), the leftover is deleted
with the existing removeHostTree; WSL checkouts stay with the distro.

Nothing creates trash any more: the scheduling queue, rename/restore
helpers and the trash_rename span are gone. The startup sweep stays to
drain entries older releases left behind, and now removes each emptied
trash root so the obligation ends.

* fix(worktrees): let Git delete Windows checkouts with long paths enabled

Removal now always runs Git's own recursive delete, and worktree creation
checks out with core.longpaths on Windows, so a deep checkout Orca created
could fail to delete with "Filename too long" (#6433). The Windows recovery
then finishes the delete but keeps the branch. Pass the same command-scoped
core.longpaths option to `git worktree remove` so Git can delete what it
created.

Also point the CI shard timing entry at the renamed real-git removal suite.

* fix(worktrees): keep an inherited GIT_ASK_YESNO out of the worktree delete

Git for Windows asks $GIT_ASK_YESNO whether to retry when a file stays
locked during a recursive delete. Orca's git env inherits the user's
environment, so an inherited value would run an arbitrary prompt program
in the middle of a removal. Drop it for the removal call only.

* perf(worktrees): run worktree deletes under their own limit, outside git admission

`git worktree remove` now deletes the whole checkout in Git's own process,
which takes 20-35 s on a large tree. It took a general git admission slot at
status tier for that whole time, and that cap is as small as two slots on a
machine with six or fewer cores, so two deletes blocked every status read.

Deletes now skip general admission and queue under their own limit of two
per host instead: two concurrent deletes already saturate one disk, and more
only slow each other down. Leftover cleanup runs inside the same slot.

* fix(worktrees): delete removed checkouts in the background and mark them removing

Since the checkout is deleted by `git worktree remove` in Git's own process,
a large delete takes 20-35 s. Answering the request only after that made web
and mobile (30 s), paired desktop (60/180 s) and the CLI (60 s) report a
failure for a delete that was still going, and mobile silently re-showed the
row.

The request now does everything that can refuse (lock, cleanliness, archive
hook, watcher/terminal gate, terminal stop, shared-link unlink), records the
removal in an in-memory table on the host and answers `removing: true`. The
delete, branch cleanup and metadata purge run after it in the same order as
before, and the watcher/terminal gate stays held until they finish.

- Listings mark rows in the table `removing` for clients that advertise
  `worktree.background-removal.v1` (the desktop renderer, paired desktop and
  web), and leave them out for everyone else (older clients, mobile, the
  CLI), which already dropped the row when the request answered.
- The outcome (removed, with any preserved branch, or the error) rides the
  existing worktrees-changed event as an optional field, sent after the row
  has left the table.
- A repeat delete while Git runs joins it. A create at the same path or with
  the same branch is refused with "Cleanup is pending; try again shortly";
  create's name search skips the path, so generated names move on.
- Nothing is persisted: after a quit or crash Git still lists the checkout
  and it can be deleted again. WSL checkouts still delete inline.
- `orca worktree rm` says the checkout is still being deleted.

* fix(worktrees): keep the existing Deleting card until the host's Git finishes

The host now answers a local worktree delete on acceptance and deletes in the
background. The renderer keeps the existing delete state set until the host
publishes how it ended:

- The delete that asked waits for the outcome on the worktrees-changed event
  (local IPC or the paired runtime's client event), then runs the same
  teardown, preserved-branch toast and card error an inline delete did. If
  that event is lost to a dropped connection, a listing that shows the row
  gone after it was marked removing finishes the wait, and one that shows it
  back without the marker fails it.
- Any other renderer (a reload, a paired desktop, web) sets the same delete
  state from the host's `removing` marker and clears it when the marker goes.
  A failure the host publishes lands on that card's existing error.
- Web advertises `worktree.background-removal.v1` so the host sends it the
  marker; paired desktop does through the Electron capability list.

No new component, style or state: the card reads the delete state it always
did. A host that predates this answers when done without `removing`, and the
renderer takes that as finished, as before.

* test(worktrees): type the removal harness and projection for the node typecheck

* fix(worktrees): don't fail a delete retry with an earlier attempt's buffered failure

A background removal's outcome that reached this renderer with no waiter (another client's
delete, a host-marked card, or one already settled from listings) was buffered for 60 s and
consumed by the next delete of the same workspace, so retrying a failed delete failed at once
with the old error while the host was deleting. Drop the buffered outcome before sending the
request; only an outcome that arrives after it can belong to it.

* fix(worktrees): let only a gap in host events settle a background delete from listings

Git unlists the checkout before the host deletes the branch, cleans the push target and purges
metadata, and the worktree-directory watcher refetches within 250 ms. The renderer read the
missing row as a finished delete, so the waiter resolved without the preserved branch (no
toast) and a failure in those last steps showed as success; the real outcome was then dropped.
The listing fallback exists only for a lost outcome event, so it now applies only after this
host's event stream had a gap: a new subscription or a replay after reconnect.

* perf(worktrees): let a bulk delete start each same-repo checkout delete once the host accepts the last

A bulk delete ran one worktree at a time per repo (#2259, for packed-refs and ref-lock races in
branch cleanup). With Git now deleting each checkout for 20-35 s before the request settles, N
worktrees in one repo took N times that. The renderer now queues same-repo deletes only until
the host accepts each one; a parent still waits for its nested children to finish. The host
serializes the branch cleanup step per repo itself, which also covers removals started by
different clients.

* test(worktrees): pin the host platform in the mocked removal suites so they pass on Windows

Removal now passes -c core.longpaths=true on Windows, so the exact-argv
assertions and command-keyed mocks never matched there (17 failures on a
Windows host). Pin darwin as the add-worktree suites already do, and drive
the one Windows-specific case through the same spy.

* test(worktrees): type the blocked git remove result instead of a broad object

The anti-slop static-analysis gate rejects `object` parameters.

* test(worktrees): clear the changed-code quality gate in the removal suites

Merge the duplicate node:fs import, build the mock child without a cast, read
worktrees:list rows through one typed helper, and give the remaining casts a SAFETY line.

* fix(worktrees): record each background delete durably and finish it after a quit or crash

A quit mid-delete left git to finish the checkout on its own while the branch
delete and metadata purge never ran; a crash left a normal-looking row. Each
accepted local removal now writes a record beside the profile state before git
starts, clears it on success or failure, and the host runs the same delete
again for any record left at startup, re-deriving what remains from git and
disk. An orderly quit stops the checkout delete without waiting for it.

* test(worktrees): type the interrupted-removal assertions for the node typecheck

* fix(worktrees): finish an interrupted delete that already removed the checkout's .git file

Quit stops git worktree remove mid-delete, and Git deletes the checkout's .git
file wherever it falls in directory order. Git then refuses the checkout
("validation failed ... .git does not exist") on every retry, so the startup
finish failed and the row could never be deleted from Orca. A registered
checkout this record owns that has lost its .git file now finishes like an
unregistered one: leftover files, prune, then the branch.

* fix(worktrees): let Git finish an interrupted delete, and never take a different checkout

A quit or crash that stops `git worktree remove` after it deleted the checkout's
.git file left a registered checkout Git refuses to remove. The previous fix
deleted that leftover inside Orca's process, which is the bulk delete this
change exists to avoid (and on Windows the leftover can be most of the
checkout). The startup finish now rewrites the missing .git file from Git's
own admin entry for that path and lets `git worktree remove --force` delete
it. `git worktree repair` is not used: it also re-points every other
registered path, including a checkout another repository now owns there.
Orca deletes the leftover itself only when no admin entry claims the path.

The startup finish forces, so it now leaves the path alone when the checkout
there is not the one recorded: a registered worktree on a different branch or
head, or a `.git` at a path Git already unregistered. The record is dropped and
the card shows why.

The record write before Git starts is now bounded (2 s, logged when exceeded)
so a stalled disk cannot hold the delete, and the outcome is published before
the record's clear reaches disk.

* test(worktrees): compare worktree paths by value and tear down with Windows lock retries

Git prints forward slashes in `git worktree list` on Windows, so the real-Git
removal suites never found a joined path there: positive checks failed and
negative ones passed without proving anything. They now compare Git's parsed
rows by value. Teardown uses the shared retrying removeTree, since Windows can
hold the deleted checkout busy for a moment after Git exits. Adds a
relative-path worktree case for the .git restore (skipped before Git 2.48).

* fix(worktrees): reply to a worktree delete when it has finished, not on a broadcast event

A current client's delete request now waits for the host's background delete and gets its real
result (removed, a preserved branch, or the error) as the reply, the way it did before the delete
moved off the request. A request that arrives while the delete runs joins it and gets the same
result. Every other view keeps reading the host's `removing` marker: the row leaving means the
delete finished, and the row listed again without the marker shows "The delete did not finish.
Try again." on a card that view had marked Deleting. A request whose reply is lost (a timeout or a
dropped connection) settles the same way from a fresh listing instead of reporting a failure.

Clients without the background-removal capability (mobile, the CLI, older desktops) are still
answered on acceptance and have rows under removal left out of their listings.

This removes the outcome on worktreesChanged and everything it needed: the renderer's outcome
waiters, early-outcome buffer and TTL, per-host event-gap generations, the request pre-registration,
and the accept callback bulk delete used. Bulk delete runs same-repo deletes in parallel only on
this machine, whose host serializes branch cleanup per repo; SSH and paired hosts stay serialized.

* test(worktrees): type the pending-removal host id in the background-removal suite

* fix(worktrees): answer a delete request even when a concurrent removal of the same worktree replaced its record

The desktop app's removal and the runtime removal (CLI, paired clients) coalesce separately, so
both can be accepted for one worktree. The second replaced the first's record, and the first
delete then finished without resolving the request waiting on it, leaving the desktop card on
Deleting indefinitely. Each delete now settles the request it was started for.

* fix(worktrees): run same-repo removal archive hooks and teardown one at a time on the host

Local bulk delete now sends same-repo removals in parallel, so their archive hooks, terminal
teardown and preflight ran at once; a hook that writes refs can race the repo's ref locks
(#2259). The host now serializes each local removal up to acceptance per repo, for every
client; Git's checkout delete still runs in parallel under the delete limit.

* fix(runtime): keep waiting worktree deletes out of a host's foreground call slots

worktree.rm now replies only after Git deletes the checkout (up to minutes), so on paired
desktop and web each waiting delete held one of the host's 8 foreground call slots, and a
bulk delete queued listing refreshes and every other foreground call behind it. Deletes now
run in their own lane with the same bound; the 2-slot background lane stays for status polls.

* fix(worktrees): join a same-worktree delete accepted while a removal waited its repo turn

The desktop app and the runtime (CLI, paired clients, web) check for a running delete before
they queue for the repo's acceptance turn. A delete of the same worktree from the other path,
accepted while this one queued, was missed: this request re-ran the archive hook, stopped the
terminals again and started a second `git worktree remove` on the directory Git was deleting.
The queued acceptance now re-checks and joins the running delete.

* fix(worktrees): fence a resumed delete's checkout from startup, and drop rows a listing read before the delete finished

A delete a quit or crash interrupted took its terminal and file-watcher gate only when the resume
job ran, after the first window was shown; session restore could open a shell or watcher inside the
half-deleted checkout first, and on Windows that handle can fail the resumed git delete. Loading the
records now fences each recorded path, and the resumed job takes the fence over in the same tick it
takes its own gate.

A listing that read git's registration before a delete finished, and replied after the removal
record cleared, returned the row unmarked, so other views briefly showed "The delete did not
finish". Listings now capture the pending removals before reading git and leave out a row whose
delete finished successfully since; a row whose delete failed stays listed as before.

* test(worktrees): keep git's auto-maintenance out of the real-git removal suite

CI's Git 2.55 failed the file-pool test in teardown with ENOTEMPTY on the scratch repo's
objects/pack after the test body passed: the 3,000-file commit's detached auto-maintenance was
still writing a pack. The scratch repo now disables auto-maintenance and auto-gc.

* fix(worktrees): one archive-hook approval covers a same-repo bulk delete again

Local same-repo deletes now start together, so each queued its trust prompt with a state snapshot
taken before the first prompt was answered; approving the first still showed the same prompt once
per remaining worktree. The queued check now reads the store when its turn comes.
2026-09-30 16:32:20 -07:00