Commit Graph
6 Commits
Author SHA1 Message Date
Neil 7ea01279cd feat(search): bundle ripgrep for local, WSL, and SSH search (#22396)
* feat(search): bundle ripgrep for local, WSL, and SSH search

Ship @vscode/ripgrep-universal's prebuilt rg for all six relay platforms in
every desktop artifact. Local and WSL searches spawn the bundled binary and
drop the git ls-files / git grep fallbacks; SSH deploys upload the remote's
binary once per ripgrep version and the relay prefers it over PATH rg.

* fix(search): address bundled ripgrep review findings

- Key the SSH ripgrep cache on the binary's content hash; a package bump is the only update step
- glibc verifier: read arch tokens below the slice root and accept static ELFs (arm64 release blocker)
- Ship ripgrep/PCRE2/musl license notices; bundle rg with orcad
- Packaged builds never spawn a bare rg; report fd pressure as transient
- SSH: install rg before sweep/GC, size-validate installs, back off instead of disabling on launch failure
- Scope Dependabot to @vscode/ripgrep-universal; revert unrelated lockfile churn

* chore(search): drop bundled-ripgrep reference doc; assert full packaging layout parity

* refactor(search): one entry point for spawning the bundled ripgrep

Local Quick Open, Quick Open path search, the Explorer name filter, and
runtime text search each repeated the same three steps: resolve the bundled
command, spread in the WSL distro, spread in the WSL shell expression. Fold
that into spawnBundledRipgrep so one place owns the rule that a bare 'rg'
must never reach spawn, and simplify the resolver's command/packaged checks.

Restore the AGENTS.md ripgrep rule dropped alongside its reference doc in
63f4dac, and note why the relay's availability probe may spawn a bare 'rg'.

No behaviour change; verified by the existing suites plus a new test that
pins the local, WSL-routed, and distro-routed-but-Windows-output cases.

* refactor(search): drop the local install-ripgrep path; enforce the rg rule

Bundling rg removed the local git/readdir fallback, so nothing can produce
the "install ripgrep on the host running the Quick Open scan" guidance any
more -- only a remote host an upload never reached still reaches the capped
listing. Drop the host parameter, the renderer's local branch and its
translation key, and the relay wrapper that existed only to pass 'remote'.

Add a ratchet test for bare 'rg' spawns, since the AGENTS.md rule alone had
nothing enforcing it. Its one allowlist entry is the relay's PATH probe,
which asks about PATH by definition. Verified the guard catches a planted
offender rather than passing vacuously.

Also stop chaining the remote cleanup sweep behind the ripgrep upload: on a
cold host that is a multi-MB transfer, and stale upload stages and
superseded version dirs were left on the remote for its whole duration. The
two touch different trees, so they now run concurrently.

* test(ssh): pin that the cleanup sweep does not wait on the ripgrep upload

* fix(search): derive rg spawn types instead of importing node:child_process

A type-only import still counts against the child_process ratchet, whose pin
and allowlist only ever shrink. Derive both types from wslAwareSpawn instead.

* fix(search): surface an unreachable WSL workspace instead of an empty result

Inside `bash -c`, a failed `cd` exits 1 -- the same code ripgrep uses for "no
matches" -- so a WSL workspace whose directory had gone away reported an empty
listing as a successful scan. main did not have this hole: checkRgAvailable ran
the same `cd` wrapper first and settled on `code === 0`, diverting to the git
fallback that this PR deletes. The WSL wrapper now takes an optional
cwdFailureExitCode; rg passes 97, and all four close handlers reject with a
clear error before the unavailable check can blame the install.

Also from review:
- Bound the fire-and-forget ripgrep upload with deploySignal. The controller
  aborts only on the deploy timeout, never on success, so this cancels a
  still-running upload when the deploy gives up.
- Run the stale-stage sweep before the installed check rather than inside its
  else branch. Once rg was installed every later deploy took the PRESENT path,
  so a stage orphaned by a dropped connection was never collected again.
- Note in orcad-remote-deploy.ts why wiring it up needs ripgrep work first:
  build-orcad.mjs copies only the build host's rg, and orcad reports
  isPackaged() === true, so a remote of another platform would find nothing.

ssh-relay-deploy.test.ts sat at the max-lines cap, so any edit to it failed the
gate. Split the four Windows named-pipe deploys into their own file (926 -> 737
+ 333); both are now well clear of it.

* fix(search): name the unreachable root in every handler, not three of four

Round-two review caught that the missing-cwd branch in scanRipgrepPaths sat
AFTER isRipgrepUnavailableExit, which classifies any code above 2 as a broken
install -- so for exit 97 it was dead code and Quick Open still told the user to
reinstall Orca. Reordered; all four handlers now check it first.

Also from review:
- A vanished workspace makes spawn fail with ENOENT, which read as a damaged
  install on every local path. Confirm the cwd with isRipgrepSpawnCwdUsable --
  the guard the relay already applies -- before blaming the binary. The async
  continuation re-checks `resolved`, because finish() drops its argument once
  settled and the rejected promise would otherwise go unhandled.
- bundledRipgrepCommand returned a bare 'rg' for an arch outside the bundled
  set, bypassing the guard that exists so Windows cannot resolve a bare name
  against the repo cwd. A packaged app now always names an absolute path.

Drop ci-shards/unit-assignment.json, a 9,425-line CI artifact swept in from
reproducing a shard locally, and gitignore the directory that produced it.

The "rg genuinely cannot start" test pointed at a synthetic /repo, which the
new guard correctly reports as unreachable; it now resolves to a real root so
it still tests what its name says.

* fix(search): let the error handler own the spawn-failure verdict

A failed spawn emits 'error' and THEN 'close' with a negative code. The cwd
check added in the error handler did not settle, so the close handler settled
first -- synchronously, with the reinstall message -- and won the race every
time. The branch was not merely flaky, it was unreachable in all four handlers:
it is guarded by pid === undefined, which is exactly the case that always
produces a following close(code < 0). Verified against a real spawn: 3/3 runs
give error(ENOENT) -> close(-2). The error handler now detaches 'close' before
the probe, so it owns the outcome.

The probe also had no rejection handler, so a probe that rejected left the
search unsettled forever -- a hang, not just a wrong message. It now falls back
to the prior verdict rather than inventing one.

Tests: filesystem-search-rg-timeout and orca-runtime-files-search already cover
error-first and close-first, but against synthetic roots that the new guard
correctly calls unreachable; they now resolve to a real root, keeping each
test's stated intent. Added a Quick Open case for the vanished-workspace path
and confirmed it fails with the old ordering.

* test(search): cover exit code 97 in all four ripgrep close handlers

Round-four review found the missing-cwd branch had zero handler coverage: no
test anywhere emitted close(97), only -2/0/1/2/127. Ordering was correct, but
guarded by source-line order alone -- and that exact ordering was wrong in
three of four handlers two commits ago. Each suite now drives close(97) through
its real handler and expects the unreachable-root message.

Verified the tests earn their place: neutering the missing-cwd check fails
exactly four tests, one per handler.

Also drop a Reflect.get the anti-slop gate rejects, in favour of `in` narrowing.

* docs(search): stop claiming the close handler always wins the race

The previous commit asserted close "would beat this threadpool round-trip every
time", from an n=3 sample that measured event ordering -- which was never in
dispute -- rather than probe-vs-close. Two later measurements disagree with each
other: 50/50 close-first here, 30/50 probe-first in review. Either way it is a
race on a sub-millisecond margin, and the detach is what makes the verdict
deterministic.

Why this wording matters: "close wins every time" is an argument for deleting
the detach as a guard against an impossible race. No test would catch that --
the suites emit error and close in the same synchronous tick.

* chore(search): ship the jemalloc and libunwind notices the Linux rg needs

The statically linked Linux builds carry jemalloc (BSD-2-Clause) and LLVM
libunwind (Apache-2.0 WITH LLVM-exception) in addition to PCRE2 and musl, and
both require their notice on binary redistribution. Confirmed with `strings`:
their symbols are present in linux-x64 and linux-arm64 and absent from the
darwin and win32 builds. Texts taken from the upstream canonical sources.

extraResources already copies the whole licenses directory, so these ship
without a packaging change.

* fix(relay): stop spawning a bare rg, name unreachable roots, collect old builds

Three gaps the reviews surfaced on the remote side, all pre-existing on main.

Bare `rg` on Windows remotes. Both relay spawn sites pass the user's repo as
cwd, and CreateProcessW searches the cwd before PATH -- the same hijack the
desktop side already fixes. The relay now walks PATH itself and spawns an
absolute rg.exe, skipping relative PATH entries because those resolve against
the cwd. No rg on PATH yields null, which callers treat as "ripgrep
unavailable" rather than handing spawn a bare name. POSIX keeps the bare name:
execvp never consults the cwd, so there is nothing to resolve and nothing to
gain. With the last probe converted, the bare-spawn ratchet allowlist is empty.

Empty results for an unreachable root. settleLaunchFailure resolved an empty,
successful-looking scan when the root was gone but PATH rg existed, and the
git/readdir chain never engaged because it only triggers on
RipgrepUnavailableError. Both relay paths now reject naming the root, matching
local workspaces. Missing-rg keeps precedence over a missing root, because only
that verdict engages the fallback chain -- two tests pinned that deliberately
and it would have been wrong to flip it.

Unbounded ~/.orca-remote/ripgrep/. Nothing collected this tree; the relay's
version GC only matches `relay-*`, so every rg bump left another ~5 MB per host
forever. The probe command now also drops sibling builds older than two weeks,
sparing the current one and live upload stages, on POSIX and PowerShell alike.
Two weeks because a client pinned to an older build may still be using it; the
cost of collecting one early is that client re-uploading once.

* fix(relay): probe the rg that failed, and close the drive-relative PATH hole

Five review findings against the previous commit, all reproduced first.

The launch-failure classifier probed PATH rg, but the spawn that failed was the
bundled binary. On the normal remote setup -- no rg on PATH, which is why Orca
uploads one -- the probe failed and a moved workspace was reported as a missing
ripgrep, telling the user to install what Orca already ships. So the fix was
inert on exactly the hosts the uploader exists for. It now takes a candidate
list and asks the binary that actually failed first, then PATH.

path.win32.isAbsolute accepts `\tools` and `/tools`: rooted, but carrying no
drive, so they resolve against whatever drive the process is on. The probe
would have validated one against the relay's drive while the spawn, running
with the user's repo as cwd, resolved it against the repo's -- the same
cwd-dependence this lookup removes, narrowed from directory to drive. A real
drive letter or UNC root is now required.

probeRipgrepVersion had lost the timeout's kill in the rewrite, leaking a live
process and a ref'd handle per launch failure -- for a hang, which is the very
case the bundled-rg back-off exists for. It also spawned without windowsHide,
which would flash a console; fixing that made an allowlist entry stale, so the
entry is gone and the pin ratchets down 63 -> 62.

`windowsPathRipgrep ??= …` never memoised a miss, because null is nullish. The
caching was inverted against cost: a hit stops at the first directory, a miss
stats every one, and only the miss was repeated -- per spawn.

The bare-spawn ratchet claimed "nothing in production spawns a bare rg", which
is false on POSIX. It now also matches PATH_RIPGREP_COMMAND at a spawn site,
and the comment states plainly what a textual guard cannot see: the POSIX bare
name reaches spawn as a parameter, and is safe because execvp ignores the cwd.

The drive-rooted predicate is tested directly rather than through the
filesystem -- a temp dir on a POSIX CI host has no drive letter to exercise
win32 semantics with, so the filesystem test could never have caught this.

* test(mobile): repin the session closure past #22452's two shared modules

Merging main brought the closure to 4220 against a pin of 4218. The two extra
modules are `src/shared/agent-turn-outcome.ts` and `src/shared/main-agent-status.ts`
from #22452, which the status projection this route already reaches import.
That change was src/shared-only, so the mobile job never ran on it -- the same
way the structured tool line slipped past, as the ledger above already records.

Repinned here because this PR's file set is what next made the job run, not
because this PR reaches either module. Verified: of the 28 source files this
branch changes, none appear anywhere in the route's 4220-module closure.

* fix(search): preserve remote binaries and complete runtime packaging

* test(relay): pin the probe's env now that it inherits the relay's PATH

8d6759a threaded the relay env into probeRipgrepVersion -- correctly, since the
probe decides whether a launch failure was the binary or the root and so has to
resolve the same rg the failed spawn would have. It left the assertion that
pins the probe's spawn arguments behind, which is what CI caught.

Asserting buildRelayCommandEnv() rather than loosening the match to any object:
under process.env the probe could resolve a different rg, or none, which is the
regression the change exists to prevent.

* feat(ssh): collect remote ripgrep builds by reference, not by age

Nothing collected `~/.orca-remote/ripgrep/`: the version GC matches only
`relay-*`, so every change to the shipped bytes left another ~5 MB on every SSH
host, permanently. The age window this replaces was the wrong instrument --
a directory's mtime is when it was written, not when it was last used, so it
cannot tell a superseded build from the one a live relay was launched against.
Deleting the latter is not graceful degradation: without a PATH ripgrep remote
text search rejects outright, and listing drops to the capped walk this PR
exists to remove.

So the question is reference. Each relay directory now records the build it
runs against in `.ripgrep-ref`, written only once that binary is confirmed
present, and the GC collects a build only when no installation names it.

The discipline is ssh-relay-native-deps-cache-gc.ts': anything the pass cannot
account for blocks the whole pass. A relay directory with no readable marker is
an older Orca's, possibly running right now against a binary it never recorded,
so the pass declines rather than guessing. Those directories are removed by the
version GC in time, which is what makes their builds collectable -- hence
running after it, not beside it. Deletion is the same tombstone, recheck under
the rename, then remove, so a deploy that takes a reference mid-pass gets its
tree restored. Windows has no pass yet, matching the native-deps cache's gate.

One test note: the first version of the "unaccountable blocks the pass" test
passed against a deliberately broken guard, because the tombstone recheck
masked its absence. The test now puts a readable recheck behind an unreadable
first scan, which is the only shape that fails when that guard is removed.

Recording the reference lives inside ensureRemoteBundledRipgrep rather than at
the call site: it is the same concern, and it keeps the deploy's ripgrep
surface to one call for the tests that mock it to protect their exec queues.

* feat(ssh): collect Windows remotes too, and ship the Rust crate notices

Three items previously left documented-but-open.

Windows remote accumulation. The cache GC was POSIX-gated, so the leak did not
go away -- it moved to the platform with the larger binary (rg.exe is 5.43 MB on
win32-x64, against 4.77 MB for linux-arm64). The PowerShell dialect now does the
same reference scan: entries and references carry token prefixes, because
PowerShell writes every uncaptured value to stdout and an untokenised listing
would feed Remove-Item whatever a cmdlet happened to emit.

Verified on a real Windows host rather than a mock: the listing emits its
ENTRY/LIST_OK tokens, a relay directory carrying a marker yields REF <entry>,
and a relay directory without one yields REFS_ERR -- the safety path, on the
real interpreter.

Rust crate notices. The crate set was read out of the shipped binary's symbols
and the licence identifiers taken from crates.io rather than assumed. Where a
crate offers the Unlicense, Orca elects it: a public-domain dedication carries
no notice obligation, and that covers eight of them. The four that do not offer
it get their MIT text reproduced. encoding_rs carries a BSD-3-Clause notice for
its WHATWG-derived encoding data that is joined by AND, not OR, so electing MIT
does not discharge it.

Release-only validation, corrected rather than repeated. Linux AppImage/deb/rpm
already runs in CI's package job on every PR, and Windows signing was already
rehearsed on this branch. macOS notarization is the only item a release must
still exercise, and the exposure is narrow: notarization requires signatures on
Mach-O binaries, and of the six bundled builds only the two darwin ones are
Mach-O -- `file` reports ELF for linux and PE32+ for win32 -- so signIgnore
excludes only files the notary never asks about.

orcad-artifacts.test.ts caught the new notice file missing from the standalone
runtime's shipped list, which is exactly the gap that test exists to catch: a
notice committed to the repo but never actually shipped.

* fix(search): protect relay cache references and handle failed spawns

* fix(ripgrep): close review gaps and repair deployment fixtures

* test(mobile): refresh merged session module census

* fix(ssh): preserve ripgrep caches with empty legacy references

* test(mobile): assert bundle boundaries instead of global module count
2026-09-24 17:25:48 -07:00
Neil e86cba888b build: reduce native dependency installs to the host platform (#20420)
* Reduce native dependency installs to the host platform

* Remove install policy documentation

* Guard cross-arch packaging and scope release installs to the runner

electron-builder only logs a warning for a missing extraResources source,
so a host-only install silently shipped a foreign-arch slice without its
natives — `pnpm build:mac` on Apple Silicon produced an x64 DMG with no
sherpa-onnx-darwin-x64 and no @parcel/watcher-darwin-x64. The previous
beforePack hook covered only win32.

- Add assertPackagedNativeVariantsInstalled, an arch-aware check over the
  target's sherpa-onnx, @parcel/watcher, and (on Windows) node-gyp addons.
  beforePack now runs it for every platform, with remedies split: another
  architecture comes from install:release, the os:win32 addons need a
  Windows host.
- Drop --os from the release installs. Every packaging job already runs on
  a runner whose OS matches its target, so only the macOS lanes need extra
  breadth, and only on CPU for their x64+arm64 config. Windows and Linux
  packaging return to a plain host-only install.
- Add --frozen-lockfile to install:release so a bare run cannot rewrite
  the lockfile.
- Restore the install policy reference doc and the CONTRIBUTING note, plus
  the rationale comments dropped from the runtime contract test.
- Gate the packaging-closure assertions on whether the Windows addons are
  installed rather than on the host OS, so a cross-arch install exercises
  them off Windows too.
- Make the workflow contract test read `run:` steps as well as retry-action
  commands, and enforce host-only scoping on the non-macOS packaging lanes.
- Remove the unreferenced install measurement script; its numbers live in
  the policy doc.

* Track the install policy doc and index it from AGENTS.md

docs/** is ignored behind a per-file allow-list, so the new reference doc
was only committed via git add -f and future edits would be skipped. Add
it to the allow-list and give it an AGENTS.md entry like every other
tracked reference doc, so the host-only install rule is discoverable
before someone packages a second architecture.

* Route Windows-lane removals through the retrying helper

Adding these four specs to the PR Windows lane pulled them into the
windows-lane-tree-removal-boundary ratchet, which failed on 20 raw
recursive removals. On Windows a bare rmSync races a handle the OS has
not released, throwing EPERM after the assertions already passed and
reporting a green test as a lane failure.

* Adapt the packaging guard to the vendored Windows registry addon

main vendored windows-native-registry as the workspace package
@orca/windows-registry (#20438). A workspace link resolves on every
host, so including it in the installed-Windows-addons checks proved
nothing. @vscode/windows-process-tree is the only os: win32 npm addon
left, so it alone decides whether the win32 resource plan resolves.
2026-09-12 21:25:03 -07:00
254a07bb05 fix(release): port three release-gate fixes to main so cuts stop re-inheriting them (#19945)
* fix(release): prune optional natives before the linux arch floor check

The linux-arm64 release build failed packaging, and retried three times:

    [verify-linux-glibc-floor] 1 bundled native binary is built for the
    wrong architecture (target arm64):
      resources/node_modules/@parcel/watcher-linux-x64-glibc/watcher.node
      is x64, expected arm64 (from its own path)

`afterPack` ran `verifyLinuxGlibcFloor` at its top, before
`prunePackagedRuntimeNodeModules`. A cross-build intentionally installs
every optional native variant, so at that point the arm64 slice still
carries the x64 `@parcel/watcher` package that the prune exists to drop.
The check was reading a file that was never going to ship.

Move the floor check below the prune so it inspects the binaries actually
packed. Same ordering as 35f1b0ebb9 on the v1.4.197 release branch, which
shipped green and was never merged back to main.

(cherry picked from commit 6400598212)
(cherry picked from commit a883e17116)

* fix(release): restore the minified telemetry constant fallback

The macOS, Windows, and both Linux release builds all failed on "Verify
telemetry constants present in app.asar", blocking publish-release:

    ::error::BUILD_IDENTITY constant missing or unexpected value in
    dist/mac-arm64/Orca.app/Contents/Resources/app.asar

The verifier printed no context sample, because the string
`BUILD_IDENTITY` does not occur anywhere in the shipped bundle at all.
#17527 added `minify: 'oxc'` to the main bundle, so oxc renames the
module-local `const BUILD_IDENTITY` to a short identifier.
`BUILD_IDENTITY_RE` keys off the literal name and therefore cannot match a
production bundle. `MINIFIED_TELEMETRY_RE` matches the adjacent injected
identity/key pair instead, which survives renaming.

#11019 created these patterns without that fallback, and its own comment
predicted this exact failure if minification were ever enabled on the main
bundle. The fallback was written on the v1.4.197 release branch in
35f1b0ebb9 and never merged back to main, so main has never been able to
verify a minified bundle.

Verified against a real local bundle built with the release env vars
(ORCA_BUILD_IDENTITY=stable, ORCA_POSTHOG_WRITE_KEY=phc_...): the bundle
does not contain "BUILD_IDENTITY"; both the narrowed and the widened
BUILD_IDENTITY_RE fail; MINIFIED_TELEMETRY_RE matches and recovers
identity=stable plus the write key.

(cherry picked from commit 39c3e44cd1)
(cherry picked from commit d0bbe4475d)

* fix(skills): stop asserting readdir order in the skill root walk test

The Windows skill-sharing release gate — which blocks publish-release —
failed on this test, wedging the 1.4.198 cut. The implementation is
correct; the assertion was not.

`findSkillFiles` pushes results in `readdir` order and dedupes visited
directories by realpath, so of the 32 junctions pointing at one target
exactly one survives. The assertion hardcoded both the array order and
which link won. `readdir` order is filesystem-dependent: APFS and ext4
return SKILL.md first, while NTFS enumerates its filename index
alphabetically on the uppercased name, where "LINK00" sorts before
"SKILL.MD" (L=0x4C < S=0x53). Windows therefore returned the same two
paths in the opposite order.

Assert the contract instead: the real file is present, exactly one
link-routed path survives dedup, and nothing else does.

(cherry picked from commit 9a832e9bda)
(cherry picked from commit 505b85cd0d)

* test(release): pin ported release gate fixes

* style(skills): apply pinned test formatting

---------

Co-authored-by: Neil <neil@stably.ai>
Co-authored-by: Merge Sim <sim@local>
2026-09-10 19:30:47 -07:00
Brennan BensonandMerge Sim 6933fd70d7 fix(packaging): ship Claude agent SDK with desktop builds (#19042)
* fix(packaging): include Claude agent SDK at runtime

* test(packaging): cover spaced runtime imports

* fix(packaging): verify every emitted main file for bare runtime imports

The packaged-main verifier read two fixed entry files, but rolldown hoists
modules shared by two entries into out/main/chunks. jsonc-parser is reached
only from a chunk today, so nothing verified it, and the agent-hooks entry
the list names contributes no coverage at all. An import that migrates into
a chunk would silently stop being checked -- the same blindness that let the
missing Claude agent SDK ship.

Scan every out/main/**/*.js entry in the asar instead, keeping the two
required-file assertions as a build-integrity check. Measured against the
shipped 1.4.198 app: 93 entries in 72ms, reporting the absent SDK and
nothing else.

Also tighten the specifier match with a (?<![.\w]) lookbehind. Orca has
three registry methods of its own named require(), two taking a string key,
so a minified registry.require('public-a') otherwise reads as a bare module
specifier and fails packaging with a confusing error -- a risk the wider
file set would have multiplied. The lookbehind drops nothing real: detection
over the shipped bundle is identical with and without it.

* test(packaging): cover the missing packaged main entry assertion

The required-file check had no test, so the refactor that split it out of
the scanning loop could have dropped it silently. Removing the assertion
now fails this case.

* docs(packaging): name the embedded-source-string limit of the main scan

ssh-relay-deploy builds a probe script for the REMOTE host as a string, and
its require("node-pty") / require("@parcel/watcher") survive into
out/main/index.js, where this scan counts them as desktop-main imports. Both
are packaged, so it is benign today, but a remote-only dependency added to
that script would fail desktop packaging with a false message -- and the two
obvious fixes (ship the remote dep, or weaken the guard) are both wrong.
Separating an embedded string from real code needs a parser.

* test(packaging): pin the exact import shape oxc emits for the SDK

The fixture only carried the spaced `import (` variant, so nothing pinned
the form a shipped build actually contains. Use the real emitted shape --
`p??=import(`@anthropic-ai/claude-agent-sdk`)`, no space, backticks, and the
`??=` that precedes it -- and keep the spaced variant on the second entry so
both stay covered.

* fix(packaging): keep the main scan able to see a spread require

The `(?<![.\w])` lookbehind also rejected `[...require("pkg")]`, because the
third dot of a spread satisfies it. That trade is not symmetric: excluding a
member call costs a loud release-build failure if it ever misfires, but
excluding a real specifier is this guard going blind -- the failure mode the
whole verifier exists to prevent. Readmit a dot that ends a spread.

Zero occurrences in the shipped bundle today, so this was latent. The chunk
test's asar mock now also emits directory nodes, because real listPackage does
and extractFile throws on them -- that makes the `.js` anchor's load-bearing
role something the tests can actually catch.

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-06 02:27:57 -07:00
Neil 1c4c6b7fec perf(startup): stop queueing window creation behind the proxy apply and i18n (#18436)
* perf(startup): stop queueing window creation behind the proxy apply and i18n

Three independent, measured startup wins, all free:

1. Park the initial Chromium proxy apply on `mainProcessState` instead of
   awaiting it mid-`initializeReadyFoundation`. `setProxy` still starts at the
   identical moment; the default-session request guard (which holds, not
   cancels) is what actually fences fetchers on it, so only window creation
   stops waiting. Runtime launch still awaits it before the desktop relay and
   before every headless-serve fetcher.
2. Run `initializeMainProcessI18nAndMenu` concurrently with
   `initializeMainProcessRuntimeLaunch`. Nothing in window creation reads a
   translated string or the native menu.
3. Load `emojibase-data` in main through `createRequire` on first use instead
   of a static import, keeping 166 KB of JSON off `out/main/index.js` and its
   ~2 ms parse off every launch. The renderer keeps its eager copy unchanged.

out/main/index.js 7,210,071 -> 7,040,147 bytes. No renderer behaviour changes.

* fix(packaging): ship the emoji shortcode dataset main lazily requires

app.asar carries no node_modules, so main's bare requires resolve only out of
Resources/node_modules. emojibase-data is a devDependency and is not in the
packaged runtime allowlist, so the new createRequire in
deferred-emoji-shortcode-dataset.ts threw MODULE_NOT_FOUND in every packaged
build — breaking sanitizeWorktreeName, and with it workspace creation.

Copy the single 166 KB dataset (not the 49 MB package root) into
Resources/node_modules, and gate every createRequire'd bare specifier in
src/main against the packaged resource plan. verifyPackagedMainRuntimeDeps
cannot catch these: the bundler renames the require binding.

* test(proxy): fail CI when a main-process fetcher escapes the default-session guard

The hoist relies on installElectronProxyRequestGuard(session.defaultSession) holding every app-owned request until the persisted proxy lands. Nothing enforced that every fetcher actually lands on defaultSession. Two source-anchored rules do now: no net.fetch/net.request may name a session/partition, and every non-net .fetch( call site is counted against an allowlist.

* test(proxy): close the shorthand and chained-receiver holes in the fetch call-site audit

The audit caught `net.request({ session: x })` and `ident.fetch(`, but not the two
shapes a real regression is just as likely to take: the `{ url, session }` shorthand
that both `net.request` overloads accept, and a receiver with no bare identifier
(`session.fromPartition(...).fetch(`, `ctx.session.fetch(`). Rule 1 now also matches
the shorthand key; rule 2 scans every `.fetch(` and excludes only a literal
`net`/`globalThis`/`global` receiver. Audited counts are unchanged (2/2/1).

* fix(startup): scope the deferred emoji loader to the projects that own it

TS6307: the composite web project lists src/main/ipc/worktree-logic.ts, which
now imports the deferred dataset loader, and the shared lazy test reached into
src/main from a project that has no src/main files. Add the loader to
tsconfig.tc.web.json and move the cross-project case into a src/main test.

Also close the last two review gaps: gate the runtime-RPC startup failure
dialog (the only launch-phase translateMain reader) on a published i18n
barrier so a concurrent i18n phase cannot leave a non-English user with the
English fallback, and let the fetch call-site audit match `net.fetch (url)`.
2026-09-03 21:19:06 -07:00
Neil f37d2fec97 fix(linux): land the reviewed Linux packaging stack on main (#18100)
* fix(linux): give the CLI one entrypoint by extracting the AppImage once

* refactor(linux): trim AppImage CLI registration seams

* test(cli): assert registration lock serialization

* fix(linux): fence AppImage terminal shim mounts

* fix(linux): accept extracted AppImage runtimes with APPDIR only

* docs(linux): make headless AppImage extraction runnable

* refactor(linux): import bundled launcher directly

* fix(linux): reclaim superseded AppImage payloads and packaged symlinks

Pruning removed 3215 of 3216 files from a superseded generation and always
stranded resources/app.asar, leaking ~105 MB per version update. Electron's
asar shim reports a *.asar file as a directory, so the recursive remove tried
to rmdir a real file and failed with ENOTEMPTY; the .catch(() => {}) hid it.
Reproduced end to end on Ubuntu 24.04: 519M -> 623M across one update, and
519M again once the payload is actually reclaimed.

removeExtractedAppImagePayload holds process.noAsar for the removal, counted
so overlapping removals cannot hand the shim back early, and the prune site
now warns with the path instead of swallowing the rejection. All three
removal sites use it -- staging cleanup and displaced roots leaked the same
way.

Also reclaim symlinks left by a packaged deb/rpm install, which the
extracted-cache-only rule turned into a hard conflict on a deb -> AppImage
migration, and name the remedy in the conflict error.

* fix(linux): bound the CLI registration lock wait

`retries: 1000` caps the attempt count, not elapsed time, so at up to 1s per
attempt an IPC-driven registration could hang ~16 minutes against a wedged
holder with no feedback.

A legitimate holder is bounded by the extraction timeout, so wait that plus
slack and then fail with a message naming the lock file, rather than hanging.
`maxRetryTime` is forwarded verbatim to the `retry` package by proper-lockfile.

* fix(linux): stop re-extracting the AppImage on inode metadata churn

The extracted-payload cache key hashed ctime alongside dev/ino/size/mtime.
ctime moves on any inode metadata write -- `chmod +x`, which every AppImage
user is told to run, plus `chown`, an ACL or SELinux relabel, and a backup
restore -- none of which alter a byte of the payload.

Measured on Ubuntu 24.04: `chmod +x` leaves dev, ino, size and mtime
identical and moves ctime alone, so the key changed and the next launch paid
a full ~519 MB re-extraction and a multi-second stall to rebuild a payload it
already had, then pruned the old generation.

Key on content identity instead. An in-place content change moves mtime and
almost always size; a replacement moves the inode. The existing
replace-in-place test still passes.

* fix(linux): stop CLI commands from falling through to Chromium startup

* refactor(cli): remove redundant command membership check

* test(cli): cover command-named project selectors

* fix(cli): redirect the open-url command before startup

* test(linux): cover AUR serve wrapper flags

* fix(linux): tighten CLI launch detection

* fix(linux): respect CLI flag value boundaries

* fix(linux): strip injected Chromium switches from CLI args

* fix(linux): report a missing display instead of dying in uv_close

* refactor(linux): read display locks without a preflight race

* fix(linux): preserve unverified external displays

* chore: format reliability gate manifest

* test(packaging): split runtime resource checks

* fix(linux): fail serve when no display is available

* fix(linux): do not treat a lockless X socket as a dead display

An X server writes its lock beside its socket and both survive a crash
(verified against Xvfb under SIGKILL), so a socket with no lock was never
left by a crashed server. It is an endpoint published from elsewhere: a
container bind-mounting only /tmp/.X11-unix, WSLg, or a foreign PID
namespace. Declaring those dead made the desktop gate exit(1) on displays
that work, with no workaround, and the serve gate refuse to start.

Liveness now splits by ownership. A foreign DISPLAY trusts a lockless
socket; Orca's own :99 does not, because removeStaleDisplayArtifacts
unlinks the lock before the socket and so manufactures that state itself --
adopting it would resurrect the orphan-socket bug and stop the cleanup from
self-healing. The stale-lock rejection is unchanged.

Also correct four doc statements this behaviour falsified.

* fix(linux): fail closed when a stale socket blocks the Xvfb rebind

Readiness only checked that /tmp/.X11-unix/X99 exists. A stale socket we
could not unlink still exists after our own Xvfb refused to bind, so Orca set
DISPLAY to a dead server and Chromium died in Ozone init.

Measured on Ubuntu 24.04 against the pre-fix build: with a leftover :99
socket and no lock, serve exits 139 (SIGSEGV), the socket inode is unchanged
before and after, and no lock is recreated -- it neither cleaned up nor
respawned. To a user that is a crash, not a misconfiguration.

This is reachable in the documented topology, where orca-xvfb.service has no
User= and runs as root while serve runs as User=orca: /tmp is sticky, so the
orca uid cannot unlink a root-owned socket, rmSync fails, and Xvfb exits with
the display already active.

Readiness now requires the display to actually be live -- our socket plus a
lock naming a running process -- so the same state reports an unusable
display and exits 1 with the existing diagnosis.

* fix(linux): recognise abstract X sockets and inherited Wayland fds

Two display setups this gate could not prove were refused outright, and on the
desktop path that is app.exit(1) with no workaround.

An X server may bind only the abstract namespace (`@/tmp/.X11-unix/X0`), which
leaves no filesystem socket to stat. Abstract addresses are kernel-owned and
vanish the moment the owner exits, so an entry in /proc/net/unix is proof of a
live server -- no lock file needed and no stale entry possible. Verified on
Ubuntu 24.04, where 139 such addresses were present.

WAYLAND_SOCKET is an already-connected fd handed over by the compositor, so
there is no path to stat and WAYLAND_DISPLAY may be unset entirely. Its
presence is the display.

Both are consulted only after the filesystem-socket check fails, so no
existing verdict changes.

* fix(linux): never treat Orca's own display number as a foreign endpoint

Recognising a lockless X socket as live is correct for an endpoint published
from elsewhere -- a container bind mount, WSLg -- because an X server writes
its lock beside its socket and both survive a crash. It is wrong for
VIRTUAL_DISPLAY_NUMBER, because Orca's own teardown unlinks the lock before
the socket and so manufactures that exact state.

The managed branch was already strict, but a caller that sets DISPLAY=:99
explicitly takes the foreign path and skipped it, accepting a dead display
left by Orca's own interrupted cleanup. Route the managed number through the
strict probe on both paths.

Found by an adversarial audit of the asymmetry introduced earlier in this
branch; the documented systemd topology is unaffected because its Xvfb writes
a real lock.

* test(linux): add a packaged-artifact contract for the CLI launch paths

* test(linux): avoid buffered serve readiness detection

* test(linux): signal AppImage serve owner directly

* test(linux): tolerate readiness timeout boundary

* test(linux): add startup margin to shutdown oracle

* ci(linux): give package contracts timeout headroom

* fix(ci): route all Linux packaging contract changes

* test(linux): poll shutdown readiness without tail leaks

* test(linux): bound shutdown cleanup grace

* test(linux): assert on CLI output, not the harness's own control lines

run-cli-case.sh echoes `RESULT status=N case=<name>`, and the two cases named
*-skills asserted `expectOutput: 'skills'`. That substring was satisfied by
the case name in the harness's own line, so 2 of 8 cases asserted nothing
about the command -- gutting `skills` entirely would still have gone green.

Control lines are now excluded before matching, and both cases assert the
rendered help header, which only real help output produces. Verified on an
Ubuntu 24.04 host: 8/8 still pass against a stack-tip AppImage.

Also register the gate in reliability-gates.jsonc, which #15085 added a CI
Docker gate without. Red/green is recorded from a stock release AppImage
failing 4 of 8, three of them at status 133 (SIGTRAP).

* fix(linux): require static AppImage runtimes (#17319)

* test(linux): reject a wrong-architecture native binary at packaging time

Cross-building the arm64 slice on an x64 host silently packed an x86-64
`pty.node` -- the rebuild logged "Forcing native rebuild for linux-arm64" and
shipped the host's binary anyway. Every gate here inspects symbol versions,
which are perfectly valid on the wrong architecture, so nothing noticed.

Observed on a Raspberry Pi 5: the packaged app loaded, then failed with
"Failed to load native module: pty.node", and the launch contract reported
3 of 8 cases crashed rather than naming the cause. Swapping in the aarch64
`pty.node` took the same build to 8/8.

Compare ELF `e_machine` against the slice being packaged and fail with the
offending path. Checked before the glibc pass, because a wrong-architecture
binary's symbol versions are valid but meaningless and would send the reader
down the wrong path.

Release CI builds arm64 on a native runner, so this guards local and future
cross-builds rather than a shipped artifact.

* test(linux): judge per-arch vendored binaries against their own path

The first CI run of the architecture gate failed the x64 package job on
`@parcel/watcher-linux-arm64-glibc/watcher.node`. That binary is arm64 on
purpose: the package ships every architecture and its loader picks the match,
so its presence in an x64 build is correct.

Judge a binary against the architecture its own path names, falling back to
the slice when the path names none. That keeps the case this gate exists for
-- `bin/linux-arm64-*/node-pty.node` holding an x86-64 binary, which is what
shipped to a Raspberry Pi 5 -- while letting multi-arch dependencies through.

Dry-run over the real dependency tree flags nothing for either target arch.

* fix(linux): move deb/rpm update installation outside Orca (#17318)

* fix(linux): complete deb/rpm package metadata

* fix(linux): preserve CLI link during package upgrades

* docs(linux): document local RPM build prerequisites

* fix(linux): move deb/rpm update installation outside Orca

* fix(updater): preserve Linux recovery across stale events

* fix(updater): fence stale downloaded events by active target

* fix(updater): preserve active Linux package recovery

* test(linux): keep workflow order assertion in scope

* test(updater): assert stale recovery stays silent

* fix(updater): preserve Linux package recovery after checks

* refactor(updater): keep Linux marker message with status

* fix(linux): describe the right manual update path for deb/rpm hosts

A remote host installed from .deb or .rpm now reports
manual-service-update-required, and the guidance told the operator to
"update through the service manager that starts this server" -- which is
correct for unsupported-headless-serve but wrong for a package install,
where nothing about the remedy involves the service manager.

Say both, keyed on how the host was installed.

* docs(linux): document orcad update restart safety

* docs(linux): scope restart census omissions

* docs(linux): use absolute service CLI launcher

* fix(serve): validate in-process serve options before startup (#17683)

* fix(linux): stop offering updates a distro-managed install cannot apply (#17918)

Closes #17702.

The resources/package-type marker is authoritative but never checked against
the host, so any repackager that unpacks Orca's .deb -- AUR, Nix, a container
rebuild -- inherits `deb` verbatim. Install feasibility was then computed
after a ~165 MB download, so those users got check -> download -> a card
promising an install command -> a dead end.

Validate the marker against the host: a deb/rpm marker with no matching
package manager in the trusted directories means a package manager owns this
install. This reuses the exact lists and resolver that
buildLinuxPackageInstallCommand already loops over, so a false positive is
impossible by construction -- any host flagged here would have failed with
no-package-manager after the download anyway. The gate only moves that
verdict earlier. Verified across Debian 12, Ubuntu 24.04, Arch, Fedora 40 and
openSUSE Leap: no false positive on a real deb host, correct on every
repackaging host.

The release is still reported, because the user does want to know 1.4.194
exists and to update through their distro; only the download path is closed.
`externallyManaged` is an additive optional field on the existing `available`
status, so older paired clients decode it unchanged. downloadUpdate() refuses
authoritatively, since main owns this verdict rather than the card, and
unwinds any pinned-build state first -- a Linux pinned jump resolves to
'release', and stranding isPinnedBuildActive would silently kill every
background check for the rest of the process.

Note the fix the issue suggests cannot work: electron-updater builds a
PacmanUpdater whose doDownloadUpdate looks for a .pacman asset Orca does not
publish, then dereferences undefined.

* style(cli): restore prettier wrapping on install error copy

* test(linux): re-pin the child-process ratchets and the batch-shim allowlist after the merge
2026-09-02 03:08:01 -07:00