Files
orca/config/scripts/verify-linux-glibc-floor.cjs
T
Neil 7ea01279cd feat(search): bundle ripgrep for local, WSL, and SSH search (#22396)
* feat(search): bundle ripgrep for local, WSL, and SSH search

Ship @vscode/ripgrep-universal's prebuilt rg for all six relay platforms in
every desktop artifact. Local and WSL searches spawn the bundled binary and
drop the git ls-files / git grep fallbacks; SSH deploys upload the remote's
binary once per ripgrep version and the relay prefers it over PATH rg.

* fix(search): address bundled ripgrep review findings

- Key the SSH ripgrep cache on the binary's content hash; a package bump is the only update step
- glibc verifier: read arch tokens below the slice root and accept static ELFs (arm64 release blocker)
- Ship ripgrep/PCRE2/musl license notices; bundle rg with orcad
- Packaged builds never spawn a bare rg; report fd pressure as transient
- SSH: install rg before sweep/GC, size-validate installs, back off instead of disabling on launch failure
- Scope Dependabot to @vscode/ripgrep-universal; revert unrelated lockfile churn

* chore(search): drop bundled-ripgrep reference doc; assert full packaging layout parity

* refactor(search): one entry point for spawning the bundled ripgrep

Local Quick Open, Quick Open path search, the Explorer name filter, and
runtime text search each repeated the same three steps: resolve the bundled
command, spread in the WSL distro, spread in the WSL shell expression. Fold
that into spawnBundledRipgrep so one place owns the rule that a bare 'rg'
must never reach spawn, and simplify the resolver's command/packaged checks.

Restore the AGENTS.md ripgrep rule dropped alongside its reference doc in
63f4dac, and note why the relay's availability probe may spawn a bare 'rg'.

No behaviour change; verified by the existing suites plus a new test that
pins the local, WSL-routed, and distro-routed-but-Windows-output cases.

* refactor(search): drop the local install-ripgrep path; enforce the rg rule

Bundling rg removed the local git/readdir fallback, so nothing can produce
the "install ripgrep on the host running the Quick Open scan" guidance any
more -- only a remote host an upload never reached still reaches the capped
listing. Drop the host parameter, the renderer's local branch and its
translation key, and the relay wrapper that existed only to pass 'remote'.

Add a ratchet test for bare 'rg' spawns, since the AGENTS.md rule alone had
nothing enforcing it. Its one allowlist entry is the relay's PATH probe,
which asks about PATH by definition. Verified the guard catches a planted
offender rather than passing vacuously.

Also stop chaining the remote cleanup sweep behind the ripgrep upload: on a
cold host that is a multi-MB transfer, and stale upload stages and
superseded version dirs were left on the remote for its whole duration. The
two touch different trees, so they now run concurrently.

* test(ssh): pin that the cleanup sweep does not wait on the ripgrep upload

* fix(search): derive rg spawn types instead of importing node:child_process

A type-only import still counts against the child_process ratchet, whose pin
and allowlist only ever shrink. Derive both types from wslAwareSpawn instead.

* fix(search): surface an unreachable WSL workspace instead of an empty result

Inside `bash -c`, a failed `cd` exits 1 -- the same code ripgrep uses for "no
matches" -- so a WSL workspace whose directory had gone away reported an empty
listing as a successful scan. main did not have this hole: checkRgAvailable ran
the same `cd` wrapper first and settled on `code === 0`, diverting to the git
fallback that this PR deletes. The WSL wrapper now takes an optional
cwdFailureExitCode; rg passes 97, and all four close handlers reject with a
clear error before the unavailable check can blame the install.

Also from review:
- Bound the fire-and-forget ripgrep upload with deploySignal. The controller
  aborts only on the deploy timeout, never on success, so this cancels a
  still-running upload when the deploy gives up.
- Run the stale-stage sweep before the installed check rather than inside its
  else branch. Once rg was installed every later deploy took the PRESENT path,
  so a stage orphaned by a dropped connection was never collected again.
- Note in orcad-remote-deploy.ts why wiring it up needs ripgrep work first:
  build-orcad.mjs copies only the build host's rg, and orcad reports
  isPackaged() === true, so a remote of another platform would find nothing.

ssh-relay-deploy.test.ts sat at the max-lines cap, so any edit to it failed the
gate. Split the four Windows named-pipe deploys into their own file (926 -> 737
+ 333); both are now well clear of it.

* fix(search): name the unreachable root in every handler, not three of four

Round-two review caught that the missing-cwd branch in scanRipgrepPaths sat
AFTER isRipgrepUnavailableExit, which classifies any code above 2 as a broken
install -- so for exit 97 it was dead code and Quick Open still told the user to
reinstall Orca. Reordered; all four handlers now check it first.

Also from review:
- A vanished workspace makes spawn fail with ENOENT, which read as a damaged
  install on every local path. Confirm the cwd with isRipgrepSpawnCwdUsable --
  the guard the relay already applies -- before blaming the binary. The async
  continuation re-checks `resolved`, because finish() drops its argument once
  settled and the rejected promise would otherwise go unhandled.
- bundledRipgrepCommand returned a bare 'rg' for an arch outside the bundled
  set, bypassing the guard that exists so Windows cannot resolve a bare name
  against the repo cwd. A packaged app now always names an absolute path.

Drop ci-shards/unit-assignment.json, a 9,425-line CI artifact swept in from
reproducing a shard locally, and gitignore the directory that produced it.

The "rg genuinely cannot start" test pointed at a synthetic /repo, which the
new guard correctly reports as unreachable; it now resolves to a real root so
it still tests what its name says.

* fix(search): let the error handler own the spawn-failure verdict

A failed spawn emits 'error' and THEN 'close' with a negative code. The cwd
check added in the error handler did not settle, so the close handler settled
first -- synchronously, with the reinstall message -- and won the race every
time. The branch was not merely flaky, it was unreachable in all four handlers:
it is guarded by pid === undefined, which is exactly the case that always
produces a following close(code < 0). Verified against a real spawn: 3/3 runs
give error(ENOENT) -> close(-2). The error handler now detaches 'close' before
the probe, so it owns the outcome.

The probe also had no rejection handler, so a probe that rejected left the
search unsettled forever -- a hang, not just a wrong message. It now falls back
to the prior verdict rather than inventing one.

Tests: filesystem-search-rg-timeout and orca-runtime-files-search already cover
error-first and close-first, but against synthetic roots that the new guard
correctly calls unreachable; they now resolve to a real root, keeping each
test's stated intent. Added a Quick Open case for the vanished-workspace path
and confirmed it fails with the old ordering.

* test(search): cover exit code 97 in all four ripgrep close handlers

Round-four review found the missing-cwd branch had zero handler coverage: no
test anywhere emitted close(97), only -2/0/1/2/127. Ordering was correct, but
guarded by source-line order alone -- and that exact ordering was wrong in
three of four handlers two commits ago. Each suite now drives close(97) through
its real handler and expects the unreachable-root message.

Verified the tests earn their place: neutering the missing-cwd check fails
exactly four tests, one per handler.

Also drop a Reflect.get the anti-slop gate rejects, in favour of `in` narrowing.

* docs(search): stop claiming the close handler always wins the race

The previous commit asserted close "would beat this threadpool round-trip every
time", from an n=3 sample that measured event ordering -- which was never in
dispute -- rather than probe-vs-close. Two later measurements disagree with each
other: 50/50 close-first here, 30/50 probe-first in review. Either way it is a
race on a sub-millisecond margin, and the detach is what makes the verdict
deterministic.

Why this wording matters: "close wins every time" is an argument for deleting
the detach as a guard against an impossible race. No test would catch that --
the suites emit error and close in the same synchronous tick.

* chore(search): ship the jemalloc and libunwind notices the Linux rg needs

The statically linked Linux builds carry jemalloc (BSD-2-Clause) and LLVM
libunwind (Apache-2.0 WITH LLVM-exception) in addition to PCRE2 and musl, and
both require their notice on binary redistribution. Confirmed with `strings`:
their symbols are present in linux-x64 and linux-arm64 and absent from the
darwin and win32 builds. Texts taken from the upstream canonical sources.

extraResources already copies the whole licenses directory, so these ship
without a packaging change.

* fix(relay): stop spawning a bare rg, name unreachable roots, collect old builds

Three gaps the reviews surfaced on the remote side, all pre-existing on main.

Bare `rg` on Windows remotes. Both relay spawn sites pass the user's repo as
cwd, and CreateProcessW searches the cwd before PATH -- the same hijack the
desktop side already fixes. The relay now walks PATH itself and spawns an
absolute rg.exe, skipping relative PATH entries because those resolve against
the cwd. No rg on PATH yields null, which callers treat as "ripgrep
unavailable" rather than handing spawn a bare name. POSIX keeps the bare name:
execvp never consults the cwd, so there is nothing to resolve and nothing to
gain. With the last probe converted, the bare-spawn ratchet allowlist is empty.

Empty results for an unreachable root. settleLaunchFailure resolved an empty,
successful-looking scan when the root was gone but PATH rg existed, and the
git/readdir chain never engaged because it only triggers on
RipgrepUnavailableError. Both relay paths now reject naming the root, matching
local workspaces. Missing-rg keeps precedence over a missing root, because only
that verdict engages the fallback chain -- two tests pinned that deliberately
and it would have been wrong to flip it.

Unbounded ~/.orca-remote/ripgrep/. Nothing collected this tree; the relay's
version GC only matches `relay-*`, so every rg bump left another ~5 MB per host
forever. The probe command now also drops sibling builds older than two weeks,
sparing the current one and live upload stages, on POSIX and PowerShell alike.
Two weeks because a client pinned to an older build may still be using it; the
cost of collecting one early is that client re-uploading once.

* fix(relay): probe the rg that failed, and close the drive-relative PATH hole

Five review findings against the previous commit, all reproduced first.

The launch-failure classifier probed PATH rg, but the spawn that failed was the
bundled binary. On the normal remote setup -- no rg on PATH, which is why Orca
uploads one -- the probe failed and a moved workspace was reported as a missing
ripgrep, telling the user to install what Orca already ships. So the fix was
inert on exactly the hosts the uploader exists for. It now takes a candidate
list and asks the binary that actually failed first, then PATH.

path.win32.isAbsolute accepts `\tools` and `/tools`: rooted, but carrying no
drive, so they resolve against whatever drive the process is on. The probe
would have validated one against the relay's drive while the spawn, running
with the user's repo as cwd, resolved it against the repo's -- the same
cwd-dependence this lookup removes, narrowed from directory to drive. A real
drive letter or UNC root is now required.

probeRipgrepVersion had lost the timeout's kill in the rewrite, leaking a live
process and a ref'd handle per launch failure -- for a hang, which is the very
case the bundled-rg back-off exists for. It also spawned without windowsHide,
which would flash a console; fixing that made an allowlist entry stale, so the
entry is gone and the pin ratchets down 63 -> 62.

`windowsPathRipgrep ??= …` never memoised a miss, because null is nullish. The
caching was inverted against cost: a hit stops at the first directory, a miss
stats every one, and only the miss was repeated -- per spawn.

The bare-spawn ratchet claimed "nothing in production spawns a bare rg", which
is false on POSIX. It now also matches PATH_RIPGREP_COMMAND at a spawn site,
and the comment states plainly what a textual guard cannot see: the POSIX bare
name reaches spawn as a parameter, and is safe because execvp ignores the cwd.

The drive-rooted predicate is tested directly rather than through the
filesystem -- a temp dir on a POSIX CI host has no drive letter to exercise
win32 semantics with, so the filesystem test could never have caught this.

* test(mobile): repin the session closure past #22452's two shared modules

Merging main brought the closure to 4220 against a pin of 4218. The two extra
modules are `src/shared/agent-turn-outcome.ts` and `src/shared/main-agent-status.ts`
from #22452, which the status projection this route already reaches import.
That change was src/shared-only, so the mobile job never ran on it -- the same
way the structured tool line slipped past, as the ledger above already records.

Repinned here because this PR's file set is what next made the job run, not
because this PR reaches either module. Verified: of the 28 source files this
branch changes, none appear anywhere in the route's 4220-module closure.

* fix(search): preserve remote binaries and complete runtime packaging

* test(relay): pin the probe's env now that it inherits the relay's PATH

8d6759a threaded the relay env into probeRipgrepVersion -- correctly, since the
probe decides whether a launch failure was the binary or the root and so has to
resolve the same rg the failed spawn would have. It left the assertion that
pins the probe's spawn arguments behind, which is what CI caught.

Asserting buildRelayCommandEnv() rather than loosening the match to any object:
under process.env the probe could resolve a different rg, or none, which is the
regression the change exists to prevent.

* feat(ssh): collect remote ripgrep builds by reference, not by age

Nothing collected `~/.orca-remote/ripgrep/`: the version GC matches only
`relay-*`, so every change to the shipped bytes left another ~5 MB on every SSH
host, permanently. The age window this replaces was the wrong instrument --
a directory's mtime is when it was written, not when it was last used, so it
cannot tell a superseded build from the one a live relay was launched against.
Deleting the latter is not graceful degradation: without a PATH ripgrep remote
text search rejects outright, and listing drops to the capped walk this PR
exists to remove.

So the question is reference. Each relay directory now records the build it
runs against in `.ripgrep-ref`, written only once that binary is confirmed
present, and the GC collects a build only when no installation names it.

The discipline is ssh-relay-native-deps-cache-gc.ts': anything the pass cannot
account for blocks the whole pass. A relay directory with no readable marker is
an older Orca's, possibly running right now against a binary it never recorded,
so the pass declines rather than guessing. Those directories are removed by the
version GC in time, which is what makes their builds collectable -- hence
running after it, not beside it. Deletion is the same tombstone, recheck under
the rename, then remove, so a deploy that takes a reference mid-pass gets its
tree restored. Windows has no pass yet, matching the native-deps cache's gate.

One test note: the first version of the "unaccountable blocks the pass" test
passed against a deliberately broken guard, because the tombstone recheck
masked its absence. The test now puts a readable recheck behind an unreadable
first scan, which is the only shape that fails when that guard is removed.

Recording the reference lives inside ensureRemoteBundledRipgrep rather than at
the call site: it is the same concern, and it keeps the deploy's ripgrep
surface to one call for the tests that mock it to protect their exec queues.

* feat(ssh): collect Windows remotes too, and ship the Rust crate notices

Three items previously left documented-but-open.

Windows remote accumulation. The cache GC was POSIX-gated, so the leak did not
go away -- it moved to the platform with the larger binary (rg.exe is 5.43 MB on
win32-x64, against 4.77 MB for linux-arm64). The PowerShell dialect now does the
same reference scan: entries and references carry token prefixes, because
PowerShell writes every uncaptured value to stdout and an untokenised listing
would feed Remove-Item whatever a cmdlet happened to emit.

Verified on a real Windows host rather than a mock: the listing emits its
ENTRY/LIST_OK tokens, a relay directory carrying a marker yields REF <entry>,
and a relay directory without one yields REFS_ERR -- the safety path, on the
real interpreter.

Rust crate notices. The crate set was read out of the shipped binary's symbols
and the licence identifiers taken from crates.io rather than assumed. Where a
crate offers the Unlicense, Orca elects it: a public-domain dedication carries
no notice obligation, and that covers eight of them. The four that do not offer
it get their MIT text reproduced. encoding_rs carries a BSD-3-Clause notice for
its WHATWG-derived encoding data that is joined by AND, not OR, so electing MIT
does not discharge it.

Release-only validation, corrected rather than repeated. Linux AppImage/deb/rpm
already runs in CI's package job on every PR, and Windows signing was already
rehearsed on this branch. macOS notarization is the only item a release must
still exercise, and the exposure is narrow: notarization requires signatures on
Mach-O binaries, and of the six bundled builds only the two darwin ones are
Mach-O -- `file` reports ELF for linux and PE32+ for win32 -- so signIgnore
excludes only files the notary never asks about.

orcad-artifacts.test.ts caught the new notice file missing from the standalone
runtime's shipped list, which is exactly the gap that test exists to catch: a
notice committed to the repo but never actually shipped.

* fix(search): protect relay cache references and handle failed spawns

* fix(ripgrep): close review gaps and repair deployment fixtures

* test(mobile): refresh merged session module census

* fix(ssh): preserve ripgrep caches with empty legacy references

* test(mobile): assert bundle boundaries instead of global module count
2026-09-24 17:25:48 -07:00

503 lines
19 KiB
JavaScript

const { readdirSync, openSync, readSync, closeSync } = require('node:fs')
const { spawnSync } = require('node:child_process')
const { join, relative } = require('node:path')
// Why: v1.4.150 shipped a Linux build whose node-pty pty.node required
// GLIBC_2.34 (openpty/forkpty were relocated into libc by glibc's
// libutil/libpthread merge), so the app crashed on startup on Ubuntu 20.04
// (glibc 2.31) — the runner image silently bumped the build-host glibc. This
// gate fails Linux packaging if any bundled native binary requires a glibc (or
// libstdc++) symbol version newer than stock Ubuntu 20.04 ships, so a future
// runner bump or dependency change cannot reintroduce the regression unnoticed.
// See docs/reference/linux-glibc-compatibility.md.
const MIN_GLIBC = Object.freeze([2, 31])
// The symbol-version families this gate checks, each with the highest version
// node stock Ubuntu 20.04 provides. glibc is the #9902 launch-crash axis;
// libstdc++ (GLIBCXX_/CXXABI_) is the same crash class for C++ native modules
// against the system libstdc++ (Orca does not bundle one).
const VERSION_FLOORS = Object.freeze([
Object.freeze({ prefix: 'GLIBC_', floor: MIN_GLIBC }),
Object.freeze({ prefix: 'GLIBCXX_', floor: Object.freeze([3, 4, 28]) }),
Object.freeze({ prefix: 'CXXABI_', floor: Object.freeze([1, 3, 12]) })
])
const FLOOR_LABEL = 'Ubuntu 20.04 (glibc 2.31 / libstdc++ GLIBCXX_3.4.28)'
// Why: the sherpa-onnx speech prebuilt is a third-party manylinux binary that
// already requires GLIBCXX_3.4.29 (GCC 11 / Ubuntu 21.10+, 22.04 LTS). It loads
// lazily in the speech worker (src/main/speech/stt-worker.ts), never at app
// launch, so it cannot cause the #9902 startup crash. Exempt it from the
// libstdc++ floor (its glibc is still gated) rather than fail the release on a
// pre-existing, non-launch condition — speech needs libstdc++ >= GCC 11.
const LIBSTDCXX_FLOOR_EXEMPT = /(?:^|[/\\])sherpa-onnx/
// VER_FLG_WEAK: a version need whose references are all weak. The loader
// tolerates its absence (resolves to null and the caller's fallback runs)
// instead of refusing to load, so a weak need must not count as a requirement.
const VER_FLG_WEAK = 0x2
/** Parse a "2.34" / "3.4.28" version string into a numeric tuple. */
function parseGlibcVersion(versionStr) {
return versionStr.split('.').map((part) => Number.parseInt(part, 10))
}
/** Compare two numeric version tuples; missing trailing parts are 0. */
function compareGlibcVersions(a, b) {
const length = Math.max(a.length, b.length)
for (let i = 0; i < length; i += 1) {
const diff = (a[i] ?? 0) - (b[i] ?? 0)
if (diff !== 0) {
return diff < 0 ? -1 : 1
}
}
return 0
}
/**
* Parse `objdump -p` "Version References" (the ELF `.gnu.version_r` section)
* into the version nodes this binary requires from each shared library. This is
* the authoritative load-time requirement list: unlike the dynamic symbol table
* (`objdump -T`), it also captures symbol-less ABI markers such as
* `GLIBC_ABI_DT_RELR` (packed relative relocations, glibc 2.36+) that still
* block loading on an older glibc. Each entry: `0xHASH 0xFLAGS <n> <NAME>`.
*/
function parseVersionNeeds(objdumpOutput) {
const needs = []
let library = null
let inSection = false
for (const line of objdumpOutput.split('\n')) {
if (line.startsWith('Version References:')) {
inSection = true
continue
}
if (!inSection) {
continue
}
// Any new non-indented line ends the Version References block.
if (!/^\s/.test(line)) {
inSection = false
continue
}
const libraryMatch = line.match(/^\s+required from (\S+):/)
if (libraryMatch) {
library = libraryMatch[1]
continue
}
const entryMatch = line.match(/^\s+0x[0-9a-fA-F]+\s+0x([0-9a-fA-F]+)\s+\d+\s+(\S+)/)
if (entryMatch) {
const flags = Number.parseInt(entryMatch[1], 16)
needs.push({ library, name: entryMatch[2], weak: (flags & VER_FLG_WEAK) !== 0 })
}
}
return needs
}
/**
* Whether a version node is newer than the floor Ubuntu 20.04 provides. Numeric
* nodes (`GLIBC_2.34`, `GLIBCXX_3.4.29`) compare by version. Any non-numeric
* glibc node is rejected: `GLIBC_ABI_DT_RELR` is a 2.36+ marker, and
* `GLIBC_PRIVATE` is not a stable ABI contract — its symbols differ across
* glibc releases, so a binary needing one can fail to load on the floor even
* though the version node itself exists (a well-formed addon needs neither).
* Named libstdc++ nodes (`CXXABI_TM_1`, `GLIBCXX_LDBL_*`) ship on 20.04.
* Families we do not gate (`GCC_`, `NSS_`) return false.
*/
function isVersionNodeAboveFloor(name) {
for (const { prefix, floor } of VERSION_FLOORS) {
if (!name.startsWith(prefix)) {
continue
}
const rest = name.slice(prefix.length)
if (/^[0-9]+(?:\.[0-9]+)*$/.test(rest)) {
return compareGlibcVersions(parseGlibcVersion(rest), floor) > 0
}
// Non-numeric suffix: reject every glibc node (ABI markers and PRIVATE).
return prefix === 'GLIBC_'
}
return false
}
function isLibstdcxxNode(name) {
return name.startsWith('GLIBCXX_') || name.startsWith('CXXABI_')
}
/**
* Version needs from `filePath` that would prevent loading on the floor OS.
* `sherpa-onnx` is exempt from the libstdc++ floor (see LIBSTDCXX_FLOOR_EXEMPT)
* but its glibc needs are still checked.
*/
function findFloorViolations(needs, filePath = '') {
const exemptLibstdcxx = LIBSTDCXX_FLOOR_EXEMPT.test(filePath)
return needs.filter(
(need) =>
!need.weak &&
isVersionNodeAboveFloor(need.name) &&
!(exemptLibstdcxx && isLibstdcxxNode(need.name))
)
}
// On stock Ubuntu 20.04 (glibc 2.31) these symbols live ONLY in these DSOs —
// glibc kept openpty/forkpty in libutil until the 2.34 merge. A binary that
// imports them must keep the DSO in DT_NEEDED or they will not resolve on the
// floor. This guards config/patches/node-pty@1.1.0.patch's forced
// `-l:libutil.so.1`: if a toolchain change ever dropped that ldflag, the pinned
// openpty@GLIBC_2.2.5 would still resolve from libc's compat alias at build time
// (so the version-floor check passes) yet fail to load on 20.04. libpthread
// (pthread_sigmask) is intentionally omitted — the Node/Electron host always
// loads it, so it resolves regardless of this addon's DT_NEEDED.
const RELOCATED_SYMBOL_PROVIDERS = Object.freeze({
openpty: 'libutil.so.1',
forkpty: 'libutil.so.1'
})
/**
* Relocated symbols the binary imports whose providing DSO is absent from
* DT_NEEDED — meaning they resolve at build time but not on the floor OS.
*/
function findMissingProviderDeps(importedSymbols, neededLibraries) {
const missing = []
for (const [symbol, library] of Object.entries(RELOCATED_SYMBOL_PROVIDERS)) {
if (importedSymbols.has(symbol) && !neededLibraries.has(library)) {
missing.push({ symbol, library })
}
}
return missing
}
// ELF e_machine values for the Linux slices we package. Names match electron-builder's Arch enum.
const ELF_MACHINE_BY_ARCH = Object.freeze({ x64: 0x3e, arm64: 0xb7 })
const ARCH_BY_ELF_MACHINE = Object.freeze({ 0x3e: 'x64', 0xb7: 'arm64' })
/**
* ELF `e_machine`, or null when the file is not a readable little-endian ELF.
*
* Why this is checked at all: cross-building an arm64 package on an x64 host can silently pack an
* x86-64 `pty.node` into the arm64 slice — the rebuild logs a forced arm64 rebuild and still ships
* the host's binary. Every other gate here inspects symbol versions, which are perfectly valid on
* the wrong architecture, so nothing noticed. Observed on a Raspberry Pi 5: the app loaded, then
* failed with "Failed to load native module: pty.node".
*/
function readElfMachine(filePath) {
let fd
try {
fd = openSync(filePath, 'r')
const header = Buffer.alloc(20)
if (readSync(fd, header, 0, 20, 0) !== 20) {
return null
}
// EI_DATA (offset 5) must be ELFDATA2LSB for a little-endian e_machine read.
if (header[5] !== 1) {
return null
}
return header.readUInt16LE(18)
} catch {
return null
} finally {
if (fd !== undefined) {
closeSync(fd)
}
}
}
// Arch tokens that appear in vendored per-architecture package/directory names.
const ARCH_TOKEN_PATTERN = /(?:^|[^a-z0-9])(arm64|aarch64|x64|x86_64)(?:[^a-z0-9]|$)/i
const ARCH_BY_TOKEN = Object.freeze({ arm64: 'arm64', aarch64: 'arm64', x64: 'x64', x86_64: 'x64' })
/**
* The architecture a path advertises, or null when it advertises none.
*
* Why this matters: some dependencies ship every architecture and let their loader pick
* (`@parcel/watcher-linux-arm64-glibc/watcher.node` is arm64 on purpose inside an x64 build). Those
* must be judged against the arch their own path declares, not against the slice.
*/
function declaredArchFromPath(filePath) {
const match = ARCH_TOKEN_PATTERN.exec(filePath)
return match ? ARCH_BY_TOKEN[match[1].toLowerCase()] : null
}
function findArchViolation(filePath, targetArch, rootDir) {
// A path that names an architecture is judged against that name, so a per-arch vendored package
// is fine while `bin/linux-arm64-.../node-pty.node` holding an x86-64 binary is still caught.
// Why relative: the arm64 slice's own dir (`linux-arm64-unpacked`) must not declare every file arm64.
const declared = declaredArchFromPath(rootDir ? relative(rootDir, filePath) : filePath)
const expectedArch = declared ?? targetArch
const expected = ELF_MACHINE_BY_ARCH[expectedArch]
if (expected === undefined) {
return null
}
const machine = readElfMachine(filePath)
if (machine === null || machine === expected) {
return null
}
return {
machine,
actual: ARCH_BY_ELF_MACHINE[machine] ?? `0x${machine.toString(16)}`,
expectedArch,
declared: declared !== null
}
}
function isElfFile(filePath) {
let fd
try {
fd = openSync(filePath, 'r')
const header = Buffer.alloc(4)
const bytesRead = readSync(fd, header, 0, 4, 0)
return bytesRead === 4 && header[0] === 0x7f && header.toString('latin1', 1, 4) === 'ELF'
} catch {
return false
} finally {
if (fd !== undefined) {
closeSync(fd)
}
}
}
/** Recursively collect ELF native binaries (`.node`, `.so[.N]`, executables). */
function collectNativeBinaries(rootDir) {
const binaries = []
const walk = (dir) => {
let entries
try {
entries = readdirSync(dir, { withFileTypes: true })
} catch {
return
}
for (const entry of entries) {
const fullPath = join(dir, entry.name)
if (entry.isSymbolicLink()) {
continue
}
if (entry.isDirectory()) {
walk(fullPath)
continue
}
if (!entry.isFile()) {
continue
}
// Why: .node/.so are always native; extensionless files (the Electron
// executable, chrome-sandbox) are checked via the ELF magic so we cover
// every launch-critical binary without objdump-ing app.asar or assets.
const looksNative = entry.name.endsWith('.node') || /\.so(\.\d+)*$/.test(entry.name)
if (looksNative || !entry.name.includes('.')) {
if (isElfFile(fullPath)) {
binaries.push(fullPath)
}
}
}
}
walk(rootDir)
return binaries.sort()
}
function resolveObjdump(explicitPath) {
const candidates = [explicitPath, 'objdump', 'llvm-objdump'].filter(Boolean)
for (const candidate of candidates) {
const probe = spawnSync(candidate, ['--version'], { encoding: 'utf8', env: cLocaleEnv() })
if (!probe.error && probe.status === 0) {
return candidate
}
}
return null
}
// Why: GNU objdump localizes its section headers ("Version References:") via
// gettext, and the parser anchors on the English text. Force the C locale so
// output stays deterministic on non-English packaging hosts (LC_ALL=C also
// disables LANGUAGE-based message translation).
function cLocaleEnv() {
return { ...process.env, LC_ALL: 'C', LANG: 'C' }
}
/**
* Run objdump with one flag on `filePath`. Fail-closed: a spawn error, non-zero
* exit, or signal throws, because a silently-unreadable binary (truncated,
* corrupt, or an objdump that cannot decode its format) would let a too-new
* binary slip past the gate.
*/
function runObjdump(objdumpPath, flag, filePath) {
const result = spawnSync(objdumpPath, [flag, filePath], {
encoding: 'utf8',
maxBuffer: 64 * 1024 * 1024,
env: cLocaleEnv()
})
if (result.error) {
throw new Error(
`[verify-linux-glibc-floor] could not run objdump on ${filePath}: ${result.error.message}`
)
}
if (result.signal || result.status !== 0) {
throw new Error(
`[verify-linux-glibc-floor] objdump ${flag} failed for ${filePath} ` +
`(status ${result.status}, signal ${result.signal ?? 'none'}): ${(result.stderr || '').trim()}`
)
}
return result.stdout || ''
}
/** DT_NEEDED shared-library names from `objdump -p` (` NEEDED <lib>`). */
function parseNeededLibraries(objdumpOutput) {
const needed = new Set()
for (const line of objdumpOutput.split('\n')) {
const match = line.match(/^\s+NEEDED\s+(\S+)/)
if (match) {
needed.add(match[1])
}
}
return needed
}
/** Undefined (imported) dynamic symbol base names from `objdump -T` (`*UND*`). */
function parseImportedSymbols(objdumpOutput) {
const imported = new Set()
for (const line of objdumpOutput.split('\n')) {
if (!line.includes('*UND*')) {
continue
}
// The symbol name is the final token; strip any @VERSION suffix.
const token = line.trim().split(/\s+/).pop()
if (token) {
imported.add(token.split('@')[0])
}
}
return imported
}
/** Version needs + DT_NEEDED from a single `objdump -p` (fail-closed). */
function readDynamicInfo(filePath, objdumpPath) {
const output = runObjdump(objdumpPath, '-p', filePath)
return {
versionNeeds: parseVersionNeeds(output),
neededLibraries: parseNeededLibraries(output)
}
}
/** Imported (undefined) dynamic symbols from `objdump -T` (fail-closed). */
function readImportedSymbols(filePath, objdumpPath) {
try {
return parseImportedSymbols(runObjdump(objdumpPath, '-T', filePath))
} catch (error) {
// Why: a statically linked binary (bundled ripgrep) has no dynamic symbol table to import from.
if (error instanceof Error && error.message.includes('not a dynamic object')) {
return new Set()
}
throw error
}
}
/**
* Fail Linux packaging if any bundled native binary under `rootDir` requires a
* glibc/libstdc++ symbol version newer than the floor OS. No-op is not allowed
* on Linux: a missing objdump throws, because a silent skip would defeat the
* regression gate on exactly the host where it matters.
*/
function verifyLinuxGlibcFloor(rootDir, options = {}) {
const binaries = collectNativeBinaries(rootDir)
const targetArch = options.targetArch
if (binaries.length === 0) {
console.log(`[verify-linux-glibc-floor] OK — no bundled native binaries under ${rootDir}`)
return
}
// Why: resolve objdump only once there is something to inspect, so a fixture
// with no ELF binaries does not fail on a host that lacks binutils.
const objdumpPath = resolveObjdump(options.objdumpPath)
if (!objdumpPath) {
throw new Error(
'[verify-linux-glibc-floor] objdump not found. Install binutils on the Linux ' +
'packaging host so the glibc-floor gate can inspect bundled native binaries.'
)
}
// Why before the glibc pass: a wrong-architecture binary's symbol versions are valid but
// meaningless, so reporting a floor violation for it would send the reader down the wrong path.
const archOffenders = binaries
.map((filePath) => ({ filePath, violation: findArchViolation(filePath, targetArch, rootDir) }))
.filter(({ violation }) => violation !== null)
if (archOffenders.length > 0) {
const detail = archOffenders
.map(
({ filePath, violation }) =>
` ${relative(rootDir, filePath) || filePath} is ${violation.actual}, expected ` +
`${violation.expectedArch}${violation.declared ? ' (from its own path)' : ''}`
)
.join('\n')
throw new Error(
`[verify-linux-glibc-floor] ${archOffenders.length} bundled native binar` +
`${archOffenders.length === 1 ? 'y is' : 'ies are'} built for the wrong architecture ` +
`(target ${targetArch}), so the app will fail to load them at runtime:\n${detail}\n` +
'Cross-building a Linux slice can pack the host architecture despite a forced rebuild; ' +
'build this slice on a native runner.'
)
}
const offenders = []
for (const filePath of binaries) {
const { versionNeeds, neededLibraries } = readDynamicInfo(filePath, objdumpPath)
const floorViolations = findFloorViolations(versionNeeds, filePath)
// Only pay for `objdump -T` when a relocated-symbol provider is not already
// in DT_NEEDED (the common, healthy case short-circuits without it).
const providerViolations = Object.values(RELOCATED_SYMBOL_PROVIDERS).some(
(library) => !neededLibraries.has(library)
)
? findMissingProviderDeps(readImportedSymbols(filePath, objdumpPath), neededLibraries)
: []
if (floorViolations.length > 0 || providerViolations.length > 0) {
offenders.push({ filePath, floorViolations, providerViolations })
}
}
if (offenders.length > 0) {
const detail = offenders
.map(({ filePath, floorViolations, providerViolations }) => {
const reasons = []
if (floorViolations.length > 0) {
const nodes = [...new Set(floorViolations.map((v) => v.name))].sort()
const libraries = [...new Set(floorViolations.map((v) => v.library).filter(Boolean))]
reasons.push(
`needs ${nodes.join(', ')}${libraries.length > 0 ? ` (from ${libraries.join(', ')})` : ''}`
)
}
for (const { symbol, library } of providerViolations) {
reasons.push(`imports ${symbol} but ${library} is not in DT_NEEDED`)
}
return ` ${relative(rootDir, filePath) || filePath} ${reasons.join('; ')}`
})
.join('\n')
throw new Error(
`[verify-linux-glibc-floor] ${offenders.length} bundled native binar${offenders.length === 1 ? 'y' : 'ies'} ` +
`will not load on ${FLOOR_LABEL}, so the app will crash on startup there:\n${detail}\n` +
'See docs/reference/linux-glibc-compatibility.md — rebuild the offending module against an older ' +
'toolchain or pin the relocated symbols (as config/patches/node-pty@1.1.0.patch does).'
)
}
console.log(
`[verify-linux-glibc-floor] OK — ${binaries.length} bundled native binaries all load on ${FLOOR_LABEL}`
)
}
module.exports = {
MIN_GLIBC,
ELF_MACHINE_BY_ARCH,
readElfMachine,
declaredArchFromPath,
findArchViolation,
VERSION_FLOORS,
FLOOR_LABEL,
RELOCATED_SYMBOL_PROVIDERS,
parseGlibcVersion,
compareGlibcVersions,
parseVersionNeeds,
parseNeededLibraries,
parseImportedSymbols,
isVersionNodeAboveFloor,
isLibstdcxxNode,
findFloorViolations,
findMissingProviderDeps,
collectNativeBinaries,
readDynamicInfo,
readImportedSymbols,
verifyLinuxGlibcFloor
}