* fix(ports): stop joining an undefined resourcesPath on a non-Electron host `resolveWorkerEntryPath` branched on `isPackaged` alone and joined `process.resourcesPath`. orcad reports `isPackaged` true — correctly, it is a production build, and ~15 consumers read it that way to gate HTTPS-only skill downloads and the real CLI name — but `process.resourcesPath` is Electron-only and `undefined` under plain Node. So the packaged branch threw `TypeError [ERR_INVALID_ARG_TYPE]: The "path" argument must be of type string` where a clean "worker unavailable" was the honest outcome. The type said `resourcesPath: string`, which is how it went unnoticed; it is now `string | undefined`, so the compiler carries the fact. A host with no Electron resources tree has no asar to look in, so it falls back to the module directory and lets the caller report a missing worker. Found by the item 1 agent while auditing the same `isPackaged` defect class in the watcher. Verified in both directions: reverting the guard reproduces the TypeError. * feat(orcad): prove node-pty loads before anything requires it Of the two ways node-pty fails, only one is catchable. A missing module throws MODULE_NOT_FOUND. A module built against the wrong libc or Node ABI is refused by the dynamic loader, and in the worst case takes the process down before any handler exists — that is #9902, which crashed the desktop app on Ubuntu 20.04 before a window appeared. There was no libc or ABI precondition anywhere in the tree. So orcad now proves the load in a CHILD process, from main.ts, before anything requires node-pty. Whatever the child does — throw, abort, die on a signal — is data rather than our own death, and the operator gets a sentence naming the host's libc, Node ABI and prebuild slot plus the command to run. Proven-unloadable exits 78 (EX_CONFIG), so a supervisor does not restart an unequippable host forever. A probe that never answered is unverifiable, not blocked: refusing to boot on an inconclusive signal would take down hosts that work. The child dlopens the file node-pty would have chosen, before requiring the package. node-pty's loader walks several directories and rethrows only the LAST error, so a refused binary reads as "Cannot find module ./prebuilds/..." — which sends the operator to install a module that is already there. It also reports through stdout: node echoes the whole -e source above a stack trace, and matching tokens against stderr made the probe's own source text answer for the verdict. Verdicts reach clients as a terminal_unavailable degradation alongside the existing browser_unavailable one, through the same cause-registry shape. degradations[].code is now an open vocabulary; clients already render only `message`. Prebuilds are compiled from PATCHED sources — the patch IS the glibc-floor fix, so an upstream tarball reproduces #9902 — into linux-{x64,arm64}-{glibc,musl} and darwin-{x64,arm64} slots. libc is in the slot name because node-pty's loader falls back to prebuilds/<platform>-<arch> and cannot tell glibc from musl. orcad installs the matching slot at boot, so a host with no compiler serves terminals. The relay's five pure toolchain-diagnosis functions moved to a transport-free module so the Node bundle can reuse them without dragging ssh2 in behind them; the relay keeps its API by re-export. macOS gets `xcode-select --install` rather than the cross-distro apt/dnf/pacman/apk menu, every line of which is wrong there. * test(orcad): pin the node-pty precondition to ground truth, not a prepared host CI's test shard runs `vitest` directly, so `ensure-native-runtime --runtime=node` never prepares node-pty for the Node ABI — `degraded` is the correct verdict there, and asserting 'ok' encoded an environment the shard does not have. Asserting whatever it returned would be vacuous, so the expectation is now derived from an independent require() of node-pty. Verified it still bites: forcing the precondition to always report 'ok' fails the suite. * feat(orcad): run the terminal daemon, and the ops contract around it orcad declared `canRecoverPersistentLocalPtys: () => false` because it did not run the terminal daemon, so every restart, update and rollback SIGKILLed every running terminal — on the host whose selling point is that work survives the client going away. That is the one property `ssh-execution-boundary.md` recommends the peer model for. Item 4 — the daemon: - Port the launch path off electron: `daemon-init.ts`, `daemon-host-relocation.ts` and `observability/logs-directory.ts` now read the `AppEnvironment` port. Relocation additionally asks whether the app root is an asar archive rather than whether the build is packaged, so a Node host answering `isPackaged() === true` no longer walks into an Electron-only NSIS-escape path (same precedent as `parcel-watcher-entry-path.ts`). - `build-orcad.mjs` emits `daemon-entry.js` beside `orcad.js`, scans the forked children's metafiles for electron/node:sqlite, and load-checks the child under plain Node. - orcad spawns and adopts the daemon; shutdown disconnects and never kills it. `canRecoverPersistentLocalPtys` now reads the live provider and is false under degraded routing, where fresh terminals would die with the process. Item 3 — the ops contract (docs/reference/orcad-operations.md): - Bind policy: `--bind`, default loopback, pinned so neither `orca serve`'s wide default nor the connected-device widen can override it, and so a paired client cannot rebind the listener from outside. - Instance lock on the data root before profile load, scoped to the runtime role so it never refuses a restart that a live daemon makes worthwhile. - Supervision: exit codes a supervisor can act on (78 = do not retry), second-signal escalation, a shutdown deadline, and crash-loop containment on daemon respawn. - Health in the readiness payload: build hash, Node ABI, and a PTY self-test that spans both processes — the daemon spawns a real PTY in its own process and the verdict crosses its socket. Both bundle load-checks now assert on exit codes: these bundles are minified onto one line, so Node's uncaught-exception report echoes every string literal in the bundle and the previous message match passed against a bundle that never loaded. * feat(orcad): deploy, activate and roll back a versioned orcad install Plan items 6 and 7 from docs/design/shipping-orcad.html. Install reuses the relay's transaction verbatim — per-version lock, staged SFTP write, .install-complete sentinel, stale-lock recovery — under a parameterized namespace, so orcad-<v>/ sits beside relay-<v>/ permanently (§06). Parameterizing GC is the trap that creates: each model now collects only its own directories, enforced twice (prefix-scoped remote listing plus a local ownership re-check), and a client picks its model from how the host is registered, never from what it finds on disk. Activation is separate from installation, because a versioned directory selects nothing. A candidate is launched, publishes orca_server_ready, and only becomes active if its cross-process health payload passes: right build hash, listening, daemon live, PTY self-test green. A rejected candidate is stopped and the incumbent restarted, so a careful deploy cannot cause the outage it was being careful about. Update and rollback are shaped by the daemon. An update restarts orcad, the daemon outlives it, and the surviving daemon was forked from the outgoing bundle — so live terminals defer the update rather than proceed, and GC pins the active version, the rollback target and the live daemon's bundle. Orca's persisted state carries no schema version, so rollback restores a pre-activation snapshot rather than trusting backward-readability; the point past which it is unsafe is the first terminal created after activation, which the snapshot cannot describe and the surviving daemon still owns. Running the generated shell for real found two bugs the text assertions missed: tar members re-quoted inside a shell variable captured nothing, and kill -0 reports a zombie as alive. * test(orcad): assert the precondition is self-consistent, not environment-shaped The real-host case cannot predict a status: CI's shard runs vitest directly, so node-pty is never built for the Node ABI and 'degraded' is correct there, while a prepared checkout gives 'ok'. The previous attempt used require('node-pty') as ground truth, which resolves the JS wrapper while the native binding loads lazily — it proved strictly less than the precondition checks, and failed CI for exactly that reason. What is invariant on a host with node-pty installed: never 'blocked', and never a degraded verdict carrying an unestablished reason. The injected-input tests keep the logic coverage. * fix(orcad): drop an eslint-disable the rule no longer needs * test(orcad): separate slot placement from the load verdict Both remaining CI failures were the same shape: tests reaching into node_modules for a pty.node that only exists after `ensure-native-runtime --runtime=node`, which CI's shard never runs because it invokes vitest directly. Slot *placement* is the logic worth checking on every host, so it now uses a synthetic payload and asserts the verdict stays honest about not loading. The three assertions that genuinely need a Node-ABI binding are gated on it existing. Verified: breaking slot installation fails both placement tests; with the real pty.node hidden the file is 17 passed / 3 skipped instead of ENOENT. * test(orcad): gate the load-dependent cases on a real load, not on the file existing CI ships a pty.node built for Electron's ABI, so existsSync was true while require still failed — the gate ran exactly the tests that host can never satisfy. It now probes the binding in a child process, so a bad one cannot take the runner down. The self-consistency assertion also allowed too little: 'blocked' is the honest verdict for a corrupt binding, alongside 'ok' on a prepared host and 'degraded' on an unprepared one. What stays invariant is that anything other than 'ok' names an established cause, so a terminal is never declined for a reason nobody worked out. Verified against all three host states: prepared (19 passed), unprepared, and a corrupt binding (17 passed / 3 skipped, no failures). * test(orcad): gate on the whole premise — binding AND spawn-helper CI has a loadable pty.node but no spawn-helper, and a slot without the helper is legitimately 'degraded'. So the previous gate let a test run whose premise ('a complete slot yields ok') that host cannot satisfy. Verified in both states: with the helper present 19 pass; with it removed the load-dependent cases skip (17 passed / 3 skipped) instead of failing. * fix(orcad): preserve degradation types after rebase
7.1 KiB
Linux glibc Compatibility
Orca's Linux builds target stock Ubuntu 20.04 and newer — glibc 2.31 and
libstdc++ GLIBCXX_3.4.28 (also Debian 11, RHEL 9), on both x64 and arm64.
Packaging enforces this floor automatically; keep it in mind when adding or
upgrading native dependencies. (The optional speech feature is the one
exception — see below.)
Why this needs attention
A native module (.node) links against the glibc of the machine that compiled
it. Our release CI compiles node-pty from source on GitHub's ubuntu-latest
runner, whose glibc rises over time as the image is bumped. A binary compiled on
a newer glibc can reference symbol versions that do not exist on an older target,
and the dynamic loader then refuses to load it:
/lib/x86_64-linux-gnu/libc.so.6: version `GLIBC_2.34' not found (required by .../pty.node)
Because the Orca main process loads node-pty at startup, that failure crashes the whole app before a window appears — this is exactly what shipped in v1.4.150 and broke launch on Ubuntu 20.04 (#9902).
The specific trap is glibc's 2.32–2.34 "libpthread/libutil merge", which moved several long-stable functions into libc under brand-new symbol versions:
| Symbol | New version | node-pty use |
|---|---|---|
pthread_sigmask |
GLIBC_2.32 |
reset child signal mask |
openpty |
GLIBC_2.34 |
allocate the pty |
forkpty |
GLIBC_2.34 |
fork the shell |
Electron itself (glibc 2.25) and the other bundled native modules
(sherpa-onnx, @parcel/watcher, both prebuilt on old glibc) stay well under
the floor, so node-pty was the sole blocker.
How we keep the floor
1. Pin the relocated symbols (the fix).
config/patches/node-pty@1.1.0.patch
adds a .symver shim in src/unix/pty.cc that binds openpty, forkpty, and
pthread_sigmask to their pre-merge version node — GLIBC_2.2.5 on x64,
GLIBC_2.17 on arm64 (each architecture's baseline glibc). glibc still ships
those as compatibility aliases, so the reference resolves on both new build hosts
and old targets.
The catch: gcc defaults to --as-needed and, since the pinned symbols now
resolve from libc's compat aliases at build time, it drops libutil/libpthread
from DT_NEEDED. On the target those libraries are where the symbols actually
live, so the patch's binding.gyp ldflags force
-Wl,--no-as-needed,-l:libutil.so.1,-l:libpthread.so.0 back into DT_NEEDED.
The shim is guarded by #if defined(__linux__); macOS and Windows are untouched.
2. Gate packaging (the regression guard).
config/scripts/verify-linux-glibc-floor.cjs
runs in the electron-builder afterPack hook for Linux. It reads every bundled
native binary's version needs (objdump -p "Version References" — the
authoritative load-time list, which also captures symbol-less markers like
GLIBC_ABI_DT_RELR) and fails the build if any strong GLIBC_/GLIBCXX_/
CXXABI_ node is newer than stock Ubuntu 20.04 provides, naming the file and the
offending node. Weak needs are ignored (the loader tolerates them). It also
asserts the flip side of the .symver fix: any binary that imports
openpty/forkpty must keep libutil.so.1 in DT_NEEDED — otherwise the
pinned openpty@GLIBC_2.2.5 resolves from libc's compat alias at build time (so
the version check passes) yet fails to load on 20.04, where those functions live
only in libutil. A future runner bump, a new native dependency, or a dropped
ldflag therefore fails the release build instead of shipping a Linux app that
crashes on launch.
The gate is a static invariant, not an integration test. The load path was verified by hand for this fix (real Ubuntu 20.04, x64 + arm64:
requirenode-pty and spawn a shell). A CI smoke test that loads the packagedpty.nodein a glibc-2.31 container and spawns a shell is the recommended follow-up — it would make the load path self-verifying and stay valid even if the build ever moves to an old-glibc sysroot.
The one carve-out is the sherpa-onnx speech prebuilt, which already requires
GLIBCXX_3.4.29 (GCC 11). It loads lazily in the speech worker
(src/main/speech/stt-worker.ts), never at app launch, so it is exempt from the
libstdc++ floor — its glibc needs are still checked. Speech-to-text therefore
needs a host with libstdc++ from GCC 11+ (Ubuntu 21.10 / 22.04 LTS or newer); the
app itself still launches on stock 20.04.
3. Check before loading, on hosts that ship without a compiler (orcad).
The two gates above protect the packaged desktop app, where the binary is built and
verified by the same pipeline. orcad is deployed to hosts Orca never built on, so it
adds a runtime precondition
(src/main/orcad/node-pty-precondition.ts),
run from main.ts before anything requires node-pty. It loads the addon in a child
process, so a binary the loader refuses — or one that aborts outright — is data rather
than this process's death, and the operator gets a sentence naming the host's libc, its
Node ABI, its prebuild slot and the command to run. A proven-unloadable binary exits 78
(EX_CONFIG) instead of reaching the require; a probe that never answered is reported
as unverifiable and boots anyway, because a silent probe is not evidence. Whatever it
finds is published in status.get's degradations[] under terminal_unavailable.
4. Ship the binary, built from patched sources.
config/scripts/build-orcad-prebuilds.mjs
(pnpm run build:orcad-prebuilds, after build:orcad) compiles node-pty for the current
host and files it under out/orcad/prebuilds/<slot>/, where a slot is
linux-{x64,arm64}-{glibc,musl} or darwin-{x64,arm64}. libc is part of the slot name
because node-pty's own loader falls back to prebuilds/<platform>-<arch> and cannot tell
glibc from musl — a glibc binary parked there is loaded on Alpine and dies at dlopen.
The script refuses to compile a tree where config/patches/node-pty@1.1.0.patch is not
applied: without the patch the prebuilt is a #9902 crash shipped as an artifact rather
than a first-connect error. CI runs it once per slot inside the matching container
(--slot= forces the label), merges the trees, and --require-slots fails a release with
a hole in the matrix.
Adding or upgrading a native dependency
-
Prefer packages that ship prebuilt binaries compiled against an old toolchain (manylinux /
glibc 2.17-class), like@parcel/watcher. -
For a module we compile from source, if the gate flags it, either pin the offending symbols the way node-pty does, or build it in an old-glibc container.
-
To check locally on a Linux host, list what a binary requires (skipping the weak
0x02-flagged needs the loader tolerates):objdump -p path/to/module.node | sed -n '/Version References/,/^$/p'No strong
GLIBC_node may exceed2.31, and noGLIBCXX_/CXXABI_node may exceed3.4.28/1.3.12— what stock Ubuntu 20.04 ships.