mirror of
https://github.com/stablyai/orca.git
synced 2026-09-25 08:02:31 +00:00
* fix(ports): stop joining an undefined resourcesPath on a non-Electron host `resolveWorkerEntryPath` branched on `isPackaged` alone and joined `process.resourcesPath`. orcad reports `isPackaged` true — correctly, it is a production build, and ~15 consumers read it that way to gate HTTPS-only skill downloads and the real CLI name — but `process.resourcesPath` is Electron-only and `undefined` under plain Node. So the packaged branch threw `TypeError [ERR_INVALID_ARG_TYPE]: The "path" argument must be of type string` where a clean "worker unavailable" was the honest outcome. The type said `resourcesPath: string`, which is how it went unnoticed; it is now `string | undefined`, so the compiler carries the fact. A host with no Electron resources tree has no asar to look in, so it falls back to the module directory and lets the caller report a missing worker. Found by the item 1 agent while auditing the same `isPackaged` defect class in the watcher. Verified in both directions: reverting the guard reproduces the TypeError. * feat(orcad): prove node-pty loads before anything requires it Of the two ways node-pty fails, only one is catchable. A missing module throws MODULE_NOT_FOUND. A module built against the wrong libc or Node ABI is refused by the dynamic loader, and in the worst case takes the process down before any handler exists — that is #9902, which crashed the desktop app on Ubuntu 20.04 before a window appeared. There was no libc or ABI precondition anywhere in the tree. So orcad now proves the load in a CHILD process, from main.ts, before anything requires node-pty. Whatever the child does — throw, abort, die on a signal — is data rather than our own death, and the operator gets a sentence naming the host's libc, Node ABI and prebuild slot plus the command to run. Proven-unloadable exits 78 (EX_CONFIG), so a supervisor does not restart an unequippable host forever. A probe that never answered is unverifiable, not blocked: refusing to boot on an inconclusive signal would take down hosts that work. The child dlopens the file node-pty would have chosen, before requiring the package. node-pty's loader walks several directories and rethrows only the LAST error, so a refused binary reads as "Cannot find module ./prebuilds/..." — which sends the operator to install a module that is already there. It also reports through stdout: node echoes the whole -e source above a stack trace, and matching tokens against stderr made the probe's own source text answer for the verdict. Verdicts reach clients as a terminal_unavailable degradation alongside the existing browser_unavailable one, through the same cause-registry shape. degradations[].code is now an open vocabulary; clients already render only `message`. Prebuilds are compiled from PATCHED sources — the patch IS the glibc-floor fix, so an upstream tarball reproduces #9902 — into linux-{x64,arm64}-{glibc,musl} and darwin-{x64,arm64} slots. libc is in the slot name because node-pty's loader falls back to prebuilds/<platform>-<arch> and cannot tell glibc from musl. orcad installs the matching slot at boot, so a host with no compiler serves terminals. The relay's five pure toolchain-diagnosis functions moved to a transport-free module so the Node bundle can reuse them without dragging ssh2 in behind them; the relay keeps its API by re-export. macOS gets `xcode-select --install` rather than the cross-distro apt/dnf/pacman/apk menu, every line of which is wrong there. * test(orcad): pin the node-pty precondition to ground truth, not a prepared host CI's test shard runs `vitest` directly, so `ensure-native-runtime --runtime=node` never prepares node-pty for the Node ABI — `degraded` is the correct verdict there, and asserting 'ok' encoded an environment the shard does not have. Asserting whatever it returned would be vacuous, so the expectation is now derived from an independent require() of node-pty. Verified it still bites: forcing the precondition to always report 'ok' fails the suite. * feat(orcad): run the terminal daemon, and the ops contract around it orcad declared `canRecoverPersistentLocalPtys: () => false` because it did not run the terminal daemon, so every restart, update and rollback SIGKILLed every running terminal — on the host whose selling point is that work survives the client going away. That is the one property `ssh-execution-boundary.md` recommends the peer model for. Item 4 — the daemon: - Port the launch path off electron: `daemon-init.ts`, `daemon-host-relocation.ts` and `observability/logs-directory.ts` now read the `AppEnvironment` port. Relocation additionally asks whether the app root is an asar archive rather than whether the build is packaged, so a Node host answering `isPackaged() === true` no longer walks into an Electron-only NSIS-escape path (same precedent as `parcel-watcher-entry-path.ts`). - `build-orcad.mjs` emits `daemon-entry.js` beside `orcad.js`, scans the forked children's metafiles for electron/node:sqlite, and load-checks the child under plain Node. - orcad spawns and adopts the daemon; shutdown disconnects and never kills it. `canRecoverPersistentLocalPtys` now reads the live provider and is false under degraded routing, where fresh terminals would die with the process. Item 3 — the ops contract (docs/reference/orcad-operations.md): - Bind policy: `--bind`, default loopback, pinned so neither `orca serve`'s wide default nor the connected-device widen can override it, and so a paired client cannot rebind the listener from outside. - Instance lock on the data root before profile load, scoped to the runtime role so it never refuses a restart that a live daemon makes worthwhile. - Supervision: exit codes a supervisor can act on (78 = do not retry), second-signal escalation, a shutdown deadline, and crash-loop containment on daemon respawn. - Health in the readiness payload: build hash, Node ABI, and a PTY self-test that spans both processes — the daemon spawns a real PTY in its own process and the verdict crosses its socket. Both bundle load-checks now assert on exit codes: these bundles are minified onto one line, so Node's uncaught-exception report echoes every string literal in the bundle and the previous message match passed against a bundle that never loaded. * feat(orcad): deploy, activate and roll back a versioned orcad install Plan items 6 and 7 from docs/design/shipping-orcad.html. Install reuses the relay's transaction verbatim — per-version lock, staged SFTP write, .install-complete sentinel, stale-lock recovery — under a parameterized namespace, so orcad-<v>/ sits beside relay-<v>/ permanently (§06). Parameterizing GC is the trap that creates: each model now collects only its own directories, enforced twice (prefix-scoped remote listing plus a local ownership re-check), and a client picks its model from how the host is registered, never from what it finds on disk. Activation is separate from installation, because a versioned directory selects nothing. A candidate is launched, publishes orca_server_ready, and only becomes active if its cross-process health payload passes: right build hash, listening, daemon live, PTY self-test green. A rejected candidate is stopped and the incumbent restarted, so a careful deploy cannot cause the outage it was being careful about. Update and rollback are shaped by the daemon. An update restarts orcad, the daemon outlives it, and the surviving daemon was forked from the outgoing bundle — so live terminals defer the update rather than proceed, and GC pins the active version, the rollback target and the live daemon's bundle. Orca's persisted state carries no schema version, so rollback restores a pre-activation snapshot rather than trusting backward-readability; the point past which it is unsafe is the first terminal created after activation, which the snapshot cannot describe and the surviving daemon still owns. Running the generated shell for real found two bugs the text assertions missed: tar members re-quoted inside a shell variable captured nothing, and kill -0 reports a zombie as alive. * test(orcad): assert the precondition is self-consistent, not environment-shaped The real-host case cannot predict a status: CI's shard runs vitest directly, so node-pty is never built for the Node ABI and 'degraded' is correct there, while a prepared checkout gives 'ok'. The previous attempt used require('node-pty') as ground truth, which resolves the JS wrapper while the native binding loads lazily — it proved strictly less than the precondition checks, and failed CI for exactly that reason. What is invariant on a host with node-pty installed: never 'blocked', and never a degraded verdict carrying an unestablished reason. The injected-input tests keep the logic coverage. * fix(orcad): drop an eslint-disable the rule no longer needs * test(orcad): separate slot placement from the load verdict Both remaining CI failures were the same shape: tests reaching into node_modules for a pty.node that only exists after `ensure-native-runtime --runtime=node`, which CI's shard never runs because it invokes vitest directly. Slot *placement* is the logic worth checking on every host, so it now uses a synthetic payload and asserts the verdict stays honest about not loading. The three assertions that genuinely need a Node-ABI binding are gated on it existing. Verified: breaking slot installation fails both placement tests; with the real pty.node hidden the file is 17 passed / 3 skipped instead of ENOENT. * test(orcad): gate the load-dependent cases on a real load, not on the file existing CI ships a pty.node built for Electron's ABI, so existsSync was true while require still failed — the gate ran exactly the tests that host can never satisfy. It now probes the binding in a child process, so a bad one cannot take the runner down. The self-consistency assertion also allowed too little: 'blocked' is the honest verdict for a corrupt binding, alongside 'ok' on a prepared host and 'degraded' on an unprepared one. What stays invariant is that anything other than 'ok' names an established cause, so a terminal is never declined for a reason nobody worked out. Verified against all three host states: prepared (19 passed), unprepared, and a corrupt binding (17 passed / 3 skipped, no failures). * test(orcad): gate on the whole premise — binding AND spawn-helper CI has a loadable pty.node but no spawn-helper, and a slot without the helper is legitimately 'degraded'. So the previous gate let a test run whose premise ('a complete slot yields ok') that host cannot satisfy. Verified in both states: with the helper present 19 pass; with it removed the load-dependent cases skip (17 passed / 3 skipped) instead of failing. * fix(orcad): preserve degradation types after rebase
191 lines
10 KiB
Markdown
191 lines
10 KiB
Markdown
# Running orcad
|
|
|
|
`orcad` is the Orca runtime served from plain Node. This is the contract between it and
|
|
whatever supervises it: what it binds, what it owns on disk, who restarts what, and what its
|
|
readiness payload actually proves.
|
|
|
|
Design background: `docs/design/shipping-orcad.html` §00c and §04.
|
|
|
|
## Two long-lived processes, not one
|
|
|
|
A deployment is **orcad** plus **the terminal daemon**.
|
|
|
|
| | orcad | terminal daemon |
|
|
| --- | --- | --- |
|
|
| Started by | the supervisor | orcad, detached |
|
|
| Owns | RPC, git, worktrees, persistence | every local PTY |
|
|
| Lifetime | one supervised run | **outlives orcad** |
|
|
| Endpoint | `ws://<bind>:<port>` | `<data-root>/daemon/daemon-v<N>.sock` |
|
|
|
|
The daemon outliving orcad is the property the whole peer model is recommended for
|
|
(`docs/reference/ssh-execution-boundary.md`): daemon-backed PTYs stay `live` across a runtime
|
|
restart, so a restart, an update or a rollback does not destroy running work. Everything
|
|
below exists to keep that true.
|
|
|
|
**Consequence for supervision:** orcad's shutdown path calls `disconnectDaemon()`, never
|
|
`shutdownDaemon()`. A supervisor that reaps orcad's whole process group — systemd's
|
|
`KillMode=control-group` — kills the daemon too and turns every restart back into data loss.
|
|
Use `KillMode=mixed` (the default) or `process`, and never `--send-sigkill` on the group.
|
|
|
|
## Bind policy
|
|
|
|
`--bind <literal-ip>`, **default `127.0.0.1`**.
|
|
|
|
Only literal IPs are accepted; hostnames are refused because DNS would decide which
|
|
interface got bound. `localhost` maps to `127.0.0.1`. `0.0.0.0` / `::` are the explicit
|
|
opt-ins to network reach, and the startup log says so on every launch.
|
|
|
|
The bind is **pinned**, not defaulted. Two things widen the desktop's listener on their own —
|
|
`orca serve`'s wide default, and a startup where some device has connected before — and an
|
|
unattended host's exposure must be exactly what the operator asked for on every launch. A
|
|
mobile pairing offer, which normally rebinds to all interfaces, is refused while the bind is
|
|
pinned to loopback and reports `network_exposure_failed` rather than advertising an endpoint
|
|
nothing can reach.
|
|
|
|
Under the shipping design a client reaches a remote orcad over an SSH local port-forward, so
|
|
loopback is the correct default and the pairing credential travels over SSH.
|
|
|
|
## Data root and the instance lock
|
|
|
|
The data root is `$ORCA_USER_DATA`, else `$XDG_DATA_HOME/Orca`, else `~/.orca`.
|
|
|
|
Before the profile index or the store is touched, orcad takes `<data-root>/orcad.lock`.
|
|
It refuses to start when:
|
|
|
|
| Code | Meaning |
|
|
| --- | --- |
|
|
| `orcad_data_root_wrong_owner` | the root is owned by another uid (POSIX) |
|
|
| `orcad_data_root_shared` | the root is group/world accessible and could not be tightened |
|
|
| `orcad_instance_lock_held` | another live orcad owns this root |
|
|
| `orcad_instance_lock_foreign_identity` | the lock belongs to a different identity |
|
|
| `orcad_data_root_unusable` | the root cannot be created, stat'd or written |
|
|
|
|
A root that is merely too permissive and that we own is tightened to `0700` rather than
|
|
refused — orcad stores credentials there unsealed (no OS keyring on this host), so the goal
|
|
is a private root, and refusing when we could just fix it helps nobody. We refuse when the
|
|
permissions are not ours to fix. Windows is exempt from the owner and mode checks: ACLs are
|
|
not expressible as a POSIX mode, and `statSync().mode` there reports a synthesized one.
|
|
|
|
A dead holder's record is reclaimed (PID plus process start time, so a recycled PID does not
|
|
read as alive). A record belonging to a different identity is never reclaimed.
|
|
|
|
**The lock scopes one role — who is the runtime.** It deliberately says nothing about the
|
|
daemon, which lives under `<data-root>/daemon` and fences its own endpoint with its own PID
|
|
record. A lock that asked "is any process using this root" would refuse exactly the restarts
|
|
a live daemon makes worthwhile.
|
|
|
|
## Supervision
|
|
|
|
### Who supervises orcad
|
|
|
|
An external supervisor (systemd, launchd, a process manager). orcad conforms to it:
|
|
|
|
- **Readiness.** One JSON line on stdout (`--json`), `type: "orca_server_ready"`, published
|
|
after the listener is bound and the daemon verdict is in. There is no separate readiness
|
|
socket; the line is the signal. Set the supervisor's start timeout generously — the daemon
|
|
launch has its own retries and can take tens of seconds on a cold host.
|
|
- **Shutdown.** `SIGTERM` or `SIGINT` starts a graceful stop. A **second** signal exits
|
|
immediately with code 1 rather than being swallowed — a supervisor's second signal means
|
|
its first deadline elapsed, and waiting silently is what turns a stop into a `SIGKILL`,
|
|
the one teardown that skips the daemon handoff. orcad also imposes its own 15s deadline
|
|
and exits 1, so the failure stays attributable instead of arriving as an unlogged kill.
|
|
- **Exit codes.**
|
|
|
|
| Code | Meaning | Supervisor should |
|
|
| --- | --- | --- |
|
|
| 0 | clean shutdown | restart per policy |
|
|
| 1 | startup or shutdown failure | restart with backoff |
|
|
| 78 | configuration fault (bind address, data root, instance lock) | **not** restart |
|
|
|
|
78 is `EX_CONFIG`. Put it in systemd's `RestartPreventExitStatus`: restarting on a data
|
|
root owned by someone else is a restart-spin, not a recovery.
|
|
- **Logs.** orcad writes human-readable diagnostics to **stderr** and its readiness contract
|
|
to **stdout**; the supervisor owns capture and rotation. The daemon, being detached, writes
|
|
its own NDJSON lifecycle log to `<data-root>/logs/daemon.log` (suppressed by
|
|
`ORCA_DIAGNOSTICS_DISABLED=1`). Rotation of that file is not implemented — see
|
|
[What is not covered](#what-is-not-covered).
|
|
|
|
### orcad supervising the daemon
|
|
|
|
- **Launch.** Forked detached from `daemon-entry.js` beside `orcad.js`, with its own PID
|
|
record, token and socket under `<data-root>/daemon`.
|
|
- **Adoption before spawn.** A daemon already answering the endpoint is adopted, not
|
|
replaced, unless it is unhealthy, foreign, or built from a superseded bundle *and* owns no
|
|
live sessions. Replacing a healthy daemon kills its PTYs, so code freshness always defers
|
|
to live work.
|
|
- **Restart.** The adapter respawns the daemon on death, transparently to callers.
|
|
- **Crash-loop containment.** At most **5 launches per 60s rolling window** per orcad run;
|
|
past that, launches are refused with `daemon_crash_loop` and terminals fail with that
|
|
message instead of the process forking forever. The window slides, so a repaired host
|
|
recovers without restarting orcad. An operator-initiated daemon restart clears it — that
|
|
is the deliberate "try again".
|
|
- **No macOS login-session watch.** That watch retires the daemon when the spawning GUI login
|
|
session dies. An orcad daemon must survive its SSH session ending.
|
|
- **Shutdown.** orcad never stops the daemon. A daemon that was never adopted retires itself
|
|
after its adoption window; an adopted one stays resident (see Decommissioning).
|
|
|
|
### Decommissioning
|
|
|
|
The daemon outliving orcad is deliberate, so stopping orcad does **not** leave the host with
|
|
zero Orca processes. A daemon that has been adopted stays resident after its runtime
|
|
disconnects — that is what makes the next start a reattach rather than a cold restore. To
|
|
retire a host completely, stop orcad and then stop the daemon named by
|
|
`health.terminalDaemon.pid`, or delete the data root and let the endpoint go stale.
|
|
|
|
## Health
|
|
|
|
The readiness payload carries a `health` object:
|
|
|
|
```
|
|
buildHash sha256 (16 hex) of the running orcad bundle — build identity that a version
|
|
string cannot give, so a rollback that did not replace the file is visible
|
|
buildVersion ORCA_VERSION
|
|
nodeVersion / nodeAbi process.versions.node / .modules — the ABI native addons must match
|
|
platform / arch / pid
|
|
terminalDaemon:
|
|
state live | degraded | absent
|
|
ownsFreshSessions whether NEW terminals are daemon-owned, i.e. survive an orcad restart
|
|
pid the live daemon's pid, from its own PID record
|
|
buildVersion the build the LIVE daemon was forked from (may legitimately predate
|
|
this orcad after an update — reporting orcad's version for both would
|
|
hide exactly that)
|
|
entryPath / protocolVersion
|
|
selfTest { ok, coverage, verdict, durationMs }
|
|
```
|
|
|
|
### What the self-test proves
|
|
|
|
`selfTest` runs `checkDaemonHealth` against the daemon's socket. It is green only when the
|
|
daemon **opened its socket, completed the protocol handshake, and ran `ptySpawnHealth` — a
|
|
real short-lived PTY spawned inside the daemon's own process**. It therefore spans both
|
|
processes: orcad drives it, the daemon performs it, the verdict crosses the socket.
|
|
|
|
- `coverage: 'pty-spawn'` — the full round trip above.
|
|
- `coverage: 'handshake'` — **win32 only**, where `checkPtySpawnHealth` returns without
|
|
spawning anything. A green verdict there covers the handshake and nothing more. It is
|
|
reported separately rather than folded into `ok` so nobody reads it as a PTY round trip.
|
|
|
|
`state` is `live` only when the self-test passed **and** `ownsFreshSessions` is true. A
|
|
daemon that answers but has fallen back to local spawning for new terminals is `degraded`,
|
|
because those terminals die with orcad. A daemon that answered and then failed its spawn
|
|
probe is also `degraded`, not `absent`: it still holds live sessions, and calling those
|
|
exited would be the verdict `ssh-execution-boundary.md` forbids guessing.
|
|
|
|
## What is not covered
|
|
|
|
Named here so nothing reads as implemented that is not:
|
|
|
|
- **A continuous health endpoint.** `health` is published once, in the readiness payload. A
|
|
supervisor's periodic liveness/readiness probe needs an HTTP or RPC surface over the same
|
|
`collectOrcadHealth()`; that surface does not exist yet.
|
|
- **libc slot.** §04 asks for it in the health payload. It belongs to the native strategy
|
|
(plan item 5), which owns libc detection; there is no honest value to publish until then.
|
|
- **`degradations[]`.** Plan item 2's contract, not this one.
|
|
- **Credential administration** (list / revoke / rotate devices, expiring pending offers,
|
|
structured security logging) — §04, not delivered here.
|
|
- **Pinned-port fail-closed.** A pinned `--port` still falls back to an OS-assigned port on
|
|
conflict.
|
|
- **Reconciling `webClientUrl` with reachability** under the loopback default.
|
|
- **State-schema rollback rules.**
|
|
- **Daemon log rotation.** `<data-root>/logs/daemon.log` grows unbounded.
|