mirror of
https://github.com/stablyai/orca.git
synced 2026-10-08 08:02:32 +00:00
cf71ae4cb6f8202a7cc8a424aa84d127c845cbe2
5
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
88f2f01061 |
fix(daemon): escape the terminal daemon into its own systemd scope so a service restart no longer kills every live PTY (#19430)
* fix(daemon): escape the terminal daemon into its own systemd scope so a service restart no longer kills every live PTY Root cause: daemon-launched-child.ts forks the detached terminal daemon with detached: true, which escapes the POSIX process group (setsid) but never the systemd cgroup. Every PTY the daemon owns is itself an undetached direct child of the daemon (native-pty-spawn.ts). Under a combined systemd unit (Type=simple, KillMode=mixed, per docs/reference/headless-linux-server.md), a systemctl restart/stop SIGKILLs every process still in the cgroup at the stop timeout -- the daemon and every live terminal -- even though the codebase already has a fully-built adoption/reattachment path for a surviving daemon (orcad-entry.ts's refreshRestoredOrchestrationAuthority + reconcileLegacyWorkerTerminals, gated on daemonOwnsFreshPersistentPtys()). That path never fires today because the daemon never survives long enough. Fix: when systemd is actually supervising the process and the OS user has a reachable systemd --user manager (isDurableDaemonScopeSupported(), Linux only), launch the daemon via systemd-run --user --scope so it lands in a cgroup that is a sibling of the service unit's cgroup, not a descendant of it. A systemctl restart of the combined unit then never reaches it. Any failure of the scoped launch (no reachable bus, D-Bus policy rejection, etc.) falls back transparently to the existing plain fork() launch, so every platform/environment without this capability is unaffected. The daemon self-detects its own resulting cgroup scope via /proc/self/cgroup (detectOwnCgroupScopeUnit()) rather than trusting the launcher's intent, and publishes it as cgroupUnit in its pid record and orcad's health/readiness payload (health.terminalDaemon.cgroupUnit), so a running deployment can be observed to confirm the fix actually engaged. No new session registry is added: the existing daemon pid-record + adoption protocol (publishDaemonPidFile, daemon-pid-record-quarantine.ts's dead-record reclaim, refreshRestoredOrchestrationAuthority) already implements durable, crash-safe reattachment for a surviving daemon -- it was simply never exercised against a full unit restart before now. Proven via a systemd-in-Docker recovery test: a live PTY session's shell process, its daemon, and the daemon's cgroup scope were all confirmed unchanged across a real systemctl restart of a Type=simple/KillMode=mixed unit, while the main process pid changed (confirming the unit actually restarted) and the new process's health payload recognized the surviving daemon as adopted and live. A fresh write into the same PTY post-restart reached the same running shell. Ordinary terminal create/work/release and the #18789/#18790 worker-release reap-fix regression tests are unaffected. Fixes stablyai/orca#19408 * fix(daemon): probe the real per-UID XDG_RUNTIME_DIR before trusting the process's own env isDurableDaemonScopeSupported()/buildDurableDaemonScopeCommand() trusted the current process's own XDG_RUNTIME_DIR env var first, falling back to /run/user/<uid> only when that var was unset entirely. On mtl-02, orca-serve@factory.service's RuntimeDirectory= hardening directive makes systemd export XDG_RUNTIME_DIR=/run/orca_serve/factory into the unit's process -- a private scratch dir that shares the env var's name but has nothing to do with the user session bus. /proc/<pid>/environ on that host confirmed exactly that path plus DBUS_SESSION_BUS_ADDRESS=disabled:, while the real bus was reachable the whole time at /run/user/985 (confirmed via systemctl --user is-system-running with that dir exported by hand). The probe treated the hardened override as authoritative, found no bus socket there, and reported unsupported on every launch -- so the cgroup-escape fix from #19408/#19430 never actually engaged on real hardware, even though tonight's factory deployment picked it up. Fix: resolveUserRuntimeDir() now always tries the conventional /run/user/<uid> path first (computed independently via getuid(), never trusted from env), checking for a genuinely connectable bus socket via statSync(...).isSocket() rather than a bare existsSync. It falls back to the process's own XDG_RUNTIME_DIR only when that canonical path has no reachable bus -- covering hosts that legitimately have no /run/user/<uid> at all but do have a working bus wherever their own environment points. buildDurableDaemonScopeCommand() now explicitly sets XDG_RUNTIME_DIR to whichever path this resolution picked, rather than inheriting the spread env's (possibly hardened-wrong) value. Both isDurableDaemonScopeSupported() and buildDurableDaemonScopeCommand() gained an injectable canonicalRuntimeDir parameter (defaulting to the real computed path) so tests can exercise the hardened-override scenario deterministically with a real, connectable AF_UNIX socket fixture instead of the live host's actual runtime directory. Docker's stock jrei/systemd-ubuntu test container never had this hardening directive, so this gap was structurally invisible to the container-based verification in #19430 -- only caught against real mtl-02 hardware. * fix(daemon): report the daemon's own pid over the ready handshake, not systemd-run's The launcher used to infer the daemon's identity pid from the immediate spawned child (`child.pid`). On the durable-scope path that child is `systemd-run --user --scope`, not the daemon, so the launcher was asserting an identity it had no authority over. `DaemonReadyIdentity` now carries a required `pid` populated from `process.pid` inside the daemon itself, and `daemon-launched-child.ts` takes `launchedIdentity.pid` from that self-report. Both sides of the `holdDaemonAdoptionLease` pid comparison therefore originate inside the daemon process, which is the idiom this branch already uses for cgroup membership (`detectOwnCgroupScopeUnit` reads `/proc/self/cgroup` rather than trusting what the launcher intended). Note on the reported consequence: `systemd-run --scope` registers its *own* pid on the transient scope unit and then `execvpe()`s the target command -- same pid, no intermediate process -- so adoption did not in fact fail on systemd >= 206 (verified against systemd 255.4-1ubuntu8.17 and current main, `src/run/run.c` `start_transient_scope()`). The fix stands on its own merits: it removes a silent dependency on that exec-vs-fork implementation detail, which a `systemd-run` shim earlier in PATH or any future systemd change would have broken with no diagnostic. `terminateLaunchedDaemonChild` was audited and deliberately left on `child.pid`: for the same execve-preserves-pid reason that pid is either still systemd-run mid-scope-setup (killing it correctly aborts the launch) or already the daemon, so it targets the right process either way. Regression coverage: `daemon-launched-child-identity.test.ts` pins the identity source, and `daemon-ready-identity.test.ts` gains pid-validation cases. Ready-message fixtures across the `daemon-init-*` suites were updated for the now-mandatory field. Addresses: https://github.com/stablyai/orca/pull/19430#discussion_r3953722704 https://github.com/stablyai/orca/pull/19430#discussion_r3954346518 * test(daemon): assert cgroupUnit in the pid-file parse contract `parseDaemonPidFile` returns `cgroupUnit` on every branch as of the durable-scope commit on this branch, but five exhaustive `toEqual` assertions in daemon-health.test.ts still described the pre-scope shape, so they failed on the branch independently of any later change. Adds the field to those expectations. Deliberately not relaxed to `toMatchObject`: asserting the full parsed shape is what makes these tests catch a field silently dropped from the pid-file contract. * refactor(daemon): resolve the canonical user runtime dir at one point The per-UID path cannot change for a live process, so compute it once into a module const instead of threading the same default call through three signatures, and drop the try/catch around a getuid() that cannot throw once it exists. Trims the module prose to the non-obvious facts and corrects the pid-file record comment: an unscoped daemon writes null; only records no daemon wrote are absent. * test(daemon): clean up the cgroup-scope fixtures and assert a verdict The cgroup fixture tracked only the file it wrote, leaking one temp dir per case. Drains both fixture lists with splice so the pop-may-be-undefined guards go away, and replaces a not-throw/typeof-boolean pair with the verdict it was circling: no resolvable runtime dir means unsupported. * refactor(daemon): share the detached child options across both launch paths cwd, detached and stdio were repeated in the fork and systemd-run branches, which left the two comments explaining them hovering over the env block instead. Names them once so each branch carries only its own delta. * refactor(daemon): validate the ready pid like every other field typeof-first narrows the value, so the two 'as number' casts the isSafeInteger check needed disappear and the pid guard reads like the startedAtMs guard below it. * fix(daemon): don't retry the launch unscoped after losing the endpoint race A scoped attempt that lost the endpoint to another daemon was retried unscoped: a second doomed fork, a misleading 'cgroup-scope launch failed' warning, and the same DaemonEndpointUnavailableError the caller was already going to adopt on. Rethrows it instead, since no launch mode can win a race that is already lost. Also drops a private alias for DaemonChildSpawnOptions and the two 'as number' casts on child.pid in the startup-failure cleanup. * fix(daemon): unlink the pid record by the pid the daemon published The record holds the daemon's self-reported pid, so match on that rather than on the immediate child's, which is the systemd-run wrapper's until it execs. * fix(daemon): route the scope launch through the child-process chokepoint The two files this PR added imported `node:child_process` directly, which `child-process-import-boundary.test.ts` fails on deterministically: the offender count went 155 -> 157 against a pin of exactly 155. Raising the pin or listing the files is what that test explicitly forbids, and the allowlist's own note says a split "moved the import, it did not add one" -- so the fix is to get both new files off the module and put the count back at 155. - `daemon-cgroup-scope.ts`: the `systemd-run --version` probe now uses `runProcessSync` instead of `execFileSync`, so it gets the shared spawn decisions. Kept synchronous deliberately: `launchDaemonChild` attaches the readiness listener in the same tick it is called, and an await before the spawn moves the child past that tick. A non-zero exit is data rather than a throw here, so the verdict now checks `code === 0 && !timedOut`. - `daemon-launched-child-spawn.ts`: the scoped launch uses `spawnProcess`, and the long-standing unscoped launch keeps `fork` semantics through a new `forkProcess`. - `src/shared/child-process/fork-process.ts`: the fork arm of the chokepoint. `spawnProcess` cannot express a Node child with an IPC channel started from a module path under an overridden `execPath`, and the existing launch tests are written against `fork`'s contract, so a spawn rewrite would have changed module resolution, `execPath` and `execArgv` at once. It passes `windowsHide: true` -- the flag every other call site in that directory sets, reachable via an assertion because `ForkOptions` omits it -- which keeps `windows-console-visibility.test.ts` at its pin of 65 too. Both ratchets pass with both pins and both allowlists untouched. Docs: `orcad-operations.md` and `headless-linux-server.md` still described the limitation this PR removes as permanent. Both now describe the durable-scope survival path and its preconditions (systemd as PID 1, a reachable user bus / `loginctl enable-linger`, `systemd-run` on PATH), and scope the old text to the unscoped-fallback case, pointing at `health.terminalDaemon.cgroupUnit` as the way to tell the two apart on a running host. * fix(daemon): seal the cgroup capability probe from the host and correct KillMode=mixed docs The capability probe consulted the host's own /run/systemd/system marker and spawned the real systemd-run binary, so the hermetic unit tests could only pass on a systemd host (and fail closed otherwise, even with faked bus sockets). - Thread systemdBootPath and runVersionProbe as test seams through isDurableDaemonScopeSupported, defaulting to the real boot marker and systemd-run --version probe in production. - Narrow the injected probe to the ProcessResult slice it consumes. - Cover: no-systemd-boot, non-zero probe exit, and probe-timeout cases. - Correct KillMode=mixed semantics in the docs: the cgroup-wide SIGKILL fires the instant the main process exits, not after TimeoutStopSec; document the Docker-container caveat and add KillMode=mixed to the multi-service template. * fix(daemon): satisfy assertion checks in scoped launch * fix(daemon): satisfy anti-slop and console guards * test(serve): update shutdown docs assertions for daemon scope * fix(daemon): migrate adopted legacy scopes * docs: qualify restart safety by daemon scope * docs(daemon): qualify Upgrade restart prose with durable scope caveat Align the Upgrade section in docs/reference/headless-linux-server.md with the earlier preservation section and docs/reference/orcad-operations.md: a service restart terminates live processes only when running under the unscoped fallback, and stops should be treated as destructive unless health.terminalDaemon.cgroupUnit names an orca-daemon-*.scope. Update the shutdown workflow test assertion in config/scripts/headless-serve-shutdown-workflow.test.mjs to match. * fix(daemon): harden legacy scope migration --------- Co-authored-by: Lesley Murfin <260182349+LesleyMurfin@users.noreply.github.com> Co-authored-by: m4air <m4air@Mac.localdomain> |
||
|
|
97aa5ff19b |
fix(mobile): open native chat when a new worktree launches a default agent (#19850)
* refactor(agent-launch): make the launch-mode decision surface-neutral
`decideWorkerStartMode` was the only shared answer to "structured chat session
or terminal agent?", but it lived in an orchestration-named module and spoke
orchestration's vocabulary, so the other launch surfaces could not call it.
Move the decision to `main/agent-launch/agent-launch-mode` unchanged and leave
`orchestration-worker-start-mode` as the adapter that supplies the noun.
A worker is not a special kind of launch; it is the same launch with a dispatch
attached. Naming the receipt's subject is the only thing orchestration actually
contributed, so that is the only thing the adapter keeps: "worker" in both
sentences, plus the `--terminal` wording, which reads as nonsense anywhere a
`--terminal` flag does not exist. Both are pinned, because they are asserted.
No behavior change. The receipts are byte-identical for every reachable case,
proven by running the new pin against both implementations.
Also pins the wording, which nothing was holding. The existing suites assert
`toContain` fragments ('terminal agent', 'cannot create') and the CLI suite
asserts a receipt handed to it by a mock rather than one this code produced;
all six files stayed green against a deliberately corrupted vocabulary. A
dispatch receipt is the only place a structured-to-terminal downgrade explains
itself, so the whole sentence is the contract, not a fragment of it.
* feat(agent-launch): add the launch intent and the one executor that runs it
The sequencing around the launch decision was duplicated per surface, and the
duplicate is where the bug lives. A new worktree was created agent-first, so
its startup terminal WAS the agent and the structured branch below it could
never be reached — every new-worktree launch was a PTY regardless of the user's
default. Orchestration fixed that for itself in #19431; mobile and the CLI
still have it.
`executeAgentLaunch` inverts the order once, for everyone. When the preference
is structured the worktree is created with NO startup agent, the executing host
is then asked whether it can host a session for the workspace that now exists,
and only then is a surface created. The host verdict cannot be hoisted above
creation: `agentSession.createSupport` only answers for a workspace it can
resolve, which is why the decision stays in two halves.
Agent-first creation is deliberately preserved for PTY launches — it is what
sequences the agent's startup command behind the setup runner, so wait-for-setup
comes for free there.
What actually differs per surface is only how a surface is built (an
orchestration worker's session takes a dispatch hold and a mailbox a plain
launch must not take), so that is injected as a factory rather than branched on.
The intent also strips the reserved agent fields from a migrated create payload:
a caller moving off `worktree.create` passes its existing params, and a stale
`startupAgent` in there would re-create the very path this replaces.
Tests assert order and arguments, not just the resulting mode. Reintroducing
agent-first creation reddens 4 of 11.
* feat(agent-launch): expose the launch executor as the agent.launch RPC
Adds `agent.launch` — one host-side method that decides structured-vs-terminal and
creates the surface — wired to the real runtime factories: `createManagedWorktree`
for the workspace, forking on `startupAgent` exactly as the orchestration worker
path does; `createStructuredAgentSessionForWorktree` for a chat session; and
`createTerminal` for a PTY agent. Allowlisted for mobile, which is the surface the
routing gap was reported on.
`worktree.create` is untouched. Its `startupAgent` keeps meaning "spawn a PTY agent"
verbatim, because it answers with `agentTerminalHandle` only on that path: a host
that quietly routed it to a structured session would hand every older client a
response with no handle and no error. All new behaviour sits behind
`agent.launch.v1`, which the host now advertises and a remote client must negotiate,
so a client that does not gets today's behaviour unchanged.
* feat(mobile): route workspace creates through agent.launch
Picking an agent on the mobile create sheet always produced a terminal, even
when the user's default was native chat, because all three create paths put
`startupAgent` on `worktree.create`. That means "create the worktree
agent-first", so its startup terminal IS the agent and the structured branch
below it is unreachable — while the same phone's in-workspace "+" button opened
a chat.
The blank, branch and new-branch creates now send the same payload through
`agent.launch` and let the host settle the surface. `worktree.create` is
untouched, and a host that does not advertise `agent.launch.v1` (read from the
existing `status.get` probe) keeps today's path exactly.
Work-item creates stay on `worktree.create`: they pre-fill the issue/PR URL as
an unsent `startupDraft`, which a structured session cannot hold yet, so routing
them would submit the URL as a first turn.
* fix(agent-launch): drop the deleted draft-prompt blocker from the reason map
main removed the draft-prompt blocker in #19681 (a structured session now holds
an unsent draft), so the exhaustive Record no longer typechecks.
* chore(agent-launch): carry a SAFETY rationale on the agent placement cast
The type-assertion gate landed after this branch's base, so the new file's
copy of the worker-start cast is now a changed-code finding.
* chore(agent-launch): carry agent.launch through main's RPC typing and casting gates
The typed-method contract, the generated params catalog and the
`assertionStyle: never` casting scan all landed after this branch's base.
- AGENT_LAUNCH_METHODS kept an `RpcMethod[]` annotation, which widened its
method name to `string` and broke assignability; every sibling infers instead.
- `agent.launch` binds a schema under src/main, so it joins the catalog's
RPC_METHODS_WITHOUT_SHARED_PARAMS and the parity gate's hand-listed twin.
- The now-typed methods make most test casts unnecessary; the few that remain
carry the line-specific SAFETY rationale the casting gate requires.
* test(mobile): supply the agent-launch fixture the create-submit recording needs
The golden RPC recordings landed upstream while this branch was out, so they
first met agent.launch here. Three things had to happen, and only one of them is
a fixture bump.
1. workspace-settings-mounts.ts mounts useNewWorkspaceCreateSubmit against a
fixture model that throws on any member it was not given. This PR added a
required getAgentLaunchSupport, so the submit aborted with "Missing model
fixture" before it ever issued the create, and three cleanup checkpoints
vanished. That read like a product regression and was not one. Supplying the
member restores the recording byte-for-byte; it is pinned false for the same
reason the cutover probe is, so the baseline stays on worktree.create.
2. Editing that adapter moves adapterSha256 for the twelve settings goldens it
mounts. Their recordings are unchanged - header only, by design: the digest
is per-golden so editing a module fails exactly the goldens that mounted it.
3. Five goldens changed behaviourally, and both changes are this PR's:
the capability probe now reports agentLaunch, and a create whose reply
carries no worktree returns "Failed to create workspace" instead of throwing
a TypeError off an unguarded result.worktree read. The launch route needs
that guard, since a receipt can arrive without a worktreeId.
* refactor(mobile): decode the launch receipt instead of asserting its shape
The changed-code quality gate refuses type assertions, and the eight it flagged
were worth removing rather than suppressing.
The production one was the point. readAgentLaunchCreateOutcome asserted the RPC
payload into Partial<AgentLaunchResult> and then runtime-checked it anyway, so
the assertion bought nothing and claimed a contract the host had not proven. It
now narrows with `in` and validates each hop, which is the same nullability
question readCreateResult already answers on the sibling path - a launch receipt
can legitimately arrive without a worktreeId. AgentLaunchCreateOutcome ties
worktreeId to the shared contract so a change there fails this reader's
typecheck rather than passing a differently-typed field through.
The test fakes claimed a whole RpcClient via `as unknown as RpcClient` while
implementing one member. They now build a typed literal, matching the pattern in
use-mobile-structured-agent-options.test.ts. The read sites cast params and then
read one field; they now assert the payload with toMatchObject, which removes
the cast and pins more of the shape than the cast did.
Also pins the warning passthrough, which nothing covered: a terminal launch that
seats the workspace but cannot start the pty reports why, and the absent, blank,
non-string and structured-surface cases report nothing. Writing that test caught
a real drop I had introduced in the reader.
* ci(mobile): re-run Mobile Checks when a shared capability changes
Mobile Checks is path-filtered to mobile/**, but mobile imports the negotiated
capability names straight from src/shared/protocol-version.ts and records the
whole capability read verbatim in its goldens. So a capability added desktop-side
rewrites a mobile fixture while never triggering the suite that would catch it.
That is what happened here: #19849 introduced agent.launch.v1 and Mobile Checks
never ran on it. Verified at the run level rather than by check name - the
window-free check-runs API on
|
||
|
|
aee98ccaa0 |
fix(browser): make the browser identity one process-wide choice (#13822) (#20767)
* feat(browser): process-wide browser identity, chosen before ready
Electron resolves worker identity from a single process-global default, so two
coherent identities cannot coexist in one process. This makes clean/native one
app-wide decision read before `ready`, instead of a per-profile one that leaves
documents on one identity and every worker request on the other.
Both identities are load-bearing, measured across four origins at five reps:
the cleaned identity clears an embedded Turnstile widget and WhatsApp's browser
check where native is refused; native clears a full-page Cloudflare interstitial
that the cleaned identity never clears.
Base commit only: removing the per-profile field, its settings surface, and the
migration notice follow.
* test(browser): cover cross-context UA wire identity
* refactor(browser): make user agent identity app-wide
* test(browser): repair process identity wire fixture
* Fix browser identity startup migration failures
* WIP: rescue in-flight reduced-design work from a dead worker
Worker ctx_cb5b1262d7fe stopped ~2h ago mid-implementation (last heartbeat
2026-09-14T22:48:06Z) leaving this uncommitted. Committed unverified to make it
recoverable; not reviewed, not necessarily green.
* fix(browser): repair the rescued identity work so it typechecks
Finishes the interrupted edits in
|
||
|
|
5631aa00dd |
feat(orcad): items 2–7 — degradation, natives, daemon, ops, deploy (#16398)
* fix(ports): stop joining an undefined resourcesPath on a non-Electron host `resolveWorkerEntryPath` branched on `isPackaged` alone and joined `process.resourcesPath`. orcad reports `isPackaged` true — correctly, it is a production build, and ~15 consumers read it that way to gate HTTPS-only skill downloads and the real CLI name — but `process.resourcesPath` is Electron-only and `undefined` under plain Node. So the packaged branch threw `TypeError [ERR_INVALID_ARG_TYPE]: The "path" argument must be of type string` where a clean "worker unavailable" was the honest outcome. The type said `resourcesPath: string`, which is how it went unnoticed; it is now `string | undefined`, so the compiler carries the fact. A host with no Electron resources tree has no asar to look in, so it falls back to the module directory and lets the caller report a missing worker. Found by the item 1 agent while auditing the same `isPackaged` defect class in the watcher. Verified in both directions: reverting the guard reproduces the TypeError. * feat(orcad): prove node-pty loads before anything requires it Of the two ways node-pty fails, only one is catchable. A missing module throws MODULE_NOT_FOUND. A module built against the wrong libc or Node ABI is refused by the dynamic loader, and in the worst case takes the process down before any handler exists — that is #9902, which crashed the desktop app on Ubuntu 20.04 before a window appeared. There was no libc or ABI precondition anywhere in the tree. So orcad now proves the load in a CHILD process, from main.ts, before anything requires node-pty. Whatever the child does — throw, abort, die on a signal — is data rather than our own death, and the operator gets a sentence naming the host's libc, Node ABI and prebuild slot plus the command to run. Proven-unloadable exits 78 (EX_CONFIG), so a supervisor does not restart an unequippable host forever. A probe that never answered is unverifiable, not blocked: refusing to boot on an inconclusive signal would take down hosts that work. The child dlopens the file node-pty would have chosen, before requiring the package. node-pty's loader walks several directories and rethrows only the LAST error, so a refused binary reads as "Cannot find module ./prebuilds/..." — which sends the operator to install a module that is already there. It also reports through stdout: node echoes the whole -e source above a stack trace, and matching tokens against stderr made the probe's own source text answer for the verdict. Verdicts reach clients as a terminal_unavailable degradation alongside the existing browser_unavailable one, through the same cause-registry shape. degradations[].code is now an open vocabulary; clients already render only `message`. Prebuilds are compiled from PATCHED sources — the patch IS the glibc-floor fix, so an upstream tarball reproduces #9902 — into linux-{x64,arm64}-{glibc,musl} and darwin-{x64,arm64} slots. libc is in the slot name because node-pty's loader falls back to prebuilds/<platform>-<arch> and cannot tell glibc from musl. orcad installs the matching slot at boot, so a host with no compiler serves terminals. The relay's five pure toolchain-diagnosis functions moved to a transport-free module so the Node bundle can reuse them without dragging ssh2 in behind them; the relay keeps its API by re-export. macOS gets `xcode-select --install` rather than the cross-distro apt/dnf/pacman/apk menu, every line of which is wrong there. * test(orcad): pin the node-pty precondition to ground truth, not a prepared host CI's test shard runs `vitest` directly, so `ensure-native-runtime --runtime=node` never prepares node-pty for the Node ABI — `degraded` is the correct verdict there, and asserting 'ok' encoded an environment the shard does not have. Asserting whatever it returned would be vacuous, so the expectation is now derived from an independent require() of node-pty. Verified it still bites: forcing the precondition to always report 'ok' fails the suite. * feat(orcad): run the terminal daemon, and the ops contract around it orcad declared `canRecoverPersistentLocalPtys: () => false` because it did not run the terminal daemon, so every restart, update and rollback SIGKILLed every running terminal — on the host whose selling point is that work survives the client going away. That is the one property `ssh-execution-boundary.md` recommends the peer model for. Item 4 — the daemon: - Port the launch path off electron: `daemon-init.ts`, `daemon-host-relocation.ts` and `observability/logs-directory.ts` now read the `AppEnvironment` port. Relocation additionally asks whether the app root is an asar archive rather than whether the build is packaged, so a Node host answering `isPackaged() === true` no longer walks into an Electron-only NSIS-escape path (same precedent as `parcel-watcher-entry-path.ts`). - `build-orcad.mjs` emits `daemon-entry.js` beside `orcad.js`, scans the forked children's metafiles for electron/node:sqlite, and load-checks the child under plain Node. - orcad spawns and adopts the daemon; shutdown disconnects and never kills it. `canRecoverPersistentLocalPtys` now reads the live provider and is false under degraded routing, where fresh terminals would die with the process. Item 3 — the ops contract (docs/reference/orcad-operations.md): - Bind policy: `--bind`, default loopback, pinned so neither `orca serve`'s wide default nor the connected-device widen can override it, and so a paired client cannot rebind the listener from outside. - Instance lock on the data root before profile load, scoped to the runtime role so it never refuses a restart that a live daemon makes worthwhile. - Supervision: exit codes a supervisor can act on (78 = do not retry), second-signal escalation, a shutdown deadline, and crash-loop containment on daemon respawn. - Health in the readiness payload: build hash, Node ABI, and a PTY self-test that spans both processes — the daemon spawns a real PTY in its own process and the verdict crosses its socket. Both bundle load-checks now assert on exit codes: these bundles are minified onto one line, so Node's uncaught-exception report echoes every string literal in the bundle and the previous message match passed against a bundle that never loaded. * feat(orcad): deploy, activate and roll back a versioned orcad install Plan items 6 and 7 from docs/design/shipping-orcad.html. Install reuses the relay's transaction verbatim — per-version lock, staged SFTP write, .install-complete sentinel, stale-lock recovery — under a parameterized namespace, so orcad-<v>/ sits beside relay-<v>/ permanently (§06). Parameterizing GC is the trap that creates: each model now collects only its own directories, enforced twice (prefix-scoped remote listing plus a local ownership re-check), and a client picks its model from how the host is registered, never from what it finds on disk. Activation is separate from installation, because a versioned directory selects nothing. A candidate is launched, publishes orca_server_ready, and only becomes active if its cross-process health payload passes: right build hash, listening, daemon live, PTY self-test green. A rejected candidate is stopped and the incumbent restarted, so a careful deploy cannot cause the outage it was being careful about. Update and rollback are shaped by the daemon. An update restarts orcad, the daemon outlives it, and the surviving daemon was forked from the outgoing bundle — so live terminals defer the update rather than proceed, and GC pins the active version, the rollback target and the live daemon's bundle. Orca's persisted state carries no schema version, so rollback restores a pre-activation snapshot rather than trusting backward-readability; the point past which it is unsafe is the first terminal created after activation, which the snapshot cannot describe and the surviving daemon still owns. Running the generated shell for real found two bugs the text assertions missed: tar members re-quoted inside a shell variable captured nothing, and kill -0 reports a zombie as alive. * test(orcad): assert the precondition is self-consistent, not environment-shaped The real-host case cannot predict a status: CI's shard runs vitest directly, so node-pty is never built for the Node ABI and 'degraded' is correct there, while a prepared checkout gives 'ok'. The previous attempt used require('node-pty') as ground truth, which resolves the JS wrapper while the native binding loads lazily — it proved strictly less than the precondition checks, and failed CI for exactly that reason. What is invariant on a host with node-pty installed: never 'blocked', and never a degraded verdict carrying an unestablished reason. The injected-input tests keep the logic coverage. * fix(orcad): drop an eslint-disable the rule no longer needs * test(orcad): separate slot placement from the load verdict Both remaining CI failures were the same shape: tests reaching into node_modules for a pty.node that only exists after `ensure-native-runtime --runtime=node`, which CI's shard never runs because it invokes vitest directly. Slot *placement* is the logic worth checking on every host, so it now uses a synthetic payload and asserts the verdict stays honest about not loading. The three assertions that genuinely need a Node-ABI binding are gated on it existing. Verified: breaking slot installation fails both placement tests; with the real pty.node hidden the file is 17 passed / 3 skipped instead of ENOENT. * test(orcad): gate the load-dependent cases on a real load, not on the file existing CI ships a pty.node built for Electron's ABI, so existsSync was true while require still failed — the gate ran exactly the tests that host can never satisfy. It now probes the binding in a child process, so a bad one cannot take the runner down. The self-consistency assertion also allowed too little: 'blocked' is the honest verdict for a corrupt binding, alongside 'ok' on a prepared host and 'degraded' on an unprepared one. What stays invariant is that anything other than 'ok' names an established cause, so a terminal is never declined for a reason nobody worked out. Verified against all three host states: prepared (19 passed), unprepared, and a corrupt binding (17 passed / 3 skipped, no failures). * test(orcad): gate on the whole premise — binding AND spawn-helper CI has a loadable pty.node but no spawn-helper, and a slot without the helper is legitimately 'degraded'. So the previous gate let a test run whose premise ('a complete slot yields ok') that host cannot satisfy. Verified in both states: with the helper present 19 pass; with it removed the load-dependent cases skip (17 passed / 3 skipped) instead of failing. * fix(orcad): preserve degradation types after rebase |
||
|
|
34c160442f | Fix headless Linux serve pairing readiness (#9785) |