Files
orca/docs/reference/orcad-operations.md
T
OrcaWinandm4air 5f308bfa9c revert: take the 26 Phase 3 (#16741 port) PRs back out of main (#24559)
* Revert "feat(orcad): source-side dormant export of a relay-hosted SSH target (#16741 T6-8) (#24519)"

This reverts commit 783101b304.

* Revert "feat(ssh): update, roll back, recover and stop a managed orcad server (#16741 T6-5 follow-up) (#24463)"

This reverts commit 38c2d1dcb9.

* Revert "feat(ssh): deploy and pair an empty managed orcad server over SSH (#16741 T6-5) (#24453)"

This reverts commit 8b76683b40.

* Revert "fix(ssh): orcad GC honors the activation journal; readiness requires proven daemon coverage (#16741 T6 follow-up) (#24451)"

This reverts commit d3f8c5063b.

* Revert "feat(ssh): remote orcad stop by request file and journaled decommission (#16741 T6-4) (#24449)"

This reverts commit 43d9b43d3f.

* Revert "feat(orcad): supervisable server: stop requests, managed stop receipts and a lifetime that keeps its lock on failed teardown (#16741 T6-3) (#24433)"

This reverts commit b093d3ab20.

* Revert "feat(ssh): crash-safe orcad activation, rollback and recovery (#16741 T6-2) (#24423)"

This reverts commit 1a9ac0e955.

* Revert "feat(runtime): SSH access links for paired servers in a downgrade-safe sidecar (#16741 T5-1+T5-2) (#24420)"

This reverts commit 99db2bfae4.

* Revert "feat(relay): capability-gated owner reset with a durable preparation journal (#16741 T3 R1) (#24418)"

This reverts commit 34a582bd39.

* Revert "feat(ssh): track connection-manager drains, test probes and provider continuations (#16741 T2 P3+P8a) (#24407)"

This reverts commit d53063d2b1.

* Revert "feat(daemon): idle retirement, session census and recovery-only provider (#16741 T2 P4b) (#24409)"

This reverts commit ff212dbbef.

* Revert "feat(ssh): add pty.resumeClient and split SSH PTY process listing (#16741 T2 P5+P6) (#24414)"

This reverts commit 92cb71765e.

* Revert "feat(relay): await owned watcher and agent children on shutdown (#16741 T2 P1) (#24400)"

This reverts commit 6b36e4f85b.

* Revert "feat(session): retry failed renderer session writes and verify local folder PTYs (#16741 T2 P9) (#24406)"

This reverts commit d23ecef301.

* Revert "feat(ssh): remote orcad primitives on the pinned Node runtime (#16741 T6-1) (#24419)"

This reverts commit dd87ae578d.

* Revert "fix(runtime): fence runtime-environment subscriptions and status probes by identity (#16741 T5-3) (#24421)"

This reverts commit ece9e4d2e3.

* Revert "feat(orcad): migration manifest and dormant-state contracts (#16741 T6-7) (#24422)"

This reverts commit 3fbdaba262.

* Revert "feat(ssh): wire SshConnection through the work and transport close ledgers (#16741 T2 P2) (#24401)"

This reverts commit 4e8edc8872.

* Revert "feat(profiles): carry markdown frontmatter visibility in project transfers (#16741 T2 P7) (#24405)"

This reverts commit 60c93263cc.

* Revert "fix(runtime): project the PTY incarnation onto mobile session tabs (#24413)"

This reverts commit 99e0303572.

* Revert "feat(daemon): tag daemon stream data with the PTY incarnation id (#16741 T2 P4a) (#24402)"

This reverts commit 817af768b0.

* Revert "feat(ssh): port the SSH connection work ledger and transport close ledger (#16741 T2) (#24210)"

This reverts commit c9918931c8.

* Revert "feat(relay): fence and drain file and git response streams on shutdown (#24185)"

This reverts commit dc08ffeba9.

* Revert "refactor(runtime-rpc): extract the Node WebSocket lifecycle; opt-in pinned port (#24186)"

This reverts commit a789233bbb.

* Revert "feat(relay): route relay handlers through work admission; producer publication drain (#24181)"

This reverts commit 0b812bd698.

* Revert "feat(relay): land the #16741 T1 seam (work drain, publication drain, release gate) (#24156)"

This reverts commit 3aa2d3af7c.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
2026-10-02 00:52:32 -07:00

266 lines
16 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Running orcad
`orcad` is the Orca runtime served from plain Node. This is the contract between it and
whatever supervises it: what it binds, what it owns on disk, who restarts what, and what its
readiness payload actually proves.
## Two long-lived processes, not one
A deployment is **orcad** plus **the terminal daemon**.
| | orcad | terminal daemon |
| ---------- | -------------------------------- | ------------------------------------- |
| Started by | the supervisor | orcad, detached |
| Owns | RPC, git, worktrees, persistence | every local PTY |
| Lifetime | one supervised run | detached from orcad, not its service |
| Endpoint | `ws://<bind>:<port>` | `<data-root>/daemon/daemon-v<N>.sock` |
orcad detaches the daemon and calls `disconnectDaemon()`, never `shutdownDaemon()`. The
built-in remote deployment path stops only the recorded orcad PID, so the daemon and its PTYs
survive. The successor adopts the current endpoint and routes supported previous protocol
versions through legacy adapters. This makes a PID-scoped update, rollback or restart
non-destructive to live work.
Process detachment is not service isolation. A daemon that orcad launches directly, and every
PTY it owns, remain in the same systemd service cgroup. `KillMode=mixed` does **not** preserve
them: it sends the graceful stop signal only to the main process, then sends `SIGKILL` to every
process remaining in the cgroup the moment that main process exits — `TimeoutStopSec` never gets
the chance to apply. `KillMode=control-group` is destructive too. `KillMode=process` leaves
service-owned processes unmanaged and is not a supported preservation mechanism.
Service-restart survival therefore requires a separately supervised cgroup, and orcad now asks
for one: on Linux it launches the daemon through `systemd-run --user --scope`, which places the
daemon and its PTYs in their own transient `orca-daemon-<launch-nonce>.scope` unit under the
user slice instead of the caller's service cgroup. A stop or restart of the service unit then
leaves that scope — and the live terminals in it — running, and the successor adopts the
endpoint as it always has.
A newly launched private daemon scope also follows the daemon's own lifetime. A small
detached shell holds an input pipe from the daemon; after that pipe closes and `/proc`
confirms the daemon PID is gone, it asks the user manager to stop that exact scope.
Systemd sends remaining processes SIGTERM and escalates after five seconds. This includes
children that double-forked or called `setsid` and can no longer be found by parent PID.
Disconnecting or restarting the runtime does not close the pipe: the daemon owns it.
The cleanup only arms on a fresh scoped launch with a matching launch nonce. Adopted
legacy scopes can contain GUI processes and are never armed retroactively. Unscoped
launches and children deliberately moved into another systemd unit remain outside this
cleanup. `nohup`, `disown`, and `tmux` alone do not move a process out of its cgroup, so
those children now end when their terminal daemon dies. Work intended to outlive that
daemon needs its own service or scope.
The scope is requested only where it can work. All of these must hold:
- **Linux with systemd as PID 1** (`/run/systemd/system` exists).
- **A reachable user bus** — a connectable `bus` socket in the per-UID runtime dir
(`/run/user/<uid>`, or whatever `XDG_RUNTIME_DIR` points at). For a service account that is
not otherwise logged in, that means `loginctl enable-linger <user>`; a unit whose
`RuntimeDirectory=` hardening moves `XDG_RUNTIME_DIR` off the per-UID path is handled, because
the real per-UID path is probed first.
- **`systemd-run` on `PATH`** and answering `--version`.
Any of those missing, or a `StartTransientUnit` call that fails anyway, falls back to the
direct launch — and in that unscoped fallback case the paragraph above still describes reality:
the daemon shares the service cgroup and a combined-unit stop ends live terminals. Read the
`cgroupUnit` field in the daemon health payload to tell the two cases apart on a running host;
it is populated from `/proc/self/cgroup`, so it reports the isolation the daemon actually has
rather than what the launcher intended.
## Bind policy
`--bind <literal-ip>`, **default `127.0.0.1`**.
Only literal IPs are accepted; hostnames are refused because DNS would decide which
interface got bound. `localhost` maps to `127.0.0.1`. `0.0.0.0` / `::` are the explicit
opt-ins to network reach, and the startup log says so on every launch.
The bind is **pinned**, not defaulted. Two things widen the desktop's listener on their own —
`orca serve`'s wide default, and a startup where some device has connected before — and an
unattended host's exposure must be exactly what the operator asked for on every launch. A
mobile pairing offer, which normally rebinds to all interfaces, is refused while the bind is
pinned to loopback and reports `network_exposure_failed` rather than advertising an endpoint
nothing can reach.
Under the shipping design a client reaches a remote orcad over an SSH local port-forward, so
loopback is the correct default and the pairing credential travels over SSH.
## Data root and the instance lock
The data root is `$ORCA_USER_DATA`, else `$XDG_DATA_HOME/Orca`, else `~/.orca`.
Before the profile index or the store is touched, orcad takes `<data-root>/orcad.lock`.
It refuses to start when:
| Code | Meaning |
| -------------------------------------- | ------------------------------------------------------------- |
| `orcad_data_root_wrong_owner` | the root is owned by another uid (POSIX) |
| `orcad_data_root_shared` | the root is group/world accessible and could not be tightened |
| `orcad_instance_lock_held` | another live orcad owns this root |
| `orcad_instance_lock_foreign_identity` | the lock belongs to a different identity |
| `orcad_data_root_unusable` | the root cannot be created, stat'd or written |
A root that is merely too permissive and that we own is tightened to `0700` rather than
refused — orcad stores credentials there unsealed (no OS keyring on this host), so the goal
is a private root, and refusing when we could just fix it helps nobody. We refuse when the
permissions are not ours to fix. Windows is exempt from the owner and mode checks: ACLs are
not expressible as a POSIX mode, and `statSync().mode` there reports a synthesized one.
A dead holder's record is reclaimed (PID plus process start time, so a recycled PID does not
read as alive). A record belonging to a different identity is never reclaimed.
**The lock scopes one role — who is the runtime.** It deliberately says nothing about the
daemon, which lives under `<data-root>/daemon` and fences its own endpoint with its own PID
record. A lock that asked "is any process using this root" would refuse exactly the restarts
a live daemon makes worthwhile.
## Supervision
### Process-scoped and cgroup-wide stops
The built-in remote updater performs a PID-scoped stop and keeps the daemon's install version
pinned while it owns sessions. A combined-unit systemd stop or restart is different: unless the
daemon holds a durable cgroup scope of its own (see
[Two long-lived processes, not one](#two-long-lived-processes-not-one)), it reaps the daemon and
every live terminal after the graceful window. Treat a stop as destructive unless
`health.terminalDaemon.cgroupUnit` names an `orca-daemon-*.scope` on that host.
Before a cgroup-wide stop, obtain a fresh `orca-ide terminal list --json` result using the same OS
account and home as the daemon. Invoke the installer's absolute launcher path so `sudo`'s
`secure_path` cannot hide a per-user registration (for example,
`sudo -Hu orca /home/orca/.local/bin/orca-ide terminal list --json`). Replace both `orca` and
`/home/orca` with the service account and home used by the unit; an extracted deployment may use
its absolute `resources/bin/orca-ide` launcher instead. A safe empty census is untruncated, has an explicit `hostScope`, covers every
execution host affected by the stop, and lists no terminals on those hosts. Every
`omittedHostIds` entry must be explicitly accounted for outside the target service's execution
boundary. A separately paired runtime is outside that boundary; local execution and SSH hosts
reached through this runtime are not. An affected or unknown omission, missing scope,
truncation, a failed request or lost contact makes the result `unverifiable`: defer the stop. Do
not admit new work after the census. Orca does not yet provide an atomic census-and-stop fence.
### Who supervises orcad
An external supervisor (systemd, launchd, a process manager). orcad conforms to it:
- **Readiness.** One JSON line on stdout (`--json`), `type: "orca_server_ready"`, published
after the listener is bound and the daemon verdict is in. There is no separate readiness
socket; the line is the signal. Set the supervisor's start timeout generously — the daemon
launch has its own retries and can take tens of seconds on a cold host.
- **Shutdown.** `SIGTERM` or `SIGINT` starts one graceful stop. Repeated signals share
that stop because a supervisor may signal both the launcher and its child. A 15s deadline
exits with code 1 if teardown stalls. The bundled runtime also stops gracefully if its
launcher's IPC channel closes. On POSIX, both the launcher and runtime ignore `SIGHUP`,
so terminal hangups do not stop a headless host. Use `SIGTERM` or `SIGINT` to stop it.
- **Exit codes.**
| Code | Meaning | Supervisor should |
| ---- | ------------------------------------------------------------ | -------------------- |
| 0 | clean shutdown | restart per policy |
| 1 | startup or shutdown failure | restart with backoff |
| 78 | configuration fault (bind address, data root, instance lock) | **not** restart |
78 is `EX_CONFIG`. Put it in systemd's `RestartPreventExitStatus`: restarting on a data
root owned by someone else is a restart-spin, not a recovery.
- **Logs.** orcad writes human-readable diagnostics to **stderr** and its readiness contract
to **stdout**; the supervisor owns capture and rotation. The daemon, being detached, writes
its own NDJSON lifecycle log to `<data-root>/logs/daemon.log` (suppressed by
`ORCA_DIAGNOSTICS_DISABLED=1`). Rotation of that file is not implemented — see
[What is not covered](#what-is-not-covered). orcad records every trace span it emits
(git commands, worktree paths, terminal spawns, structured-chat failures and the rest) to
`<data-root>/logs/orcad.trace.ndjson`, rotated at 10 MB × 10 files, private to its user and
redacted for secret-shaped strings. It stays on the host: a desktop's diagnostics bundle does
not collect it. `ORCA_DIAGNOSTICS_DISABLED=1` turns it off, and a logs folder orcad cannot
open leaves it off with one stderr warning rather than stopping orcad.
### orcad supervising the daemon
- **Launch.** On Linux, through `systemd-run --user --scope` so the daemon gets its own
transient cgroup and survives a service-unit restart; everywhere else, and wherever that
scope is unavailable, forked detached. Either way it runs `daemon-entry.js` beside
`orcad.js` with its own PID record, token and socket under `<data-root>/daemon`.
- **Adoption before spawn.** A daemon already answering the endpoint is adopted, not
replaced, unless it is unhealthy, foreign, or built from a superseded bundle _and_ owns no
live sessions. Replacing a healthy daemon kills its PTYs, so code freshness always defers
to live work.
- **Restart.** The adapter respawns the daemon on death, transparently to callers.
- **Crash-loop containment.** At most **5 launches per 60s rolling window** per orcad run;
past that, launches are refused with `daemon_crash_loop` and terminals fail with that
message instead of the process forking forever. The window slides, so a repaired host
recovers without restarting orcad. An operator-initiated daemon restart clears it — that
is the deliberate "try again".
- **No macOS login-session watch.** That watch retires the daemon when the spawning GUI login
session dies. An orcad daemon must survive its SSH session ending.
- **Shutdown.** orcad never stops the daemon. A daemon that was never adopted retires itself
after its adoption window; an adopted one stays resident (see Decommissioning).
### Decommissioning
After a PID-scoped stop, an adopted daemon stays resident so the next orcad can reattach. A
combined-unit systemd stop also leaves a scope-isolated daemon resident, but kills one that
fell back to the service cgroup. To retire a process-scoped deployment, apply the census rule
above, stop orcad, then stop the daemon named by `health.terminalDaemon.pid`.
Only report it `exited` after verification on the execution host; loss of contact is
`unverifiable`.
## Health
The readiness payload carries a `health` object:
```
buildHash sha256 (16 hex) of the running orcad bundle — build identity that a version
string cannot give, so a rollback that did not replace the file is visible
buildVersion ORCA_VERSION
nodeVersion / nodeAbi process.versions.node / .modules — the ABI native addons must match
platform / arch / pid
terminalDaemon:
state live | degraded | absent
ownsFreshSessions whether NEW terminals are daemon-owned; this supports PID-scoped
restart recovery, not supervisor or service-cgroup isolation
pid the live daemon's pid, from its own PID record
buildVersion the build the LIVE daemon was forked from (may legitimately predate
this orcad after an update — reporting orcad's version for both would
hide exactly that)
entryPath / protocolVersion
selfTest { ok, coverage, verdict, durationMs }
```
### What the self-test proves
`selfTest` runs `checkDaemonHealth` against the daemon's socket. It is green only when the
daemon **opened its socket, completed the protocol handshake, and ran `ptySpawnHealth` — a
real short-lived PTY spawned inside the daemon's own process**. It therefore spans both
processes: orcad drives it, the daemon performs it, the verdict crosses the socket.
- `coverage: 'pty-spawn'` — the full round trip above.
- `coverage: 'handshake'` — **win32 only**, where `checkPtySpawnHealth` returns without
spawning anything. A green verdict there covers the handshake and nothing more. It is
reported separately rather than folded into `ok` so nobody reads it as a PTY round trip.
`state` is `live` only when the self-test passed **and** `ownsFreshSessions` is true. A
daemon that answers but has fallen back to local spawning for new terminals is `degraded`,
because those terminals die with orcad. A daemon that answered and then failed its spawn
probe is also `degraded`, not `absent`: it still holds live sessions, and calling those
exited would be the verdict `ssh-execution-boundary.md` forbids guessing.
## What is not covered
Named here so nothing reads as implemented that is not:
- **A continuous health endpoint.** `health` is published once, in the readiness payload. A
supervisor's periodic liveness/readiness probe needs an HTTP or RPC surface over the same
`collectOrcadHealth()`; that surface does not exist yet.
- **Supervision of an unscoped fallback daemon.** When the durable `systemd-run --user --scope`
launch is unavailable (see [above](#two-long-lived-processes-not-one)) orcad and its daemon
share one service cgroup, and a combined-unit stop cannot preserve live terminals. There is
no mechanism that re-isolates such a daemon after the fact.
- **libc slot.** There is no honest health value to publish until native libc detection owns
it.
- **`degradations[]`.** The readiness contract does not publish this collection yet.
- **Credential administration** (list / revoke / rotate devices, expiring pending offers,
structured security logging).
- **Pinned-port fail-closed.** A pinned `--port` still falls back to an OS-assigned port on
conflict.
- **Reconciling `webClientUrl` with reachability** under the loopback default.
- **State-schema rollback rules.**
- **Daemon log rotation.** `<data-root>/logs/daemon.log` grows unbounded.