mirror of
https://github.com/stablyai/orca.git
synced 2026-10-04 08:02:09 +00:00
* Revert "feat(orcad): source-side dormant export of a relay-hosted SSH target (#16741 T6-8) (#24519)" This reverts commit783101b304. * Revert "feat(ssh): update, roll back, recover and stop a managed orcad server (#16741 T6-5 follow-up) (#24463)" This reverts commit38c2d1dcb9. * Revert "feat(ssh): deploy and pair an empty managed orcad server over SSH (#16741 T6-5) (#24453)" This reverts commit8b76683b40. * Revert "fix(ssh): orcad GC honors the activation journal; readiness requires proven daemon coverage (#16741 T6 follow-up) (#24451)" This reverts commitd3f8c5063b. * Revert "feat(ssh): remote orcad stop by request file and journaled decommission (#16741 T6-4) (#24449)" This reverts commit43d9b43d3f. * Revert "feat(orcad): supervisable server: stop requests, managed stop receipts and a lifetime that keeps its lock on failed teardown (#16741 T6-3) (#24433)" This reverts commitb093d3ab20. * Revert "feat(ssh): crash-safe orcad activation, rollback and recovery (#16741 T6-2) (#24423)" This reverts commit1a9ac0e955. * Revert "feat(runtime): SSH access links for paired servers in a downgrade-safe sidecar (#16741 T5-1+T5-2) (#24420)" This reverts commit99db2bfae4. * Revert "feat(relay): capability-gated owner reset with a durable preparation journal (#16741 T3 R1) (#24418)" This reverts commit34a582bd39. * Revert "feat(ssh): track connection-manager drains, test probes and provider continuations (#16741 T2 P3+P8a) (#24407)" This reverts commitd53063d2b1. * Revert "feat(daemon): idle retirement, session census and recovery-only provider (#16741 T2 P4b) (#24409)" This reverts commitff212dbbef. * Revert "feat(ssh): add pty.resumeClient and split SSH PTY process listing (#16741 T2 P5+P6) (#24414)" This reverts commit92cb71765e. * Revert "feat(relay): await owned watcher and agent children on shutdown (#16741 T2 P1) (#24400)" This reverts commit6b36e4f85b. * Revert "feat(session): retry failed renderer session writes and verify local folder PTYs (#16741 T2 P9) (#24406)" This reverts commitd23ecef301. * Revert "feat(ssh): remote orcad primitives on the pinned Node runtime (#16741 T6-1) (#24419)" This reverts commitdd87ae578d. * Revert "fix(runtime): fence runtime-environment subscriptions and status probes by identity (#16741 T5-3) (#24421)" This reverts commitece9e4d2e3. * Revert "feat(orcad): migration manifest and dormant-state contracts (#16741 T6-7) (#24422)" This reverts commit3fbdaba262. * Revert "feat(ssh): wire SshConnection through the work and transport close ledgers (#16741 T2 P2) (#24401)" This reverts commit4e8edc8872. * Revert "feat(profiles): carry markdown frontmatter visibility in project transfers (#16741 T2 P7) (#24405)" This reverts commit60c93263cc. * Revert "fix(runtime): project the PTY incarnation onto mobile session tabs (#24413)" This reverts commit99e0303572. * Revert "feat(daemon): tag daemon stream data with the PTY incarnation id (#16741 T2 P4a) (#24402)" This reverts commit817af768b0. * Revert "feat(ssh): port the SSH connection work ledger and transport close ledger (#16741 T2) (#24210)" This reverts commitc9918931c8. * Revert "feat(relay): fence and drain file and git response streams on shutdown (#24185)" This reverts commitdc08ffeba9. * Revert "refactor(runtime-rpc): extract the Node WebSocket lifecycle; opt-in pinned port (#24186)" This reverts commita789233bbb. * Revert "feat(relay): route relay handlers through work admission; producer publication drain (#24181)" This reverts commit0b812bd698. * Revert "feat(relay): land the #16741 T1 seam (work drain, publication drain, release gate) (#24156)" This reverts commit3aa2d3af7c. --------- Co-authored-by: m4air <m4air@Mac.localdomain>
266 lines
16 KiB
Markdown
266 lines
16 KiB
Markdown
# Running orcad
|
||
|
||
`orcad` is the Orca runtime served from plain Node. This is the contract between it and
|
||
whatever supervises it: what it binds, what it owns on disk, who restarts what, and what its
|
||
readiness payload actually proves.
|
||
|
||
## Two long-lived processes, not one
|
||
|
||
A deployment is **orcad** plus **the terminal daemon**.
|
||
|
||
| | orcad | terminal daemon |
|
||
| ---------- | -------------------------------- | ------------------------------------- |
|
||
| Started by | the supervisor | orcad, detached |
|
||
| Owns | RPC, git, worktrees, persistence | every local PTY |
|
||
| Lifetime | one supervised run | detached from orcad, not its service |
|
||
| Endpoint | `ws://<bind>:<port>` | `<data-root>/daemon/daemon-v<N>.sock` |
|
||
|
||
orcad detaches the daemon and calls `disconnectDaemon()`, never `shutdownDaemon()`. The
|
||
built-in remote deployment path stops only the recorded orcad PID, so the daemon and its PTYs
|
||
survive. The successor adopts the current endpoint and routes supported previous protocol
|
||
versions through legacy adapters. This makes a PID-scoped update, rollback or restart
|
||
non-destructive to live work.
|
||
|
||
Process detachment is not service isolation. A daemon that orcad launches directly, and every
|
||
PTY it owns, remain in the same systemd service cgroup. `KillMode=mixed` does **not** preserve
|
||
them: it sends the graceful stop signal only to the main process, then sends `SIGKILL` to every
|
||
process remaining in the cgroup the moment that main process exits — `TimeoutStopSec` never gets
|
||
the chance to apply. `KillMode=control-group` is destructive too. `KillMode=process` leaves
|
||
service-owned processes unmanaged and is not a supported preservation mechanism.
|
||
|
||
Service-restart survival therefore requires a separately supervised cgroup, and orcad now asks
|
||
for one: on Linux it launches the daemon through `systemd-run --user --scope`, which places the
|
||
daemon and its PTYs in their own transient `orca-daemon-<launch-nonce>.scope` unit under the
|
||
user slice instead of the caller's service cgroup. A stop or restart of the service unit then
|
||
leaves that scope — and the live terminals in it — running, and the successor adopts the
|
||
endpoint as it always has.
|
||
|
||
A newly launched private daemon scope also follows the daemon's own lifetime. A small
|
||
detached shell holds an input pipe from the daemon; after that pipe closes and `/proc`
|
||
confirms the daemon PID is gone, it asks the user manager to stop that exact scope.
|
||
Systemd sends remaining processes SIGTERM and escalates after five seconds. This includes
|
||
children that double-forked or called `setsid` and can no longer be found by parent PID.
|
||
Disconnecting or restarting the runtime does not close the pipe: the daemon owns it.
|
||
|
||
The cleanup only arms on a fresh scoped launch with a matching launch nonce. Adopted
|
||
legacy scopes can contain GUI processes and are never armed retroactively. Unscoped
|
||
launches and children deliberately moved into another systemd unit remain outside this
|
||
cleanup. `nohup`, `disown`, and `tmux` alone do not move a process out of its cgroup, so
|
||
those children now end when their terminal daemon dies. Work intended to outlive that
|
||
daemon needs its own service or scope.
|
||
|
||
The scope is requested only where it can work. All of these must hold:
|
||
|
||
- **Linux with systemd as PID 1** (`/run/systemd/system` exists).
|
||
- **A reachable user bus** — a connectable `bus` socket in the per-UID runtime dir
|
||
(`/run/user/<uid>`, or whatever `XDG_RUNTIME_DIR` points at). For a service account that is
|
||
not otherwise logged in, that means `loginctl enable-linger <user>`; a unit whose
|
||
`RuntimeDirectory=` hardening moves `XDG_RUNTIME_DIR` off the per-UID path is handled, because
|
||
the real per-UID path is probed first.
|
||
- **`systemd-run` on `PATH`** and answering `--version`.
|
||
|
||
Any of those missing, or a `StartTransientUnit` call that fails anyway, falls back to the
|
||
direct launch — and in that unscoped fallback case the paragraph above still describes reality:
|
||
the daemon shares the service cgroup and a combined-unit stop ends live terminals. Read the
|
||
`cgroupUnit` field in the daemon health payload to tell the two cases apart on a running host;
|
||
it is populated from `/proc/self/cgroup`, so it reports the isolation the daemon actually has
|
||
rather than what the launcher intended.
|
||
|
||
## Bind policy
|
||
|
||
`--bind <literal-ip>`, **default `127.0.0.1`**.
|
||
|
||
Only literal IPs are accepted; hostnames are refused because DNS would decide which
|
||
interface got bound. `localhost` maps to `127.0.0.1`. `0.0.0.0` / `::` are the explicit
|
||
opt-ins to network reach, and the startup log says so on every launch.
|
||
|
||
The bind is **pinned**, not defaulted. Two things widen the desktop's listener on their own —
|
||
`orca serve`'s wide default, and a startup where some device has connected before — and an
|
||
unattended host's exposure must be exactly what the operator asked for on every launch. A
|
||
mobile pairing offer, which normally rebinds to all interfaces, is refused while the bind is
|
||
pinned to loopback and reports `network_exposure_failed` rather than advertising an endpoint
|
||
nothing can reach.
|
||
|
||
Under the shipping design a client reaches a remote orcad over an SSH local port-forward, so
|
||
loopback is the correct default and the pairing credential travels over SSH.
|
||
|
||
## Data root and the instance lock
|
||
|
||
The data root is `$ORCA_USER_DATA`, else `$XDG_DATA_HOME/Orca`, else `~/.orca`.
|
||
|
||
Before the profile index or the store is touched, orcad takes `<data-root>/orcad.lock`.
|
||
It refuses to start when:
|
||
|
||
| Code | Meaning |
|
||
| -------------------------------------- | ------------------------------------------------------------- |
|
||
| `orcad_data_root_wrong_owner` | the root is owned by another uid (POSIX) |
|
||
| `orcad_data_root_shared` | the root is group/world accessible and could not be tightened |
|
||
| `orcad_instance_lock_held` | another live orcad owns this root |
|
||
| `orcad_instance_lock_foreign_identity` | the lock belongs to a different identity |
|
||
| `orcad_data_root_unusable` | the root cannot be created, stat'd or written |
|
||
|
||
A root that is merely too permissive and that we own is tightened to `0700` rather than
|
||
refused — orcad stores credentials there unsealed (no OS keyring on this host), so the goal
|
||
is a private root, and refusing when we could just fix it helps nobody. We refuse when the
|
||
permissions are not ours to fix. Windows is exempt from the owner and mode checks: ACLs are
|
||
not expressible as a POSIX mode, and `statSync().mode` there reports a synthesized one.
|
||
|
||
A dead holder's record is reclaimed (PID plus process start time, so a recycled PID does not
|
||
read as alive). A record belonging to a different identity is never reclaimed.
|
||
|
||
**The lock scopes one role — who is the runtime.** It deliberately says nothing about the
|
||
daemon, which lives under `<data-root>/daemon` and fences its own endpoint with its own PID
|
||
record. A lock that asked "is any process using this root" would refuse exactly the restarts
|
||
a live daemon makes worthwhile.
|
||
|
||
## Supervision
|
||
|
||
### Process-scoped and cgroup-wide stops
|
||
|
||
The built-in remote updater performs a PID-scoped stop and keeps the daemon's install version
|
||
pinned while it owns sessions. A combined-unit systemd stop or restart is different: unless the
|
||
daemon holds a durable cgroup scope of its own (see
|
||
[Two long-lived processes, not one](#two-long-lived-processes-not-one)), it reaps the daemon and
|
||
every live terminal after the graceful window. Treat a stop as destructive unless
|
||
`health.terminalDaemon.cgroupUnit` names an `orca-daemon-*.scope` on that host.
|
||
|
||
Before a cgroup-wide stop, obtain a fresh `orca-ide terminal list --json` result using the same OS
|
||
account and home as the daemon. Invoke the installer's absolute launcher path so `sudo`'s
|
||
`secure_path` cannot hide a per-user registration (for example,
|
||
`sudo -Hu orca /home/orca/.local/bin/orca-ide terminal list --json`). Replace both `orca` and
|
||
`/home/orca` with the service account and home used by the unit; an extracted deployment may use
|
||
its absolute `resources/bin/orca-ide` launcher instead. A safe empty census is untruncated, has an explicit `hostScope`, covers every
|
||
execution host affected by the stop, and lists no terminals on those hosts. Every
|
||
`omittedHostIds` entry must be explicitly accounted for outside the target service's execution
|
||
boundary. A separately paired runtime is outside that boundary; local execution and SSH hosts
|
||
reached through this runtime are not. An affected or unknown omission, missing scope,
|
||
truncation, a failed request or lost contact makes the result `unverifiable`: defer the stop. Do
|
||
not admit new work after the census. Orca does not yet provide an atomic census-and-stop fence.
|
||
|
||
### Who supervises orcad
|
||
|
||
An external supervisor (systemd, launchd, a process manager). orcad conforms to it:
|
||
|
||
- **Readiness.** One JSON line on stdout (`--json`), `type: "orca_server_ready"`, published
|
||
after the listener is bound and the daemon verdict is in. There is no separate readiness
|
||
socket; the line is the signal. Set the supervisor's start timeout generously — the daemon
|
||
launch has its own retries and can take tens of seconds on a cold host.
|
||
- **Shutdown.** `SIGTERM` or `SIGINT` starts one graceful stop. Repeated signals share
|
||
that stop because a supervisor may signal both the launcher and its child. A 15s deadline
|
||
exits with code 1 if teardown stalls. The bundled runtime also stops gracefully if its
|
||
launcher's IPC channel closes. On POSIX, both the launcher and runtime ignore `SIGHUP`,
|
||
so terminal hangups do not stop a headless host. Use `SIGTERM` or `SIGINT` to stop it.
|
||
- **Exit codes.**
|
||
|
||
| Code | Meaning | Supervisor should |
|
||
| ---- | ------------------------------------------------------------ | -------------------- |
|
||
| 0 | clean shutdown | restart per policy |
|
||
| 1 | startup or shutdown failure | restart with backoff |
|
||
| 78 | configuration fault (bind address, data root, instance lock) | **not** restart |
|
||
|
||
78 is `EX_CONFIG`. Put it in systemd's `RestartPreventExitStatus`: restarting on a data
|
||
root owned by someone else is a restart-spin, not a recovery.
|
||
|
||
- **Logs.** orcad writes human-readable diagnostics to **stderr** and its readiness contract
|
||
to **stdout**; the supervisor owns capture and rotation. The daemon, being detached, writes
|
||
its own NDJSON lifecycle log to `<data-root>/logs/daemon.log` (suppressed by
|
||
`ORCA_DIAGNOSTICS_DISABLED=1`). Rotation of that file is not implemented — see
|
||
[What is not covered](#what-is-not-covered). orcad records every trace span it emits
|
||
(git commands, worktree paths, terminal spawns, structured-chat failures and the rest) to
|
||
`<data-root>/logs/orcad.trace.ndjson`, rotated at 10 MB × 10 files, private to its user and
|
||
redacted for secret-shaped strings. It stays on the host: a desktop's diagnostics bundle does
|
||
not collect it. `ORCA_DIAGNOSTICS_DISABLED=1` turns it off, and a logs folder orcad cannot
|
||
open leaves it off with one stderr warning rather than stopping orcad.
|
||
|
||
### orcad supervising the daemon
|
||
|
||
- **Launch.** On Linux, through `systemd-run --user --scope` so the daemon gets its own
|
||
transient cgroup and survives a service-unit restart; everywhere else, and wherever that
|
||
scope is unavailable, forked detached. Either way it runs `daemon-entry.js` beside
|
||
`orcad.js` with its own PID record, token and socket under `<data-root>/daemon`.
|
||
- **Adoption before spawn.** A daemon already answering the endpoint is adopted, not
|
||
replaced, unless it is unhealthy, foreign, or built from a superseded bundle _and_ owns no
|
||
live sessions. Replacing a healthy daemon kills its PTYs, so code freshness always defers
|
||
to live work.
|
||
- **Restart.** The adapter respawns the daemon on death, transparently to callers.
|
||
- **Crash-loop containment.** At most **5 launches per 60s rolling window** per orcad run;
|
||
past that, launches are refused with `daemon_crash_loop` and terminals fail with that
|
||
message instead of the process forking forever. The window slides, so a repaired host
|
||
recovers without restarting orcad. An operator-initiated daemon restart clears it — that
|
||
is the deliberate "try again".
|
||
- **No macOS login-session watch.** That watch retires the daemon when the spawning GUI login
|
||
session dies. An orcad daemon must survive its SSH session ending.
|
||
- **Shutdown.** orcad never stops the daemon. A daemon that was never adopted retires itself
|
||
after its adoption window; an adopted one stays resident (see Decommissioning).
|
||
|
||
### Decommissioning
|
||
|
||
After a PID-scoped stop, an adopted daemon stays resident so the next orcad can reattach. A
|
||
combined-unit systemd stop also leaves a scope-isolated daemon resident, but kills one that
|
||
fell back to the service cgroup. To retire a process-scoped deployment, apply the census rule
|
||
above, stop orcad, then stop the daemon named by `health.terminalDaemon.pid`.
|
||
Only report it `exited` after verification on the execution host; loss of contact is
|
||
`unverifiable`.
|
||
|
||
## Health
|
||
|
||
The readiness payload carries a `health` object:
|
||
|
||
```
|
||
buildHash sha256 (16 hex) of the running orcad bundle — build identity that a version
|
||
string cannot give, so a rollback that did not replace the file is visible
|
||
buildVersion ORCA_VERSION
|
||
nodeVersion / nodeAbi process.versions.node / .modules — the ABI native addons must match
|
||
platform / arch / pid
|
||
terminalDaemon:
|
||
state live | degraded | absent
|
||
ownsFreshSessions whether NEW terminals are daemon-owned; this supports PID-scoped
|
||
restart recovery, not supervisor or service-cgroup isolation
|
||
pid the live daemon's pid, from its own PID record
|
||
buildVersion the build the LIVE daemon was forked from (may legitimately predate
|
||
this orcad after an update — reporting orcad's version for both would
|
||
hide exactly that)
|
||
entryPath / protocolVersion
|
||
selfTest { ok, coverage, verdict, durationMs }
|
||
```
|
||
|
||
### What the self-test proves
|
||
|
||
`selfTest` runs `checkDaemonHealth` against the daemon's socket. It is green only when the
|
||
daemon **opened its socket, completed the protocol handshake, and ran `ptySpawnHealth` — a
|
||
real short-lived PTY spawned inside the daemon's own process**. It therefore spans both
|
||
processes: orcad drives it, the daemon performs it, the verdict crosses the socket.
|
||
|
||
- `coverage: 'pty-spawn'` — the full round trip above.
|
||
- `coverage: 'handshake'` — **win32 only**, where `checkPtySpawnHealth` returns without
|
||
spawning anything. A green verdict there covers the handshake and nothing more. It is
|
||
reported separately rather than folded into `ok` so nobody reads it as a PTY round trip.
|
||
|
||
`state` is `live` only when the self-test passed **and** `ownsFreshSessions` is true. A
|
||
daemon that answers but has fallen back to local spawning for new terminals is `degraded`,
|
||
because those terminals die with orcad. A daemon that answered and then failed its spawn
|
||
probe is also `degraded`, not `absent`: it still holds live sessions, and calling those
|
||
exited would be the verdict `ssh-execution-boundary.md` forbids guessing.
|
||
|
||
## What is not covered
|
||
|
||
Named here so nothing reads as implemented that is not:
|
||
|
||
- **A continuous health endpoint.** `health` is published once, in the readiness payload. A
|
||
supervisor's periodic liveness/readiness probe needs an HTTP or RPC surface over the same
|
||
`collectOrcadHealth()`; that surface does not exist yet.
|
||
- **Supervision of an unscoped fallback daemon.** When the durable `systemd-run --user --scope`
|
||
launch is unavailable (see [above](#two-long-lived-processes-not-one)) orcad and its daemon
|
||
share one service cgroup, and a combined-unit stop cannot preserve live terminals. There is
|
||
no mechanism that re-isolates such a daemon after the fact.
|
||
- **libc slot.** There is no honest health value to publish until native libc detection owns
|
||
it.
|
||
- **`degradations[]`.** The readiness contract does not publish this collection yet.
|
||
- **Credential administration** (list / revoke / rotate devices, expiring pending offers,
|
||
structured security logging).
|
||
- **Pinned-port fail-closed.** A pinned `--port` still falls back to an OS-assigned port on
|
||
conflict.
|
||
- **Reconciling `webClientUrl` with reachability** under the loopback default.
|
||
- **State-schema rollback rules.**
|
||
- **Daemon log rotation.** `<data-root>/logs/daemon.log` grows unbounded.
|