* fix(linux): give the CLI one entrypoint by extracting the AppImage once
* refactor(linux): trim AppImage CLI registration seams
* test(cli): assert registration lock serialization
* fix(linux): fence AppImage terminal shim mounts
* fix(linux): accept extracted AppImage runtimes with APPDIR only
* docs(linux): make headless AppImage extraction runnable
* refactor(linux): import bundled launcher directly
* fix(linux): reclaim superseded AppImage payloads and packaged symlinks
Pruning removed 3215 of 3216 files from a superseded generation and always
stranded resources/app.asar, leaking ~105 MB per version update. Electron's
asar shim reports a *.asar file as a directory, so the recursive remove tried
to rmdir a real file and failed with ENOTEMPTY; the .catch(() => {}) hid it.
Reproduced end to end on Ubuntu 24.04: 519M -> 623M across one update, and
519M again once the payload is actually reclaimed.
removeExtractedAppImagePayload holds process.noAsar for the removal, counted
so overlapping removals cannot hand the shim back early, and the prune site
now warns with the path instead of swallowing the rejection. All three
removal sites use it -- staging cleanup and displaced roots leaked the same
way.
Also reclaim symlinks left by a packaged deb/rpm install, which the
extracted-cache-only rule turned into a hard conflict on a deb -> AppImage
migration, and name the remedy in the conflict error.
* fix(linux): bound the CLI registration lock wait
`retries: 1000` caps the attempt count, not elapsed time, so at up to 1s per
attempt an IPC-driven registration could hang ~16 minutes against a wedged
holder with no feedback.
A legitimate holder is bounded by the extraction timeout, so wait that plus
slack and then fail with a message naming the lock file, rather than hanging.
`maxRetryTime` is forwarded verbatim to the `retry` package by proper-lockfile.
* fix(linux): stop re-extracting the AppImage on inode metadata churn
The extracted-payload cache key hashed ctime alongside dev/ino/size/mtime.
ctime moves on any inode metadata write -- `chmod +x`, which every AppImage
user is told to run, plus `chown`, an ACL or SELinux relabel, and a backup
restore -- none of which alter a byte of the payload.
Measured on Ubuntu 24.04: `chmod +x` leaves dev, ino, size and mtime
identical and moves ctime alone, so the key changed and the next launch paid
a full ~519 MB re-extraction and a multi-second stall to rebuild a payload it
already had, then pruned the old generation.
Key on content identity instead. An in-place content change moves mtime and
almost always size; a replacement moves the inode. The existing
replace-in-place test still passes.
* fix(linux): stop CLI commands from falling through to Chromium startup
* refactor(cli): remove redundant command membership check
* test(cli): cover command-named project selectors
* fix(cli): redirect the open-url command before startup
* test(linux): cover AUR serve wrapper flags
* fix(linux): tighten CLI launch detection
* fix(linux): respect CLI flag value boundaries
* fix(linux): strip injected Chromium switches from CLI args
* fix(linux): report a missing display instead of dying in uv_close
* refactor(linux): read display locks without a preflight race
* fix(linux): preserve unverified external displays
* chore: format reliability gate manifest
* test(packaging): split runtime resource checks
* fix(linux): fail serve when no display is available
* fix(linux): do not treat a lockless X socket as a dead display
An X server writes its lock beside its socket and both survive a crash
(verified against Xvfb under SIGKILL), so a socket with no lock was never
left by a crashed server. It is an endpoint published from elsewhere: a
container bind-mounting only /tmp/.X11-unix, WSLg, or a foreign PID
namespace. Declaring those dead made the desktop gate exit(1) on displays
that work, with no workaround, and the serve gate refuse to start.
Liveness now splits by ownership. A foreign DISPLAY trusts a lockless
socket; Orca's own :99 does not, because removeStaleDisplayArtifacts
unlinks the lock before the socket and so manufactures that state itself --
adopting it would resurrect the orphan-socket bug and stop the cleanup from
self-healing. The stale-lock rejection is unchanged.
Also correct four doc statements this behaviour falsified.
* fix(linux): fail closed when a stale socket blocks the Xvfb rebind
Readiness only checked that /tmp/.X11-unix/X99 exists. A stale socket we
could not unlink still exists after our own Xvfb refused to bind, so Orca set
DISPLAY to a dead server and Chromium died in Ozone init.
Measured on Ubuntu 24.04 against the pre-fix build: with a leftover :99
socket and no lock, serve exits 139 (SIGSEGV), the socket inode is unchanged
before and after, and no lock is recreated -- it neither cleaned up nor
respawned. To a user that is a crash, not a misconfiguration.
This is reachable in the documented topology, where orca-xvfb.service has no
User= and runs as root while serve runs as User=orca: /tmp is sticky, so the
orca uid cannot unlink a root-owned socket, rmSync fails, and Xvfb exits with
the display already active.
Readiness now requires the display to actually be live -- our socket plus a
lock naming a running process -- so the same state reports an unusable
display and exits 1 with the existing diagnosis.
* fix(linux): recognise abstract X sockets and inherited Wayland fds
Two display setups this gate could not prove were refused outright, and on the
desktop path that is app.exit(1) with no workaround.
An X server may bind only the abstract namespace (`@/tmp/.X11-unix/X0`), which
leaves no filesystem socket to stat. Abstract addresses are kernel-owned and
vanish the moment the owner exits, so an entry in /proc/net/unix is proof of a
live server -- no lock file needed and no stale entry possible. Verified on
Ubuntu 24.04, where 139 such addresses were present.
WAYLAND_SOCKET is an already-connected fd handed over by the compositor, so
there is no path to stat and WAYLAND_DISPLAY may be unset entirely. Its
presence is the display.
Both are consulted only after the filesystem-socket check fails, so no
existing verdict changes.
* fix(linux): never treat Orca's own display number as a foreign endpoint
Recognising a lockless X socket as live is correct for an endpoint published
from elsewhere -- a container bind mount, WSLg -- because an X server writes
its lock beside its socket and both survive a crash. It is wrong for
VIRTUAL_DISPLAY_NUMBER, because Orca's own teardown unlinks the lock before
the socket and so manufactures that exact state.
The managed branch was already strict, but a caller that sets DISPLAY=:99
explicitly takes the foreign path and skipped it, accepting a dead display
left by Orca's own interrupted cleanup. Route the managed number through the
strict probe on both paths.
Found by an adversarial audit of the asymmetry introduced earlier in this
branch; the documented systemd topology is unaffected because its Xvfb writes
a real lock.
* test(linux): add a packaged-artifact contract for the CLI launch paths
* test(linux): avoid buffered serve readiness detection
* test(linux): signal AppImage serve owner directly
* test(linux): tolerate readiness timeout boundary
* test(linux): add startup margin to shutdown oracle
* ci(linux): give package contracts timeout headroom
* fix(ci): route all Linux packaging contract changes
* test(linux): poll shutdown readiness without tail leaks
* test(linux): bound shutdown cleanup grace
* test(linux): assert on CLI output, not the harness's own control lines
run-cli-case.sh echoes `RESULT status=N case=<name>`, and the two cases named
*-skills asserted `expectOutput: 'skills'`. That substring was satisfied by
the case name in the harness's own line, so 2 of 8 cases asserted nothing
about the command -- gutting `skills` entirely would still have gone green.
Control lines are now excluded before matching, and both cases assert the
rendered help header, which only real help output produces. Verified on an
Ubuntu 24.04 host: 8/8 still pass against a stack-tip AppImage.
Also register the gate in reliability-gates.jsonc, which #15085 added a CI
Docker gate without. Red/green is recorded from a stock release AppImage
failing 4 of 8, three of them at status 133 (SIGTRAP).
* fix(linux): require static AppImage runtimes (#17319)
* test(linux): reject a wrong-architecture native binary at packaging time
Cross-building the arm64 slice on an x64 host silently packed an x86-64
`pty.node` -- the rebuild logged "Forcing native rebuild for linux-arm64" and
shipped the host's binary anyway. Every gate here inspects symbol versions,
which are perfectly valid on the wrong architecture, so nothing noticed.
Observed on a Raspberry Pi 5: the packaged app loaded, then failed with
"Failed to load native module: pty.node", and the launch contract reported
3 of 8 cases crashed rather than naming the cause. Swapping in the aarch64
`pty.node` took the same build to 8/8.
Compare ELF `e_machine` against the slice being packaged and fail with the
offending path. Checked before the glibc pass, because a wrong-architecture
binary's symbol versions are valid but meaningless and would send the reader
down the wrong path.
Release CI builds arm64 on a native runner, so this guards local and future
cross-builds rather than a shipped artifact.
* test(linux): judge per-arch vendored binaries against their own path
The first CI run of the architecture gate failed the x64 package job on
`@parcel/watcher-linux-arm64-glibc/watcher.node`. That binary is arm64 on
purpose: the package ships every architecture and its loader picks the match,
so its presence in an x64 build is correct.
Judge a binary against the architecture its own path names, falling back to
the slice when the path names none. That keeps the case this gate exists for
-- `bin/linux-arm64-*/node-pty.node` holding an x86-64 binary, which is what
shipped to a Raspberry Pi 5 -- while letting multi-arch dependencies through.
Dry-run over the real dependency tree flags nothing for either target arch.
* fix(linux): move deb/rpm update installation outside Orca (#17318)
* fix(linux): complete deb/rpm package metadata
* fix(linux): preserve CLI link during package upgrades
* docs(linux): document local RPM build prerequisites
* fix(linux): move deb/rpm update installation outside Orca
* fix(updater): preserve Linux recovery across stale events
* fix(updater): fence stale downloaded events by active target
* fix(updater): preserve active Linux package recovery
* test(linux): keep workflow order assertion in scope
* test(updater): assert stale recovery stays silent
* fix(updater): preserve Linux package recovery after checks
* refactor(updater): keep Linux marker message with status
* fix(linux): describe the right manual update path for deb/rpm hosts
A remote host installed from .deb or .rpm now reports
manual-service-update-required, and the guidance told the operator to
"update through the service manager that starts this server" -- which is
correct for unsupported-headless-serve but wrong for a package install,
where nothing about the remedy involves the service manager.
Say both, keyed on how the host was installed.
* docs(linux): document orcad update restart safety
* docs(linux): scope restart census omissions
* docs(linux): use absolute service CLI launcher
* fix(serve): validate in-process serve options before startup (#17683)
* fix(linux): stop offering updates a distro-managed install cannot apply (#17918)
Closes #17702.
The resources/package-type marker is authoritative but never checked against
the host, so any repackager that unpacks Orca's .deb -- AUR, Nix, a container
rebuild -- inherits `deb` verbatim. Install feasibility was then computed
after a ~165 MB download, so those users got check -> download -> a card
promising an install command -> a dead end.
Validate the marker against the host: a deb/rpm marker with no matching
package manager in the trusted directories means a package manager owns this
install. This reuses the exact lists and resolver that
buildLinuxPackageInstallCommand already loops over, so a false positive is
impossible by construction -- any host flagged here would have failed with
no-package-manager after the download anyway. The gate only moves that
verdict earlier. Verified across Debian 12, Ubuntu 24.04, Arch, Fedora 40 and
openSUSE Leap: no false positive on a real deb host, correct on every
repackaging host.
The release is still reported, because the user does want to know 1.4.194
exists and to update through their distro; only the download path is closed.
`externallyManaged` is an additive optional field on the existing `available`
status, so older paired clients decode it unchanged. downloadUpdate() refuses
authoritatively, since main owns this verdict rather than the card, and
unwinds any pinned-build state first -- a Linux pinned jump resolves to
'release', and stranding isPinnedBuildActive would silently kill every
background check for the rest of the process.
Note the fix the issue suggests cannot work: electron-updater builds a
PacmanUpdater whose doDownloadUpdate looks for a .pacman asset Orca does not
publish, then dereferences undefined.
* style(cli): restore prettier wrapping on install error copy
* test(linux): re-pin the child-process ratchets and the batch-shim allowlist after the merge
11 KiB
SSH Execution Boundary
How Orca splits work between your machine and an SSH host, what survives a disconnect, and how to keep unverifiable distinct from exited. Nothing under docs/ stated this before; agents and humans were inferring it from error strings and getting it wrong.
The rule
The execution host owns everything that touches execution — tools, credentials, identity, environment, processes, and artifacts. The client owns the UI, transport, and Orca control-plane state, but is not authoritative for execution state.
Two consequences, both non-negotiable:
- No silent substitution. An operation on a remote
repoPathmust never fall back to running on the client. A missing SSH provider is not permission to answer locally — a local run can answer for the wrong repository. - No asserting what you cannot observe. Loss of contact is not evidence of
exited. Reportunverifiable, neverexited.
The vocabulary is fixed: live / unverifiable / exited, taken from the incumbent UnstoppedPtyVerdict. Do not introduce synonyms, and never collapse unverifiable into either neighbour. exited requires positive evidence of absence from the host that owns the process; a transport failure can only ever produce unverifiable.
Rule 1 is stated at src/main/source-control/repo-default-branch.ts:76-78, src/main/repo-worktrees.ts:45-48, OrcaRuntimeService.probeWorktreeDrift in src/main/runtime/orca-runtime.ts, and src/renderer/src/lib/connection-context.ts:22-24. It is enforced throughout src/main/runtime/orca-runtime-git.ts by the guard that throws SSH_GIT_PROVIDER_UNAVAILABLE_MESSAGE whenever target.connectionId is set and no provider is registered — grep that constant for the current call sites rather than trusting a count.
src/main/runtime/unstopped-pty-verification.ts:12-16 is the reference implementation of rule 2: it keeps live / unverifiable / exited as three distinct verdicts, and treats "we could not ask" as its own answer.
What runs where
| Concern | Executes on | Notes |
|---|---|---|
| PTYs, agent CLIs | remote | children of the detached relay daemon, not of the ssh channel |
| git (status, diff, log, fetch, push, commit, branch, worktree) | remote | via src/relay/git-handler.ts |
| filesystem, watching, search | remote | |
repo setup hooks (--setup) |
remote | identical policy to local |
| commit-message / PR-field AI generation | remote | uses the remote agent CLI and its auth |
gh / GitHub API, glab / GitLab |
client | inconsistent with the rule; PRs carry the client's identity |
the orca CLI inside a remote terminal |
client runtime | control plane only — your files and processes stay remote; see below |
Survival: what a disconnect does not do
By default, remote work survives your machine going away. The relay is a detached daemon (nohup … </dev/null &), its handler in src/relay/relay.ts ignores SIGHUP, the PTY is its child rather than the ssh channel's, and quitting Orca is a detach, not a dispose (src/main/ssh/ssh-relay-session.ts:901-915). Sleep additionally pushes graceTimeSeconds: 0 to un-bound any running grace window.
Two ways remote work can actually stop:
- A bounded grace period. The shipped default is
0= keep alive until reset. If "keep terminals alive until reset" is unchecked, the configurable range is 60s–7d and the form defaults to 24h. The countdown starts when the client disconnects, after which the relay SIGKILLs every PTY. Note the asymmetry: sleep protects you, but ordinary disconnect and app quit do not. No command reports which setting is in effect for a target, so at N hours since disconnect you cannot tell "unlimited" from "24h with 7 left" — treat the remote asunverifiable, notexited. - Host-acknowledged explicit user action — End Remote Terminals, Reset Relay, removing the target, or closing the tab. When the host cannot acknowledge the request, closing a tab or removing a target may clear only client state; the remote verdict remains
unverifiable.
Reconnect re-attaches to the same live PTYs and replays a bounded buffer (REPLAY_BUFFER_MAX, a 102,400-code-unit tail). Output beyond that while you were away is lost to the client even though the process was never interrupted: the transcript is truncated; the work stays live.
Updating Orca strands relay-backed terminals
There is a third outcome that is neither of the two above, and the vocabulary matters: the work does not stop, it becomes permanently unreachable.
The relay's install directory — and therefore its socket path — is namespaced by a content hash of the relay bundle (computeRemoteRelayDir in src/main/ssh/ssh-relay-versioned-install.ts, consumed by resolveRemoteInstallState in src/main/ssh/ssh-relay-deploy.ts), and the daemon refuses any client whose bundle hash differs (handleDaemonHandshakeFrame in src/relay/relay-handshake.ts, exit EXIT_CODE_VERSION_MISMATCH 42). Two builds whose relay protocol is byte-identical still refuse each other. So the first reconnect after an app update deploys a new relay at a path the incumbent was never listening on, and cannot reach it even in principle. Every PTY the incumbent owns is unverifiable — running, unreachable, and never exited. The client's leases are attempted against a relay that never minted their ids, expired on the not-found answer (handlePtyReattachFailure in src/main/ssh/ssh-relay-session.ts), and the pane falls back to a cold-restore agent resume — or to a bare shell when no resumable provider session was captured for it. The old relay keeps its directory pinned against GC, because its socket really is live (hasLiveRelaySocket in src/main/ssh/remote-install-gc.ts). See #13852.
The peer model does not have this failure, and that is the concrete reason behind "One host, one model" below. The daemon's endpoint is namespaced by a semantic protocol version rather than a build (daemon-v<N>.sock, from getDaemonSocketPath in src/main/daemon/daemon-spawner.ts), every earlier protocol version stays attachable (PROTOCOL_VERSION in src/main/daemon/daemon-protocol-version.ts), and a daemon holding live sessions is preserved across a version change instead of replaced (shouldPreserveDaemonWithLiveSessions in src/main/daemon/daemon-replacement-preflight.ts).
Control plane
On an SSH host, orca is a shim (~/.orca-relay/bin/orca) that proxies back to the client's runtime over the relay socket. Your repository, processes, and files remain remote — only the control plane is on the client. This is correct for an SSH target, but it has a consequence worth stating plainly:
When the client disconnects, every
orca …command run on the SSH host fails withNo owning Orca client is connected to the relay. The PTY stayslive; its control plane does not.
Orchestration state (Runs, Tasks, Dispatches, mailboxes) is client-resident for the same reason. An agent on an SSH host should not depend on orca for anything it must finish while you are away. Commit and push early — unpushed work on a remote box is unavailable to the client until it reconnects.
Distinguishing unverifiable from exited
A verdict needs evidence from the host that owns the process. Apply these tests in order.
Was the signal produced by the owning host, or by the client's own bookkeeping? Absence from a client-side set, a lookup that threw, a socket that closed, a command that timed out — none of these observe the process. They are unverifiable by construction, whatever the field is named.
Did every remote PTY on that target go quiet at once? A transport drop takes them all together. Simultaneous silence across a host indicates a lost link, not simultaneous death.
Does the termination event match the current identity? A host-delivered exit for the live PTY incarnation and provider generation, while its siblings still report, establishes exited. A stale event, an event for a superseded incarnation, or one quiet terminal with no host evidence does not.
Is a returned status actually a claim of success? An operation that reports failure may have succeeded, and one that reports success may not have run — check the durable state it should have changed rather than trusting the return.
Anything short of positive host evidence is unverifiable. Reporting it as exited is the error this document exists to prevent: it orphans live work and can cold-start a duplicate over the same worktree.
Reading artifacts instead of process state
Artifacts are stronger evidence than liveness signals, but they answer a narrower question than they appear to.
A matching commit from git ls-remote --heads origin <branch> or a PR head lookup proves that commit reached the remote — not that the current run pushed it, and not that the latest work was included. An absent result proves nothing was found, not that nothing was pushed: the ref may have been deleted, the PR closed, or the query may simply have failed.
A listing is only evidence about the hosts it actually covered. When a result does not name its scope, an empty answer is not evidence that nothing is running elsewhere. A clean local worktree says nothing at all about the remote one.
One host, one model
An SSH host and a paired runtime (orca environment) imply opposite boundaries: the first is a dumb execution host driven by your client, the second is a peer that owns its own control plane. Registering the same machine both ways splits its worktrees across two identities, makes terminal list return different sets depending on --environment, and reliably confuses both humans and agents. Pick one per machine.
For work that must continue while you are offline, use the peer/headless-runtime model on the remote host instead of the direct-SSH model. Its control plane is host-local, and its daemon-backed PTYs can stay live across a PID-scoped runtime restart so the runtime can reattach. A service manager that reaps the runtime's cgroup, or an explicit daemon shutdown, makes them exited; see Running orcad. Do not register the same machine through both models. A detached agent process outside Orca can also survive a control-plane outage, but it has no stdin, so its instructions cannot be amended mid-run.