mirror of
https://github.com/stablyai/orca.git
synced 2026-09-23 00:02:29 +00:00
<!-- orca-pr-loc -->
<!-- Programmatic LoC summary. Do not edit by hand; rewritten on every commit. -->
| | Files | Added | Deleted | Net |
| :--- | ---: | ---: | ---: | ---: |
| Test | 19 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$962 | $\color{#cf222e}{\Huge{\mathbf{−}}}$136 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$826 |
| Prod | 18 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$295 | $\color{#cf222e}{\Huge{\mathbf{−}}}$116 | $\color{#1a7f37}{\Huge{\mathbf{+}}}$179 |
<!-- /orca-pr-loc -->
## Symptom
Live 2026-09-05 (Orca 1.4.198 client, Ubuntu host): both relay processes `kill -STOP`ped for 20 s, then `-CONT`. The client redeployed while the host was frozen. Its fresh daemon lost the socket bind (`Socket path already in use`) but had **already rewritten** `relay-<id>.sock.credential`. The surviving daemon kept its in-memory credential, so every later `--connect` got `Endpoint credential mismatch; closing socket`, then `Grace started … timeoutMs=0 … ptys=1, clients=0` every ~20 s, forever. Only a manual `kill -TERM` cleared it. Receipts: `review-archive/orchestration-v3-pr16904/smoke-receipts-t012b/E16,E17,E18,E24`.
Three independent defects kept the wedge alive; each is fixed at its own seam.
## Fix
**1. The relay daemon owns credential publication (race-free under two concurrent starters).**
`relay-daemon.ts` binds the socket first, then publishes via the new `src/relay/relay-endpoint-credential-publication.ts`: adopt a valid pre-existing file (older clients still pre-write), else mint 32 random bytes and write temp+rename at 0600. A start that loses the bind exits inside `listen()` and never reaches the file. Why this option and not restore-on-loss or a client-side write: the only process that can *prove* ownership is the one whose `listen()` succeeded, and that proof is atomic with the bind. The client-side pre-write (`ssh-relay-endpoint-credential.ts`) and the launch-command `chmod 600`/`icacls` are removed on POSIX and Windows. The racing test also exposed that macOS reports a mid-bind collision as `EEXIST` rather than `EADDRINUSE`; `relay-socket-ownership.ts` now treats both as "held or stale".
**2. The client distinguishes "no daemon" from "daemon present but not answering", and never rewrites.**
A credential refusal is now typed on the wire: the daemon replies `orca-relay-handshake-credential-mismatch` (same frame type, no new opcode) and the bridge exits **43**; `waitForSentinel` maps it to `RelayCredentialMismatchError`, which the takeover treats as handshake-refusal evidence exactly like exit 42. A relay that holds the endpoint but **never refused** (the stalled-host shape: kernel backlog accepts the probe, handshake gets no answer) is now `RelayEndpointUnresponsiveError`, routed to the relay-lost backoff instead of the terminal Reset Relay path. Silence is not a decision (`docs/reference/ssh-execution-boundary.md`).
**2b. Deploy honours the verdict.** The 40 s live run exposed that the `--connect` catch block in `deployAndLaunchRelay` predates the incumbent probe and swallowed both verdicts as "probe failed, launch fresh", so a fresh daemon was still launched over the live one (it lost the bind by luck, which is exactly the collision in the incident). Held and Unresponsive now propagate; the session backs off on Unresponsive and surfaces Reset Relay on Held. Red-first in `ssh-relay-deploy-incumbent-verdict.test.ts`.
**3. The daemon cannot be wedged by a rotated file, because nothing can rotate it.**
The credential lives in the content-hashed relay dir, and after (1) the only writer is the daemon that owns the socket, so the "file changed under a live daemon" state the incident depended on is no longer reachable in-product. The credential is therefore fixed for the daemon's lifetime, as a plain secret should be. A hand-edited file is refused with the typed reply until restored (tested). Startup adoption of a pre-written file applies an owner-only + same-uid rule (review finding): anything else is replaced by a fresh mint. An earlier revision of this PR also re-read the file on mismatch and adopted it; that was removed as unreachable machinery that turned the credential into a per-handshake file-ownership check.
**3b. Fail closed between bind and publication.** A client that arrives after `listen()` resolves but before the credential is set is refused, not admitted as `unproved`. Nothing can be delivered in that window today; the guard makes the boundary structural instead of an event-loop ordering fact. Red-first in `relay-reconnect-listener-credential-gate.test.ts`.
**Wire compat.** New optional handshake reply only; an old `--connect` hits `Unknown handshake type` and exits 1 pre-sentinel, which it already treated as a generic failure. New daemon adopts an old client's pre-written file; new client still passes `--credential-file` so an old daemon reads it as before. Absence of exit 43 is never used as evidence.
**Also.** `terminal create` on a reconnecting SSH host now says what to do instead of a bare `No PTY provider for connection "<id>"` (prefix preserved; the renderer matches it).
## Tests (red first)
- `src/relay/subprocess.test.ts`: two `--detached` starts race one socket + credential file → exactly one reaches the sentinel, loser exits 1 with `Socket path already in use`, file valid + 0600, a `--connect` reading it reaches `relay.status` and reports the winner's pid. Red before (both starters died: daemon required a pre-existing file), green 6/6 after.
- `src/relay/relay-endpoint-credential-publication.test.ts`: mints after bind; adopts a pre-written 0600 file; replaces a pre-written 0644 file with a fresh mint; refuses a stale credential with exit 43 while still serving the real one, and keeps refusing a rewritten file until it is restored.
- `src/relay/relay-reconnect-listener-credential-gate.test.ts`: a client in the bind-to-publish window is refused and never attached; after publication the right credential is accepted and a wrong one refused; a daemon launched without a credential file is not gated. Red without the guard.
- `ssh-relay-deploy-incumbent-verdict.test.ts`: live-but-silent incumbent → `RelayEndpointUnresponsiveError`, refused → `RelayEndpointHeldError`, and in neither case is `--detached` launched; a failed `test -S` probe still launches fresh. Red 2/3 without the deploy change.
- `ssh-relay-deploy-helpers.test.ts` (exit 43), `ssh-relay-endpoint-takeover.test.ts` (refused → Held even with no `lsof`; silent → Unresponsive, nothing unlinked or signalled), `ssh-relay-session-terminal-error.test.ts` (Unresponsive → `onRelayLost`, not terminal). Deploy/namespace/native-deps tests updated to assert the client writes **no** credential.
## Live proof
New `tests/e2e/ssh-docker-relay-stall-credential.spec.ts` (claimed in `run-ssh-docker-e2e.mjs` and PR source routing), two cases: `kill -STOP` every relay pid in the container, send input during the freeze, hold **20 s** (the incident's duration, which races the mux liveness timeout) or **40 s** (past it for sure), `kill -CONT`; assert status back to `connected`, same pty, same daemon pid, same credential inode and content, relay.log did not shrink (a relaunch truncates it) and has zero `Endpoint credential mismatch` / `Socket path already in use` lines, in-stall input delivered at most once.
Run output (local, fixture image `orca-e2e-ssh-relay:3a864c665ba2cefd`, `ORCA_E2E_SSH_DOCKER=1 SKIP_BUILD=1 ORCA_E2E_FORWARD_APP_LOGS=1 … --project electron-headless --workers=1`, head `c2c20fd994`; re-run identically on the final head after the credential-lifetime change, 2 passed (1.7m), same annotations, and the bind-to-publish refusal never fired):
```
✓ keeps the same daemon and credential across a 20s relay freeze (38.3s)
relay-processes-stopped: 2 relay-processes-continued: 2
bridge-pids-before-after: 480 -> 480
socket-clients-accepted-before-after: 1 -> 1
in-stall-input-delivered: 1
✓ backs off and reattaches, never relaunching, across a 40s relay freeze (57.5s)
relay-processes-stopped: 2 relay-processes-continued: 4
bridge-pids-before-after: 480 -> 1202
socket-clients-accepted-before-after: 1 -> 3
in-stall-input-delivered: 1
2 passed (1.6m)
```
Client log in the 40 s case shows the new path end to end: `Relay channel lost … reconnect attempt 1/6` → `Socket probe result: "ALIVE"` → `Socket reconnect failed … Relay failed to start within 10s` → `Relay endpoint incumbent: … verdict=live evidence=accepted-connection holders=unenumerable` → `Failed to re-establish relay … A relay still owns … but did not answer the handshake … Orca will retry` → `reconnect attempt 2/6` → `Reconnected to existing relay via socket`. The 20 s case never left the frozen bridge (same bridge pid, one accept), so it exercises the "silence is not death" side of the same race. The 20 s case passed 6/6 across the session; the 40 s case was red on the prior head (`Socket path already in use` + `Startup failed: listen EADDRINUSE` in relay.log from the swallowed verdict) and is green after 2b. Before the fix the same injection produced a fresh daemon that rewrote the credential and a survivor refusing every client.
The `relay-processes-continued` count exceeds `stopped` in the 40 s case because the timed-out client's `--connect` bridge and the loser-side processes are parked behind the frozen listener when `CONT` runs; they exit on their own once it resumes.
## Gates
`pnpm test src/relay src/main/ssh` 332 files / 3884 tests pass · `pnpm typecheck:tsc:node` clean · `check:code-quality:changed` 0 findings · `check:react-doctor:changed` 0 findings · `pr-e2e-gate-contract.test.mjs` 42 pass · no lint disables or max-lines bumps added.
## Noted, not fixed here
- `terminal list` `orphaned:false` / `terminal close` `ptyKilled:true` for a pane whose relay is gone (`orca-runtime-stop-explicitly-closed-tab-ptys.ts`): different seam, `@ts-nocheck` characterization-covered file.
- On a host with no `lsof`, a stalled relay still cannot be enumerated as the holder; it is now retried rather than declared held, but a relay frozen past the backoff budget still ends in the existing "reconnect manually" banner.
177 lines
6.0 KiB
TypeScript
177 lines
6.0 KiB
TypeScript
import { installRelayLogRotation } from './rotating-log-writer'
|
|
import { readLaunchVersion } from './relay-handshake'
|
|
import type { RelayLaunchOptions } from './relay-launch-options'
|
|
import { RELAY_EMPTY_DETACHED_STARTUP_GRACE_MS, RELAY_IDLE_GRACE_MS } from './relay-launch-options'
|
|
import { relayLogLine } from './relay-diagnostic-log'
|
|
import { RelayPrimaryChannel } from './relay-primary-channel'
|
|
import { RelayRuntimeServices } from './relay-runtime-services'
|
|
import { RelayAgentHookRuntime } from './relay-agent-hook-runtime'
|
|
import { RelaySocketOwnership } from './relay-socket-ownership'
|
|
import { RelayReconnectListener } from './relay-reconnect-listener'
|
|
import { RelayGraceLifecycle } from './relay-grace-lifecycle'
|
|
import {
|
|
publishRelayEndpointCredential,
|
|
restrictWindowsRelayEndpointCredential
|
|
} from './relay-endpoint-credential-publication'
|
|
import { SKILL_RELAY_CAPABILITIES } from './skill-install-handler'
|
|
|
|
export async function runRelayDaemon(options: RelayLaunchOptions): Promise<void> {
|
|
if (options.detached && options.logFile) {
|
|
installRelayLogRotation(options.logFile)
|
|
}
|
|
|
|
const socketOwnership = new RelaySocketOwnership(options.sockPath)
|
|
let fatalPtyHandler: RelayRuntimeServices['ptyHandler'] | null = null
|
|
process.on('uncaughtException', (error) => {
|
|
relayLogLine(`[relay] Uncaught exception: ${error.message}\n${error.stack}`)
|
|
try {
|
|
fatalPtyHandler?.forceKillAllPtyProcesses()
|
|
} catch (reapError) {
|
|
// Why log rather than swallow: exit must still win, but this line is the only
|
|
// forensic trace a crashed remote daemon leaves behind for an orphaned shell.
|
|
relayLogLine(
|
|
`[relay] Fatal PTY reap failed: ${reapError instanceof Error ? reapError.message : String(reapError)}`
|
|
)
|
|
}
|
|
socketOwnership.cleanup()
|
|
process.exit(1)
|
|
})
|
|
process.on('unhandledRejection', (reason) => {
|
|
relayLogLine(`[relay] Unhandled rejection: ${String(reason)}`)
|
|
})
|
|
|
|
const primaryChannel = new RelayPrimaryChannel()
|
|
const launchVersion = readLaunchVersion()
|
|
const runtime = new RelayRuntimeServices(
|
|
primaryChannel.dispatcher,
|
|
options.graceTimeMs,
|
|
launchVersion
|
|
)
|
|
fatalPtyHandler = runtime.ptyHandler
|
|
let reconnectListener: RelayReconnectListener | null = null
|
|
const agentHooks = new RelayAgentHookRuntime(
|
|
primaryChannel.dispatcher,
|
|
runtime.ptyHandler,
|
|
options.sockPath,
|
|
options.endpointDir
|
|
)
|
|
const lifecycle = new RelayGraceLifecycle({
|
|
dispatcher: primaryChannel.dispatcher,
|
|
ptyHandler: runtime.ptyHandler,
|
|
detached: options.detached,
|
|
emptyDetachedStartupGraceMs: RELAY_EMPTY_DETACHED_STARTUP_GRACE_MS,
|
|
idleRelayGraceMs: RELAY_IDLE_GRACE_MS,
|
|
readSocketClientCount: () => reconnectListener?.clientCount ?? 0,
|
|
hasAcceptedSocketClient: () => reconnectListener?.hasAcceptedClient ?? false,
|
|
ownsSocketPath: () => socketOwnership.owned,
|
|
disposeOwnedProcesses: () => runtime.disposeOwnedProcesses(),
|
|
disposeRuntime: () => {
|
|
primaryChannel.dispatcher.dispose()
|
|
runtime.disposeHandlers()
|
|
agentHooks.stop()
|
|
socketOwnership.closeAndCleanup()
|
|
}
|
|
})
|
|
|
|
await agentHooks.start()
|
|
reconnectListener = new RelayReconnectListener(
|
|
primaryChannel.dispatcher,
|
|
socketOwnership,
|
|
launchVersion,
|
|
options.credentialFile,
|
|
{
|
|
detachPrimaryInput: () => primaryChannel.detachInput(),
|
|
cancelGrace: (reason) => lifecycle.cancel(reason),
|
|
onLastClientClosed: () => {
|
|
if (!primaryChannel.isAlive) {
|
|
lifecycle.start('socket client closed')
|
|
}
|
|
}
|
|
}
|
|
)
|
|
const startedAt = Date.now()
|
|
registerRelayStatus(
|
|
primaryChannel,
|
|
runtime,
|
|
reconnectListener,
|
|
socketOwnership,
|
|
lifecycle,
|
|
options,
|
|
startedAt
|
|
)
|
|
|
|
try {
|
|
// Why this order: the bind is the only proof of endpoint ownership. A start that loses it
|
|
// exits inside start() and never reaches the credential file, so racing starters cannot
|
|
// rotate the secret a surviving daemon enforces.
|
|
await reconnectListener.start()
|
|
reconnectListener.setEndpointCredential(publishRelayEndpointCredential(options.credentialFile))
|
|
agentHooks.publishEndpointFile()
|
|
} catch (error) {
|
|
relayLogLine(
|
|
`[relay] Startup failed: ${error instanceof Error ? error.message : String(error)}`
|
|
)
|
|
process.exit(1)
|
|
return
|
|
}
|
|
if (options.credentialFile) {
|
|
void restrictWindowsRelayEndpointCredential(options.credentialFile)
|
|
}
|
|
|
|
primaryChannel.startOutputFailureHandling()
|
|
if (options.detached) {
|
|
lifecycle.start('detached startup')
|
|
} else {
|
|
primaryChannel.startInput({
|
|
onData: () => lifecycle.cancel('stdin data'),
|
|
onDisconnect: (reason) => {
|
|
if ((reconnectListener?.clientCount ?? 0) === 0) {
|
|
lifecycle.start(reason)
|
|
}
|
|
}
|
|
})
|
|
}
|
|
lifecycle.installProcessLifecycle()
|
|
primaryChannel.writeSentinel()
|
|
if (options.detached) {
|
|
primaryChannel.detachPrimaryClient()
|
|
}
|
|
}
|
|
|
|
function registerRelayStatus(
|
|
primaryChannel: RelayPrimaryChannel,
|
|
runtime: RelayRuntimeServices,
|
|
reconnectListener: RelayReconnectListener,
|
|
socketOwnership: RelaySocketOwnership,
|
|
lifecycle: RelayGraceLifecycle,
|
|
options: RelayLaunchOptions,
|
|
startedAt: number
|
|
): void {
|
|
primaryChannel.dispatcher.onRequest('relay.status', async () => ({
|
|
capabilities: SKILL_RELAY_CAPABILITIES,
|
|
pid: process.pid,
|
|
uptimeMs: Date.now() - startedAt,
|
|
detached: options.detached,
|
|
stdoutAlive: primaryChannel.isAlive,
|
|
memory: process.memoryUsage(),
|
|
ptys: { active: runtime.ptyHandler.activePtyCount },
|
|
ptySourceCredit: {
|
|
enabled: true,
|
|
session: runtime.ptyConsumerSessionAdapter.getDebugSnapshot(),
|
|
publication: runtime.ptySourcePublication.getDebugSnapshot()
|
|
},
|
|
socket: {
|
|
path: options.sockPath,
|
|
owned: socketOwnership.owned,
|
|
listening: socketOwnership.server?.listening ?? false,
|
|
clients: reconnectListener.clientCount,
|
|
acceptedConnections: reconnectListener.acceptedConnections
|
|
},
|
|
grace: {
|
|
active: runtime.ptyHandler.graceTimerActive,
|
|
deadlineAt: lifecycle.deadlineAt,
|
|
reason: lifecycle.reason
|
|
}
|
|
}))
|
|
}
|