mirror of
https://github.com/stablyai/orca.git
synced 2026-09-22 00:02:31 +00:00
* fix(terminal): fence daemon endpoint ownership * fix(terminal): clean failed daemon PID claims * fix(terminal): close daemon ownership review gaps * test(daemon): release startup IPC in boot smoke * test(daemon): mirror production stdio in boot smoke * fix(daemon): exit after rpc shutdown cleanup * fix(terminal): make the socket name the daemon endpoint authority The reported failure was a live daemon hosting PTYs that nothing could reach: terminals acknowledged input and never ran it, listings diverged from reality, and restarting the app never helped because the detached helper survived. The ownership fence added for it could not fire in the sequence that produces the split brain. libuv unlinks the pathname a server bound to when that server closes, with no ownership check. A daemon that lost its endpoint name therefore deleted whichever socket then sat at that path — including a live replacement's — stranding a daemon that still hosted every session. Bind a private same-directory name and hard-link it into place instead: libuv can only ever unlink our own bind name, the exclusive link is a kernel-enforced endpoint claim, and the canonical name is removed only under an inode ownership check. The bind name replaces the basename rather than extending it, so it cannot overflow sun_path. killStaleDaemon removed the PID record unconditionally immediately before every fork, so the exclusive PID claim was always uncontested at bind time. It also unlinked a live daemon's endpoint whenever a connect probe merely timed out, and treated a `ps` timeout as proof of PID recycling. Now only positive evidence of a dead endpoint authorizes reclaiming it, SIGKILL is confirmed rather than assumed, and a daemon that cannot be proven stopped keeps its record and endpoint while the launcher refuses to fork beside it. A daemon whose endpoint was taken over now retires itself, draining rather than killing, so an unreachable orphan stops being permanent. A repaired PID record re-derives entryPath, appVersion and the Linux incarnation markers from the authenticated owner instead of dropping them; without appVersion a healthy daemon read as a permanently stale bundle and, on Windows, went unpinned against daemon-host pruning. Repair failure now fails open — abandoning a healthy daemon over a pid file write cost every persistent terminal on the machine. Also: treat only ENOENT as an unclaimed record so a Windows file lock is not reported as an ownership conflict; settle start() before close() so an accepted connection cannot defer it forever; sweep abandoned claim and bind names; and type the endpoint-identity seam so a rename cannot silently disable the fence. Adds a real-process handover smoke that reproduces the failure with two daemons racing one endpoint, and wires it into the native-smoke job. * fix(daemon): retire only on proven endpoint ownership loss The ownership watchdog read a null identity for any stat failure, so a transient EACCES or EIO on the runtime directory would retire a daemon that was still serving every terminal on the machine. Distinguish "the entry is gone" from "the probe failed" and act only on the former. Also require the loss to persist across two polls: a replacement publishes by unlink-then-link, and a single observation can land in that gap. * fix(daemon): source repaired ownership metadata from the authenticated hello Adversarial review found three defects in the previous two commits. Re-deriving entryPath from the owner's command line truncated it at the first space. A command line is a single space-joined string, so `C:\Program Files\Orca\...` and `/Applications/Orca 2.app/...` came back as `"C:\Program` and `/Applications/Orca`. getDaemonLaunchIdentity treats a present entryPath as authoritative, so a healthy daemon read as `different_app_path` and was killed and re-forked — worse than the missing-metadata case the derivation was added to fix. Carry entryPath and appVersion as optional fields on the daemon hello identity instead: the daemon already has both from its own argv, and per docs/reference/remote-wire-compatibility.md a new optional field is safe because every reader falls back when it is absent. This also removes a synchronous `ps` spawn from the Electron main thread during startup. `start()` rolled back the PID record even when it never published one. Losing the endpoint link now runs that path, and the ownership-checked unlink briefly renames the incumbent's record aside — enough to strand a live daemon's ownership. Roll back only what we actually wrote. publishDaemonSocketPath read its identity from the canonical name after linking, so a concurrent unlink returned null: no ownership watchdog and no endpoint cleanup on any shutdown path. Read it from the bound name before linking, which shares the inode. Refusing to fork beside an unconfirmed daemon left the user with no daemon at all and no in-app recovery, since restart re-entered the same fence. We have just proved something answers the endpoint, so adopt it in degraded mode: live sessions keep working, fresh terminals run locally. SIGTERM is also individually guarded now — an EPERM fell into the blanket catch and reported "nothing alive", authorizing the very duplicate this fence exists to prevent. Also reset the ownership-loss streak on an inconclusive probe so the confirmations are consecutive, and sweep scratch names before the launch so a failed launch still reclaims them.