Files
orca/src
Brennan Benson 61bcca9fee fix: stop Codex and OpenCode helper servers when Orca quits or crashes (STA-9254, 2 of 2) (#25753)
* fix(supervisor): add a one-shot lifetime that runs past stdin end, and relay a provider's last output

A one-shot CLI (claude -p, codex exec -) reads its request until stdin ends, so the
supervisor's session rule (stdin end means the owner is closing) would stop it about
1 s into its answer. The one-shot lifetime passes stdin end through and leaves only
the owner-death watch and explicit signals to stop it.

The supervisor also exited as soon as its provider did, dropping output still in the
pipes when the owner reads slowly. It now waits, bounded, for the provider's output
to be relayed before exiting.

* fix(text-generation): run agent one-shots under the provider supervisor on POSIX

The Claude model-list probe, model discovery for every agent, and commit message,
pull request and branch name generation all spawned the agent CLI as a plain child
of Orca. If Orca quit, crashed or was killed while one was running, nothing stopped
it, and a CLI that hung kept running after Orca was gone.

They now run under the provider supervisor that native chat already uses, in its
one-shot lifetime, so Orca's exit stops the agent's whole process group however
Orca exits. A timeout or cancel asks the supervisor to stop (SIGTERM) and only
tears the tree down once it has had its full stop time; killing the supervisor
first would orphan the agent's group. A missing binary now reaches Orca as the
supervisor's exit 127, which is mapped back to the existing not-found message.
Windows and WSL keep spawning the agent directly.

* fix(supervisor): carry the provider argv on the supervisor's argv, map every spawn error back, and share one stop ladder

- The supervisor read the provider's command and arguments from one base64 JSON env string.
  Linux caps a single env string at 128 KiB, so an argv prompt of about 90-120 KiB, which
  passes the 120 KiB per-argument guard, failed execve with E2BIG. The provider argv now
  follows the supervisor script's '--' as real arguments; the env keeps only small fields.
- A supervisor that cannot start its provider reports Node's spawn error line and exits 127.
  That line now becomes the same error a direct spawn emits, so ENOENT still reads as
  'not found on PATH' and EACCES or any other spawn error reads as 'failed to start'.
- stopSupervisedProvider is the one ask, wait and force ladder: it gives a supervisor its full
  stop time before forcing. Agent one-shots, the Codex app-server close and the Claude child
  exit proof now share it, with the same requests, bounds and forced steps as before.

* fix(supervisor): report every provider spawn failure on one marked stderr line

Node throws most spawn failures (ENOEXEC, ENOTDIR, ELOOP, EPERM, ...) instead of emitting
them, and the supervisor had no catch, so it died with exit 1 and a stack trace that the
user saw as the agent's failure. A thrown or emitted spawn failure now exits 127 with one
marked line carrying whether it was thrown, its code and its message. Orca reads only the
last stderr line, so a runtime warning printed earlier cannot hide it, and maps it to the
message a direct spawn gave: thrown is 'could not be started', ENOENT is 'not found on
PATH', and any other emitted error is 'failed to start'.

* fix(supervisor): keep the user's Node options away from the supervisor and hand them to the provider

The supervisor runs Electron in Node mode, which honours NODE_OPTIONS and
NODE_REPL_EXTERNAL_MODULE. A user value such as a --require of a missing file stopped
the supervisor from starting, breaking even native CLIs like Codex that never load it.
The launch now takes both out of the supervisor's environment, carries them in its spec,
and restores them for the provider only, so a Node-based CLI still gets them.

* fix(text-generation): file a forced agent one-shot teardown under its own breadcrumb site

The forced tree teardown recorded every self-initiated kill as the Codex app-server's, so
a forced commit-message or model-discovery stop read as a Codex teardown in crash
breadcrumbs. The teardown now takes the caller's site; source-control stops pass the
site their Windows tree kill already uses.

* test(text-generation): cover the supervised stop under timeout and output limit; name the direct-child suites

The commit-message suites that drive fake children spawn them directly, the unsupervised
shape Windows and WSL use, so they now say so. The supervised POSIX stop gets its own
compositions: a timed-out Codex generation settles at once but holds the Codex home until
its supervisor has stopped (faithful fake, fake timers), and an agent that floods past the
output limit is stopped through its real supervisor with no process left behind.

* fix(codex): run short-lived app-server sessions under the provider supervisor on POSIX

The Codex model-list probe, the hook trust grant and the session index heal each
start a short-lived `codex app-server` as a plain child of Orca, and counted on it
exiting when its input closes. A wedged Codex (a cold model/list waiting on the
network, say) kept running after Orca quit, crashed or was killed.

These sessions now run under the provider supervisor that native chat's Codex
connection already uses, in its session lifetime, so Orca's exit stops the
server's whole process group. The session's end and its deadline share the one
stop ladder: end its input (and SIGTERM a session past its deadline), give the
supervisor its full stop time, and only then tear the tree down. A missing binary
reaches Orca as the supervisor's exit 127 and is mapped back to the spawn error,
so the trust-grant telemetry still reads it as a missing binary. Windows and WSL
keep spawning the server directly, with the same timings as before.

The supervisor now takes only the command line it starts, so the CLI build, which
also runs these sessions, no longer pulls in the native chat connection's types.

* refactor(supervisor): share the stop of a supervised child process

Agent one-shots stop their supervisor with SIGTERM through the shared stop ladder,
watching the child's own exit. That adapter moves into one helper so the Codex
backfill recovery can use it with its own stop request, instead of a copy.

* fix(codex): supervise the app-server that keeps a Codex index backfill alive

While Codex rebuilds its session index, Orca keeps a read-only `codex app-server`
running for up to an hour so Codex can finish. It was a plain child of Orca with
its input held open, so a quit, crash or kill left it running.

On POSIX it now runs under the provider supervisor in its session lifetime, so
Orca's exit stops its whole process group. Stopping it (done, aborted, or given
up) ends its input, as a Codex connection close does; the supervisor then
SIGTERMs the group and SIGKILLs it after the grace, and the tree is torn down
only if the supervisor outlives its full stop time. Windows and WSL keep the
direct spawn and the drain-first probe termination.

The spawn and stop of that process move into their own module.

* fix(opencode): stop the launch model preflight server with Orca on POSIX

Before an OpenCode launch, Orca starts `opencode serve` to read the configured
agent and models, then stops it. The server was detached into its own process
group and never exits when its input ends, so if Orca quit, crashed or was killed
during that preflight (up to 10 s), nothing ever stopped it.

On POSIX the server now runs under the provider supervisor in its one-shot
lifetime: the preflight's closed input does not stop it, and Orca's exit does.
The preflight's teardown asks the supervisor to stop (SIGTERM) and forces the
tree only after the supervisor's full stop time; signalling or SIGKILLing the
supervisor's own group would orphan the server's. A descendant that ignores
SIGTERM is now killed with the group rather than left running once the pipes
close. Windows keeps the direct spawn and its existing teardown.

* test(codex): pin when supervised session and backfill stops escalate

A session past its deadline is SIGTERMed through its supervisor rather than
waiting out the stdin-end grace, and its tree is torn down only after the
supervisor's full stop time; Windows keeps its deadline kill and 1.5 s close
wait. A supervised backfill app-server is stopped by ending its input, with the
same full stop time before any teardown.

* test(text-generation): run the direct-child suites on the Windows path and cover supervised discovery

The commit-message suites that drive fake children mocked the supervisor away on POSIX, so
they asserted a direct root SIGKILL that production no longer takes there. They now pin the
platform to Windows (with an empty PATH, so host installs cannot answer a bare agent name)
and assert the Windows kill, taskkill included. The three tests that check the host's own
discovery spawn shape run on the host and read the agent argv past the supervisor's '--'.
Model discovery gets its supervised composition: a timed-out Codex discovery settles at once
but holds the Codex home until its supervisor has stopped.

* fix(supervisor): show a supervised spawn failure in native chat as the spawn error it was

Native chat's exit errors carry the provider's stderr tail into Details. Under the supervisor
a missing CLI left the supervisor's internal spawn-failure report there instead of Node's own
'spawn <cmd> ENOENT'. The report, its parser and a display formatter now live in one module;
the Codex app-server and Claude stream-json exit errors pass the tail through the formatter,
which turns a report back into the spawn error and leaves any other stderr unchanged.

* fix(supervisor): report a spawn that failed without a pid instead of crashing on its missing pipes

When the provider spawn fails outright (EMFILE, ENFILE), Node emits 'error' later and leaves
the child with no pid and no stdio. Piping stdin into the missing pipe threw first, so the
supervisor died with exit 1 and a stack trace and never wrote its spawn-failure report. The
pipes are now wired only for a provider that started.

* refactor(supervisor): share the stop of a supervised child process

Agent one-shots stop their supervisor with SIGTERM through the shared stop ladder, watching
the child's own exit. That adapter moves into one helper beside the ladder, so other
supervised children can use it with their own stop request instead of a copy. The caller's
breadcrumb site still reaches the forced teardown. Same request, wait and force as before.

* fix(codex,opencode): file forced backfill and preflight teardowns under their own breadcrumb sites

A forced teardown of the Codex backfill app-server is filed under
'codex-state-db-backfill-recovery', and one of the OpenCode launch model
preflight under 'opencode-launch-model-preflight', instead of the generic
'codex-app-server-teardown'.

* test(codex): write the stand-in pid report atomically

A loaded host let the test read the pid file between its creation and its
write (Unexpected end of JSON input); the stand-in now renames it into place.

* fix(supervisor): give a session provider its stdin end and grace when its owner dies

The owner-death watch went straight to the group SIGTERM and cancelled any stdin-end grace,
so when Orca quit or crashed a session provider such as the Codex app-server never saw the
EOF that lets it finish writing its state (auth.json, the state database). A session whose
owner is gone now closes as an owner's stdin end does: the provider's stdin is ended, it
gets the stdin-end grace, and only then the SIGTERM and SIGKILL ladder. A one-shot already
had its EOF at the end of its request, so its owner's death still stops it at once.

* fix(opencode): stop the preflight server through the shared supervised stop, and trust only a proven stop

The preflight's own stop wrapper waited on the supervisor's pipes and counted a
forced teardown as proof, though the teardown reports success even when it found
no descendants to check, and a supervisor that failed to reap its group exits 1
with its pipes closed. It now uses stopSupervisedChildProcess with its breadcrumb
site, and counts the server stopped only when the supervisor ended on its stop
signal or relayed the server's own exit. A forced stop, or an exit of 1, returns
no context, as an unverified stop did before.

* chore(codex): state the supervised stop time in the probe and trust grant deadline comments

* test(codex): assert the signals a supervised stop sends itself, not a SIGKILL the mocked teardown never could

* fix(supervisor): kill the rest of the provider group once the provider exits on a stop

A requested stop waited out the whole SIGTERM grace for the provider's group even after the
provider itself had exited, so a SIGTERM-ignoring helper it left behind held every stop for
up to 3 s. Under a stop, the rest of the group is now SIGKILLed as soon as the provider has
exited, the same rule its own exit already follows.

* test(text-generation): cover a supervised Codex discovery past its output limit

Model discovery's supervised stop was covered only under timeout. A Codex discovery that
floods past the output limit now runs through a real supervisor: it settles with the
too-much-data error, its agent is stopped through the supervisor, and the next discovery on
the same Codex home starts only after that agent is gone. The direct-child suites' headers
now list exactly the supervised cases that are covered.

* fix(supervisor): close a provider whose owner is gone the way its owner closes it

Owner death gave every session provider the stdin-end grace, so after an Orca crash a
Claude session, whose close is a stdin end plus SIGTERM, could keep working on its turn
for a second with nobody watching. The spawn spec now names the provider's close request:
'stdin-end' (the Codex app-server drains and exits on EOF, then gets its grace) or
'stdin-end-and-sigterm' (Claude; the default). A gone owner gets that same request. One
constant per provider feeds both its spawn spec and its owner-side close, through one
requestProviderClose, so the two cannot drift. One-shots still stop at once.

* fix(codex): close short-lived sessions and the backfill app-server the way a Codex connection closes

Codex finishes its writes and exits on its stdin end, so the Codex connection's
supervisor closes it by ending stdin, and an owner that is gone now gets that
same close. The short-lived sessions and the backfill app-server now name the
same close request, through one shared constant, in their spawn spec and in
their own close: Orca quitting or crashing gives them the stdin end and its
grace before SIGTERM, instead of an immediate SIGTERM. A session past its
deadline still adds a SIGTERM, since it is wedged.

* test(codex,opencode): an owner's death drains Codex before SIGTERM, and the preflight stop no longer waits out the grace

The stand-in now records when its stdin ended and when SIGTERM arrived. A
SIGKILLed owner leaves a Codex session or backfill app-server its stdin-end grace
before SIGTERM, and the OpenCode preflight ends within 1 s of its server's SIGTERM
even with a descendant that ignores SIGTERM.

* test(codex): give the trust-grant deadline tests room for a supervised start on a loaded host

At 500 ms a loaded full run hit the deadline before the supervisor had started
the stub, which never wrote the pid the test reads.

* fix(supervisor): keep the SIGTERM grace for a session's group after its provider exits

Killing the rest of the group the moment the provider exited under a stop also reached
native chat's closes, so an MCP server, a tool's child or a dev server still in Claude's or
Codex's group was SIGKILLed mid-cleanup instead of getting the rest of the SIGTERM grace.
The early group kill now applies only to one-shots, where the saved wait was the point;
a session's stop is back to waiting out the grace for its group.

* test(claude): pin that Claude's spawn passes its close request explicitly

The spawn-spec assertion matched the default close request, so dropping Claude's explicit
request still passed. The test now checks that the spec is built with the exit-proof ladder's
own constant.

* refactor(codex): give the Codex app-server close request its own module

Other Codex app-server spawns will name the same close request as the connection does.
Holding it in its own small module lets them import it without the connection itself.

* build(cli): list the Codex close request and the provider supervisor in the CLI project

The command-line build runs the short-lived Codex app-server session, which will name the
same close request as the Codex connection. Listing the close request, the provider
supervisor it takes its type from, and the spawn-failure report the supervisor uses lets the
CLI project typecheck that import without pulling the connection in.

* test(codex): give a stand-in 20 s to report its pids on a loaded host

At a load average near 50 both owner-death tests timed out waiting for the
bundled owner's stand-in at 10 s; they pass alone.

* test(codex): start the deadline tests' clocks past the stand-in's start, and cover a server that ignores SIGTERM

The session and trust-grant deadline tests ran a 2-4 s deadline from spawn, so a
loaded host could stop the stand-in before it wrote its pid. The session test's
deadline now outlasts its pid-read budget, and the trust-grant deadlines are 8 s.
The trust-grant comment said a wedged server may ignore everything but SIGKILL,
but that stub dies on its stdin end; a new POSIX case pins that a server ignoring
both its stdin end and SIGTERM is SIGKILLed after the SIGTERM grace.

* test(text-generation): check the ENOEXEC start failure only where Node reports one

On Linux, glibc's execvp hands an executable that is not a program to /bin/sh, so both a
direct and a supervised spawn run it and it exits 127; only macOS throws ENOEXEC. The
not-a-program case now runs on macOS only; the path-through-a-file case (ENOTDIR) still
covers a thrown start failure everywhere.

* test(wsl): follow the backfill's wsl.exe spawn into its new process module

The WSL invocation boundary lists files that spawn wsl.exe directly. The
backfill recovery's spawn moved into codex-state-db-backfill-recovery-process.ts,
so the entry moves with it; the count is unchanged.

* test(codex): keep factory child_process mocks loadable now that Codex stops reach the process-table reader

The backfill recovery and Codex sessions now stop through the shared supervised
teardown, whose process-table reader binds execFile when it loads. Five rate-limit
fetcher tests mocked node:child_process with only spawn and failed at load: they
now mock the backfill recovery, as their sibling fetcher tests already do. The
account add-login tests' child_process mocks gain an execFile stub.

* fix(opencode): count no forced preflight stop as proof the server is gone

A forced teardown walks and group-kills the supervisor's tree, but the server
leads its own detached group, so a teardown verdict of 'exited' does not cover
members left in the server's group. Only a supervisor that reaped the group
itself, by its own stop or relaying the server's exit, now counts as a proven
stop. The Windows session-stop test now says why it sees no direct kill.
2026-10-06 17:24:19 -07:00
..