mirror of
https://github.com/stablyai/orca.git
synced 2026-10-09 08:02:35 +00:00
* fix(supervisor): add a one-shot lifetime that runs past stdin end, and relay a provider's last output A one-shot CLI (claude -p, codex exec -) reads its request until stdin ends, so the supervisor's session rule (stdin end means the owner is closing) would stop it about 1 s into its answer. The one-shot lifetime passes stdin end through and leaves only the owner-death watch and explicit signals to stop it. The supervisor also exited as soon as its provider did, dropping output still in the pipes when the owner reads slowly. It now waits, bounded, for the provider's output to be relayed before exiting. * fix(text-generation): run agent one-shots under the provider supervisor on POSIX The Claude model-list probe, model discovery for every agent, and commit message, pull request and branch name generation all spawned the agent CLI as a plain child of Orca. If Orca quit, crashed or was killed while one was running, nothing stopped it, and a CLI that hung kept running after Orca was gone. They now run under the provider supervisor that native chat already uses, in its one-shot lifetime, so Orca's exit stops the agent's whole process group however Orca exits. A timeout or cancel asks the supervisor to stop (SIGTERM) and only tears the tree down once it has had its full stop time; killing the supervisor first would orphan the agent's group. A missing binary now reaches Orca as the supervisor's exit 127, which is mapped back to the existing not-found message. Windows and WSL keep spawning the agent directly. * fix(supervisor): carry the provider argv on the supervisor's argv, map every spawn error back, and share one stop ladder - The supervisor read the provider's command and arguments from one base64 JSON env string. Linux caps a single env string at 128 KiB, so an argv prompt of about 90-120 KiB, which passes the 120 KiB per-argument guard, failed execve with E2BIG. The provider argv now follows the supervisor script's '--' as real arguments; the env keeps only small fields. - A supervisor that cannot start its provider reports Node's spawn error line and exits 127. That line now becomes the same error a direct spawn emits, so ENOENT still reads as 'not found on PATH' and EACCES or any other spawn error reads as 'failed to start'. - stopSupervisedProvider is the one ask, wait and force ladder: it gives a supervisor its full stop time before forcing. Agent one-shots, the Codex app-server close and the Claude child exit proof now share it, with the same requests, bounds and forced steps as before. * fix(supervisor): report every provider spawn failure on one marked stderr line Node throws most spawn failures (ENOEXEC, ENOTDIR, ELOOP, EPERM, ...) instead of emitting them, and the supervisor had no catch, so it died with exit 1 and a stack trace that the user saw as the agent's failure. A thrown or emitted spawn failure now exits 127 with one marked line carrying whether it was thrown, its code and its message. Orca reads only the last stderr line, so a runtime warning printed earlier cannot hide it, and maps it to the message a direct spawn gave: thrown is 'could not be started', ENOENT is 'not found on PATH', and any other emitted error is 'failed to start'. * fix(supervisor): keep the user's Node options away from the supervisor and hand them to the provider The supervisor runs Electron in Node mode, which honours NODE_OPTIONS and NODE_REPL_EXTERNAL_MODULE. A user value such as a --require of a missing file stopped the supervisor from starting, breaking even native CLIs like Codex that never load it. The launch now takes both out of the supervisor's environment, carries them in its spec, and restores them for the provider only, so a Node-based CLI still gets them. * fix(text-generation): file a forced agent one-shot teardown under its own breadcrumb site The forced tree teardown recorded every self-initiated kill as the Codex app-server's, so a forced commit-message or model-discovery stop read as a Codex teardown in crash breadcrumbs. The teardown now takes the caller's site; source-control stops pass the site their Windows tree kill already uses. * test(text-generation): cover the supervised stop under timeout and output limit; name the direct-child suites The commit-message suites that drive fake children spawn them directly, the unsupervised shape Windows and WSL use, so they now say so. The supervised POSIX stop gets its own compositions: a timed-out Codex generation settles at once but holds the Codex home until its supervisor has stopped (faithful fake, fake timers), and an agent that floods past the output limit is stopped through its real supervisor with no process left behind. * fix(codex): run short-lived app-server sessions under the provider supervisor on POSIX The Codex model-list probe, the hook trust grant and the session index heal each start a short-lived `codex app-server` as a plain child of Orca, and counted on it exiting when its input closes. A wedged Codex (a cold model/list waiting on the network, say) kept running after Orca quit, crashed or was killed. These sessions now run under the provider supervisor that native chat's Codex connection already uses, in its session lifetime, so Orca's exit stops the server's whole process group. The session's end and its deadline share the one stop ladder: end its input (and SIGTERM a session past its deadline), give the supervisor its full stop time, and only then tear the tree down. A missing binary reaches Orca as the supervisor's exit 127 and is mapped back to the spawn error, so the trust-grant telemetry still reads it as a missing binary. Windows and WSL keep spawning the server directly, with the same timings as before. The supervisor now takes only the command line it starts, so the CLI build, which also runs these sessions, no longer pulls in the native chat connection's types. * refactor(supervisor): share the stop of a supervised child process Agent one-shots stop their supervisor with SIGTERM through the shared stop ladder, watching the child's own exit. That adapter moves into one helper so the Codex backfill recovery can use it with its own stop request, instead of a copy. * fix(codex): supervise the app-server that keeps a Codex index backfill alive While Codex rebuilds its session index, Orca keeps a read-only `codex app-server` running for up to an hour so Codex can finish. It was a plain child of Orca with its input held open, so a quit, crash or kill left it running. On POSIX it now runs under the provider supervisor in its session lifetime, so Orca's exit stops its whole process group. Stopping it (done, aborted, or given up) ends its input, as a Codex connection close does; the supervisor then SIGTERMs the group and SIGKILLs it after the grace, and the tree is torn down only if the supervisor outlives its full stop time. Windows and WSL keep the direct spawn and the drain-first probe termination. The spawn and stop of that process move into their own module. * fix(opencode): stop the launch model preflight server with Orca on POSIX Before an OpenCode launch, Orca starts `opencode serve` to read the configured agent and models, then stops it. The server was detached into its own process group and never exits when its input ends, so if Orca quit, crashed or was killed during that preflight (up to 10 s), nothing ever stopped it. On POSIX the server now runs under the provider supervisor in its one-shot lifetime: the preflight's closed input does not stop it, and Orca's exit does. The preflight's teardown asks the supervisor to stop (SIGTERM) and forces the tree only after the supervisor's full stop time; signalling or SIGKILLing the supervisor's own group would orphan the server's. A descendant that ignores SIGTERM is now killed with the group rather than left running once the pipes close. Windows keeps the direct spawn and its existing teardown. * test(codex): pin when supervised session and backfill stops escalate A session past its deadline is SIGTERMed through its supervisor rather than waiting out the stdin-end grace, and its tree is torn down only after the supervisor's full stop time; Windows keeps its deadline kill and 1.5 s close wait. A supervised backfill app-server is stopped by ending its input, with the same full stop time before any teardown. * test(text-generation): run the direct-child suites on the Windows path and cover supervised discovery The commit-message suites that drive fake children mocked the supervisor away on POSIX, so they asserted a direct root SIGKILL that production no longer takes there. They now pin the platform to Windows (with an empty PATH, so host installs cannot answer a bare agent name) and assert the Windows kill, taskkill included. The three tests that check the host's own discovery spawn shape run on the host and read the agent argv past the supervisor's '--'. Model discovery gets its supervised composition: a timed-out Codex discovery settles at once but holds the Codex home until its supervisor has stopped. * fix(supervisor): show a supervised spawn failure in native chat as the spawn error it was Native chat's exit errors carry the provider's stderr tail into Details. Under the supervisor a missing CLI left the supervisor's internal spawn-failure report there instead of Node's own 'spawn <cmd> ENOENT'. The report, its parser and a display formatter now live in one module; the Codex app-server and Claude stream-json exit errors pass the tail through the formatter, which turns a report back into the spawn error and leaves any other stderr unchanged. * fix(supervisor): report a spawn that failed without a pid instead of crashing on its missing pipes When the provider spawn fails outright (EMFILE, ENFILE), Node emits 'error' later and leaves the child with no pid and no stdio. Piping stdin into the missing pipe threw first, so the supervisor died with exit 1 and a stack trace and never wrote its spawn-failure report. The pipes are now wired only for a provider that started. * refactor(supervisor): share the stop of a supervised child process Agent one-shots stop their supervisor with SIGTERM through the shared stop ladder, watching the child's own exit. That adapter moves into one helper beside the ladder, so other supervised children can use it with their own stop request instead of a copy. The caller's breadcrumb site still reaches the forced teardown. Same request, wait and force as before. * fix(codex,opencode): file forced backfill and preflight teardowns under their own breadcrumb sites A forced teardown of the Codex backfill app-server is filed under 'codex-state-db-backfill-recovery', and one of the OpenCode launch model preflight under 'opencode-launch-model-preflight', instead of the generic 'codex-app-server-teardown'. * test(codex): write the stand-in pid report atomically A loaded host let the test read the pid file between its creation and its write (Unexpected end of JSON input); the stand-in now renames it into place. * fix(supervisor): give a session provider its stdin end and grace when its owner dies The owner-death watch went straight to the group SIGTERM and cancelled any stdin-end grace, so when Orca quit or crashed a session provider such as the Codex app-server never saw the EOF that lets it finish writing its state (auth.json, the state database). A session whose owner is gone now closes as an owner's stdin end does: the provider's stdin is ended, it gets the stdin-end grace, and only then the SIGTERM and SIGKILL ladder. A one-shot already had its EOF at the end of its request, so its owner's death still stops it at once. * fix(opencode): stop the preflight server through the shared supervised stop, and trust only a proven stop The preflight's own stop wrapper waited on the supervisor's pipes and counted a forced teardown as proof, though the teardown reports success even when it found no descendants to check, and a supervisor that failed to reap its group exits 1 with its pipes closed. It now uses stopSupervisedChildProcess with its breadcrumb site, and counts the server stopped only when the supervisor ended on its stop signal or relayed the server's own exit. A forced stop, or an exit of 1, returns no context, as an unverified stop did before. * chore(codex): state the supervised stop time in the probe and trust grant deadline comments * test(codex): assert the signals a supervised stop sends itself, not a SIGKILL the mocked teardown never could * fix(supervisor): kill the rest of the provider group once the provider exits on a stop A requested stop waited out the whole SIGTERM grace for the provider's group even after the provider itself had exited, so a SIGTERM-ignoring helper it left behind held every stop for up to 3 s. Under a stop, the rest of the group is now SIGKILLed as soon as the provider has exited, the same rule its own exit already follows. * test(text-generation): cover a supervised Codex discovery past its output limit Model discovery's supervised stop was covered only under timeout. A Codex discovery that floods past the output limit now runs through a real supervisor: it settles with the too-much-data error, its agent is stopped through the supervisor, and the next discovery on the same Codex home starts only after that agent is gone. The direct-child suites' headers now list exactly the supervised cases that are covered. * fix(supervisor): close a provider whose owner is gone the way its owner closes it Owner death gave every session provider the stdin-end grace, so after an Orca crash a Claude session, whose close is a stdin end plus SIGTERM, could keep working on its turn for a second with nobody watching. The spawn spec now names the provider's close request: 'stdin-end' (the Codex app-server drains and exits on EOF, then gets its grace) or 'stdin-end-and-sigterm' (Claude; the default). A gone owner gets that same request. One constant per provider feeds both its spawn spec and its owner-side close, through one requestProviderClose, so the two cannot drift. One-shots still stop at once. * fix(codex): close short-lived sessions and the backfill app-server the way a Codex connection closes Codex finishes its writes and exits on its stdin end, so the Codex connection's supervisor closes it by ending stdin, and an owner that is gone now gets that same close. The short-lived sessions and the backfill app-server now name the same close request, through one shared constant, in their spawn spec and in their own close: Orca quitting or crashing gives them the stdin end and its grace before SIGTERM, instead of an immediate SIGTERM. A session past its deadline still adds a SIGTERM, since it is wedged. * test(codex,opencode): an owner's death drains Codex before SIGTERM, and the preflight stop no longer waits out the grace The stand-in now records when its stdin ended and when SIGTERM arrived. A SIGKILLed owner leaves a Codex session or backfill app-server its stdin-end grace before SIGTERM, and the OpenCode preflight ends within 1 s of its server's SIGTERM even with a descendant that ignores SIGTERM. * test(codex): give the trust-grant deadline tests room for a supervised start on a loaded host At 500 ms a loaded full run hit the deadline before the supervisor had started the stub, which never wrote the pid the test reads. * fix(supervisor): keep the SIGTERM grace for a session's group after its provider exits Killing the rest of the group the moment the provider exited under a stop also reached native chat's closes, so an MCP server, a tool's child or a dev server still in Claude's or Codex's group was SIGKILLed mid-cleanup instead of getting the rest of the SIGTERM grace. The early group kill now applies only to one-shots, where the saved wait was the point; a session's stop is back to waiting out the grace for its group. * test(claude): pin that Claude's spawn passes its close request explicitly The spawn-spec assertion matched the default close request, so dropping Claude's explicit request still passed. The test now checks that the spec is built with the exit-proof ladder's own constant. * refactor(codex): give the Codex app-server close request its own module Other Codex app-server spawns will name the same close request as the connection does. Holding it in its own small module lets them import it without the connection itself. * build(cli): list the Codex close request and the provider supervisor in the CLI project The command-line build runs the short-lived Codex app-server session, which will name the same close request as the Codex connection. Listing the close request, the provider supervisor it takes its type from, and the spawn-failure report the supervisor uses lets the CLI project typecheck that import without pulling the connection in. * test(codex): give a stand-in 20 s to report its pids on a loaded host At a load average near 50 both owner-death tests timed out waiting for the bundled owner's stand-in at 10 s; they pass alone. * test(codex): start the deadline tests' clocks past the stand-in's start, and cover a server that ignores SIGTERM The session and trust-grant deadline tests ran a 2-4 s deadline from spawn, so a loaded host could stop the stand-in before it wrote its pid. The session test's deadline now outlasts its pid-read budget, and the trust-grant deadlines are 8 s. The trust-grant comment said a wedged server may ignore everything but SIGKILL, but that stub dies on its stdin end; a new POSIX case pins that a server ignoring both its stdin end and SIGTERM is SIGKILLed after the SIGTERM grace. * test(text-generation): check the ENOEXEC start failure only where Node reports one On Linux, glibc's execvp hands an executable that is not a program to /bin/sh, so both a direct and a supervised spawn run it and it exits 127; only macOS throws ENOEXEC. The not-a-program case now runs on macOS only; the path-through-a-file case (ENOTDIR) still covers a thrown start failure everywhere. * test(wsl): follow the backfill's wsl.exe spawn into its new process module The WSL invocation boundary lists files that spawn wsl.exe directly. The backfill recovery's spawn moved into codex-state-db-backfill-recovery-process.ts, so the entry moves with it; the count is unchanged. * test(codex): keep factory child_process mocks loadable now that Codex stops reach the process-table reader The backfill recovery and Codex sessions now stop through the shared supervised teardown, whose process-table reader binds execFile when it loads. Five rate-limit fetcher tests mocked node:child_process with only spawn and failed at load: they now mock the backfill recovery, as their sibling fetcher tests already do. The account add-login tests' child_process mocks gain an execFile stub. * fix(opencode): count no forced preflight stop as proof the server is gone A forced teardown walks and group-kills the supervisor's tree, but the server leads its own detached group, so a teardown verdict of 'exited' does not cover members left in the server's group. Only a supervisor that reaped the group itself, by its own stop or relaying the server's exit, now counts as a proven stop. The Windows session-stop test now says why it sees no direct kill.