Files
orca/package.json
T
NeilandOrca 2e8cf589de fix(daemon): stop killing live coding agents when the daemon can't report its sessions (#13928)
* fix(daemon): stop killing a wedged daemon that still owns live agent PTYs

A daemon too busy to answer listSessions was indistinguishable from a dead
one: getAliveDaemonSessionCount() returns null ("could not verify"), the
preserve gate required `!== null && > 0`, so the run fell through to
killStaleDaemon() and every running coding agent died with it. The sibling
replace branches all preserve on null; this one alone collapsed "can't tell"
into "empty", which src/main/daemon/AGENTS.md already forbids.

Give the decision an out-of-band second opinion. inspectDaemonPtyOwnership()
reads the OS process table — never the daemon socket, which is exactly what
failed — and reports whether the daemon's own process still has live PTY
descendants. Under preserveWhenOwningLivePtys, that evidence vetoes the
signal and the launcher adopts the daemon in degraded mode instead.

The veto is opt-in so it cannot make a daemon unkillable: only the
failed_health_check path enables it. Manage Sessions -> Restart still kills.
Only positive evidence preserves, so a wedged daemon with nothing to lose is
still replaced (#8689).

Also stop replacing silently: the verdict now prints on the post-kill truth,
which stays quiet on a cold start because nothing was killed.

Co-authored-by: Orca <help@stably.ai>

* fix(daemon): survive preserving a daemon too wedged to be adopted

Adversarial review found the veto's own success path could not complete
against the daemon it exists to protect. Preserving routed through
holdDaemonAdoptionLease(), which opens a hello — the exact operation a
wedged daemon cannot answer — so it threw, aborted initDaemonPtyProvider,
and left no spawner. restartDaemon() throws without one, so the user lost
the documented Manage Sessions -> Restart remedy on top of having no
daemon: strictly worse than the data loss being fixed.

A still-listening endpoint means wedged, not gone, so keep a lease-free
handle instead. The lease only cancels the adoption watchdog, which never
fires on a daemon that owns sessions. Degraded mode likewise tolerates a
lease and a session discovery it cannot complete.

Three more from the same review:

- The veto keyed on reason === 'failed_health_check', but a daemon that
  answered listSessions with 0 lands in that same branch and must stay
  replaceable. Key on liveSessionCount === null, which is what the option
  actually documents.
- Zombies are not evidence of live work. A wedged daemon cannot reap, so
  its exited agents linger as <defunct> and would read as "still running"
  — a false positive correlated with the wedge itself. Enumerate through
  the process table's stat column and exclude them; sample twice so a
  resolver probe or health-check shell cannot masquerade as an agent.
- Restore the "did anything answer?" log guard alongside the confirmed
  kill, so a daemon that self-retires before the kill is still announced.

Co-authored-by: Orca <help@stably.ai>

* test(daemon): kill the mutations that let the PTY veto ship as a no-op

Mutation testing found four survivors — changes that break the fix while
every test stays green:

- Swapping the POSIX reader to the 500ms-cached one passed. It is not just
  a staleness hazard: inside the TTL both sampling attempts receive the same
  array, collapsing the two-sample confirmation to one. Pin the fresh reader.
- Replacing killStaleDaemon's default inspector with one that never reports
  live PTYs — the veto disabled in production — passed, because every veto
  test injects the hook. Exercise the real seam.
- Adding the veto to cleanupDaemonForProtocol passed, which is verbatim the
  failure its own doc warns about: a user-initiated restart of a daemon
  owning live PTYs would refuse, then throw. Pin that call's arity.
- The ppid-cycle fixture put the cycle outside the daemon's subtree, so the
  walk never entered it and deleting the visited guard passed.

Also bound the Windows enumeration, which had no budget of its own: two CIM
queries with a wmic fallback can stall the launch path for tens of seconds.
Blind is a safe answer there; hanging is not.

Co-authored-by: Orca <help@stably.ai>

* fix(daemon): require session-leader evidence and stop preserving daemons that can never be adopted

Round-three review found the veto too eager in three ways, each of which
traded the original data loss for a whole-session degrade or a permanently
daemon-less app.

Evidence was "any non-zombie descendant", justified by re-sampling to weed
out transients. The two samples are taken back to back — one ps fork apart —
so nothing transient is ever weeded out, and a hung helper (often the very
reason the daemon is wedged) reads as an agent. Use the structural signal
instead: a PTY child is a session leader, because forkpty calls setsid, and
no helper the daemon forks ever is. Re-sampling now only retries blindness.

A 'rejected' daemon answered and refused the handshake, so it can never be
adopted; preserving it repeated the same failed adoption on every launch,
forever. And the veto read the process table, not the socket, so it could
fire on a daemon whose endpoint was already gone — adoption then threw,
init aborted, and the app was left with no spawner and no working Restart.
Gate on both: only preserve what could still be reached.

Also: releasing the launcher's temporary lease after the permanent lease
failed reopened the adoption gap that ordering exists to close, and the
tolerance added to discoverDaemonSessions was dead code — nothing on that
path rejects.

Co-authored-by: Orca <help@stably.ai>

* refactor(daemon): decide occupancy before the kill, not inside it

Three review rounds each found a new failure state in the previous shape,
which was the design telling us something. The safety rule — never destroy
running work — was replicated across the launcher's branches instead of
being decided once, and the last round added it to one more branch behind a
boolean. Policy had been put inside a mechanism: killStaleDaemon grew an
input flag to disable its new veto, an output back-channel to report it, and
a caller-side re-derivation of the classification the flag had lost. One
structural error, one symptom per layer it crossed.

Name the question instead. resolveDaemonOccupancy answers occupied | empty |
unknown, asking the daemon first (authoritative both ways) and falling back
to the process table only to RAISE the answer to occupied. OS evidence can
prove work exists; it can never prove absence, so it never licenses a kill.

The launcher now decides before it destroys anything, so killStaleDaemon
goes back to being only "make this pid go away" — no options, nothing for a
future caller to forget to disable, and Manage Sessions -> Restart cannot be
vetoed because there is no veto left to hit.

Holding is a real outcome now. A daemon that owns live terminals but cannot
answer a handshake gets mode 'held': no adoption attempt, no lease, no fork
beside it. That deletes the lease-free-handle fallback, the try/catch around
preserve, and the tolerated-lease branch in init that existed only because
the correct outcome had no representation.

Two defects this removes outright:

- The endpoint check used a local boolean probe that returns false on
  timeout, so under load — and unconditionally on Windows named pipes, where
  a busy server answers ERROR_PIPE_BUSY — the guard disabled itself in
  exactly the conditions it was written for. Use the canonical three-valued
  probe, whose own docs say absence of proof is not proof of death.
- A daemon that answered and refused the handshake was killed with its
  agents. It is now held like any other occupied daemon.

Also folds the replacement verdict onto pendingReplacement, retiring a pair
of mutable launcher locals whose only job was moving one warning past the
kill.

Co-authored-by: Orca <help@stably.ai>

* fix(daemon): make the occupancy residual total

The module's contract is that 'unknown' is where every unanswerable question
lands, but a throwing dependency escaped instead — routing a failed
observation into the launch path rather than onto the safe residual. Latent
today because both real implementations swallow their own failures, which is
exactly the kind of thing that stops being true during a refactor.

Co-authored-by: Orca <help@stably.ai>

* fix(daemon): stop counting the daemon's own probe PTYs as hosted work

Round-four review found the evidence proving the wrong thing. The filter
excluded the daemon's plain subprocesses on the grounds that only a PTY child
is a session leader — but the daemon opens PTYs for its own health probe and
conpty warmup, and forkpty makes those session leaders too. The comment's own
premise refuted its exclusion list. A daemon hosting zero user terminals could
be held on the strength of its stuck probe child, and since the held daemon
also had no sessions, dropping our authenticated pair let it retire and take
the very state we were protecting. Exclude them by exact command, on both
platforms.

Two more from the same review:

- pty-spawn-unhealthy is only reachable after a successful hello, so that
  daemon is adoptable. It was routed to 'held' — which never adopts — purely
  because the count had come from the process table. Check it first; hold now
  requires an unreachable daemon.
- The grace loop rescanned the process table every pass, though it is waiting
  for IPC and the table cannot change its answer in five seconds. Ask the
  daemon during the wait and read the table once, after. raiseOccupancy-
  WithProcessEvidence makes that split explicit, and can only ever raise.

Bound the wait by wall clock too: the retry count alone never bounded it, and
startup fails open at 60s by abandoning the daemon provider outright, which
would trade a wedged daemon for no daemon and a Restart that throws.

Co-authored-by: Orca <help@stably.ai>

* fix(daemon): hold a hello-rejected daemon instead of adopting one that refused us

Found while reviewing why a mutation looked equivalent. Gating the hold on
health === 'unreachable' left 'rejected' — a daemon that answered and refused
the handshake — falling through to preserveDaemon(), whose adoption opens the
very hello it just refused. That throws, and the throw costs the app its
daemon and its Restart remedy. Killing it instead is no better: it can still
be hosting running agents.

Neither of those daemons can complete a handshake, so neither may be adopted,
and both must be held. Gate on that rather than on one of its two causes.

Adds regression tests for the round-four fixes: the self-spawned probe
exclusion is exact-match on both platforms, an adoptable pty-spawn-unhealthy
daemon is never routed to a mode that cannot adopt, evidence can only raise a
verdict, and the grace budget stays under the startup fail-open cap.

Co-authored-by: Orca <help@stably.ai>

* fix(daemon): decide adoptability from what the daemon can do now, in one place

Round five found the same question — can this daemon complete a hello right
now? — answered in three places with three different conclusions, because
'held' had been added as a fifth branch rather than as the classification the
other branches route through. Two of those answers were wrong, and both ended
in no daemon at all, which is the outcome 'held' exists to prevent.

- A daemon whose adoption hello had just failed was handed back tagged
  'degraded-new-pty-fallback'. Init skips the lease only for 'held', so it
  reopened the same connection, threw, and aborted startup — leaving the
  agents alive but unreachable and Manage Sessions -> Restart throwing.
- The pty-spawn-unhealthy arm ran first and claimed a successful hello proved
  adoptability, but that reading is from before the grace window. A daemon
  that answered at t=0 and went silent through thirty seconds of retries took
  that arm and threw the same way. Ask whether it is answering now, first.

The budget was a comment with a Date.now() beside it: one occupancy probe
could cost 50s, because the client's default is a 5s hello per connection
step plus a 30s request timeout. Bound the probe explicitly, start the clock
before the first one, and size the window so the whole path — health check,
pid verification, loop overshoot and process-table read — fits under the
startup fail-open with room to spare.

Also folds the four sibling branches onto the same occupancy resolution.
getAliveDaemonSessionCount was byte-identical to countLiveSessionsOverIpc, so
one concept had two implementations and only one of them had been fixed.

The Windows probe exclusions could never match: the warmup spawns COMSPEC, an
absolute path, against an exact-equality test on 'cmd.exe /c exit'. Compare
the program by basename and keep the argv tail exact.

Co-authored-by: Orca <help@stably.ai>

* fix(daemon): stop the local fallback answering for sessions it does not own

Closing a held daemon's pane reported success while the agent kept running.
Unrouted ids resolve to the in-process fallback, whose shutdown returns
silently for an id it has never heard of and whose write and resize are
no-ops — and while a daemon is held nothing ever enumerates its sessions, so
every one of them is unrouted. The pane vanished, the orphan outlived the
app, and typing into a stuck terminal disappeared without a word.

Only attach was fenced against that route. Extend the same rule to the
operations that change or feed a session: route to the fallback only when it
genuinely owns the pty, and otherwise say the session cannot be reached.

The error type is load-bearing. pty:kill treats "Session not found" as proof
the pty is already gone and synthesizes an exit, so reusing that error would
have reproduced the lie one layer down. TerminalSessionOwnerUnverifiedError
means "still there, we cannot reach its host", which is reported as a failed
close and keeps ownership for a retry.

Co-authored-by: Orca <help@stably.ai>

* test(daemon): pin that a held session cannot be closed by a provider that never had it

Covers the held-daemon routing fence, including the coupling that is
invisible from the routing file: the thrown error must not match pty:kill's
already-gone predicate, or the close is swallowed into a synthesized exit and
the orphan is hidden again. A rename would otherwise reintroduce the bug
silently.

Co-authored-by: Orca <help@stably.ai>

* fix(daemon): bound the POSIX evidence read and re-verify the pid it describes

Round-six follow-ups, none destructive.

The process-table read had a deadline on Windows but not on POSIX, where the
shared reader's ps timeout does not cover queueing behind an in-flight scan —
so the one step that runs after the grace window could still outlast it.

The pid handed to that read was verified before the grace window, which is
long enough for the daemon to die and its pid to be recycled onto a shell
with children. Verify it where it is used instead, and only when there is
still something to raise: the common case now skips the identity probe
altogether, which also takes a few seconds off the worst-case launch.

Two renderer call sites killed PTYs without handling rejection. That was
harmless while an unreachable session was answered by a silent no-op; now
that it honestly rejects, pane teardown and repo removal would log an
unhandled rejection every time — exactly when the daemon is already sick.

Deliberately not taken from that review: giving the endpoint-occupied catch
the same held fallback as the failed-health path. That path arrives with
occupancy unknown or empty, so holding there would swallow a real launch
failure to protect nothing.

Splits the repro script, which had grown past the line limit, into the
sequence it proves and the two things it proves it with: process-table
inspection, and the static assertions on the launcher's hold decision.

Co-authored-by: Orca <help@stably.ai>

* test(daemon): hold the classification budget to the whole path, not one term

The previous assertion compared the grace window to the fail-open cap, which
passed while the real path ran to roughly twice the cap — a single probe cost
50s against a 5s assumption, and the terms on either side of the loop were
never counted at all.

Sum the declared budgets instead: health check, grace window, the one probe
that always runs past a ceiling tested at loop entry, and the evidence read
on both platforms. Raising any of them now has to face this, and the spare
time the kill ladder and fork still need afterwards is stated rather than
assumed.

Lives outside the launcher's own spec because that file mocks daemon-health,
which would shadow the constants being held to account.

Co-authored-by: Orca <help@stably.ai>

* fix(daemon): do not read a stranded login wrapper as a hosted terminal

Raised by a colleague's handoff on the same-day reports. macOS wraps every
terminal in /usr/bin/login for TCC attribution, and #13764 shows the wrapper
can outlive the shell it wrapped — leaving a session leader that hosts
nothing. One affected host had accumulated enough of them to reach swap
pressure.

That is exactly the evidence this change treats as proof of live work, so a
daemon whose sessions had all ended would have been held indefinitely on the
strength of the corpses, on precisely the hosts where the problem is worst.
Same class as the daemon's own probe PTYs: a session leader is necessary
evidence, not sufficient. A wrapper still doing its job has the shell it
exec'd beneath it.

Co-authored-by: Orca <help@stably.ai>

* test(daemon): pin the daemon's PTY spawn sites so the exclusion list cannot silently rot

The ownership evidence discounts the PTYs the daemon opens for itself, and that
list is only safe while it is complete — a self-spawned PTY nobody excluded
reads as user work and holds a daemon that owns nothing. The list grew one
reviewer at a time, which is the wrong mechanism for a correctness invariant.

Pin the input rather than the list. The daemon has exactly three PTY spawn
sites: the user's terminal, the spawn health probe, and the Windows conpty
warmup. A fourth now fails this test until someone decides which side it
belongs on.

Co-authored-by: Orca <help@stably.ai>

* fix(daemon): stop the evidence going blind on the host it exists to protect

Readiness review, section 04. Every agent pane already drives the shared
process-table reader on its own cadence, so the uncached read queues behind
them — and the host with the most agents to lose is the one likeliest to blow
the deadline on queueing alone. Both attempts return unknown and the daemon is
killed anyway, which is the original bug wearing the fix as a costume.

Fall back to the TTL-cached table, which on that host is always warm for
exactly the reason the uncached read is always queued. A table a few hundred
milliseconds old still answers whether this daemon has children, and the
failure directions are not symmetric: over-holding costs one degraded launch
that self-heals, under-counting ends running agents.

The same review found the launch budget still overran the 60s startup
fail-open — by ~7s on Windows — and that the test guarding it under-counted
the path it was written to bound, for the second time. It omitted the identity
probe before the evidence read and the endpoint check that ends the grace
loop. Both are now summed, the headroom requirement covers the kill ladder and
fork that follow a replace verdict, and the grace window and Windows probe
deadline are sized to fit.

Not taken from that review: reusing the verified pid inside killStaleDaemon to
drop the duplicate probe. That second verification is what fences the signal to
this incarnation, and a seconds-old result is exactly the pid-reuse hazard it
exists to prevent.

Co-authored-by: Orca <help@stably.ai>

* fix(daemon): confirm emptiness before it authorizes a kill

Readiness review, sections 02/03/05/06 — no P0 or P1 in any of them. These are
the P2s worth taking.

The important one: on macOS a terminal contributes exactly one session leader,
the login wrapper, because the shell it forks is in the same session and shows
S+ rather than Ss. I had assumed the shell counted too. It does not — so a
wrapper that looks childless in a single snapshot makes its whole terminal
invisible, and that snapshot cannot tell a wrapper whose shell has gone from
one whose shell has not yet appeared. Emptiness is the answer that authorizes a
kill, so it now costs a second read; 'owns-live-ptys' still needs none. The
fixtures said Ss where a real shell says S+, which is why the tests never
noticed.

Also from that review: a fabricated row was cast to ProcessTableRow to reuse a
command-only predicate, which is sound only while that predicate reads nothing
else — narrowed to Pick<'command'> so the compiler keeps it honest. The
pty:signal listener relied on its provider staying async to convert a routing
refusal into a rejection; it is an ipcMain.on listener with nothing above it, so
it now catches synchronously too. And the repro's teardown signalled remembered
pids a minute after phase 1 waited for them to die — re-verify by tag first,
since signalling a recycled pid is the mistake the script exists to study.

Co-authored-by: Orca <help@stably.ai>

* fix(daemon): stop guessing at Windows occupancy instead of guessing better

The readiness review found the Windows branch reading a wedged daemon's
orphaned conpty hosts as live terminals. ClosePseudoConsole only runs on the
daemon's own JS thread, so a daemon too wedged to answer is also too wedged to
reap them, and they accumulate exactly when this code runs. A wedged, empty
Windows daemon would then be held forever — #8689 re-opened, and a regression
from main rather than a missing protection.

The tempting fix is another exclusion. That would be the sixth revision to what
counts as a live PTY, each one added because a reviewer found something that
looks like a session and is not, and each one trading safety for availability
in a fix whose entire purpose is the opposite trade. The list is the problem.

POSIX has a real signal: forkpty makes a hosted terminal a session leader, which
nothing the daemon forks for itself ever is. Windows has no equivalent, so its
branch could only ever count descendants and subtract guesses. Delete it and
answer 'unknown' — Windows keeps exactly the behaviour it has on main, and the
protection is claimed only where it can be justified.

Also stops a blind confirming read from upgrading an unconfirmed emptiness into
a verdict. Emptiness is what authorizes a kill; a read that saw nothing
corroborates nothing.

Co-authored-by: Orca <help@stably.ai>

* refactor(daemon): clear the debris the Windows removal left behind

Behaviour-preserving. Windows now abstains once at the entry point rather than
twice inside a retry loop that had nothing to retry, which also retires the
platform check further down that could no longer be false. The self-spawn
matcher kept backslash splitting and .exe stripping for a branch that no longer
exists, and the launcher's grace loop repeated its own IPC call and stacked two
explanations above the wrong statement.

Co-authored-by: Orca <help@stably.ai>

* fix(daemon): stop the budget cut spending Windows' only protection

Round seven found the one thing this PR must never do: kill a session that main
would have kept.

Main's grace loop probed with a non-shared 5s connect budget, so a wedged
daemon got roughly a minute to come back. Bounding the probes and adding a
wall clock cut that to about twelve seconds — a good trade on POSIX, where a
daemon that outlasts the window is still protected by process-table evidence,
and a bad one on Windows, which has no such evidence and now has nothing else.
A Windows daemon wedged for half a minute while hosting agents was adopted by
main and is killed by this branch. The fail-open cannot rescue it either:
ensureRunning() is not abortable, so the launcher runs to completion.

Size the window per platform instead, against what each actually spends:
Windows pays no evidence read and no identity probe to feed one, so it can
afford far more grace, and grace is worth more where it is the only thing
there. Both numbers come from the budget test rather than taste.

Three more from the same review:

- The evidence read applied its deadline twice, once to the fresh table and
  again to the cached fallback, so an attempt could cost double what the launch
  budget was told. Share one deadline across both.
- Two tests described protection the code no longer delivers: one asserted ~60s
  of grace the wall clock had already retired, the other passed only because its
  mocked probes are free and would fail against real ones. Say what the code
  actually promises, and freeze the clock where the point is retry depth.
- The self-spawned PTY inventory promised more than it inspects. It sees direct
  node-pty calls in one directory; the macOS login-session probe reaches a PTY
  through expect(1) and is caught by the stranded-wrapper filter instead. Scope
  the claim, since that indirection is the shape the next escape will take.

Co-authored-by: Orca <help@stably.ai>

* refactor(daemon): spend the launch budget against a clock instead of a sum

Round eight found the fourth term missing from the hand-written budget — the
launcher's own adoption connect, which runs before the health check on the
non-shared five-second path. The three before it were an identity probe, an
endpoint probe, and an evidence deadline applied twice. Every one of them
passed the test meant to catch exactly that, because the test could only check
the terms someone had remembered to add.

So stop summing. The classification now runs against a deadline and stops when
it expires, and the test asserts only that the deadline leaves room for the
kill ladder and the fork that follow it. A budget that has to be remembered is
a budget that will be wrong; this one cannot be, because nothing has to be
counted.

That also retires the platform-split grace window, which existed to hand
Windows more of a sum nobody could total correctly.

The same review found the regression it was compensating for was never the
window. main gave each probe up to fifty seconds — five per connection step,
thirty for the request — where this branch gave eight for both together. A
daemon whose handshake needs more than four seconds therefore answered none of
the probes, however many it got, and on Windows nothing else can speak for it.
Splitting the two budgets fixes the case the window never could: connecting
stays tight, because a daemon that cannot handshake is wedged and worth
re-asking cheaply, while a daemon that did handshake is demonstrably alive and
its count is worth waiting for.

Also stops a Date.now spy leaking out of a failed test and freezing the clock
for the rest of the file.

Co-authored-by: Orca <help@stably.ai>

* fix(daemon): make the classification clock actually bound the work it names

The clock introduced in the previous commit gated the probes but not the two
steps after them. The identity re-check and the process-table read ran on their
own deadlines, outside the ceiling, so the launcher could still spend its whole
budget on probes and then take another ten seconds — the same overrun the sum
used to produce, arrived at from the other end.

Hold that time back from every probe instead. A probe is only started when the
clock can still fund a handshake after the reserve, and its budget is what
remains minus the reserve, so no probe can eat it however long the daemon takes
to answer. Worst case is now the ceiling by construction rather than by
addition.

The reserve has a test asserting it is large enough for what it covers, which
failed on its first run and caught that ten seconds was not.

Co-authored-by: Orca <help@stably.ai>

* fix(daemon): ask the wedged daemon a question it can actually answer

Round nine found the re-verification was stricter than the check that triaged
the daemon onto this path. The launcher gets here because a three-second health
check — one socket, one hello — timed out. It then re-asked with two sockets
and two hellos inside four shared seconds, and repeated that identical question
up to twelve times. A daemon that consistently needs five seconds fails every
one of them, so the retries could only ever agree with the check that sent it
here. main re-asked with five seconds per connection step and thirty for the
answer, and kept the sessions this branch destroyed.

Retries and patience solve different problems. Keep the cheap probes, which
catch a daemon that recovers on its own, then spend what is left of the clock
on one tolerant ask — the only question that can disagree with the triage. It
is skipped when the endpoint is provably gone, since a cold start arrives here
too and has nothing to wait for.

Two more from the same review. The clock claimed to cover the launcher's own
adoption connect and started after it, so the fourth term that went missing
from the sum was still uncounted; it now starts above that connect and bounds
it. And Windows was holding back twelve seconds for an identity check and a
process-table read it never performs — the reserve is zero where the steps it
reserves for do not run.

Retuned the ceiling to leave the kill ladder and fork real margin rather than
half a second, with the packaged-Windows host copy named as what the margin is
for.

Co-authored-by: Orca <help@stably.ai>

* fix(daemon): ask the tolerant question while the clock can still fund the answer

The launcher asked the cheap question first and the patient one last. That is
backwards. This path is only reached because a 3s health check timed out, so
every 4s probe re-asks on a stricter budget than the one that triaged the daemon
here — it can only ever agree. The one ask that could disagree ran last, by which
point the clock could fund its handshake but not its answer, and a daemon that
answered in 12s was read as dead and replaced along with its agents.

Three changes, one idea: never make an ask you cannot afford to hear out.

- The patient ask goes first, with every millisecond the answer does not need.
- Cheap retries follow it, and stop once the clock cannot fund both halves.
  Funded to knock but not to listen is not an ask.
- The adoption connect is capped. It is not a classification step — it acquires
  a lease preserveDaemon() re-establishes anyway — but on a daemon that accepts
  the socket and never completes hello it would spend the entire classification
  budget, leaving nothing for the probes that protect that daemon's sessions.

The test double now forwards the connect budget instead of dropping it, so what
the launcher was willing to wait for is observable at all.

* fix(daemon): stop three follow-ups rotting where review already found them

A truncated paste now says so. A routing throw partway through a paste was not a
PtyWriteUnavailableError, so no pty:writeUnavailable reached the renderer and the
pane never re-attached — the remaining chunks simply vanished with nothing to
attribute the gap to. It stays deliberately distinct from SessionNotFoundError,
which isPtyAlreadyGoneError matches and synthesizes into an exit the session never
had; a test now pins both directions so neither drifts.

The launch-budget spec no longer fails on Windows. The evidence reserve is zero
there by design — neither guarded step runs without a session-leader signal — but
the assertion demanded eleven seconds of it unconditionally, so `pnpm test` broke
on any Windows dev machine. PR CI never saw it: only the WSL boundary spec runs on
windows-2022. The identity ceiling it reserves against is imported now instead of
being a 3_000 someone would have to remember to change.

And the grace-retry comment described the design that preceded the clock: ~5s
probes, ~60s of grace, a number worth raising. The clock binds first and usually
permits far fewer, so raising it alone buys nothing.

* fix(daemon): stop killing a daemon we merely failed to observe

Ten review rounds each found a different band where this branch replaced a daemon
that origin/main would have kept, and every fix bought a new one. The reason is
arithmetic, not carelessness: matching the old tolerance for a single probe costs
about 25s, the classification clock has 15s to give, and funding the difference
puts startup past the 60s fail-open once the kill ladder and the fork are paid.
No assignment of those numbers is safe.

So the residual stops being lethal. daemon-occupancy.ts always said it — "'unknown'
is the residual, and it is not permission" — while daemon-init.ts fell through from
unknown to killStaleDaemon. That fall-through is what made every millisecond of
budget a correctness parameter. Now an unclassifiable daemon is held in degraded
mode: existing terminals keep working, fresh ones run locally, and being wrong
costs a degraded session instead of somebody's agent.

Two exclusions, both about never holding something unrecoverable. A proven-dead
endpoint is a cold start or a corpse, and holding one would hand every first launch
a provider pointed at no daemon. 'rejected' answered and refused, so it can never be
adopted and its sessions can never be reattached.

The cost is real and deliberate: a wedged-but-empty daemon is no longer replaced at
launch, so #8689 degrades to "restart it from Manage Sessions". Which only works if
the user knows — and degraded mode was computed, plumbed through preload, and read
by nothing. It now renders where the Restart button already lives, and says what it
actually costs: new terminals close when you quit, and restarting ends whatever the
host is still holding.

Rejected on the way here: a background reclassifier (a permanently wedged daemon
never answers, and recovery is already handled by degraded-daemon-fresh-spawn-
routing.ts), letting accumulated process-table reads license a kill (its errors are
systematic, so more samples agree rather than converge), and sidelining the daemon
onto a renamed socket (verified working at the syscall level, then abandoned: the
daemon's own endpoint-ownership watch reads the moved entry as lost and retires
itself, precisely when it recovers).

* test(daemon): pin the hold that no longer depends on a clock

The repro proved the launcher holds a daemon it can see is occupied. The protection
that now matters most is the one for a daemon it cannot see at all, and nothing
asserted it. Adds the unknown-hold to the static assertions: that it exists, that it
excludes 'rejected' and a proven-dead endpoint, and that every killStaleDaemon call
site in the file is downstream of it.

Verified by mutation — dropping either exclusion fails the assertion.

* fix(daemon): repair what round eleven found, including a fix that did nothing

The patient ask was not patient. With twelve seconds reserved for the evidence
read, `max(CONNECT, probeBudget - REQUEST)` resolved to exactly CONNECT — the
tolerant ask got the cheap ask's four seconds, and the grace loop's gate needed
19s of a budget that only ever held 11s, so it never ran at all. Both shipped
green because mocked probes consume no wall clock, so the budget never binds in
a test. Two arithmetic tests now compute against realistic elapsed time, which is
where the defect actually lived.

The reservation was backwards anyway. Evidence can only raise 'unknown' to an
uncounted 'occupied', and both now hold the daemon, so the read changes a log
line and nothing else — while starving the one probe whose counted answer still
reaches preserveDaemon() and full daemon mode. It is opportunistic now: if the
probes spent the clock, it is skipped and the verdict stays 'unknown', which
holds exactly as an evidence-raised 'occupied' would have.

Recovery was a one-way flip on a two-way condition. Once a health check promoted
fresh spawns back to the daemon, nothing ever demoted them, so a daemon that
wedged again cost a hello timeout plus a full launcher re-classification for
every new terminal, for the rest of the session. A failed spawn now routes back
and re-arms the cooldown. The class had no tests at all; it has six.

The notice was wrong twice. It claimed terminals already open keep working — in
held mode discovery runs over the same IPC the daemon is failing, so its sessions
are never routed and attach refuses rather than answering on its behalf. They are
running, but unreachable. And it named half the cost of Restart: runRestartDaemon
shuts down the local fallback sessions too, so the terminals it had just called
safe die as well. Also fixed: the amber-on-amber body text failed AA at 3.94:1,
the keys were in a namespace no sibling uses, and the scale did not match the
notice it renders beside.

The banner armed a destructive button with state that never refreshed. It now
refetches on focus, like the sibling notice that solved this first.

Also: the launch mode type in the test harness omitted 'held', so no test could
describe the launch the hold produces; held mode's routing through the degraded
provider was unpinned, and deleting it left every test green; and pty:signal kept
a try/catch for a synchronous throw that async routing cannot produce.

* fix(daemon): stop offering a remedy that cannot work, and say why two modules stay

The degraded warning told the user to restart the daemon. When something other
than an Orca daemon holds the endpoint, restarting clears nothing: killStaleDaemon
only kills a process whose identity matches the pid record, and a foreign holder
matches none, so the next launch is identical. The message now says the daemon is
unreachable rather than asserting what it owns, and names the second remedy. The
notice says "usually clears this" for the same reason.

That case is also now written down as a known residual: an endpoint that accepts
connections but never speaks the protocol reads as an incumbent on every launch,
so it stays degraded with no auto-recovery, where before it was replaced.

The rest is comments, because three separate deletion proposals landed on this
code in one review cycle and each was a regression. The evidence read is not
redundant with the unknown hold: the occupied branch has no proven-dead check and
the unknown hold does, so it is the only thing between a kill and a daemon whose
socket entry vanished while it still hosts agents. Its children scan is not
redundant with pid verification either — a verified-live pid alone would also hold
a childless daemon, which is the one #8689 case still safe to replace. And grace
retries are worth more since 'unknown' stopped killing, not less: a counted
'occupied' reaches full adoption where the alternative is a degraded hold.

Each now says which case dies if it is removed. Reviewers reaching for the delete
key three times in a row is the code failing to explain itself, not excess.

* docs(daemon): record what a budget raise would owe before it is safe

The classification budget serves two verdicts with opposite time-costs. Reaching
"don't kill" slowly is free — the daemon survives however long it took. Reaching
'empty' slowly is not, because the kill ladder and the fork still have to fit
before the fail-open. At 34s that case cannot arise; at 44s it can, and an overrun
there is the worst branch on offer: daemon killed, replacement forked then
discarded, no provider installed, Restart broken.

So the raise is not a number change, it is a number change plus a guard: hold
rather than replace when the headroom left cannot fund the ladder and the fork.
Safe precisely because that path has proven the daemon empty, so holding costs no
agents. Written down next to the warning against tuning the budget, because the
next person to want a bigger number will read that warning and need this one.

Also recorded: the launcher closure has no access to the startup abort signal, so
the cheap version of that guard is not available without threading it through the
spawner.

* fix(daemon): delete a retry loop that could never run, at any budget

Round twelve proved the loop unreachable by algebra rather than by tracing:

  remaining = B - E - max(CONNECT, (B - E) - REQUEST) = REQUEST,
  whenever B - E > CONNECT + REQUEST

The patient connect takes every millisecond the answer does not need, so what
survives it is always exactly OCCUPANCY_REQUEST_BUDGET_MS — and the gate wanted
CONNECT + REQUEST. That holds for every ceiling, which also settles the raise I
had been holding open: at 44s the remainder is still exactly the request budget,
so ten more seconds of worst-case startup would have funded zero retries. Funding
one honestly needs ~71s against a 60s fail-open.

So WEDGED_DAEMON_GRACE_RETRIES = 11 documented patience the launcher did not have,
and no number could give it. Deleted, with the derivation left where the loop was
so the next person does not re-derive it from scratch.

Little is lost. A 4s retry cannot reach a daemon needing longer than 4s to answer,
which is the entire wedge population, while the one patient ask waits ~12s. The
only case retries caught and this does not is a daemon recovering within seconds
of being asked — and DegradedDaemonFreshSpawnRouter.recover() already returns it to
full daemon service on the next spawn, off the startup clock.

The budget tests could not have caught any of this: they recompute the expression
from imported constants and never execute the launcher, so collapsing the patient
connect back to the cheap constant — the round-eleven defect exactly — left all
1475 green. There is now a test that watches the launcher spend it, verified by
mutation, and the arithmetic ones say plainly that they are not the guard.

Also corrected: the evidence-read comment claimed the read only affects a log line.
It decides the verdict wherever the unknown hold declines to — it has neither the
proven-dead check nor the rejected check — so it is what holds a daemon whose socket
vanished, and what holds a hello-rejected daemon still hosting agents.

* docs(daemon): cost the deferred guard honestly, and say why it is unreachable

Two corrections to the note, both of which change what it tells the next person.

The reason the guard's case cannot arise at 34s is structural, not a lucky
margin: an `empty` verdict means the daemon answered, so it resolved fast by
construction, and proven-dead means nothing is listening, so the probe and the
ladder both short-circuit. The path that actually spends the budget is the wedge
that never answers — and that one now ends in a hold, paying neither the ladder
nor the fork. Long path and expensive tail are disjoint. Raising the budget is
precisely what re-couples them, by extending how late an `empty` may arrive.

And the guard was costed as a signature change through DaemonSpawner, which is
wrong. createOutOfProcessLauncher is a factory called where `signal` is already in
scope; a third parameter closed over there leaves the launcher's call signature
untouched. Overstating the price invites skipping the guard rather than paying it.

Also recorded: two terms this budget does not bound at all — the healthy branch,
which never consults the clock and still ends in a cleanup and a fork, and the
unbounded daemon-host copy on packaged Windows.

* docs(daemon): correct an overstated claim about the deleted retry loop

The deletion was justified as "the loop could never run, at any budget." That is
true of three wedge shapes and false of a fourth: a connect that fails fast leaves
the budget nearly whole, and while refused and missing endpoints are caught by the
proven-dead guard, the EPERM/EMFILE class reads 'unknown' and would have passed
the gate.

The deletion still stands — retrying an fd-exhausted or permission-denied connect
fails identically the second time, and recover() restores full daemon service on
the next spawn once the condition clears — but a comment that overstates its own
reach is how the next person concludes the reasoning was never checked.

* test(daemon): restore two guards a range deletion swallowed

Deleting the obsolete grace-loop test took out the two tests either side of it —
the ones pinning that the hold declines a proven-dead endpoint and a rejected
daemon. Both exclusions went unpinned in the same commit that removed the loop,
and the suite stayed green, because nothing else covers either path.

Found by mutation rather than by reading: removing `health !== 'rejected'` from
the hold left all 1472 passing. The pre-existing rejected test does not cover it —
its second client answers listSessions, so occupancy resolves to 'empty' and the
replace path is reached without the exclusion ever being consulted. The restored
test keeps the daemon unreachable so the verdict stays 'unknown', which is the
only state where the exclusion decides anything.

Also pins the evidence gate, the other survivor: the threshold must cover an
identity ps plus two ownership probes, or a read started at the last moment the
gate allows finishes past the ceiling the kill ladder and fork are sized against.

Mutation results now: patient connect collapsed -> caught; evidence gate -> caught;
hold removed -> caught; proven-dead exclusion -> caught; rejected exclusion ->
caught; fresh-spawn revert -> caught.

* fix(settings): make the degraded copy the copy users actually see

Two user-facing fixes were no-ops. translate() resolves from en.json, and the
catalog only ever gained the string it was first synced with — the sync script adds
missing keys and never updates changed defaults, which the extraction gate reports
as "inline defaults differ" and then passes anyway. So editing the inline default
changed the source and nothing else. Caught by rendering the component and reading
what came out, not by reading the diff.

What was stale in the catalog, and is now corrected there:

The notice promised the panes reconnect on their own once the host recovers. They
do not. TerminalErrorToast already tells the user "Reopen this pane to retry",
because nothing re-attaches a pane whose owner could not be verified — the session
is left untouched, which is the point, but recovery is a user action. Third claim
of mine in this PR that was stronger than the code.

And "Restarting the host clears this" still overstated the foreign-endpoint case,
where killStaleDaemon matches no pid record and clears nothing.

Also removes the components.settings.DaemonDegradedNotice.* namespace, left behind
when the keys were renamed to the auto.* convention every sibling uses. It was dead
weight carrying the oldest copy of all three strings.

* docs(daemon): 'held' no longer means what its type said it meant

The mode was introduced for a daemon that demonstrably owns live terminals and
cannot answer a handshake. It is now also what an unclassifiable daemon gets, where
the whole point is that we could not establish what it owns. A type whose comment
asserts the one fact the branch could not determine is the same overclaim this PR
has been correcting elsewhere.

Also un-exports ENDPOINT_PROBE_TIMEOUT_MS: it was widened for a test that no longer
references it, and nothing outside the module reads it.

* fix(daemon): stop a lost spawn reply from shadowing a live agent

`!mapped` was standing in for "this is a fresh spawn," and it is not the same
question. The mapping is only recorded after a reply arrives, so a spawn that names
a session and then loses its reply — the daemon created it, the answer timed out —
is indistinguishable from a genuinely new one. Demoting there sent the retry to the
fallback, which answers with a fresh local shell under the same id while the agent
keeps running on the daemon. The pane binds to the shell; the agent is orphaned.

That is the symptom this PR exists to remove, arriving through a door the PR opened
itself. Reachability today looks nil — the only sessionId-bearing spawn in ipc/pty.ts
carries attachOnly, which routes elsewhere — but the guard was unsound rather than
merely unused, and "no caller does that yet" is not a property anyone maintains.

Now an identified session pins to the provider that may already own it, and only an
anonymous spawn moves the shared route. Anonymous spawns are what demotion was for:
nothing can shadow them, and they are the ones paying a hello timeout plus a full
re-classification per terminal.

Found by GPT-5.6-Sol reviewing this file in isolation. Two tests added; reverting to
the old guard fails one.

* docs(daemon): name the three paths that can still kill a daemon

Adversarial review found all three; none is a regression against the pre-hold
behaviour, and none should be closed by weakening the evidence rules.

'unknown' plus a proven-dead endpoint still kills when process evidence cannot
answer. The probe proves the directory entry is gone, not the process — a socket
entry can vanish while the daemon still hosts agents. Evidence covers that on POSIX
because it runs for any 'unknown' rather than only a live endpoint, so the gap is
the blind cases: clock spent, pid unverifiable, ps unreadable. It is not reachable
on Windows at all, where a named pipe vanishes with its process, so a dead endpoint
there implies no agents to lose.

'unknown' plus 'rejected' still kills, and the reviewer is right that inability to
adopt is not inability to preserve — those agents keep running, unreachable. Killing
stays the choice because a daemon that can never be adopted and is never replaced
leaves the app permanently degraded with no route back, but that is a judgement, not
a proof, and it is now written as one.

And the verdict is not atomic with the kill: an 'empty' answer can go stale if
another instance creates a session first. Pre-existing, and narrowed rather than
widened here — the window now opens only after the daemon has reported zero sessions
itself.

* docs(daemon): record the TOCTOU fix that was built, tested and reverted

shutdownIfIdle is the right instrument and the daemon already implements it
atomically: sole authenticated client, nothing in flight, zero sessions, listener
closed before the acknowledgement. Asking it immediately before the kill closes the
window that a re-read of listSessions can only move.

It is not landing here. Gating every empty-verdict replacement on a new round trip
means every failure of that round trip has to mean hold, which trades a rare race
for a common failure mode and makes #8689 worse whenever the call is merely slow.
It also flipped two endpoint-identity tests from rejecting to resolving, and an
unexplained behaviour change is not something to merge at commit forty-two of a
change whose whole subject is unintended consequences.

Written down with the mechanism intact so the next person starts from a working
design rather than rediscovering it.

* fix(repro): restore the 91 lines a bad deletion took out of the repro

Removing readSourceConstant matched a docblock far earlier in the file and deleted
everything between, taking the imports and three functions with it. The script
still parsed, still linted, and still passed `node --check` — it failed only when
run, with `existsSync is not defined`, which is why it went unnoticed for two
commits. The only end-to-end proof in this PR had been dead that whole time.

Restored from before the deletion and the intended edit re-applied by itself. Now
runs green: phase 1 kills a wedged daemon with real agents attached and confirms
they die; phase 2 puts the identical wedge through the decision and shows the
daemon unsignalled, both agents alive, and everything back with sessions intact on
SIGCONT; phase 3 asserts the launcher holds — now including the unknown-hold, which
is the branch this PR turns on.

Lesson worth keeping: a syntax check is not a test. Running it is.

* fix(daemon): stop the emptiness confirmation reading the same snapshot twice

The second sample exists because one snapshot cannot tell a login(1) wrapper whose
shell has not appeared yet from a wrapper whose shell has gone. But when the fresh
read misses its deadline — the busy host this evidence exists to protect — both
attempts fell through to the same TTL-cached table. Two agreeing samples, one
observation, and the window being excluded is shorter than the cache.

The confirming read is now denied the cached fallback. If it cannot get a fresh
table it answers 'unknown', which holds. The first sample keeps the fallback,
because there the cache protects the answer worth protecting: a stale table still
shows that a daemon has children, and going blind there is what got them killed.

Also records three limits of this evidence that review surfaced and that no code
change should paper over — a reparented orphan outside the descendant tree, the
Windows abstention, and an argv match that cannot establish executable identity.
Each only fails to raise a verdict, so each costs a hold not taken rather than a
kill licensed.

* fix(daemon): stop discarding Linux terminals to solve a macOS problem

The stranded-login(1) exclusion ran on every POSIX host. Orca only wraps terminals
in login(1) for TCC attribution on macOS, so off darwin the pattern can only match
a user's own login — and one still prompting for credentials has no child yet,
which is precisely the shape the exclusion throws away. A Linux daemon hosting that
terminal read as childless, and a childless daemon is one nothing protects from the
endpoint-dead path. Now scoped to darwin, where the problem it solves lives.

And a count that is not a count is no longer read as emptiness. `counted > 0` maps
NaN, -1 and 1.5 to 'empty', which is the single verdict that licenses a kill; the
listSessions dep is injectable, so reaching it never required asking a daemon
anything. Non-integers and negatives now resolve to 'unknown'.

Both mutation-verified. Noting for whoever runs the next mutation pass: the first
attempt at the login mutation silently failed to apply because the pattern had been
reflowed by the formatter, and a mutation that does not apply looks exactly like a
test suite that caught it. Assert the pattern matched.

* docs(daemon): bound what the degraded owner check actually covers

Review confirmed the five destructive operations are guarded and that the error
taxonomy holds — TerminalSessionOwnerUnverifiedError cannot be reclassified as a
gone session anywhere in production, so no fake exit is ever synthesized from it.

It also found six methods that still route raw, and they are worth naming rather
than leaving for the next reader to rediscover: flow control, background state,
buffer clear, startup authority and the per-session queries. For an unresolved
daemon id those reach the fallback silently. None can destroy a session, which is
why they are not being changed at this point in this branch, but a buffer clear
that reports success while the daemon's history survives is a real lie.

Recorded with the reason not to fix them casually: acknowledgeDataEvent is invoked
directly from an ipcMain.on listener and setPtyBackgrounded from a synchronous
callback, so a throwing owner check added there without changing the call sites
turns a silent misroute into an escaping exception.

* fix(daemon): restore the demotion my shadow fix made unreachable

The review is right. Gating demotion on the absence of a sessionId assumed fresh
spawns are anonymous, and they are not: every production fresh spawn mints an id
before reaching the provider (ipc/pty.ts assigns spawnOptions.sessionId on all
three paths). So the branch only ever ran in tests, and a daemon that recovered and
wedged again kept every later terminal pointed at itself — each paying a hello
timeout plus a full launcher re-classification, each failing anyway. That is the
cost the demotion existed to remove, reintroduced while fixing something else.

The mistake was treating one condition as two questions. Pinning protects THIS id,
which the daemon may already have created before losing the reply, so a retry must
never be answered locally under the same name. Demoting protects the NEXT terminal,
which is a different session and cannot be shadowed by this one. They are
independent, and both now happen.

`attachOnly` is the honest discriminator for the second: an attach that names a
session never reaches this router, so anything arriving here without it is a fresh
terminal whatever id it carries.

Both halves mutation-verified separately — restoring the sessionId guard fails
three tests, dropping the pin fails two — including a test shaped like what
ipc/pty.ts actually sends, which is what the old test suite never had.

* fix(repro): stop the script doing the exact thing it exists to warn about

Both findings are right, and the first is pointed: this script demonstrates that
signalling a pid you have not re-verified can kill someone else's work, and its own
teardown did that twice.

The staged daemon is killed on purpose in phase 1. Once Node reaps that child its
pid is free for reuse, and teardown signalled the remembered number anyway; it now
signals only while the child object still reports no exit.

The markers were half-verified. Each proves its own identity by tag before being
signalled, but its session leader was killed on a number remembered from staging a
minute earlier — and the leader is the one pid here that can be recycled while its
child lives on under a new parent. The leader is now re-read from the live marker
and signalled only when the marker still claims it.

Second finding: isRealUserDaemon matched a hardcoded macOS userData path, so on
Linux a real daemon could never be recognised — the assertion that this run harmed
nothing was inert on the platform where nobody would notice. Now matched per
platform, with a loose fallback rather than a silent false.

Repro re-run end to end: all three phases pass, pre-existing daemons still running,
real userData daemon untouched.

* fix(daemon): only demote and only pin when the failure earns it

Both guards in the fresh-spawn catch were too broad, and narrowing them is the one
thing worth keeping from the restart-ownership branch.

Demotion fired on any spawn failure. A rejected cwd or a bad profile says nothing
about whether the daemon is reachable, and costing the whole session its daemon
persistence over one of those degrades terminals the daemon would have served fine.
It now requires the failure to look like an unreachable daemon.

Pinning fired on any failure too. The pin exists for a request that was sent and
whose answer was lost — that is the only shape that can hide a session the daemon
already created. A failure that never reached it created nothing, so pinning that
id would strand later attempts on a host holding nothing of theirs. It now requires
an error that could have been dispatched.

The distinction is sharper than the code it replaces: a hello timeout demotes but
does not pin, because a handshake that never completed cannot have created a
session. The old code pinned it anyway.

Test doubles now raise DaemonProtocolError rather than plain Errors, which is what
the client actually produces and what these predicates are written against. Both
narrowings mutation-verified.

* test(daemon): prove the recovery the degraded notice promises

The readiness review flagged one claim it could neither confirm nor refute: the
notice tells the user "reopening a pane retries, and works once it does", and
nothing pinned that. It mattered because held mode is exactly the case where no
route was ever recorded — discovery ran over the same IPC the daemon was failing —
so recovery cannot come from a cached route. It has to come from the next attach
re-inventorying a provider whose failure cooldown has expired.

It does. While wedged the resolver refuses rather than letting the fallback answer
with a fresh shell, and once the daemon answers again the same attach reattaches
the original session. Verified by mutation: with the daemon left wedged, the test
fails.

Fourth user-facing claim in this branch checked against the code rather than
assumed. The previous three were wrong.

---------

Co-authored-by: Orca <help@stably.ai>
2026-08-13 02:12:20 -07:00

323 lines
20 KiB
JSON

{
"name": "orca",
"version": "1.4.178-rc.2",
"description": "Next-gen IDE for parallel agentic development",
"homepage": "https://github.com/stablyai/orca",
"author": "stablyai",
"bin": {
"orca": "./out/cli/index.js",
"orca-dev": "./config/scripts/orca-dev.mjs"
},
"main": "./out/main/index.js",
"scripts": {
"format": "oxfmt --write .",
"lint": "oxlint && pnpm run audit:code-quality:native && pnpm run audit:code-quality:type-aware && pnpm run check:reliability-gates && pnpm run check:max-lines-ratchet && pnpm run verify:bundled-skill-guides && pnpm run verify:skill-bundle-manifest && pnpm run verify:localization-catalog && pnpm run verify:localization-extraction && pnpm run verify:localization-coverage",
"audit:code-quality": "pnpm run audit:code-quality:native && pnpm run audit:code-quality:type-aware && pnpm run audit:react-doctor",
"audit:code-quality:native": "oxlint --config config/oxlint-code-quality-native-plugins.json src config tests mobile --deny-warnings",
"audit:code-quality:type-aware": "oxlint --type-aware --config config/oxlint-code-quality-type-aware.json src config tests --deny-warnings",
"audit:react-doctor": "pnpm dlx react-doctor@0.9.1 . --yes --no-supply-chain --no-telemetry --blocking none",
"audit:dead-code": "pnpm dlx knip@5.88.1 --config config/knip.json",
"check:code-quality:changed": "node config/scripts/check-changed-code-quality.mjs",
"check:react-doctor:changed": "node config/scripts/check-react-doctor-changed.mjs",
"check:zustand-selector-fanout": "node config/scripts/zustand-selector-fanout-benchmark.mjs --check",
"doctor": "pnpm dlx react-doctor@0.9.1 . --no-telemetry",
"lint:react-doctor": "oxlint --config config/oxlint-react-doctor.json",
"lint:react-doctor:changed": "node config/scripts/lint-react-doctor-changed.mjs",
"prepare": "husky",
"test": "node config/scripts/ensure-native-runtime.mjs --runtime=node && vitest run --config config/vitest.config.ts",
"test:repro:remote-agent-session": "pnpm run build:cli && pnpm run build:electron-vite && node config/scripts/remote-agent-session-authority-repro.mjs",
"test:repro:daemon-replacement-live-agent-pty-preservation": "pnpm run build:electron-vite && node config/scripts/daemon-replacement-live-agent-pty-preservation-repro.mjs",
"check:reliability-gates": "node config/scripts/check-reliability-gates.mjs",
"check:max-lines-ratchet": "node config/scripts/check-max-lines-ratchet.mjs",
"check:feature-wall-assets": "node config/scripts/check-feature-wall-assets.mjs",
"generate:bundled-skill-guides": "node config/scripts/generate-bundled-skill-guides.mjs --write",
"verify:bundled-skill-guides": "node config/scripts/generate-bundled-skill-guides.mjs --check",
"generate:skill-bundle-manifest": "node config/scripts/generate-skill-bundle-manifest.mjs --write",
"verify:skill-bundle-manifest": "node config/scripts/generate-skill-bundle-manifest.mjs",
"verify:macos-entitlements": "node config/scripts/verify-macos-entitlements.mjs",
"vendor:feature-wall-assets": "node config/scripts/vendor-feature-wall-assets.mjs",
"tc:node": "pnpm run typecheck:node",
"tc:cli": "pnpm run typecheck:cli",
"tc:web": "pnpm run typecheck:web",
"tc": "pnpm run typecheck",
"typecheck:node": "tsc --noEmit -p config/tsconfig.node.json",
"typecheck:cli": "tsc --noEmit -p config/tsconfig.tc.cli.json",
"typecheck:web": "tsc --noEmit -p config/tsconfig.tc.web.json",
"typecheck": "tsc --noEmit -p config/tsconfig.node.json && tsc --noEmit -p config/tsconfig.tc.cli.json && tsc --noEmit -p config/tsconfig.tc.web.json",
"typecheck:tsc:node": "tsc --noEmit -p config/tsconfig.node.json --composite false",
"typecheck:tsc:cli": "tsc --noEmit -p config/tsconfig.cli.json --composite false",
"typecheck:tsc:web": "tsc --noEmit -p config/tsconfig.web.json --composite false",
"typecheck:tsc": "tsc --noEmit -p config/tsconfig.node.json --composite false && tsc --noEmit -p config/tsconfig.cli.json --composite false && tsc --noEmit -p config/tsconfig.web.json --composite false",
"ensure:electron-runtime": "node config/scripts/ensure-native-runtime.mjs --runtime=electron",
"start": "pnpm run ensure:electron-runtime && electron-vite preview",
"dev": "pnpm run ensure:electron-runtime && node config/scripts/run-electron-vite-dev.mjs",
"dev-stable-name": "pnpm run ensure:electron-runtime && node config/scripts/run-electron-vite-dev.mjs --stable-name",
"dev:web": "vite --config vite.web.config.ts --host 127.0.0.1",
"build:relay": "node config/scripts/build-relay.mjs",
"build:computer-macos": "node config/scripts/build-computer-macos.mjs",
"build:notification-status-macos": "node config/scripts/build-notification-status-macos.mjs",
"build:native": "node config/scripts/build-native-for-platform.mjs",
"smoke:computer": "node config/scripts/computer-use-smoke.mjs",
"verify:computer-native": "node config/scripts/verify-computer-native.mjs",
"verify:cli-bin": "node config/scripts/verify-cli-bin.mjs",
"verify:built-skills-cli": "node config/scripts/verify-skills-cli-runtime.cjs out",
"verify:localization-catalog": "node config/scripts/verify-localization-catalog.mjs",
"sync:localization-catalog": "node config/scripts/verify-localization-catalog.mjs --fix",
"verify:localization-extraction": "node config/scripts/verify-localization-extraction.mjs",
"verify:localization-coverage": "node config/scripts/audit-localization-coverage.mjs --check",
"audit:localization": "node config/scripts/audit-localization-coverage.mjs",
"build:cli": "tsc -p config/tsconfig.cli.json --outDir out --composite false --incremental false && node config/scripts/verify-cli-bin.mjs --fix-executable --fix-package-json && node config/scripts/install-dev-cli.mjs",
"test:repro:skills-cli-runtime": "pnpm run build:cli && pnpm run build:electron-vite && pnpm run verify:built-skills-cli",
"build:electron-vite": "node config/scripts/run-electron-vite-build.mjs",
"build:electron-vite:parallel": "node config/scripts/run-electron-vite-targets-in-parallel.mjs",
"build:web": "node config/scripts/run-vite-web-build.mjs && node config/scripts/verify-web-build.mjs",
"build:web-from-renderer": "node config/scripts/project-renderer-web-client.mjs && node config/scripts/verify-web-build.mjs",
"build:desktop": "pnpm run typecheck && pnpm run build:relay && pnpm run build:cli && pnpm run build:electron-vite && pnpm run verify:built-skills-cli && pnpm run build:web-from-renderer",
"build": "pnpm run build:desktop && pnpm run build:native",
"build:release": "pnpm run build:relay && pnpm run build:native && pnpm run verify:computer-native && pnpm run build:cli && pnpm run build:electron-vite && pnpm run verify:built-skills-cli && pnpm run build:web-from-renderer",
"postinstall": "node config/scripts/rebuild-native-deps.mjs",
"rebuild:electron": "node config/scripts/rebuild-native-deps.mjs",
"rebuild:node": "pnpm rebuild node-pty",
"build:unpack": "pnpm run build && pnpm run ensure:electron-runtime && electron-builder --config config/electron-builder.config.cjs --dir",
"build:win": "pnpm run build:desktop && pnpm run ensure:electron-runtime && electron-builder --config config/electron-builder.config.cjs --win",
"build:icons": "bash resources/icon-source/generate.sh",
"build:mac": "pnpm run build:desktop && pnpm run build:computer-macos && pnpm run build:notification-status-macos && pnpm run ensure:electron-runtime && node config/scripts/build-mac-local.mjs",
"build:mac:release": "node config/scripts/verify-macos-release-env.mjs && ORCA_MAC_RELEASE=1 pnpm run build:desktop && ORCA_MAC_RELEASE=1 pnpm run build:computer-macos && ORCA_MAC_RELEASE=1 pnpm run build:notification-status-macos && pnpm run ensure:electron-runtime && ORCA_MAC_RELEASE=1 electron-builder --config config/electron-builder.config.cjs --mac",
"build:linux": "pnpm run build:desktop && pnpm run ensure:electron-runtime && electron-builder --config config/electron-builder.config.cjs --linux AppImage deb",
"test:e2e": "pnpm run ensure:electron-runtime && npx playwright test --config tests/playwright.config.ts --project=electron-headless",
"test:e2e:multi-client-navigation": "node config/scripts/run-multi-client-navigation-e2e.mjs",
"test:e2e:floating-mobile-emulator": "pnpm run ensure:electron-runtime && npx playwright test tests/e2e/floating-mobile-emulator-tab.spec.ts --config tests/playwright.config.ts --project electron-headless --workers=1",
"test:e2e:terminal-rendering-golden": "pnpm run ensure:electron-runtime && npx playwright test tests/e2e/terminal-raw-emoji-table-scroll-restore.spec.ts tests/e2e/terminal-webgl-atlas-budget.spec.ts --grep @terminal-rendering-golden --config tests/playwright.config.ts --project electron-headless --workers=1",
"test:e2e:posix-profile-index-golden": "pnpm run ensure:electron-runtime && npx playwright test tests/e2e/golden-posix-fresh-startup.spec.ts tests/e2e/golden-posix-profile-index-fsync.spec.ts --grep @posix-profile-index-golden --config tests/playwright.config.ts --project electron-headless --workers=1",
"test:e2e:windows-fresh-startup-golden": "pnpm run ensure:electron-runtime && npx playwright test tests/e2e/golden-windows-fresh-startup.spec.ts --grep @windows-fresh-startup-golden --config tests/playwright.config.ts --project electron-headless --workers=1",
"test:e2e:terminal-rendering-release-evidence": "pnpm run ensure:electron-runtime && npx playwright test tests/e2e/terminal-opencode-emoji-table-rendering.spec.ts tests/e2e/terminal-long-table-scroll-restore.spec.ts --config tests/playwright.config.ts --project electron-headless --workers=2",
"test:e2e:terminal-perf": "pnpm run ensure:electron-runtime && npx playwright test tests/e2e/terminal-typing-latency.spec.ts tests/e2e/terminal-foreground-redraw-freeze.spec.ts tests/e2e/terminal-output-scheduler.spec.ts tests/e2e/terminal-hidden-tui-visual-restore.spec.ts tests/e2e/artificial-opencode-terminal-load.spec.ts --config tests/playwright.config.ts --project electron-headless --workers=2",
"test:e2e:terminal-perf:scale": "pnpm run ensure:electron-runtime && node config/scripts/run-terminal-scale-perf-e2e.mjs",
"test:e2e:terminal-perf:scale:report": "pnpm run ensure:electron-runtime && node config/scripts/run-terminal-scale-perf-report-gate.mjs",
"test:e2e:terminal-perf:check-report": "node config/scripts/check-terminal-perf-report-budgets.mjs",
"test:e2e:terminal-perf:summarize": "node config/scripts/summarize-terminal-perf-report.mjs",
"test:e2e:terminal-perf:html-report": "node config/scripts/generate-terminal-perf-html-report.mjs",
"test:e2e:ssh-docker-perf": "node config/scripts/run-ssh-docker-perf-e2e.mjs",
"test:e2e:ssh-docker-watcher-isolation": "node config/scripts/run-ssh-docker-watcher-isolation-e2e.mjs",
"test:e2e:ssh-docker-terminal-parking": "node config/scripts/run-ssh-docker-terminal-parking-e2e.mjs",
"test:e2e:ssh-docker-reconnect-pane-cardinality": "node config/scripts/run-ssh-docker-reconnect-pane-cardinality-e2e.mjs",
"test:e2e:ssh-docker-maxsessions-identity": "node config/scripts/run-ssh-docker-maxsessions-identity-e2e.mjs",
"test:e2e:ssh-docker-two-hosts-isolation": "node config/scripts/run-ssh-docker-two-hosts-isolation-e2e.mjs",
"test:e2e:nested-runtime-ssh": "node config/scripts/run-nested-runtime-ssh-e2e.mjs",
"test:e2e:source-control-scale": "pnpm run ensure:electron-runtime && npx playwright test tests/e2e/source-control-large-file-count.spec.ts --config tests/playwright.config.ts --project electron-headless --workers=1",
"win-update-e2e": "node tests/tools/win-update-e2e/run.mjs",
"win-crash-survival-e2e": "node tests/tools/win-crash-survival-e2e/run.mjs",
"test:e2e:ssh-codex-artifacts-repro": "node config/scripts/run-ssh-codex-artifacts-repro-e2e.mjs",
"test:e2e:headful": "pnpm run ensure:electron-runtime && npx playwright test --config tests/playwright.config.ts --project electron-headful",
"test:e2e:terminal-ime-native": "node config/scripts/run-terminal-ibus-hangul-e2e.mjs",
"test:e2e:computer": "vitest run --config tests/e2e/vitest.config.ts",
"bench:idle-cpu": "pnpm run ensure:electron-runtime && node config/scripts/run-idle-cpu-benchmark.mjs",
"bench:macos-computer-helper-owner-loss": "node config/scripts/macos-computer-helper-owner-loss-benchmark.mjs",
"bench:startup": "pnpm run ensure:electron-runtime && node tests/tools/benchmarks/startup-time-bench.mjs",
"bench:daemon-coldstart": "pnpm run ensure:electron-runtime && node tests/tools/benchmarks/daemon-coldstart-bench.mjs",
"bench:wsl-hook-relay-reattach": "pnpm run build:relay && node config/scripts/wsl-hook-relay-reattach-benchmark.mjs",
"bench:wsl-git-shell": "node config/scripts/wsl-git-shell-benchmark.mjs",
"bench:hang-watchdog-memory": "pnpm run ensure:electron-runtime && node config/scripts/hang-watchdog-memory-benchmark.mjs",
"bench:main-thread-jank": "pnpm run ensure:electron-runtime && node tests/tools/benchmarks/main-thread-jank-bench.mjs",
"bench:worktree-deletion": "node tests/tools/benchmarks/worktree-deletion-dev-bench.mjs",
"bench:zustand-selector-fanout": "node config/scripts/zustand-selector-fanout-benchmark.mjs",
"bench:worktree-refresh-churn": "node --disable-warning=MODULE_TYPELESS_PACKAGE_JSON config/scripts/worktree-refresh-churn-benchmark.mjs",
"bench:multi-workspace-typing": "pnpm run ensure:electron-runtime && node config/scripts/run-multi-workspace-typing-bench.mjs",
"bench:ai-vault-typing": "pnpm run ensure:electron-runtime && node config/scripts/run-ai-vault-typing-bench.mjs",
"bench:cold-park-reveal": "pnpm run ensure:electron-runtime && node tests/tools/benchmarks/terminal-cold-park-reveal-bench.mjs",
"bench:cold-park-resource": "pnpm run ensure:electron-runtime && node tests/tools/benchmarks/terminal-cold-park-resource-bench.mjs",
"bench:compare": "node config/scripts/compare-benchmark-artifacts.mjs",
"test:e2e:remote-bulk-open-freeze": "pnpm run ensure:electron-runtime && pnpm exec playwright test tests/e2e/remote-session-bulk-open-freeze-repro.spec.ts --config tests/playwright.config.ts --project electron-headless --workers=1",
"test:e2e:daemon-session-liveness": "pnpm run ensure:electron-runtime && pnpm exec playwright test --config tests/playwright.daemon-liveness.local.config.ts",
"test:e2e:ssh-docker-bulk-open-freeze": "node config/scripts/run-ssh-docker-bulk-open-freeze-e2e.mjs",
"repro:live-remote-bulk-open-freeze": "node config/scripts/live-remote-bulk-open-freeze-repro.mjs",
"repro:live-remote-realistic-freeze": "node config/scripts/live-remote-realistic-freeze-repro.mjs"
},
"dependencies": {
"@electron-toolkit/preload": "^3.0.2",
"@electron-toolkit/utils": "^4.0.0",
"@floating-ui/dom": "1.7.6",
"@linear/sdk": "^82.1.0",
"@parcel/watcher": "^2.5.6",
"@xterm/addon-serialize": "0.15.0-beta.287",
"@xterm/headless": "6.1.0-beta.287",
"agent-browser": "~0.27.0",
"electron-updater": "^6.8.9",
"i18next": "^26.3.1",
"jsonc-parser": "^3.3.1",
"node-pty": "^1.1.0",
"posthog-node": "^5.33.3",
"psl": "1.15.0",
"qrcode": "^1.5.4",
"react-i18next": "^17.0.8",
"serve-sim": "^0.1.40",
"sherpa-onnx": "1.12.37",
"ssh2": "^1.17.0",
"tweetnacl": "^1.0.3",
"ws": "^8.21.0",
"yaml": "^2.8.4",
"zod": "~4.4.3"
},
"devDependencies": {
"@dnd-kit/core": "^6.3.1",
"@dnd-kit/sortable": "^10.0.0",
"@electron-toolkit/tsconfig": "^2.0.0",
"@electron/rebuild": "^4.2.0",
"@monaco-editor/react": "^4.7.0",
"@playwright/test": "^1.59.1",
"@sanity/diff-match-patch": "^3.2.0",
"@stablyai/playwright-test": "^2.1.14",
"@tailwindcss/vite": "^4.2.4",
"@tanstack/react-virtual": "^3.13.24",
"@testing-library/jest-dom": "^6.9.1",
"@testing-library/react": "^16.3.2",
"@testing-library/user-event": "^14.6.1",
"@tiptap/extension-code-block-lowlight": "^3.22.5",
"@tiptap/extension-details": "^3.22.5",
"@tiptap/extension-image": "^3.22.5",
"@tiptap/extension-link": "^3.22.5",
"@tiptap/extension-mathematics": "3.22.5",
"@tiptap/extension-placeholder": "^3.22.5",
"@tiptap/extension-table": "3.22.4",
"@tiptap/extension-table-cell": "3.22.4",
"@tiptap/extension-table-header": "3.22.4",
"@tiptap/extension-table-row": "3.22.4",
"@tiptap/extension-task-item": "^3.22.5",
"@tiptap/extension-task-list": "^3.22.5",
"@tiptap/markdown": "^3.22.5",
"@tiptap/pm": "^3.22.5",
"@tiptap/react": "^3.22.5",
"@tiptap/starter-kit": "^3.22.5",
"@types/node": "^25.6.0",
"@types/qrcode": "^1.5.6",
"@types/react": "^19.2.17",
"@types/react-dom": "^19.2.3",
"@types/ssh2": "^1.15.5",
"@types/ws": "^8.18.1",
"@vitejs/plugin-react": "^5.2.0",
"@xterm/addon-fit": "0.12.0-beta.287",
"@xterm/addon-ligatures": "0.11.0-beta.287",
"@xterm/addon-search": "0.17.0-beta.287",
"@xterm/addon-unicode11": "0.10.0-beta.287",
"@xterm/addon-web-links": "0.13.0-beta.287",
"@xterm/addon-webgl": "0.20.0-beta.286",
"@xterm/xterm": "6.1.0-beta.287",
"class-variance-authority": "^0.7.1",
"clsx": "^2.1.1",
"cmdk": "^1.1.1",
"dompurify": "^3.4.13",
"electron": "^43.1.0",
"electron-builder": "^26.15.3",
"electron-builder-squirrel-windows": "^26.15.3",
"electron-vite": "^5.0.0",
"emoji-picker-react": "^4.19.1",
"emojibase-data": "17.0.0",
"happy-dom": "^20.9.0",
"html-to-image": "^1.11.13",
"husky": "^9.1.7",
"i18next-cli": "1.65.0",
"katex": "^0.16.45",
"lint-staged": "^16.4.0",
"lowlight": "^3.3.0",
"lucide-react": "^0.577.0",
"marked": "^17.0.1",
"mermaid": "^11.16.1",
"monaco-editor": "^0.55.1",
"oxfmt": "^0.52.0",
"oxlint": "^1.77.0",
"oxlint-plugin-react-doctor": "0.9.1",
"oxlint-tsgolint": "7.0.2001",
"pdfjs-dist": "^6.2.108",
"pngjs": "^7.0.0",
"radix-ui": "^1.6.2",
"react": "^19.2.7",
"react-colorful": "^5.7.0",
"react-dom": "^19.2.7",
"react-grab": "^0.1.33",
"react-markdown": "^10.1.0",
"react-refresh": "^0.18.0",
"rehype-highlight": "^7.0.2",
"rehype-katex": "^7.0.1",
"rehype-raw": "^7.0.0",
"rehype-sanitize": "^6.0.0",
"rehype-slug": "^6.0.0",
"remark-breaks": "^4.0.0",
"remark-frontmatter": "^5.0.0",
"remark-gfm": "^4.0.1",
"remark-math": "^6.0.0",
"remark-parse": "^11.0.0",
"shadcn": "^4.13.1",
"sonner": "^2.0.7",
"tailwind-merge": "^3.5.0",
"tailwindcss": "^4.2.4",
"tw-animate-css": "^1.4.0",
"typescript": "^7.0.2",
"typescript-api": "npm:typescript@6.0.3",
"unified": "^11.0.5",
"vite": "npm:rolldown-vite@7.3.1",
"vitest": "^4.1.5",
"vscode-oniguruma": "^2.0.1",
"vscode-textmate": "^9.3.2",
"zustand": "^5.0.14"
},
"optionalDependencies": {
"sherpa-onnx-darwin-arm64": "1.12.37",
"sherpa-onnx-darwin-x64": "1.12.37",
"sherpa-onnx-linux-arm64": "1.12.37",
"sherpa-onnx-linux-x64": "1.12.37",
"sherpa-onnx-win-x64": "1.12.37",
"windows-native-registry": "3.2.2"
},
"lint-staged": {
"*.{ts,tsx,js,jsx,mjs,mts,cts}": [
"oxlint",
"oxlint --config config/oxlint-react-doctor.json",
"oxfmt --write"
],
"*.{json,css}": [
"oxfmt --write"
]
},
"engines": {
"node": "24"
},
"packageManager": "pnpm@10.24.0+sha512.01ff8ae71b4419903b65c60fb2dc9d34cf8bb6e06d03bde112ef38f7a34d6904c424ba66bea5cdcf12890230bf39f9580473140ed9c946fef328b6e5238a345a",
"pnpm": {
"overrides": {
"monaco-editor>dompurify": "3.4.13"
},
"supportedArchitectures": {
"os": [
"current",
"darwin",
"linux",
"win32"
],
"cpu": [
"current",
"x64",
"arm64"
]
},
"onlyBuiltDependencies": [
"@parcel/watcher",
"cpu-features",
"esbuild",
"node-pty",
"sherpa-onnx"
],
"patchedDependencies": {
"node-pty@1.1.0": "config/patches/node-pty@1.1.0.patch",
"@xterm/addon-ligatures@0.11.0-beta.287": "config/patches/@xterm__addon-ligatures@0.11.0-beta.287.patch",
"@xterm/addon-webgl@0.20.0-beta.286": "config/patches/@xterm__addon-webgl@0.20.0-beta.286.patch",
"@xterm/addon-serialize@0.15.0-beta.287": "config/patches/@xterm__addon-serialize@0.15.0-beta.287.patch",
"@xterm/xterm@6.1.0-beta.287": "config/patches/@xterm__xterm@6.1.0-beta.287.patch"
}
},
"reactDoctor": {
"rules": {
"react-doctor/js-combine-iterations": "off"
}
}
}