Commit Graph
1411 Commits
Author SHA1 Message Date
Neil 0a81211994 fix(cursor): read desktop login outside the main thread (#24572)
* fix(cursor): read desktop login outside the main thread

* fix(cursor): await native worker retirement before respawning
2026-10-02 03:48:33 -07:00
Neil 13ecf051c3 Reuse prepared Windows native builds in SSH CI (#24555)
* ci: reuse qualified Windows server slots for SSH host tests

* ci: reuse prepared relay addons after an exact native cache hit
2026-10-02 03:31:57 -07:00
Neil efbf651c7b Reduce CI setup costs and fixture failures (#24537)
* Let scheduled CI warmers wait and measure WebRTC startup

* Measure a smaller daemon shutdown fixture image

* Counterbalance WebRTC startup and verify retained fixture files

* Record CI fixture measurements and remove temporary pilots

* Clarify fixture build dependency cleanup evidence

* Make coalesced snapshot fixture delivery deterministic

* test: type the PTY write delay observer
2026-10-02 02:46:02 -07:00
Neil 6153fbcfe4 Reduce redundant headless server CI work (#24527)
* ci: avoid unrelated headless server qualification

* ci: skip headless detection for ineligible draft PRs

* ci: preserve cross-host qualification and skip supplied prerequisites

* ci: include Windows server cache validation in change detection
2026-10-02 01:42:41 -07:00
OrcaWinandm4air 5f308bfa9c revert: take the 26 Phase 3 (#16741 port) PRs back out of main (#24559)
* Revert "feat(orcad): source-side dormant export of a relay-hosted SSH target (#16741 T6-8) (#24519)"

This reverts commit 783101b304.

* Revert "feat(ssh): update, roll back, recover and stop a managed orcad server (#16741 T6-5 follow-up) (#24463)"

This reverts commit 38c2d1dcb9.

* Revert "feat(ssh): deploy and pair an empty managed orcad server over SSH (#16741 T6-5) (#24453)"

This reverts commit 8b76683b40.

* Revert "fix(ssh): orcad GC honors the activation journal; readiness requires proven daemon coverage (#16741 T6 follow-up) (#24451)"

This reverts commit d3f8c5063b.

* Revert "feat(ssh): remote orcad stop by request file and journaled decommission (#16741 T6-4) (#24449)"

This reverts commit 43d9b43d3f.

* Revert "feat(orcad): supervisable server: stop requests, managed stop receipts and a lifetime that keeps its lock on failed teardown (#16741 T6-3) (#24433)"

This reverts commit b093d3ab20.

* Revert "feat(ssh): crash-safe orcad activation, rollback and recovery (#16741 T6-2) (#24423)"

This reverts commit 1a9ac0e955.

* Revert "feat(runtime): SSH access links for paired servers in a downgrade-safe sidecar (#16741 T5-1+T5-2) (#24420)"

This reverts commit 99db2bfae4.

* Revert "feat(relay): capability-gated owner reset with a durable preparation journal (#16741 T3 R1) (#24418)"

This reverts commit 34a582bd39.

* Revert "feat(ssh): track connection-manager drains, test probes and provider continuations (#16741 T2 P3+P8a) (#24407)"

This reverts commit d53063d2b1.

* Revert "feat(daemon): idle retirement, session census and recovery-only provider (#16741 T2 P4b) (#24409)"

This reverts commit ff212dbbef.

* Revert "feat(ssh): add pty.resumeClient and split SSH PTY process listing (#16741 T2 P5+P6) (#24414)"

This reverts commit 92cb71765e.

* Revert "feat(relay): await owned watcher and agent children on shutdown (#16741 T2 P1) (#24400)"

This reverts commit 6b36e4f85b.

* Revert "feat(session): retry failed renderer session writes and verify local folder PTYs (#16741 T2 P9) (#24406)"

This reverts commit d23ecef301.

* Revert "feat(ssh): remote orcad primitives on the pinned Node runtime (#16741 T6-1) (#24419)"

This reverts commit dd87ae578d.

* Revert "fix(runtime): fence runtime-environment subscriptions and status probes by identity (#16741 T5-3) (#24421)"

This reverts commit ece9e4d2e3.

* Revert "feat(orcad): migration manifest and dormant-state contracts (#16741 T6-7) (#24422)"

This reverts commit 3fbdaba262.

* Revert "feat(ssh): wire SshConnection through the work and transport close ledgers (#16741 T2 P2) (#24401)"

This reverts commit 4e8edc8872.

* Revert "feat(profiles): carry markdown frontmatter visibility in project transfers (#16741 T2 P7) (#24405)"

This reverts commit 60c93263cc.

* Revert "fix(runtime): project the PTY incarnation onto mobile session tabs (#24413)"

This reverts commit 99e0303572.

* Revert "feat(daemon): tag daemon stream data with the PTY incarnation id (#16741 T2 P4a) (#24402)"

This reverts commit 817af768b0.

* Revert "feat(ssh): port the SSH connection work ledger and transport close ledger (#16741 T2) (#24210)"

This reverts commit c9918931c8.

* Revert "feat(relay): fence and drain file and git response streams on shutdown (#24185)"

This reverts commit dc08ffeba9.

* Revert "refactor(runtime-rpc): extract the Node WebSocket lifecycle; opt-in pinned port (#24186)"

This reverts commit a789233bbb.

* Revert "feat(relay): route relay handlers through work admission; producer publication drain (#24181)"

This reverts commit 0b812bd698.

* Revert "feat(relay): land the #16741 T1 seam (work drain, publication drain, release gate) (#24156)"

This reverts commit 3aa2d3af7c.

---------

Co-authored-by: m4air <m4air@Mac.localdomain>
2026-10-02 00:52:32 -07:00
Neil 8ff6296bc7 Speed up serializer checks and keep native caches stable (#24476)
* Reuse serializer oracle cells and isolate native cache policy

* Preserve native cache post-save paths and record hosted oracle gain

* Record native cache reuse and separate cancel-test startup budget
2026-10-01 21:43:56 -07:00
Brennan Benson 444f1952c7 ci: run every cross-version wire test, picked up by folder so new ones can't be skipped (#24499)
* ci(cross-version-wire): run the whole directory so no compatibility test is left out

Three cross-version tests ran in no CI job because the job named its files by hand.
Run the directory instead, ratchet that every file kept out of the unit shards
runs in some PR job, and re-run the job when the modules the newly running
tests guard change.

* test(cross-version): give the orchestration downgrade test its siblings' 120 s budget

* ci(unit-exclusion): count only merge-gating jobs, and require each excluded file's job to fire on it

The coverage check counted any pr.yml job, including e2e, terminal IME and Windows WSL, which are
left out of verify.needs and so cannot block a merge. It now reads verify.needs and the reusable
workflows those jobs call.

It also only proved that some step names each excluded file, not that the job runs when the file
changes. The structured-session zsh login-shell test runs only in shell_contracts, whose path
trigger matched neither it, its harness nor its subject, so a PR touching only those ran it
nowhere. The check now asserts a change to each excluded file fires a gating job that names it,
and the shell trigger gains those three paths.

* ci(cross-version-wire): trigger on the turn-outcome vocabulary and the schema version-skew resolver

A change confined to src/shared/agent-turn-outcome (the arms a newer host publishes) or to
orchestration-schema-version-skew (how current code reopens a downgraded database) skipped the
job whose tests guard exactly those contracts. Also corrects the publish/read direction in the
turn-end comment.

* test(cross-version): state why the orchestration downgrade test needs 120 s

* test(ci): glob the unit tree once for the unit-exclusion coverage checks
2026-10-01 20:53:51 -07:00
Neil f69052e113 Reuse qualified Windows server builds and dependency verification records (#24448) 2026-10-01 16:19:40 -07:00
OrcaWinandm4air 1a9ac0e955 feat(ssh): crash-safe orcad activation, rollback and recovery (#16741 T6-2) (#24423)
Journal every orcad activation and rollback under a host fence so an interrupted one recovers to exactly the slot the activation record names. D7: planOrcadUpdate and assessOrcadRollback refuse a restart whose incoming build cannot attach the live terminal daemon's protocol. POSIX-only and inert: no production caller.

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 12:34:11 -07:00
Neil 197ea3a3b3 Free PR CI capacity by avoiding repeated setup and real-time test waits (#24355)
* Reduce repeated PR setup and transcript timing waits; add hosted comparisons

* Align parallelism contract with Node-only external rebuild toolchain

* Record hosted coverage and launch package, store, and cancellation comparisons

* Apply hosted Windows setup savings and remove measured test waits

* Keep measured PR package gains and remove completed comparison jobs

* Report measured test counts with precise units
2026-10-01 11:51:43 -07:00
Zun 35c8887a76 fix(i18n): correct Korean working and shell labels (#24341) 2026-10-01 14:22:05 -04:00
Brennan Benson 7176648759 fix(native-chat): every chat action press is its own action (re-land #23916 on main) (#24301)
* refactor(native-chat): a stopped child ends on the one reading of its stop

The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.

* feat(native-chat): the host says it accepts a send before any agent has it

The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.

* refactor(native-chat): an attach never opens a journal of its own

The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.

* fix(native-chat): a moved fence resends nothing on a host that accepts first

The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.

The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.

* refactor(native-chat): a child's end says whether the user or the host stopped it

The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.

* fix(native-chat): a chat whose only work is a queued message is not offered for resume

A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.

* fix(native-chat): the conversation outlives its agent

Opening a chat no longer starts its agent. A conversation is reached through one host
accessor that opens its journal at rest, and a send is what starts the agent, through
the delivery loop. One idle sweep, every five minutes, stops an agent that has been
quiet for thirty minutes and owes no work, then drops an open journal handle that is
only a cache. Its record, tab, status row and readers stay.

- hold and release are no-ops; hold still builds the host for shipped mobile builds.
- The holders, the holds, the release clock and the exit respawn are deleted.
- Options, the model list, the goal and the context meter answer at rest; a model pick
  at rest is recorded as intent for the next start.
- Compact, rewind, clear and goal changes start the agent first. A send does too when
  a rewind is still in doubt after the conversation opens.
- Orchestration routes mail and group addresses on ownership (the record plus the chat
  tab), not on whether the process runs. An open dispatch keeps its worker running.
- The restart continuation is a send; Resume all holds each slot until the message is
  handed over or rejected.
- A read error never replaces a loaded transcript, and shows the host's own words.

* test(native-chat): type the queued-message fixtures in the resume-offer tests

* fix(native-chat): a start that dies while a message waits on it is that message's failed start

Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.

* fix(native-chat): a request that failed reads as failed

A structured chat whose only message the agent's start refused read as a
green finish, and a cancelled structured turn did too: the host published a
verdict only for turn records, and structured rows carried no `interrupted`.

The host projection now reads the session's latest request: its turn's
outcome, or `failure` for a send the agent or its start refused. A send
that was withdrawn, or left undelivered by a restart or a close, fails
nobody and makes nothing listable. The ingest publishes `interrupted` as the
hook lanes do, and every reader decodes the verdict through one accessor, so
a failure reads Failed on the dot, the rollups, history and `worktree ps`,
behaves like a cancellation in every clean-finish policy, and notifies as
"failed".

* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now

* test(native-chat): a verdict change republishes the mobile status projection

* refactor(native-chat): the store's retention trigger keeps its flag compare

A verdict change always moves the completion clock the same check already
reads, so a second verdict compare there caught nothing new.

* test(native-chat): a user message the provider journaled keeps its session listed

* test(native-chat): pin what a failed start settles, and what a resume offer names

A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.

* test(native-chat): the failed-start pins fail on what the message became, not on a timeout

* fix(native-chat): a restart offer ends when the chat's agent starts again

The offer used to end only when the chat's newest user message changed,
because opening a chat started its agent and that start could not be told
apart from real activity. Opening a chat starts nothing now, so the host
reads the fact it already publishes: a chat's status row goes from not
host-owned to host-owned exactly when its agent is started. At that edge the
offer and any failure record for the chat are withdrawn, unless the start is
a resume action's own (its continuation is the oldest undelivered message).

A continuation and a message racing to be first are decided at acceptance:
the continuation is refused, quietly and with nothing filed, when any other
message was accepted since the restart. A failed continuation start leaves
the offer retryable, and each resume action sends its own message id.

Deleted: the newest-user-message comparison, its journal reader, the
continuation filter, and the failure ledger's own "answered by the chat"
check. The marker still carries its message id for one release, so the
previous build can read it.

* fix(runtime): end a transcript stream when its client unsubscribes

Desktop: the IPC subscription controller was dropped as soon as the streaming
handler returned, which for most streams is right after it binds. A later
runtime:unsubscribe then found nothing to abort, so the host kept the subscriber
and derived and sent every publish to a channel no one listened to. The controller
now lives until the renderer unsubscribes, resubscribes the same id, or goes away.

Mobile: disposing an agentSession.subscribe stream now sends agentSession.unsubscribe
with the stream's frame id, so the host ends that subscriber and leaves a sibling
stream on the same socket running. The direct path now passes the frame id the relay
path already passed.

* fix(native-chat): a late provider-session update keeps a failed recovery record failed

A provider-session heartbeat that rewrites a completed recovery record kept
its interrupted flag but dropped the outcome it was copied with, so a live
failed checkpoint read as a clean finish until the next status write.

* test(orchestration): the preamble's host stub is typed, not cast

The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.

* test(native-chat): the terminal-bell check asserts the renamed verdict field

The bell notification test still checked for agentInterrupted, which no
longer exists, so it could not catch a verdict leaking into a bell dispatch.

* fix(native-chat): a failed turn ranks like a completion for attention

Attention readers (completion time, Smart Sort, sticky retention, Cmd+J
Recent) now demote only a turn the user stopped. A failure is news the
user has not seen, so it keeps its completion time, ranks in the Done
class, stays retained after its pane goes away, and a retained failure
reads failed in the worktree rollup instead of done. Clean-finish
policy (hibernation, pane ownership, the value moment) still treats a
failure like a stop.

The retention trigger compares verdicts again: success -> failure no
longer moves the completion clock.

* fix(native-chat): one fact ends a restart offer: the chat moved on since the restart

The offer is live while no other message has been accepted in the chat since the
restart and its agent has not proved a start since. The offer list, the resume's
reservation check and the continuation's acceptance check all read that one fact,
so a message whose start then failed withdraws the offer too, and a stale click
finds nothing to act on.

The fact is read off the conversation's open handle, which the restart closed, so
it is retired durably whenever it may have changed: a message accepted, a start
proven. A close and reopen within the same run therefore cannot bring the offer
back. A continuation rejected before it reached the agent does not count, so a
retry after a failed start still runs.

Deleted: the quit-time gate on withdrawal, which changed nothing because the
withdrawal and the quit's own offer write share one queue; the per-action
"withdrawn" flag and the separate acceptance check it paired with.

* test(native-chat): an older build reads the restart offer this build records

The offer lives in a file the previous release reads after a downgrade. Pin that
against the pinned release's own capsule, and run the lane when the marker or the
capsule changes.

* fix(native-chat): read a restart offer against where the journal stood when it was taken

"Since the restart" was read off the conversation's open handle, which the idle
sweep closes: after a reopen, a message the user had already sent looked older
than the handle and the withdrawn offer came back.

The offer now records the journal position (epoch and sequence) at the moment
it is taken, and a message accepted after that position, or a journal on another
epoch, means the chat moved on. That is derived from the journal, so it holds
across any number of closes and reopens. An older build's offer has no position;
only a start withdraws it. Because the message half is now durable, the offer is
no longer rewritten in the recovery file on every accepted message; a proven
start still writes it, since only the host that saw the start knows of it.

* test(native-chat): wait for the listing's retire write before reading the recovery file

* refactor(native-chat): every journal row states which turn it belongs to

Rows gain a turn scope stated by the write that creates them: the open root
turn, or the conversation. A queued message takes its scope from its handover.
Rows stored before scopes existed are placed on replay by the root turn open
when they were created, so no persisted state is needed for them. Rewind keeps
each retained row's scope and producer, so a subagent's row stays its own.

* fix(native-chat): keep the terminal-backed chat's read error over its local echoes

Messages winning over a read error is right for the structured chat, whose read retries and whose
messages came from the transcript. The terminal-backed view assembles its list from local echoes
too (a launch prompt, a pending send), so a failed read there showed only those bubbles and no
error. Only the structured pane now keeps messages over an error.

* fix(native-chat): a start retries the exit settlement a failed journal write left owed

An agent exit whose journal settlement write failed releases the lease latched until a retry lands.
Reopening the chat used to be that retry; with reveal now only opening the journal, nothing retried
it before the next app launch, and every send was refused. The start the send needs now runs the
retry first, where the attach would.

* fix(native-chat): a failed main agent reads failed while its subagents still work

The verdict is now read from the main agent's own state, not the folded
row: a main agent that is done and failed has a verdict even while its
subagents keep the row working. Without mainAgent (history, worktree ps,
older hosts) the old combined-done rule stands.

Display marks the verdict through agentVerdictDisplayMark: a failure
outranks every combined state on the agent's dot, label, tab badge,
dashboard and activity rows; a stop marks only a done row, so a
successful or stopped main agent with live subagents still reads
working. Subagent rows keep their own state. The worktree card, terminal
tab and Cmd+J rollups share one pane fold and rank a pending question,
then failed, then working, monitoring, interrupted and done.

worktree ps publishes the main agent's outcome on a working row, and the
mobile mirror reads it. The store's change check, the paired-client
mirror's equality and its epoch now see a verdict change on a working
row, which otherwise moves no state or clock and left the worktree card
reading working. Clean-finish policy is unchanged: a working row is never
hibernated and has no completion time.

* perf(native-chat): answer the owner check without opening the chat

Worktree activation calls agentSession.handoffStatus for every chat tab in the worktree, and the
answer comes from the session record alone. Reaching it through the accessor opened each resting
chat's journal (a full read, the crash-boundary write and a restored status publish), then kept it
open for the idle window. It now checks the record and the adapter's support, as before this series,
and opens nothing.

* fix(native-chat): a read waiting on the session lock opens nothing once quit began

The accessor checked for quit before queueing the open, so a read queued behind a session task ran
its open after teardown had begun and indexed a journal no teardown step would close. The check now
runs at the open itself.

* test(native-chat): pin stated turn scopes, the upcast of unscoped rows, and rewind attribution

* fix(native-chat): /compact is a message the chat sends, run as a turn of its own

The conversation command RPC now accepts /compact into the queue like any
send and answers once it is handed over. The delivery loop opens the command's
own turn, starts the provider on it, and waits for the provider's end off the
session's queue, so messages typed meanwhile are held and delivered after it,
even when it fails. It settles by re-reading the journal: a child that died
meanwhile already wrote the verdict. Stop ends the command at once. The 180 s
completion window, the unconfirmed row and the recovery of an older build's
compaction record are gone; that record no longer gates anything. On Codex the
provider turn the command opens is claimed into the command's turn.

* fix(native-chat): read a failed resume's chat before calling it retryable

Whether a failed resume is retryable is the offer's own rule: the chat has not moved on since the
restart, read from its journal. The failure list read it only for a chat already open, so once the
idle sweep closed a chat the user had moved on in, its failure showed Retry again, and the click did
nothing. The list now opens the failed chats first, as the offer list does.

* test(native-chat): type the provider event sink the settlement test reaches for

* docs(native-chat): the worktree ps outcome comment no longer claims old hosts send it

The field is new: an old host sends no outcome at all, so a reader falls
back to interrupted. The removed clause said old hosts send it on done
rows, which never shipped.

* fix(native-chat): say the structured read keeps trying only where it does

The structured pane's "Orca keeps trying to load it" line never showed: the view state filled in an
untranslated fallback whenever the read error had no text, and the empty state prefers any message.
The view state now leaves the message out, so the structured pane shows that line and the
terminal-backed pane its own translated one. Mobile's structured lane does not resubscribe after an
error frame, so it no longer makes the claim.

* fix(native-chat): rows group under the turn their record names, not the one above them

Each row's turn is the turn its stated scope names, anchored on the entry
that opened it, or on the turn itself when the provider opened it unasked.
So /compact groups its own rows and the previous turn is untouched, a message
typed into a running turn joins it, and a provider-resumed turn folds under
its own Worked-for. A row reporting how a turn ended, an error or the
compaction separator, never folds. Desktop and mobile read the same keys; a
host that states no scope keeps today's positional grouping.

* test(native-chat): await the send's settlement instead of polling for the start

The at-rest send tests polled for the provider start with vi.waitFor's one-second default, which a
loaded machine outran. They now await the host's own settlement of the message.

* docs(native-chat): the status-store listing rule names provider-journaled user messages

* fix(native-chat): a restart offer resumes any time after the quit, and knows its own continuations

The continuation's message id was dated by the quit, and the ledger refuses a new id dated more than
a day back, so Resume or Retry a day after quitting was always refused (on main too). It is now
dated by the resume action.

Telling a rejected continuation from the user's own message read the operation ledger, whose rows
expire after about a day; after that a failed resume stopped being retryable. The offer now
records the continuation each action sends on its own capsule entry, bounded to the newest 16, so
the ids end with the offer. The ledger read is deleted.

* fix(native-chat): a /compact is not a request the sidebar, notifications or restart resume report

The sidebar's prompt, preview, verdict and instant, the turn-completion feed,
and the restart-resume marker read past a conversation command and its turn to
the last real request, so a /compact neither notifies nor re-dates the row,
and a command in flight is never offered as work to resume. An older client
shown a command's turn in the legacy form names the session's own agent.

* fix(orchestration): route no mail to a structured worker its orchestration released

A structured worker is routed on ownership, and a resting worker's lease is released, so ownership
held while its chat tab stayed listed. A worker the coordinator abandoned and then released, found
at rest by the release, therefore still took peer mail and @worktree: broadcasts, and each one
restarted its agent. Routing now also reads the orchestration's own resource row: once it is
released, direct mail, group addressing and worker-show's addressable answer drop the worker, as
they would a terminal worker whose terminal closed. The chat tab stays, and nothing new is stored.

* fix(native-chat): a failed retry names the user's prompt, not Orca's continuation

A resume's continuation is written to the chat before its start, so after a failed attempt the chat's
newest user message is that rejected continuation. A second failure then showed Orca's own restart
text as the chat's prompt. A retry now keeps the prompt its first failure named.

* test(native-chat): pin what a conversation command's admission refuses at rest and at handover

* test(native-chat): tests merged from the base state which turn their rows belong to

* fix(native-chat): a refused send notifies failed through the completion feed

The host's completion feed followed only the newest turn, so a send the
agent or its start refused, which creates no turn, read Failed on its row
but sent no notification. The feed now follows the session's latest
request, read from the projection the status feed already makes for the
commit: a turn keeps its id, a refused send is named by its journal item
key. It announces only while the session is idle, as the row reports a
verdict, so queued sends refused one commit at a time notify once, and a
withdrawn send falls back to a request already announced.

* fix(orchestration): read the released row optionally, as the authority does

worker-show's observation called the row lookup directly, which a runtime double without it threw on
and failed the structured tab-retirement release.

* chore(native-chat): one import per module and no unexplained casts in the turn-scope changes

* test(claude): pin which turn a Claude row joins, including a subagent's after the turn ends

* fix(native-chat): the status bar drops a restart offer the chat moved on from

The renderer re-read the host's restart offer only when a failed chat showed activity, so after a
message withdrew a pending offer the host answered no chats while the status bar kept counting one,
and clicking it opened nothing. The same watch now covers pending offers: a status change in an
offered chat asks the host again, once.

* fix(native-chat): a refused steer is read from the turn its handover named

The latest-request reader decided whether a refused send had joined a running turn by comparing
host clocks: its handover time against the previous turn's end. The handover row now states the
turn it delivered into, so the reader reads that instead and the clock comparison goes. A journal
written before handover rows stated a turn is scoped on replay from the turn open when each row
was written, which can differ from the clock reading only when a send and a turn's end share a
millisecond.

* fix(mobile): the native-chat controller contract carries the turn journal

The controller and overlay already pass nativeChatTurnJournal, but the
contract type never declared it, so mobile failed to typecheck.

* fix(native-chat): the live turn is the running turn, not the newest user row

A turn the provider opened on its own (a background wake, a resumed turn)
anchors on its own record, but the list still treated the newest user row
as the live turn. While such a turn ran, the settled user turn before it
lost its duration and the running turn's own rows were drawn as settled,
so its tool calls lost their live state.

nativeChatTurnMembership now answers both questions from the turn record:
each row's turn, and the live turn (the running root turn's anchor, else
the newest user row, which is also all an unscoped host has). Desktop and
mobile key liveness, the timing clock and the live status's row on it.

* test(native-chat): a turn the provider opened keeps its own clock

Pins that the local turn clock follows the live turn, so a wake after a
settled turn does not restart that turn's clock when no host durations
are recorded.

* fix(native-chat): a running turn no message opened draws its status on no row

Its live status belongs to the transcript-tail indicator alone. Once it
settles, its duration draws at its first row as before; a running turn a
message opened still draws on that message.

* fix(native-chat): every copy of a row carries the main agent's own status

History entries, sleep records and `worktree ps` rows carried a flattened
top-level `outcome`, copied under different gates and without the main agent's
clock. They now carry `mainAgent` (state, outcome, stateStartedAt), the type
the live row already persists and sends, and every copy site takes it with
`interrupted` through one function, `agentVerdictFields`.

- The accessor reads `mainAgent` then the legacy flag; the mobile mirror
  matches it line for line.
- Sleep records admit `mainAgent` with `normalizeMainAgentStatusField`, so a
  malformed value drops the field, never the record.
- Mobile dates a main agent that failed under live subagents by its own clock,
  as desktop does, and its row equality compares `mainAgent`.
- The activity feed reads a history entry's own `mainAgent` instead of
  rebuilding one; the sync key and history equality compare it.

* test(native-chat): pin the worktree ps verdict across host and phone versions

Pairs the real v1.4.212 host and phone row reader with this build: an old phone
reads a new host's rows by `interrupted`, a new phone reads an old host's rows
(no `mainAgent`) the same way, and a new phone reads a failure under live
subagents as Failed, dated by `mainAgent.stateStartedAt`. The release checkout
now carries the phone's self-contained row reader, and the lane runs when the
`worktree ps` row producers change.

* test(mobile): name the parity table's row for its role

* test(native-chat): a roster of idle or finished children does not keep an agent awake

The sweep reads owed background work through the shared child-work liveness that upstream's
release clock adopted; a child that went idle or finished is not work the agent still owes.

* fix(native-chat): a request that settles while the user is asked something notifies once

The completion edge waited for an idle session, and a pending prompt (including a
subagent's approval) is not idle. Structured chat has no other attention producer,
so a main turn that finished while a subagent waited on the user sent nothing
until the prompt was answered.

The edge now waits only on owed work (a running turn or an unanswered send), which
the projection reports even beneath a pending prompt. A request that settles with
a prompt pending announces once; the renderer words it "needs input" from the
host status mirror's `attention`, and answering the prompt keeps the same request
identity, so it does not announce again. The wire shape is unchanged.

* fix(orchestration): a task dispatched into a resting structured worker keeps it running

The sweep's open-dispatch check read only the worker-start dispatch that owns the worker's terminal
resource, so a task later dispatched to the same worker (orchestration dispatch --to, which writes a
dispatch with no worker row) did not count: after thirty quiet minutes the worker was stopped while
that task was open, and its coordinator read exited. Any unsettled dispatch addressed to the worker's
process incarnation now counts, derived from the existing rows.

* fix(native-chat): a command's wait ends when its child does

The delivery loop waited for a /compact only on the adapter's compaction
tracker, which learns of the child's end only on some exit paths: a Codex
exit or close, and a Claude close, never reach it. The wait then never
ended, so nothing queued behind the command was delivered again, Stop had
no child to answer through, and the tracker's leftover entry refused the
next /compact.

Every way a child ends passes endProviderChild, so the host now offers a
per-child end signal there. The loop races the tracker against it (the
dead-generation settlement has already written the command's verdict),
and on that end asks every adapter to release the command, so a later
command runs and no later provider turn is claimed into the dead one.
The adapters' own exit-time releases were unreachable (Codex) or covered
one path of several (Claude), and are removed.

The Codex RPC test harness moves to its own module so the exit can be
driven through the real adapter's connection callback.

* fix(native-chat): keep refusing sends during a command on an older host

An older host's controller still refuses a send while a conversation
command runs, so dropping the client's block turned every message typed
during /compact into a 'not sent' row with Retry there. The block stays
for hosts that do not run the command as a send-path turn, and goes only
for those that do.

The signal is one the client already holds: a host that runs /compact on
the send path states a turn scope on every journal row it writes, the
same fact turn membership uses to tell it from an older host. Both now
read it from one predicate. On an empty conversation, or one whose rows
all predate the upgrade, the signal is absent until the command's own
entry streams in, so that brief window keeps the old local refusal; no
capability or wire field is added.

* docs(native-chat): comments stop describing the hold this PR removed

Eight comments still justified orderings and teardown choices by a viewer or dispatch hold that
pinned the provider child. Nothing holds any more; the orderings stand for the binding's redrive
subscription and parked mail, and a chat's agent runs from a send until the idle sweep rests it.
Comment-only.

* fix(native-chat): the completion says when the user is being asked

A request that settles while a prompt waits on the user was worded "needs input"
from the renderer's status-feed mirror. Remote clients receive the status and
completion streams over separate sockets, so they can arrive in either order and
the wording could be wrong both ways.

The host already knows at emit time, so the completion now carries an optional
`awaitingUser: true` in that case and omits it otherwise. The renderer words the
notification from that field alone and no longer reads the status mirror. Old
clients ignore the field and word by outcome; old hosts never send it.

* fix(native-chat): a restart offer keeps the start its own continuation made

Whose start ended an offer was decided at read time, from whether the offer's continuation was
still the queued message. Once the provider refused that continuation, the child it had started
read as someone else's start, so the offer ended and its failure showed no Retry. The delivery
loop now records which queued message a start is for on the in-memory child, and the child's end
carries it; the offer counts a start as its own when that message is one of its continuations.

* fix(native-chat): a rewound turn still names the message that opened it

A Codex rewind rebuilds the epoch without submissions, so each sent message survives only under
its provider key. The kept turn records still named the submission key, so each turn anchored on
itself and its rows grouped apart from the message that opened it. The rewind now renames the
turn's opener along with the message.

* fix(native-chat): Stop ends only the command it names

Stop on a command turn abandoned whatever compaction the session had pending, so a late Stop for
an earlier /compact cancelled the one running now. The tracker now ends a command only when the
Stop names its turn, and the cancel reply reports whether it did.

* fix(native-chat): an agent gets a full idle window after its owed work ends

The sweep measured quiet only from the last journal row, so once a subagent, command, monitor or
dispatch that had outlived the window ended, the agent was stopped at the next tick. A child can
read done before the lead's wake-up turn writes anything, and stopping in that gap loses the
wake-up. The sweep now counts owed work it observes as activity, which gives the agent the full
window afterwards, as the release clock it replaced did.

* test(claude): the options-read fixture runs a live child

The fixture marked its conversation running with a hasProviderChild field the
session type does not have, so the read took the at-rest path and refused a
session with no record. It now carries a child, which is what the read checks.

* test(native-chat): host tests reach its collaborators through a typed seam

The rest-test rig and three test files read the host's private members with
Reflect.get and cast the result. The host now exposes one test-only accessor,
collaboratorsForTests(), and the subscribers class a subscriberCountForTests()
beside its existing retainedActivityCountForTests(), so the tests are checked
against the real types and the casts are gone.

* fix(worktree-status): a departed agent's failure yields to live work on the worktree card

A retained failed agent has no expiry, so ranking it with a live failure pinned the card to Failed over other panes' live work. It now ranks below working, monitoring and permission, and above every finished outcome.

* refactor(orchestration): one owner answers a structured worker's custody

Routing, group addressing, worker-show and the idle sweep each composed their own reading of
whether orchestration still holds a structured worker, so each new obligation or retirement state
had to be added to every reader. structured-worker-custody now derives both answers from the
worker-terminal list state coordinators see in worker-list: addressable is owned and not released,
and owed work is an active custody or an unsettled task dispatched to the same incarnation. The
owner's state is read through the remote dispatch attachment too, as the terminal transfer lookup
already does. Behaviour is unchanged; a settled worker awaiting its coordinator still rests.

* refactor(orchestration): owed work is an open dispatch on the worker's incarnation

A supervised worker's own dispatch context stays open exactly while the worker is active, so the
separate active-custody branch only repeated it. Owed work is now one fact, which also states the
policy that a worker awaiting its coordinator's decision may rest, and both custody decisions are
written once at the top of the module.

* docs(agent-status): a departed agent's failure ranks below live work on the worktree card

* fix(native-chat): a restart offer knows its continuations by a tag in their id

The offer recorded each continuation id in a list on its capsule entry, capped at 16, and a running
action's id in memory. Both could disagree with the journal: past the cap an old rejected
continuation read as the chat moving on, and a crash during a retry restored the failure's older
entry, which lacked the retry's id. Each continuation id now carries a tag derived from the offer
(its teardown and chat), then the action's own part, so any continuation of this offer, queued or
rejected, is recognised from the journal row and the marker alone. The persisted list, its cap and
the in-memory action map are deleted; the agent-start withdrawal keeps an offer whose own
continuation the start was for, read against the stored marker.

* test(runtime): the legacy-worker reveal test judges its stale snapshot inside the wait

The tui-idle probe reads through readTerminal, which now awaits the structured
worker check before the PTY read, so the probe's snapshot request starts a
microtask later. vi.waitFor missed it on its first check and polled again at
50 ms, the same moment the wait's own 50 ms timeout fired. The stale snapshot
then resolved after the wait had already timed out, so the test passed without
judging it, and the rejection landed before any handler was attached. Vitest
reported that as an unhandled error and failed the shard.

Polling every 1 ms sees the request within a few ms, so the snapshot is judged
while the wait is still pending.

* fix(native-chat): a message held behind /compact is drawn where it was handed over

A message typed while /compact runs was drawn above the compaction's result, between
itself and its own answer. The reducer kept every item at the sequence and timestamp of
the row that created it, and a queued message is created at acceptance, long before the
command it waits behind writes its result. The phone orders by that sequence and the
desktop by that timestamp, so both put the message first.

A queued message now takes its position from its handover row, the same row that already
states its turn scope. Everything the agent did before the handover, a command it waited
behind included, draws above it. This holds for every held message, not only /compact's,
and needs no client change: every client, older builds included, reads the position the
host publishes. A live batch already carries the item when its dispatch row lands, and
history pages cut the reduced timeline by sequence, so paging stays contiguous.

* fix(native-chat): a phone's send during /compact answers without waiting out the compaction

A client that predates accepted-send replies, which is every phone build, has its send
reply held until the host hands the message over. A message sent during /compact is not
handed over until the compaction ends, so the phone's 15 s request timeout fired first
and showed the message as unconfirmed.

That wait now also ends once the message is queued behind a running command. This is
read from the journal's running turn and needs no new state. Every other wait still
ends at the handover: behind a starting child or an ordinary turn, and for restart
resume, the command front door and orchestration, which keep the plain handover point.

* perf(native-chat): a rewind places provider items with one pass over the merged rows

A Codex rewind gives each provider item the old epoch never held the turn record for its
provider turn. It found that record by scanning every merged row, restoring each row's
body, once per provider item. That is quadratic, and it runs on the host's main thread
up to the journal's 10,000-row cap, twice per rewind. A rewind record written before
rows carried their scope holds no scope for any provider item, so it paid the full cost.

The merge now indexes turn records by provider turn id once, keeping the first match as
the scan did, and each provider item looks its record up.

* fix(native-chat): a view never restarts a chat whose last start failed

A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.

* fix(native-chat): a message waiting behind /compact is drawn after it until it is sent

A message sent while /compact runs is placed where it was handed over. It was still
drawn where it was accepted until then. /compact writes its result one step before the
handover, so for that step the waiting message sat above the compaction's separator.

A message the host accepted but has not handed over is not part of the conversation
yet, so both clients now draw it after everything the agent has done. The shared
projection moves it to the end, which is the order the phone draws. The desktop ranks
it with the other not-yet-sent rows, after the streaming preview. At handover it takes
its place from its handover row, which is also after the separator, so it never
appears above the compaction it waited for.

* fix(native-chat): the idle sweep reads owed work every tick

Owed work counted as activity, but the sweep read it only once the idle window had elapsed, so it
refreshed the clock at most once a window. Work that ended just before the next read left the
agent to be stopped at that read, moments after the work ended, which is the gap the refresh was
meant to cover. The sweep now reads owed work on every tick for a started agent, so the window
always runs from the last tick that saw work owed.

* fix(native-chat): a continuation handed to the agent stays sent

The offer read its own continuation as not reaching the agent while its dispatch was pending, which
also covered one already handed over and still unanswered. When the wait for that answer ended first,
the failure it filed read as retryable, and a retry sent a second continuation to an agent that may
have acted on the first. Only a continuation still queued, or rejected, is now read as unsent.

* test(native-chat): start the child the loop waits on with an attach, not a second view

A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.

* fix(native-chat): settle a gone generation's turn wherever a conversation opens

A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.

* test(native-chat): prove the next child's start settles the turn an earlier child left

The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.

* test(native-chat): count a failed start's rows by row, not by text

Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.

* test(cross-version): load the phone row readers without mobile's toolchain

Vite transforms a file against its nearest tsconfig, and mobile/tsconfig.json
extends expo/tsconfig.base.json, which the root-only cross-version lane never
installs. The worktree ps verdict suite imported the current phone row reader
from mobile/ directly, so CI failed with TSConfckParseError before any test ran.

The harness now imports a copy of the working-tree reader placed under the
checkout cache, where the root tsconfig applies, as it already does for the
release checkout's copy. Both readers are still the real files.

* test(cross-version): keep the checkout path-guard message and justify the copy import's cast

* fix(native-chat): a command ends only by its own provider answer or its child's end

Stop no longer settles a conversation command. It interrupts it like any turn,
and when the provider cannot take that (Codex has not opened the command's turn
yet, or Claude refuses the interrupt) it stops the child, whose dead-generation
settlement writes the verdict.

The pending command now lives on the provider child's own session instead of an
adapter-wide map keyed by session, so it dies with the child and nothing has to
release it. Claude's /compact is sent under a uuid the slot records, and only a
root result naming that input (or naming none) ends it; its outcome is read with
the ordinary result reading, so a stopped /compact is a cancellation.

* fix(native-chat): a command's settle answers its message before ending its turn

The two writes are not one batch. Writing the message's answer first means a
crash between them leaves a running command turn, which the stale-turn sweep
already settles, instead of an ended turn whose message reads as in flight
forever. The settle now writes only while the command turn is still running.

* fix(native-chat): "Worked for" counts from the handover, not the send

A message held behind /compact, or behind a cold start, used to count the wait
as the agent's work, although its row is drawn at the handover. Every handed-over
submission's turn, the command's own included, now starts at the handover row's
instant, falling back to the send time for a host that recorded none.

* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget

* test(native-chat): the interrupted create's own retry continues again

The merge of main's lease-latch fix replaced that test's retry of the interrupted create, under its
own operation id, with a fresh start whose result nothing read. That fresh start passes with the
released-reservation continuation deleted, so the case the fix exists for went untested. The retry
and its assertion are main's again.

* docs(native-chat): three comments that still had views starting agents

A start with nothing queued now comes from a command, goal change or rewind; an interrupted compaction
left alone would refuse every send, so no agent would ever start to finish it; and a current host
raises the unattached read refusal only once quit began, with the attach window belonging to an older
host.

* test(native-chat): pin the open's and the send's start and row counts, however the view binds

Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.

* fix(native-chat): a second Stop on a command ends its child; one compaction verdict for every provider

A Stop's note now names itself in its key, so a later Stop on a command still
running reads, from the journal, that the provider was already asked and never
answered, and stops the child instead of interrupting again. Nothing is held in
memory for it.

Adds the rule both translators will read a compaction's end by: only a
compaction the provider reported is a success; none after Orca's interrupt is a
cancellation; anything else is a failure. A real Claude capture, pinned as a
fixture, is why: a stopped /compact ends in the same success result as a
finished one.

* test(native-chat): a reader's open settles the turn a failed exit settlement left running

An exit whose settlement write failed leaves its turn running in the open journal. PR 1's open now
settles it, and this pins the two reads that reach it here: a reader reopening a chat the idle
sweep closed, and a read that opens the chat before the restart restore reaches it.

* test(native-chat): the view-start test's starting window outlasts two subscriptions on a loaded runner

A subscription reads the conversation before it returns, so under load the two views took longer
than the create child's 300 ms start, which then exited before the test checked that it had not.
The child now takes a second to fail.

* fix(native-chat): settle a gone generation's turn at every open but an acquisition's

The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.

* test(native-chat): hold the create's start open until the views bind

The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.

* refactor(native-chat): the provider's translator ends a command's turn; the loop holds no command state

A conversation command is now a turn of the provider child's own journal
pipeline. The adapter-wide tracker, its promise and the loop's settle step are
gone.

- Codex: the translator claims the provider turn that carries the command, scopes
  its rows to the command's turn, and writes the command's end in the same batch
  that settles that turn. Codex's own compaction marker is the success row.
- Claude: the command's turn is the translator's open turn until the result that
  answers the /compact input ends it. The command's own frames, such as the
  continuation summary, its echo and "Compaction canceled.", draw nothing.
- Both read the end with the one compaction rule: success needs the provider's
  report of the compaction; none after Orca's interrupt is a cancellation.
- The message resolves at the provider's receipt, as any send does: the Codex
  ack, or the Claude slash-command waiter on its result. The host writes a
  command's end only when the provider never took it.
- The delivery loop stops while a command's turn runs, and every journal commit
  re-wakes it through the session's serialize, so an end that lands while a step
  decides to stop is never lost. A child that ends first is settled with it.

* test(native-chat): pin a command's end to real /compact frames and to each path it threads

The captured /compact frames drive the Claude translator's command turn: a
finished compaction ends as a success with only the separator drawn; a stopped
one ends as a cancellation with no failure row, and the next send answers in its
own turn; a result naming another input ends nothing. The command's end is
checked at each point the ordinary result path threads through: the reopen latch
after a failure, the settling of a child still working, the context facts the
result reports, and the provider's own error row.

On the host: a message held behind a command is handed over when the command
ends just as the loop stops for it, a refused command settles as a failure and
the loop moves on, and a Claude child that exits mid-command settles the command
and hands what waited to a fresh child.

* test(native-chat): tests merged from the base state which turn their rows belong to

* refactor(native-chat): drop the child-end waiter nothing waits on

A command no longer waits for its child here: its turn ends from the provider's frames or from
that child's settlement, and the delivery loop is woken by the commit. The waiter and its test
were left from the earlier shape.

* fix(native-chat): a command holds the queue only while its child runs it

The delivery loop stopped whenever the journal showed a command's turn running. When the
command's child ended and its settlement could not be written, that turn stayed running with
no child to end it, and the loop's gate kept it from ever starting the next child, which is
what settles a gone generation's leftovers. Every later send was held for good, and Stop had
no child to end.

The gate now holds only while the conversation has a child: with none, the command belongs to
a gone generation, and the loop's start settles it like any turn a dead child left running.

* fix(native-chat): a Claude /compact succeeds only on its compaction boundary

The command's evidence counted Claude's `compact_result: 'success'` status as the compaction
done. That status comes before the boundary that replaces the history, so a Stop landing
between the two read as a finished compaction even though no boundary was ever written. Only
the boundary now counts, as the rule for both providers states; the capture's finished
compaction carries one, so it still reads as a success.

* fix(native-chat): a Claude child's exit says why the turn it ended stopped

When a Claude child exited mid-/compact, the command showed "Worked for 0s" and no reason. The
child's translator ends its open turn the moment the exit is reported, stamped with the exit's
instant, so by the time the exit settlement ran nothing was running. The settlement recognises a
turn the exit already ended by that same instant, but the Claude lifecycle event dropped it on the
way to the host, which then used its own clock, matched nothing, and wrote no row. When the clocks
did agree, the row was scoped to the running turn, of which there was none, so it landed outside
the turn it explained.

The exit's instant now reaches the host, and the exit row belongs to the turn the exit ended:
still running, or ended by the translator at that instant.

* fix(native-chat): a message waiting behind /compact draws below its live activity

A message sent while /compact runs waits on the host until the command ends. Both clients moved
it to the end of the transcript rows, but the running turn's live activity line ("Compacting the
conversation") draws after every row, so the waiting message sat between the command and its own
live status.

A row that is queued, and not what the live turn is for, now draws after that live activity: on
desktop outside the transcript window, below the activity line; on the phone in the list footer,
below the live status. A message whose own start is pending still draws above the activity that
start reports.

* fix(native-chat): only a running command holds a message below its live activity

A message is accepted, then handed over a moment later, and in between it reads as waiting. Every
message waiting behind a live turn drew below that turn's activity line, so an ordinary message
sent while the agent was working crossed below "Thinking" and jumped back up once it was handed
over, on desktop and phone. Only a conversation command's turn holds the queue on the host.

A message now waits below the live activity only while the running turn is one a command opened,
read from the entry that opened it. The phone test also typechecks, which the mobile test ratchet
requires.

* test(codex): the claim test names its notification params as a record

* test(native-chat): a read that reaches a crashed chat before the startup reconcile settles its turn

On desktop the chat on screen at relaunch reads before startup reconciles the leases, while the
dead process's lease still reads live. The open settles the turn it left running anyway, and the
restore that follows finds it settled.

* refactor(native-chat): drop the composer's second error formatter

After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.

* test(native-chat): pin the reason on a message rejected while its chat was closed

The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.

* docs(native-chat): drop the removed dispatch hold from six comments

A worker's session no longer takes a dispatch hold, and no release clock
rests a chat by visibility; the agent-launch comments, the abandon test,
the teardown test and the refusal census still said so.

* test(native-chat): rest the owner-status chat through the idle sweep, not a hold

The activation-gate test from #22808 put its chat at rest by holding and
releasing it, and passed the release-clock grace. This branch deleted both,
so the case threw before it reached its assertions. It now moves the host's
clock past the idle window and lets the sweep stop the agent and close the
conversation, then asserts the same owner answer and activation gate.

* fix(native-chat): show the structured pane's retrying line when a read fails

The read transport always hands the pane the host's words, so the error
state's "Orca keeps trying to load it" line, which showed only when there
were none, was never seen: the pane showed the host's text twice, as its
subtitle and on the status line under it. The structured pane now always
says its read keeps retrying, and the host's text stays on the status line.
The terminal-backed chat is unchanged.

* fix(native-chat): a send the provider never received after a restart has no verdict

Restart reconciliation rejects a crash-stranded send that is absent from a
trustworthy provider history with reason 'not_delivered'. Nobody failed that
send, but the verdict allowlist did not name it, so after a crash the chat
read Failed, was listed, and could notify "failed". Give the reason a shared
constant (persisted value unchanged), add it to the no-verdict set, and treat
it as an internal marker so the Retry row no longer shows the raw string.

* fix(native-chat): a failed Codex compaction's late completion writes no turn of its own

Codex ends a failed turn with an error and then still completes it as failed.
The error settled the compaction and released its claim on the provider turn,
so the completion read that turn as an ordinary one and wrote a stray record.
The claim now lasts until the completion, which adds nothing to a command the
error already ended.

* test(native-chat): the mid-command exit case resumes its next child as a real one does

The case's fake started every child as a newly created thread with the same generation. The
store refuses a created link once the conversation has a thread, so the next child's start
failed and wrote its own error row, which landed before or after the case read the journal.
The next child now resumes the thread under its own generation, and the case reads the
journal once the waiting message is delivered, which also proves the loop moved on.

* fix(native-chat): a /clear that never committed no longer locks the chat

A /clear wrote a durable "prepared, outcome unknown" record before starting
the replacement conversation. When that start was refused without a definite
answer (or Orca died), the record stayed forever, and while it did the chat
refused every send, /compact, a new /clear and rewind. Its only exit was a
rerun under the same operation id, which only the renderer held.

The record guarded nothing the process does not already know: a clear in
flight holds the session's serialize for its whole run and the command
controller refuses sends meanwhile, and the replacement's id and start
operation are pure functions of the clear's operation id. So the clear now
writes nothing durable before its commit, the gates refuse only a committed
clear (an older build's prepared record is inert), and a clear with no
committed answer reruns: a same-op retry re-attaches the same replacement,
a new op id runs a fresh clear.

A crash between the replacement's start and the commit leaves a replacement
record nothing points at. Verified: it has no tab, is not in the
replacement list, and a restart opens and starts nothing for it (restore
reads only the visible tab index); restart reconciliation releases its lease
like any dead owner's. In a live process its agent is stopped by the idle
sweep like any quiet agent. Session History lists provider transcripts and
only annotates them with an owner, so it can list this only if the provider
wrote a transcript for a thread that never got a message. Its record stays
on disk, as every closed chat's does; the store deletes none.

* fix(native-chat): a Codex rewind the provider did not keep no longer fails every attach

When Codex acknowledged a revert and Orca stopped before proving it, the
rewind stayed prepared with providerApplied set. On the next attach,
recovery read the provider's history, found the target turn still there
(provider-refused), and threw, because that settlement was limited to
reverts never sent. The throw ran inside the attach, so every attach, and
every send that needs one, failed for good.

The journal is replaced only once the provider proves the revert, so both
the provider and the journal still hold the target turn: settling the
rewind refused is consistent whether or not the provider acknowledged it.

* test(native-chat): a clear retried after a crash starts no second replacement

The replacement's id is the only thing that keeps a retried clear from leaving a second one, and no test held it across a restart.

* chore(native-chat): the clear rerun comment claims only the stable replacement id

* test(native-chat): wait for a send's background start before the refusal oracle removes its store

An accepted send wakes the delivery loop, which starts the agent in the background. The oracle's teardown disposed the loop but did not wait for that start, so its lease write could create a temp file in the store directory while the directory was being removed, failing the test with ENOTEMPTY about one run in four. The teardown now drains tracked starts before it closes the journals.

* fix(native-chat): a start a message waited on gets one failure row, the delivery loop's

When a queued message's start failed, two writers could report it under the same row: the delivery loop, when the adapter settled the start without proving it, and the exit settlement, when the child's exit landed. The last one won, so the chat's row could name a different cause than the one the message was rejected with, or be written twice.

The exit settlement now writes the start's row only when no message is queued and the loop has not already recorded that start. A start for a command, goal change or rewind, with nothing queued, still gets its row from the exit.

* fix(native-chat): a /compact whose start failed says to run /compact again

The failure-words context named only /clear as a command to retry, so a
/compact whose agent failed to start read "Send your message to try again."
on its row, its rejected message and the command reply. The context now
carries any conversation command; the host derives it from the oldest
message still waiting on the provider, which is the one a failed start
fails first, and the /compact reply names it directly.

* fix(native-chat): a Codex /compact ends only on its turn's completion, below Codex's own error row

Since only turn/completed ends a Codex turn, Codex's turn-ending `error` is a row
inside the still-open command turn, and the failed completion that follows it is
the command's end: completed, outcome failure, at the completion's receipt time.
The command's own "Compaction failed" row was written on that completion too, so a
failed /compact read its reason twice.

The command turn now notes when Codex's turn-ending error for the turn it carries
was written as a row, and its end then adds no second row. A retried stream error
ends nothing and is not counted. The flag that let the error end the command and
kept the claim until the completion is gone with the error-driven end.

A test replays the captured failed compaction from the real app-server through a
claimed command turn.

* test(native-chat): main's crash-turn test states its row's turn, and a dead /compact settles on its recorded exit

Two tests the main merge brought together:
- The crash-turn test from #23456 writes a turn record through the event sink
  without options; every row here states its turn scope, and a turn record's is
  the thread.
- The /compact whose exit settlement could not be written no longer stays running
  until the next start: main now settles an open chat from the exit it recorded, so
  the command reads interrupted before the next message, which is then delivered.

* test(native-chat): main's new journal tests state each row's turn

The crash-turn, stale-turn and sink-queue tests main added wrote rows without a
turn scope, which every item write now states. Rows written inside a running
turn name that turn; the sink-queue batch and a send handed over with no live
turn name the thread.

* fix(native-chat): draw a queued turn's message after the earlier turn's rows

A message sent while A runs is written to the journal when it is sent.
When the provider queues it (Claude answers it after A), A's remaining
rows - its last tool run and its answer - are written after that
message, and the message's own turn opens only after them. Grouping put
those rows in A's turn, but the transcript still drew them in journal
order, below B's bubble and bar, where A's answer read as B's reply. This
is the residual #23671 left open.

A message that opened a turn now draws after the earlier turns' rows the
journal wrote after it, just before its own turn's rows
(nativeChatTurnDrawOrder, returned by nativeChatTurnMembership as
drawOrder). Desktop and mobile both draw in that order. A steer, and a
message that has opened no turn yet, stay where they were written. It
applies on hosts that state turn scopes and, through journal order, on
older ones.

* test(native-chat): run #23026's Stop tests against #23059's command turns

Two of #23026's tests call APIs #23059 changed, and failed after the
merge:

- codex-structured-conversation-stop: a compaction now goes through
  adapter.compact with the command run the host wrote (#23059), not a
  bare turn id, and answers with the provider's receipt. With the command
  claimed, a Stop that names no turn while the compaction's provider turn
  has not opened still interrupts nothing.
- main-agent-working-agreement: a provider row states its turn scope
  (#23059's appendItem contract); the retry and subagent rows are
  conversation-scoped.

* fix(native-chat): typecheck main's Stop and restore-grouping code against #23059

A Stop's compaction interrupt reads the narrowed requested turn, and the
restore-grouping test states whether each row reports its turn's outcome.

* fix(native-chat): say a /clear cut off by a restart left the chat unchanged

A /clear retried under the same operation after Orca restarted could not reuse the new conversation its first try started, and its row said "Codex couldn't start. Run /clear again." The agent did not fail to start: the earlier try was cut off. The row now reads "This /clear didn't finish, so the chat is unchanged. Run /clear again to start fresh.", from a new clearUnfinished failure fact written through agentSessionFailureWords.

The clearUnconfirmed and conversationCommandUnconfirmed reasons stay, with their words, for older hosts that still send them.

* fix(native-chat): a retried /clear finishes onto the conversation its earlier try started

When an earlier try of the same /clear started its replacement conversation and a restart or the
idle sweep has since stopped it, the retry could not replay that settled start and reported the
chat unchanged. That replacement is a fresh conversation at rest, so the retry now commits onto it
and its first message starts its agent. A replacement whose start definitely failed still reads
that failure, and one Orca can't prove stopped still commits nothing. The clearUnfinished failure
kind this made unnecessary is removed.

* refactor(native-chat): stop recording that Codex acknowledged a rewind

A refused rewind recovery now settles as refused whether or not Codex acknowledged the revert,
so nothing reads providerApplied any more. Stop writing it and drop the hook that wrote it.
Records that still carry the field load as before; the schema ignores the extra key.

* fix(native-chat): a /clear retried under a new operation id finishes the same replacement

A /clear's replacement id came from the client's operation id, so a retry the client sent
under a fresh id started a second replacement and orphaned the first. The host now derives
it from this caller's oldest /clear since its last commit whose replacement start reached
the operation ledger, so any retry from that caller finishes the same replacement, including
after a restart. A /clear after a committed one starts a new replacement. Another caller's
/clear is refused only while such a replacement is running or not proven stopped. An older
client that resends the same operation id still lands on the same replacement.

* fix(native-chat): a /clear retry never repeats a failed start or waits on an unproven stop

A retry under a new operation id could pick an earlier try whose replacement start had already
failed, replay that failure and commit it again, so a user who had since signed in was told
they were still signed out. Such a try is now skipped, and the retry starts afresh.

Another window's /clear was refused while the first window's leftover replacement was merely
not proven stopped. Nothing but the first window's own retry would settle that, so the refusal
could last until its ledger row expired a day later. It now waits only on a replacement whose
agent is running.

* fix(native-chat): a /clear retry finishes only a replacement that started

A retry picked an earlier try whose replacement start never answered, because a crash left
that start unsettled. Replaying it could only repeat "couldn't start" or, with the old agent
unproven, refuse every /clear from that window. Only a start that succeeded left a
conversation to finish; any other try is skipped and the retry starts afresh.

* refactor(native-chat): a record's identity fields are built in one place

A created record and a founded one (a conversation no agent has run yet, at
rest) share who and where the agent is and how it launches. The founding
builder is used by the /clear commit that follows.

* feat(native-chat): the store commits a /clear and its new conversation in one write

commitConversationClear founds the at-rest replacement from the cleared
record's identity and writes the committed marker and tab move in the same
transaction, so neither can land without the other. It refuses to overwrite
an existing record under the replacement id.

* fix(native-chat): /clear starts nothing; the new chat's first message starts its agent

/clear used to start the new conversation's agent before it committed, so it
could fail on that start ("Run /clear again"), and a crash between the start
and the commit left a running conversation nothing pointed at. #23524 then
needed a ledger scan to find an earlier try's replacement, a nonce half of the
derived ids, a refusal of another window's /clear while a leftover agent ran,
and a check for a start that had already finished.

Now /clear opens the chat for writing (it no longer starts an at-rest chat's
agent either) and makes one store write: the at-rest replacement under a
random id, the committed marker, and the tab move. The first message in the
new chat starts its agent through the existing send and delivery path, fresh
because its handle chain is empty. A failed start shows on that message with
the typed failure and a Retry, and a conversation no agent ever ran now reads
"couldn't start" rather than "couldn't restart".

Deletes clearTryToFinish, otherCallersClearIsLive, the attach block and the
committed start-failure branch, and the tests of that retry machinery.

* test(native-chat): drop the /clear retry wording test; no start runs for a /clear now

* test(native-chat): another window and a phone read a /clear's replacement from the host

Both list the replacement the committed marker names, under the chat's tab,
and each one's session list shows it with nothing unread until its first
message runs. A reader that recomputed the id from the operation turns this
red.

* test(native-chat): a never-started replacement closes as settled

A worktree delete closes every chat in it and asks the user to force any it
cannot prove stopped. A replacement no agent has run is released, so its
close settles like any at-rest chat's.

* fix(native-chat): /clear settles an interrupted Codex rewind the way a send does

/clear moved from starting the chat's agent to only opening the
conversation. A Codex rewind cut off mid-way on a chat at rest can only be
settled by its agent, so /clear was refused as "rewind unconfirmed" every
time until the user happened to send a message. It now prepares like a send
or /compact: the agent starts only when such a rewind is in doubt.

* fix(native-chat): a chat whose agent is not running keeps its `/` commands

Claude reports its skills and project commands only from a running process,
and the host served the `/` menu only from the running agent. Now that
/clear starts nothing, the new chat's menu lost those entries until its
first message; a chat stopped by the idle sweep already did.

The host now keeps, in memory, the list a running agent last reported for
its launch (provider, host, workspace, account and launch arguments) and
serves it to a chat of the same launch whose agent is not running. A new
report replaces it; nothing is stored on disk, so a relaunch still shows
the short menu until the agent reports again, and no list is ever served
across accounts, workspaces or hosts.

* test(native-chat): queued drafts around /clear follow what a /clear now is

Three queue tests from #23726 are red on main 29c49aec31 itself:
- Two expected an older build's unconfirmed ("prepared") /clear to hold
  sends and drafts back. #23524 made such a clear inert, because it
  changed nothing; Send-now and the drain now go past it, like any send.
- One held /clear in flight by holding the new conversation's start,
  which /clear no longer makes. It now holds the one store write, and
  still sees a send refused and no draft left on either conversation.

* fix(native-chat): /clear opens the new conversation under its own lock

Carrying queued drafts after a /clear opened the new conversation while
holding only the old conversation's lock. Every other open runs under the
opened conversation's own lock, so that a concurrent reader cannot open a
second handle on the same transcript. The carry now opens it through the
same entry point everything else uses. A clear with no drafts still opens
nothing.

* test(native-chat): queued drafts reach a /clear replacement that never started

Under lazy start the new conversation has no agent when /clear answers.
Pin that the drafts have already moved there by then, on a queue paused
"cleared", and that the user's first message starts the agent and goes
ahead of them. Red with the carry removed.

* fix(native-chat): /clear stops the old agent before it records the clear

/clear wrote its marker first and the RPC handler stopped the old
conversation's agent afterwards. It now stops the agent first, under the
same lock, then writes the marker, so nothing the old agent does can land
after the clear. A write that fails leaves the chat usable: its next
message starts the agent again, as after an idle stop.

The stop releases the lease, which moves its fence, so the marker is
written at the fence the record holds after the stop.

* fix(native-chat): give every chat write press its own operation id

A client kept an operation id per method and payload across presses, so a later identical press replayed an earlier action instead of running: an option picked again stayed on the other one, a goal set again stayed cleared, and a second Stop or phone Stop did nothing. Every press now mints its own id, on desktop and phone, and no client keeps a retry id.

The host answers a harmless repeat from what the chat records: the same prompt answer again returns the resolution it holds, and a Cancel of a prompt already cancelled answers ok. A second Stop of the same turn while the first is on its way joins it.

* fix(native-chat): answer a /clear pressed again after it committed with that clear

Every press now carries its own operation id, so a /clear that reaches the host after this caller's /clear already committed (a double press, or a retype after a lost answer) no longer replays the first id. It started the cleared conversation's agent and then refused it with "This conversation has been cleared."

The host now answers such a /clear from the committed record: the same caller, on a conversation whose tab moved to its replacement, gets that clear's result, replacement included, before admission starts anything. No second clear runs. Another window, or a cleared conversation reopened from history, still reads "cleared".

* test(native-chat): scope the prompt-cancel tests' journal rows to the thread

* test(native-chat): give the repeated prompt Cancel its own host test file

* chore(reliability): drop the deleted mobile id-retention test from the gates that ran it

* fix(native-chat): a Claude chat at rest reads its `/` menu from Claude's folders

Claude reports its skills and custom commands only while it runs, so a
chat whose Claude was not running (right after /clear, after the idle
stop, or after a relaunch) offered only the built-in `/` menu. The last
commit on this branch kept the last reported list in memory, which could
not survive a relaunch and could only repeat what a running Claude had
said.

The host now reads that surface where Claude itself reads it, on the
host that runs the chat, without starting anything: the workspace's and
the account's custom command folders (every `*.md` below them, named by
path with `:` between folders), the skills Orca's existing skill
discovery finds for Claude there, and the built-in commands Orca knows
Claude has. A running Claude's own report still wins. One scan answers
for a workspace and account for 10 seconds; a scan that finds something
new is pushed to the panes showing those chats. A chat run by another
host is never answered from this host's folders. The in-memory list is
removed.

* fix(native-chat): a Claude chat at rest keeps /model, /effort, /clear and /compact in its menu

Once the host sends a `/` list for a chat, the menu shows that list
instead of its own. The at-rest list started from Claude's text-driven
commands, which is empty, so a Claude chat whose Claude was not running
lost /model, /effort, /clear and /compact from its menu. It now starts
from exactly the menu a chat at rest showed before, then adds what the
folders hold.

A scan that never answers also held the next one off for good, freezing
the menu until relaunch; one unanswered for 30 seconds is now given up
on and the next read scans again.

* fix(native-chat): an at-rest `/` scan keeps any newer answer and never piles up

Giving up on a scan after 30 seconds threw away every scan that took
longer than that, so a slow folder never updated the menu, and a folder
that stayed hung started another stuck walk every 30 seconds, each
holding one of Node's few file-system threads.

Each scan is now numbered and its answer is kept whenever it is newer
than the one already kept, however long it took; an older answer that
lands after a newer one changes nothing. A new scan starts beside an
overdue one, but never more than two run at once.

* fix(native-chat): a failed at-rest `/` scan no longer discards an older answer

A failed scan marked itself as the newest answer, so an older scan that
answered after it was ignored and the menu stayed without custom
commands until the next scan. A failure now only starts the 10-second
wait. Test pins that the wait runs from when a scan lands, a failed one
included.

* test(native-chat): open the at-rest command host on main's shared journal database

* test(native-chat): a repeated draft press under its own id is answered by the draft's state

* fix(native-chat): the host joins a second Stop of a turn it is still stopping

A client joins it itself only against a host older than this, which runs both
and writes a false 'already finished' row for the second.

* fix(mobile): keep the Stop's fence narrowed, and type the held card answer

* test(native-chat): open the repeat-press host on the shared journal database

* fix(native-chat): a second Stop of a turn an earlier Stop answered for adds no row

Read from the earlier Stop's note on that turn in the journal, so it holds
however late the second Stop lands and from whichever client, and across a
host restart. Replaces the in-memory join. The capability now says the host
answers a repeated Stop quietly; clients still join a Stop on its way only
for a host without it.

* fix(native-chat): a /compact or /clear pressed again is answered from the one it repeats

A /compact from the same caller joins its last /compact while that one waits
or runs, and is answered by it once it compacted with nothing sent since,
read from the journal, so a retry after a lost reply never compacts twice.
The same caller's second /clear waits for the first and is answered by the
committed clear. Another caller or another command is still refused.

* Revert "fix(native-chat): a /compact or /clear pressed again is answered from the one it repeats"

This reverts commit 25541b7fc0.

* fix(native-chat): a Stop writes its one note only when it stopped something, keyed by the turn

A Stop that cancelled nothing writes no row, and the note is keyed by the turn
it stopped, so another Stop of that turn rewrites it. The earlier-note lookup
goes. A /compact or /clear pressed again while one runs is refused as before,
and one pressed after it ended runs.

* chore: the repeated-Stop capability's comments say what it now means

* refactor(agent-session): the transaction queue opens the session store file

AgentSessionRecordStore.open hardened permissions, loaded the file, marked every
lease unreconciled, built the transaction queue and persisted a pending rewrite.
That is the queue's load lifecycle, and the queue already applies the same
unreconciled rule when it reloads an externally changed file. Move it to
AgentSessionStoreTransactionQueue.open beside fromLoadedStore, and drop the
exported wrapper that existed only for the store's open.

No behavior change; the store's public API is unchanged. Brings the record
store back under max-lines after commitConversationClear.

* test: open the store on the journal database and pass the close cause, where main's tests still used the old calls

#24006 moved the store into the journal database and #23684 added a test on the
old open call; main's close now takes a cause. Four tests catch up.

* refactor(native-chat): a provider's at-rest commands are one adapter member

The at-rest `/` surface's read and change listener travel together as
`atRestCommands`, which the Claude catalog already is, so the adapter types
stay within their line limit after main's growth.

* test: the startup-reconcile tab close passes the close cause main now requires

* test(native-chat): the command-start test's client mock knows the repeated-Stop capability

* fix(native-chat): a Stop keeps the fence it was pressed at while the client checks the host

The desktop's capability check before a named Stop could let the runtime move
underneath, so the Stop went out against the new one; it now carries the fence
read at the press.

* test(native-chat): a Stop keeps its pressed fence against either kind of host

* fix(native-chat): keep the repeated-Stop rules through #24235's session-ending Stop

A Stop that ends the provider's session stopped its turn, so it writes its note even when the
provider declined the interrupt; a repeated Cancel of a cancelled prompt is answered at the prompt
Cancel's new entry point too.
2026-10-01 10:09:21 -07:00
OrcaWinandm4air 6d1a97ef98 fix(ssh): launch the Windows relay outside sshd's job so standard users work (#24224)
* fix(ssh): launch the Windows relay outside sshd's job without WMI

Win32-OpenSSH kills a session's job on close but allows breakaway. relay.js
gains a one-shot launcher mode that starts the detached relay with
CREATE_BREAKAWAY_FROM_JOB through the staged process-tree addon, so a standard
user no longer needs a WMI Remote Enable grant. WMI stays as the fallback for a
relay without the addon, and a refusal there is named. The Windows SSH-host
lanes drop their WMI grant and assert the breakaway route and adoption.

* fix(ssh): find runtime holds without WMI on a standard-user Windows host

The store GC read held runtimes through Get-CimInstance Win32_Process, which
WMI refuses to a standard user's SSH logon, so the pass kept every runtime.
On a refusal it now reads this account's own process image paths through
Get-Process.

* build(relay): ship the Windows relay launcher addon in every desktop package

macOS and Linux packages carried Windows relays without windows-process-tree.node,
so a legacy-runtime relay they uploaded to a Windows SSH host could not launch
outside sshd's job and fell back to WMI, which a standard user is refused.

A reusable Windows job now compiles the x64 and arm64 addons once and uploads
them; release-cut, release-mac-build, and the hourly/daily/adhoc mac builds
download them before build:release and require both arches. Staging now rejects
a binary with the wrong PE machine, the ReadProcessMemory import, or no
spawnOutsideJob export, so a stale pre-launcher build cannot ship.

* ci(ssh): run the Windows SSH-host lanes when the relay process-tree build scripts change

The staging and gyp-rebuild scripts decide which windows-process-tree addon the
relay ships, so a change to either must re-prove the Windows host cells.

* test(ci): find the mac orcad-template download by artifact name

The release mac job now also downloads the relay Windows process-tree addons, so
the first download-artifact step is no longer the template's.

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 05:32:09 -07:00
8afa1db50c feat(ssh): rung B glibc 2.17 compat runtime; gate remote vault on host node:sqlite (#24148)
* feat(ssh): wire rung B to the glibc 2.17 compat runtime; gate rung C vault on full node:sqlite

- COMPAT_RELAY_RUNTIMES lists linux-x64-glibc217; rung B plans the compat slot and compat
  pinned Node when glibc is below 2.28 or rung A refused with libc_floor/missing_lib.
- The relay version folds the compat runtime's executable hash; refusals are cached per runtime.
- The orcad template stages an optional linux-x64-glibc217 target (base package + compat
  node-pty slot + compat runtime marker); the verifier and materializer accept it.
- node-pty slot loader falls back to the compat slot when the default slot is missing or
  needs a newer glibc.
- Runtime store GC keeps the compat pin beside the default one on every relay connect.
- hasNodeSqliteReaderApi (DatabaseSync + backup) gates relay session search and the relay
  OpenCode reader, which now names the host Node version in its unavailable reason; the SSH
  vault reader installs the compat Node on old-glibc hosts and uploads nothing when no
  pinned Node can run.
- Rung D: a remembered noexec reports home_noexec and never advises installing Node.

* fix(ssh): re-prove a replayed noexec after rung D so allowing exec recovers the host

* fix(ssh): keep the rung B compat runtime pinned in the relay-connect store GC

* test(ssh): mock deployment-target facts in the Windows OpenCode runtime tests

* ci(ssh): build the glibc 2.17 compat slot for the hostile-host matrix; CentOS 7 lands on rung B

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>
2026-10-01 05:32:05 -07:00
OrcaWinandm4air f9940d5354 ci(ssh): macOS SSH-host lane for the pinned relay; fix uploads under a symlinked root (#24179)
* test(ssh): upload a root reached through a symlinked parent

The upload-root realpath fix landed with #24180; this keeps macoshost's case
where the root is passed explicitly beneath a symlinked parent.

* ci(ssh): macOS hostile-host lane on a loopback user-level sshd

Adds local-sshd cells for darwin-arm64 (macos-14) and darwin-x64
(macos-15-intel): a non-root sshd on 127.0.0.1 logs in as the runner user
with SetEnv PATH=<shims>:/usr/bin:/bin:/usr/sbin:/sbin and an empty HOME, so
no rc file restores Homebrew. The driver asserts rung A, terminal echo,
cached runtime reuse, GC keeping the in-use runtime, no toolchain or xattr
calls, and that the SFTP-uploaded Node carries no quarantine and runs as
uploaded. Docker cells are unchanged; each machine runs only cells it can host.

* test(ssh): fail a hostile-host run that would skip every named or hostable cell

A cell named for the wrong OS or arch was silently skipped, so a macOS job on a
mismatched runner went green having deployed nothing. Named cells must now be
hostable here, and a gated run must select at least one cell.

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 04:43:46 -07:00
OrcaWinandm4air 0ad77ea2f7 ci(ssh): Windows SSH-host lanes (inbox + preview OpenSSH) for the pinned relay (#24180)
* ci(ssh): import the private Windows OpenSSH provisioning harness

Copied unchanged from origin/OrcaWin/np-windows-ssh-provider-diagnostic
(config/ci/windows-ssh-provider/preview-ssh/ at 1242f3c4c8, commits 78b3857a0d,
c24adccff0, 069aa7b38b): a private LocalSystem sshd service on 127.0.0.1 for a
dedicated standard user, either the Microsoft-signed Win32-OpenSSH
10.0.0.0p2-Preview ZIP (archive and every binary pinned by sha256) or the inbox
OpenSSH.Server capability binaries. The following commits extend it for the
pinned-Node relay host lanes.

* test(ssh): run hostile-host cells through a host-agnostic driver

The Docker matrix drove the relay deploy and inspected the container with
inline docker exec calls, so no other host could reuse it. Split it into:

- ssh-hostile-host-test-harness.ts: the deploy, ladder observation, terminal
  echo (per-shell probe), runtime reuse and GC-keeps-in-use assertions, now
  also capturing every command the deploy sent the host.
- ssh-hostile-host-observer.ts: how a driver inspects the host outside SSH;
  docker exec for containers, the local filesystem for a loopback host.
- a legacy_opt_out outcome: the ladder never runs and nothing enters the
  pinned store, whatever the host-Node path does.

Launched cells now also check the runtime's sha256 on the host and that every
slot file (the Windows bundled ConPTY pair included) landed in the relay dir.

* ci(ssh): Windows SSH-host lanes for the pinned-Node relay

Phase 2 exit gate, Windows half: the real deployAndLaunchRelay through a real
SshConnection against Win32-OpenSSH on 127.0.0.1, on windows-2022 (x64) and
windows-11-arm (arm64), for both the inbox OpenSSH.Server capability and the
Microsoft-signed 10.0.0.0p2-Preview release (ZIP; archive and each binary
pinned by sha256 and Authenticode, as in the imported harness).

Builds on the provisioning harness from
origin/OrcaWin/np-windows-ssh-provider-diagnostic (previous import commit):
- one private standard account per cell, so every cell starts from an empty
  runtime store;
- -HiddenTools: the private sshd service's own Environment carries a PATH
  without any machine PATH entry holding node/npm/compilers, led by logging
  .cmd shims; a session probe fails the job if node.exe still resolves;
- DefaultShell set per cell by invoke-pinned-relay-cells.ps1 and restored at
  cleanup (dispatch proven per cell via %COMSPEC%).

Cells (src/main/ssh/ssh-windows-host-cells.ts): pinned-cmd (stock sshd),
pinned-powershell (DefaultShell = Windows PowerShell) expect rung A on the
pinned node.exe with the relay self-test passing, terminal echo, runtime
reuse, GC keeping the in-use runtime, stage identity through node.exe and no
Add-Type in any decoded session command; legacy-opt-out expects the ladder
never to run and an untouched pinned store.

* fix(ssh-ci): tolerate absent-drive PATH entries and retry Windows userData teardown

Join-Path throws on a machine PATH entry naming a drive the runner lacks,
which would abort provisioning before any cell ran; the toolchain split now
probes with [IO.File]::Exists over [IO.Path]::Combine, and the self-test
covers an absent drive. The hostile-host harness removes its throwaway
userData with removeTreeSync so a transient Windows lock cannot fail the
lane's afterAll.

* fix(ssh-ci): stop the account list rebinding the typed -Accounts param

PowerShell variable names are case-insensitive, so $accounts=[List[hashtable]] assigned into
the [int]$Accounts parameter and every Windows host job died before provisioning. Rename the
list and make the provisioning self-test reject script-scope assignments that shadow a param.

* fix(ssh-ci): hide the host toolchain by ACL, since sessions ignore the service PATH

Win32-OpenSSH builds a session's PATH from the machine and user registry values, so the private
service's Environment never reached SSH sessions and host node.exe stayed visible. Deny the private
accounts the toolchain PATH directories, put the logging shims on each account's own PATH, and
record failing sshd and client log lines so a refused login is diagnosable from the receipt.

* fix(ssh): resolve the upload root before checking entries stay inside it

uploadDirectory compared each entry's realpath against the root as given, so a root reached
through a symlink, junction or Windows 8.3 short name (C:\Users\RUNNER~1 in TEMP) rejected every
entry as escaped and the pinned runtime upload never started.

* fix(ssh-ci): fail cells on a vitest failure and give each account its own keys file

The cells script read $LASTEXITCODE under the workflow's GetNewClosure callback, which sees a
stale captured copy, so failed cells reported exit 0 and the job passed. Read the global value.
Inbox sshd 8.1 checks authorized_keys with read_ok=0, refusing a file other accounts can read;
use one keys file per account via %u.

* fix(ssh-ci): keep the account name in inbox mode and surface the WMI launch gap

The inbox binary-verification loop reused $name, so later SSH and SFTP probes logged in as
'sftp-server.exe'. Before the cells run, probe whether a standard SSH user can call WMI
Win32_Process.Create (the Windows relay launch path); when refused, warn and grant the cell
accounts Remote Enable on root\cimv2 for the run so the remaining assertions execute.

* test(ssh): keep the first terminal session answering keepalives through GC

The hostile-host driver disposed the first session's multiplexer before the GC and reconnect
steps, so a slow Windows GC let the relay reap the silent owner as 'local' and the reconnect then
waited out the full owner grace. Keep the session live until the connection closes, as the app
does, and resend the terminal probe until the shell evaluates it: ConPTY PowerShell can drop
typeahead sent before its first prompt.

* ci(ssh): keep each cell's relay logs in the receipts

* fix(relay): detach an ended socket client as peer-closed before destroying it

The listener destroyed a socket on 'end' but detached its client only on 'close'. A relay write
in that window failed with 'Relay socket is closed', and the dispatcher closed the client as
'local', so its PTY owner kept the full 30s grace instead of the peer-closed floor and a quick
reconnect was refused. The Windows host lanes logged this race on the named-pipe endpoint.

* test(relay): drive the peer-end listener test with a real dispatcher instead of a cast stub

The stub was an unchecked 'as unknown as RelayDispatcher' that failed the changed-code casting gate.

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 04:03:07 -07:00
OrcaWinandm4air 554f7f4ce5 feat(packaging): ship the orcad server template in desktop builds (#24155)
* build(orcad): merge per-runner prebuild slot trees into one matrix

Each node-server lane builds only its own node-pty slot. Release CI needs
their union before `build:orcad-prebuilds --require-slots` and the
template build can run; merge-orcad-prebuilds.mjs verifies every lane's
files against its own manifest, refuses duplicate slots and mismatched
node-pty/N-API/Node-header builds, then writes one merged manifest.

* build(orcad): keep agent-browser out of the desktop deployment template

The template rides inside every desktop build (design D2). Seven ~10 MB
agent-browser binaries would be ~76 MB, more than the rest of the template;
design D2's package contents never listed it, and a slot without one
already reports no headless browser. ORCAD_OMIT_AGENT_BROWSER=1 skips the
copy; standalone build:orcad still includes it.

* feat(packaging): ship the orcad deployment template in desktop builds

Design D2: the server JS and every target's addons ship inside the app,
as out/relay does; the ~120 MB Node runtimes stay excluded and are
downloaded on demand. electron-builder copies out/orcad-template to
Resources/orcad-template on every desktop OS, which is the first path
materializeOrcadArtifact tries (process.resourcesPath).

Platform signing rewrites native bytes the template manifest hashes:
- macOS: the tree is signIgnored (codesign rejects its ELF/PE payloads);
  afterPack signs the darwin targets' Mach-O files with the app identity,
  as notarization requires, then reseals only those manifest entries.
- Windows: SignPath signs after packaging, so release CI reseals from the
  inner-signing list (packaged-orcad-template.cjs --reseal-signed).
Every other file must still match the build's hashes; afterPack verifies.

ORCA_REQUIRE_ORCAD_TEMPLATE=1 makes a missing template fail beforePack and
afterPack; without it a build ships none and SSH relays keep the legacy
path. verify-packaged-orcad-template.test.mjs's "unused, excluded"
contract is reversed on purpose.

* ci(release): build the orcad template from qualified lanes and package it

node-server-tests.yml becomes callable with a ref and build_template.
With build_template, each lane that owns a release slot (macOS, Windows,
the glibc 2.28 and Alpine lanes, and the glibc 2.17 compat lane) uploads
its qualified out/orcad-prebuilds, the Windows lane also uploads both
process-table addons, and desktop_template merges them, gates the full
matrix plus the compat slot with --require-slots, runs
build:orcad-template and uploads the orcad-template artifact.

release-cut calls it at the release tag beside the other gates. The
build and build-mac jobs wait for it, download it into out/orcad-template
(the mac workflow from the parent run), and require it via
ORCA_REQUIRE_ORCAD_TEMPLATE. The Windows signing staging skips the
template's Linux/macOS payloads, and a reseal step records SignPath's
bytes before the installer rebuild. A template-scoped concurrency group
keeps a release call and main's push runs from cancelling each other.

* test(orcad): keep the packaged-lookup imports clear of the compat-slot import edits

* ci(orcad): let a rerun lane replace its template artifacts

upload-artifact v4 refuses a second upload under an existing name in the same
run, so rerunning a flaky node-server lane during a release would fail at the
upload instead of re-qualifying the slot.

* ci(node-server): build the template's Windows addons before the lane switches to Node 18

The addon build script imports TypeScript, which Node 18 cannot load, so every
build_template run (release-cut included) failed on windows-2022.

* fix(build): ship the orcad template's shared node_modules

electron-builder's extraResources filter always drops the root node_modules of
a source directory, so packaged apps lost orcad-template/node_modules and the
afterPack verify failed. Copy it through its own resource entry.

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 04:01:26 -07:00
14d4bb2e2a fix(ssh): Windows hosts without Add-Type staging; runtime-store GC on Windows (#24149)
* fix(ssh): collect the pinned-Node runtime store on Windows hosts

Windows SSH hosts now run runtime-store GC instead of skipping it: one
PowerShell inventory reads .runtime-ref-node-<sha> and .runtime-node refs from
every version dir, and one Get-CimInstance Win32_Process query filtered on an
image path under runtimes\ adds process holds (never by image name; a failed
query keeps everything). Stale upload stages are swept with the same rule as
POSIX. Promotion and the post-upload hold check now take the store lock on
Windows too, and the lock's own commands run unwrapped there.

Windows relay version-dir liveness now honours .relay-pid (design D5): a live
PID answers ALIVE before any pipe is touched, a dead one (ESRCH) plus refusing
pipes is exited, anything else is unverifiable. The runtime probe adopts a
pinned node.exe an earlier vault reader left without a .verified marker after
running it.

* fix(ssh): Windows stage fencing and vault runtime go through the verified node.exe

Upload-stage file identity on Windows no longer compiles an Add-Type P/Invoke
helper when the relay runs on Orca's verified pinned node.exe: the stage
commands run a fixed fs.lstatSync(..., {bigint:true}) script through it. It
prints the legacy helper's vol:high:low lowercase hex, and identity files are
compared after normalising hex spelling, so old and new clients recover each
other's stages. Host-Node relays keep the legacy helper; the choice is
documented in windows-edr-posture.md.

The Windows OpenCode vault reader now installs the pinned runtime through
ensureRemoteOrcadNodeRuntime (official zip, host-side extraction, .verified,
store lock) instead of uploading a client-extracted node.exe, and the relay dir
gains a .runtime-ref-node-<sha> so store GC keeps the runtime the vault uses.

* test(ssh): run the Windows stage-identity and store-GC tests on the Windows lane

The legacy/node.exe identity compatibility test and the Win32_Process hold path
were gated to win32 but no CI lane ran them. Add both files to the Windows
package lane and a real running-node.exe hold test.

* test(ssh): tear down Windows-lane temp trees through removeTreeSync

* test(ssh): grant the store lock to the Windows OpenCode runtime setup test

The Windows promote now runs under runtimes/.store-lock, so the mocked host
must answer the lock's CreateNew step.

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: m4air <m4air@Mac.localdomain>
2026-10-01 03:25:42 -07:00
OrcaWinandm4air 6aed05471c ci(ssh): hostile-host matrix for the relay runtime ladder (#24146)
* fix(ssh): classify a musl host missing libstdc++ as missing_lib, not wrong_libc

musl's loader follows each missing-library line with one 'Error relocating ... symbol
not found' per unresolved symbol, and the relocation pattern was checked first. Check
missing libraries before relocation errors; the ld-linux/ld-musl interpreter case stays
wrong_libc.

* build(orcad): allow a partial deployment template for CI

build-orcad-template --targets a,b builds and verifies only the named slots, so a CI job
that can fill just the x64 Linux prebuild slots can still materialize rung A/C addons.
Without the flag every target is still built and verified.

* ci(ssh): hostile-host matrix for the relay runtime ladder

Drives the real client-side relay deploy against Docker sshd targets and asserts the
design D6 rung each lands on: Debian 10 and AlmaLinux 8 (glibc 2.28) and Alpine (musl)
on rung A; Alpine without libstdc++ refused missing_lib down to D; Ubuntu 22.04 with a
host Node 20 and a noexec home straight to D (home_noexec); CentOS 7 (glibc 2.17)
refused libc_floor at A and C, falling to a host-npm path with no Node; and a
no-egress Debian 10 still on rung A. Launched cells also prove the terminal echoes,
no npm or compiler ran, a second connect reuses the uploaded runtime, and runtime GC
keeps the in-use runtime while collecting an idle one.

New workflow ssh-hostile-hosts.yml runs on dispatch and on path-filtered PRs.

* test(ci): pin the hostile-host workflow to the headless-server builder images

The matrix builds its runtime slots in copies of the node-server lanes' Alpine
and manylinux images; this contract fails when NODE_RUNTIME_PIN or either
builder digest moves in one workflow and not the other.

* test(ssh): reconnect as the same client and retry a grace-held PTY owner in the hostile-host matrix

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 03:24:17 -07:00
Neil bd90da7a5b ci: share PR planning setup and reuse the static native cache (#24329) 2026-10-01 02:46:18 -07:00
OrcaWinandm4air 53fd2dea0b feat(ssh): relay runtime fallback ladder, telemetry and host runtime setting (#24133)
* feat(ssh): complete the relay runtime fallback ladder (D6 rungs B slot, C, D)

Rung C runs the relay on the host's Node >= 18 with Orca's prebuilt N-API
addons and no npm (addon-only probe mode). Rung B is a data-driven slot chosen
only when a compat runtime is listed. Rung D fails the connect with a
classified reason carried as a TerminalUnavailableCause. The ladder steps
down only on classified refusals; unanswered probes throw. The rung decision
is persisted per host keyed by (glibc, runtime hash, Orca major), and
ssh_remote_runtime_resolved reports it once per host per session.

* feat(settings): SSH host runtime choice (Auto | Orca-managed Node | Host Node)

* docs(telemetry): describe ssh_remote_runtime_resolved

* fix(ssh): let a passing rung C disprove a remembered noexec; allow glibc-less compat runtimes

A remembered rung A noexec was re-persisted even after rung C self-tested addons from the same
~/.orca-remote tree, so rung A stayed skipped until the key changed. Rung B's evaluator also could
never match a musl compat runtime.

* test(ssh): import node:fs once in the host-node addon test

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 02:08:26 -07:00
OrcaWinandm4air ddd4927a0b build(orcad): server node-pty slots at glibc 2.28, plus a glibc 2.17 compat slot (#24134)
* build(orcad): build server glibc slots on glibc 2.28 and add the glibc 2.17 compat slot

Design D6: the default linux-{x64,arm64}-glibc node-pty slots now build in
manylinux_2_28 (digest-pinned) and are gated at glibc 2.28 / GLIBCXX_3.4.25
through a floor profile on verify-linux-glibc-floor.cjs; the desktop keeps
its Ubuntu 20.04 (2.31) default.

Adds the opt-in linux-x64-glibc217 compat target: NODE_RUNTIME_COMPAT_ASSETS
pins the unofficial glibc-217 Node (update/check pin scripts cover it, outside
SERVER_TARGETS), and a new CI lane builds the compat slot in manylinux2014
with static libstdc++, gates it at glibc 2.17 with no shared C++ runtime in
DT_NEEDED, and smokes it under the glibc-217 Node.

* refactor(node-runtime-pin): route compat lookups through isCompatServerTarget; keep the glibc doc's slot-name paragraph intact

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 01:10:57 -07:00
OrcaWinandm4air 8c2cd7d331 feat(ai-vault): read remote OpenCode history with the pinned Node; remove Bun (#24128)
* feat(ai-vault): read OpenCode history with the pinned Node instead of Bun

SSH hosts whose Node lacks node:sqlite (or its backup(), which 22.13-22.15
omit) now get the pinned Node in the shared ~/.orca-remote/runtimes/node-<sha>
store orcad uses: POSIX hosts receive the official archive and extract and
hash-verify it on the host; Windows hosts receive the verified node.exe the
client extracted, promoted by host Node with the same hash check. WSL distros
use the same layout and checks under ~/.cache/orca/runtimes/.

The Bun release pin table and its materializer are deleted. Old relays keep
reading their vault-sqlite/<sha>/bun references; nothing deletes those files.
An unconfirmed runtime upload now keeps its stage instead of removing it.

* refactor(sqlite): drop the Bun SQLite adapter; node:sqlite is the only backend

Nothing outside Electron runs on Bun any more (design D4), so SyncDatabase
loses its Bun branch, and bun-sqlite-database, bun-sqlite-statement and
bun-readonly-wal go, with the relay's bun:sqlite external. The profile-state
backup worker admits Electron or an entry that exists, and startup errors
name the pinned Node. The D7 cross-runtime gate still runs Bun 1.4.2, now
reaching Bun's SQLite through its node:sqlite.

* test(native-chat): drop the Bun SQLite driver case now that node:sqlite is the only backend

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 01:10:31 -07:00
OrcaWinandm4air 6593d7d194 feat(orcad): run orcad on the pinned Node instead of Bun (#24110)
* ci(daemon): gate PRs on daemon protocol crossing from the newest release

Lands daemon-protocol-facts.mjs from the Windows update diagnostic branch with a
stricter parser, and adds check-daemon-protocol-crossing.mjs (rule R1): the working
tree must attach the newest release tag's daemon. Rollback crossing is reported only.
Runs in the cross-version-wire job, which already has full tags; tag selection moves
to config/scripts/stable-release-tags.mjs so both use one rule.

* feat(persistence): run profile backups in the worker whenever its entry is bundled

* refactor(orcad): make profile and native preflight runtime-neutral

The profile preflight parser now takes the expected runtime identity from the
caller (shipped callers pass the pinned Bun identity), and the native
preflight is renamed to orcad-runtime-native-preflight with neutral wording.

* feat(runtime): pin the Node 24.21.0 server runtime with an offline CI check

Add src/shared/node-runtime-pin.ts (NODE_RUNTIME_PIN, SERVER_TARGETS,
NODE_RUNTIME_ASSETS for all 8 server targets plus the headers tarball),
generated by config/scripts/update-node-runtime-pin.mjs from the nodejs.org
and unofficial-builds SHASUMS. check-node-runtime-pin.mjs verifies, with no
network, that the pin tracks the locked Electron, matches engines.node's
major, and covers exactly SERVER_TARGETS; it runs in the static analysis job.

ORCAD_BUN_TARGETS consumers now read SERVER_TARGETS so there is one target
list; orcad's Bun runtime and build output are unchanged.

* test(persistence): skip plain-Node backup selection tests in the Bun profile suite

* fix(runtime): reject a pinned archive that belongs to another target

* ci(daemon): fail PRs that swap a runtime launcher and bump the daemon protocol

D7.1 R3: hosting orcad or the daemon on another runtime is not a protocol change,
so one PR must not do both. The launcher file list lives in the check script; the
allow-runtime-launcher-protocol-bump label overrides it.

* feat(orcad): select pinned-Node slots by a .runtime-node marker

D7.1 R5: a Node slot names its shared runtimes/node-<sha256>/node through
.runtime-node instead of .build-target, so Bun-era clients read it as a legacy
slot rather than exiting 78 on a missing bun-runtime. Nothing builds the marker yet.

* fix(runtime): load the Node pin without the typeless-module warning

check-node-runtime-pin.mjs now requires the pin and takes nodeDistArchiveName from
its own module, so it no longer loads the update script's build graph.

* fix(orcad): resolve Node slots to the design's runtimes/node-<sha>/bin/node layout

* feat(orcad): 8-slot node-pty prebuilds against the pinned Node headers at N-API 8

- build-orcad-prebuilds.mjs adds win32-x64/arm64 (conpty.node, the vendored
  conpty.dll/OpenConsole.exe, upstream's N-API conpty_console_list.node), compiles
  in a scratch copy against the hash-verified pinned headers (node.lib pinned per
  Windows arch) with NAPI_VERSION=8, rejects post-8 node_api_* imports, and writes
  a schema 2 manifest with per-file sha256, N-API level and the glibc need.
- --require-slots [slots] verifies files against hashes; --smoke loads the slot
  under the pinned Node and spawns a PTY; --print-slot names the host slot.
- The slot installer gates on N-API, libc, arch, glibc and file hashes instead of
  the exact NODE_MODULE_VERSION, and installs nested files (conpty/).
- bun-profile-tests.yml builds, verifies and smokes each runner's slot.

* fix(orcad): scope node-pty's glibc .symver pins to glibc on musl prebuild slots

musl's unversioned libc cannot satisfy openpty@GLIBC_* references at link
time, so the Alpine slot compile would fail. Pin the staged pty.cc guard to
__GLIBC__ and assert both musl transforms against the installed patch.

* feat(orcad): run orcad on the pinned Node instead of Bun

A packaged orcad slot now references the pinned Node 24.21.0 by its
executableSha256 (`.runtime-node`, `.server-target`) instead of carrying
bun-runtime, and ships node-pty from the slot's prebuild, only its own
ripgrep, and no Windows Bun PTY gate. The runtime lives beside the slots
at runtimes/node-<sha>/bin/node (node.exe on Windows, upstream name).

- build:orcad (build-orcad-node.mjs) builds the host slot's prebuild when
  missing and places the pinned runtime; the template is schema 3 with
  per-target files.
- handoffToBundledOrcad() resolves the slot's runtime reference and checks
  process.versions.node against the pin; a host Node >= 18 still hands off.
  Startup preflight keys on running as that runtime; callers expect 'node'.
- orcad and its daemon use node-pty (ConPTY + windows-pty-job on Windows);
  the Bun PTY sources, gate entry and canUseBunPty branches are removed.
- SSH deploy uploads the official archive once per pin, extracts and
  hash-checks it on the host, and self-tests it before publishing. Bun
  slots stay launchable for rollback; Node slots never use host Node.
- The runtime materializer is generic over pinned assets; the Bun wrapper
  remains only for the OpenCode vault reader (design Phase 2).
- Cross-runtime test: a profile DB written by Bun 1.4.2 (WAL left by
  SIGKILL) opens and backs up under the pinned Node, and the reverse.

No daemon PROTOCOL_VERSION change (design D7.1 R3).

* docs(ci): name the headless lanes after the pinned Node, drop Bun shard timings

Design D10: ci-demand-rollout.md and ci-runner-efficiency.md follow the
bun-profile-tests.yml -> node-server-tests.yml rename; shard timings drop the
deleted Bun PTY tests and follow the renamed ones.

* chore(ci): count the runtime archive download as a runtime launcher path

* fix(orcad): pin the macOS C++ standard for node-pty prebuilds

The official Node headers' config.gypi sets clang: 0, so common.gypi skips its
gnu++20 xcode_settings and Apple clang 15 (macos-14 runners) compiles
node-addon-api as C++98.

* fix(orcad): resolve the preflight's slot through realpath, as the handoff does

A symlinked orcad.js handed off to its real slot's pinned Node, but the
startup and profile preflights read the symlink's directory, found no
runtime marker there, and silently skipped the readiness check.

* refactor(ssh): drop materializeCachedNodeRuntime, which nothing calls

Deploys upload the verified official archive (design D5); no client path
needs an extracted Node executable cached by digest.

* test(orcad): gate the Bun-to-Node upgrade and Node-to-Bun rollback with live terminals

Design D7.1 R1/R3/R4 and D7.2. The last Bun orcad and this checkout's Node
slot are installed side by side under ~/.orca-remote, launched and stopped
with the client's own deploy commands, and share one data root. Each
direction proves the incoming orcad adopts the outgoing runtime's daemon
(same PID, same shell, output continues), opens its profile database and
backs it up with its own shipped worker, and that GC keeps the slot the
live daemon was forked from.

The node-server Linux lanes provide Bun 1.4.2 and build that Bun orcad
from main, and run with --cross-runtime. --artifact and --cross-runtime
now make their tests fail on a missing input instead of skipping.

* ci(node-server): pin node:24.21.0-alpine by its multi-arch index digest

* test(ssh): name the runtime archive fixture after its role

* test(node-server): load node-pty from the packaged slot in artifact runs

The node-server lane installs dependencies without building node-pty, and
Linux has no upstream prebuild, so the real-PTY failed-I/O teardown test
(picked up by the pty-subprocess selector) could not load pty.node. In
--artifact runs, alias node-pty to out/orcad's shipped slot so the test
exercises the addon orcad actually runs under the pinned Node.

* fix(orcad): let the Windows profile preflight exit after its PTY probe

On Windows, node-pty keeps the conout worker thread and pseudoconsole alive
until kill(), even after the shell exits. The PTY health probe never killed a
cleanly exited probe, so the packaged preflight printed its readiness line
and then hung until the build's 30s timeout, reported with an empty stderr.

- The probe kills its PTY on Windows after exit and uses the bundled ConPTY
  the daemon spawns with.
- The preflight exits once stdout is flushed; its owner reads to EOF.
- Preflight failures now report code, signal, timeout, stdout and stderr.

* test(node-server): load the slot's node-pty in the real-PTY test, not by alias

A vite alias redirected only ESM imports of node-pty; windows-pty-job and
local-pty-utils resolve it through require, so Windows loaded two conpty.node
copies and the Git Bash job-membership proof read an empty job. The failed-I/O
teardown test now loads node-pty through a fixture that picks the packaged slot
in artifact lanes.

The pty-subprocess selector was a prefix that also pulled in its POSIX-host
sibling unit tests, which pr.yml runs and which were never qualified on
Windows. Select the directory plus the two sibling files that belong here.

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 00:39:00 -07:00
OrcaWinandm4air d2dfc79764 ci(daemon): runtime-launcher protocol ratchet and Node slot marker (#24108)
* ci(daemon): gate PRs on daemon protocol crossing from the newest release

Lands daemon-protocol-facts.mjs from the Windows update diagnostic branch with a
stricter parser, and adds check-daemon-protocol-crossing.mjs (rule R1): the working
tree must attach the newest release tag's daemon. Rollback crossing is reported only.
Runs in the cross-version-wire job, which already has full tags; tag selection moves
to config/scripts/stable-release-tags.mjs so both use one rule.

* feat(persistence): run profile backups in the worker whenever its entry is bundled

* refactor(orcad): make profile and native preflight runtime-neutral

The profile preflight parser now takes the expected runtime identity from the
caller (shipped callers pass the pinned Bun identity), and the native
preflight is renamed to orcad-runtime-native-preflight with neutral wording.

* feat(runtime): pin the Node 24.21.0 server runtime with an offline CI check

Add src/shared/node-runtime-pin.ts (NODE_RUNTIME_PIN, SERVER_TARGETS,
NODE_RUNTIME_ASSETS for all 8 server targets plus the headers tarball),
generated by config/scripts/update-node-runtime-pin.mjs from the nodejs.org
and unofficial-builds SHASUMS. check-node-runtime-pin.mjs verifies, with no
network, that the pin tracks the locked Electron, matches engines.node's
major, and covers exactly SERVER_TARGETS; it runs in the static analysis job.

ORCAD_BUN_TARGETS consumers now read SERVER_TARGETS so there is one target
list; orcad's Bun runtime and build output are unchanged.

* test(persistence): skip plain-Node backup selection tests in the Bun profile suite

* fix(runtime): reject a pinned archive that belongs to another target

* ci(daemon): fail PRs that swap a runtime launcher and bump the daemon protocol

D7.1 R3: hosting orcad or the daemon on another runtime is not a protocol change,
so one PR must not do both. The launcher file list lives in the check script; the
allow-runtime-launcher-protocol-bump label overrides it.

* feat(orcad): select pinned-Node slots by a .runtime-node marker

D7.1 R5: a Node slot names its shared runtimes/node-<sha256>/node through
.runtime-node instead of .build-target, so Bun-era clients read it as a legacy
slot rather than exiting 78 on a missing bun-runtime. Nothing builds the marker yet.

* fix(runtime): load the Node pin without the typeless-module warning

check-node-runtime-pin.mjs now requires the pin and takes nodeDistArchiveName from
its own module, so it no longer loads the update script's build graph.

* fix(orcad): resolve Node slots to the design's runtimes/node-<sha>/bin/node layout

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-10-01 00:10:57 -07:00
OrcaWinandm4air 2a83c9536f ci(daemon): gate PRs on daemon protocol crossing from the newest release (#24089)
Lands daemon-protocol-facts.mjs from the Windows update diagnostic branch with a
stricter parser, and adds check-daemon-protocol-crossing.mjs (rule R1): the working
tree must attach the newest release tag's daemon. Rollback crossing is reported only.
Runs in the cross-version-wire job, which already has full tags; tag selection moves
to config/scripts/stable-release-tags.mjs so both use one rule.

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-30 23:23:28 -07:00
OrcaWinandm4air 49a83deaef refactor(orcad): make profile backup and preflight runtime-neutral (#24088)
* feat(persistence): run profile backups in the worker whenever its entry is bundled

* refactor(orcad): make profile and native preflight runtime-neutral

The profile preflight parser now takes the expected runtime identity from the
caller (shipped callers pass the pinned Bun identity), and the native
preflight is renamed to orcad-runtime-native-preflight with neutral wording.

* test(persistence): skip plain-Node backup selection tests in the Bun profile suite

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-30 23:23:23 -07:00
OrcaWinandm4air 3135fbbf49 feat(runtime): pin the Node 24.21.0 server runtime with an offline CI check (#24087)
* feat(runtime): pin the Node 24.21.0 server runtime with an offline CI check

Add src/shared/node-runtime-pin.ts (NODE_RUNTIME_PIN, SERVER_TARGETS,
NODE_RUNTIME_ASSETS for all 8 server targets plus the headers tarball),
generated by config/scripts/update-node-runtime-pin.mjs from the nodejs.org
and unofficial-builds SHASUMS. check-node-runtime-pin.mjs verifies, with no
network, that the pin tracks the locked Electron, matches engines.node's
major, and covers exactly SERVER_TARGETS; it runs in the static analysis job.

ORCAD_BUN_TARGETS consumers now read SERVER_TARGETS so there is one target
list; orcad's Bun runtime and build output are unchanged.

* fix(runtime): reject a pinned archive that belongs to another target

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-30 22:57:10 -07:00
OrcaWinandm4air 3fe4b18dae fix(ai-vault): require node:sqlite backup support in remote SQLite probes (#24086)
* fix(ai-vault): require the full SyncDatabase node:sqlite surface in host SQLite probes

The SSH and WSL OpenCode probes admitted any Node with DatabaseSync, so
Node 22.13-22.15 hosts (no backup export) skipped the pinned-runtime
fallback. Share one admission predicate with isSqliteAvailable() and
embed its source in both probe scripts.

* build(cli): list the node:sqlite admission predicate in the CLI project

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-30 22:57:01 -07:00
Jinwoo Hong 1762a138f7 feat(mobile): slide the page's host stack on push and Back (#24268)
expo-router's Stack on web renders native-stack's web view, which flips display and ignores animation. The page's host stack now keeps expo-router's StackRouter under its public Navigator and draws the slide with the Web Animations API; a popped screen stays mounted until it has slid out. Native is a pure move.
2026-10-01 01:19:35 -04:00
Jinwoo Hong e9ec63168f fix(mobile): hold-to-dictate, repeat keys and the browser long-press survive the page's long-press (#24277)
On the OTA page a held press died ~500 ms in: the WebView's long-press selected nearby text and that selection's selectionchange/touchcancel ended the press. Page text is now unselectable unless it opts in (as native), hold surfaces declare onLongPress, the browser pane refuses contextmenu termination, and the chat mic's swapped icons no longer steal the touch target.
2026-10-01 01:18:24 -04:00
Brennan Benson 12b8ef8c0b fix(worktree): update local main safely, once per branch, alongside the checkout (#23698)
* fix(worktree): retry local main refresh through git lock contention and skip false alarms

* fix(worktree): overlap the local main refresh with the checkout and run one refresh per repo at a time

* fix(worktree): skip the local base refresh when the create makes that branch itself

Creating a workspace named feature-x from origin/feature-x runs `worktree add -b feature-x`,
which now overlaps the refresh. The refresh's drift probe could see refs/heads/feature-x
missing and its presence probe then see it (the add just wrote it), which reported
"not fast-forward" and showed a sticky "Local feature-x was not refreshed" warning.
`-b` refuses an existing branch, so there is nothing to refresh in that case: skip it on
the local, prepared-checkout and SSH create paths.

The SSH overlap tests move to their own file so the existing suite stays under the line limit.

* fix(worktree): say plainly what happens after the local base refresh queue wait expires

* test(worktree): prove SSH local base refreshes of one repo run one at a time

* test(worktree): drop type assertions from the SSH refresh overlap test mocks

* fix(worktree): fast-forward local main with one host-owned merge --ff-only per branch

Moves the whole local base refresh into one shared routine that runs on the
execution host (main process for local and WSL repos, the relay for SSH), so
the app no longer keeps a second copy of the checks, queue and retry.

A checked-out branch now moves with merge --ff-only (hooks, auto-gc and
autostash off) instead of status-then-reset --hard, which silently overwrote an
untracked file the new commit adds and could discard an edit or a commit made
after the check. A free branch moves with a compare-and-swap update-ref that
writes a reflog message. Status reads no longer take index.lock.

Creates of one branch share one run plus at most one trailing run; a create
waits at most 30 s and never starts a competing mutation. The failure toast is
keyed by repo and branch because every create that joined a run reports the
same fact.

* fix(worktree): fast-forward local main even when the repo requires signed merges

With merge.verifySignatures=true, the owner-checkout fast-forward refused an
unsigned origin/main tip, so every create warned "Local main was not
refreshed" where the old reset moved main. The new workspace is already
created from that same unsigned commit, and a branch that is not checked out
moves without a signature check, so the refusal protected nothing. Turn the
setting off for this one merge, like the hooks, gc and autostash overrides.

* fix(worktree): clear git read caches when a shared local main update lands late

The update of local main can finish after a create stopped waiting for it, so
the shared run now invalidates git read caches itself. The index.lock real-git
test also no longer reads the developer's global git config.

* fix(worktree): never overwrite an ignored file when fast-forwarding local main

A plain `git merge --ff-only` silently replaces an ignored file (for example a
local `.env`) at a path the new commit starts tracking. Pass
`--no-overwrite-ignore` so git refuses instead and the create reports the
checkout as having local changes. Supported on the fast-forward path since
well before Git 2.25.

Also make the relay test for one-refresh-per-branch hold the first merge until
the second request has reached the relay, so it fails without the coalescing.

* fix(worktree): keep the local main update a plain fast-forward whatever the user's merge settings say

A per-branch mergeOptions such as '-s ours' or '--squash', or pull.twohead=ours,
made the update create a merge commit that dropped upstream, or stage upstream
without moving main, while reporting success. The command now clears the
branch's mergeOptions and passes the strategy and signature choice on the
command line, which beats any config. After the move Orca confirms local main
is exactly the target before reporting it updated. The exact command also runs
in the Git 2.25 compatibility suite.

* fix(worktree): make the Git 2.25 fast-forward contract pass in CI and rerun on every change to it

The new real-Git contract for the local main fast-forward wrote a post-merge hook into .git/hooks, which does not exist when the repo is created by the uninstalled Git 2.25.5 build CI uses (no templates), so the Git compatibility check failed. Create the directory first.

The Git compatibility check also did not run when only the fast-forward module changed, so a later edit to its merge arguments (for example a flag Git 2.25 lacks) would skip the one check that tests them. Add the module to the check's paths.

* fix(worktree): answer every create from a local main update toward its own base

Creates from different remotes' main (origin/main and upstream/main) shared one queued update
per repo and branch, which ran only the latest caller's target: a create could get no result
for its own base, or a false "not refreshed" warning computed for another remote's main.

The per-branch runner now queues one run per distinct target, still one at a time per branch,
and only callers toward the same target share a queued run. Applied in the app and the relay.

* test(worktree): record the third create's result in the mixed-remote burst tests and update the toast id rationale

* chore(worktree): correct the toast id rationale
2026-09-30 18:30:22 -07:00
Brennan Benson 24540300f0 fix(agent-status): preserve hook presence when process checks cannot answer (step 1 of 3) (#23947)
* fix(agent-status): admit hook process presence on the execution host

* fix(agent-status): restrict process checks to real hook ingress

* fix(agent-status): keep presence checks from causing false exits or losing real ones

- Pin the macOS process start time to UTC on both the hook and the host so a
  shell TZ or a time-zone change cannot turn a live Claude into an exit.
- An unanswered process check falls back to the foreground confirmation, so
  Codex, SSH and Windows panes still leave the agent state on a real exit and
  the Codex late-completion recovery still runs.
- A nested agent that inherits the pane key cannot take over the pane's
  presence; a retired session is replaced by the next session even when its
  SessionStart was lost.
- Drop the unused Windows process read (no Windows hook captures an identity
  yet) so shared code no longer imports main-process modules.
- Skip the capture outside Orca panes, gate relay re-checks to real title
  changes, and list the capture module in the CLI project.

* refactor(agent-status): own pane presence by the agent's process, not its session

Presence now exists only when a hook carries the agent's process identity,
and only that process's evidence changes it: its SessionEnd ends the pane,
its /clear and /resume keep it, and hooks from any other process (a nested
agent inheriting the pane key) update status without taking ownership.
Hooks without an identity (Windows, sessions started before the capture)
behave exactly as before, so a nested agent can no longer end a pane it
does not own. An ended owner stops answering 'exited', and the host probes
the owner only when another process reports in the pane.

* fix(agent-status): close round-3 review gaps in presence handling

- Relay retries and transcript polls schedule against the row the relay
  cached, and identity-less events pass through the transition unchanged,
  so SSH Grok replies and Codex transcript polls deliver again.
- A live process check restores the runtime's agent status and releases
  queued orchestration mail, like a foreground read that finds the agent.
- An answered foreground read naming a non-agent (wsl.exe, tmux) is still
  an exit; only silence is not.
- A suspended (Ctrl-Z) agent is unverifiable, not live, so the foreground
  read decides as before.
- An agent of another type started mid-turn cannot own or end the pane.
- Replayed spool hooks check each pane once, and SessionEnd ends presence
  only for reasons that end the process.

* refactor(agent-status): record the pane's owning agent even without a process id

Ownership is now decided only by comparing the recorded owner with the
sender, never from the row's agent type or turn state (which identity
resolution rewrites). The first agent hook in an empty pane claims it; a
live owner keeps it against any other agent (a nested claude -p, a Codex
started inside Claude or the reverse); only the owner's own proven process
can end it. An owner no hook identified never ends from a hook and is not
probed, so those panes behave as before.

* test(agent-status): read the optional process id in the relay presence test

* fix(agent-status): other agents' SessionEnd hooks settle their status again

Only an admitted exit (Claude's process-ending SessionEnd, or a host-proved
exit) is marked ended on the event, and ownership keys on that marker, not
the hook name. Devin, Qoder, CodeBuddy and Copilot SessionEnd hooks are
ordinary status updates again, locally and through the relay.

* test(agent-status): cover the owner's own SessionEnd through the relay

* fix(agent-status): rows without a process identity keep today's command-finished cleanup

The renderer's command-finished cleanup kept every row when its shell check
could not answer, which left Codex, hookless-agent and old-relay rows over
SSH showing done after a real exit. It now asks the host whether the pane's
agent process can be checked: only a pane with an identified, running owner
keeps its row on an unanswered check; every other pane drops it exactly as
before. The drop stays armed while the host answers, so a new command still
cancels it, and a missing or failing answer (web client, older host) keeps
today's behaviour.

* fix(agent-status): panes without an identified owner keep today's exit confirmation

The title-driven exit confirmation applied 'silence is never an exit' to
every pane. It now applies only when the pane has an identified owner whose
process cannot be checked right now; a pane with no process identity
(Codex, hookless agents, Windows, old relays, sessions started before the
update, or an owner that already ended) confirms exits exactly as before.
2026-09-30 18:21:46 -07:00
Jinwoo Hong 2ab179ccf7 fix(emulator): take serve-sim 0.1.47 so the iOS simulator works on Xcode 27 (#24228)
* fix(emulator): take serve-sim 0.1.47 so the iOS simulator works on Xcode 27

serve-sim 0.1.40's helper binary hard-linked SimulatorKit at its pre-Xcode 27
path, so every iOS emulator start failed with a dyld error on Xcode 27.
0.1.47 replaces that binary with a napi addon that finds SimulatorKit in
either location, and runs the stream helper as a node process.

Match the new helper process (serve-sim.js ... --exit-on-simulator-shutdown)
while still matching legacy serve-sim-bin helpers left by an older Orca, and
mark the package's new standalone helpers executable.

* refactor(emulator): drop redundant serve-sim chmod; the package already ships helpers executable

serve-sim publishes its helpers as 755 and pnpm, fs.cpSync and electron-builder
all preserve the mode, so the executables list and every chmod over it are dead.

* fix(emulator): give the materialized serve-sim runtime its node dependencies

serve-sim 0.1.47's entry imports `ws`. Packaged macOS builds run serve-sim from
a copy under userData, and that copy had no node_modules, so every iOS emulator
attach failed with ERR_MODULE_NOT_FOUND. Dev builds were unaffected because
they run serve-sim straight from pnpm's node_modules.

Lay the copy out as <version>/node_modules/serve-sim and link each dependency
the package declares to the bundle's installed sibling, so transitive deps keep
resolving from the bundle. A runtime in the old flat layout, or one whose links
dangle because the app moved, is rebuilt instead of reused.
2026-09-30 21:04:31 -04:00
Brennan Benson 6f2a7d05c9 fix(worktrees): let git delete removed checkouts so chat sends never wait behind them (#23837)
* fix(worktrees): delete removed checkouts in git, not in Orca's file pool

Local worktree removal renamed the checkout into a sibling trash root and
deleted it in the background with a recursive fs.rm in the main process.
That queued one request per entry on libuv's shared 4-thread file pool, so
for minutes every other async fs call in the main process (the agent-session
store behind chat sends, file explorer reads) waited behind the delete.

`git worktree remove` now deletes the checkout inline in git's own process
again, so the card stays in its Deleting state for the length of the delete
while Orca's file pool stays free. No timeout applies to the call, so a
large delete is never killed halfway.

If git reports success but the path still exists (Git for Windows leaves
junctions and their parent directories in place), the leftover is deleted
with the existing removeHostTree; WSL checkouts stay with the distro.

Nothing creates trash any more: the scheduling queue, rename/restore
helpers and the trash_rename span are gone. The startup sweep stays to
drain entries older releases left behind, and now removes each emptied
trash root so the obligation ends.

* fix(worktrees): let Git delete Windows checkouts with long paths enabled

Removal now always runs Git's own recursive delete, and worktree creation
checks out with core.longpaths on Windows, so a deep checkout Orca created
could fail to delete with "Filename too long" (#6433). The Windows recovery
then finishes the delete but keeps the branch. Pass the same command-scoped
core.longpaths option to `git worktree remove` so Git can delete what it
created.

Also point the CI shard timing entry at the renamed real-git removal suite.

* fix(worktrees): keep an inherited GIT_ASK_YESNO out of the worktree delete

Git for Windows asks $GIT_ASK_YESNO whether to retry when a file stays
locked during a recursive delete. Orca's git env inherits the user's
environment, so an inherited value would run an arbitrary prompt program
in the middle of a removal. Drop it for the removal call only.

* perf(worktrees): run worktree deletes under their own limit, outside git admission

`git worktree remove` now deletes the whole checkout in Git's own process,
which takes 20-35 s on a large tree. It took a general git admission slot at
status tier for that whole time, and that cap is as small as two slots on a
machine with six or fewer cores, so two deletes blocked every status read.

Deletes now skip general admission and queue under their own limit of two
per host instead: two concurrent deletes already saturate one disk, and more
only slow each other down. Leftover cleanup runs inside the same slot.

* fix(worktrees): delete removed checkouts in the background and mark them removing

Since the checkout is deleted by `git worktree remove` in Git's own process,
a large delete takes 20-35 s. Answering the request only after that made web
and mobile (30 s), paired desktop (60/180 s) and the CLI (60 s) report a
failure for a delete that was still going, and mobile silently re-showed the
row.

The request now does everything that can refuse (lock, cleanliness, archive
hook, watcher/terminal gate, terminal stop, shared-link unlink), records the
removal in an in-memory table on the host and answers `removing: true`. The
delete, branch cleanup and metadata purge run after it in the same order as
before, and the watcher/terminal gate stays held until they finish.

- Listings mark rows in the table `removing` for clients that advertise
  `worktree.background-removal.v1` (the desktop renderer, paired desktop and
  web), and leave them out for everyone else (older clients, mobile, the
  CLI), which already dropped the row when the request answered.
- The outcome (removed, with any preserved branch, or the error) rides the
  existing worktrees-changed event as an optional field, sent after the row
  has left the table.
- A repeat delete while Git runs joins it. A create at the same path or with
  the same branch is refused with "Cleanup is pending; try again shortly";
  create's name search skips the path, so generated names move on.
- Nothing is persisted: after a quit or crash Git still lists the checkout
  and it can be deleted again. WSL checkouts still delete inline.
- `orca worktree rm` says the checkout is still being deleted.

* fix(worktrees): keep the existing Deleting card until the host's Git finishes

The host now answers a local worktree delete on acceptance and deletes in the
background. The renderer keeps the existing delete state set until the host
publishes how it ended:

- The delete that asked waits for the outcome on the worktrees-changed event
  (local IPC or the paired runtime's client event), then runs the same
  teardown, preserved-branch toast and card error an inline delete did. If
  that event is lost to a dropped connection, a listing that shows the row
  gone after it was marked removing finishes the wait, and one that shows it
  back without the marker fails it.
- Any other renderer (a reload, a paired desktop, web) sets the same delete
  state from the host's `removing` marker and clears it when the marker goes.
  A failure the host publishes lands on that card's existing error.
- Web advertises `worktree.background-removal.v1` so the host sends it the
  marker; paired desktop does through the Electron capability list.

No new component, style or state: the card reads the delete state it always
did. A host that predates this answers when done without `removing`, and the
renderer takes that as finished, as before.

* test(worktrees): type the removal harness and projection for the node typecheck

* fix(worktrees): don't fail a delete retry with an earlier attempt's buffered failure

A background removal's outcome that reached this renderer with no waiter (another client's
delete, a host-marked card, or one already settled from listings) was buffered for 60 s and
consumed by the next delete of the same workspace, so retrying a failed delete failed at once
with the old error while the host was deleting. Drop the buffered outcome before sending the
request; only an outcome that arrives after it can belong to it.

* fix(worktrees): let only a gap in host events settle a background delete from listings

Git unlists the checkout before the host deletes the branch, cleans the push target and purges
metadata, and the worktree-directory watcher refetches within 250 ms. The renderer read the
missing row as a finished delete, so the waiter resolved without the preserved branch (no
toast) and a failure in those last steps showed as success; the real outcome was then dropped.
The listing fallback exists only for a lost outcome event, so it now applies only after this
host's event stream had a gap: a new subscription or a replay after reconnect.

* perf(worktrees): let a bulk delete start each same-repo checkout delete once the host accepts the last

A bulk delete ran one worktree at a time per repo (#2259, for packed-refs and ref-lock races in
branch cleanup). With Git now deleting each checkout for 20-35 s before the request settles, N
worktrees in one repo took N times that. The renderer now queues same-repo deletes only until
the host accepts each one; a parent still waits for its nested children to finish. The host
serializes the branch cleanup step per repo itself, which also covers removals started by
different clients.

* test(worktrees): pin the host platform in the mocked removal suites so they pass on Windows

Removal now passes -c core.longpaths=true on Windows, so the exact-argv
assertions and command-keyed mocks never matched there (17 failures on a
Windows host). Pin darwin as the add-worktree suites already do, and drive
the one Windows-specific case through the same spy.

* test(worktrees): type the blocked git remove result instead of a broad object

The anti-slop static-analysis gate rejects `object` parameters.

* test(worktrees): clear the changed-code quality gate in the removal suites

Merge the duplicate node:fs import, build the mock child without a cast, read
worktrees:list rows through one typed helper, and give the remaining casts a SAFETY line.

* fix(worktrees): record each background delete durably and finish it after a quit or crash

A quit mid-delete left git to finish the checkout on its own while the branch
delete and metadata purge never ran; a crash left a normal-looking row. Each
accepted local removal now writes a record beside the profile state before git
starts, clears it on success or failure, and the host runs the same delete
again for any record left at startup, re-deriving what remains from git and
disk. An orderly quit stops the checkout delete without waiting for it.

* test(worktrees): type the interrupted-removal assertions for the node typecheck

* fix(worktrees): finish an interrupted delete that already removed the checkout's .git file

Quit stops git worktree remove mid-delete, and Git deletes the checkout's .git
file wherever it falls in directory order. Git then refuses the checkout
("validation failed ... .git does not exist") on every retry, so the startup
finish failed and the row could never be deleted from Orca. A registered
checkout this record owns that has lost its .git file now finishes like an
unregistered one: leftover files, prune, then the branch.

* fix(worktrees): let Git finish an interrupted delete, and never take a different checkout

A quit or crash that stops `git worktree remove` after it deleted the checkout's
.git file left a registered checkout Git refuses to remove. The previous fix
deleted that leftover inside Orca's process, which is the bulk delete this
change exists to avoid (and on Windows the leftover can be most of the
checkout). The startup finish now rewrites the missing .git file from Git's
own admin entry for that path and lets `git worktree remove --force` delete
it. `git worktree repair` is not used: it also re-points every other
registered path, including a checkout another repository now owns there.
Orca deletes the leftover itself only when no admin entry claims the path.

The startup finish forces, so it now leaves the path alone when the checkout
there is not the one recorded: a registered worktree on a different branch or
head, or a `.git` at a path Git already unregistered. The record is dropped and
the card shows why.

The record write before Git starts is now bounded (2 s, logged when exceeded)
so a stalled disk cannot hold the delete, and the outcome is published before
the record's clear reaches disk.

* test(worktrees): compare worktree paths by value and tear down with Windows lock retries

Git prints forward slashes in `git worktree list` on Windows, so the real-Git
removal suites never found a joined path there: positive checks failed and
negative ones passed without proving anything. They now compare Git's parsed
rows by value. Teardown uses the shared retrying removeTree, since Windows can
hold the deleted checkout busy for a moment after Git exits. Adds a
relative-path worktree case for the .git restore (skipped before Git 2.48).

* fix(worktrees): reply to a worktree delete when it has finished, not on a broadcast event

A current client's delete request now waits for the host's background delete and gets its real
result (removed, a preserved branch, or the error) as the reply, the way it did before the delete
moved off the request. A request that arrives while the delete runs joins it and gets the same
result. Every other view keeps reading the host's `removing` marker: the row leaving means the
delete finished, and the row listed again without the marker shows "The delete did not finish.
Try again." on a card that view had marked Deleting. A request whose reply is lost (a timeout or a
dropped connection) settles the same way from a fresh listing instead of reporting a failure.

Clients without the background-removal capability (mobile, the CLI, older desktops) are still
answered on acceptance and have rows under removal left out of their listings.

This removes the outcome on worktreesChanged and everything it needed: the renderer's outcome
waiters, early-outcome buffer and TTL, per-host event-gap generations, the request pre-registration,
and the accept callback bulk delete used. Bulk delete runs same-repo deletes in parallel only on
this machine, whose host serializes branch cleanup per repo; SSH and paired hosts stay serialized.

* test(worktrees): type the pending-removal host id in the background-removal suite

* fix(worktrees): answer a delete request even when a concurrent removal of the same worktree replaced its record

The desktop app's removal and the runtime removal (CLI, paired clients) coalesce separately, so
both can be accepted for one worktree. The second replaced the first's record, and the first
delete then finished without resolving the request waiting on it, leaving the desktop card on
Deleting indefinitely. Each delete now settles the request it was started for.

* fix(worktrees): run same-repo removal archive hooks and teardown one at a time on the host

Local bulk delete now sends same-repo removals in parallel, so their archive hooks, terminal
teardown and preflight ran at once; a hook that writes refs can race the repo's ref locks
(#2259). The host now serializes each local removal up to acceptance per repo, for every
client; Git's checkout delete still runs in parallel under the delete limit.

* fix(runtime): keep waiting worktree deletes out of a host's foreground call slots

worktree.rm now replies only after Git deletes the checkout (up to minutes), so on paired
desktop and web each waiting delete held one of the host's 8 foreground call slots, and a
bulk delete queued listing refreshes and every other foreground call behind it. Deletes now
run in their own lane with the same bound; the 2-slot background lane stays for status polls.

* fix(worktrees): join a same-worktree delete accepted while a removal waited its repo turn

The desktop app and the runtime (CLI, paired clients, web) check for a running delete before
they queue for the repo's acceptance turn. A delete of the same worktree from the other path,
accepted while this one queued, was missed: this request re-ran the archive hook, stopped the
terminals again and started a second `git worktree remove` on the directory Git was deleting.
The queued acceptance now re-checks and joins the running delete.

* fix(worktrees): fence a resumed delete's checkout from startup, and drop rows a listing read before the delete finished

A delete a quit or crash interrupted took its terminal and file-watcher gate only when the resume
job ran, after the first window was shown; session restore could open a shell or watcher inside the
half-deleted checkout first, and on Windows that handle can fail the resumed git delete. Loading the
records now fences each recorded path, and the resumed job takes the fence over in the same tick it
takes its own gate.

A listing that read git's registration before a delete finished, and replied after the removal
record cleared, returned the row unmarked, so other views briefly showed "The delete did not
finish". Listings now capture the pending removals before reading git and leave out a row whose
delete finished successfully since; a row whose delete failed stays listed as before.

* test(worktrees): keep git's auto-maintenance out of the real-git removal suite

CI's Git 2.55 failed the file-pool test in teardown with ENOTEMPTY on the scratch repo's
objects/pack after the test body passed: the 3,000-file commit's detached auto-maintenance was
still writing a pack. The scratch repo now disables auto-maintenance and auto-gc.

* fix(worktrees): one archive-hook approval covers a same-repo bulk delete again

Local same-repo deletes now start together, so each queued its trust prompt with a state snapshot
taken before the first prompt was answered; approving the first still showed the same prompt once
per remaining worktree. The queued check now reads the store when its turn comes.
2026-09-30 16:32:20 -07:00
Brennan Benson cfa43e7eab fix(codex): opening a terminal no longer strips Codex hooks from the real ~/.codex (#23552)
* fix(codex): a real-home restore leaves a file alone once someone else changed it

Orca writes ~/.codex/hooks.json (and a trust rebase writes config.toml), then
runs a Codex trust session for up to 10 s, then restores the original bytes if
the session fails. The restore wrote unconditionally, so a save that landed
during the session, from the user or another Orca, was silently reverted.

Each restore now compares first: it writes the original back only while the
file still holds the generation Orca's mutation left, and otherwise logs and
leaves it alone. This covers the real-home install and opt-out sweep
(restoreRealHomeHooksJson), the legacy sweep's hooks restore, and config.toml
rollback (restoreCodexTrustConfig).

For hooks.json the generation is the exact bytes Orca wrote. For a config.toml
that a trust rebase changed it is the file as the rebase left it. When Codex
itself wrote config.toml inside the session that just failed, Orca never knew
those bytes, so that rollback compares against the file as the session settled.
The next commit keeps other Orca instances out of that window; a user edit made
during such a session can still be rolled back.

* fix(codex): serialize real-home Codex writes across Orca instances

Every Orca on one HOME (a dev and a packaged app, or an offline CLI) writes the
same ~/.codex/hooks.json, config.toml and ~/.orca/agent-hooks/codex-hook.sh.
The per-file lane that orders capture, mutate and restore was in-process only,
so another instance could write inside this one's restore window, or undo it.

The lane for the user's real config.toml now also holds the existing
crash-safe managed-hook install lock (~/.orca/managed-hook-install.lock, the
one relay installers take for the same home). It is taken only by the
outermost acquire, because the lock file is not reentrant and grants and trust
rebases nest inside an install. Managed-home installs, the real-home install
and opt-out sweep, and the legacy sweep all enter through it. Compare-and-swap
on restore stays as the backstop.

A lock that cannot be taken within its 10 s wait fails that install, which is
already best effort: launch prep logs it, and the real-home lane falls back to
the managed lane until its retry.

* fix(codex): opening a terminal no longer strips the shared Codex entry from ~/.codex

Every Orca instance on one HOME writes the same status-hook entry into the
user's ~/.codex/hooks.json, with its trust in config.toml. Launch prep runs on
every pane spawn, and under a managed Codex account it ran the legacy system
sweep. That sweep matched Orca entries by script file name, so it removed the
current shared entry and the trust blocks the grant ledger recorded. On a live
laptop hooks.json went 4139 -> 18 bytes about 150 ms before a new pane opened.
With hooks off, the real-home lane's launch prep swept the same way.

Now nothing automatic removes the current entry or its trust:
- The legacy sweep removes only an enumerated list of retired command forms
  that no build writes any more (#1019's double-quoted form, #1536's
  exec-guarded form, and Windows' per-userData bare path), plus their trust.
- ensureRealHomeCodexHookState with hooks off writes nothing; that covers
  launch prep, session resume and startup.
- Only the user's explicit opt-out (codexHookService.remove()) strips the entry
  and its ledger-recorded trust from the real home.
- The sweep-suppression gate existed only to stop the sweep from deleting the
  current entry, so it is deleted with its main-process wiring.

Startup with hooks off already skipped the real-home install; with this change
the first pane's launch prep with hooks off also leaves ~/.codex untouched.

* fix(codex): a pane's prepare-codex only repairs a home its own HOME's app installed

On macOS a pane starts through login(1), so it gets the user's real HOME even
when its Orca app runs with another one. The pane's `codex()` preflight
installed hooks in the CLI process with that real HOME: it rewrote
~/.orca/agent-hooks/codex-hook.sh, promoted trust into the real config.toml,
and wrote the real HOME's script path into the app's managed home.

The preflight now acts only when the managed home's hooks already run this
process's own shared script, which proves the app that installed them shares
its HOME. Otherwise it writes nothing; the app installed the home at spawn.

Why not a no-op: the preflight was added (#14326) because trust can go stale
between opening a pane and typing `codex`, for example in a pane that survives
an app update, and Codex then stops in hook review. For a same-HOME pane it
still repairs that. Why keep promotion: the install drops runtime trust the
system config does not back, so skipping promotion would delete approvals the
user gave inside Orca-launched Codex.

* test(agent-hooks): await every installer in the refresher coverage test

The test fired each managed installer without awaiting it and read
~/.orca/agent-hooks straight after. Codex's install now takes the
cross-process real-home lock before it writes its script, so the script
landed after the read. Await the installers, and stub Codex's trust sessions
so the awaited install cannot start a real `codex app-server`.

* fix(codex): retire the two real-home command forms the list missed

The real-home lane wrote two Codex hook forms into ~/.codex that no build
writes any more and that the enumerated retired list did not name:
- POSIX, #9501 until #10885: the file-guarded form draining with a bare `cat`.
- Windows, #9501 until #10221 took Windows off the real-home lane: the encoded
  PowerShell launcher for a non-cmd-safe script path.

The file-name sweep removed both before; the enumerated sweep left them in
place, trusted, still passing the script's exit status to Codex. Both now
match as frozen literals.

Also corrects the startup ordering comment: the real-home install runs first
so its in-slot upgrade lands before the managed install's sweep retires the
prior command; nothing re-arms a legacy sweep any more.

* fix(codex): take the real-home lock only when a write is needed

The previous commit made every entry to the real-home config lane take the
cross-process lock. That lane runs on every pane spawn and every typed
`codex` preflight, so the steady state paid an owner probe (a `ps` spawn on
macOS) and could wait up to 10 s behind another instance's trust session,
even though it wrote nothing.

Each real-home writer now compares the desired state with the files on disk
first, without the lock. Only when a write is needed does it take the lock,
re-read and recheck, then write:
- real-home install: the planned hooks.json, the shared script and the
  ledger-recorded grant are compared; the locked path re-plans from disk.
- legacy sweep: locks only when a retired entry is present; the sweep re-reads.
- approval promotion: locks only when there is something to promote; the
  promotions are recomputed under the lock.
- the shared ~/.orca/agent-hooks script: locks only when its bytes differ.
The explicit opt-out always takes the lock. The lock is reentrant through
async context, since grants and rebases nest inside an install, so the
config-lane option the previous commit added is removed.

* fix(codex): a shared script without its exec bit is not the steady state

The compare-first check matched the shared ~/.orca/agent-hooks script on bytes
alone. writeManagedScript also restores 0755 on every call, and the POSIX hook
guard skips a script that is not executable, so a script whose mode was lost
(a dotfiles restore, a plain copy) now stayed that way: every Codex hook
drained stdin and reported nothing until an app restart refreshed the script.

The check now also requires the mode the writer sets, so that case takes the
lock and the write path repairs it.

* test(codex): the retired encoded launcher never matches today's shared one

The shared encoded Windows launcher is still current for other agents, so the
comment claiming today's launcher is never encoded was wrong. What keeps the
retired matcher off it is the exact payload: since #14825 the shared launcher
prefixes its payload and drops -ExecutionPolicy Bypass. Pin that with a case.

* fix(codex): the pane step recognises its own script under a home path with an apostrophe

The same-HOME check looked for the script path wrapped in bare single quotes,
but both hook writers escape an apostrophe inside the quotes. A home such as
C:\Users\O'Brien never matched, so the pane-step repair never ran there.

* fix(codex): the trust-RPC escape hatch still keeps the real home off its lane

The no-write check reported a recorded grant as current, so with
ORCA_DISABLE_CODEX_TRUST_RPC set the real-home lane stayed in use. The grant
itself refuses before reading its ledger; the check now does the same.

* fix(codex): the shared script write no longer waits on the real-home lock

The write is atomic and skips identical bytes; waiting behind another
instance's trust session could only fail a pane's managed-home install.

* fix(codex): an in-Orca approval survives a launch that cannot get the real-home lock

The install drops runtime trust the system config does not back, so a
promotion skipped for want of the lock lost the approval for good. It
now writes unlocked, as it did before the lock existed.

* refactor(codex): take the cross-process real-home lock back out

The lock fixed no observed failure. The three that were observed each have
their own fix in this series: the legacy sweep matches only frozen retired
command forms, hooks-off launch prep writes nothing, and a pane's
prepare-codex repairs only a home its own HOME's app installed. The lock
instead brought its own defects: a steady-state spawn waiting behind another
instance's trust session, a compare-first split to avoid that, a script
write and an approval promotion that could fail for want of the lock.

Removed, with their tests: the real-home write lock and its async-context
reentrancy, the plan/compare split that kept it off steady-state spawns, the
compare-first legacy sweep, the locked approval promotion and its unlocked
fallback, the compare-first shared script write (writeManagedScript already
skips identical bytes and restores the exec bit), and the CLI tsconfig
entries the lock pulled in.

Kept: the retired-forms matcher, the hooks-off no-op, removal only on an
explicit opt-out, the pane own-script check, and the compare-and-swap
rollbacks. Every instance now writes identical bytes idempotently.

* fix(codex): an opt-out that cannot read hooks.json keeps Orca's trust and ledger

The opt-out swept the real-home entry, then dropped Orca's ledger-proven trust
whenever a ledger existed, even when the sweep could not read hooks.json. The
entry could still be there, now untrusted, and the ledger that proves ownership
was gone for the retry. Drop that trust only after a sweep that read the file.

* refactor(agent-hooks): one predicate for whether an agent's status hooks are on

"Global switch on and this agent not turned off" was spelled out separately
in the startup controls, the settings reconcile, the retained-home
reconcile, the WSL preflight RPC, the CLI preflight and the OpenCode plugin
selection. They now share one function, in a module light enough for the
CLI's per-launch Codex preflight to load. The PTY spawn env derives the
Codex flag from the switch and opt-out list it already carries, the same way
it does for OpenCode and Pi, instead of receiving a second copy.

* fix(codex): launch and resume prep honour Codex's per-agent hook opt-out

Turning Codex off in the per-agent hook settings removes Orca's Codex hook
entry, but launch prep and session resume read only the global hooks switch,
so the next Codex launch or resume wrote the entry straight back into the
real ~/.codex or the account's home. Both now read the per-agent predicate,
which the PTY spawn env and startup already honoured.

* fix(codex): turning Codex off per agent clears the real ~/.codex entry

While the real-home lane owns ~/.codex/hooks.json, the legacy system-home
sweep stands down. That gate read only the global switch, so turning Codex
off per agent ran remove() with the sweep still suppressed and left Orca's
entry in the real ~/.codex. The gate now reads the per-agent predicate, the
same as turning every hook off.

* test(codex): cover the system ~/.codex sweep gate for Codex turned off

The gate that lets the legacy system-home sweep run was an inline closure in
startup, so reverting it to the global switch left CI green. It is now a
pure function beside the gate it feeds, with a table test and a remove()
test on a seeded ~/.codex: turning Codex off strips Orca's entry and keeps
user hooks; with Codex on the entry stays.

* fix(cli): keep the agent-status hooks predicate loadable by the packaged CLI

The CLI's prepare-codex handler imported the predicate from src/main, but
the Electron build rebuilds out/main from its declared entries only, so the
packaged `orca agent hooks` commands could not load it (package jobs and the
CLI bundle-parity test were red). The predicate reads only settings, so it
now lives in src/shared, which the CLI compiles itself.

* feat(codex): every Orca build writes one frozen Codex hook command

The Codex hook command was built from this build's wrapper, so two builds
on one HOME disagreed about the bytes of the shared ~/.codex entry and kept
rewriting it, with a Codex trust session each time.

The command is now fixed per form and carries its form number:
- POSIX: one command with no path in it. It runs the shared script only in
  an Orca pane with hooks on (pane key and hook port set), drains stdin
  everywhere else, and always exits 0. A branch for a per-build script root
  is written now and stays dormant until Orca sets ORCA_AGENT_HOOK_ROOT, so
  that change will not move these bytes.
- Windows: the bare forward-slash path to the shared .cmd, which runs under
  PowerShell 7 and 5.1, Codex's hook hosts. A profile path that is not one
  PowerShell token gets a plain PowerShell form with the same branches.

The literals live in the form module, so a change to the shared hook
constants cannot move them; goldens pin the bytes. Every form keeps
`agent-hooks/codex-hook.*` in plain text, so older builds still recognize it.

* fix(codex): one main-process owner adds the real-home entry; nothing restores files

Each Orca writer of ~/.codex decided what Orca's entry must be from its own
build and instance, then removed or reverted whatever differed: launch prep
rewrote any Orca-shaped entry to this build's command and stripped Orca
entries from events this build does not use, and a failed trust session
restored hooks.json and config.toml from snapshots. With several instances
and builds on one HOME, every disagreement became a deletion or a revert.

The main process is now the one writer, and its writes are add-only:
- A launch or resume adds Orca's frozen entry to an event that has none and
  leaves every Orca entry it finds, so a running older build is never fought.
- App start also converts an older Orca form to the frozen command, once, in
  its own slot: one hooks.json write (one .bak) and one trust grant per home.
- A newer form is never rewritten or appended beside, and Orca entries in
  events this build does not use are kept.
- After a failed trust grant, only an entry this call wrote that is still
  untrusted is withdrawn, putting back the handler it replaced. Both files
  are re-read, so a concurrent edit, or the identical entry another Orca
  trusted meanwhile, survives.

Deleted: the compare-and-swap hooks.json restore, the config.toml snapshot
restore after a grant session and after a user-trust re-key, and the
rollback module. A grant session writes trust only at Orca's own keys, and
every caller settles those keys itself. A failed re-key of moved user hooks
now keeps the write and reports it; Codex lists those hooks for review.

* fix(codex): the pane CLI asks the app to prepare its Codex home

`orca agent hooks prepare-codex` ran Codex's install inside the pane. That
process can have the real HOME (login(1)) and runs outside the app's
in-process queues, so it was a second writer of ~/.codex and ~/.orca beside
the app. A check that the home ran "its own script" guarded it.

The pane step now only asks the app, over the same kind of local RPC the WSL
pane step already uses (agentHooks.prepareCodexForPane). The app checks that
the pane's CODEX_HOME is one its own userData owns, reads its own hooks
setting, and installs on its own queue. An app that is not running, or is
too old to know the method, makes the step a no-op, as it is on WSL. The
own-script check and the CLI's settings read are gone, and the preflight
module leaves the CLI bundle.

* fix(codex): delete the pane step on native hosts

The previous commit had `orca agent hooks prepare-codex` ask the app to
prepare the pane's Codex home. The case it existed for (#14326, a pane that
survives an app update with stale hook trust) did not reproduce, and no other
desktop agent host writes agent config from a terminal or launch wrapper.

- Deleted: the agentHooks.prepareCodexForPane RPC method, its params and
  catalog entry, and prepareManagedCodexHomeBeforeShellLaunch with its module,
  tests and CLI build entry.
- `agent hooks prepare-codex` is a no-op on native hosts. It stays for one
  release so shell wrappers from older builds, which still call it, exit 0.
- WSL panes are unchanged: they still ask the app over
  agentHooks.prepareCodexForWslPane.

The shell wrappers and ORCA_CODEX_LAUNCH_PREFLIGHT stay, because WSL panes
use the same wrappers and variable (forwarded through WSLENV). A native pane
still starts the CLI once per `codex` it runs; skipping that is a follow-up.

* test(codex): a failed trust session keeps concurrent edits to both files

QA case 9 at host level, on a real file system in a temp HOME: Codex's trust
session fails after another writer saved hooks.json and config.toml.

- Both saves survive, and no Orca entry is left that Codex would list for
  review: this call's entry is withdrawn.
- A failed one-time conversion puts the older Orca entry back in its slot and
  keeps both saves.

Both tests fail on the previous head, which restored config.toml from a
snapshot and left the untrusted entries in hooks.json. Removing the
withdrawal turns both red.

* feat(codex): read whether an Orca entry's stored trust is still current

A Codex release that changes how it hashes a hook leaves Orca's stored trust
stale: the entry is present, but Codex lists it as modified. Checking only
whether the entry is missing cannot see that.

readOrcaEntryTrust sorts a present entry into four states:
- trusted: the stored hash is the current one;
- untrusted: there is no stored hash;
- stale: the stored hash is not the current one;
- disabled: the user turned the entry off.

The caller can pass Codex's current hash, for example one a grant recorded.
The failed-grant withdrawal now uses it, and also keeps an entry the user
turned off. Nothing re-grants on 'stale' yet.

* fix(codex): a slow Codex start retries on the next launch, never for minutes

On a loaded Mac a cold `codex app-server` took over 10 s (QA case 4). The
grant timed out, the entry was withdrawn, and a 5-minute cooldown in both the
grant and the real-home install then refused every retry.

- The native session deadline is 30 s, the same as WSL's.
- A timeout starts no cooldown in the grant or in the real-home install. The
  next launch retries. Other failures keep their cooldown.
- Launches that queue behind a slow session share one follow-up run, so a
  launch waits for at most two sessions, not one per earlier launch.

Tests: a 15 s cold start still grants and keeps the entry; after a timeout,
the next launch runs a session at once; four queued launches run two
sessions. Each is red on the previous head, and each mechanism was removed in
turn to confirm its test turns red.

* fix(codex): Orca's automatic writes never move a user hook

Codex keys a hook's trust by its position in hooks.json. App start's collapse
of Orca duplicates removed every Orca entry and appended one at the end. That
moved any user hook that followed a removed entry, so the write waited on a
session to re-key the moved hook's trust.

App start now:
- converts the first Orca entry that sits in a plain slot to the frozen
  command, in place;
- drops any other Orca entry only when that moves no user hook;
- keeps a duplicate that a user hook follows, and trusts every frozen copy,
  so none is listed for review;
- appends only when no frozen entry is left.

Tests check user positions and user trust blocks byte-for-byte for each
automatic write: add-missing (append), the one-time conversion (in place),
a trailing duplicate, a duplicate before a user hook, and older duplicates
normalized to one entry. The three collapse cases fail on the previous head.
Removing the position check, or the in-place conversion, turns its tests red.
Only the explicit opt-out still removes an entry that user hooks follow.

* fix(codex): removing an Orca entry never waits on a Codex session

Removing an Orca entry from ~/.codex/hooks.json moves every user hook behind
it up a slot, and Codex keys trust by slot. The retired-form sweep, the
opt-out and a failed-grant withdrawal all asked a `codex app-server` session
to list the old trust before writing, and to re-key it afterwards. A timeout
there threw before the write and latched a 5-minute cooldown, so a slow cold
start blocked the retired-form sweep at boot (QA case 4).

Each moved hook's [hooks.state] block now moves to its new key, body bytes
unchanged, straight after the hooks.json write. Codex hashes a hook's content,
not its position or its file path, so the moved block stays exactly as valid
as it was: a trusted hook stays trusted, an untrusted one stays untrusted, and
one the user turned off stays off. No removal waits on or depends on a
session. A failed config.toml write keeps the hooks write and logs.

Deleted: the inspect and repair sessions, their client, and their cooldown.
The generation guards on the hooks.json writes stay, for other processes.

Tests: the retired sweep removes the retired entry and carries the trust of
the user hook behind it while every Codex session times out (red on the
previous head); the opt-out carries an appended user hook's trust; the move
carries trusted, disabled and untrusted states byte for byte. Removing the
move turns all of them red.

* fix(codex): a Codex launch never waits on Codex's approval of Orca's entry

A launch on the real-home lane awaited Codex's trust grant for the entry it
had just added. A cold `codex app-server` on a loaded Mac took over 10 s, so
the launch could wait that long, and a failure then latched a 5-minute
cooldown.

- Codex's approval runs in the background, with a 30 s cold-start budget.
- A launch uses the real home only when the ledger shows trust is already
  current. Otherwise it goes to the managed home at once, and the next launch
  picks up the finished grant.
- A launch that arrives while a grant runs does no work and does not queue
  behind it.
- A resume into the real home has no managed home to fall back to. It waits
  for the grant, but no longer than the 10 s a launch always could.
- A background grant that times out starts no cooldown; the next launch
  retries. Any other failure backs off for 10 s instead of 5 minutes.
  Success is what the ledger remembers.
- A failed grant still withdraws only what that install added and is still
  unapproved. The log now says how many entries it took back and when the
  next try comes.

Managed-home grants keep their 10 s deadline and stay on launch prep, as
before; they fall back to Orca-computed trust.

Tests:
- A 15 s start: the launch returns in under a second on the managed home, a
  second launch starts no session, the grant lands in the background, and the
  next launch uses the real home.
- A timeout sets no cooldown, withdraws its adds and logs it.
- Another failure retries after 10 s, not before.
- A resume waits only as long as allowed.
- Case 9 checks the log line and the retry.

Making the launch await the grant, a 10 s budget, either timeout cooldown, and
a 5-minute backoff were each tried, and each turns its test red.

* fix(codex): move a hook's trust only when every stored key has the known shape

Orca now edits Codex's trust store directly when a removal moves a user
hook. Three safeguards keep that honest:

- Fail safe. If any [hooks.state] key in config.toml does not have the
  shape `<path>:<event>:<group>:<handler>`, nothing moves and Codex asks the
  user to review. That shape was checked unchanged from Codex 0.141 to 0.158.
- Targeted. The file is read immediately before the atomic rename, and only
  the moved keys' blocks change. Every other byte stays, and no snapshot is
  restored.
- Verbatim. Each block's body moves as Codex wrote it, including fields
  Orca does not know. No hash is ever computed, and a hook with no block
  gets none.

Tests:
- An unknown key shape stops every move.
- Everything except the moved block survives byte for byte, and the moved
  body keeps an unknown field.
- In case 9, a hook the user approved during the failed session keeps its
  approval when the withdrawal moves it, beside the concurrent project edit.

Removing the shape check, or writing a computed block instead of the stored
body, turns these tests red.

* refactor(codex): keep only the trust read the failed-grant withdrawal uses

A capture across Codex 0.141, 0.150 and 0.158, switching in all six
directions, showed Orca's entry keeps the same hash and stays trusted. A
Codex upgrade does not make its trust stale, so nothing needs to re-grant
on staleness.

readOrcaEntryTrust keeps the four states the withdrawal needs, but loses
the parameter that let a caller pass a different current hash, and the test
for a Codex that hashes differently.

* fix(codex): native panes no longer start the Orca CLI before each codex

The pane step is a no-op on native hosts, but native panes still carried
ORCA_CODEX_LAUNCH_PREFLIGHT, so every `codex` typed in a pane started the
Orca CLI for nothing. Only a packaged Windows build's WSL pane now gets the
variable; the app prepares every native Codex home itself.

The resolver loses the dev-launcher path and its userDataPath option, which
only native panes used.

Tests: a native macOS, Linux and Windows pane gets no preflight, packaged or
not, even with the bundled CLI present; a WSL pane still gets the verified
absolute launcher. Letting native panes through again turns them red.

* chore(cli): say when the native prepare-codex no-op can go

Native pane wrappers from builds up to v1.4.216 still call it. It can be
deleted once no supported build's wrapper does.

* test(codex): check the WSL launcher path instead of asserting it

* fix(codex): a launch no longer waits behind the background real-home approval

The background grant ran its whole codex app-server session inside the shared
~/.codex/config.toml lane, and on a cold host its session was also the shared
capability probe. A launch sent to the managed home then waited on both: the
managed install and the project-trust write queue on that lane, and the
managed install's own grant waited for the probe. On a cold app-server that
was up to 30 s per launch.

The lane was held across the session only to protect the retired
capture-and-restore. Codex writes its own records, so the lane is now taken
only around Orca's own pre-grant write. The background grant runs its session
without publishing it as the shared probe, and the whole grant is bounded by
its deadline, so a hang outside the session cannot leave the lane 'granting'.

* fix(codex): a failed re-grant no longer strips Codex's own approval of Orca's entries

Before each trust session, the grant deleted every Orca record whose hash
matched the one Orca computes. That exists because a managed home's fallback
writes Orca-computed trust under both Windows path-separator spellings, and
Codex rewrites only its own spelling, so the other copy would linger. On
failure the managed and WSL fallbacks write that trust back, and before this
fold a snapshot restore covered it.

The real ~/.codex has neither: Orca never writes computed trust there (the
real-home lane does not run on Windows at all), so a matching record there is
Codex's own approval. After a ledger miss (another Orca profile, a Codex
update, a lost ledger) and a failed session, nothing put it back, and every
Orca entry showed "Hooks need review".

The clear now runs only for homes whose fallback writes that trust.

* fix(codex): a real-home resume spawns only once Orca's entry is approved or withdrawn

A resume that must run in ~/.codex waited at most 10 s for the background
approval, then spawned anyway. On a cold app-server that left Codex beside an
unapproved Orca entry, so the resumed pane showed hook review.

The resume now waits for the grant to settle. Settled means Codex approved the
entry, or the grant failed and withdrew its own unapproved write; the grant's
deadline bounds the wait (30 s, the cold-start budget), and a failed approval
never fails the resume.

Why this over the alternatives:
- Spawning at 10 s keeps the review prompt this fold exists to remove.
- Withdrawing at 10 s from the resume races the still-running session: Codex
  can write the frozen entry's hash after the withdrawal, and for a converted
  entry that marks the older command Orca put back as modified.
- A resume cannot use the managed home: the session lives in ~/.codex.
So the only states that cannot race Codex are the grant's own settle. The cost
is a longer worst case on a cold app-server (up to the 30 s deadline, plus any
managed-home install that holds the config.toml lane); a warm approval takes
seconds, and an approved entry costs no wait.

* fix(codex): keep the 5-minute trust cooldown for launch-path grants

The fold shortened the host's trust-grant cooldown from 5 minutes to 10
seconds for every grant. That was meant for the background ~/.codex approval,
which blocks no launch. The managed-home and WSL grants run inline on the
launch path, so with a hung app-server every launch more than 10 s after the
last failure paid the full inline timeout again (10 s native, 30 s WSL).

Cooldowns are now kept per lane: inline grants keep 5 minutes, the background
grant retries after 10 s, and neither lane's failure cools the other down. A
success, or a proven-missing surface, still clears both. The real-home
install's own retries (an unreadable hooks.json, unknown keys) are back on the
5-minute interval they had before the fold.

The cooldown moves to its own module so the grant stays within the file limit.

* fix(codex): a failed grant withdraws the exact copy it wrote

The withdrawal re-found "this call's" entry by command, taking the first
frozen handler in the event. When app start converted a later slot while an
earlier frozen copy sat in a matcher group (which conversion skips), a failed
grant acted on that earlier copy: it put the older command into it, or skipped
it, and left the converted, unapproved copy in place.

Each write now records where its handler landed, after any duplicate drops,
and the withdrawal acts only on that slot. A copy that has since moved is left
alone; the next launch's grant retries it.

* fix(codex): the failed-grant withdrawal checks hooks.json is unchanged before writing

The install and the retired-form sweep both refuse to replace ~/.codex/hooks.json
if it changed since they read it. The withdrawal did not: a save landing
between its read and its atomic replace was lost. The window is small, since
the withdrawal is synchronous, but it now carries the same guard.

* refactor(codex): drop rationale left over from the snapshot restore; name the trust-move module for what it does

Comments on the config.toml lanes still justified them by a grant's
capture-and-restore window, which the fold deleted, and the trust-write
deadline still counted a grant session holding the lane. They now give the
reason that remains: Orca's own multi-step reads and writes, and managed-home
installs that hold the lane across their inline grant.

codex-user-hook-trust-rebase no longer rebases through Codex; it moves stored
trust records, so it is now codex-user-hook-trust-moves.

The grant test that pinned two sessions on one config.toml to run one at a
time is removed: its reason was an interleaved capture and restore. Callers
that write config.toml around a grant hold their own lane, which the nested
installer test still covers.

* build(cli): list the trust-grant cooldown module in the CLI program

The CLI's agent-hooks handler loads the hook controls, which reach the Codex
trust grant; the CLI project is composite, so every module in that graph must
be listed.

* docs(codex): say which Windows hosts each hook command form runs under

Codex runs a hook under the turn's shell (PowerShell 7 or 5.1 in every
captured session) and, with no single local turn shell, under %COMSPEC% /C.
The bare forward-slash path ran under all three in the Windows host census.
The PowerShell form used for a profile path with a space does not parse under
cmd.exe; no form valid in all three hosts has been run for such a path, so the
form stays and the gap is stated here and in the PR.

* test(codex): type the withdrawal seam without an assertion

* fix(codex): a real-home resume starts at once, trusting Orca's entries for that process

A resume that must run in ~/.codex waited for Codex's background approval of
Orca's newly written hook entry: up to 30-40 s on a cold app-server. That made
the user's resume wait on bookkeeping, and the alternatives (start at 10 s with
Codex's hook review showing, or withdraw the entry and race Codex's own write)
were worse.

Codex reads hook trust from its session-flag config layer as well as the user's
config.toml, merged per key, and has since hook trust shipped. So the resume no
longer waits. When Orca's own frozen entries in ~/.codex are untrusted (or hold
a stale hash), the resume command carries
`-c hooks.state={'<key>'={trusted_hash='<hash>'},...}` for exactly those entries:
the key under both the logical and the real path of ~/.codex (Codex keys an
explicit CODEX_HOME by its real path), and the hash of that entry's content, so
it can trust nothing else at that slot. The user's hooks are never included,
nothing is written, and the background approval still runs for later plain
`codex` launches. An approved entry adds nothing; a Codex known to lack hook
trust gets nothing.

One inline table, because Codex splits a `-c` key on every `.` and the key holds
`.codex/hooks.json`. TOML literal strings keep `"` out of Windows native-argument
quoting. The flag goes before `resume <id>`, quoted for the pane's shell (portable
Unix, PowerShell or cmd), in the launch command and in the setup-sequenced copy of
it; a cmd line whose path cmd would expand, or a key with an apostrophe, is left
unchanged. SSH and WSL resumes get no preparation, so no local path reaches them.

* Revert "fix(codex): a real-home resume starts at once, trusting Orca's entries for that process"

This reverts commit 1bd30651d6.

* fix(codex): a real-home resume starts at once, without waiting for approval

A resume into the real ~/.codex waited until the background approval settled,
up to its 30 s deadline on a cold app-server: bookkeeping for later launches
gating the resume the user asked for. It now starts at once. If the approval is
still running, that first resume can show Codex's hook review once; the
approval then lands and later resumes and plain codex launches are trusted.

Trusting Orca's entries per process was the alternative, but the resume command
is typed into the pane's shell, and hook settings stay out of typed commands.

* test(codex): read real-home hook groups with the installer's own type

* fix(codex): a background approval is bounded only by its session's own deadline

Review loop 2, L3. grantWithinDeadline raced a second 30 s timer against
the background approval. Loop 1 added it so that a hang upstream of the
session could not leave the lane 'granting' forever.

That hang cannot happen. The only caller is the native real-home grant
(its plan is always host 'native'; the real-home lane is off on Windows,
so WSL never reaches it). Everything before the session is synchronous
there: command resolution and binary stamp, the ledger read, the
state-db backfill check, the capability and cooldown checks, and
runUnshared awaits no shared probe. A synchronous hang would freeze the
main thread, which no timer can rescue. The session itself starts a kill
timer right after spawn (runCodexAppServerSession), with the same 30 s,
and it kills the app-server tree when it fires.

So the outer timer was a second copy of that bound. Because it started
first, it won by the spawn time. It then settled the lane and cleared
backgroundGrant while the app-server was still alive, and the next
launch could start a second concurrent session. It abandoned the
session rather than cancelling it. Deleted, not moved: the session's
own timer is the one bound, and it cancels.

Test: codex-real-home-slow-app-server.test.ts "runs one session at a
time, ended by its own deadline". The fake session starts its timer
after a simulated spawn, as the real one does. A launch at 30 s finds
the session still running and starts none; the lane settles when the
session times out. It replaces the "settles a grant that never answers"
test, whose never-answering session could not time out at all.

* fix(codex): a background approval's retry has one schedule, the real-home lane's

Review loop 2, L4. A non-timeout background failure set two 10 s
schedules for one failure: the real-home lane's installRetryAfterMs,
which gates ensure, and a `<host>#background` cooldown in the grant
module. ensure's gate always tripped first, so the second one was
consulted only after something reset the first (turning hooks off).
Then it answered 'retry-cached', which wrote the entry into
~/.codex/hooks.json only to withdraw it again: churn, not protection.

Background plans now neither start nor consult a grant-module cooldown.
The real-home lane (installRetryAfterMs) is the one source of truth for
when a background approval runs again, and its 10 s interval moves into
codex-real-home-background-grant.ts, the module that sets it. The
cooldown module is back to one host-keyed map for launch-path grants,
with the same 5-minute interval as main. A success or a proven-missing
surface from either lane still clears the host's cooldown.

Tests:
- codex-hook-trust-grant.test.ts "neither starts nor waits on a
  cooldown for a background grant": two failing background grants each
  run a session and leave no cooldown; an inline failure still cools
  down inline grants and not the background one.
- codex-real-home-slow-app-server.test.ts "has one retry schedule:
  turning hooks off and on after a failure retries at once": after a
  failed approval, hooks off then on runs a session and installs,
  instead of a retry-cached write-and-withdraw.

* fix(codex): hooks turned off and on during an approval re-add Orca's entry

Review loop 2, L1. ensure returned at once whenever a background
approval was running, whatever the lane. Turning hooks off during an
approval sets the lane to 'removed' (usable), so turning them back on
returned 'removed' without re-adding the entry. Launches in that window
spawned in ~/.codex with no Orca hook and got no status for their
lifetime, for up to 30 s, until the approval settled and a later launch
re-added it.

ensure now returns early only while the lane is 'granting', which is
what the early return exists for: a launch never waits on Codex's
approval and uses the managed home until it lands. Any other lane runs
the normal add-missing install.

That install can start a second approval while the first is still
running. Approvals are now chained, so Codex still runs one session at
a time, and a finished approval clears the handle only if it is still
the latest one (before, an older approval's finally could clear a newer
one's handle). The older approval's result is already dropped by the
lane generation check.

Test: codex-real-home-slow-app-server.test.ts "re-adds the entry when
hooks go off and on during an approval, one session at a time". While
the approval hangs: opt-out removes the entry; re-enable re-adds every
entry, keeps launches on the managed home, and starts no second
session; once Codex answers, the lane is installed and every entry is
approved.

* test(codex): a launch during the real-home approval shows what it waits on

Review loop 2, M2. The launch test's fake Codex failed every
managed-home session at once with ENOENT, so the managed home's own
approval was an instant "unsupported" fallback, and the test could not
show that a launch sent to the managed home still waits on that home's
inline approval when its ledger misses (first use, a Codex update, a
lost ledger), up to 10 s, as on main.

Now the managed-home session behaves like a real one:
- "settles on the managed home with its hooks and the project trust
  written": the managed app-server answers; two launches settle in
  under 2 s while the real-home approval hangs, and the second launch
  finds the managed approval in its ledger (one managed session).
- new "waits up to the managed home's own 10 s approval when that home
  is cold too": the managed session fails at its own deadline, as the
  real one does. The first launch is still pending at 9.999 s and
  settles on the managed home at 10 s; the request asked for 10 s. The
  next launch settles at once, because the failed inline approval cools
  down for 5 minutes.

No product change.

* refactor(codex): the managed and WSL installs own their pre-approval trust clear

Review loop 2, L7. Before a Codex approval session, a managed or WSL home
clears the approvals Orca itself computed, because on Windows its
fallback writes them under both path spellings and Codex's canonical key
may not overwrite the other one. The fallback writes them back if the
session fails. ~/.codex has no such fallback, so there the clear would
only delete Codex's own records (loop-1 H2). The grant module carried
this as a plan flag, fallbackWritesSelfComputedTrust, and took the
config.toml lane around the clear itself.

The reviewer proposed moving the clear into the two callers. A literal
move, clearing before the grant call, is NOT behaviour-neutral, so this
does not do that:
- The grant first checks its ledger, which compares the stored hash
  with the one Codex recorded. Codex's hash equals Orca's computed one
  (the premise of readOrcaEntryTrust), so a clear before that check
  deletes exactly the record the ledger proves. Every managed launch
  would then miss the ledger and run an inline session (up to 10 s).
- Checked, not inferred: with the clear moved before the call in the
  managed install, codex-launch-during-real-home-grant.test.ts "settles
  on the managed home..." fails (2 managed sessions instead of 1).
  Log: ~/orca-qa/codex-real-home-leak/fb6/l7-literal-move.log

What this does instead: each caller passes its clear as the grant's
`beforeSession` step, which the grant runs only when a session will
actually run (after a ledger miss, and not on a cooldown or cached
fallback), exactly where the flag ran it. So:
- the flag and its "never set for the real home" rule are gone; the
  real-home grant passes no step, so the grant module has no path left
  that deletes a trust record in ~/.codex;
- the grant module's own lane acquisition around the clear is gone. It
  was always a pass-through: both callers already hold that file's
  lane (the managed install holds the runtime and system lanes, the
  WSL install holds its config.toml lane) across the whole grant.

No behaviour change. The loop-1 probes still pass as fixed: trust-strip
prints every entry trusted after a failed re-grant, and lane-hold
prints managedInstall=settled projectTrust=settled.

Tests (codex-hook-trust-grant.test.ts):
- "removes equivalent Windows fallback keys before the RPC writes
  canonical trust" now passes the managed caller's step;
- new "runs the caller's pre-session step only when a session runs":
  the step runs once for a session and not on the ledger hit after it.

* chore(codex): comments stop describing a lock held across the session, or a rollback

Review loop 2, L6 comment sweep (comments and one test name only):
- codex-trust-config-concurrent-launch.test.ts: the test named "does not
  let a failing launch roll back a concurrent launch" said the per-file
  lane was the only thing left and that the doomed run's rollback must
  not resurrect the file. There is no lane across a session and no
  rollback now. Retargeted to what it covers: "leaves a concurrent
  grant's records in place when a sibling grant fails" (a restore would
  still turn it red).
- codex-trust-grant-ledger.ts: "a grant session blocks launch prep" is
  true only of inline grants; the background one still costs an
  app-server start. The drift clause no longer says "before the pane
  launches", which is false for the real home.
- agent-trust-write-deadline.ts: a stray hard wrap.
The install.ts:105 comment was fixed with L1. A sweep of src/main/codex,
src/main/startup, src/main/agent-hooks, the trust presets and the CLI
handlers for rollback, restore, rebase, capture/restore, and a lane held
across a grant or session found nothing else stale; the remaining "no
restore" comments state the current rule.

* fix(codex): a real-home resume waits for the one running approval, up to its 30 s limit

Review loop 2, M1; coordinator ruling. A resume into ~/.codex has no
managed home to fall back to. 5a737261d8 let it start at once beside an
Orca entry still awaiting Codex's approval. Codex's TUI then shows a
full-screen hook-review picker before the session and waits for keys:
"Trust all and continue" also trusts the user's own unreviewed hooks,
and "Continue without trusting" leaves that session with no Orca status
for its whole life, because Codex does not reload hooks when Orca's
approval lands later. Panes restored at app start after an update hit
it too, since the start-time conversion leaves every entry awaiting
approval.

The resume now waits, but only while Orca's entry in ~/.codex is
written and a grant is approving it (lane 'granting'). Every resume
waits on that same in-flight grant: ensure never starts a second one
while the lane is 'granting', so panes restored together share one
session. The bound is the grant's own session limit (30 s). The grant
settles only after Codex approved the entry, or after it withdrew its
own unapproved adds, so the resumed session starts either trusted or
with no Orca entry: never beside an unapproved one, and no picker. On
a withdrawal that session has no Orca status, as on main after its
10 s wait. A failed approval never fails the resume.

Tests (codex-launch-during-real-home-grant.test.ts):
- "waits for a warm approval, and spawns with the entries approved";
- "spawns at the approval session limit with Orca entries withdrawn"
  (fake timers: pending at 29.999 s, spawns at 30 s with no Orca entry);
- "makes panes restored together wait on one approval session" (three
  resumes, one session, all settle once it lands).
codex-launch-per-agent-hook-opt-out.test.ts: a resume into ~/.codex
awaits the approval; a resume into a managed account home does not.

* fix(codex): repeated background approval timeouts back off, growing to 5 minutes

Review loop 2, M3; coordinator ruling. A timeout of the ~/.codex
approval starts no cooldown, so the next launch retries at once. On a
host where codex app-server never starts within 30 s, every launch then
wrote Orca's entry into ~/.codex/hooks.json, withdrew it again, and
started another 30 s session, for the rest of the process: an unbounded
retry with no exit.

After 3 timeouts in a row the retry now waits 10 s, then 1 minute, then
5 minutes for every later one. The first two timeouts still retry on the
next launch, so a slow cold start is not punished. Any other outcome
ends the streak (a success, or any other failure, which keeps its own
10 s wait). The streak lives only in memory, so every app start begins
at zero and a slow boot can never latch.

Tests (codex-real-home-slow-app-server.test.ts):
- "backs off after three timeouts in a row, growing to 5 minutes, and a
  success resets it": the first two timeouts retry at once, then 10 s,
  1 min, 5 min, 5 min; after a success, a fresh approval gets two
  immediate retries again and a 10 s backoff after the third;
- "keeps trying after timeouts during a slow first start, once the app
  server answers": three timeouts, then the next attempt at 10 s
  installs.

* refactor(codex): one approval at a time, decided under the config.toml lane

The real-home check kept a lane label, a generation stamp, a promise chain of
ensures and a chain of approvals, and decided from the label at call time.
Concurrent resumes from any state other than 'granting' each started their own
approval (N x 30 s), a chained approval ran a plan an earlier failure had
withdrawn, a hooks-off check during an approval released a waiting resume beside
unapproved entries, and an app-start conversion during an approval was dropped.

Now each check is one step under the real config.toml lane: an approval in
flight answers 'approving' (unusable), hooks off answers 'removed', an open
retry window answers 'unavailable', and otherwise the unchanged install runs and
starts at most one approval. The approval settles under the lane: it withdraws
its own unapproved adds on failure, sets the retry, and derives the verdict from
the settings and the outcome, then runs an owed conversion. A resume waits only
while an approval runs and an unapproved Orca entry is on disk. The opt-out
sweep moves verbatim into its own module.

* fix(codex): only a success or app start resets the approval timeout streak

The ruling is that three timeouts in a row back off, and the count resets on
success and at app start. A non-timeout failure or an unexpected error also
reset it, so a host alternating those with timeouts never backed off.

* fix(codex): a Windows profile path the shells cannot carry bare runs through cmd.exe

The Windows hook command was the bare forward-slash script path, or, for a
profile path that is not one PowerShell word, a PowerShell script. That script
cannot parse under cmd.exe, which Codex uses when a session has no single local
turn shell, so such a profile got no status there.

A path of only letters, digits and _ . : / ~ - stays bare. Any other path,
including one with a space, & ^ $ ` ' ! ( ) or a non-ASCII character, is written
as cmd --% /d /c @"<path>", which ran under PowerShell 7, Windows PowerShell 5.1
and cmd.exe for each of those characters with a real Codex 0.158.0. The choice
depends only on the path, so every build on a machine writes the same bytes. A
machine holding the earlier PowerShell spelling converts it once at app start.

* build(cli): list the real-home hook sweep module in the CLI program

* fix(codex): the Windows cmd spelling names the system cmd.exe and turns off delayed expansion

A profile path the shells cannot carry bare was written as
cmd --% /d /c @"<path>". Under Codex's cmd.exe host the outer cmd.exe resolves
a bare `cmd` from the hook's working directory first, so a repo holding
cmd.bat (or .cmd, .com, .exe) at the session cwd would run on every hook event.
And with delayed expansion turned on in the registry, a `!` in the path was
dropped.

The spelling is now <SystemRoot>/System32/cmd.exe --% /d /v:off /c @"<path>",
unquoted (PowerShell reads a quoted first token as an expression) and with
forward slashes. The Windows directory comes from %SystemRoot% when written,
else from the directory above %ComSpec%'s System32, so both give the same bytes;
if neither is a drive-absolute path it can spell unquoted, it is C:/Windows,
which is still absolute. The bytes stay a pure function of the profile path and
that directory, so every build on a machine writes the same command. Safe
profile paths keep the bare path. Older Orca forms, including the bare-cmd
spelling, convert once; the new spelling is never swept as retired.

* refactor(codex): an approval's settle runs no deferred conversion

An app-start conversion that arrived while an approval ran was remembered and
run by that approval's settle. The settle then rewrote an older entry in place,
unapproved, and started a second approval inside the same wait that releases
every resume, so a resume could start beside an entry Codex would put up for
review.

That path could not happen: the only conversion caller is app start, and it is
the process's first check, so no approval can be running when it arrives. The
deferral and the settle's second check are deleted. A conversion that met an
approval would now be skipped until the next start, and the test for this case
pins that the settle writes nothing new and runs one session.

* fix(codex): an approval's settle keeps a failed opt-out's verdict and ends only its own flight

With hooks read off, an approval's settle always concluded 'removed', which the
routing check treats as usable. If an opt-out during that approval could not
read hooks.json, it had concluded 'unavailable' because the entry may still be
there, and the settle overwrote that. The settle now keeps 'unavailable' when
hooks are off; the next hooks-off check or opt-out re-derives it as before.

The settle's fallback when it cannot run now clears the running approval only
if it is still its own, and the routing check's comment states its rule: never
usable while an approval runs.

* fix(codex): spell the system cmd.exe with backslashes

Under Codex's cmd.exe host the outer cmd.exe hands the typed program text to
the child verbatim, and cmd.exe scans its whole command line for switches, so a
forward-slash C:/Windows/System32/cmd.exe is read as switches: the hook never
runs ("The syntax of the command is incorrect.") and /d is lost. Measured live
on Windows; both PowerShell hosts rewrite argv0 and were unaffected. The script
path after @" keeps forward slashes.

* chore(codex): say why the cmd.exe path is absolute, as measured on Windows

* test(codex): Windows managed-install tests expect the frozen command

They still asserted main's PowerShell text and a backslash bare path; they only
run on Windows, so nothing here caught it. Also correct the /v:off comment: a
lone ! is never dropped, only a !NAME! pair expands.

* ci: run the Codex managed-install tests in the Windows job

Its Windows-only cases skip everywhere else, so nothing ran them; three of them
still asserted a command this branch no longer writes.

* ci: a change to the Codex managed-install tests starts the Windows job

Also say what the missing-script case asserts: a non-zero exit, which
PowerShell reports as 1.

* chore(codex): name the hook trust key pattern for what it matches

* test(codex): the managed-install tests remove folders with the retrying helper

Now that they run in the Windows lane, a raw recursive rm there can throw EPERM
after the assertions pass.

* refactor(codex): one Codex hook-trust key pattern for the trust move and #23958's carry

* test(codex): the trust move carries a block in Codex's quoted spelling and leaves no second table
2026-09-30 14:26:23 -07:00
Neil cfe4c633eb fix(mobile): publish the Android APK's size and checksum with the release (#24037)
An APK that fails to install with a missing certificate or a package-parse error
is usually a download that died near the end: the signature block sits in the
last ~100 KB of a 133 MB file, so a truncated APK looks complete and carries no
signature at all. The release published neither a size nor a digest, so there
was no way to tell that apart from a bad build without deriving both from the
asset by hand.

The release now uploads app-release.apk.sha256 next to the APK in
`sha256sum -c` format (binary marker, so Git Bash cannot translate line endings
while hashing) and puts the exact byte size and digest in the release body,
naming `shasum -a 256 -c` for readers on macOS.

The upload path rewrites the body too: --clobber replaces the APK, so a digest
left over from the previous build would describe a file nobody can download,
and a reader comparing against it would reject a good APK. Both paths reserve
the section's own length out of the release-body cap before truncating, so the
section always survives and the body always fits; MAX_RELEASE_BODY_LENGTH is
exported from the desktop release script rather than restated.

Refs #24011, #12248, #11444.
2026-09-30 12:14:13 -07:00
Brennan Benson cb363444f3 feat(native-chat): register Codex default-mode helpers as subagents (#22619)
* feat(native-chat): register Codex default-mode helpers as subagents

Codex's default multi-agent mode announces a helper only as the
collabAgentToolCall that spawned it; it sends no subAgentActivity. The
roster and the background-task tracker registered children only from
subAgentActivity, so such a helper had no record, no strip row and no
roster row, and its commands read as the session's own.

One announcement reader now yields a child from either wire shape, and
both the journal roster and the tracker register through it into the
same executions, so a session sending both keeps one child per thread.
A finished closeAgent ends the helper's running turn as stopped through
the executions, beside the child's own turn and thread frames.

Each collab call renders as a tool row (spawn_agent, wait_agent,
close_agent, ...) naming the helper the way the roster does, with what
the helper said back as its output, instead of the raw provider row.

* fix(native-chat): name a spawn row by its prompt until the roster holds its helper

Also pin that the roster row appears when the helper's first turn arrives
before the spawn call finishes.

* test(native-chat): a spawn call keeps its own row beside the roster row it starts

* fix(native-chat): pair a structured tool row's output with its own call

A run paired results to calls by position alone, so once one call finished
with no output (a Codex spawn_agent row) every later output drew under the
call before its own. A structured row carries its call and output together,
so the projected result now names its call id and pairing honors it,
falling back to position for results that name none.

* fix(native-chat): the Codex roster row follows the executions, so every child ending settles it

The subagent-group row was rewritten only on a child's turn/started and
turn/completed. A child turn ended any other way — a fatal error, its
thread closing, its caller closing it — settled the strip and the host
record through the executions but left the transcript row reading working
until the session ended.

The executions now say when a child's execution changes, and the roster
revises its row from that, whichever frame changed it. handleTurn still
re-derives to hand back its write admission; the revision is idempotent.

* fix(native-chat): a restored Codex call row names its helper, not its thread id

History replay never runs the live item router, so the roster never learned
the helpers a restored thread had spawned, and a restored wait_agent,
close_agent or send_input row labelled its helper with the raw thread id.
Replay now registers each announced helper for its name and membership only:
register starts no execution, so no strip entry, record or roster row claims
the helper runs until a live turn of its own says so.

The replayed-item handling moves into the restore module beside the replay
that calls it. Also name the roster's render contract in handleItem, and say
why the collab tool-name map is not a spelling fix.

* fix(native-chat): register a Codex helper whose spawn call failed but created its thread

Codex reports a spawn as failed when the helper it created errored at birth,
yet the call still names the thread it created, and that thread can run. The
spawn was registered only when the call completed, so such a helper existed
nowhere and its shell read as the session's own bare command: the original
default-mode bug. A spawn now registers whenever it names a receiver; one in
progress, or one that created nothing, names none. A failed closeAgent still
ends nothing, because the helper was not closed.

* refactor(native-chat): each Codex thread's latest token total gets its own home

Every thread's running total, member or not, with the rule that the newest
frame replaces the last and the recency-ordered cap, moves out of the roster
into codex-thread-token-totals.ts. The roster still selects its children's
totals at write time. With the producer linkage the roster now builds, it
was past the size limit.

* refactor(native-chat): the messages a session list quotes get their own module

The readers of a session's newest own prompt and own assistant prose move from
the status projection into structured-agent-session-latest-messages.ts, and
are re-exported so their consumers keep one import site. With the status
clock the projection now carries, a tool result naming its call put it past
the size limit.

* test(native-chat): a Codex helper's command is a live record until its process reports its exit

A helper's in-turn shell is a live command record owned by the helper and is also its open
Bash operation; its exit removes the record. A caller's closeAgent kills the helper's
processes, and each still reports its exit on the helper's thread, so the close leaves the
command to that exit rather than ending it itself.

* fix(native-chat): every reader of a tool run pairs a result with the call it names

The folded desktop run now pairs a result with the call it names, but mobile's tool run, the desktop edit cards and the task lists still paired by position through `pairToolBlocks`. In a Codex default-mode session the new `spawn_agent` row finishes with no output, so on mobile the helper's reply drew under `spawn_agent` and `wait_agent` showed none, and a desktop edit card could take the next command's output as its own.

`pairToolBlocks` now follows the same rule as `pairNativeChatToolResults`: a result that names its call answers that call, and one that names none answers the oldest unanswered call as before.

* fix(native-chat): a Codex helper's turn ends on its own frames, not on its caller's closeAgent

A finished `closeAgent` call ended the helper's running turn as `stopped`. But the call's status is not the close's outcome: Codex sets it from the helper's own agent status, so a close that errors on a running helper still reports `completed`. Orca then marked a helper that was still running as cancelled in the strip, the host record and the roster row, and because the first ending a turn gets stands, its real ending could never correct it.

A close that works does not need the edge either. Codex's shutdown of the helper aborts its running turn, and the app-server reports that as the helper's own `turn/completed` with status `interrupted`, which Orca already maps to `stopped`. So a helper's turn now ends only on its own frames (its `turn/completed`, a fatal `error`, `thread/closed`) or the session ending, and the close is only its call row.

This also drops the child-work evidence re-keying that existed only for the close: every remaining frame names the thread that sent it.

* fix(native-chat): a tool result that names a call is never given to a different call

A result that named a call id with no unanswered match fell back to the oldest unanswered call. On mobile, a run shows at most six calls, so a command past that window whose output named its own call gave that output to an earlier call that finished with none, such as `spawn_agent`.

A result that names its call now answers only that call and otherwise stays unpaired. A result that names no call still answers the oldest unanswered one. Both pairing readers share the one rule through `answeredToolCallIndex`.

* test(native-chat): name the Codex default-mode test file after the subagents it covers

* fix(native-chat): a Codex helper is known from any call that names it, not only its spawn

A helper whose spawn Orca never saw (history replay after compaction dropped the spawn, or a resume or message to a helper from earlier) stayed unregistered, so its shell read as the session's own command: the bug this PR fixes, in another shape.

Every collab call now announces each helper it names, skipping a receiver the finished call reports notFound. The spawn's prompt still labels a helper when it was seen; otherwise the helper has no label and reads as any unnamed subagent does. A send_input, like the spawn, names the parent turn the helper's next run belongs to.

The session's own thread is excluded once, in the reader, instead of at each of its three callers.

* fix(native-chat): every finished Codex collab call row carries what the call reports

A client that predates result call ids (an older mobile app, or an older desktop reading a newer host) projects the journal itself and pairs each tool result with the oldest unanswered call. This PR publishes collab calls as tool rows, and a finished spawn_agent, send_input, close_agent or resume_agent on a running helper had no output, so such a client drew each later output in the run under the call before its own (the parent's shell output under spawn_agent, the helper's reply under the shell).

Every finished call now has an output taken from the item: a wait's reply from each helper that finished (an errored helper's error), and for every other call the helper's reported status in Codex's own words (Pending init, Running, Completed, ...). A close or resume no longer shows the helper's last reply as if the call returned it. A wait whose end names no helper (it timed out, or v2) reads Finished waiting; any other call with no state reads its own status.

A call also keeps naming the helpers its started item named: Codex ends a timed-out wait with no receivers, so its row lost the helper's name when it finished.

* refactor(native-chat): the desktop run's tool pairing is a view of pairToolBlocks

Desktop runs and every other reader (mobile, edit cards, task lists, the ask row) each had their own pairing loop sharing one index rule. The desktop's pairNativeChatToolResults now reads the pairs pairToolBlocks makes, so a run is paired by one loop and a new reader cannot add a third. No behaviour change.

* test(native-chat): type the positional-client pairs without assertions

* fix(native-chat): a finished Codex collab call row says what the call did, in Codex's own words

A close_agent row read `Running` and a spawn_agent row `Pending init`: each showed the helper's status snapshot from before the call took effect, so a close that stopped its helper read as though it had not worked.

Each finished call's output now follows Codex's own client: spawn reads `Spawned` (or `Agent spawn failed` when no helper was created), send_input `Sent input`, close `Closed`, and resume the helper's status summary. A wait keeps each finished helper's reply, with other statuses in Codex's summary wording (`Error - <message>`). The row label already names the helper, so the output does not repeat it.

* docs(native-chat): the collab call reader says which calls show the helper's snapshot as output

Since spawn, send_input and close rows say what the call did, the reader's header was wrong to claim every call row shows the reported snapshot as its output: only a wait or resume does.

* fix(native-chat): a Codex spawn call no longer claims it ran an agent

Codex's spawn_agent call ends as soon as the helper starts, so counting it as running an agent drew "Ran 1 agent" directly above "Kicked off 1 subagent · working" while the helper was still working, and "Ran 1 agent · 1 failed" when Stop cancelled a spawn before any helper existed. Claude's Task call lasts as long as its subagent, so it keeps the agent category; the Codex spawn row is now a plain tool call and the roster row alone stands for the helper.

* test(native-chat): a Codex helper's row settles on the failed completion that follows a fatal error

Codex follows every turn-ending `error` with the turn's failed `turn/completed`,
on a helper's thread as on the primary, and only that completion ends the turn.
The roster test now sends both and checks the row, strip and record stay
working through the error and settle on the completion.

* fix(native-chat): a Codex helper's section opens while the parent spawns, waits on or messages it

A running chat holds a subagent's section open while the parent's newest row delegates to that subagent. For Codex that rule only knew the raw collab status row; this PR writes each collab call as a tool row (spawn_agent, wait_agent, send_input, close_agent, ...) naming its helpers in `input.agents`, so a default-mode helper's section stayed shut for its whole run while Claude's opened.

The delegation reader now reads a Codex collab tool row as a delegation to the first helper it names, the same rule the raw row keeps for journals written before this change; a call naming no helper stays ordinary output. The row names and the `agents` key live in one shared module the host writes from and the reader reads, instead of a second list.

The new test drives the real adapter's rows through the transcript projection to the delegation; the collab frame harness moves to a fixture shared with the positional-clients test.

* fix(native-chat): a Codex collab call with a long prompt still names its helper

A collab call row bounded its whole input as one value, so a spawn or send_input whose prompt passed the 16 KB journal limit was stored as a clipped wrapper: the row lost the helper's name and thread ids, and with them the label and the delegation that opens the helper's section.

The prompt is now clipped on its own with the journal's inline-text bound, marker included; the helper's name, ids, model and effort are bounded as before and always survive.
2026-09-30 11:36:57 -07:00
Jinwoo Hong 9afd1101ff fix(orchestration): stop minting and printing the dispatch capability (#23994)
* fix(orchestration): authorize worker reports without the dispatch capability

Worker lifecycle reports and questions no longer depend on the per-dispatch
capability token that lives only in the agent's conversation. The host now:

- ignores capability_hash/capability_revoked_at for authorization on every
  row and checks the exact worker process instead (ask gains that check);
- refuses a report whose calling terminal is provably another orchestration
  party (a Run coordinator or another Dispatch's worker), treating env that
  names no live pane here as absent;
- applies one worker-state rule locally and remotely: a stop in flight
  refuses, while stop_unknown and start_unknown accept and settle.

Minting and printing the flag are unchanged, so an older host and older
preambles keep working.

* fix(orchestration): stop minting the dispatch capability

Dispatches no longer mint a per-Dispatch token, and preambles, the bundled
skill guide and the ask resume hint stop printing --dispatch-capability. The
consumer-generation bump and delivery fence that minting carried stay, now
as setDispatchConsumer. Readers that inferred meaning from capability_hash
read what they meant instead: worker-show's injected stage comes from the
attached consumer, and a failed start copies custody identity only when no
authority was ever attached. The CLI keeps accepting and forwarding the
flag for older hosts.

Cancelling a Task is recorded as failed with a reason; task-list now shows
that reason and the guide and task-update notes document the recipe.

* fix(orchestration): name the fenced party without implying which Dispatch it owns

* refactor(orchestration): one worker report rule, fence only a different party

- One module owns the unproven/settleable worker states and the refusal rule; local send records it, ask and remote throw it. A stale process is worker_identity_changed on every path.
- The caller fence passes the worker's own terminal when its --from handle went stale.
- Document the shared-tmux-server limit; drop the dead dispatch_capability_invalid rejection member; tests assert dispatch state, not the capability column.

* refactor(orchestration): drop setDispatchConsumer and the dead capability retention

- dispatch --inject no longer re-points the row createDispatchContext just wrote; worker-show reports every worker-less Dispatch as context_only, since Orca keeps no record of the paste. Tests re-point through a fixture.
- failWorkerStart always records when the lifecycle closed; nothing authorizes on it.
- Restore the ask resume hint's echo of a passed --dispatch-capability: an old host checks it before --resume.
- Move the cancellation convention to #23983.

* test(orchestration): drop capability-era assertions other tests already cover

* refactor(orchestration): drop the host-side capability field and no-op test fixtures

- RpcRequest and the SSH bridge stop carrying orchestrationCapability; the CLI's wire field stays for older hosts.
- Fixtures pass identity to createRootDispatch instead of re-pointing to the same values; drop absence checks for a flag that can no longer be produced.

* test(orchestration): cover a current process whose terminal moved to another pane

* chore(orchestration): finish the capability cleanup in test stubs and skill wording

* test(orchestration): drop needless response casts; mark the db stub cast safe
2026-09-30 13:26:06 -04:00
Neil a781a602a8 test: retire duplicate cases that replay an owner across a re-export or provider shim (#24114)
Resolves 208 candidate pairs where the same case title appears verbatim in two or more
files, produced by a repo-wide scan calibrated against a known positive. 46 case
declarations removed across 32 files, 798 lines gone. No file deleted whole, no
production code touched.

The headline result is the measurement, not the deletions: across the three buckets that
reported in detail, the signal ran roughly 86% false-positive (3/42, 9/42, and the rest).
It has good recall and poor precision, and it reorders a reading queue rather than
replacing one. Calibrating a detector against a known positive proves recall, not
precision.

What the deletions were:

- Duplicate invocation through a re-export shim. `native-chat-tool-summary.ts` is a
  ten-line `export {...} from '../../../../shared/native-chat-tool-summary'`, and
  `agent-status.ts:161` is `export { isExplicitAgentStatusFresh } from
  './pane-agent-evidence'`. Cases on the shim side were byte-equivalent to the owner's
  with no rendering or transport hop.
- Provider-local replays of a shared helper: three `repository-ref` providers that are
  each `createRemoteRefProbeCache(parseXRef)` and contribute nothing to transient
  handling; two `local-pty` and `daemon/session` tables replaying
  `shell-startup-output-scanner`, whose owner additionally checks every split point.
- A reader-side replay of store policy. `runtime-worktree-agent-rows-structured.test.ts`
  asserted an attention-to-blocked mapping; the reader contains zero `attention` or
  `blocked` tokens and copies `state` through. The mapping lives in
  `structuredAgentSessionAgentStatus`. Consistent with
  `docs/reference/agent-status-store.md`: readers keep only presentation policy.
- Constructor-only subclass duplication: the shared capability-cache case is covered by
  `codex-app-server-capability-cache.test.ts`, whose ten cases include the identical
  title plus all four risks `docs/reference/git-compatibility.md` names — first fallback,
  later cached call, concurrent probes, per-host isolation.
- A private predicate duplicated at a real boundary, varying only a path passed straight
  into the shared predicate.

Why most pairs were KEPT, because the false positives are principled rather than noise:

- Two independent execution hosts. `src/relay/git-handler-*` and `src/main/git/*` are
  separate Git implementations that cannot import each other and hold separate capability
  caches, exactly as the compatibility doc requires; the repo already ships
  `status-branch-line-total-relay-parity.test.ts` to pin the duality deliberately. Neither
  side's argv, timeout or cache regression is visible to the other.
- Deliberately duplicated production siblings: Codex vs Claude (different account fields,
  different CLIs, different wire protocols), gitea vs bitbucket (`/pulls/42` vs
  `/pullrequests/42`), gl-utils vs gh-utils (separate in-flight maps). Same contract
  shape, different implementations — an identical title is the correct naming.
- Shared-predicate consumers: one side tests the predicate, the other tests a caller's
  wiring to it. A caller that forgot to call the predicate passes the shared test.

In a codebase with intentional provider and host symmetry, identical test titles are
expected, and the signal cannot distinguish "copied" from "parallel by design" because
both produce the same prose. Only reading both bodies separates them.

Verified: 6,968 desktop test files pass; the three modified mobile files pass (39 cases);
`check-reliability-gates.mjs` 140 gates; nothing under
`mobile/src/test-support/rpc-recording/` or `mobile/rpc-foundation/goldens/` touched.

62 local failures across 12 files were each accounted for and none is caused by this
change: `browser-manager-tab-identity`, `browser-manager-viewport-ownership`,
`session-scanner-codex-workers` and `managed-hook-script-refresh` all fail identically in
a pristine `origin/main` worktree; five `mobile-web-app-*-render` tests need Playwright
browsers this machine lacks; `structured-agent-session-restart-ownership` and
`ssh-remote-commands` pass in isolation and fail only under concurrent load.
2026-09-30 03:30:23 -07:00
Neil d2dbe2c385 fix(windows): replace the managed CLI launcher with a native one (#24094)
* docs(security): add the antivirus clearance path for future releases

Every AV false positive here has been handled one vendor and one shipped
version at a time. Document the programs that clear future releases instead --
signer and product enrollment rather than per-build sample submission -- and add
a script that reports an RC's current detection state by hash, so a verdict is
found before users meet it in an issue report.

Hash lookup only by default; --upload transmits the artifact and stays manual.

* fix(windows): replace the managed CLI launcher with a native one

resources\bin\orca.exe was a csc-compiled MSIL assembly: a small, freshly
compiled .NET image in a user-writable directory that mutates environment
variables and proxies a child process. That is the shape .NET dropper
heuristics are trained on, and every verdict against it named the family --
MSILHeracles from two vendors, Wacatac!ml from a third. Signing the file does
not change its shape, so signing never cleared it.

Rebuild it in Rust. Same resolution, same environment contract, same argv
passthrough that keeps newline-bearing orchestration bodies intact (#8374), and
the child still inherits our environment block rather than an explicit map, so
a block carrying both PATH and Path survives (#12046). The PE now carries
publisher, version, icon and an asInvoker manifest from build.rs.

Refs #23383

* ci(windows): install the Rust toolchain before building the CLI launcher

The hosted runners happen to ship cargo, but a real Windows dev box does not --
verified on our own Windows QA host, where cargo and rustc were both absent.
Relying on the image means a future image change fails deep inside
electron-builder's native hook instead of at an obvious step.
2026-09-30 02:49:03 -07:00
Neil b99462ac1c test: retire mobile, cloud, config and e2e cases their input cannot reach (#24077)
Completes the first pass over every test area in the repository. Sweep over
`mobile/src`, `config/scripts`, `cloud/`, and `tests/` (1,494 files in scope, with
the 24 files under `mobile/src/test-support/rpc-recording/` deliberately excluded).
31 case declarations removed across 17 files, 2 test files deleted, 356 lines gone.

What went, by pattern:

- Cross-boundary replays of a shared helper. A whole mobile file re-ran
  `extractPendingAsk`/`parseAskFromStatus`/`formatAskAnswer`, all owned by
  `src/shared/native-chat-ask.test.ts`, `native-chat-ask-fifo.test.ts` and the
  renderer's interactive-prompt suite — one case title was verbatim identical to the
  owner's, and the owners' inputs are supersets. The mobile file imported the shared
  module directly and exercised no mobile transport, lifecycle or rendering.
- A case whose input cannot reach the behavior its title names: "arms it on Android
  while the drawer is open", where `use-back-claim.ts` has zero
  Platform/OS references, so flipping the mocked OS changes only shadow styles.
- Identity copiers, including one asserting `prSidebarRenderBranch(state) ===
  state.kind` against a production body that is `return state.kind`. The function
  stays; it has three live callers.
- A test of the runtime rather than the product: a case asserting Node's own
  `EventEmitter` crash contract on a bare emitter, with zero production code in the
  path. The guard it documents is exercised behaviourally by the case after it.
- Duplicate invocations, one of them provable rather than eyeballed: with
  `MODULE_SCOPE_ENV_WRITER_PIN = 0`, `files.size <= 0` is strictly implied by the
  sibling's `expect(offenders).toEqual([])`, since a non-empty `offenders` forces
  `files.size >= 1`. The pin's own doc says it may only ever be decreased from 0, so
  it could never become a meaningful bound either. Its policy guidance survives as a
  comment; the file's real ratchet and its regex self-test both stay.
- Expected values produced by the test's own arithmetic, and a p95 case strictly
  implied by a sibling that already pins exact p95 and exact max over a wider range.

One production line goes: the `export` keyword on `assignmentCleanupSteps` in
`cloud/apps/relay/src/assignment-cleanup-steps.ts`. The function itself stays and is
still called internally; only the test-only export was orphaned.

Kept deliberately: everything a gate cites, checked by case title and not only by
file path; a gate-cited case that does not deliver its claim (reported instead — see
below); a cross-version wire cell whose ledger is never invoked, left under the
raised bar for wire coverage; and every limit, bound, quota and provenance guard.

Nothing under `mobile/src/test-support/rpc-recording/` or
`mobile/rpc-foundation/goldens/` was touched — those bytes feed a `recorderSha256`
digest pinning 398 golden recordings.

Verified: `mobile` vitest over the modified mobile files (8 files, 50 cases);
`mobile/scripts/check-tests-typecheck-ratchet.mjs` OK (898 files in program, 125
grandfathered, none @ts-nocheck); relay suite 799 passed; `check-reliability-gates.mjs`
140 gates; both deleted files confirmed absent from the gate manifest,
`cloud/package.json` and `mobile/tests-typecheck-baseline.txt`.

Seven local failures were investigated and none is caused by this change: five
`mobile-web-app-*-render` tests drive `playwright-core` chromium/webkit and need
browsers this machine lacks, `release-checkout.unit.test.ts` needs cross-version git
refs, and `e2e-worker-env-isolation.unit.test.ts` fails identically with its HEAD
content restored — it recurses `tests/e2e` with symlink-following `statSync` and no
depth guard.
2026-09-30 02:05:30 -07:00
d68a5be13a fix(claude): run Windows hooks without shell operators (#23944)
* fix(claude): run Windows hooks without shell operators

Keep neutral replies inside the managed entry and payload scripts, repair missing files from managed registrations, and stop using Git Bash discovery to guess Claude's hook shell.

Co-authored-by: latte271 <junghyeyun27@gmail.com>
Co-authored-by: Bing.Z <zzb@gxsmjx.com>

* fix(claude): keep the Windows hook refresh async and its scripts after uninstall

- List windows-hook-files.ts in the CLI project so typecheck passes.
- Refresh the entry/payload pair only from a surviving entry, with an async
  existence check, so startup refresh stays off the main thread on Windows.
- Keep both scripts on uninstall like every other agent; a Claude session
  still holding old settings keeps answering instead of erroring per event.
- A payload that exists but cannot start falls through to the neutral reply,
  and the missing-payload branch exits early for background jobs.
- Update the EDR posture reference for the operator-free command.

* test(claude): run the Windows hook host legs for real

The live Windows host legs never ran: runProcessSync cannot take a string
stdin (it forces encoding 'buffer'), so every leg threw before starting a
host. Use async runProcess, pass PATHEXT (without it Windows PowerShell 5.1
prints nothing and exits 0 for a .cmd path), and name the host in each
assertion. Drop the POSIX pwsh leg: its drive-mapping shim proved nothing
about Windows, and the Windows legs cover both PowerShell hosts.

---------

Co-authored-by: latte271 <junghyeyun27@gmail.com>
Co-authored-by: Bing.Z <zzb@gxsmjx.com>
2026-09-30 01:07:09 -07:00
Brennan Benson 59c05d32b0 fix(rate-limits): stop driving a hidden Codex TUI to read usage (#23806)
* fix(rate-limits): stop driving a hidden Codex TUI to read usage

When the headless Codex usage call failed, Orca opened a hidden
interactive Codex, typed /status and pressed Enter without reading the
screen, then killed it after 15 s. If Codex showed its "Update available"
prompt, that Enter picked "Update now", Codex started its installer, and
the 15 s kill interrupted it, leaving the global install broken.

Drop the hidden-terminal fallback. When the headless call fails with a
non-sign-in error, read the same usage from the HTTP endpoint Orca
already calls on WSL and for the 5-hour window, so a usage check can no
longer answer any Codex startup screen.

Fixes #17415

* test(rate-limits): pin the RPC error when the Codex HTTP fallback fails

Also drop comments that still described the removed hidden-terminal fallback.

* test(rate-limits): use real Response objects in Codex fetcher tests

The changed-code gate rejects the new type assertions this PR added.
2026-09-30 01:04:43 -07:00
Brennan Benson 444e0b1cf9 fix(codex): recognise Codex's quoted spellings in config.toml, and repair Orca's duplicates (#22592) (#23958)
* fix(codex): recognise Codex's quoted project-trust spellings in config.toml (#22592)

Codex's settings screen writes project trust as ["projects"."/p"] and
"trust_level" = "trusted". Orca's matchers only knew the bare spelling, so a
trust write appended a second [projects."/p"] table (or a second trust_level
line) and every codex command then failed with "duplicate key". The config
mirror kept both spellings in Orca-managed homes for the same reason.

- Project table headers are now read through the existing TOML key-path
  parser, so bare, quoted, literal-quoted, mixed and spaced spellings are the
  same table for trust writes and the managed-home mirror/dedupe.
- trust_level is found by decoded key, in both the trust writer and the
  mirror's trust reader, and an existing key is rewritten, never duplicated.
- On the next trust write, a table older Orca appended (exactly
  [projects."<p>"] holding only trust_level = "trusted") that duplicates the
  user's table, or the bare line it inserted under a quoted "trust_level", is
  removed; the user's table wins and the atomic writer keeps config.toml.bak.
  Any other duplicate, or a repair that would still leave one, leaves the file
  untouched and logs once.

* build(cli): list the new Codex trust modules in the CLI project

* fix(codex): recognise Codex's quoted hooks.state spellings and repair Orca's copies (#22592)

Codex writes hook trust as ["hooks"."state"."<key>"] (and the parent as
["hooks"."state"]). Orca's hook-trust writer, parent-table check and mirror
only knew the bare spelling, so a hook-trust write appended a bare copy and
the file failed to parse with "Cannot declare ... twice".

- The hooks.state header, parent-table and mirror checks now use the TOML
  key-path parser, like project tables.
- The duplicate repair now also removes Orca's own hooks.state tables (an
  exact [hooks.state."<k>"] with only enabled + trusted_hash, or an empty
  [hooks.state]) that repeat a table in another spelling, and runs on hook
  trust writes too, so a file with both project and hook duplicates is fully
  repaired. The Orca-shaped copy is removed whichever order the two tables are
  in, only when exactly one other table (the user's) remains; anything else is
  left untouched and logged once.

* fix(codex): carry plain-Codex plugin and project hook trust into Orca's Codex homes (#22592)

Codex keeps hook trust in $CODEX_HOME/config.toml under hooks.state, keyed
by the hook's source. Plugin keys (`id@mkt:path`) and project keys
(`<repo>/.codex/...`) are the same in every home, but the mirror dropped
every hooks.state table from ~/.codex, so Codex inside Orca asked users to
re-trust plugin and project hooks they had already trusted in plain Codex.

- classifyHookTrustKey splits keys into home-scoped (the home's own
  hooks.json/config.toml, re-keyed by install as before) and shared.
- The mirror now carries shared hook trust from ~/.codex in every spelling.
  A key the managed home already holds keeps the managed copy, a key
  repeated in ~/.codex is carried once, and the parent [hooks.state] table
  is never copied, so the result never declares a table twice.
- mergeSystemCodexConfigIntoRuntime moves to codex-config-mirror-merge.ts
  to keep codex-config-mirror.ts under the line limit.
- Tests cover plugin/project carry in each spelling, user-hook keys staying
  out, repeated launches, managed-copy precedence, Windows key spellings,
  parent tables, and user-hook trust re-keying (trusted_hash and enabled)
  from every ~/.codex spelling.

* fix(codex): carry session_end and interrupt hook trust into Orca's Codex homes (#22592)

The shared-trust classifier parsed hook keys with Orca's own trust-key
parser, which only knows the ten events Orca installs hooks for. Keys for
Codex's session_end and interrupt events did not parse, so their plugin and
project trust was treated as home-scoped and left out of the managed home.

The classifier now reads the source path from Codex's key shape
`{source}:{event}:{group}:{handler}` for any event label. A key without
that shape is still never carried. Tests cover both events for plugin and
project keys in both spellings, user-layer keys for both events, and five
unattributable key shapes.
2026-09-30 00:09:26 -07:00
Brennan Benson 5b3366f78e test: unit tests can no longer write a developer's real agent or Orca settings (#23979)
* fix(agent-trust): write per-user trust under the home the launched agent reads

The Cursor, Copilot, Qoder and Antigravity writers and the local Codex config list
resolved ~ with os.homedir() at write time, so any test that reached them wrote into
the developer's real ~/.codex, ~/.cursor, ~/.copilot or ~/.gemini. Each writer now
takes the home, derived once from the launch env (HOME, or USERPROFILE on Windows,
else this host's home) by launchedAgentHome, which the SSH relay already used.

* test: give tests that wrote the real agent or Orca home a temp one

The structured Codex adoption replay pre-trusted /repos/workspace-1 in the real
~/.codex/config.toml; it now runs with a temp HOME and userData. The Codex
session-resume and WSL hook tests created Orca's managed Codex home under the live
userData, and the Claude Agent Teams tests wrote their tmux shim into ~/.orca; each
now runs against a temp userData or HOME.

* test: fail any unit test that writes the real agent or Orca home

A vitest setup file wraps the node:fs mutating calls and refuses a target under the
account's real ~/.codex, ~/.claude(.json), ~/.orca, ~/.cursor, ~/.copilot, ~/.gemini,
~/.qoder or Orca userData, found through os.userInfo() so a test that swaps HOME
cannot hide it. The refusal is recorded and rethrown after the test, since trust
writers swallow errors. Reads are untouched. It stands down only while an opted-in
real-agent suite's own switch is set. It also unsets what an Orca terminal exports
toward the live app (userData, Codex and Claude homes, and the Codex launch preflight
CLI, which a shell test would otherwise run), so a local run matches CI.

* test: type the guarded fs call from its narrowed original
2026-09-30 00:02:35 -07:00
Neil d414033400 fix(packaging): stop shipping the relay bundles inside app.asar (#24027)
resources/relay is the only relay copy a packaged build resolves, but out/relay
was also packed into app.asar — 14.2MB of unreachable duplicate. Kaspersky
flagged app.asar as a compound object precisely because relay.js was inside it,
so one script-heuristic verdict on relay.js gutted the whole install. Excluding
it decouples app.asar from that verdict and drops the duplicate bytes.
2026-09-29 23:30:46 -07:00
Brennan Benson e594cb06af test(mobile): record RPC goldens without a pinned commit, and check recorded requests against the desktop's params rules (#23732)
* test(mobile): add rpc:diff to decode what a recording change moved

The RPC recording goldens are content-addressed JSON, so their raw git diff is
pool hashes. `pnpm --dir mobile rpc:diff [<base>]` decodes both sides and prints,
per golden, the checkpoint, field and JSON path that moved with both values,
grouped across checkpoints, plus added and removed goldens. `--summary <file>`
appends a Markdown report capped for GitHub's step-summary limit.

It reads any pooled format, so it can prove the next commit's format change
moves no recorded value. Checkpoints are matched by occurrence because an id can
repeat within one golden.

This commit adds files under the recorder directory, which moves the header
digest every golden pins; the next commit removes that header.

* test(mobile): record RPC goldens without a pinned commit or input digests

Every golden carried a pinned `baseline` commit plus digests of the recorder,
its mount adapter and its scenario, and the record script refused to run unless
the product tree matched the pin. So every behaviour change repinned to its own
branch commit and rewrote all ~790 files, the squash made that commit
unreachable, and main's pin job stayed red until a hand-made repin pull request
landed (22 of them in 12 days). The digests could only fail when an input moved
and the recording did not, which is exactly the change that carries no
information; every run already re-derives each golden from the current tree and
compares it.

Format 6 keeps the format version, operation, family, named deltas, the value
pool and the recording. Removed: the pin and fence, the three digest modules and
their test, the pin guard and its CI job, and the dead scenario `version` field
(the manifest reader now refuses `baseline` and `version` with a message).

- `pnpm --dir mobile rpc:record [<golden-id>...] [--prune]` records all or some
  goldens; orphans are listed, and deleted only with `--prune`. Every derived
  test title now starts with its golden id so an id selects it.
- `compareGolden` reports every difference in one failure (identity fields by
  name, the checkpoint list, each checkpoint/field/path grouped), keeps the
  final byte compare, and ends with the command to re-record that golden.
- `unhandled-recording.test.ts` now drives a detached rejection through
  `runRecording` into a checkpoint and the cleanup checkpoint; no golden carries
  one, and disconnecting the capture passed every suite before.
- Seam rules that existed only to keep a digest honest are gone; the
  mutant-reachability, register-completeness and one-exposure rules stay.
- CI: `Mobile tests on main` runs the whole mobile suite on every merge that
  touches mobile/, src/shared/, the root lockfile or the host RPC paths, since
  `verify` never runs on main. A new `Mobile RPC Recording Replay` workflow
  replays the recordings on pull requests that touch src/shared/ or the root
  lockfile without touching mobile/. `verify` writes the `rpc:diff` report to
  the job summary.

Proof: `rpc:diff` against the parent reports no recorded behaviour moved; each
golden only loses its ten header lines.

* test(mobile): check every recorded request against the host's params contract

The goldens script the host's replies, so a scenario could record a success
for a request the real host would refuse, and a desktop change that tightens a
params schema moved no golden at all.

`recorded-request-params.test.ts` parses every distinct request the corpus puts
on the wire with the host dispatcher's own `parseRpcRequestParams` and the
schema `rpc-params-catalog.generated.ts` binds to that method. It fails on a
method the host lacks, params it refuses, params sent to a method that takes
none (the dispatcher never reads them), and keys the schema silently strips
unless an inventory entry gives the reason; a stale entry fails too. Each rule
is also shown firing on a made-up request, since the corpus has no instance of
three of them. It imports the desktop dispatcher, so it sits beside the other
Node-side tests outside the RN test program, and the params-contract boundary
now exempts test files, which are never bundled.

It found twelve requests the host would refuse, all from invented fixture
values, not product code, fixed at their source:
- git.branchDiff sent `base-oid`/`head-oid`/`merge-base` where the host needs
  full object ids (diff-review and source-control adapters, and the branch
  compare replies in the manifest that feed them);
- an iOS push registration without `apnsEnvironment`, which a real iOS token
  always carries (`push-token.ts`); the adapter now defaults to `production`;
- `settings.update` given Linear's `assigned` filter as a GitHub preset, which
  the product type forbids; the scenario now picks `my-issues`;
- GitLab `projectRef` as a string where the host and the product type take
  `{ host, path }` (7 methods, 5 adapters and the manifest).

46 goldens move, and a decoded comparison of every one of them shows no change
other than those substitutions; `rpc:diff` lists them.

* ci(mobile): detect a mobile change without a SIGPIPE-prone grep pipe

Under the runner's pipefail, grep -q exiting on its first match SIGPIPEs git
diff on a long file list, so a large pull request touching mobile/ read as
uncovered and replayed the recordings a second time.

* test(mobile): drop comments that still describe the golden header and digests

Eleven adapters justified an import rule by the header a golden no longer
carries, and that rule's test is gone. The census failure now names the
rpc:record and --prune commands.

* test(mobile): refuse a golden that keeps a key no recording writes

Decoding dropped unknown top-level keys, so an old header left behind by a
hand-resolved merge conflict passed every compare unseen.

* ci(mobile): summarize RPC recording changes after a failed test step too

* test(mobile): stream rpc:record output instead of capturing it

A captured run stayed silent for its whole duration and clipped its tail,
where the failure summary sits, past 8 MB.

* test(ci): let the Ruby-gate contract skip the always-run RPC summary step

fef088d8f4 gave the summary step an `if: ${{ !cancelled() }}`, and this test lists every gated step in `verify` and expects each to be gated on the Ruby scope.

* test(mobile): replay only a golden file that is exactly what rpc:record writes

Replay compared two re-encodings of decoded values, so anything decoding drops (a leftover header
key, a hand edit) sat in the committed file uncompared; a key allow-list covered one case of that.
Replay now passes only if the file text equals the formatted golden for the run, sharing one
formatter with writeGolden, and keeps the field-level report as the failure message. The allow-list
goes; the value-based compareGolden stays for the bridged run, which has no file.

* test(mobile): end a corrupt or hand-edited golden's failure with the re-record command

A hand edit to a pooled value failed in decode with only "Golden value <hash> does not hash to its
pool key": no golden id and no command to fix it. readGolden now prefixes parse and decode failures
with the golden id and ends them with the rpc:record command. The rpc:diff header also said it
always exits 0; it exits non-zero when git or a golden cannot be read, and now says so.

* ci(mobile): run Mobile Checks on every src/shared and root lockfile change

Replaces the replay-only workflow: mobile imports hundreds of shared modules, so a shared edit can
move a golden or break mobile's typecheck, and the full job catches both before merge. A root
lockfile-only change skips the Ruby release checks, which read no root Node dependency.

* test(mobile): list or prune orphaned goldens even when the recording run fails

Orphans come from the manifest, not the run, so a failed or timed-out rpc:record still reports
them; the exit code stays non-zero. README: say what a failed replay reports (first differing
path per field, capped groups) and what rpc:diff compares with and without a base.

* test(mobile): pin the RPC recording goldens to LF so a CRLF checkout still replays

Replay now requires the committed golden text to equal exactly what rpc:record writes, which is
LF. A Windows checkout with core.autocrlf=true converted every golden to CRLF and failed all 790
with "holds the same recording but is not the file rpc:record writes for it".
2026-09-29 23:26:55 -07:00
Neil fb52c0602a fix(terminal): release xterm's DEC 2026 render hold instead of waiting out its 1s timeout (#23920)
* fix(terminal): release xterm's DEC 2026 render hold instead of waiting out its 1s timeout

xterm paints nothing while DEC mode 2026 (synchronized output) is open and only
force-flushes after 1000ms. Codex wraps every draw in mode 2026, so any byte gap
or chunk split that loses the closing \x1b[?2026l freezes the pane for a full
second and then repaints in one burst.

Orca never emitted \x1b[?2026l anywhere, and three paths could destroy a TUI's:
the per-PTY pending cap drops buffered output wholesale (mode 2031 was already
salvaged there, 2026 was not), main sliced pending data at a blind 16KB offset
that can land inside an open frame or sever the 8-byte marker, and the renderer's
backlog warnings replace a queued tail that may hold the close.

- salvage the 2026 latch across dropped output, mirroring the existing 2031
  salvage, and append the release on both delivery sites
- ground 2026 in RESET_AFTER_BYTE_GAP and the replay baseline, and in both
  backlog warnings, so every drop path is self-healing
- make main's 16KB flush split frame-aware instead of a blind byte offset
- lift the synchronized-output scanner into shared/ so main and the renderer
  use one implementation

Closing a frame early costs one premature repaint; leaving it open costs a
second of blank screen, so the asymmetry favours always closing.

Also adds the reproduction this needed: the pre-existing typing bench observes
the xterm BUFFER, which the parser fills while rendering is held, so it scored
these freezes as fast echoes.

* fix(terminal): stop the renderer's queue drain cutting inside an open DEC 2026 frame

takeQueuedChunk sliced a queued chunk at a blind byte offset to fit the 16KB
coalescing budget, which can strand a frame's closing \x1b[?2026l in the residual
until a later drain. Same defect as main's flush split, same fix: reuse the
frame-aware split helper.

Usually masked because the drain coalesces adjacent chunks and reassembles what
main split, but not when the budget boundary falls inside a frame.

* fix(relay): keep the SSH path's bounded slice outside an open DEC 2026 frame

pty-handler split pending output at a byte offset with a surrogate-pair guard but
no synchronized-output awareness, so a frame straddling the 16KB wire slice had
its closing \x1b[?2026l stranded in the remainder — the same defect just fixed on
the local path, on the path AGENTS.md requires us to consider.

Placed before the surrogate guard so that guard keeps the final say, and floored
at 2 so frame alignment can never walk a healthy slice into the guard's
decrement and then into the chunkChars <= 0 pause-and-retry path.

Also drops a dead `splitAt === 0` branch in takeQueuedChunk: both callers pass a
positive limit and the helper never returns 0 for one.

The two new split tests were each confirmed to fail without their fix.

* test(terminal): sweep the DEC 2026 split helper over escape-sequence shapes and every limit

Covers OSC 52, DCS, repeated open/close markers and limits 1..len+3, asserting the
result never exceeds the limit, never reaches 0, and stays byte-exact. Also pins
that a buffer beginning inside an open frame degrades to the blind offset rather
than doing something worse, and documents that callers do not thread latch state.

* fix(terminal): ground DEC 2026 on the daemon slice, the recovery replays, and the process boundary

Four more sites could strand the latch, found by sweeping every path that drops,
splits, or replays terminal bytes.

- daemon-stream-data-batcher: the 64KB bulk-write slice used a surrogate-only
  clamp, and its remainder is HELD until 'drain' — "seconds for multi-MB
  backlogs" per the file's own note. A frame straddling that boundary parked its
  \x1b[?2026l behind the hold, blanking the pane past xterm's 1s timeout once per
  frame for as long as the backlog lasted. This is the default daemon-backed pane
  path, so it is the one users actually hit. The new
  clampToSafeBulkWriteSplitIndex frame-aligns first and surrogate-clamps last,
  and lives in daemon-stream-data-split alongside the policy it belongs to.
- replay-data-drain and remote-runtime-terminal-binary-snapshots wrote a bare
  \x1b[2J\x1b[3J\x1b[H, which does not clear mode 2026 — so on the SSH/remote
  reconnect path, the very event most likely to sever a frame, the whole replay
  could paint nothing.
- ipc-pty-attach: trimIncompleteTerminalControlTail can cut a half-written
  \x1b[?2026l while its opening marker survives in the replayed prefix.
- PROCESS_BOUNDARY_GROUND: the "process that armed these modes is gone" ground
  omitted 2026, the last unexplained gap in that file. A disable, so it still
  satisfies the recovery barrier's ownership scan (only ?25h may be an enable).

Recovery-path expectations updated where they pin the emitted bytes. Deliberately
NOT touched: apply-reattach-payload and ssh-snapshot-prepaint already ground via
buildSnapshotReplayPrologue.

Still unfixed, deferred with reason: terminal-output-frame-chunks.ts splits the
remote wire on accumulated UTF-8 byte width and needs a different shape than the
char-index helper; desktop clients reassemble in main's pending buffer, so the
exposure is mobile/web only.

* fix(terminal): emit the DEC 2026 release before the mode-2031 tail, and stop claiming the drop path writes it

Two corrections from adversarial review of the earlier commits.

1. Ordering bug I introduced. getDroppedMode2031RendererData ends with
   `state.tail`, which extractPrivateModeScanTail deliberately retains as an
   INCOMPLETE private-mode sequence so the next chunk can resolve it. Appending the
   2026 release after it put an ESC behind a dangling CSI, aborting it and silently
   losing whatever mode spanned the drop boundary. The release now goes first.

2. The drop-path release does not reach xterm in the dominant case, and the comment
   now says so instead of implying otherwise. live-data-callback's droppedOutput
   branch discards `data` and salvages only queries
   (salvageRendererQueriesFromDiscardedRestoreData handles CPR/DA1/OSC colour;
   \x1b[?2026l is not a query), so for hidden panes and visible panes outside
   foreground-restore backpressure the synthesized release was dropped. The grounded
   snapshot replay releases the latch instead.

   I tried writing it through writePtyOutputToXterm there and reverted: it consumes
   the pending hidden-output snapshot and broke
   pty-connection-hidden-snapshot-resize-signals ("re-restores a skipped alt frame"),
   so the release rides the restore rather than perturbing that state machine.
   Residual gap, documented: a cap-dropped pane whose restore never arrives.

The salvage is still load-bearing on the fall-through path, so it stays.

* fix(terminal): release DEC 2026 on the reattach clears, floor the split, and correct the freeze framing

Remaining findings from adversarial review.

- apply-reattach-payload's three bare-clear branches (:63 daemon snapshot, :229
  relay replay, :269 cold restore) had no release anywhere in their sequence: I
  checked all seven POST_REPLAY_* profiles reachable via chooseReattachReplayReset
  and none contains \x1b[?2026l. Only the buildMainModelSnapshotReplayWrites branch
  was grounded, so covering the streamed replay path and not the main reattach path
  was inconsistent. Verified no production code matches these clear strings — the
  three test updates are mock equality, and each was confirmed to fail without the
  source change.
- clampToSafeBulkWriteSplitIndex could return 0 (('\u{1F600}aaaa', 1) — alignment
  returns 1, the surrogate clamp decrements to 0), which would leave a zero-length
  slice that never shifts the batcher's queue entry and spin its drain loop.
  Unreachable from today's only caller, but it is exported with an unstated
  precondition. Floored at 1.
- Frame alignment could halve per-PTY flush throughput: main re-queues the
  remainder with eligibleRound = round + 1, so the shortfall cannot be refilled in
  the same round, and aligned size is floor(W/F)*F — 50% worst case in the 8-16KB
  band, which is exactly the full-screen redraw burst that reaches the pending cap.
  Alignment is now rejected below half the window, preferring throughput and
  letting the reset profiles release the latch.

Framing corrected throughout: bufferRows records a row range and clears nothing, so
the pane freezes on its last painted frame — it does not go blank. The real trade is
"stale but coherent for <=1s" versus "immediate partial frame", and
RESET_AFTER_BYTE_GAP (written alone, with no repaint behind it in the same write) is
the one site that can newly flash a partial frame. Said so at the constant instead
of implying the release is free.

* fix(terminal): rename the shape-flagged symbols the anti-slop audit rejects

CI's anti-slop gate rejects "shape" in symbol names as structural rather than
domain language: `shapes` -> `outputSamples`, and
`writeCodexShapedEchoProbeScript`/`codexShapedEchoProbeScript` ->
`writeCodexEchoProbeScript`/`codexEchoProbeScript`.
2026-09-29 20:27:30 -07:00
Neil 45c63a66e9 test: delete the source-grep tests an earlier detector's regex missed (#23976)
A rebuilt detector found 195 source-grep candidates where the original found 111.
The gap was one over-specific regex: the first scanner required a literal `.ts`
path inside `readFileSync(...)`, so every test that built its path from variables
(`join(dirname, '..', 'foo.tsx')`) was invisible to it. Roughly 84 files of a
pattern an earlier wave reported as cleared had in fact survived.

Deleted whole, every case asserting on production source text:
- `app-startup-routing.test.ts` (27 cases) — exact import statements
  (`"import('../components/UpdateCard').then"`), relative-path spelling, and
  `indexOf` source ordering. A file move or a `lazy()` refactor breaks it.
- `pull-request-page-host-boundary.test.ts` (13) — `toContain` on whole argument
  expressions concatenated across 20+ component files.
- `SmartWorkspaceNameField-source-boundaries.test.ts` (7) — placeholder copy, a
  Tailwind class string, and `not.toContain` on an already-deleted symbol.
- `github-project-repo-list-load.test.ts` (9) — `indexOf` statement ordering
  inside `loadTasks`.
- `github-enterprise-slug-routing-boundary.test.ts` (4) —
  `toContain('host: githubProjectHost(parsed?.slug.host)')`.
- `web-viewport-shell.test.ts` (3) — a regex demanding exact CSS selector-list
  ordering and whitespace.
- `agent-catalog-links.test.ts` (1) — restates two `homepageUrl` literals straight
  out of `agent-catalog.ts` with nothing in between.

Trimmed, keeping only what nothing else can reach:
- `desktop-startup-ordering.test.ts` 549 -> 66 lines, retaining the three cases
  named as `assertionRefs` by the `ssh-filesystem.stream-inactivity-lifecycle` and
  `agent-browser.owner-boundary-cleanup` gates; 15 source-order greps went.
- `ResourceUsageStatusSegment.session-polling.test.ts` keeps its census that no
  `setInterval` exists and `listSessions()` is called exactly once — an added poll
  multiplies a global daemon scan and no behavioral test sees it. The
  `indexOf('if (!open)')` ordering pair and four `not.toContain` lines went.
- `agent-skill-installed-command-callers.test.ts` 231 -> 86, keeping the
  `readdirSync` census that discovers every `<AgentSkillSetupPanel` caller and
  asserts set equality against the allowlist, so a new panel host cannot silently
  show a default Update action.

Also in this wave, from the renderer lib/runtime sweep: 22 cases whose routing
signal the production path never reads — verified by mutation, stripping
`connectionId`, the WSL preference and the UNC path from four of them left all 29
tests passing — plus braille-spinner rows collapsed onto one regex range, copied
`WELL_KNOWN_LABELS` rows, and a whole `resolveAiVaultResumeStartupShell` describe
whose four darwin/linux fixtures all return before the login shell is read.

`config/reliability-gates.jsonc` drops the two `app-startup-routing.test.ts`
references; the manifest still validates for 140 gates.
2026-09-29 19:55:50 -07:00