Commit Graph
492 Commits
Author SHA1 Message Date
Neil e6fbbdf684 perf(ci): cache pnpm verification records on Linux (#23568)
* perf(ci): pilot pnpm verification record caching on Linux

* test(ci): review pnpm verification record in mobile cache audit
2026-09-28 01:50:30 -07:00
Neil 080c562898 perf(ci): diff against the merge commit's first parent so PR checkouts can be shallow (#23562)
Every changed-path gate asked git for `--merge-base "$BASE_SHA" "$HEAD_SHA"`,
which needs the event payload's base SHA to be in the local graph. That is the
only reason two jobs cloned all 8127 refs' history. On a pull_request checkout
HEAD is already the merge commit, so its first parent is the base side and no
merge base has to be computed. config/scripts/git-pull-request-diff-base.mjs
resolved that for the two Node gates; the workflow's inline gates now use the
same helper through a small CLI rather than open-coding it.

code_paths gates all 22 jobs, so its checkout is charged to the start of every
one of them: measured 20.7s to 1.6s, keeping blob:none because its sparse tree
is ~7 files and leaves no blobs to refetch. Static analysis drops the filter
instead, since populating all 30,226 files makes blob:none force a second
promisor fetch: 23s to ~11s.

Verified on a real merge ref. At depth 50 the old and new forms produce
identical changed-file sets. At depth 2 the new form still works and the old one
fails with `fatal: bad object`, which is the failure a stale base would have
caused once the checkout stopped being complete.

Also drops the dead resolveBase + merge-base prelude in the changed-code gate,
whose result resolvePullRequestDiffBase already discarded on every PR.
2026-09-28 00:37:26 -07:00
Brennan BensonandClaude 7a24d3d335 fix(native-chat): the conversation outlives its agent (#22835)
* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* fix(native-chat): name a chat write by its target, not the owner generation

A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.

Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.

Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.

* fix(native-chat): every journal append reaches the chats that are open

A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.

A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.

* test(native-chat): an epoch replacement reaches the open chat

* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

* perf(native-chat): a publish behind a delivered commit reads nothing

Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.

* test(native-chat): state why the teardown test's fake journal is safe to cast

* docs(native-chat): say mutation admission checks only the writer lease

* docs(native-chat): drop the send rebase from comments that still described it

* fix(native-chat): a message is accepted, then delivered

A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".

A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.

Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.

A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.

Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.

* fix(native-chat): settle queued messages only for the child that ended

A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.

A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.

The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.

* fix(native-chat): an adoption that fails to import keeps the conversation open

The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.

* perf(native-chat): the recovering open reads the journal once

Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.

* fix(native-chat): an attach that fails after indexing its child leaves no child behind

A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.

* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer

The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.

A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.

* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down

The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.

* fix(native-chat): a message rejected while its chat was closed reads as not sent

A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.

* test(orchestration): name why the readiness settlement fakes are cast

* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent

* docs(native-chat): drop the fence from the admission the send effects run behind

* docs(native-chat): give the fence move on release the reason that still holds

* docs(native-chat): stop citing a write fence check in launch and mailbox comments

Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.

* refactor(native-chat): the provider child is its own record

A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.

- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
  patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
  own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
  the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
  dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
  is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
  the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.

* fix(native-chat): the delivery loop alone settles a message its start or child failed

A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.

- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
  reads how it ended: a Stop continues; anything else writes one failure row and rejects every
  queued message with the same words, then stops. A child still starting whose start the adapter
  says did not land fails the same way. The exit, eviction and the settlement retry only settle
  the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
  nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
  failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
  closed, with or without a child, and a start the loop already has in flight is waited for so the
  child it produces is stopped rather than left behind.

* refactor(native-chat): a stopped child ends on the one reading of its stop

The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.

* feat(native-chat): the host says it accepts a send before any agent has it

The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.

* refactor(native-chat): an attach never opens a journal of its own

The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.

* fix(native-chat): a moved fence resends nothing on a host that accepts first

The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.

The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.

* refactor(native-chat): a child's end says whether the user or the host stopped it

The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.

* fix(native-chat): a chat whose only work is a queued message is not offered for resume

A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.

* fix(native-chat): the conversation outlives its agent

Opening a chat no longer starts its agent. A conversation is reached through one host
accessor that opens its journal at rest, and a send is what starts the agent, through
the delivery loop. One idle sweep, every five minutes, stops an agent that has been
quiet for thirty minutes and owes no work, then drops an open journal handle that is
only a cache. Its record, tab, status row and readers stay.

- hold and release are no-ops; hold still builds the host for shipped mobile builds.
- The holders, the holds, the release clock and the exit respawn are deleted.
- Options, the model list, the goal and the context meter answer at rest; a model pick
  at rest is recorded as intent for the next start.
- Compact, rewind, clear and goal changes start the agent first. A send does too when
  a rewind is still in doubt after the conversation opens.
- Orchestration routes mail and group addresses on ownership (the record plus the chat
  tab), not on whether the process runs. An open dispatch keeps its worker running.
- The restart continuation is a send; Resume all holds each slot until the message is
  handed over or rejected.
- A read error never replaces a loaded transcript, and shows the host's own words.

* test(native-chat): type the queued-message fixtures in the resume-offer tests

* fix(native-chat): a start that dies while a message waits on it is that message's failed start

Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.

* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now

* test(native-chat): pin what a failed start settles, and what a resume offer names

A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.

* test(native-chat): the failed-start pins fail on what the message became, not on a timeout

* fix(native-chat): a restart offer ends when the chat's agent starts again

The offer used to end only when the chat's newest user message changed,
because opening a chat started its agent and that start could not be told
apart from real activity. Opening a chat starts nothing now, so the host
reads the fact it already publishes: a chat's status row goes from not
host-owned to host-owned exactly when its agent is started. At that edge the
offer and any failure record for the chat are withdrawn, unless the start is
a resume action's own (its continuation is the oldest undelivered message).

A continuation and a message racing to be first are decided at acceptance:
the continuation is refused, quietly and with nothing filed, when any other
message was accepted since the restart. A failed continuation start leaves
the offer retryable, and each resume action sends its own message id.

Deleted: the newest-user-message comparison, its journal reader, the
continuation filter, and the failure ledger's own "answered by the chat"
check. The marker still carries its message id for one release, so the
previous build can read it.

* fix(runtime): end a transcript stream when its client unsubscribes

Desktop: the IPC subscription controller was dropped as soon as the streaming
handler returned, which for most streams is right after it binds. A later
runtime:unsubscribe then found nothing to abort, so the host kept the subscriber
and derived and sent every publish to a channel no one listened to. The controller
now lives until the renderer unsubscribes, resubscribes the same id, or goes away.

Mobile: disposing an agentSession.subscribe stream now sends agentSession.unsubscribe
with the stream's frame id, so the host ends that subscriber and leaves a sibling
stream on the same socket running. The direct path now passes the frame id the relay
path already passed.

* test(orchestration): the preamble's host stub is typed, not cast

The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.

* fix(native-chat): one fact ends a restart offer: the chat moved on since the restart

The offer is live while no other message has been accepted in the chat since the
restart and its agent has not proved a start since. The offer list, the resume's
reservation check and the continuation's acceptance check all read that one fact,
so a message whose start then failed withdraws the offer too, and a stale click
finds nothing to act on.

The fact is read off the conversation's open handle, which the restart closed, so
it is retired durably whenever it may have changed: a message accepted, a start
proven. A close and reopen within the same run therefore cannot bring the offer
back. A continuation rejected before it reached the agent does not count, so a
retry after a failed start still runs.

Deleted: the quit-time gate on withdrawal, which changed nothing because the
withdrawal and the quit's own offer write share one queue; the per-action
"withdrawn" flag and the separate acceptance check it paired with.

* test(native-chat): an older build reads the restart offer this build records

The offer lives in a file the previous release reads after a downgrade. Pin that
against the pinned release's own capsule, and run the lane when the marker or the
capsule changes.

* fix(native-chat): read a restart offer against where the journal stood when it was taken

"Since the restart" was read off the conversation's open handle, which the idle
sweep closes: after a reopen, a message the user had already sent looked older
than the handle and the withdrawn offer came back.

The offer now records the journal position (epoch and sequence) at the moment
it is taken, and a message accepted after that position, or a journal on another
epoch, means the chat moved on. That is derived from the journal, so it holds
across any number of closes and reopens. An older build's offer has no position;
only a start withdraws it. Because the message half is now durable, the offer is
no longer rewritten in the recovery file on every accepted message; a proven
start still writes it, since only the host that saw the start knows of it.

* test(native-chat): wait for the listing's retire write before reading the recovery file

* fix(native-chat): keep the terminal-backed chat's read error over its local echoes

Messages winning over a read error is right for the structured chat, whose read retries and whose
messages came from the transcript. The terminal-backed view assembles its list from local echoes
too (a launch prompt, a pending send), so a failed read there showed only those bubbles and no
error. Only the structured pane now keeps messages over an error.

* fix(native-chat): a start retries the exit settlement a failed journal write left owed

An agent exit whose journal settlement write failed releases the lease latched until a retry lands.
Reopening the chat used to be that retry; with reveal now only opening the journal, nothing retried
it before the next app launch, and every send was refused. The start the send needs now runs the
retry first, where the attach would.

* perf(native-chat): answer the owner check without opening the chat

Worktree activation calls agentSession.handoffStatus for every chat tab in the worktree, and the
answer comes from the session record alone. Reaching it through the accessor opened each resting
chat's journal (a full read, the crash-boundary write and a restored status publish), then kept it
open for the idle window. It now checks the record and the adapter's support, as before this series,
and opens nothing.

* fix(native-chat): a read waiting on the session lock opens nothing once quit began

The accessor checked for quit before queueing the open, so a read queued behind a session task ran
its open after teardown had begun and indexed a journal no teardown step would close. The check now
runs at the open itself.

* fix(native-chat): read a failed resume's chat before calling it retryable

Whether a failed resume is retryable is the offer's own rule: the chat has not moved on since the
restart, read from its journal. The failure list read it only for a chat already open, so once the
idle sweep closed a chat the user had moved on in, its failure showed Retry again, and the click did
nothing. The list now opens the failed chats first, as the offer list does.

* test(native-chat): type the provider event sink the settlement test reaches for

* fix(native-chat): say the structured read keeps trying only where it does

The structured pane's "Orca keeps trying to load it" line never showed: the view state filled in an
untranslated fallback whenever the read error had no text, and the empty state prefers any message.
The view state now leaves the message out, so the structured pane shows that line and the
terminal-backed pane its own translated one. Mobile's structured lane does not resubscribe after an
error frame, so it no longer makes the claim.

* test(native-chat): await the send's settlement instead of polling for the start

The at-rest send tests polled for the provider start with vi.waitFor's one-second default, which a
loaded machine outran. They now await the host's own settlement of the message.

* fix(native-chat): a restart offer resumes any time after the quit, and knows its own continuations

The continuation's message id was dated by the quit, and the ledger refuses a new id dated more than
a day back, so Resume or Retry a day after quitting was always refused (on main too). It is now
dated by the resume action.

Telling a rejected continuation from the user's own message read the operation ledger, whose rows
expire after about a day; after that a failed resume stopped being retryable. The offer now
records the continuation each action sends on its own capsule entry, bounded to the newest 16, so
the ids end with the offer. The ledger read is deleted.

* fix(orchestration): route no mail to a structured worker its orchestration released

A structured worker is routed on ownership, and a resting worker's lease is released, so ownership
held while its chat tab stayed listed. A worker the coordinator abandoned and then released, found
at rest by the release, therefore still took peer mail and @worktree: broadcasts, and each one
restarted its agent. Routing now also reads the orchestration's own resource row: once it is
released, direct mail, group addressing and worker-show's addressable answer drop the worker, as
they would a terminal worker whose terminal closed. The chat tab stays, and nothing new is stored.

* fix(native-chat): a failed retry names the user's prompt, not Orca's continuation

A resume's continuation is written to the chat before its start, so after a failed attempt the chat's
newest user message is that rejected continuation. A second failure then showed Orca's own restart
text as the chat's prompt. A retry now keeps the prompt its first failure named.

* fix(orchestration): read the released row optionally, as the authority does

worker-show's observation called the row lookup directly, which a runtime double without it threw on
and failed the structured tab-retirement release.

* fix(native-chat): the status bar drops a restart offer the chat moved on from

The renderer re-read the host's restart offer only when a failed chat showed activity, so after a
message withdrew a pending offer the host answered no chats while the status bar kept counting one,
and clicking it opened nothing. The same watch now covers pending offers: a status change in an
offered chat asks the host again, once.

* test(native-chat): a roster of idle or finished children does not keep an agent awake

The sweep reads owed background work through the shared child-work liveness that upstream's
release clock adopted; a child that went idle or finished is not work the agent still owes.

* fix(orchestration): a task dispatched into a resting structured worker keeps it running

The sweep's open-dispatch check read only the worker-start dispatch that owns the worker's terminal
resource, so a task later dispatched to the same worker (orchestration dispatch --to, which writes a
dispatch with no worker row) did not count: after thirty quiet minutes the worker was stopped while
that task was open, and its coordinator read exited. Any unsettled dispatch addressed to the worker's
process incarnation now counts, derived from the existing rows.

* docs(native-chat): comments stop describing the hold this PR removed

Eight comments still justified orderings and teardown choices by a viewer or dispatch hold that
pinned the provider child. Nothing holds any more; the orderings stand for the binding's redrive
subscription and parked mail, and a chat's agent runs from a send until the idle sweep rests it.
Comment-only.

* fix(native-chat): a restart offer keeps the start its own continuation made

Whose start ended an offer was decided at read time, from whether the offer's continuation was
still the queued message. Once the provider refused that continuation, the child it had started
read as someone else's start, so the offer ended and its failure showed no Retry. The delivery
loop now records which queued message a start is for on the in-memory child, and the child's end
carries it; the offer counts a start as its own when that message is one of its continuations.

* fix(native-chat): an agent gets a full idle window after its owed work ends

The sweep measured quiet only from the last journal row, so once a subagent, command, monitor or
dispatch that had outlived the window ended, the agent was stopped at the next tick. A child can
read done before the lead's wake-up turn writes anything, and stopping in that gap loses the
wake-up. The sweep now counts owed work it observes as activity, which gives the agent the full
window afterwards, as the release clock it replaced did.

* test(claude): the options-read fixture runs a live child

The fixture marked its conversation running with a hasProviderChild field the
session type does not have, so the read took the at-rest path and refused a
session with no record. It now carries a child, which is what the read checks.

* test(native-chat): host tests reach its collaborators through a typed seam

The rest-test rig and three test files read the host's private members with
Reflect.get and cast the result. The host now exposes one test-only accessor,
collaboratorsForTests(), and the subscribers class a subscriberCountForTests()
beside its existing retainedActivityCountForTests(), so the tests are checked
against the real types and the casts are gone.

* refactor(orchestration): one owner answers a structured worker's custody

Routing, group addressing, worker-show and the idle sweep each composed their own reading of
whether orchestration still holds a structured worker, so each new obligation or retirement state
had to be added to every reader. structured-worker-custody now derives both answers from the
worker-terminal list state coordinators see in worker-list: addressable is owned and not released,
and owed work is an active custody or an unsettled task dispatched to the same incarnation. The
owner's state is read through the remote dispatch attachment too, as the terminal transfer lookup
already does. Behaviour is unchanged; a settled worker awaiting its coordinator still rests.

* refactor(orchestration): owed work is an open dispatch on the worker's incarnation

A supervised worker's own dispatch context stays open exactly while the worker is active, so the
separate active-custody branch only repeated it. Owed work is now one fact, which also states the
policy that a worker awaiting its coordinator's decision may rest, and both custody decisions are
written once at the top of the module.

* fix(native-chat): a restart offer knows its continuations by a tag in their id

The offer recorded each continuation id in a list on its capsule entry, capped at 16, and a running
action's id in memory. Both could disagree with the journal: past the cap an old rejected
continuation read as the chat moving on, and a crash during a retry restored the failure's older
entry, which lacked the retry's id. Each continuation id now carries a tag derived from the offer
(its teardown and chat), then the action's own part, so any continuation of this offer, queued or
rejected, is recognised from the journal row and the marker alone. The persisted list, its cap and
the in-memory action map are deleted; the agent-start withdrawal keeps an offer whose own
continuation the start was for, read against the stored marker.

* test(runtime): the legacy-worker reveal test judges its stale snapshot inside the wait

The tui-idle probe reads through readTerminal, which now awaits the structured
worker check before the PTY read, so the probe's snapshot request starts a
microtask later. vi.waitFor missed it on its first check and polled again at
50 ms, the same moment the wait's own 50 ms timeout fired. The stale snapshot
then resolved after the wait had already timed out, so the test passed without
judging it, and the rejection landed before any handler was attached. Vitest
reported that as an unhandled error and failed the shard.

Polling every 1 ms sees the request within a few ms, so the snapshot is judged
while the wait is still pending.

* fix(native-chat): a view never restarts a chat whose last start failed

A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.

* fix(native-chat): the idle sweep reads owed work every tick

Owed work counted as activity, but the sweep read it only once the idle window had elapsed, so it
refreshed the clock at most once a window. Work that ended just before the next read left the
agent to be stopped at that read, moments after the work ended, which is the gap the refresh was
meant to cover. The sweep now reads owed work on every tick for a started agent, so the window
always runs from the last tick that saw work owed.

* fix(native-chat): a continuation handed to the agent stays sent

The offer read its own continuation as not reaching the agent while its dispatch was pending, which
also covered one already handed over and still unanswered. When the wait for that answer ended first,
the failure it filed read as retryable, and a retry sent a second continuation to an agent that may
have acted on the first. Only a continuation still queued, or rejected, is now read as unsent.

* test(native-chat): start the child the loop waits on with an attach, not a second view

A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.

* fix(native-chat): settle a gone generation's turn wherever a conversation opens

A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.

* test(native-chat): prove the next child's start settles the turn an earlier child left

The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.

* test(native-chat): count a failed start's rows by row, not by text

Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.

* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget

* test(native-chat): the interrupted create's own retry continues again

The merge of main's lease-latch fix replaced that test's retry of the interrupted create, under its
own operation id, with a fresh start whose result nothing read. That fresh start passes with the
released-reservation continuation deleted, so the case the fix exists for went untested. The retry
and its assertion are main's again.

* docs(native-chat): three comments that still had views starting agents

A start with nothing queued now comes from a command, goal change or rewind; an interrupted compaction
left alone would refuse every send, so no agent would ever start to finish it; and a current host
raises the unattached read refusal only once quit began, with the attach window belonging to an older
host.

* test(native-chat): pin the open's and the send's start and row counts, however the view binds

Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.

* test(native-chat): a reader's open settles the turn a failed exit settlement left running

An exit whose settlement write failed leaves its turn running in the open journal. PR 1's open now
settles it, and this pins the two reads that reach it here: a reader reopening a chat the idle
sweep closed, and a read that opens the chat before the restart restore reaches it.

* test(native-chat): the view-start test's starting window outlasts two subscriptions on a loaded runner

A subscription reads the conversation before it returns, so under load the two views took longer
than the create child's 300 ms start, which then exited before the test checked that it had not.
The child now takes a second to fail.

* fix(native-chat): settle a gone generation's turn at every open but an acquisition's

The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.

* test(native-chat): hold the create's start open until the views bind

The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.

* test(native-chat): a read that reaches a crashed chat before the startup reconcile settles its turn

On desktop the chat on screen at relaunch reads before startup reconciles the leases, while the
dead process's lease still reads live. The open settles the turn it left running anyway, and the
restore that follows finds it settled.

* refactor(native-chat): drop the composer's second error formatter

After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.

* test(native-chat): pin the reason on a message rejected while its chat was closed

The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.

* docs(native-chat): drop the removed dispatch hold from six comments

A worker's session no longer takes a dispatch hold, and no release clock
rests a chat by visibility; the agent-launch comments, the abandon test,
the teardown test and the refusal census still said so.

* test(native-chat): rest the owner-status chat through the idle sweep, not a hold

The activation-gate test from #22808 put its chat at rest by holding and
releasing it, and passed the release-clock grace. This branch deleted both,
so the case threw before it reached its assertions. It now moves the host's
clock past the idle window and lets the sweep stop the agent and close the
conversation, then asserts the same owner answer and activation gate.

* fix(native-chat): show the structured pane's retrying line when a read fails

The read transport always hands the pane the host's words, so the error
state's "Orca keeps trying to load it" line, which showed only when there
were none, was never seen: the pane showed the host's text twice, as its
subtitle and on the status line under it. The structured pane now always
says its read keeps retrying, and the host's text stays on the status line.
The terminal-backed chat is unchanged.

* test(native-chat): wait for a send's background start before the refusal oracle removes its store

An accepted send wakes the delivery loop, which starts the agent in the background. The oracle's teardown disposed the loop but did not wait for that start, so its lease write could create a temp file in the store directory while the directory was being removed, failing the test with ENOTEMPTY about one run in four. The teardown now drains tracked starts before it closes the journals.

* fix(native-chat): a start a message waited on gets one failure row, the delivery loop's

When a queued message's start failed, two writers could report it under the same row: the delivery loop, when the adapter settled the start without proving it, and the exit settlement, when the child's exit landed. The last one won, so the chat's row could name a different cause than the one the message was rejected with, or be written twice.

The exit settlement now writes the start's row only when no message is queued and the loop has not already recorded that start. A start for a command, goal change or rewind, with nothing queued, still gets its row from the exit.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-27 23:46:14 -07:00
Neil 9179b93ebf ci: reduce repeated runner work and validate affected-test selection (#23540)
* ci: stage heavy checks and measure affected-test selection

* fix(ci): exercise the real Git boundary in unit selection planning

* Harden review cancellation and CI demand reporting
2026-09-27 23:25:20 -07:00
Neil 67581dd090 perf(ci): stop the orcad smoke idling 15s and the static job fetching mobile packages it skips (#23541)
The shutdown race in the orcad terminal smoke never cleared its losing timer, so
the process sat on a live 15s timer after PASS had already printed. Measured
locally: 21.4s -> 6.52s, with the round trip and the shutdown assertion intact.

Static analysis also asked for the mixed root+mobile pnpm store (537 MB, 8.6s to
restore) on every run, while installing mobile dependencies only when the diff
needs them. Most runs paid 216 MB for packages they never linked.
2026-09-27 22:51:59 -07:00
Brennan BensonandClaude 85067494a1 fix(native-chat): a request that failed reads as failed (#22944)
* refactor(native-chat): remove the unused terminal handoff

No client ever called agentSession.requestHandoff or mounted the handoff
chrome. Delete the handoff coordinator, the terminal-owner runtime, the
proof write path and the unmounted UI. Keep agentSession.handoffStatus,
which released desktop clients read for worktree activation, and let
records an older build left mid handoff reconcile through the ordinary
restart and recovery paths.

* fix(native-chat): never let the pre-stop snapshot hold a chat's stop

Eviction now drains delivered events before quit's resume-offer snapshot. An
unbounded wait there sits ahead of the provider stop, so a sink whose journal
write stalls kept the child running until the step deadline aborted the
eviction. The offer is advisory: bound the drain and stop the child regardless.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop helpers only the terminal handoff called

`claudeAuthEnvCarriedForward`, `isPathWithinDirectory` and
`queryWindowsProcessRowsFresh` lost their last caller with the handoff. The
fresh-scan tests now go through `queryWindowsProcessDescendants({ fresh: true })`,
the teardown path that still depends on that contract.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(native-chat): stop citing the removed handoff in lifecycle comments

Six comments still named the handoff coordinator, a handoff suspend, or a
terminal-owned session as live participants in the flows they describe.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stalled snapshot drain without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): pin that a start dead before proving owes no settlement

The removed restart handoff test pinned this branch; nothing else did.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): keep the owner-status read behind an in-flight attach

The handoff removal dropped the per-session queue from `handoffStatus`, so a
read landing mid-start reported the reservation (no owner) instead of the
settled chat owner, and shipped desktop clients blocked worktree activation on
it. The read is queued again, as it was before the removal.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(terminal): remove the agent-session PTY write gate

The gate only refused a write when a PTY had been bound to a chat session, and the
only code that ever bound one was the terminal handoff this branch removes. With it
gone, every admit/readmit returned "admitted" unconditionally, so the checks on the
renderer write path, the runtime controller backstop, terminal.send, agent prompts,
preview input and orchestration pointers, the refusal fields on terminal.send and
worker-start receipts, the plugin and CLI refusal copy, and the adopted-pane
orchestration routing could no longer run. Ordinary writes take the same path in
the same order as before.

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(native-chat): drop the transcript helpers only the handoff called

appendLegacyTranscriptMessages fed the terminal transcript catch-up and
proveClaudeTranscriptBranch backed the terminal owner's exit proof. Both lost
their last caller with the handoff. Their tests now go through the live entry
points instead: the roster bounds through the legacy import, the pinned-read and
growth tests through the ancestry replay the history window uses, and the marker
rules through the string proof in their own file rather than the session-file
resolver's.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(native-chat): stop calling a starting chat "mid-handoff"

A send refused because the chat's owner is not settled showed "The session is
mid-handoff (<stage>)." in the composer. With the handoff gone, the stages that
reach it are a chat that is still starting, or one whose previous agent process
has not yet been confirmed stopped. The message now says which of the two it is.
The refusal code is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(native-chat): type the stand-in roster decoder without a cast

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(codex): name the pinned rollout lookup for what it does

With the terminal handoff gone, the module named codex-tui-rollout-proof holds
only the pinned rollout lookup that structured Codex launches use to resume a
thread, so the name described code that no longer exists. Rename the module and
its options type. Also drop a mobile allowlist assertion that pinned the
removed agentSession.requestHandoff method, which no longer exists to allow.

* refactor(native-chat): type the owner-status reply as the host sends it

The handoffStatus reply type still listed the terminal handoff's fields and
states (terminal placement, host label, proof retry, queued and waiting phases,
the to-terminal direction). No host writes them any more and the only client
reader parses the reply as unknown, so they described nothing. The reply on the
wire is unchanged.

* refactor(native-chat): normalize terminal-handoff lease values once at decode

Nothing in this build writes a terminal owner (`runtimeKind: 'tui'`) or the
handoff's `preparing` / `old-owner-stopped` stages, but the in-memory types
still admitted them, so readers across the host kept branches for values no
path produces and the compiler could not point at them.

The store now validates the on-disk shape, which still accepts those values so
an older record is not quarantined, and maps them once while parsing:

- `preparing` and `old-owner-stopped` become `recovering`
- a `tui` lease becomes `native`; when it records a process it also becomes
  `conflicted`, the claim every build probes but never stops. A plain native
  owner would be stopped by restart recovery, here and in older builds.

Revisions are taken over the normalized state on both sides of every compare,
and the mapped record reaches disk with the store's first transaction, the
same way the tab-id backfill does.

The in-memory types narrow to what this build writes, and the branches that
existed only for the removed values go. Structured-worker identity keeps its
verdict for a former terminal owner by refusing a conflicted claim rather
than a non-native kind.

* refactor(native-chat): stop threading the owner kind through a reservation

A reservation only ever names a native owner now, so the request no longer
carries a kind and the reserved lease records `native` directly. The attach
params keep `runtimeKind`: agentSession.ensure and create accept it, and the
operation fingerprint stored in the ledger covers it.

* test(native-chat): pin the legacy-lease rewrite with a transaction that changes nothing else

Hiding a tab also committed the visibility index, so the no-op transaction
wrote the file even when its open-time revision was wrong. Committing the index
first leaves the pending rewrite as the only reason to write.

* fix(native-chat): name a chat write by its target, not the owner generation

A write carried the fence of the last frame the pane read, and the host refused it
unless that fence was still current. An idle release and the restart after it each
move the fence, and the release publishes nothing, so a send after a release was
refused "Expected runtime fence 1; the session is at 3", and a Stop queued behind a
cold start was refused as stale.

Every write already names what it acts on: a send its conversation, a cancel its
turn, a prompt answer its item revision, a rewind its epoch; an option is
last-writer-wins. So admission stops comparing the client's fence, and the rebase
that papered over one restart (admitAtResumedFence, resumedFromFence) goes with it.
The writer-lease check stays, and so does the attach's compare-and-swap.

Frames now stamp the fence read when each frame is sent instead of a copy each
subscriber kept, which went stale on the same release.

* fix(native-chat): every journal append reaches the chats that are open

A journal write and its delivery to open readers were two calls, and some
writers made only the first. A failed start whose lease could not be handed
back, a provider revision with no frame behind it, and eviction's settlement
were all journaled without reaching an open chat.

A journal handle now reports every durable change, and the host's session map
binds that report to the session's readers when the handle is set. Writers no
longer publish what they append; the per-writer publish calls are deleted.

* test(native-chat): an epoch replacement reaches the open chat

* test(native-chat): each row reaches an open chat once, and a live handle enters only through the map

* test(native-chat): give the legacy-lease store test a tab id so the backfill cannot supply its rewrite

The seeded record had no surface tab id, so the next open backfilled one and
that rewrite alone made the no-op transaction write. The test passed with the
legacy-lease rewrite signal removed.

* test(worktree-activation): restore the OMP surfaced-agent resume test

The handoff removal deleted it alongside the terminal-owner tests, but it
covers the surfaced-PTY block that still guards resume, including an agent
whose ownership is unknown.

* perf(native-chat): a publish behind a delivered commit reads nothing

Each commit now delivers itself, so the publish a provider frame still sends
afterwards found every reader caught up but still read rows and rebuilt the
timeline for each one. A caught-up reader now skips the read.

* test(native-chat): state why the teardown test's fake journal is safe to cast

* docs(native-chat): say mutation admission checks only the writer lease

* docs(native-chat): drop the send rebase from comments that still described it

* fix(native-chat): a message is accepted, then delivered

A send to a chat with no running agent restarted the agent inside the send
call, before the message was recorded, so the client waited for the whole
start and a failed restart refused the message. Claude held prompts sent
during startup, and those could settle as "unconfirmed".

A send is now accepted inside the session's serialized queue: one ledger row
and one submission row marked handoverRecorded, published, answered pending.
A per-session delivery loop exists while a message is queued. It starts the
agent through the same serialized attach a hold uses, waits outside the queue
for a Claude child to prove its start, and hands the oldest queued message
over as its own serialized step, writing dispatch{pending} before the adapter
call. A start it needed and did not get writes one error-tone row and rejects
every queued message with the same words; a start Stop cancelled writes none.

Settlement follows from the rows. A queued message is provably unwritten, so a
close, an eviction or an exit rejects it. A handed-over message stays in doubt.
A queued row at or below the sequence a handle found when it opened was left
by an earlier process and is rejected at open, with no latch. Stop withdraws
queued messages with no writer lease and no fence. An attach failure keeps the
conversation open, and the attach adopts its journal. Owed work counts the
loop and queued rows.

A compaction or rewind found prepared when a conversation opens was started
under a child this process no longer has, so the open settles it rather than
leaving it to refuse every send until a view attaches. The open cursor is
scoped to its epoch, because sequences restart when an epoch is replaced.

Deleted: restart-before-admission, recordFailedRestart, the fence rebase,
Claude's startup gate, the attach's forget on failure and its own crash
boundary. Clients without agent-session.accepted-send.v1 get their reply held
until the handover; the desktop and paired desktop lists advertise it.

* fix(native-chat): settle queued messages only for the child that ended

A child that proved its start and then exited before its message was handed
over left the message queued: the exit settlement returned early when nothing
else was in flight. Delivery then started another child for it, and a child
that died the same way started another, without end and without a row.

A retried settlement for an earlier generation, run by the attach that
delivery started, did the opposite: with that generation's turn unfinished it
rejected the message queued for the child being attached.

The settlement now takes the rejection for queued messages from its caller.
The unexpected exit and the eviction pass one, and it applies even with no
other work in flight; the retry for an earlier generation passes none.

* fix(native-chat): an adoption that fails to import keeps the conversation open

The attach now writes into the conversation's own open journal, but a failed
transcript import still closed it as if it were the attach's provisional one.
The conversation stayed indexed with a closed journal, so every later send
answered "could not be recorded" and every attach failed again until the app
restarted. The import now closes only a journal the attach opened for itself.

* perf(native-chat): the recovering open reads the journal once

Every conversation open now goes through the recovering open, including the
read restore of every chat at startup, which used to replay its journal once.
The recovering open replayed it twice: once to probe it and again inside the
open. The probe is now handed to the open as its load.

* fix(native-chat): an attach that fails after indexing its child leaves no child behind

A failed attach now keeps the conversation open, but a failure after
`onAttached` indexed the child (the rewind or compaction recovery, or the
attach's own success record) left that entry claiming a child the failure
path had already released. The next send found the phantom, skipped the start,
and wrote at a fence the journal had moved past, so the message stayed queued
for good. The entry now drops the released child and its event sink, and
follows the record's fence, as a failure before indexing already did.

* fix(native-chat): a withdrawn message shows no error, and a rejection outlasts the send's answer

The error strip for a message the host accepted and then did not deliver matched the entry before
the outbox reconciled, so a Stop's withdrawal, which the reconcile drops, showed "Orca could not
send your message" with nothing to retry. It now reads the reconciled entry.

A rejection the journal records before the send's own pending answer lands is final as well:
that answer no longer puts the entry back to dispatching with no Retry.

* fix(orchestration): a structured worker whose agent outlasts the preamble wait is left unknown, not torn down

The preamble waits for its submission to be delivered while the worker's agent starts. When that
wait ran out it threw operation_unknown, and the failed-start teardown then closed the session,
which rejected the very preamble the host was about to deliver. It now reports a turn start
nobody observed yet: the worker is start-unknown with its session kept, the host delivers the
preamble when the agent starts, and the worker's report settles the dispatch as for any
unobserved start. The receipt no longer suggests reading a screen a structured worker lacks.

* fix(native-chat): a message rejected while its chat was closed reads as not sent

A remount reads an entry it left dispatching as unconfirmed. When the journal had rejected it
meanwhile, as a failed start or a quit now does, the reconcile left it unconfirmed: it blocked
every later message behind a Retry and no reason, and the delivery probe, seeing the journal
already answered, never ran. The reconcile now settles it as rejected like a dispatching one.

* test(orchestration): name why the readiness settlement fakes are cast

* fix(native-chat): keep each pane's own fence on frames so a failed restart is not resent

* docs(native-chat): drop the fence from the admission the send effects run behind

* docs(native-chat): give the fence move on release the reason that still holds

* docs(native-chat): stop citing a write fence check in launch and mailbox comments

Three places still gave the removed fence check as a reason: the launch replay said admission puts the ledger ahead of the fence, the launch surface said a send must name the lease it was admitted against, and the direct-mailbox path said the lease fence decides whether delivery is safe. Admission now checks only the writer lease.

* refactor(native-chat): the provider child is its own record

A conversation now outlives any number of provider children, so the child is one record on the
conversation's entry instead of five loose fields beside its journal. It is written in one place:
indexed only once an attach has fully succeeded, and ended through one function that an exit, a
failed re-attach, a Stop and an eviction all share, matched on the child's generation and fence.

- A failed attach writes no child, so there is nothing to unwind: the field unwind and the fence
  patch after it are gone.
- Conversation writes read the record's fence, the way mutation admission already does; a child's
  own writes use its fence. The four stored-fence patches, and the settlement retry's overwrite of
  the conversation's fence, are gone.
- The owed wind-down is its own tombstone, carrying the child it is owed for, and is no longer
  dropped when an attach replaced the whole entry.
- Stop on a child still proving its start stops only the child: its lease goes back and the chat
  is told it is idle, but the journal, the holders and the readers stay. Close is that stop plus
  the conversation's close.
- The settlement retry uses the conversation's own journal, opened through the host's one open.

* fix(native-chat): the delivery loop alone settles a message its start or child failed

A queued message was settled by whichever path happened to end the child first: the loop, the
unexpected exit, eviction's work settlement, the open's leftover rule, and the startup branch that
rejected every pending row. That gave two failure rows with different tones for one start, a loop
that could hand over to a different child than the one it waited on, and a Claude start that died
while starting reading unlike every other failed start.

- The loop remembers the child it waited on. At handover, if that child is gone or replaced, it
  reads how it ended: a Stop continues; anything else writes one failure row and rejects every
  queued message with the same words, then stops. A child still starting whose start the adapter
  says did not land fails the same way. The exit, eviction and the settlement retry only settle
  the handed-over and legacy rows of the child that ended.
- One failure row, always an error, keyed by the start. A start a view began that dies with
  nothing queued writes the same row through the same builder, so a second report revises it.
- The open no longer rejects leftovers; the loop's first step does, and the open wakes it.
- `awaitStarted` answers why a start did not land, so the row says it even when the loop sees the
  failure before the exit is processed.
- Quit closes every conversation the way closing a chat does: what is still queued is rejected as
  closed, with or without a child, and a start the loop already has in flight is waited for so the
  child it produces is stopped rather than left behind.

* refactor(native-chat): a stopped child ends on the one reading of its stop

The eviction step reads a stop's result through `stopAgentSessionProviderRoot` and hands that
verdict to the child's ending, so the host never forms a second view of whether the root is gone.
Every ending carries it: a stop's comes from that reading, an exit's root is gone by definition,
and a failed re-attach passes what its release saw. The end-of-child record can therefore also
carry a stop whose root was not seen to go, which nothing ends on yet.

* feat(native-chat): the host says it accepts a send before any agent has it

The host now lists agent-session.accepted-send.v1 among its own runtime capabilities, the same
string capable clients already send. A client can then tell a host that answers a send at
acceptance, and admits a Stop with no writer before a turn starts, from an older one that still
restarts the agent inside the send. Additive: an older client ignores a capability it does not
know.

* refactor(native-chat): an attach never opens a journal of its own

The attach adopts the conversation's open journal, which outlives it, so it no longer opens one
for a direct caller either. That leaves nothing for a failed adopted import to close, and the flag
that told the two cases apart is gone. Tests that attach without a host open the conversation the
way a host does.

* fix(native-chat): a moved fence resends nothing on a host that accepts first

The outbox treated any fence change as a new owner: it dropped the answer of a send in flight,
queued that send to go out again under the same id, and unblocked a refused head. On an older
host that is how a send the restart refused, unrecorded, gets another try. On a host that records
every send before it starts an agent, a fence moves because that start ran, so the same rule
resent into every failed start. With a fence stamped on every frame, that became a loop.

The outbox now reacts to a fence change only when the host has not advertised that it accepts a
send before any agent has it. On such a host, only a Retry or a new send goes out, and a failed
start reaches the client as a rejected message it keeps with its Retry. Against an older host, or
before one has answered, the outbox behaves as it did. Desktop and paired web share this hook.

* refactor(native-chat): a child's end says whether the user or the host stopped it

The end-of-child record's cause now tells a user's Stop from the host stopping the child for a
cause of its own: `user-stop` and `host-stop` replace `stop`. The delivery loop goes on after a
user's Stop, as before, and fails the start it was waiting on after a host stop, with the one
error row and every queued message rejected, in the stop's reason when it gave one. The reason
stays description only. Stop passes `user-stop`; nothing passes `host-stop` yet.

* fix(native-chat): a chat whose only work is a queued message is not offered for resume

A message accepted while the agent was starting counts as working in the chat, and quit rejects it
as never sent. The teardown snapshot read the same working rule, so a relaunch offered to resume a
chat whose agent never had the message. The snapshot now reads only what was handed over.

* test(native-chat): type the queued-message fixtures in the resume-offer tests

* fix(native-chat): a start that dies while a message waits on it is that message's failed start

Opening a chat's tab starts an agent for the view, and a send accepted meanwhile waits on it. When
that start died, its exit wrote the start's error row and left the message queued, so the delivery
loop started a second agent into the same failure and wrote a second row. A child's end now records
where the conversation's journal stood, and the loop settles a message accepted before a failed
start ended with that start: one row, under its key, and no second start. A message sent after the
failure still gets a fresh start.

* fix(native-chat): a request that failed reads as failed

A structured chat whose only message the agent's start refused read as a
green finish, and a cancelled structured turn did too: the host published a
verdict only for turn records, and structured rows carried no `interrupted`.

The host projection now reads the session's latest request: its turn's
outcome, or `failure` for a send the agent or its start refused. A send
that was withdrawn, or left undelivered by a restart or a close, fails
nobody and makes nothing listable. The ingest publishes `interrupted` as the
hook lanes do, and every reader decodes the verdict through one accessor, so
a failure reads Failed on the dot, the rollups, history and `worktree ps`,
behaves like a cancellation in every clean-finish policy, and notifies as
"failed".

* docs(native-chat): say what an attach's open conversation and unconfirmed ids are now

* test(native-chat): a verdict change republishes the mobile status projection

* refactor(native-chat): the store's retention trigger keeps its flag compare

A verdict change always moves the completion clock the same check already
reads, so a second verdict compare there caught nothing new.

* test(native-chat): a user message the provider journaled keeps its session listed

* test(native-chat): pin what a failed start settles, and what a resume offer names

A view's child that dies while a sent message waits settles that message only when it died starting
and no child has taken its place: a proven child's crash, or a second start since, gets the message
delivered. The resume offer names the handed-over message, never a newer one still queued.

* test(native-chat): the failed-start pins fail on what the message became, not on a timeout

* fix(native-chat): a late provider-session update keeps a failed recovery record failed

A provider-session heartbeat that rewrites a completed recovery record kept
its interrupted flag but dropped the outcome it was copied with, so a live
failed checkpoint read as a clean finish until the next status write.

* test(orchestration): the preamble's host stub is typed, not cast

The preamble send now takes only what it reads of the host, the send, the settlement wait and the
record's fence, so its test builds that host with real types instead of `as never`.

* test(native-chat): the terminal-bell check asserts the renamed verdict field

The bell notification test still checked for agentInterrupted, which no
longer exists, so it could not catch a verdict leaking into a bell dispatch.

* fix(native-chat): a failed turn ranks like a completion for attention

Attention readers (completion time, Smart Sort, sticky retention, Cmd+J
Recent) now demote only a turn the user stopped. A failure is news the
user has not seen, so it keeps its completion time, ranks in the Done
class, stays retained after its pane goes away, and a retained failure
reads failed in the worktree rollup instead of done. Clean-finish
policy (hibernation, pane ownership, the value moment) still treats a
failure like a stop.

The retention trigger compares verdicts again: success -> failure no
longer moves the completion clock.

* fix(native-chat): a failed main agent reads failed while its subagents still work

The verdict is now read from the main agent's own state, not the folded
row: a main agent that is done and failed has a verdict even while its
subagents keep the row working. Without mainAgent (history, worktree ps,
older hosts) the old combined-done rule stands.

Display marks the verdict through agentVerdictDisplayMark: a failure
outranks every combined state on the agent's dot, label, tab badge,
dashboard and activity rows; a stop marks only a done row, so a
successful or stopped main agent with live subagents still reads
working. Subagent rows keep their own state. The worktree card, terminal
tab and Cmd+J rollups share one pane fold and rank a pending question,
then failed, then working, monitoring, interrupted and done.

worktree ps publishes the main agent's outcome on a working row, and the
mobile mirror reads it. The store's change check, the paired-client
mirror's equality and its epoch now see a verdict change on a working
row, which otherwise moves no state or clock and left the worktree card
reading working. Clean-finish policy is unchanged: a working row is never
hibernated and has no completion time.

* docs(native-chat): the worktree ps outcome comment no longer claims old hosts send it

The field is new: an old host sends no outcome at all, so a reader falls
back to interrupted. The removed clause said old hosts send it on done
rows, which never shipped.

* docs(native-chat): the status-store listing rule names provider-journaled user messages

* fix(native-chat): a refused send notifies failed through the completion feed

The host's completion feed followed only the newest turn, so a send the
agent or its start refused, which creates no turn, read Failed on its row
but sent no notification. The feed now follows the session's latest
request, read from the projection the status feed already makes for the
commit: a turn keeps its id, a refused send is named by its journal item
key. It announces only while the session is idle, as the row reports a
verdict, so queued sends refused one commit at a time notify once, and a
withdrawn send falls back to a request already announced.

* fix(native-chat): every copy of a row carries the main agent's own status

History entries, sleep records and `worktree ps` rows carried a flattened
top-level `outcome`, copied under different gates and without the main agent's
clock. They now carry `mainAgent` (state, outcome, stateStartedAt), the type
the live row already persists and sends, and every copy site takes it with
`interrupted` through one function, `agentVerdictFields`.

- The accessor reads `mainAgent` then the legacy flag; the mobile mirror
  matches it line for line.
- Sleep records admit `mainAgent` with `normalizeMainAgentStatusField`, so a
  malformed value drops the field, never the record.
- Mobile dates a main agent that failed under live subagents by its own clock,
  as desktop does, and its row equality compares `mainAgent`.
- The activity feed reads a history entry's own `mainAgent` instead of
  rebuilding one; the sync key and history equality compare it.

* test(native-chat): pin the worktree ps verdict across host and phone versions

Pairs the real v1.4.212 host and phone row reader with this build: an old phone
reads a new host's rows by `interrupted`, a new phone reads an old host's rows
(no `mainAgent`) the same way, and a new phone reads a failure under live
subagents as Failed, dated by `mainAgent.stateStartedAt`. The release checkout
now carries the phone's self-contained row reader, and the lane runs when the
`worktree ps` row producers change.

* test(mobile): name the parity table's row for its role

* fix(native-chat): a request that settles while the user is asked something notifies once

The completion edge waited for an idle session, and a pending prompt (including a
subagent's approval) is not idle. Structured chat has no other attention producer,
so a main turn that finished while a subagent waited on the user sent nothing
until the prompt was answered.

The edge now waits only on owed work (a running turn or an unanswered send), which
the projection reports even beneath a pending prompt. A request that settles with
a prompt pending announces once; the renderer words it "needs input" from the
host status mirror's `attention`, and answering the prompt keeps the same request
identity, so it does not announce again. The wire shape is unchanged.

* fix(native-chat): the completion says when the user is being asked

A request that settles while a prompt waits on the user was worded "needs input"
from the renderer's status-feed mirror. Remote clients receive the status and
completion streams over separate sockets, so they can arrive in either order and
the wording could be wrong both ways.

The host already knows at emit time, so the completion now carries an optional
`awaitingUser: true` in that case and omits it otherwise. The renderer words the
notification from that field alone and no longer reads the status mirror. Old
clients ignore the field and word by outcome; old hosts never send it.

* fix(worktree-status): a departed agent's failure yields to live work on the worktree card

A retained failed agent has no expiry, so ranking it with a live failure pinned the card to Failed over other panes' live work. It now ranks below working, monitoring and permission, and above every finished outcome.

* docs(agent-status): a departed agent's failure ranks below live work on the worktree card

* fix(native-chat): a view never restarts a chat whose last start failed

A Claude chat whose CLI exits during startup left one red row per start, and
every time a view bound to it (the chat opening right after its create died,
or the user switching back to it) the hold started the CLI again, so the same
launch-failure row repeated. Only a send retries a failed start now, the same
rule provider-exit recovery already applied; the rule lives in one predicate
the hold, exit recovery and the delivery loop share.

* test(native-chat): start the child the loop waits on with an attach, not a second view

A view no longer starts a child whose last start failed, so the R2 case that
waits on a child started since the failure now gets that child from a client
attach, the one non-send starter left.

* fix(native-chat): settle a gone generation's turn wherever a conversation opens

A send that opens a chat this process had not read yet (after a crash, from a
phone or the CLI) went through the delivery open, which never settled what the
dead generation left running; only the read restore and a successful acquire
did. When the send's start then failed, the turn stayed running for every
reader. The settlement now runs in the one journal open, at the crash boundary,
for every opener except an acquisition, which settles from the evidence it read
before its reserve; the read restore's separate step is gone.

* test(native-chat): prove the next child's start settles the turn an earlier child left

The R1 case lost its only settlement assertion when the latch it checked was
deleted. It now seeds the running turn the earlier child left and asserts it
ends at the exit's receipt, with the exit's row, before the message is handed
to the new child.

* test(native-chat): count a failed start's rows by row, not by text

Comparing the set of texts passed when two different rows carried the same
words, which is the duplicate the test exists to catch.

* test(cross-version): load the phone row readers without mobile's toolchain

Vite transforms a file against its nearest tsconfig, and mobile/tsconfig.json
extends expo/tsconfig.base.json, which the root-only cross-version lane never
installs. The worktree ps verdict suite imported the current phone row reader
from mobile/ directly, so CI failed with TSConfckParseError before any test ran.

The harness now imports a copy of the working-tree reader placed under the
checkout cache, where the root tsconfig applies, as it already does for the
release checkout's copy. Both readers are still the real files.

* test(cross-version): keep the checkout path-guard message and justify the copy import's cast

* test(native-chat): give the failed-start and stale-turn waits a loaded runner's budget

* test(native-chat): pin the open's and the send's start and row counts, however the view binds

Opening a fresh chat whose starts fail makes one start and one row, with two
views bound before or after the create's child died; one send makes one more
of each.

* fix(native-chat): settle a gone generation's turn at every open but an acquisition's

The journal open skipped the settlement whenever the lease read reserved or
live, to leave an acquisition's own open to the acquisition. But a lease a
crashed process left in recovery also reads live, until the next acquire
resolves it. A send that opened such a chat, from a phone or the CLI after a
crash on a host that could not prove the old owner gone, skipped the
settlement; when its start then failed, the dead turn stayed running for every
reader. The acquisition now says it is the opener, and every other open
settles, whatever the lease still claims.

* test(native-chat): hold the create's start open until the views bind

The "view binds while the create is still starting" case gave the create a
300 ms head start and asserted the views bound before it died. On a loaded
runner the holds took longer, the create's exit landed first, and the case
failed its own precondition. The create's initialize now waits on a gate the
test releases once the views are bound.

* refactor(native-chat): drop the composer's second error formatter

After the merge with main, every chat write in the composer path reports its
failure as a typed outcome worded by the refusal-notice table, so the send's
catch sees only a local throw. The {code, message} formatter this branch added
for it has no payload left to format, and its claim to be the one way a chat
words a failure is no longer true. The composer send is main's again.

* test(native-chat): pin the reason on a message rejected while its chat was closed

The reopen test checked only that the message reads as not sent; it now also
checks the Retry row carries the host's reason.

* fix(native-chat): a send the provider never received after a restart has no verdict

Restart reconciliation rejects a crash-stranded send that is absent from a
trustworthy provider history with reason 'not_delivered'. Nobody failed that
send, but the verdict allowlist did not name it, so after a crash the chat
read Failed, was listed, and could notify "failed". Give the reason a shared
constant (persisted value unchanged), add it to the no-verdict set, and treat
it as an internal marker so the Retry row no longer shows the raw string.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-27 22:23:49 -07:00
OrcaWinandm4air 27b823f934 ci: compile the E2E CLI once for all consumers (#23384)
* ci: share compiled CLI output across E2E consumers

* ci: preserve CLI setup and old-ref fallback for shared artifacts

* docs: record shared E2E CLI benchmark evidence

* docs: include final CLI reuse timing range

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-27 02:02:29 -07:00
OrcaWinandm4air c15f082031 ci: build independent Electron targets together for E2E (#23378)
* ci: reuse parallel Electron targets for E2E builds and guard cache action setup

* test: recognize the top-level cache repository preload

* docs: record E2E build timings and exact output parity

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-27 01:17:27 -07:00
OrcaWinandm4air 47cebbf5d2 ci: use ARM unit runners, overlap web builds, and reuse verifier fixtures (#23376)
* test: reuse isolated mobile bundle fixtures for verifier checks

* ci: run PR unit shards on ARM and overlap independent web builds

* docs: record controlled CI overlap and runner measurements

* test: observe WebRTC packets with the host clock

* ci: isolate Windows installer CIM probe from native test load

* docs: record native probe scheduling validation

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-27 01:06:54 -07:00
OrcaWinandm4air 25c3ac400b ci: overlap shell setup, localization extraction, and mobile route preparation (#23368)
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-27 00:28:02 -07:00
OrcaWinandm4air d8e2a694f6 ci: overlap package preparation and security scans; share localization parsing (#23364)
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 23:53:19 -07:00
OrcaWinandm4air ccd1e87287 Overlap independent CI checks with native Actions background steps (#23351)
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 23:09:40 -07:00
OrcaWinandm4air b5dec85a4e ci: reuse mobile web route analysis and skip unrelated mobile tests (#23329)
* ci: share mobile route analysis and scope mobile test runs

* ci: cover mobile web runner process dependencies

* test: verify mobile web selectors through the new runner

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 21:37:40 -07:00
OrcaWinandm4air 1882458f44 ci: scope orcad smoke and parallelize Linux packages (#23314)
* ci: scope orcad smoke and parallelize Linux package formats

* ci: validate packaging when its copy dependency changes

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 20:44:15 -07:00
OrcaWinandm4air 3eb1adec20 ci: reuse fixture setup and scope localization extraction (#23291)
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 18:46:26 -07:00
OrcaWinandm4air 7a4f83336b ci: shard SSH Docker coverage and warm Windows native caches (#23286)
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 17:47:57 -07:00
dffb3498e2 fix(ci): preserve focused Playwright file selection (#23270)
* fix(ci): preserve focused Playwright file selection

* test: isolate historical hourly build version inputs

* test: align historical package input with shared CI repair

* test: inject hourly package version without module mocking

* test: share the hourly package input contract across CI repairs

* test(ci): require focused commands on every platform

Assert each golden platform step and its focused arguments before deduplicating Playwright discovery probes.

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
Co-authored-by: Neil <neil@stably.ai>
2026-09-26 15:10:35 -07:00
OrcaWinandm4air 4b6fe95943 fix(windows): preserve relocated terminals and native process scans (#22872)
* fix(windows): ship the process-table addon to the relocated daemon host

The Windows terminal daemon runs from a copy of the app under
%LOCALAPPDATA%\Orca\daemon-host\<version>. That copy took node-pty but not
@vscode/windows-process-tree, so the daemon's bare require of the addon found
nothing and every process-table read (foreground tracking, descendant sweeps)
fell back to a powershell.exe Get-CimInstance scan (#16905).

- Copy the addon's runtime files (package.json, lib/, the .node binary) into
  the host; the ~25MB of gyp intermediates beside them are filtered out.
- Treat a host missing those files as unmaterialized, so hosts built before
  this are rebuilt, and skip relocation if the install itself lacks them.
- Log the daemon's native/CIM capability at startup and warn once when the
  process table falls back to CIM.

Revives #19525 on current main.

* test(windows): locate update-survival loss before relaunch

* test(windows): preserve daemon tree before update-survival proof

* test(windows): distinguish Electron exit from launcher close timeout

* test(windows): verify process exit when inherited pipes delay close

* test(windows): trace installer process checks in isolated survival runs

* fix(windows): probe process-query capability before installer sweep

* fix(windows): match installer probe and process-check profile behavior

* fix(windows): use NSIS separators for the process-check include

* test(windows): dismiss session-search overlay in survival harness

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 14:31:07 -07:00
681f3ca1ba perf: stream the packaged browser installer checksum in CI (#23082)
* perf: stream the packaged browser installer checksum in CI

* fix(i18n): restore diff note draft catalog entries

---------

Co-authored-by: OrcaWin <alpha-eng@stably.ai>
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 13:44:12 -07:00
OrcaWinandm4air d20cb69c48 Optimize CI follow-up workflows (#23190)
Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 02:53:37 -07:00
OrcaWinandm4air 3a081abf71 fix(persistence): reclaim Windows profile locks after PID reuse (#23122)
* fix(persistence): identify reused Windows profile-owner processes

* fix(persistence): preserve absent-owner recovery without native registry

* ci: build Windows registry before native profile identity checks

* fix(cli): include native profile-owner dependencies in typecheck

* test: register native profile owner test in Windows PR lane

* test: use resilient Windows profile-owner cleanup

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 02:26:10 -07:00
Neil d3434a2db4 Improve pull request template for issue linking
Updated the pull request template to clarify issue linking for outside contributors and maintainers.
2026-09-26 02:18:22 -07:00
OrcaWinandm4air d17a17684b Reduce redundant CI runs, pnpm uploads, and fixture startups (#23145)
* Reduce redundant CI runs, store uploads, and fixture processes

* Avoid repeating draft-independent mobile checks on readiness

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 01:17:51 -07:00
OrcaWinandm4air 9b30c7f60a ci: verify mobile disposal and balance unit-test costs (#23114)
* ci: verify mobile disposal and reduce unit scheduling costs

* docs(ci): clarify timing assignment validation

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-26 00:12:06 -07:00
OrcaWin 6fc3cdcad6 Bundle Bun for headless Orca and profile persistence (#22635)
Bundle a pinned, verified Bun runtime for headless Orca so existing Node launch commands can hand off before opening a profile. Keep desktop execution on Electron.

Add the Bun SQLite adapter and terminal backend, bounded shutdown, process inspection and cross-platform artifact qualification. Keep future managed SSH deployment separate from current production launch paths.
2026-09-25 22:49:06 -07:00
OrcaWinandm4air 38bcdf76ac perf(ci): reduce queue pressure without paid runners (#23053)
* ci: measure complete unit file costs for shard balancing

* perf(ci): reduce repeated PR setup and capture complete shard timings

* perf(ci): seed reusable main-branch native and typecheck caches

* fix(ci): stop superseded unit workflows from resisting cancellation

* perf(ci): reuse bundle fixtures and share the baseline Git build

* ci: record hosted gains and refresh main hook parity

* ci: retain default workers after performance-budget regression

* docs: record hosted mobile timing flake

---------

Co-authored-by: m4air <m4air@m4airs-Air.localdomain>
2026-09-25 22:48:28 -07:00
Brennan Benson 067975bfd1 fix(native-chat): every lease latch has a way to die (#22820)
* fix(native-chat): every lease latch has a way to die

A failed exit settlement no longer leaves the lease in recovery: the release
writes no stage and keeps the exit in its death evidence, and whatever the dead
generation left running is settled from that evidence at the next acquire or
read restore. The settlement retry flag, its disposition and every branch that
read it are gone. A reservation that recorded no process is released at
startup and after a failed start, the never-written conflicted status and the
processless proof are deleted, recovery resolution always concludes, and Codex
records its child's identity at spawn, before the handshake.

* test(native-chat): a re-create needs a release proven by death evidence

* test(codex): the child's pid is reported before the handshake

* test(native-chat): type the crash and exit fixtures without casts

* fix(native-chat): wait out a terminal owner an older build recorded, in recovery rather than manual recovery

* test(native-chat): a chat mid-turn at quit reopens idle, and an older build reads an unproven release

* test(native-chat): explain the baseline store cast

* fix(native-chat): a terminal owner's refusal names the process instead of recursing

Opening a chat whose terminal owner an older build recorded threw a stack
overflow instead of the refusal that names the process to quit.

* fix(native-chat): wait out a terminal owner recovery cannot verify instead of releasing it

A terminal agent an older build recorded keeps its PTY across an Orca
restart, so a probe that cannot answer (a start-time read that fails on a
loaded host) is not evidence its transport is gone. Releasing it let a
native child resume the same conversation beside the live terminal agent.
Only proof of its exit now ends the claim.

* ci(cross-version): run the unproven-release downgrade test

The sharded unit job excludes tests/e2e/cross-version-wire, and the
cross-version job runs an explicit list that did not name the new test,
so it never ran in CI. A change to the record validator now also starts
the job.

* refactor(native-chat): map the retired manual-recovery stage to recovering at decode

Nothing in this build writes manual-recovery, and restart reconciliation
already rewrites it. Mapping it where the other retired handoff stages are
mapped removes it from the in-memory lease type and deletes the branches
that could only see it: the acquisition refusal, the renewer skip, the
unproven-release stage check, and the handoff-status 'manual recovery is
required' answer. Older builds accept recovering, so a record written back
still loads after a downgrade.

* docs(native-chat): say what happens to a live child an ownerless reservation leaves

The reaper runs once at store open, while the unreconciled lease still
claims the child's token, so it does not stop that child on this launch.
The comment claimed it did.

* test(native-chat): name the each-case label for its role

* fix(native-chat): continue a create retried after recovery released its reservation

The client retries a create it never heard back from under the same operation id.
Recovery had released that create's reservation, so the retry was refused
agent_session_ownership_unknown while its row was pending, and
agent_session_operation_expired once the row aged out, and the chat never started.
A retry whose lease nothing holds now continues as a fresh reservation at the next
fence, which also stops the old reservation's spawn from committing.

* test(native-chat): name the refusal a replayed create used to get

* fix(native-chat): one quit-the-terminal-agent message for a chat a terminal agent holds

A chat held by a terminal agent an older build recorded frees only when that agent
exits. Sending said to reopen the chat and opening it said two runtimes claimed
it; both now say the chat is open in a terminal agent, name its process, and say
to quit it. Error codes are unchanged.

* ci: run PR checks on the rebased head

* fix(native-chat): name a terminal owner's process only when its start time can tell it from a reused pid

* test(native-chat): relaunch from the dying host's durable state, so its still-pending attach cannot race the new host
2026-09-25 21:04:31 -07:00
Brennan Benson f9356d491a fix(release): run the tag's own skill freshness inventory tests in the release gate (#22775)
The skill-sharing release gate checks out the release tag, then restores
main's copy of skill-freshness-inventory.test.ts. That file holds behaviour
tests, so a behaviour test added on main runs against a tag that predates the
behaviour: #22606 added one and failed the v1.4.211 gate on macOS and Linux.
Keep restoring skill-provider-runtime-roots.test.ts from the workflow ref.
2026-09-24 22:47:59 -07:00
Jinwoo Hong a05649de91 fix(terminal): prove an idle Git Bash prompt through its bin launcher (#22752)
* fix(terminal): prove an idle Git Bash prompt through its bin launcher

Git for Windows' bin\bash.exe is a launcher that runs usr\bin\bash.exe as a
child and waits, so an idle Git Bash pane's job always holds two pids and the
Windows shell proof never confirmed it. Accept exactly the launcher plus its
direct bash.exe child, checked against the identity process table.

* test(terminal): wait for the Git Bash prompt before asserting the hand-off job

* fix(terminal): prove a Git Bash prompt as one unbranched MSYS bash chain

Orca launches Git Bash as bin\bash.exe -c "chcp.com ...; exec \"$BASH\" ... -i",
and each MSYS exec leaves its pre-exec process alive as a stub, so an idle
pane's job is launcher -> stub -> interactive bash. Accept any job that is one
parent-to-child chain rooted at the launcher whose every later member is
bash.exe, instead of a fixed two-process shape.

* ci: register the Git Bash shell-proof win32 test in the package-test list

* fix(terminal): read the spawned shell as a path, and keep one shell map

A spawned shell path with a space (/Users/John Doe/bin/zsh) was split as a
command line, so the POSIX proof compared against "john" and never
confirmed. Local panes now keep only the spawned shell path and derive the
name from it; the Git Bash chain walk drops guards the member check already
covers.
2026-09-25 01:00:36 -04:00
Neil 7ea01279cd feat(search): bundle ripgrep for local, WSL, and SSH search (#22396)
* feat(search): bundle ripgrep for local, WSL, and SSH search

Ship @vscode/ripgrep-universal's prebuilt rg for all six relay platforms in
every desktop artifact. Local and WSL searches spawn the bundled binary and
drop the git ls-files / git grep fallbacks; SSH deploys upload the remote's
binary once per ripgrep version and the relay prefers it over PATH rg.

* fix(search): address bundled ripgrep review findings

- Key the SSH ripgrep cache on the binary's content hash; a package bump is the only update step
- glibc verifier: read arch tokens below the slice root and accept static ELFs (arm64 release blocker)
- Ship ripgrep/PCRE2/musl license notices; bundle rg with orcad
- Packaged builds never spawn a bare rg; report fd pressure as transient
- SSH: install rg before sweep/GC, size-validate installs, back off instead of disabling on launch failure
- Scope Dependabot to @vscode/ripgrep-universal; revert unrelated lockfile churn

* chore(search): drop bundled-ripgrep reference doc; assert full packaging layout parity

* refactor(search): one entry point for spawning the bundled ripgrep

Local Quick Open, Quick Open path search, the Explorer name filter, and
runtime text search each repeated the same three steps: resolve the bundled
command, spread in the WSL distro, spread in the WSL shell expression. Fold
that into spawnBundledRipgrep so one place owns the rule that a bare 'rg'
must never reach spawn, and simplify the resolver's command/packaged checks.

Restore the AGENTS.md ripgrep rule dropped alongside its reference doc in
63f4dac, and note why the relay's availability probe may spawn a bare 'rg'.

No behaviour change; verified by the existing suites plus a new test that
pins the local, WSL-routed, and distro-routed-but-Windows-output cases.

* refactor(search): drop the local install-ripgrep path; enforce the rg rule

Bundling rg removed the local git/readdir fallback, so nothing can produce
the "install ripgrep on the host running the Quick Open scan" guidance any
more -- only a remote host an upload never reached still reaches the capped
listing. Drop the host parameter, the renderer's local branch and its
translation key, and the relay wrapper that existed only to pass 'remote'.

Add a ratchet test for bare 'rg' spawns, since the AGENTS.md rule alone had
nothing enforcing it. Its one allowlist entry is the relay's PATH probe,
which asks about PATH by definition. Verified the guard catches a planted
offender rather than passing vacuously.

Also stop chaining the remote cleanup sweep behind the ripgrep upload: on a
cold host that is a multi-MB transfer, and stale upload stages and
superseded version dirs were left on the remote for its whole duration. The
two touch different trees, so they now run concurrently.

* test(ssh): pin that the cleanup sweep does not wait on the ripgrep upload

* fix(search): derive rg spawn types instead of importing node:child_process

A type-only import still counts against the child_process ratchet, whose pin
and allowlist only ever shrink. Derive both types from wslAwareSpawn instead.

* fix(search): surface an unreachable WSL workspace instead of an empty result

Inside `bash -c`, a failed `cd` exits 1 -- the same code ripgrep uses for "no
matches" -- so a WSL workspace whose directory had gone away reported an empty
listing as a successful scan. main did not have this hole: checkRgAvailable ran
the same `cd` wrapper first and settled on `code === 0`, diverting to the git
fallback that this PR deletes. The WSL wrapper now takes an optional
cwdFailureExitCode; rg passes 97, and all four close handlers reject with a
clear error before the unavailable check can blame the install.

Also from review:
- Bound the fire-and-forget ripgrep upload with deploySignal. The controller
  aborts only on the deploy timeout, never on success, so this cancels a
  still-running upload when the deploy gives up.
- Run the stale-stage sweep before the installed check rather than inside its
  else branch. Once rg was installed every later deploy took the PRESENT path,
  so a stage orphaned by a dropped connection was never collected again.
- Note in orcad-remote-deploy.ts why wiring it up needs ripgrep work first:
  build-orcad.mjs copies only the build host's rg, and orcad reports
  isPackaged() === true, so a remote of another platform would find nothing.

ssh-relay-deploy.test.ts sat at the max-lines cap, so any edit to it failed the
gate. Split the four Windows named-pipe deploys into their own file (926 -> 737
+ 333); both are now well clear of it.

* fix(search): name the unreachable root in every handler, not three of four

Round-two review caught that the missing-cwd branch in scanRipgrepPaths sat
AFTER isRipgrepUnavailableExit, which classifies any code above 2 as a broken
install -- so for exit 97 it was dead code and Quick Open still told the user to
reinstall Orca. Reordered; all four handlers now check it first.

Also from review:
- A vanished workspace makes spawn fail with ENOENT, which read as a damaged
  install on every local path. Confirm the cwd with isRipgrepSpawnCwdUsable --
  the guard the relay already applies -- before blaming the binary. The async
  continuation re-checks `resolved`, because finish() drops its argument once
  settled and the rejected promise would otherwise go unhandled.
- bundledRipgrepCommand returned a bare 'rg' for an arch outside the bundled
  set, bypassing the guard that exists so Windows cannot resolve a bare name
  against the repo cwd. A packaged app now always names an absolute path.

Drop ci-shards/unit-assignment.json, a 9,425-line CI artifact swept in from
reproducing a shard locally, and gitignore the directory that produced it.

The "rg genuinely cannot start" test pointed at a synthetic /repo, which the
new guard correctly reports as unreachable; it now resolves to a real root so
it still tests what its name says.

* fix(search): let the error handler own the spawn-failure verdict

A failed spawn emits 'error' and THEN 'close' with a negative code. The cwd
check added in the error handler did not settle, so the close handler settled
first -- synchronously, with the reinstall message -- and won the race every
time. The branch was not merely flaky, it was unreachable in all four handlers:
it is guarded by pid === undefined, which is exactly the case that always
produces a following close(code < 0). Verified against a real spawn: 3/3 runs
give error(ENOENT) -> close(-2). The error handler now detaches 'close' before
the probe, so it owns the outcome.

The probe also had no rejection handler, so a probe that rejected left the
search unsettled forever -- a hang, not just a wrong message. It now falls back
to the prior verdict rather than inventing one.

Tests: filesystem-search-rg-timeout and orca-runtime-files-search already cover
error-first and close-first, but against synthetic roots that the new guard
correctly calls unreachable; they now resolve to a real root, keeping each
test's stated intent. Added a Quick Open case for the vanished-workspace path
and confirmed it fails with the old ordering.

* test(search): cover exit code 97 in all four ripgrep close handlers

Round-four review found the missing-cwd branch had zero handler coverage: no
test anywhere emitted close(97), only -2/0/1/2/127. Ordering was correct, but
guarded by source-line order alone -- and that exact ordering was wrong in
three of four handlers two commits ago. Each suite now drives close(97) through
its real handler and expects the unreachable-root message.

Verified the tests earn their place: neutering the missing-cwd check fails
exactly four tests, one per handler.

Also drop a Reflect.get the anti-slop gate rejects, in favour of `in` narrowing.

* docs(search): stop claiming the close handler always wins the race

The previous commit asserted close "would beat this threadpool round-trip every
time", from an n=3 sample that measured event ordering -- which was never in
dispute -- rather than probe-vs-close. Two later measurements disagree with each
other: 50/50 close-first here, 30/50 probe-first in review. Either way it is a
race on a sub-millisecond margin, and the detach is what makes the verdict
deterministic.

Why this wording matters: "close wins every time" is an argument for deleting
the detach as a guard against an impossible race. No test would catch that --
the suites emit error and close in the same synchronous tick.

* chore(search): ship the jemalloc and libunwind notices the Linux rg needs

The statically linked Linux builds carry jemalloc (BSD-2-Clause) and LLVM
libunwind (Apache-2.0 WITH LLVM-exception) in addition to PCRE2 and musl, and
both require their notice on binary redistribution. Confirmed with `strings`:
their symbols are present in linux-x64 and linux-arm64 and absent from the
darwin and win32 builds. Texts taken from the upstream canonical sources.

extraResources already copies the whole licenses directory, so these ship
without a packaging change.

* fix(relay): stop spawning a bare rg, name unreachable roots, collect old builds

Three gaps the reviews surfaced on the remote side, all pre-existing on main.

Bare `rg` on Windows remotes. Both relay spawn sites pass the user's repo as
cwd, and CreateProcessW searches the cwd before PATH -- the same hijack the
desktop side already fixes. The relay now walks PATH itself and spawns an
absolute rg.exe, skipping relative PATH entries because those resolve against
the cwd. No rg on PATH yields null, which callers treat as "ripgrep
unavailable" rather than handing spawn a bare name. POSIX keeps the bare name:
execvp never consults the cwd, so there is nothing to resolve and nothing to
gain. With the last probe converted, the bare-spawn ratchet allowlist is empty.

Empty results for an unreachable root. settleLaunchFailure resolved an empty,
successful-looking scan when the root was gone but PATH rg existed, and the
git/readdir chain never engaged because it only triggers on
RipgrepUnavailableError. Both relay paths now reject naming the root, matching
local workspaces. Missing-rg keeps precedence over a missing root, because only
that verdict engages the fallback chain -- two tests pinned that deliberately
and it would have been wrong to flip it.

Unbounded ~/.orca-remote/ripgrep/. Nothing collected this tree; the relay's
version GC only matches `relay-*`, so every rg bump left another ~5 MB per host
forever. The probe command now also drops sibling builds older than two weeks,
sparing the current one and live upload stages, on POSIX and PowerShell alike.
Two weeks because a client pinned to an older build may still be using it; the
cost of collecting one early is that client re-uploading once.

* fix(relay): probe the rg that failed, and close the drive-relative PATH hole

Five review findings against the previous commit, all reproduced first.

The launch-failure classifier probed PATH rg, but the spawn that failed was the
bundled binary. On the normal remote setup -- no rg on PATH, which is why Orca
uploads one -- the probe failed and a moved workspace was reported as a missing
ripgrep, telling the user to install what Orca already ships. So the fix was
inert on exactly the hosts the uploader exists for. It now takes a candidate
list and asks the binary that actually failed first, then PATH.

path.win32.isAbsolute accepts `\tools` and `/tools`: rooted, but carrying no
drive, so they resolve against whatever drive the process is on. The probe
would have validated one against the relay's drive while the spawn, running
with the user's repo as cwd, resolved it against the repo's -- the same
cwd-dependence this lookup removes, narrowed from directory to drive. A real
drive letter or UNC root is now required.

probeRipgrepVersion had lost the timeout's kill in the rewrite, leaking a live
process and a ref'd handle per launch failure -- for a hang, which is the very
case the bundled-rg back-off exists for. It also spawned without windowsHide,
which would flash a console; fixing that made an allowlist entry stale, so the
entry is gone and the pin ratchets down 63 -> 62.

`windowsPathRipgrep ??= …` never memoised a miss, because null is nullish. The
caching was inverted against cost: a hit stops at the first directory, a miss
stats every one, and only the miss was repeated -- per spawn.

The bare-spawn ratchet claimed "nothing in production spawns a bare rg", which
is false on POSIX. It now also matches PATH_RIPGREP_COMMAND at a spawn site,
and the comment states plainly what a textual guard cannot see: the POSIX bare
name reaches spawn as a parameter, and is safe because execvp ignores the cwd.

The drive-rooted predicate is tested directly rather than through the
filesystem -- a temp dir on a POSIX CI host has no drive letter to exercise
win32 semantics with, so the filesystem test could never have caught this.

* test(mobile): repin the session closure past #22452's two shared modules

Merging main brought the closure to 4220 against a pin of 4218. The two extra
modules are `src/shared/agent-turn-outcome.ts` and `src/shared/main-agent-status.ts`
from #22452, which the status projection this route already reaches import.
That change was src/shared-only, so the mobile job never ran on it -- the same
way the structured tool line slipped past, as the ledger above already records.

Repinned here because this PR's file set is what next made the job run, not
because this PR reaches either module. Verified: of the 28 source files this
branch changes, none appear anywhere in the route's 4220-module closure.

* fix(search): preserve remote binaries and complete runtime packaging

* test(relay): pin the probe's env now that it inherits the relay's PATH

8d6759a threaded the relay env into probeRipgrepVersion -- correctly, since the
probe decides whether a launch failure was the binary or the root and so has to
resolve the same rg the failed spawn would have. It left the assertion that
pins the probe's spawn arguments behind, which is what CI caught.

Asserting buildRelayCommandEnv() rather than loosening the match to any object:
under process.env the probe could resolve a different rg, or none, which is the
regression the change exists to prevent.

* feat(ssh): collect remote ripgrep builds by reference, not by age

Nothing collected `~/.orca-remote/ripgrep/`: the version GC matches only
`relay-*`, so every change to the shipped bytes left another ~5 MB on every SSH
host, permanently. The age window this replaces was the wrong instrument --
a directory's mtime is when it was written, not when it was last used, so it
cannot tell a superseded build from the one a live relay was launched against.
Deleting the latter is not graceful degradation: without a PATH ripgrep remote
text search rejects outright, and listing drops to the capped walk this PR
exists to remove.

So the question is reference. Each relay directory now records the build it
runs against in `.ripgrep-ref`, written only once that binary is confirmed
present, and the GC collects a build only when no installation names it.

The discipline is ssh-relay-native-deps-cache-gc.ts': anything the pass cannot
account for blocks the whole pass. A relay directory with no readable marker is
an older Orca's, possibly running right now against a binary it never recorded,
so the pass declines rather than guessing. Those directories are removed by the
version GC in time, which is what makes their builds collectable -- hence
running after it, not beside it. Deletion is the same tombstone, recheck under
the rename, then remove, so a deploy that takes a reference mid-pass gets its
tree restored. Windows has no pass yet, matching the native-deps cache's gate.

One test note: the first version of the "unaccountable blocks the pass" test
passed against a deliberately broken guard, because the tombstone recheck
masked its absence. The test now puts a readable recheck behind an unreadable
first scan, which is the only shape that fails when that guard is removed.

Recording the reference lives inside ensureRemoteBundledRipgrep rather than at
the call site: it is the same concern, and it keeps the deploy's ripgrep
surface to one call for the tests that mock it to protect their exec queues.

* feat(ssh): collect Windows remotes too, and ship the Rust crate notices

Three items previously left documented-but-open.

Windows remote accumulation. The cache GC was POSIX-gated, so the leak did not
go away -- it moved to the platform with the larger binary (rg.exe is 5.43 MB on
win32-x64, against 4.77 MB for linux-arm64). The PowerShell dialect now does the
same reference scan: entries and references carry token prefixes, because
PowerShell writes every uncaptured value to stdout and an untokenised listing
would feed Remove-Item whatever a cmdlet happened to emit.

Verified on a real Windows host rather than a mock: the listing emits its
ENTRY/LIST_OK tokens, a relay directory carrying a marker yields REF <entry>,
and a relay directory without one yields REFS_ERR -- the safety path, on the
real interpreter.

Rust crate notices. The crate set was read out of the shipped binary's symbols
and the licence identifiers taken from crates.io rather than assumed. Where a
crate offers the Unlicense, Orca elects it: a public-domain dedication carries
no notice obligation, and that covers eight of them. The four that do not offer
it get their MIT text reproduced. encoding_rs carries a BSD-3-Clause notice for
its WHATWG-derived encoding data that is joined by AND, not OR, so electing MIT
does not discharge it.

Release-only validation, corrected rather than repeated. Linux AppImage/deb/rpm
already runs in CI's package job on every PR, and Windows signing was already
rehearsed on this branch. macOS notarization is the only item a release must
still exercise, and the exposure is narrow: notarization requires signatures on
Mach-O binaries, and of the six bundled builds only the two darwin ones are
Mach-O -- `file` reports ELF for linux and PE32+ for win32 -- so signIgnore
excludes only files the notary never asks about.

orcad-artifacts.test.ts caught the new notice file missing from the standalone
runtime's shipped list, which is exactly the gap that test exists to catch: a
notice committed to the repo but never actually shipped.

* fix(search): protect relay cache references and handle failed spawns

* fix(ripgrep): close review gaps and repair deployment fixtures

* test(mobile): refresh merged session module census

* fix(ssh): preserve ripgrep caches with empty legacy references

* test(mobile): assert bundle boundaries instead of global module count
2026-09-24 17:25:48 -07:00
Jinwoo Hong 128e97ffca fix(relay): drain a same-cap cell over 5 minutes, not 2 (#22584)
* fix(relay): drain a same-cap cell over 5 minutes, not 2

The 2026-09-23 c27 roll drained 2,145 hosts over the 2-minute window,
about 18 re-dials/s, while the director re-places roughly 8/s through
its single-slot sticky lane. The overflow queued behind slow
re-placements and timed out, so /v1/assign returned 503 fleet-wide for
about 5 minutes. 5 minutes is the cell's maximum pace window and keeps
the remaining 2,650-host cells near lane capacity. The transition wait
already outlasts a 5-minute window (17 min).

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* test(relay): pin the same-cap drain window contract at 5 minutes

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): keep the drain wait at lease plus the 5-minute window

The transition wait after a drain was set to the 15-minute migration
lease plus the pace window. Widening the window to 5 minutes without
moving the wait left 12 minutes for a migration that can hold for 15,
so a late-window migration would time the wave out into rollback.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-09-23 23:38:15 -04:00
Jinjing 8d6fec597b Optimize cloud-verify workflow to scan HEAD instead of all history (#22457)
* fix(cloud-verify): scan HEAD instead of all history

- Gitleaks now verifies only the checked-out revision
- Reduces scan scope and improves verification workflow performance

* ci(cloud-verify): clarify that Gitleaks scans HEAD-reachable history
2026-09-23 01:36:24 -07:00
Jinwoo Hong 483fa0aca2 fix(cloud): compare the Asia topology budget gate against the measured 500-connection default (#22386)
The topology workflow's Cloud SQL gate carried a hard-coded 400 for the
instance's tier default while the consumer contract records the value
measured on the live instance (SHOW max_connections = 500, 2026-09-16,
#21163). The gate compares the two and the first production plan run
(35815654836) failed silently on that mismatch before Terraform ran.

The verified default now lives beside the tier and version it is verified
for, as VERIFIED_DEFAULT_MAX_CONNECTIONS, so the contract and the workflow
are two independent records of the same measurement and the gate keeps
its cross-check. The test pins the new source and forbids a bare literal.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-23 00:28:44 -04:00
Jinwoo Hong bf3f95245c feat(relay): declare Asia cell c30 at the c27 shape (#22375)
* feat(relay): declare Asia cell c30 at the c27 shape

Adds production-gce-c30 in asia-east2-a at the reviewed Asia shape (6,000
request units, 3,000/60 connection limits, 16-connection pool, disabled) and
the rehome trust the other Asia cells carry.

Every Asia enumeration now knows C30. The topology, admission, and director
tools treat it as its own reviewed wave so its plan and registration never
touch the live launch cells. C30 promotion requires C27 general and fresh
staging evidence. The topology and director validators now pin the committed
production pool of 16 instead of the stale 10, which had made the topology
workflow reject the committed launch cells.

* fix(relay): plan C30 at live images and prove it with its own canary

The shared URL map pulls every cell into the C30 topology plan, so the workflow
now plans each non-target cell at the image its live template serves, and the
validator names any change to a cell outside the wave. C30 promotion runs the
same five-minute production canary and automatic rollback C27 used, with the
load report proving the canary control was placed on C30, instead of relying
on staging evidence. C30 leaves the shadow gate's fleet pool list until it
serves, rollback rejects mixed partial sets, and a budget test pins the
mixed-Asia-pool refusal.

* fix(relay): pin C30 to the production director's live image digest

C30 promotion requires the director and C30 to report one digest, so C30
takes the director's sha256:4158d8a2 (read 2026-09-22). C27-C29 keep their
committed lines; every Asia check compares only the cells named in a run.

* fix(relay): read the committed cell map from a plan, not console

terraform console evaluates every output against state, and the Relay
deployments output indexes each cell's MIG, so it fails with Invalid index
while C30 is declared but not created. Read the map from a no-refresh,
unlocked plan over the same targets instead, and refuse empty overlay input.

* fix(relay): keep console readers working and C30 migration-only until promotion

relay_gce_cell_deployments indexed each cell's MIG, backend, and template,
so once C30 is declared but not applied every production terraform console
reader printed a warning to stdout and broke its jq parse. Wrap those six
lookups in try(..., null).

Same-cap listed C30 as general, so a rollback dispatch on a migration-only
C30 would restore it with activate and skip its canary. List it with the
migration-only cells until the promotion follow-up moves it.
2026-09-22 23:41:09 -04:00
OrcaWinandm4air 632ae1320b fix(daemon): reap terminal descendants during shutdown (#22232)
Co-authored-by: m4air <m4air@m4airs-MacBook-Air.local>
2026-09-22 04:32:26 -07:00
Jinwoo Hong da1c322b00 feat(mobile): one build-time switch picks native or OTA, default native (OTA phase E1) (#22193)
* feat(mobile): one build-time constant decides native or OTA, default native

EXPO_PUBLIC_MOBILE_SHELL is read in exactly one place, mobileShellBuildKind in
preferences.ts. Expo's babel preset inlines a literal process.env member
expression at build time, so a release bundle carries the answer as a constant
and anything but the exact string 'ota' — unset, empty, a typo — is native.
Every default build is therefore the native app, unchanged.

mobileWebShellFlagCanBeOn now answers __DEV__ or an OTA build, so the ability to
mount the page comes from the build and never from storage: a native binary
installed over an OTA one, same bundle id and same data container, still refuses
a stored 'true' without reading the key. An unset key reads on only in an OTA
build; a development build keeps its opt-in, and a stored 'false' wins
everywhere so the Troubleshoot toggle can switch an OTA build back to native.

That toggle now mounts wherever the flag can be on, which is the only way back
to the native screens in an OTA build, and its label names the build kind rather
than saying "(dev)". The bundle probe row beside it stays development-only: it
fetches.

The flag census gains two rules — one module reads the switch, in the member
form Expo inlines and not the bracket form, and one named function answers the
build kind — and the build-kind fence now lists the Troubleshoot route that asks
it. Docblocks that said a store build can never mount the shell now say it
mounts only when built for OTA.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* ci(mobile): one workflow input picks the shell, and no input means native

Both release workflows gain a `shell` workflow_dispatch choice, options native
and ota, default native, and hand it to the step that bundles the JavaScript as
EXPO_PUBLIC_MOBILE_SHELL. That is the Gradle assembleRelease step on Android and
the fastlane build_and_upload step on iOS; nothing else in either file sets it.

A tag push and a schedule carry no inputs at all, so `inputs.shell || 'native'`
yields native for them — the first OTA release is a dispatch with one field
changed, and every other run is the app we ship today.

Each build step prints the value it is about to build with, read back from the
same variable rather than from a second copy of the expression, so a run's log
cannot claim a shell the build did not use.

The new contract test evaluates that expression rather than matching its text:
absent, empty and 'native' all resolve to native, 'ota' to ota, and any
expression shape it cannot evaluate is a failure rather than a pass.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* build: the desktop packages the real page, and the placeholder is retired

build:mobile-web now runs the app builder and app verifier, and both take their
output root from MOBILE_WEB_BUNDLE_DIR in the packaging guard rather than each
carrying a constant of their own — one definition of where the bundle lives, so
a drift cannot leave electron-builder's beforePack looking at an empty directory
while the builder reports a tree it wrote elsewhere. build:mobile-web:app is
gone; it was the same two commands.

src/mobile-web/ and its two scripts go with it. What the app builder shared with
them is split into three modules named for what they hold rather than for the
bundle that used to own them: mobile-web-bundle-manifest.mjs (content types, the
canonical asset serialization, buildId, hashed assets, the protocol window and
the manifest write), script-entry-detection.mjs (isDirectInvocation, whose two
failure modes are Windows paths and symlinked entries), and
mobile-web-source-line-endings.mjs (the CRLF guard, now with a required
directory rather than a default pointing at the deleted tree).

The two suites that only needed *a* valid tree on disk — the beforePack guard
and the packaged-bundle guard — build one from mobile-web-bundle-fixture-tree
instead of bundling the whole mobile graph. It goes through the same manifest
writer the page does, so a manifest shape change still reaches them.

Also retired: the placeholder's tsconfig project and its typecheck lane, its
knip entry, its electron-builder exclusion and .gitattributes pins, and the
app-bundle test that asserted the shims stayed out of a builder that no longer
exists. pr.yml's page job builds the same bundle the package job ships.

Inert for native phones: they never fetch it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* style(config): one import of node:fs/promises in the entry-detection suite

The changed-code quality gate's focused plugins read the two as a duplicate
import; the readFile line was left over from the split.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs: the comments that still describe the retired placeholder bundle

The web entry said it was built by `build:mobile-web:app` into out/mobile-web-app
and shipped by nothing. That script, that directory and that fact are all gone:
it is built by `build:mobile-web` into the packaged bundle dir, and a phone
mounts it only when the binary was built with EXPO_PUBLIC_MOBILE_SHELL=ota.

Two Windows cache keys explained themselves by naming src/mobile-web and "the
two bundle builders"; config/** now covers the builder, the verifier and the
manifest writer, and the spike's key no longer waits on a Phase C flip that has
happened. The keys themselves are unchanged.

Three scratch directories in the app-bundle suites and one in the verifier still
spelled the retired output root. Renamed to mobile-web, which is what the build
writes; they are temp subdirectory names and nothing reads them.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-22 06:18:33 -04:00
Jinwoo Hong bf8d63bc67 fix(relay): keep the backend service out of the same-cap wave; the capacity role cannot update it (#22140)
The capacity role the same-cap wave authenticates as, orcaRelayProductionCapacity,
has no compute.backendServices.update. Since #21860 added
`google_compute_backend_service.relay_gce_cell["${TARGET_CELL_ID}"]` to both of the
job's plan invocations, every wave has therefore created the new instance template,
modified the MIG, and then failed 403 on the backend, leaving the cell isolated with
its trust probe, admission restore, and shadow gate all skipped. Run 35684694704 on
production-gce-c7 is the first one that hit it in production.

Drop the backend target from both plans and restore the resume gate to exactly
`.changes == 2` (the template-and-MIG rollback-image drift) or a converged plan,
removing the backend-only resume apply #21865 added on top. A resume applies nothing
again, which is what a resume means.

The validator keeps its bound on a cell backend update, so it still reports one and
refuses anything wider, but a wave plan can no longer contain one. The drain timeout
from #21848 and the log_config from #21860 need a root apply by a principal that holds
the permission; granting the capacity role that permission is itself a root apply, so
it can follow as its own change rather than blocking every wave in the meantime.

Claude-Session: ced32ebb-7155-4413-adad-1eccd14c2010
2026-09-22 02:50:32 -04:00
Neil 35fe67b610 fix(perf): measure terminal latency with presented CI frames (#22096)
* fix(perf): present benchmark frames only on isolated CI display

* fix(perf): wait for the benchmark page before presenting its window

* docs(perf): record full scale pass with unchanged latency budgets

* test(perf): document and verify the isolated display exception
2026-09-21 16:03:40 -07:00
Jinwoo Hong c8a5580659 fix(relay): re-place hosts off a cell isolated for a roll (#21911)
* fix(relay): re-place hosts off a cell isolated for a roll

A roll isolates a cell by moving it out of the 'general' admission class; the
cell then refuses every attach with 4503. The director never noticed, because
the only liveness test it applies to a host's current cell reads
`relay_cell_runtime.ready` and the heartbeat, and an isolated cell keeps
heartbeating ready=1 for the whole drain. So every host on that cell was handed
its own dead cell, closed, and handed it back — 500-1,900 hosts looping for
13-16 minutes per cell roll, at ~6 dials each per minute, with no neighbour
absorbing anything.

The sticky lane now treats a live incumbent whose admission is 'migration-only'
— the state a roll's isolate step writes — the same way it treats a dead one:
it returns null, which means "fall through to placement". The placement lane
had the identical hole eleven lines further down, so it takes the same
predicate; without that second swap the sticky change is inert, because
placement would hand the pin straight back (a draining cell has more headroom
than anyone). An isolated incumbent skips the dead-cell fence branch: that
branch exists to prove an unreachable cell stopped serving a host, and this one
is reachable and enforces the epoch itself.

'existing-only' is deliberately untouched — those cells serve the hosts they
already hold, and only `assignmentStrandedOnUnservedCell` may release that pin.
A host with an open `relay_assignment_migrations` row keeps its pin too, so
this stays disjoint from the migration machinery.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(relay): gate re-placement on a roll-isolation marker, not on admission

Review of the first commit found the predicate wrong. `migration-only` is an
admission class, not a drain signal: an Asia `--mode rollback`, an evacuation or
forward-recovery target awaiting a separate promote dispatch, a failed same-cap
wave's re-isolate, an abandoned migration retired on its target and a rehome
settlement all park loaded cells there durably, with no migration lease and no
open migration row. All five were indistinguishable from a roll's isolate, so
the first commit would have converted `operate-relay-asia-admission --mode
rollback` from a reversible admission flip into a mass move of ~4,000 hosts —
and, because `leastLoadedCell` treated region as a preference, into us-central1.

The signal is now an explicit stamp. `relay_cell_admission` gains a nullable
`roll_isolated_at`, added through the shared schema runner's catalog pre-check
so a migrated database takes no relation lock on boot and an un-migrated one
gets a catalog-only rewrite. The same-cap isolate step is its only writer, via a
new optional `rollIsolatedCells` on the selector apply; the same UPDATE that
writes the state clears the stamp whenever a cell leaves 'migration-only', so a
restore cannot leave one behind and a failed wave's re-isolate keeps the one it
has. Every other admission writer omits the field, so its cells stay unmarked
and their hosts stay pinned. Old directors ignore the field; old callers never
send it.

Region is now a constraint rather than a preference on this path only: a
re-placement must find a general, live cell with connection headroom in the
host's own region, or the pin is kept and one
`orca_relay_sticky_replacement_deferred` event is logged. Cross-region spill is
no longer reachable here.

The fence bypass is narrowed to a live incumbent. It was always a no-op for the
intended case, and for a stamped cell that stops heartbeating while still
holding sockets it reopened split-brain; that cell now takes the dead-cell path
unchanged.

Also: the hot-path admission reader no longer throws on an unrecognised state —
it sits on every sticky dial and the rule it feeds is "move the host", so an
unreadable row has to mean "don't". And the sticky lane reads the admission row
once for both the stranded rule and the stamp instead of twice.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(relay): emit the re-placement events after the transaction commits

CodeRabbit on assignment-store.ts:1075. Both events were written where they are
decided, which is inside assignOnce's transaction. A reservation or lease write
failing after that point rolls the placement back, but a line already on stdout
cannot be rolled back with it — so the canary this PR asks an operator to read
would count re-placements that never happened, and a Postgres transaction retry
could leave a stale line behind as well.

The transaction now returns its events alongside the RelayAssignment and the
caller flushes them once it has resolved. Returning them rather than setting a
variable in the enclosing scope is what makes the retry case safe too: only the
attempt that committed can carry its events out. assign()'s signature is
unchanged; the extra shape lives entirely inside assignOnce.

orca_relay_sticky_replacement_deferred was moved the same way. It cost one more
push into the array that already existed, and it is decided inside the same
transaction, so leaving it behind would have been the odd case rather than the
cheap one.

The new test injects a failure on the first write after the decision, asserts no
event is emitted, and asserts the assignment is still on its original cell —
without that second assertion the absence would only prove the emit was early,
not that it would have been wrong. A control dial with nothing injected emits
exactly one event, so the case cannot pass on a broken harness. With the emit
put back inside the transaction, it fails.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(relay): expire the roll stamp, correct the wire note, assert the stamp landed

Delta review findings B, D and E. A (the deferral path's cost) is deliberately
not implemented; it is now written up under Follow-ups in the PR body as
required before any Asia roll, because it cannot fire in a US canary.

B, which also closes C: the stamp was written, carried and never compared to
anything. A roll isolates and restores one cell inside ~15 minutes, so a stamp
older than two hours is not a roll in progress. It is a failed wave whose
failsafe re-isolated a possibly healthy cell and is waiting on an operator — the
postmortem in this tree records gaps of hours — or an orphan left by a director
rollback whose restore wrote 'general' without the clause that clears the stamp,
which the selector's 'keep' branch would then preserve until some later park
reactivated it. Both want the same answer and it is the pre-existing one: keep
the pin. One comparison against a value already on the row.

The bound takes the caller's `now` rather than reading the clock again, so one
assign reasons about one instant; the stamp's age is now a thing that decides
whether a host moves, and two clock reads could disagree across it.

D: the comment beside the new request field claimed an updated caller reaching
an older director "is simply ignored". The schema is .strict(), so it is a 400.
That fails closed — the isolate aborts before MUTATION_STARTED is set and
nothing is written — but it is a deploy ordering constraint, and it was
undocumented. The comment now says so and the PR body's rollout notes carry it.

E: nothing read the `rollIsolated` the script already prints, so an older script
against a newer director would silently produce today's behaviour and the canary
would read as "the fix did nothing" with no way to tell that from a wrong
premise. Both isolate steps now assert it, beside the generation they already
parse.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-21 04:18:57 -04:00
Jinwoo Hong e476193bf5 chore(relay): bound the shadow health gate and apply a pending backend update on resume (#21865)
* fix(relay): bound the same-cap shadow gate and apply a resumed backend update

Two findings both adversarial reviews of tonight's merged set agree on.

The report-only shadow health gate (#21849) had `continue-on-error: true` but
no step timeout. That bounds the step's contribution to the job outcome, not
its clock. Its reads are serialised, and a failure that answers nothing slowly
— an expired credential, a project-wide Logging 429 storm — makes every read
cost its full 3 x 60 s retry budget, so the cost scales with the roll window:
roughly 8S + 2 reads for S ten-minute sub-windows. A 40-minute window is about
34 reads, or 108 minutes, against the job's `timeout-minutes: 75`. A cancelled
job cannot be absorbed by continue-on-error, fires the failure-gated cleanup
isolation on an already-restored cell, and stops the strict next-cell chain.

Give the step `timeout-minutes: 5` and the artifact upload `timeout-minutes: 2`.
A timed-out step is a failed step, which continue-on-error covers, so the job
stays green. Inside the script, stop reading after an overall four-minute
deadline and report the remaining checks unverified, so the normal outcome is a
written verdict rather than a killed process; the step timeout is then only for
a hung process. The census test pins both timeouts and that the deadline leaves
the step time to write its verdict.

The resume branch (#21860) accepted `changes == 0` with a non-empty
`backendUpdate` as complete and applied nothing, so a resumed cell silently
kept the 300-second drain and no request logging behind a green resume. That
shape means the template and MIG are converged and only this cell's reviewed
backend update is left, so apply the saved resume plan — the validator has
already bounded it to this cell's backend and neither attribute restarts an
instance — then continue as converged. Template-and-MIG drift still applies
nothing, which is what a resume means, and a stranded cell's explicit MIG
replace is unchanged.

Claude-Session: relay-same-cap-gate-timeout-and-resume

* fix(relay): raise the shadow gate bounds clear of a healthy gate's read time

A healthy gate is already minutes of serial reads on the 2-vcpu runner, so a
four-minute deadline would report unverified tails on ordinary days and stop
the shadow roll measuring the comparison it exists for. Raise both together:
the step to eight minutes and the script's own deadline to seven, keeping the
census pin that the deadline leaves the step room to write its verdict. The
job budget is unaffected: a ~14-minute cell plus eight is well inside 75.

Claude-Session: relay-same-cap-gate-timeout-and-resume
2026-09-20 21:35:09 -04:00
Jinwoo Hong 2524737ef0 chore(relay): apply the cell backend drain and request-logging settings inside each same-cap wave (#21860)
* chore(relay): target each cell's backend service from the same-cap job

The 60 s connection drain timeout merged in #21848 has no safe apply path.
A root plan scoped to the backend services alone still pulls every
`google_compute_instance_template.relay_gce_cell` in as a dependency, and
standing image drift turns all 29 into replacements, so applying it would roll
the fleet at once.

Add `google_compute_backend_service.relay_gce_cell["${TARGET_CELL_ID}"]` to
both plan invocations in the per-cell same-cap job, next to the template and
MIG it already targets, and teach the reviewed plan validator to allow exactly
one extra change: an in-place update of that one cell's backend whose only
changed attribute is `connection_draining_timeout_sec`, landing on the
constant `validate-relay-asia-topology-plan.mjs` exports. Any other attribute,
any other resource, or a backend for another cell still fails the validator.

The accepted update is reported as `connectionDrainUpdate` and kept out of
`changes`, so the apply step's stranded branch and the resume step's drift
branch keep reading the template-and-MIG count they were written against; the
resume branch additionally accepts a plan whose only pending change is that
drain update, which restarts nothing.

Claude-Session: relay-same-cap-targets-cell-backend

* fix(relay): also let the same-cap wave apply this cell's LB request logging

A read-only production plan for production-gce-c7 showed the live US cell
backends carry no `log_config` at all, while relay-gce-cells.tf has declared
`log_config { enable = true, sample_rate = var.relay_gce_cell_log_sample_rate }`
on every cell backend since the Terraform root landed in 3eec77c11a (#18413).
Nothing has applied it because every production apply since is a per-cell
targeted plan that names only the template and the MIG.

So the real canary plan's backend moves two paths, not one:
`["connection_draining_timeout_sec", "log_config.0"]`. The drain-only validator
rejected exactly that plan, which would have stranded the cell mid-wave after
the drain had already started.

Accept both, each optional, for this cell's backend only: the drain landing on
RELAY_CELL_CONNECTION_DRAIN_SECONDS, and a log_config of exactly one block with
`enable = true` and `sample_rate` equal to RELAY_CELL_LOG_SAMPLE_RATE, the
declared default of a variable no environment file overrides. Any third path,
a different sample rate, disabled logging, another cell, or a replacement still
fails. The accepted paths are reported as `backendUpdate`, which the resume
branch now reads instead of the drain-only flag.

Verified against the real production plan: the masked, three-target plan for
production-gce-c7 contains exactly that cell's template, MIG, and backend
service and nothing else, and this validator returns
`{"changes":2,"backendUpdate":["connection_draining_timeout_sec","log_config.0"]}`.

Claude-Session: relay-same-cap-targets-cell-backend
2026-09-20 20:35:52 -04:00
Jinwoo Hong 5b8ac36f41 chore(relay): add a report-only post-wave health gate to the same-cap cell job (#21849)
* feat(relay): report a post-wave health verdict on each same-cap cell, without gating on it

After a same-cap cell finishes rolling, an operator reads five things by hand
before dispatching the next cell: director 503s against the same clock hour a day
and two days earlier, whether the cell's new container announced its listener and
has stayed up, the cell's own pool pressure, the asia-east2 pool trio, and Cloud
SQL FATALs. This runs those same reads automatically and records PASS / WARN /
WOULD_BLOCK with its numbers, so its calls can be compared with the operator's
over a full roll before it is ever allowed to stop one.

It cannot fail a cell in this change. The script exits 0 on every verdict, and
the step is continue-on-error, so even a crash stays off the job's outcome and the
failure failsafe cannot fire on anything it observes. It also runs after the
restore, so no cell waits on it to go back into admission.

Cloud Logging returns only --limit entries and says nothing when it truncates, so
every count is split into sub-windows of ten minutes and a sub-window that comes
back at the limit is reported unverified rather than as a count. Windows are
always explicitly bounded: --freshness does not bind on these logs.

Claude-Session: https://claude.ai/session/ced32ebb-7155-4413-adad-1eccd14c2010

* fix(relay): bound the shadow gate's cell reads at the apply start and cap every read

Four fixes from review, all in the report-only shadow health gate.

The boot search opened at apply-completed-at, which is stamped after
`terraform apply` and `wait-until --stable`. The new container announces its
listener while the MIG is still converging, so that bound is already past the
announcement it looks for and a healthy roll read as would-block. The job now
stamps apply-started-at immediately before the apply, and the boot search opens
there; apply-completed-at is kept, recorded rather than judged, so an operator
comparing verdicts can see apply time next to boot time.

The crash query started at the newest listener timestamp, which erased any crash
before it. A crash-restart loop ends with an announcement that looks like a clean
boot, so that is exactly the case it hid: against production, the 2026-09-20 c28
crash at 20:18:10 was dropped because the listener landed at 20:18:27. It now
runs from the apply start, still scoped to the instance id the listener
identified, and that crash is counted.

A runtime-metrics read that came back at its 500-entry limit fed judgePool as
though it were a complete sample run. A truncated run has holes and the
consecutive-sample rule reads a hole as a recovery, so it now reports unverified.

gcloud reads had no timeout. continue-on-error bounds the job's outcome but not
its clock, so a stalled read could have spent the rollout's remaining minutes.
Each read now gets 60 s and a timed-out read is just a failed read.

Claude-Session: https://claude.ai/session/ced32ebb-7155-4413-adad-1eccd14c2010

* test(relay): require each shadow-gate stamp's presence before asserting its order

The ordering assertion used indexOf, which answers -1 for an absent stamp, and
-1 precedes every real offset. Deleting the apply-started-at line left the test
green, so the census could not see the fix it was written to pin.

Each stamp's presence is now asserted first, with a message naming the stamp and
the step, and presence is judged inside the step that owns the stamp rather than
anywhere in the file: a stamp written into a neighbouring step records the wrong
instant but would satisfy a whole-file match.

Control-run against a scratch copy of the job. Deleting drain-started-at,
apply-started-at, or apply-completed-at each reds with its own message, and
moving apply-started-at after terraform apply reds on the ordering assertion, so
presence and order both fail independently.

Claude-Session: https://claude.ai/session/ced32ebb-7155-4413-adad-1eccd14c2010
2026-09-20 19:10:35 -04:00
Jinwoo Hong 4f839cc8c9 chore(relay): cut the cell LB connection drain to 60 s and allow ten-cell same-cap batches (#21848)
* perf(relay): cut the cell LB drain to 60s and widen the same-cap batch to ten cells

Two independent sources of relay roll wall clock, neither of which protects a
host:

1. `connection_draining_timeout_sec` on the per-cell backend services was 300s.
   The same-cap job drains every host off the cell to a restart-safe condition
   before Terraform runs, so the LB drain only ever covers a host still
   mid-handshake. Measured 2026-09-16 over ten same-cap cell jobs, it sat as
   ~5m55s of dead time between `Apply complete` and the old VM powering off,
   inside an 8.5-minute `wait-until --stable` step. Now 60s, and pinned in the
   topology `check` block beside the other fixed-one invariants.

2. The same-cap wave capped a batch at four cells, so a 22-cell roll needed six
   batches, six single-use monitor gates, and a human handoff per batch. The
   wave workflow now declares cell_1..cell_10 with the identical serial shape
   and chaining, and the validator accepts two to ten.

The shared wave-index rule (`relay-monitor-evidence.mjs` and the relay-ops
preflight CLI) widens from 0-3 to 0-9 so the later cells can present the same
evidence; each job workflow keeps its own narrower range, so the capacity wave
stays at four. Cells remain strictly serial, one at a time behind the rollout
lease, each with its own live preflight.

Claude-Session: https://claude.ai/session/relay-roll-drain-timeout-and-batch-cap

* fix(relay): align the Asia topology plan validator with the 60s cell drain

`validate-relay-asia-topology-plan.mjs` rejected any Asia backend whose
`connection_draining_timeout_sec` was not 300, and
`cloud-deploy-relay-asia-topology.yml` targets
`google_compute_backend_service.relay_gce_cell["<cell>"]` per cell. With the
Terraform local at 60 that workflow would have failed its own plan review.

The validator's two restated topology values are now named exports, and a new
census test reads `relay-gce-cells.tf` and equates three statements of each:
the `relay_gce_topology` local, the topology `check` assert that pins it, and
the validator constant. Terraform cannot export a local to JS, so reading the
source is the only way to stop them drifting; the test was confirmed to fail
when the local alone is moved back to 300.

Repo-wide grep finds no other pin of the drain value.

Claude-Session: https://claude.ai/session/relay-roll-drain-timeout-and-batch-cap
2026-09-20 19:01:44 -04:00
Neil cb715898cd fix(release): pass the draft-verify tag on Windows pwsh (#21851)
The Windows matrix defaults to pwsh, so assert-github-release-is-draft.mjs
received an empty argv and failed with "tag is required" after the signed
installer was already uploaded. Force bash, interpolate the tag in YAML,
and fall back to env TAG.
2026-09-20 16:01:01 -07:00
Neil 1b9d218df5 fix(release): force draft publishes on tag checkouts (#21842)
Build jobs check out the release tag, so electron-builder still used
releaseType:release from older SHAs and published v1.4.206 as latest with
only Linux assets. Override publish.releaseType=draft on the CLI (workflow
YAML comes from main) and restore the draft helpers from the workflow ref.
2026-09-20 14:58:35 -07:00
Neil 72d61c459f fix(e2e): wait for terminal remount after golden worktree switch (#21837)
Mac release goldens failed after switching back to the original worktree:
sidebar aria-current landed while the store still pointed at the child tab,
so waitForActiveTerminalManager timed out. Wait for activeWorktreeId, force
the terminal tab visible, and restore this spec from the workflow ref so
older cut SHAs pick up the harness.
2026-09-20 14:28:03 -07:00
Neil 3cadcabe11 fix(release): keep GitHub releases draft until all assets exist (#21835)
electron-builder --publish always was creating a public GitHub release as
soon as the first platform uploaded, so /releases/latest could serve a
missing Windows exe. Keep the main-repo publisher on draft, pin draft
creation to the tag commit, re-draft immediately if anything flips public,
and refuse mac publish after the parent cut is cancelled.
2026-09-20 14:01:15 -07:00
Jinwoo Hong 68b11282a5 fix(relay): let the rehome evidence parser read a line the director grew (#21823)
The enable workflow reads the director's `[orca-relay] regional rehome
inventory` line out of Cloud Logging and pins the whole line with one regex.
Adding `hostNotArrivedLast24Hours` in #21813 made every healthy line stop
matching, so "Read fresh aggregate completion and abort evidence" threw
"no aggregate regional rehome inventory evidence" and the fail-closed step
disabled the durable switch at control generation 26.

The parser now requires the six original fields and tolerates further ones
in any order. Extra fields stay fenced by value shape rather than by pinning
the whole line: a field must be a bare name and a non-negative integer or
`none`, so `hostId=someone` is still not a counter and cannot ride along.
An absent count reads as null, not zero, because an older director not
reporting leaks is not the same as reporting none.

`hostNotArrivedLast24Hours` and `oldestActiveAgeMs` now reach the evidence
JSON and the operator step summary.

Two guards close the chain, each verified to fail on the regression it
exists for: a census in the relay package feeds the real formatter's output
to the real parser, and a script-side test pins the parser's output to the
fields the workflow summary renders.

Claude-Session: https://claude.ai/session/ced32ebb-7155-4413-adad-1eccd14c2010
2026-09-20 15:50:43 -04:00
Jinwoo Hong e6aa90ff36 test(mobile): certify the browser pane's golden families and render it in a page (OTA phase C, C6.5) (#21777)
* test(mobile): pin the browser pane's golden families

The half pin for C6: 4 families, 15 goldens, every verdict the one C2's
rule predicts. Measured per family with vitest `-t` over the full
787-golden corpus, with C1's 103 reproduced golden-for-golden as the
control: 6 byte-identical, 9 result-absent-settlement.

No composed `c6-page-closure.ts`: a composed table is pinned against a
route and the browser is a pane, so C7's route is what composes this
with C1's.

The derivation census does not wait for that route. `mobileWebAppRoute-
Closure` becomes one case of `mobileWebAppModuleClosure`, which takes
any entries, so the pane's own closure can be read from the module. Two
cases: the pane alone reaches exactly the pinned four, and the pane
beside `app/h/_layout` adds exactly those four and no other, with the
layout reproducing C1's 22 as the control for the difference.

Closure at this base: 48 local modules alone, 34 beyond the layout, 30
under `src/browser` and four through the web siblings. The design said
23, all under `src/browser`; it was measured before C6.2 and C6.3 added
those siblings, so the pin carries the re-measured number.

`browser.screencast` has no golden at all, so this certifies the input
path and says nothing about the frame path.

Red first: with `browser.wheel` dropped from the table, both census
cases fail naming the missing family; restored, the file's 10 cases pass
and the parity suite reports "15 goldens in 4 families, 6 byte-identical".

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): pin the frame budget against the shell's real frame

Ruling 2's pin. `binaryEventEnvelopeBytes()` sizes the mobile view's
device scale from a skeleton it builds itself, and until now its only
check was another skeleton of the same shape in the same file: two
copies of one assumption agreeing with each other.

This measures the real thing. A frame with CDP's nine metadata fields
and a real `Page.screencastFrame` timestamp, encoded by C6.1's
`encodeBridgeScreencastFrame` and serialized by the real
`BridgeHostSubscriptions`, posted through the host harness: 303 bytes
besides the image, against a bound of 516.

Held above is not enough on its own — 213 bytes of slack is room for the
shell to grow the envelope by a field the page never hears about — so
the bound is reconstructed exactly instead. Every byte of that slack is
a number this frame prints narrower than a double can; adding those back
gives 516 on the nose.

The budget cases run a generated noise image at the budgeted scale, not
a committed fixture: the worst case is the image JPEG compresses least,
and a photograph sits a tenth of the way to it. 901,161 px at 0.545
bytes per pixel is 491,132 bytes, which the shell posts at 654,857 of
the 655,360-byte cap. One envelope more and the shell drops it, which is
ruling 1 read from the budget's side.

Red first, two ways. Drop the metadata widening from the bound and three
cases fail, the sharpest being the real shell answering the frame the
page thought it could send with zero posts. Add a field to the shell's
own envelope and the reconstruction fails at 516 against 548, where the
existing suite stays green on all 14 — which is the drift this file
exists for.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* docs(mobile): record the measured frame bytes, correcting e7cd24ef10

The previous commit message says the shell posts 654,857 bytes for a
frame at the budgeted area. That number was not measured; I wrote it
from the budget arithmetic instead of reading it off the harness. The
measured value is 655,147, which is 213 under the cap rather than 503.

Nothing in the assertions changes — they compare against the cap and
the bound, never against a literal — but the figure now lives in the
file where it was measured rather than only in a message that has it
wrong.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): render the browser pane in a page and paint a real frame

The only place C6's whole frame path runs. Every other check reads one
half: the shell suites drive the host with no page, the page suites
drive the hooks with no shell, and the parity pin certifies the input
path from a recording.

Ruling 4: no route is added. `bundleMobileWebApp` already takes an
`appDir`, so this builds a one-route tree of its own and nothing under
`mobile/app` moves. The shell double grows a screencast lane to serve
it: it accepts a subscribe, posts `event.binary`, and prices each frame
the way `BridgeHostSubscriptions` does, so an over-cap frame is dropped
where the page can watch the stream survive it.

Six cases: the pane subscribes with `wantsBinary` and paints the frame
it is handed; a second frame flips the double buffer; an over-cap frame
is dropped and the next one paints on the same subscription; the grant
withheld produces the update-the-app copy and no subscribe at all; a tap
issues one `browser.mouseClick` at the centre of the source viewport;
and no request leaves the bundle's own origin.

Measured. The frame the pane asks this viewport for is 390x698, which as
noise is 201,924 base64 characters. The phone's mobile-mode frame is
780x1424 and encodes to 811,168, which is 124% of the cap and the reason
the area budget exists; the over-cap case uses 2400x2160 at 3,761,580,
574% of it. The tap maps to (194, 356) against a 390x712 source, one
device pixel off centre because the rendered width is 382.33 CSS pixels
for 390 source pixels.

One finding, recorded rather than fixed because it is not the pane's.
The page files a CSP `script-src` violation on every load, on any route:
Zod 4 feature-detects its compiled path with `new Function('')`, the
shell's `script-src 'self'` blocks it, Zod catches the throw and takes
the interpreted path. The page is correct and the report is filed
anyway. One case names it so a second `eval` is visible, and every other
case asserts no violation beyond it.

Red first, twice, both by reverting behaviour C6.2 landed. Stub out the
decode probe in `whenBrowserFrameDisplayable` and the flip case fails on
two identical frame digests. Point `updateBrowserImageSource` at the
host element instead of the surface child and the paint and flip cases
both fail. The first case's comment is corrected by the first of those:
it claimed a visible layer proved the decode-then-flip, and the frame
still paints with the probe gone, so the flip case is what proves it.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): sweep the worst-case JPEG cost instead of taking one point

The reviewer is right, and it is worse than the report says. At 0.545
bytes per pixel, 90 of 143 viewports posted a frame over the cap, and
59 of the 111 the budget claims to fit were dropped outright by the
real shell — 390x712 at scale 1.8 among them.

The old number came from one 2400x2160 frame. A single large frame is
the cheapest per pixel in the whole range, so a worst case measured
there is not a worst case anywhere else.

Swept 143 viewports, widths 320 to 1400 and heights 480 to 1600, each
encoded by Chromium at the scale the real budget picks for it. Across
the 111 the budget fits, the cost ranges 0.54470 to 0.55351 bytes per
pixel. The constant is now 0.56: that maximum plus 0.00649, about 1.2%,
for the encoder version it was not swept on. The docstring carries the
sweep, the range, the margin and the date.

0.56 is a fixed point, not a guess. Raising the constant shrinks the
budget, which lowers the scale, which moves the cost; 0.555, 0.56 and
0.565 all leave the same 31 viewports over the cap, and every one of
those sits at the scale floor of 1, where the module already declines
to go blurrier and C6 ruling 1's drop rule is the protection. The new
test asserts both halves: nothing the budget fits goes over, and the
largest viewport it cannot fit is dropped by the real shell.

The sweep lives in `config/scripts` because it needs Chromium: the
frames are CDP screencast frames, so Chromium's encoder is the oracle
and a Node JPEG library would calibrate against the wrong bytes. It
drives the real budget, the real scale function and the real
`BridgeHostSubscriptions`, and runs in about 4 seconds.

Ruling 2's block in `browser-screencast-budget-at-the-shell.test.ts`
now says plainly what it measures. It feeds `noise(area * theConstant)`,
a byte count the constant itself produced, so it can falsify the
expansion and the drop rule but never the constant. It read as if it
validated the worst case, and it did not.

Red first: put 0.545 back and the sweep fails with 59 viewports, each
naming its scale and reporting `null` — the real shell dropping the
frame rather than posting it over the cap.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): turn Zod's JIT probe off for the page, before any module

The page filed a CSP `script-src` violation on every load: Zod decides
whether it may compile by constructing `new Function('')` and reading
the throw as "no JIT here", the shell's `script-src 'self'` is exactly
that throw, and the browser reports it before Zod catches it. Zod's own
source gates the probe on `jitless` for this case.

`z.config({ jitless: true })` at the entry does not work, and the
reviewer's suggestion of putting it there was measured losing the race.
`$ZodObject` reads `allowsEval` when a schema is constructed, not when
one is parsed, so the first module-scope `z.object(...)` in the bundle
fires the probe — and esbuild evaluates the chunk holding zod and its
callers before the chunk holding any module of ours that imports zod. A
Function-constructor trap in the page put the call under `new ZodObject`
ahead of the entry's first statement.

`globalConfig` is `globalThis.__zod_globalConfig`, which zod adopts with
`??=` rather than replacing, so the banner can set the flag before any
module runs. That is where it now lives, beside the `process` shim and
under the same `MOBILE_WEB_APP_SHIMS` contract, which asserts it is
applied. Nothing is lost: the compiled path was never reachable in a
page under this policy.

The render check's `newCsp()` filter is gone. It dropped violations by
`blockedURI === 'eval'`, which would have hidden a real one, and every
case now asserts zero. The first case walks load and first paint, which
is where the second of the two reports fired. The dead `violations`
array is deleted.

Red first: blank the banner constant and four of the six cases fail,
each naming a `blockedUri: 'eval'` the filter used to swallow.

Finding, not fixed here and reported instead: the page bundles two
copies of zod, mobile's 4.4.3 and the repo root's 4.5.4, because
`src/shared/zod-salvage.ts` resolves upward. That is 808 KB of duplicate
source. Aliasing `zod` to one copy in the builder fixes it and was
measured working, but it changes which zod shared code runs in the
shipped page, which is a call to make on its own rather than inside a
CSP fix.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): give the shell double the whole canCarry rule and the acks

The double reproduced one arm of `BridgeHostSubscriptions.canCarry`, the
message cap, and silently carried anything the other two would have
refused: a window already holding its maximum frames, and a window whose
bytes the frame would push past the limit. It also ignored the page's
`ack` frames, so its window never reopened — which was invisible only
because no case streamed far enough to close it.

Both arms are in now, and the `ack` arm consumes the page's acks exactly
as the host does. The three caps are read out of `bridge-caps.ts` and
`bridge-host-subscriptions.ts` rather than retyped, the same way the
harness already reads the protocol version and the CSP, so a double
carrying a stale number is not possible. The render check's own
`640 * 1024` is gone with them.

One case for it: thirty frames of about 200 KB, roughly 6 MB through a
4 MiB window, nothing over the message cap, so a drop can only come from
the window. Every frame posts, nothing is dropped, and the page's ack
seqs are read back to show the window stayed open because the page acked
rather than because the double was generous.

The file docstring said the double answers no RPC. It serves a
screencast stream now, so it says that instead, and says what it still
is not: it decides no domain behaviour.

The dead `violations` array is gone, folded with the CSP commit.

Red first: make the `ack` arm inert, as it was before this commit, and
the case fails with `Set{'posted','dropped'}` against `Set{'posted'}`.
A first attempt at that mutation left the byte subtraction in place and
stayed green, which is the mutation being wrong rather than the case.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* ci(mobile): run the whole mobile-web-app family, not a list that goes stale

The `mobile_web_app` job hand-listed ten files. The render check this
chain added was not among them, so it would have skipped in CI — and it
was not the first: three landed censuses were already unlisted, and
their closure blocks only run with `ORCA_MOBILE_WEB_APP_DEPS_REQUIRED=1`,
so they are green in the sharded `test` job whether or not they ever
ran here. Nobody could see it.

The list is now two vitest filename filters, `config/scripts/mobile-web-
app-` and the one builder test outside that prefix. Quoted, because
vitest matches a positional as a substring against the discovered files
rather than expanding a glob: `mobile-web-app-*.test.mjs` finds nothing,
and it fails by reporting no test files rather than by running fewer.
Both forms were tried before this one was written.

It runs 18 files and 205 cases, against 10 files before. With mobile
dependencies absent, 111 of those 205 skip, which is the measure of what
only this job runs. Per file, cases CI has never run:

  browser-pane-render            7 of 7   (this chain)
  source-control-external-links  9 of 9
  source-control-keyboard        6 of 6
  source-control-text-inputs     6 of 19  (C4.2)
  frame-budget-sweep             4 of 4   (this chain)
  route-manifest                 2 of 17

Two more files the filter adds run fully in the sharded job already and
change nothing here: `browser-pane-text-inputs` (C6.4 — it censuses a
hand-written closure and never bundles, so unlike the report it was not
skipping) and `external-link-seam`.

The sweep is renamed into the family for the same reason. As
`mobile-browser-frame-budget-sweep.test.ts` it matched neither the job's
filter nor `pr-code-change-scope.mjs`'s `config/scripts/mobile-web-app-`
prefix, so a change to it alone would not have run the job that runs it.
It is also gated on the dependency check now: it needs no
react-native-web, but it launches Chromium, and that flag is what tells
the job with a browser from the one without. Unguarded it would have
failed the sharded `test` job outright.

That filter is a prefix match. `config/scripts/mobile-web-app-` and
`mobile/src/` both fire this job, and `.github/workflows/pr.yml` is in
GLOBAL_FORCE_PREFIXES, so this commit runs everything.

Also: `postedFrame` in the ruling-2 pin and in the sweep both reached
the binary lane through `?.`, so a subscribe that opened no stream read
as zero posts — indistinguishable from a dropped frame, which is the
verdict both files are about. They throw now. The render check's
restated `640 * 1024` went with the window caps in c73b405f81; the cap
is read from `bridge-caps.ts`.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* ci(mobile): record what the mobile-web-app job costs to run

The filter that replaced the hand list runs 18 files where the list ran
10, so the step's cost is now a function of what anyone names into the
family rather than of what a reviewer remembered to add. Measured on
this machine: 25-30s wall for the whole step, of which the frame-budget
sweep is 2.5s.

The sweep is the one part whose cost is a choice. It encodes 111 noise
JPEGs in Chromium, one per viewport the budget fits, so adding rows to
that set is a decision about this job's runtime and the comment says so
where someone would make it.

Found, not fixed, and reported for its own PR rather than folded here:
the page bundles two copies of zod, mobile's 4.4.3 and the repo root's
4.5.4, reached through `src/shared/zod-salvage.ts`, which resolves
upward while `mobile/src/` resolves to mobile's. That is 808 KB of
duplicate source and two module instances in the shipped page. Aliasing
`zod` in `mobileWebAppBuildOptions` fixes it and was measured working
during this chain; it is reverted and stays reverted, because it changes
which zod shared code runs in the page and that is not a call to make
inside a CI commit. The CSP fix in 71254ab3a9 does not depend on it:
`globalConfig` lives on `globalThis`, so the banner covers both copies.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): load the frame-budget sweep's mobile modules after the dependency guard

vite transforms every file under mobile/ against mobile/tsconfig.json, which extends
expo/tsconfig.base.json; the sharded test job installs no mobile dependencies, so the
sweep's static imports failed the file at load before describe.skip ran. Type-only imports
stay static; the values load in beforeAll behind mobileWebAppDependenciesPresent().

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): certify the frame budget at the quality the pane ships

Round 2: the sweep and the render check encoded fixtures at a retyped 0.72; both now read
BROWSER_FRAME_QUALITY (the sweep from the module, the render check through the harness reader),
so a quality change fails the certification instead of leaving it green. Every render case now
asserts zero CSP violations; the sweep pins the 32 viewports left at scale 1; three references to
a renamed file and a file that never existed are corrected; a shim count comment is made
count-agnostic.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): assert the frame-budget sweep against the constant, not the measured maximum

The margin above the measured 0.55351 is what an encoder drift is allowed to spend; pinning the
measurement made a drift inside the margin fail a budget that still held. The number stays in the
docstring as the sweep's record.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): move the sweep's measured-maximum note beside the assertion it explains

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-20 07:43:14 -04:00
Jinwoo Hong ac024d4f05 feat(mobile): serve the files explorer and preview from the page (OTA phase C, C3.1) (#21710)
* refactor(mobile): take the files screens' router from the handoff seam

Inside the shell's page a screen is one document standing in for one screen, so
a target the page does not render has to be handed back to the app that does.
`useRouteHandoff` is where that decision lives, and its web sibling is the only
thing that makes it; both files screens held expo-router's own `useRouter`, so
on the web the explorer's Back and the preview's Back would post nothing and a
target outside the page would paint Unmatched over the page it is on.

Natively this is the same object — `route-handoff.ts` is `useRouter()` — so no
behaviour moves here, and `back()` stays expo-router's until the navigate-back
verb lands and the seam starts wrapping it.

A census rather than a behaviour test: neither screen's own tests can see the
difference, because a push that is never handed off still works for a target
inside the page. It walks this directory, refuses a value import of
expo-router, and names the two screens that must hold a router so a walk that
found nothing fails instead of passing empty.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): let the shell stand in for the two files routes

Both route files take the index.tsx shape — flag, MobileWebShellScreen, native
screen as fallback — and both gain the `.web.tsx` sibling that shape forces.

Inert until the manifest lists these routes: the shell answers `native-route`
for a route the bundle does not name, which is what `fallback` renders, and the
flag is `__DEV__`-only besides. Listing them waits on C2.3 and C2.5.

The sibling is not a precaution. The manifest defers every route behind
`import()`, so a native-only route module is invisible until the page opens
that route; the render check now opens both and, without the siblings, painted
`expo-modules-core.requireNativeViewManager is not available on web` instead of
the screen. That is also why the two cases render the route rather than
asserting a file exists.

The file path never becomes a path segment: only `hostId` and `worktreeId` are
spelled into the pathname, encoded, and everything else — `relativePath`,
`absolutePath`, `cwd`, `pathText` — is a param, which is how a `/`, a space or a
`..` stays out of the segment vocabulary the bridge holds a route to. The
preview render case proves the round trip on `docs/my notes/readme.md`.

`mobileFilePreviewShellParams` drops a param the normalizer left `undefined`
rather than sending it empty, because the page reads these back through
useLocalSearchParams where `line: ''` and no `line` are different screens. Its
test drives the normalizer rather than a hand-written literal: the literal omits
the key entirely, so it held with the filter removed.

The preview case also records what React Native Web says out loud — BackHandler
is inert on web, so Android back inside the page skips the unsaved-draft
prompt. Named in the assertion rather than filtered out, so closing it is a
change to that line.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): ask about an unsaved draft in the screen, not through Alert

React Native Web's `Alert` is `static alert() {}`. Inside the shell's page that
made Back with an unsaved terminal-artifact draft a button that did nothing at
all: no prompt, because the dialog is a no-op, and no navigation either, because
the code took the branch that shows one. Silently, with nothing on the console.

The prompt is now a row under the header. Not `ConfirmModal`, which every other
confirm here uses: that is a `BottomDrawer`, and C1.9 has Reanimated's animated
styles never reaching the DOM node on WKWebView, so on iOS in the page the
drawer parks off-screen and Back would be dead a second way. This paints the
same on every platform with no animation behind it.

Hardware back is registered natively only. React Native Web's
`BackHandler.addEventListener` logs "BackHandler is not supported on web and
should not be used." and hands back an inert subscription, so the guard never
armed there regardless; the render check asserted that console error on main and
now asserts none. The degradation is real and stated rather than hidden: Android
back inside the page pops the native stack without asking, and the page's own
Back control is where the question lives.

The decision moved to a hook so it is testable without a screen: the prompt also
drops itself when the draft it was about is saved or reverted, which is a state
`Alert` had no way to be in.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): keep expo-haptics' DOM shim out of the page

expo-haptics has a web build, and with no `navigator.vibrate` — iOS Safari,
which is the WebView the page runs in — it fakes a haptic by appending a hidden
`<label><input type="checkbox" switch>` to `document.head`, clicking it, and
removing it, once per call. C1.9 traced a long press that never fired on the
worktree list to exactly that stray click, and the file explorer calls
`triggerSelection` on every row tap, so C3 is the first domain to fire it per
tap rather than per long press.

`haptics.web.ts` answers the same five names with nothing. A phone holding the
page is a phone whose native app is right there with the real haptics, and a
missing tap feedback is worth less than a tap that does not register.

The test reads the shipped bytes rather than the import, because that is the
claim: with the override removed the bundle carries `ariaHidden` and
`pointer: coarse`; with it, neither, nor the `setAttribute("switch"` that does
the clicking. Not `navigator.vibrate` — react-native-web's own Vibration export
calls that and touches no DOM until something invokes it, which cost this test
one wrong red before it was narrowed.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): keep the files routes native when the page could not be given one

A file path is a param, so `/`, spaces and `..` all cross safely — but
`BRIDGE_MAX_ROUTE_PARAM_CHARS` is 1024 and a Windows long path is not bounded by
anything the user cannot exceed.

The symptom is not the blank document the design predicted, and the correction
matters: `bridge-host.ts` already parses the route against the page's own schema
and drops it to `null` when it fails, so `init` arrives naming no screen and the
page paints "Update Orca to open this workspace" — a wrong message about a fine
app, over a native screen that works. Deciding before the switch instead leaves
the route native, which is where every route starts.

The schema is the predicate rather than a copy of its bounds, so the rule cannot
drift from the half that matters, which is the half the page reads. The same
call also refuses a `worktreeId` the segment rule will not route: `..` survives
`encodeURIComponent`, which is the C1.8 class.

The tests assert the schema really refuses each input before asserting the guard
does, so neither case can pass by being impossible.

This belongs in the shell beside the schema; it is in the files domain while the
contract files are the C2 lane's.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): pin what keeps a file path out of the route vocabulary

Seven shapes, one case each rather than a representative: a plain path, a space,
a dot segment, an already-encoded slash, a fragment, non-ASCII, and an absolute
path. Each is checked in the two directions a path travels — the href the shell
writes into the page's history, and the href the page would hand back — for both
the pattern accepting it and the path coming back out of the query unchanged.

The counterfactual is in the file: the same paths spelled as a segment are
refused. Without that, the cases above would hold for a rule that was never
doing any work. Mutating `stringifyRouteHref` to join its query by hand instead
of through `URLSearchParams` fails three of them.

Also fixes two new test files the tests-typecheck ratchet caught: the partial
`react-native` mock needs a typed `addEventListener`, `act` will not take a
callback that returns a value, and `findAllByType('Pressable')` does not
typecheck against `ElementType` — the neighbouring files that do it are
grandfathered, so the tag comparison goes through a helper instead.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): derive the discard prompt instead of clearing it in an effect

Both changed-code gate findings, which the lane had not run until the last
commit. React Doctor is right: the effect that cleared the prompt when the draft
went away adjusted state after a prop changed, so a save landing while the
prompt was up painted one frame still offering to discard nothing. The prompt is
now `asking && hasUnsavedDraft`, which cannot be stale by construction, and the
test that covers it passes unchanged.

The hoisted mock's `as` on a string literal is gone too: the literal narrows on
its own and the tests reassign it, so the holder is annotated instead.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* test(mobile): add the files routes to the hybrid shell flag census

The census pins every file that reads `useMobileWebShellEnabled`, because a
reader nobody listed is how a dark feature stops being dark. C3's two routes are
deliberate entries: each has a native screen behind it as `fallback`, and each
is inert until the manifest lists the route.

Found by the full mobile suite rather than by the files subset this lane had
been running per commit.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* feat(mobile): serve the files explorer and preview from the page

The last C3 commit: both routes join MOBILE_WEB_PAGE_ROUTES, and the shell
starts rendering the page for them on a phone with the dev flag on.

Grants are not the same for the two, and the difference is the point. Both take
`navigate` (Back pops the native stack, and the explorer's rows open the preview
beside it) and `storage` (the shared components the host layout renders above
them). Only the preview takes `externalLink`: a Markdown preview renders links
and `MobileMarkdown` opens them through the platform seam.

The explorer does not, and measuring is what says so rather than reading. Every
page route reaches `external-link.web.ts` — `/h/[hostId]` and agent-history
included, both granted nothing for it — because the protocol wall in the shared
host layout imports it. So closure membership is not the oracle for a grant; the
question is whether the route's own screens call it, and only the preview's do.
`MobileMarkdown` is in the preview closure and absent from the explorer's, which
the census now asserts in both directions.

Neither route writes a clipboard, so neither takes `native.clipboard.write`;
the census pins that as the absence of both `ExpoClipboard.web.js` and the
clipboard seam, with the tasks closure as the control that the probe can see one
when there is one.

The seam predicate moved into a module both censuses import rather than being
restated per series: two spellings of one rule drift, and this one is a regex.

Red-first: both manifest assertions failed on the new entries before they were
updated, and routing `MobileMarkdown` around the seam fails the preview's census
while leaving the explorer's passing, which is the asymmetry the grants encode.

Closure sizes as the page ships them, extensionless so the `.web.tsx` is what is
measured: explorer 3439 modules / 302 local / 10 under src/files, preview 3667 /
331 / 20.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): read the files route's ids as one value and key the shell on them

Two round-1 findings, both reproduced before the fix.

A repeated query key reaches `useLocalSearchParams` as an array, and the
explorer read `hostId` and `worktreeId` bare. `String(['a','b'])` is `a,b`, so
the template built `/h/host-a%2Chost-b/files/wt-1%2Cwt-2` — a single segment the
bridge's rule accepts, and the shell would open a page for a host nobody has.
Read through `firstParam` now, as the tasks and agent-history switches do. The
preview already went through `singleParam` and is unchanged.

Neither switch keyed `MobileWebShellScreen`, where `index.tsx`, `tasks.tsx` and
agent-history all do. A host captures the grants its session opened with, so a
screen reused across a route change keeps authorising frames under the grants of
the route the page has left; only a remount drops that bridge. Both are keyed on
the route pathname now, with agent-history's reason.

The new route test is the agent-history one's shape. It caught both: the array
case landed on no route at all, because `name` was an array too and the schema
refuses a non-string param value, and the two lifecycle cases saw a prop update
where a remount was owed. It also needs agent-history's `lucide-react-native`
mock, since `firstParam` lives in the source-control barrel.

`name` is now omitted when empty rather than sent as `name=`, matching the two
switches beside it: an absent label lets the panel derive its own.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): confirm a discarded draft with the app's own modal

Round-1 findings 3, 4, 5 and the minor one.

**ConfirmModal, not the bespoke row.** The row existed because C1.9 had
Reanimated's animated styles never reaching the DOM node on WKWebView, which
left every BottomDrawer parked off-screen. C1.10 (`b7c06900e2`, an ancestor of
this branch) fixed that with a dependency array on the mapper hooks, and the
drawer render check now holds it on WebKit as well as Chromium. With the reason
gone the row does not stand on its other merits: `Alert.alert` was modal on
native before the page existed, and the row quietly changed that for phones
too, so the app's own confirm is both the idiom and the closer behaviour.
`MobileFilePreviewDiscardPrompt`, its test and its thirty style keys are gone;
the hook's state machine and its tests are unchanged.

**The encoding test claimed more than it pinned.** Hand-joining the query reds
only three of the seven shapes; `docs/readme.md`, `../etc/passwd`,
`docs/日本語.md` and `/logs/run.txt` are encoding-neutral in the query, whose
pattern half is `[^#\s]*` and admits a slash, a dot segment and non-ASCII
verbatim. Rather than narrow the claim in a comment, the split is now pinned by
behaviour: each neutral shape must survive the query unencoded, each
load-bearing one must not. Moving `docs/readme.md` between the lists fails it.

**The manifest comment named one shared-layout opener and there are two.** The
New Workspace source field, which the sidebar renders on a wide layout, opens a
URL through the seam as well. Both are the shared layout's and every `/h` route
reaches both, `/h/[hostId]` included with no `externalLink`, so the tablet tap
is dead on all of them — recorded here as pre-existing rather than fixed, since
the grants do not move.

**Minor:** the dot-segment case in the guard test now asserts the schema refuses
the route before asserting the guard returns null, as the length case does.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): stop every page drawer logging a BackHandler error when it opens

Round-2 findings.

**The registration belongs to the drawer, and that is where the guard went.**
`mounted-bottom-drawer.tsx` armed `hardwareBackPress` whenever a drawer was
visible and interactive, with no platform check, so the hook's claim to have
dropped that console line held only while its prompt was closed — and every page
drawer since C1 has logged it on open. Platform-gated at the drawer now; the
hook's comment says so rather than claiming the credit.

**Nothing had ever opened a modal in a browser.** The render check next door
mounts both files routes and reads what they paint but taps nothing, so
`ConfirmModal` inside the page — a BottomDrawer, so Reanimated, a portal and a
gesture handler — was unproved. A new render file loads an editable terminal
artifact through the harness's scripted reply, edits it, taps the page's Back,
and asserts the prompt's title is up and no BackHandler line is on the console.
Red first on exactly that line; the prompt itself painted, which is also the
first proof on a browser that C1.10's fix carries a real drawer in the page. A
second case answers Stay and checks the draft survives. Its own file rather than
the render check's, which is at 482 of the 600-line cap; registered in pr.yml.

**The encoding rule was stated wrong.** Two rules decide it and neither is about
paths: the pattern's query half refuses whitespace and `#`, and
`URLSearchParams` is form-urlencoded, so it reinterprets `&`, `+` and a valid
`%XX`. `a+b.ts` reads back `a b.ts` and `a&b.ts` reads back `a`, so both are
load-bearing; `a=b.ts` and `a%b.ts` are not, because only the first `=` splits
the pair and a lone `%` begins no escape. A newline joins the load-bearing list
as the refused shape rather than the altered one.

**The web sibling read its params bare** where the native one uses `firstParam`.
Not reachable — the page only arrives through `init.route`, whose params are
already `Record<string, string>` — but the two files are meant to be one screen.

The preview keys on the pathname alone, and the comment now says why that is
enough: every caller in this tree pushes.

Closures after this: explorer 3441 / 304 / 10, preview 3666 / 330 / 19. The
explorer grew two modules because its web sibling now reaches `firstParam`.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): give the explorer the grants the preview needs, and key on the route

Bot findings, one of them a real gap.

**Pullfrog is right, and my grant oracle was half a rule.** Grants resolve once,
from the route the shell opened: `grantsForRoute` reads `session.routePathname`
and `init.grants.native` carries the answer for that session. The explorer's
rows push to the preview, and because the preview is a page route that push
stays inside the same document — no second `init`. So a preview opened that way
runs under the explorer's grants, and a Markdown link in it was refused by
`notifyExternalLink` with nothing on screen to say why. "Does the route's own
screen call it" was right for a route's own screens and wrong for the routes it
reaches in-page, so the explorer now declares `externalLink` as a transitive
grant, with the comment saying that rather than claiming it opens links. The
census pins the pair as a superset; removing the grant reds it.

**The seam regexes matched one quote style.** A double-quoted `react-native`
specifier walked past both censuses unseen. Both styles now, with the predicate
tested directly for the first time.

**The discard request outlived its draft.** `asking` stayed set after a save or
a revert, so the next edit re-showed the prompt with no Back request behind it.
The request is now dropped when the draft it was about goes, adjusted during
render rather than in an effect — the shape React Doctor named in the round-1
fold. Red first: save with the prompt up, edit again, prompt is back.

**CodeRabbit's keying comment is a correctness point, not the question I
answered.** The page learns its route exactly once, out of `init`, so a
same-path param change — another file in the same worktree — left the shell
mounted and the page still showing the file it was opened on. My comment claimed
"the screen reloads the preview from the param either way", which is true only
with the shell absent. Both switches key on the whole route now, params
included; two tests cover the same-path case and both red on a pathname-only
key.

`build-mobile-web-app-bundle.test.mjs` hit 601 of its 600-line cap on the way,
so the two manifest assertions now share one expected list instead of repeating
it. Closures unchanged: explorer 3441 / 304 / 10, preview 3666 / 330 / 19.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb

* fix(mobile): make the seam test import the module it is testing

Round 3.

**The blocker is mine and the reviewer's diagnosis is exact.** The seam
predicate test imported an absolute path into this lane's worktree. On CI that
module does not exist and it takes the whole `config/scripts` suite down; here
it resolved to the same file by accident, so the test was green against a tree
rather than against the checkout — which is why reverting the double-quote fix
left it passing and the predicate untested. Relative now, and proved: reverting
the fix in place reds both double-quoted cases, which is the first time this
test has failed for the right reason. Every file this PR touches is grepped for
`/Users/` and `orca-lanes`; none carries a path.

**Three comments outlived the grant change.** The two lists became equal when
the explorer took `externalLink`, so "longer than the explorer's" and "declared
with different grants" were both false. Corrected to what is actually true: the
lists are equal and the reasons are not — the preview has its own consumer in
`MobileMarkdown`, the explorer has none and declares the grant because its rows
push to the preview in-page.

**The duplicated serializer is pinned rather than imported.** `shellRouteHref`
lives in `page-bootstrap.ts` beside the page's RPC client and its document
channel, so a native route file importing it would pull both into the app. The
copy stays, and a test asserts the two agree on three routes; dropping the
empty-search branch reds it.

**Recorded, not fixed:** the sidebar `HostScreen` pushes to `/h/<id>/tasks`
through the handoff, which is local, so on a tablet the tasks page runs without
`native.clipboard.write` from any page route and its copy actions refuse
silently. Pre-existing since C2.1 for the worktree list and agent history. Named
in the explorer's manifest comment as the known remaining hop, with the fix
being a handoff rule in its own PR.

The equality pin needed `it.each<BridgeInitRoute>`: the inferred table is a
union whose members carry `?: undefined`, which the ratchet caught.

Claude-Session: https://claude.ai/code/session_01JNnE9qzUZMMnqpZWCqM3nb
2026-09-19 15:47:25 -04:00